← Back to home Local LLM or Claude Code? cover

Local LLM or Claude Code?

96% of Requests Failed Locally, Half of the Steps Didn't: Measured AI Agent Routing with Ollama

96% of my requests couldn't go to a local LLM. Split into single steps, half of them could

Routing work to a local LLM should make it cheaper and safer. Then I counted my own requests: 96% couldn't go local. This book breaks Claude Code's work into single steps and routes them with rules and probabilities, all measured.

Local LLM Series [Routing & Design]: deciding what stays on your machine
Read on Kindle Read sample chapters See chapter list

30+ technical books across 4 languages · Sold on Kindle in 6 countries · From a year of real production use

$4.99 Included with Kindle Unlimited Published:
ken imoto
ken imoto — Author of the Practical Claude Code & Harness Engineering series. 30+ technical books across JA/EN/PT/ES. · 7-day return window via Amazon
Other editions: 日本語

Overview

Local LLM vs Claude Code, measured on 550 requests and steps: 96% of requests failed locally, half of single steps didn't. Rules, logprobs, thresholds.

What you will be able to do

Who is this book for

Problems this book solves

Where this book stands

Why this book

How this differs from other AI books

Compared to This book's difference
Local LLM setup guides (Ollama, LM Studio) Setup guides stop once the model runs. This book measures which parts of Claude Code's work to hand it after that
"Run Claude Code fully local" guides They replace the frontier model. This book measured that 96% of whole requests fail locally and routes single steps instead
LLM routing papers (RouteLLM and others) Papers route whole queries on public benchmarks. This book measures your own requests and steps and combines step-type rules with probabilities

Table of contents

Introduction: the gatekeeper let a password through

Read a local LLM's judgment as a probability, and don't take it at face value

  1. Preface
    • What this book covers
    • What this book does not cover
    • Who this book is for
    • The measurement setup
    • How to read it
    • Shelf life
  2. Prologue: The AI Looked at a Note Containing the Production DB Password and Said "Fine" with 69% Confidence
    • It was supposed to keep secrets in
    • The ranking was perfect
    • Don't take the probability at face value
    • What I measured
    • Map of the book
Sounds good so far → Continue on Kindle

Part 1 What to hand to a local model

At the request level almost nothing goes; at the step level half does

  1. 01 Chapter 1: 96% of My Requests Couldn't Go Local
    • 1-1 What I set out to measure
    • 1-2 How I built the 100
    • 1-3 Results
    • 1-4 Raw instructions told the same story
    • 1-5 Why it turned out this way
    • 1-6 What about the 4 with secrets?
    • 1-7 Request-level routing doesn't have the material
  2. 02 Chapter 2: Break It into Steps and Half Can Go Local
    • 2-1 Counting steps
    • 2-2 Results: 97 of 200 are fine locally
    • 2-3 Template prompts were the exception
    • 2-4 Claude Code already routes steps
    • 2-5 Some steps don't need an LLM at all
    • 2-6 Count your own logs
  3. 03 Chapter 3: Decide by Rules Where You Can, and Use Probability Only for the Rest
    • 3-1 A router that asks for a probability
    • 3-2 Results
    • 3-3 What was the probability actually predicting?
    • 3-4 Decide by rules where you can
    • 3-5 The secrets gatekeeper takes the same shape
    • 3-6 My three layers
Sounds good so far → Continue on Kindle

Part 2 AI that returns only a judgment

Jev's idea of answering with a type and a probability, and its open implementations

  1. 04 Chapter 4: What Jev Changed: Answering with Types and Probabilities
    • 4-1 No text back: a type and a probability
    • 4-2 The official numbers, and the vendor's own caveats
    • 4-3 Comparing against a cheap model instead
    • 4-4 An LLM can do the same thing
    • 4-5 Jev's own patterns also say "decide the fixed parts by rules"
    • 4-6 Where Jev sits in this book
  2. 05 Chapter 5: 24 Implementations in Two Weeks: Reading Open Implementations and Independent Evaluations
    • 5-1 Jev's internals are not public
    • 5-2 The 24 fall into three approaches
    • 5-3 Running rizzo-flow on an RTX 4070
    • 5-4 Reading the independent evaluations
    • 5-5 What to take home from the independent evaluations
Sounds good so far → Continue on Kindle

Part 3 Getting probabilities out locally

logprobs, model size, and how to set the threshold

  1. 06 Chapter 6: Getting logprobs Out, and Three Traps That Fail Silently
    • 6-1 What logprobs are
    • 6-2 Make it answer in one character, then read that character's probability
    • 6-3 Trap 1: The OpenAI-compatible API silently drops the probabilities
    • 6-4 Trap 2: Thinking models don't answer in the first token
    • 6-5 Trap 3: A candidate outside the top 20 looks like probability 0
    • 6-6 Other runtimes work the same way
    • 6-7 A probability is how sure it is, not whether it's right
    • 6-8 What to check with the very first question
  2. 07 Chapter 7: How Many B Is Enough?
    • 7-1 Same 60 requests, same questions, only the model changes
    • 7-2 What the numbers say
    • 7-3 Change what's being judged, and the ranking changes
    • 7-4 How I choose
  3. 08 Chapter 8: The Threshold Isn't Always 0.5
    • 8-1 Ordering and calibration are two different properties
    • 8-2 Where the line goes depends on what you most want to avoid
    • 8-3 How to draw the line: the highest line with zero misses
    • 8-4 A zero-miss line still leaks new secrets
    • 8-5 Leave some margin
    • 8-6 Collect secrets, not just items
    • 8-7 Worksheet: drawing the line
Sounds good so far → Continue on Kindle

Part 4 Keeping secrets in

What regular expressions catch, what a model catches, and what counts as a secret for you

  1. 09 Chapter 9: Stop Secrets That Wear a Label with Regular Expressions
    • 9-1 Secrets with a fixed shape walk in wearing a label
    • 9-2 In the first experiment, the labels were unreadable
    • 9-3 What gitleaks stopped and what it didn't
    • 9-4 Catching Japanese personal data by shape and checksum
    • 9-5 False alarms come from fake labels
    • 9-6 What to take from this chapter
  2. 10 Chapter 10: A Model Stops Only What Your Definition Names
    • 10-1 Using a model to catch what regex misses
    • 10-2 Going for zero misses with the model alone
    • 10-3 Real secrets looked nothing like the evaluation set
    • 10-4 Rewrite the definition as your own secrets
    • 10-5 A model alone can't protect unlabeled secrets
    • 10-6 Write down your secrets
Sounds good so far → Continue on Kindle

Part 5 After routing, how results come back

Designing the return value, operations, and the blueprint that survived

  1. 11 Chapter 11: The Free Executor Cost the Most
    • 11-1 The experiment: a strong model plans, a cheap model does the work
    • 11-2 Result: the setup with the free worker was the most expensive on every task
    • 11-3 Why the free worker cost the most
    • 11-4 Rerunning it in Claude Code
    • 11-5 The return value is part of routing
  2. 12 Chapter 12: Return a Typed Value
    • 12-1 Three ways to return a result
    • 12-2 Result: it works on short tasks and fails on long ones
    • 12-3 The typed return value left out "what the worker did"
    • 12-4 What to have the worker return
  3. 13 Chapter 13: Operations: Sharing the GPU, Waiting, and Redrawing the Line
    • 13-1 There is only one GPU
    • 13-2 Write the code as if timeouts will happen
    • 13-3 A local worker that doesn't clean up
    • 13-4 Change the definition, redraw the line
    • 13-5 What to log
    • 13-6 Operations checklist
  4. 14 Chapter 14: The Blueprint That Survived
    • 14-1 The first blueprint and the last
    • 14-2 The flow
    • 14-3 The gate at the entrance
    • 14-4 Splitting
    • 14-5 Type rules
    • 14-6 The probability judgment
    • 14-7 Running
    • 14-8 The return value
    • 14-9 As a config file
    • 14-10 The order for moving this to your environment
Sounds good so far → Continue on Kindle

Epilogue

Shelf life and the next move

  1. Epilogue: Shelf Life and the Next Move
    • What stays
    • What changes
    • The next move
Sounds good so far → Continue on Kindle

Send whole requests to a local LLM and almost nothing goes through. I counted 100 of my own requests, and 96 had to go to a frontier model. Then I broke Claude Code’s work into single steps and counted again: 97 of 200 steps were fine on a local model.

This book builds that routing on a single RTX 4070. Step types that settle the destination are decided by rules; only the rest goes to a local LLM’s probability. Secrets are stopped at the entrance before a request is handed over, and work sent to the local model comes back as a typed value. Chapter 14 collects a blueprint and example config files you can take home.

Dive deeper with related articles

FAQ

What is "Local LLM or Claude Code?" about?
It measures how to split the work you give Claude Code between a local LLM on your machine and a frontier model in the cloud, using 550 of my own requests and steps and a single RTX 4070. Routed as whole requests, 96% had to go to the frontier model. Broken into single steps, half could stay local. Across 14 chapters plus a prologue and epilogue, it builds one design: decide by rules where you can, use a local LLM's probability only for the rest, stop secrets at the entrance, and get results back as typed values.
Do I need the Jev API?
No. The book does not use Jev (TypeSafe's judgment AI). It borrows Jev's idea, answering with a type and a probability instead of text, and builds the same kind of judgment on a local LLM with Ollama's logprobs. Part 2 covers Jev itself, how to read its official numbers, and 24 open implementations.
What hardware do I need?
My setup is one RTX 4070 (12 GB) with Ollama. qwen3.5:4b is enough as the judge and fits in 12 GB. The numbers come from my own requests, so every chapter shows how to measure again on your data.
The measurements were done on Japanese work. Does it apply to me?
The method carries over. The personal-data regular expressions in Chapters 9 and 10 are Japanese formats (for example the My Number national ID), and the book says where to swap in your own locale's formats. Requests, steps, thresholds, and return values are language-independent.
Is the code available?
Yes. The entrance gatekeeper, the local LLM judge, the threshold tool, the step router, the typed return value, and an evaluation set built from made-up values are on GitHub as local-step-router (MIT license).
Where can I buy it?
Kindle only, on Amazon. It is enrolled in KDP Select, so it is included in Kindle Unlimited. A Japanese edition is also available.

Read on Kindle

Included in Kindle Unlimited

Read on Kindle ($4.99)
Topics: Local LLMClaude CodeAI AgentsOllamaLLM Routing