Same coding agents. Smaller bill.
Probe0 runs under Claude Code, Codex and Cursor on your machine. It cuts what they cost and shows you what every call spent.
- claude-code · 41822sonnet-4.6 · $0.0732sonnet-4.618.4k / 1.2k2.9s$0.0732Upstream
- codex · 41903qwen3:8b · local · $0.0000qwen3:8b · local6.1k / 8400.8s$0.0000Local
- claude-code · 42011opus-4.7 · blockedopus-4.7--blockedCapped
1proxy slot
Your machine has one, and every tool wants it. So you pick an inspector, a cache or a router, and go without the rest.
93%of the 1-hour cache
Claude Code pays for an hour of cache it then uses for five minutes. Probe0 buys the five minutes.
92%of the bill is cache
51% reads, 41% writes, across 85,569 real turns. The money goes on buying the cache, not on your prompt.
Four ways it makes the same work cost less
It runs under the tools you already use, and any step can be turned off. On 85,569 real turns it cut the bill 12.8% with no model swaps, and 24% with swaps on.
- 01
Intercept
Install it once. Every agent you run goes through it — Claude Code, Codex, Cursor.
- 02
Route
Easy work goes to a model on your own machine. Hard work still goes to the best one. A weak local answer gets sent up automatically.
- 03
Reuse
Ask the same thing twice and the second answer is free. Two agents asking at once only pay once.
- 04
Account
Every call gets written down: model, tokens, cost, time, which agent. A spend cap pauses a run before it gets expensive.
Turn on only the parts you want
Ten switches. Each reports what it saved, so nothing stays on out of faith.
Send less
- Strip Compression
- Shrinks tool output before you pay for it.
- Caveman
- Optional terse mode. No articles, no hedges.
Send it somewhere cheaper
- Local Routing
- Whole classes of work run on your own model.
- Cloud Routing
- Easy turns go to a cheaper model in the same family.
- Effort step-down
- Same model, less thinking, on easy turns.
Don't send it at all
- Exact Cache
- Repeat questions never hit the network again.
- Semantic Cache
- Near-identical questions too, held to a strict match floor.
- Duplicate Collapse
- Two agents ask at once, you pay once.
Know what it cost
- Spend Guard
- A hard cap per run or per day. Warns, then pauses.
- Recording
- Every call, the agent behind it, cost you can check.
Pay for less output, keep the cache
Rewrite a prompt and you lose the copy your provider is already holding. Probe0 only shrinks what your tools send back, so that copy keeps paying off.
- Strip Compressionon by default
- Bash output comes back about 30% smaller, the turn it arrives. Old file reads shrink later, once the provider has stopped holding them. Worth about three points of the bill.
- Cavemanoptional
- A terse mode. Articles, linking verbs and hedges go, word by word. No saving is claimed for it.
One slot, and anyone can build in it
Long term, this is meant to be the layer other people build on.
- Only one thing can hold the slot
- Your machine has one proxy position. An inspector, or a cache, or a router. Not all three.
- So the slot should be programmable
- Probe0 takes the position and then opens it up. A module sees a request and decides what happens to it.
- A module asks for what it needs
- Read the details, change the body, reach the network — and nothing else. You approve that list when you install it.
- Everything published is reviewed
- Asking up front is what makes review possible: only what a module asked for has to be checked.
The cheapest call never leaves your machine
Plenty of agent work doesn't need the best model on the market.
- Your models, your hardware
- Point Probe0 at a model already running in Ollama or LM Studio. It becomes another tier.
- Or a cheaper cloud model
- Step down to a cheaper model in the same family instead. Requests carrying tool calls take the safe step only.
- You set the share, not a rule
- Both are a setting: how much of your traffic takes the cheaper path. Left at the default, about 11% of turns.
- Nothing leaves your machine
- The proxy, the cache and the record of every call all run here. There is no server to phone home to.
The month you didn't need that plan
Probe0 knows what you actually used. So it can tell you when you're paying for a tier above the one you need. Usually worth more than any cache will save you.
It works the other way too. Hitting your limit every week? It says so.
- Billed to provider
- $96.20
- Answered from cache
- 2,317
- Duplicate calls collapsed
- 1,842
- Cache writes avoided
- 3,106
You could drop to the tier below. You used 48% of a Max 20x plan over these 30 days.
Try it on your own agents
Private beta on macOS. Install it, point your shell at it, and it starts counting on the next run. Nothing gets uploaded.