The Sonnet and Opus setup that matters: effort, cache, verification, and cost per finished task
The cheapest model is not the model with the lowest token price
It is the one that finishes the job, passes the check, and does not make you pay for the same context five more times
Sonnet 5.5 and Opus 5.5 make that distinction unusually important. One is priced for volume. The other is priced for harder work. Both have new effort behavior, and both can be surprisingly cheap inside a well-cached agent loop
Copy your old config into either model and the result may be slower, more expensive, or a 400 error
This is the setup I would build instead
1TASK → SONNET 5.5 → CHECK → OPUS 5.5 IF NEEDED → VERIFIED RESULT2 ↘ effort ↗ ↘ cache + usage ledger ↗
I publish practical breakdowns of AI agents, workflows, and production systems on Substack
The number that should decide your stack
Most model comparisons start with dollars per million tokens
Your agent does not ship tokens. It ships completed tasks
https://x.com/claudeai/status/2102435511222890900
Again: their tests. Your architecture needs your numbers
That means the check cannot be a vague thumbs-up after reading one impressive answer. For coding, use the test that would block the merge. For extraction, compare required fields against a labeled set. For research, record whether the cited source actually supports each claim. Include the cost of tasks that never pass, not just the nice examples in your demo
And inspect the hard tail separately. If the cheapest setting handles 90% of your requests but burns half the budget on the remaining 10%, its average can hide the part of the workflow that needs a different model
What the 5.5 price sheet actually says
As of October 3, 2026, at standard Claude API rates, per million tokens:
1SONNET 5.52Fresh input $2 Output $103Cache read $0.20 Cache write $2.50 / 5m, $4 / 1h45OPUS 5.56Fresh input $4 Output $207Cache read $0.20 Cache write $5 / 5m, $8 / 1h
Both models have a 1M-token context window and a 128K-token maximum output. These are ceilings, not a reason to fill either one.
Model specifications and pricing
The odd line is cache read
Opus costs twice as much for fresh input and output, but a cached prefix costs the same $0.20 per million on both models. That does not make an Opus run equally cheap: it still pays more for new input, output, and cache writes. It does mean the model price gap can shrink in a read-heavy session
There is a second distinction people miss. Anthropic's "40% less than Opus 5" is an estimate of typical \run cost\. Opus 5.5's fresh token prices fell 20%; its cache-read price fell 60%. Those numbers are related, but they are not interchangeable

Effort is a routing decision, not a quality slider
Sonnet 5.5 supports low, medium, high, xhigh, and max. On the API, it defaults to high, in Claude apps, Anthropic says medium is the default. Opus 5.5 defaults to medium on the API. These levels are not calibrated to mean exactly what the same words meant on the previous models.
My starting map:
- Sonnet low For narrow, latency-sensitive requests with a cheap check
- Sonnet medium For well-specified coding and routine multi-step work
- Sonnet high When the medium run fails a real check, or the task has a demonstrated complexity pattern
- Opus medium For ambiguous, cross-file, long-horizon work where Sonnet is spending turns circling the problem
- Xhigh/max Only after your evals show a gain worth the extra time and tokens
This is a starting hypothesis, not a universal hierarchy. In Anthropic's reported FrontierCode results for Sonnet 5.5, xhigh scored above max. More effort is not a proof of a better outcome
Anthropic's footnote explains the counterintuitive result: at max, the model more often launched extra code review work. In two examined cases, that led to a timeout or edits beyond the task's scope.
The failure mode was not "the model did not think enough" It was spending effort in the wrong place. If your agent already passes its checks, extra review turns can become a cost and a source of new mistakes
https://x.com/edwinarbus/status/2104675431853248816
Also, do not set max_tokens low and call that optimization. The limit covers thinking and visible output. If you cut it mid-task, you can buy a truncated answer and a second run instead of saving money
https://x.com/claudeai/status/2104633115620823187
That is a compelling launch claim. A production config still needs to beat your own baseline
Run a small sweep before you invent a model router
Take 10-30 tasks you actually care about. Include easy work, ambiguous work, and the annoying failures from your logs. Give each task a verifier: tests, a structured comparison, a known answer, or a human rubric recorded before the run
Here is the smallest useful API probe. It logs the usage fields you need. Run it on each model and effort level against the same task, then attach your own pass/fail check. It is not a full agent benchmark
1import anthropic23client = anthropic.Anthropic()4response = client.messages.create(5 model="claude-sonnet-5-5", # repeat with claude-opus-5-56 max_tokens=8192,7 output_config={"effort": "medium"}, # repeat at high8 messages=[{"role": "user", "content": "Replace this with a real task."}],9)1011answer = "".join(b.text for b in response.content if b.type == "text")12usage = response.usage13print(answer)14print("fresh", usage.input_tokens, "output", usage.output_tokens)15print("cache read", usage.cache_read_input_tokens)16print("cache write", usage.cache_creation_input_tokens)
This assumes the official Anthropic Python package and an ANTHROPIC_API_KEY environment variable. It is one isolated call without caching enabled, so zero cache reads and writes are expected. The next section shows what changes that
For a real agent, sum usage across all API calls under one task ID, including retries and tool calls. Count a pass only when the verifier says the job is finished
Compare total dollars per pass before choosing your default
Keep the test honest:
- Freeze the task set and the verifier before comparing configs
- Run the same tools, permissions, context, and output requirements on every candidate
- Record pass rate, total spend, cost per pass, latency, and the longest or most expensive failures
- Count stop_reason: "max_tokens" as an incomplete attempt, not a cheap success
The short code sample above uses an 8K output cap for a single-turn probe. Do not copy that cap into a long coding agent.
Anthropic recommends much more headroom for agentic work because hidden thinking counts toward the same limit.
Set the cap for the work, then control spend with effort, caching, and a task budget rather than forcing the answer to stop mid-job
Cache the stable part of the job
Agents repeatedly send the same system instructions, tool definitions, repository map, and prior conversation. If that prefix is stable, prompt caching changes the economics more than a minor prompt rewrite
For example, 200K cached tokens read 50 times are 10M cache-read tokens. At $0.20 per million, the reads cost $2 on either 5.5 model.
On Opus 5.5, sending those same 10M tokens as fresh input would cost $40. The first five-minute cache write of 200K tokens is another $1.
This is an illustration of prefix charges only: new input, output, other writes, TTL expiry, and actual cache misses add to the bill
The practical rules:
- On the Claude API, opt in to prompt caching with top-level cache_control={"type": "ephemeral"} or explicit cache breakpoints. The probe above does neither, so its cache counters will normally stay at zero
- Put stable instructions and tools before the changing user request
- Keep the shared prefix identical across turns; verify actual cache_read_input_tokens
- Treat a model switch as a new conversation budget, not a free continuation. The cache is per-model: an Opus request cannot read the prefix Sonnet just cached
- Avoid changing top-level effort on every turn; that changes the rendered prompt and invalidates cached prefixes
On supported models, a per-message effort change can preserve the earlier cache, but it requires Anthropic's beta header and is not the same as changing the top-level output_config.
Sonnet 5.5 also has a between_tools caveat: effort cannot change mid-conversation in that mode.
Do not infer a cache hit from a fast response. Read the usage object. It separates fresh input, cache creation, and cache reads
Both 5.5 models need at least 512 tokens in a cacheable prefix. A tiny system prompt will not produce the savings in the example above. The default cache lifetime is five minutes, which fits a fast tool loop.
A one-hour write costs more and makes sense only when real sessions often pause long enough to miss the five-minute window. Measure those gaps before paying for the longer TTL
Escalate after evidence, not anxiety
Most teams build the router backwards: classify a task as "hard," send it to the expensive model, and never learn whether the cheaper path would have passed
Use the verifier as the routing signal

11 Sonnet 5.5 · chosen effort → run the task22 Verifier → accept if it passes33 Opus 5.5 · medium → retry only with failure evidence44 Verifier → accept or hand off with evidence
The check can be a test suite, schema validation, a known answer, or a reviewer. It should explain what failed.
"The answer feels weak" is a poor escalation signal, "the changed endpoint fails two integration tests" is useful
Do not blindly repeat an identical prompt. Give the next attempt the failed check, the relevant artifacts, and a specific instruction to correct the gap. Cap the ladder so an agent cannot burn its budget trying to fix a task that needs a human decision
You can test a Sonnet high retry in your offline sweep. Keep it in the live route only if it lowers cost per verified task. There is no reason to make every failure pay for two Sonnet runs before Opus
The model switch itself can break a cached prefix. Account for that when you compare the rescue path against an Opus-first route
The crossover is easy to miss. Suppose a Sonnet attempt costs $0.06 and passes 80% of your tasks.
If every failed task then costs $0.20 to finish on Opus, your illustrative average is $0.10 per finished task: $0.06 plus a $0.20 rescue on one task in five. That beats paying $0.20 for Opus on every task. But if Sonnet costs $0.14 and only passes half, the same ladder costs $0.24, even before you price the model switch. In that workload, Opus-first is cheaper and faster
Those figures are examples, not measured Claude results. Their purpose is to make the routing rule falsifiable. The ladder only earns its place if the saved Opus calls outweigh the failed Sonnet attempts, cache misses, and added latency
There is also a middle path: Anthropic's beta advisor tool. Sonnet can keep running the task and ask Opus for help on a difficult decision, instead of handing the entire job to Opus.
This is not automatically cheaper. Log how often Sonnet actually consults the advisor, what those calls cost, and whether they improve the final pass rate. If the executor rarely asks, the advisor is just an unused feature.
With these 5.5 models, the advice itself is returned encrypted to the client, so evaluate the resulting work rather than pretending you can audit the private advice text
Four leaks that raise the bill before model choice matters
Not every cost problem deserves a new router
Check these first:
- Output that keeps growing On both 5.5 models, output tokens cost five times fresh input tokens. In a conversation, a long answer can also return as context on later turns. Ask for the artifact and a short completion note, not a narrated transcript of every step. Hidden thinking is billed as output too, so a terse final answer alone will not fix an effort problem. Do not suppress the evidence you need to verify the result
- Images larger than the task needsSonnet 5.5 can process higher-resolution images than older Sonnet releases, which can raise image token counts. If the agent only needs a button label or one paragraph, crop or resize first. If it needs a dense chart or tiny UI details, keep the resolution and measure the cost instead of blindly shrinking it
- Context nobody usesTool definitions, stale logs, old search results, and a sprawling CLAUDE.md can follow every request. Put durable rules in a short stable prefix; keep transient evidence near the task that needs it. Trimming context should not delete facts the model still needs to finish correctly
- Interactive pricing for work nobody is waiting onThe Message Batches API discounts input and output by 50% on both models. It is useful for offline evaluations, document backfills, and other asynchronous jobs. It is not a replacement for a live tool loop where a person needs the next step now
The pattern is the same across all four: remove work the task does not need before buying more intelligence or lowering effort until quality breaks
The migration traps that turn a saving into a 400
Old request bodies are a bad starting point for the 5.5 family.
In particular:
- Opus 5.5 thinking is always on Remove thinking: {"type": "disabled"} and old fixed budget_tokens settings; control depth with output_config.effort
- Forced tool choice fails on both 5.5 models
tool_choice values any and tool return a 400. Use auto, specify when the tool should be used, and validate the tool result in your own code
- Thinking blocks are not text blocks Read content by type, not content [0]. In tool loops, return thinking blocks unchanged with the assistant turn
- Your UI may appear silent On Opus 5.5, between-tool progress can arrive in thinking blocks that are empty at the default display setting. If you previously rendered those notes to users, request a supported thinking display mode and render blocks by type Otherwise the agent may be working while the interface looks frozen
- Old computer-use tool versions can fail Check the current tool version before migrating a browser/computer agent
- A smaller max_tokens cap can cut off work
Thinking is included even when the text is hidden
These are API behavior changes, not prompt-writing tricks.
Opus migration guide and Sonnet migration guide
Put the contract in Claude Code, not in your head
The API is where you can measure every usage field. Claude Code is where many people will first feel the model change. The same principle applies: give the agent a bounded definition of done, then make it show the evidence
In Claude Code, /model selects the model and /effort selects the supported effort level. Check the active settings before comparing sessions. The Sonnet API default is not a reliable description of what your Claude app or Claude Code session is currently using
This is a complete, reusable CLAUDE.md starting block. Change the commands to match your project
1# Working contract23Make only the requested change. Preserve unrelated work.4Run the relevant tests after editing. Report any check you cannot run.5Stop when the requested work passes. Do not add extra features or review loops.6End with: Changed / Verified / Remaining risk.7Ask before destructive actions, publishing, or changes outside this repository.
That block will not magically make every run cheap. It makes success and failure visible. From there, you can compare a Sonnet-first workflow with an Opus-first workflow on the same tasks
The task message still has to be specific. Here is the difference between "fix the payments code" and a job an agent can actually finish
1Change: migrate the payment endpoint to the new client2Done: old client removed, endpoint tests pass, diff limited to this path3Stop: ask before deleting data or changing anything outside the repo4Report: changed files, exact checks run, remaining risk
That small contract gives the verifier something concrete to inspect. It also gives the model a reason to stop. An open-ended instruction to "review until perfect" can turn a passing change into another paid loop
For long projects, keep the checklist in a file that survives compaction. For subagents, ask the lead agent to inspect their evidence before accepting their reports. And if you only asked for ideas, tell Claude not to start building. Those are workflow boundaries, not "be smarter" prompts
The setup I would ship first
- Pick 10-30 real tasks and define a check for each
- Sweep Sonnet 5.5 at medium and high, then Opus 5.5 at medium
- Log fresh input, output, cache writes, cache reads, latency, retries, and pass/fail per task
- Keep the stable prefix cacheable and confirm hits in usage
- Route only the failures upward, with their evidence attached
- Revisit the ladder when the workload changes. A saved benchmark is not a permanent truth
If the difficult 10% repeatedly goes straight from Sonnet failure to Opus success, consider routing that recognizable task class to Opus from the start. If Sonnet high passes those same cases for less, keep them there. A router is a measured policy, not a permanent opinion about which model is smarter
The 5.5 upgrade is not just "use Sonnet for cheap work and Opus for hard work"
It is a chance to stop pricing the model and start pricing the finished job
If u read this far
-> Subscribe to my Substack
-> Join my Telegram
-> Bookmark this article
-> Follow @0xwhrrari



![Claude Code God-Tier Setup Guide for All Japanese Users [Free Copy-Paste]](/cdn-cgi/image/width=1920,quality=90,format=auto,metadata=none/https%3A%2F%2Fcms-assets.youmind.com%2Fmedia%2F1791139777975_arnfzw_HTsDkD9a0AAaAvN.jpg)

![Daily Crown Stakes [S] Prediction](/cdn-cgi/image/width=1920,quality=90,format=auto,metadata=none/https%3A%2F%2Fcms-assets.youmind.com%2Fmedia%2F1791139224864_jgbtnv_HTsRNi4awAArhBP.jpg)