YouMind
Sign in

Claude 5.5: Stop Paying for Failed Tasks

@0xwhrrari
ENGLISHOct 03, 2026
103K
118
10
38
152

TL;DR

This article provides a detailed guide on optimizing costs for Claude 5.5 models (Sonnet and Opus) by shifting focus from token prices to cost-per-successful-task. It covers effective effort routing, prompt caching strategies, and common migration pitfalls.

The Sonnet and Opus setup that matters: effort, cache, verification, and cost per finished task

The cheapest model is not the model with the lowest token price

It is the one that finishes the job, passes the check, and does not make you pay for the same context five more times

Sonnet 5.5 and Opus 5.5 make that distinction unusually important. One is priced for volume. The other is priced for harder work. Both have new effort behavior, and both can be surprisingly cheap inside a well-cached agent loop

Copy your old config into either model and the result may be slower, more expensive, or a 400 error

This is the setup I would build instead

text
1TASK → SONNET 5.5 → CHECK → OPUS 5.5 IF NEEDED → VERIFIED RESULT
2 ↘ effort ↗ ↘ cache + usage ledger ↗

I publish practical breakdowns of AI agents, workflows, and production systems on Substack

Join the newsletter here

The number that should decide your stack

Most model comparisons start with dollars per million tokens

Your agent does not ship tokens. It ships completed tasks

https://x.com/claudeai/status/2102435511222890900

Again: their tests. Your architecture needs your numbers

That means the check cannot be a vague thumbs-up after reading one impressive answer. For coding, use the test that would block the merge. For extraction, compare required fields against a labeled set. For research, record whether the cited source actually supports each claim. Include the cost of tasks that never pass, not just the nice examples in your demo

And inspect the hard tail separately. If the cheapest setting handles 90% of your requests but burns half the budget on the remaining 10%, its average can hide the part of the workflow that needs a different model

What the 5.5 price sheet actually says

As of October 3, 2026, at standard Claude API rates, per million tokens:

text
1SONNET 5.5
2Fresh input $2 Output $10
3Cache read $0.20 Cache write $2.50 / 5m, $4 / 1h
4
5OPUS 5.5
6Fresh input $4 Output $20
7Cache read $0.20 Cache write $5 / 5m, $8 / 1h

Both models have a 1M-token context window and a 128K-token maximum output. These are ceilings, not a reason to fill either one.

Model specifications and pricing

The odd line is cache read

Opus costs twice as much for fresh input and output, but a cached prefix costs the same $0.20 per million on both models. That does not make an Opus run equally cheap: it still pays more for new input, output, and cache writes. It does mean the model price gap can shrink in a read-heavy session

There is a second distinction people miss. Anthropic's "40% less than Opus 5" is an estimate of typical \run cost\. Opus 5.5's fresh token prices fell 20%; its cache-read price fell 60%. Those numbers are related, but they are not interchangeable

rari - inline image

Effort is a routing decision, not a quality slider

Sonnet 5.5 supports low, medium, high, xhigh, and max. On the API, it defaults to high, in Claude apps, Anthropic says medium is the default. Opus 5.5 defaults to medium on the API. These levels are not calibrated to mean exactly what the same words meant on the previous models.

My starting map:

  • Sonnet low For narrow, latency-sensitive requests with a cheap check
  • Sonnet medium For well-specified coding and routine multi-step work
  • Sonnet high When the medium run fails a real check, or the task has a demonstrated complexity pattern
  • Opus medium For ambiguous, cross-file, long-horizon work where Sonnet is spending turns circling the problem
  • Xhigh/max Only after your evals show a gain worth the extra time and tokens

This is a starting hypothesis, not a universal hierarchy. In Anthropic's reported FrontierCode results for Sonnet 5.5, xhigh scored above max. More effort is not a proof of a better outcome

Anthropic's footnote explains the counterintuitive result: at max, the model more often launched extra code review work. In two examined cases, that led to a timeout or edits beyond the task's scope.

The failure mode was not "the model did not think enough" It was spending effort in the wrong place. If your agent already passes its checks, extra review turns can become a cost and a source of new mistakes

https://x.com/edwinarbus/status/2104675431853248816

Also, do not set max_tokens low and call that optimization. The limit covers thinking and visible output. If you cut it mid-task, you can buy a truncated answer and a second run instead of saving money

https://x.com/claudeai/status/2104633115620823187

That is a compelling launch claim. A production config still needs to beat your own baseline

Run a small sweep before you invent a model router

Take 10-30 tasks you actually care about. Include easy work, ambiguous work, and the annoying failures from your logs. Give each task a verifier: tests, a structured comparison, a known answer, or a human rubric recorded before the run

Here is the smallest useful API probe. It logs the usage fields you need. Run it on each model and effort level against the same task, then attach your own pass/fail check. It is not a full agent benchmark

python
1import anthropic
2
3client = anthropic.Anthropic()
4response = client.messages.create(
5 model="claude-sonnet-5-5", # repeat with claude-opus-5-5
6 max_tokens=8192,
7 output_config={"effort": "medium"}, # repeat at high
8 messages=[{"role": "user", "content": "Replace this with a real task."}],
9)
10
11answer = "".join(b.text for b in response.content if b.type == "text")
12usage = response.usage
13print(answer)
14print("fresh", usage.input_tokens, "output", usage.output_tokens)
15print("cache read", usage.cache_read_input_tokens)
16print("cache write", usage.cache_creation_input_tokens)

This assumes the official Anthropic Python package and an ANTHROPIC_API_KEY environment variable. It is one isolated call without caching enabled, so zero cache reads and writes are expected. The next section shows what changes that

For a real agent, sum usage across all API calls under one task ID, including retries and tool calls. Count a pass only when the verifier says the job is finished

Compare total dollars per pass before choosing your default

Keep the test honest:

  • Freeze the task set and the verifier before comparing configs
  • Run the same tools, permissions, context, and output requirements on every candidate
  • Record pass rate, total spend, cost per pass, latency, and the longest or most expensive failures
  • Count stop_reason: "max_tokens" as an incomplete attempt, not a cheap success

The short code sample above uses an 8K output cap for a single-turn probe. Do not copy that cap into a long coding agent.

Anthropic recommends much more headroom for agentic work because hidden thinking counts toward the same limit.

Set the cap for the work, then control spend with effort, caching, and a task budget rather than forcing the answer to stop mid-job

Cache the stable part of the job

Agents repeatedly send the same system instructions, tool definitions, repository map, and prior conversation. If that prefix is stable, prompt caching changes the economics more than a minor prompt rewrite

For example, 200K cached tokens read 50 times are 10M cache-read tokens. At $0.20 per million, the reads cost $2 on either 5.5 model.

On Opus 5.5, sending those same 10M tokens as fresh input would cost $40. The first five-minute cache write of 200K tokens is another $1.

This is an illustration of prefix charges only: new input, output, other writes, TTL expiry, and actual cache misses add to the bill

The practical rules:

  • On the Claude API, opt in to prompt caching with top-level cache_control={"type": "ephemeral"} or explicit cache breakpoints. The probe above does neither, so its cache counters will normally stay at zero
  • Put stable instructions and tools before the changing user request
  • Keep the shared prefix identical across turns; verify actual cache_read_input_tokens
  • Treat a model switch as a new conversation budget, not a free continuation. The cache is per-model: an Opus request cannot read the prefix Sonnet just cached
  • Avoid changing top-level effort on every turn; that changes the rendered prompt and invalidates cached prefixes

On supported models, a per-message effort change can preserve the earlier cache, but it requires Anthropic's beta header and is not the same as changing the top-level output_config.

Sonnet 5.5 also has a between_tools caveat: effort cannot change mid-conversation in that mode.

Prompt caching documentation

Do not infer a cache hit from a fast response. Read the usage object. It separates fresh input, cache creation, and cache reads

Both 5.5 models need at least 512 tokens in a cacheable prefix. A tiny system prompt will not produce the savings in the example above. The default cache lifetime is five minutes, which fits a fast tool loop.

A one-hour write costs more and makes sense only when real sessions often pause long enough to miss the five-minute window. Measure those gaps before paying for the longer TTL

Escalate after evidence, not anxiety

Most teams build the router backwards: classify a task as "hard," send it to the expensive model, and never learn whether the cheaper path would have passed

Use the verifier as the routing signal

rari - inline image
text
11 Sonnet 5.5 · chosen effort → run the task
22 Verifier → accept if it passes
33 Opus 5.5 · medium → retry only with failure evidence
44 Verifier → accept or hand off with evidence

The check can be a test suite, schema validation, a known answer, or a reviewer. It should explain what failed.

"The answer feels weak" is a poor escalation signal, "the changed endpoint fails two integration tests" is useful

Do not blindly repeat an identical prompt. Give the next attempt the failed check, the relevant artifacts, and a specific instruction to correct the gap. Cap the ladder so an agent cannot burn its budget trying to fix a task that needs a human decision

You can test a Sonnet high retry in your offline sweep. Keep it in the live route only if it lowers cost per verified task. There is no reason to make every failure pay for two Sonnet runs before Opus

The model switch itself can break a cached prefix. Account for that when you compare the rescue path against an Opus-first route

The crossover is easy to miss. Suppose a Sonnet attempt costs $0.06 and passes 80% of your tasks.

If every failed task then costs $0.20 to finish on Opus, your illustrative average is $0.10 per finished task: $0.06 plus a $0.20 rescue on one task in five. That beats paying $0.20 for Opus on every task. But if Sonnet costs $0.14 and only passes half, the same ladder costs $0.24, even before you price the model switch. In that workload, Opus-first is cheaper and faster

Those figures are examples, not measured Claude results. Their purpose is to make the routing rule falsifiable. The ladder only earns its place if the saved Opus calls outweigh the failed Sonnet attempts, cache misses, and added latency

There is also a middle path: Anthropic's beta advisor tool. Sonnet can keep running the task and ask Opus for help on a difficult decision, instead of handing the entire job to Opus.

This is not automatically cheaper. Log how often Sonnet actually consults the advisor, what those calls cost, and whether they improve the final pass rate. If the executor rarely asks, the advisor is just an unused feature.

With these 5.5 models, the advice itself is returned encrypted to the client, so evaluate the resulting work rather than pretending you can audit the private advice text

Four leaks that raise the bill before model choice matters

Not every cost problem deserves a new router

Check these first:

  • Output that keeps growing On both 5.5 models, output tokens cost five times fresh input tokens. In a conversation, a long answer can also return as context on later turns. Ask for the artifact and a short completion note, not a narrated transcript of every step. Hidden thinking is billed as output too, so a terse final answer alone will not fix an effort problem. Do not suppress the evidence you need to verify the result
  • Images larger than the task needsSonnet 5.5 can process higher-resolution images than older Sonnet releases, which can raise image token counts. If the agent only needs a button label or one paragraph, crop or resize first. If it needs a dense chart or tiny UI details, keep the resolution and measure the cost instead of blindly shrinking it
  • Context nobody usesTool definitions, stale logs, old search results, and a sprawling CLAUDE.md can follow every request. Put durable rules in a short stable prefix; keep transient evidence near the task that needs it. Trimming context should not delete facts the model still needs to finish correctly
  • Interactive pricing for work nobody is waiting onThe Message Batches API discounts input and output by 50% on both models. It is useful for offline evaluations, document backfills, and other asynchronous jobs. It is not a replacement for a live tool loop where a person needs the next step now

The pattern is the same across all four: remove work the task does not need before buying more intelligence or lowering effort until quality breaks

The migration traps that turn a saving into a 400

Old request bodies are a bad starting point for the 5.5 family.

In particular:

  • Opus 5.5 thinking is always on Remove thinking: {"type": "disabled"} and old fixed budget_tokens settings; control depth with output_config.effort
  • Forced tool choice fails on both 5.5 models

tool_choice values any and tool return a 400. Use auto, specify when the tool should be used, and validate the tool result in your own code

  • Thinking blocks are not text blocks Read content by type, not content [0]. In tool loops, return thinking blocks unchanged with the assistant turn
  • Your UI may appear silent On Opus 5.5, between-tool progress can arrive in thinking blocks that are empty at the default display setting. If you previously rendered those notes to users, request a supported thinking display mode and render blocks by type Otherwise the agent may be working while the interface looks frozen
  • Old computer-use tool versions can fail Check the current tool version before migrating a browser/computer agent
  • A smaller max_tokens cap can cut off work

Thinking is included even when the text is hidden

These are API behavior changes, not prompt-writing tricks.

Opus migration guide and Sonnet migration guide

Put the contract in Claude Code, not in your head

The API is where you can measure every usage field. Claude Code is where many people will first feel the model change. The same principle applies: give the agent a bounded definition of done, then make it show the evidence

In Claude Code, /model selects the model and /effort selects the supported effort level. Check the active settings before comparing sessions. The Sonnet API default is not a reliable description of what your Claude app or Claude Code session is currently using

This is a complete, reusable CLAUDE.md starting block. Change the commands to match your project

markdown
1# Working contract
2
3Make only the requested change. Preserve unrelated work.
4Run the relevant tests after editing. Report any check you cannot run.
5Stop when the requested work passes. Do not add extra features or review loops.
6End with: Changed / Verified / Remaining risk.
7Ask before destructive actions, publishing, or changes outside this repository.

That block will not magically make every run cheap. It makes success and failure visible. From there, you can compare a Sonnet-first workflow with an Opus-first workflow on the same tasks

The task message still has to be specific. Here is the difference between "fix the payments code" and a job an agent can actually finish

text
1Change: migrate the payment endpoint to the new client
2Done: old client removed, endpoint tests pass, diff limited to this path
3Stop: ask before deleting data or changing anything outside the repo
4Report: changed files, exact checks run, remaining risk

That small contract gives the verifier something concrete to inspect. It also gives the model a reason to stop. An open-ended instruction to "review until perfect" can turn a passing change into another paid loop

For long projects, keep the checklist in a file that survives compaction. For subagents, ask the lead agent to inspect their evidence before accepting their reports. And if you only asked for ideas, tell Claude not to start building. Those are workflow boundaries, not "be smarter" prompts

The setup I would ship first

  • Pick 10-30 real tasks and define a check for each
  • Sweep Sonnet 5.5 at medium and high, then Opus 5.5 at medium
  • Log fresh input, output, cache writes, cache reads, latency, retries, and pass/fail per task
  • Keep the stable prefix cacheable and confirm hits in usage
  • Route only the failures upward, with their evidence attached
  • Revisit the ladder when the workload changes. A saved benchmark is not a permanent truth

If the difficult 10% repeatedly goes straight from Sonnet failure to Opus success, consider routing that recognizable task class to Opus from the start. If Sonnet high passes those same cases for less, keep them there. A router is a measured policy, not a permanent opinion about which model is smarter

The 5.5 upgrade is not just "use Sonnet for cheap work and Opus for hard work"

It is a chance to stop pricing the model and start pricing the finished job

If u read this far

-> Subscribe to my Substack

-> Join my Telegram

-> Bookmark this article

-> Follow @0xwhrrari

One-click save

Use YouMind for AI deep reading of viral articles

Save the source, ask focused questions, summarize the argument, and turn a viral article into reusable notes in one AI workspace.

Explore YouMind
For creators

Turn your Markdown into a clean 𝕏 article

When you publish your own long-form writing, images, tables, and code blocks make 𝕏 formatting painful. YouMind turns a full Markdown draft into a clean, ready-to-post 𝕏 article.

Try Markdown to 𝕏

More patterns to decode

Recent viral articles

Explore more viral articles