YouMind
تسجيل الدخول

Opus 4.8: Same Price, You Pay Double

@shmidtqq
الإنجليزية30 مايو 2026
1.2M
215
18
33
516

ليرة تركية؛ د

Anthropic's Opus 4.8 introduces significant operational upgrades including effort control and dynamic workflows. While benchmarks show modest gains, the real value lies in optimizing token usage and speed for complex coding tasks.

Same price. Most people will just swap the model and miss everything else.

shmidt - inline image

Alongside Opus 4.8, Anthropic shipped three things that change how you work in Claude Code more than the benchmark numbers do: effort control, dynamic workflows, and cheaper fast mode. The people who configure these properly get better results and spend less. Here's the breakdown.

The one number that matters most

Not the 69.2% on coding. It's that 4.8 is about 4x less likely to let flaws slip through in code it wrote itself.

Older models loved to confidently say "done, bug fixed" with thin evidence. Anthropic leaned on honesty: 4.8 slows down more often and flags where it isn't sure, instead of generating plausible code that silently breaks edge cases. Over a long session this compounds. A model that admits uncertainty on turn 15 saves you a couple hours of debugging on turn 40.

What changed: the short version

Coding.

SWE-bench Verified: 87.6% → 88.6% (near the ceiling).

SWE-bench Pro (harder, less contaminated): 64.3% → 69.2%, 10+ points ahead of GPT-5.5.

This is the headline coding number.

Terminal.

Terminal-Bench 2.1: 66.1% → 74.6% (+8.5), but GPT-5.5 still leads (78.2%). The one benchmark of seven where 4.8 loses.

Knowledge work.

GDPval-AA: 1890 Elo vs 1769 for GPT-5.5. The widest gap on the whole table.

Computer use.

OSWorld-Verified: 83.4% (part of the gain is an updated harness, which Anthropic flags openly).

Science.

GPQA Diamond: 93.6%, essentially flat (both models sit above 93%, that's saturation, not a regression).

shmidt - inline image

Anthropic itself calls the release "a modest but tangible improvement." The benchmarks are a plus. The operational changes below are the real story.

Feature 1. Effort control (the most underrated one)

Opus 4.8 defaults to High effort. But now you decide how much thinking it puts into a task. Think of it as a manual transmission for reasoning.

In claude.ai and Cowork the slider sits next to the model picker, with five levels: Low, Medium, High (Default), Extra, Max, plus a separate Thinking toggle. In Claude Code it's the same levels with slightly different names, and there's an auto mode:

Note: "Extra" in claude.ai and xhigh in Claude Code are the same level, just different labels on different surfaces.

shmidt - inline image

The money part, so there's no confusion. There is NO separate surcharge per effort level. The rate is the same at every level: $5 per 1M input tokens, $25 per 1M output. The difference isn't the per-token price, it's how many tokens the model spends. Low answers briefly and thinks little, Max thinks long and burns several times more tokens. So Max is "pricier" because of volume, not because of the rate.

A rough guide to relative spend on the same task (numbers illustrative, for intuition):

A money example. Say a High answer costs ~$0.05. The same request at Low runs around ~$0.005, and at Max it can easily hit ~$0.20-0.40. If 60% of your prompts are small ("what does this function return?"), moving them from High to Low cuts daily spend several times over, with no quality hit where it counts.

Why this hits the bill hardest. Most people will leave everything on High (or crank it to Max "to be safe") and never touch the slider. The ones who route effort per task get the same results for noticeably less. That's the most underrated feature in the release.

Feature 2. Fast mode (3x cheaper)

Fast mode runs Opus at 2.5x speed with the same quality, and it's now 3x cheaper than before.

Toggle it in Claude Code with /fast (an active session is marked with a ↯ icon). API access is gated for now (waitlist at claude.com/fast-mode).

Use fast mode for: large multi-file refactors, code generation from specs, documentation, test generation. Anywhere speed beats depth. Stay in standard for: complex debugging, architecture decisions, security review. Anywhere thinking quality beats speed.

Feature 3. Dynamic workflows (the big one)

Claude Code now writes a JavaScript orchestration script for your task and runs it in the background while your session stays responsive. The plan lives in code, not the model's context, so only the final answer comes back to your session.

What to know, from the official docs:

  • Up to 16 subagents concurrently, a hard ceiling of 1,000 total per run.
  • Resumable: if your laptop dies or you close the terminal, the run picks up where it stopped.
  • Requires Claude Code v2.1.154+. Research preview.
  • On by default on Max, Team, Enterprise. On Pro, turn it on in /config (and switch to Opus 4.8 first).
  • Subagents run in acceptEdits mode and inherit your tool allowlist.

Trigger it with /effort ultracode (a Claude Code setting: xhigh plus automatic workflow orchestration), or just describe a big task in words:

shmidt - inline image

Good for: migrations across 200+ files, full-codebase security audits, project-wide test generation, large refactors, research across multiple repos. Not good for: simple bug fixes, single-file edits, quick questions. Overkill.

On cost. Workflows consume meaningfully more tokens than a normal session. The official way to keep spend in check is to scope the task: run an audit on one folder first, see how many subagents it spawns, then calibrate the big run from there.

To disable workflows entirely:

Example Claude Code config (starter template)

Environment variables (in ~/.zshrc or ~/.bashrc):

Then comes the useful part: a settings.json with safe permissions. It lets the model read and edit code, run tests, and commit, but hard-blocks the dangerous stuff: reading .env and SSH keys, rm -rf, sudo, and git push. The agent works freely but can't break anything or leak anything.

shmidt - inline image

The ready-to-paste settings.json file is in my Telegram channel (the format is too big to read cleanly here): t.me/+JmDeelv5UCwwMTcy

shmidt - inline image

Daily command cheat sheet

Installing Opus 4.8: three ways

A. Browser or app. On claude.ai, pick Claude Opus 4.8 in the model list and set the effort slider. No install needed.

B. API. Model ID claude-opus-4-8. Price $5/$25 per 1M. Context up to 1M (200k on Microsoft Foundry), up to 128k output. Fixed-budget extended thinking isn't supported: use adaptive thinking (thinking: {type: "adaptive"}) and the effort parameter.

C. Claude Code (terminal). The native installer is now the recommended method (npm is legacy):

Then run claude, sign in through the browser, and pick Opus 4.8 with /model.

Migration checklist: 4.7 to 4.8

  • Change the model ID to claude-opus-4-8
  • Run 10-20 of your real tasks on 4.7 and 4.8 with identical prompts
  • Compare: completion rate, step count, tokens, test pass rate
  • Check prompts tuned for 4.7's behavior
  • Assign effort levels per task type
  • Test fast mode where speed matters
  • Update Claude Code to v2.1.154+ (for workflows)
  • Recheck the API bill after week one

Save card: everything in one place

Model: claude-opus-4-8 · shipped 2026-05-28

Price: $5 / $25 per 1M · fast mode $10 / $50

Context: up to 1M · up to 128k output

Key: SWE-bench Pro 64.3% → 69.2% (+10 over GPT-5.5) · SWE-bench Verified 88.6% · 4x fewer missed bugs · GDPval-AA 1890 Elo

New: effort control · fast mode 3x cheaper · dynamic workflows (16 concurrent / 1,000 per run)

Thanks for reading. If you take one thing from this, make it "route effort per task."

The rest of my notes on Claude and vibe coding live here: t.me/+JmDeelv5UCwwMTcy.

See you there.

@shmidtqq.

بنقرة واحدة حفظ

استخدم YouMind للقراءة العميقة للمقالات سريعة الانتشار بتقنية الذكاء الاصطناعي

احفظ المصدر، واطرح أسئلة مركزة، ولخص الحجة، وحوّل المقالة واسعة الانتشار إلى ملاحظات قابلة لإعادة الاستخدام في مساحة عمل واحدة تعمل بالذكاء الاصطناعي.

اكتشف YouMind
للمبدعين

حول Markdown إلى مقالة 𝕏 نظيفة

عندما تنشر كتاباتك الطويلة، فإن الصور والجداول وكتل التعليمات البرمجية تجعل تنسيق 𝕏 مؤلمًا. YouMind يحول مسودة Markdown كاملة إلى مقالة نظيفة وجاهزة للنشر 𝕏.

حاول Markdown إلى 𝕏

المزيد من الأنماط لفك التشفير

المقالات الفيروسية الأخيرة

استكشاف المزيد من المقالات الفيروسية