Which models earn their keep? We ran 14 models on real knowledge work to find out.

@andrewbusse
İNGILIZCE1 gün önce · 23 Tem 2026
166K
11
3
1
8

TL;DR

This analysis of 14 AI models reveals that benchmark leaders often fail in real-world knowledge work. It provides a framework to choose models based on judgment and frequency, categorizing them into sprinters, middleweights, and heavyweights.

The best model for a job is almost never the best model on the leaderboard.

In the AI world, Christmas comes about once a week. A new model drops, everyone stares at the same false-precision benchmark charts, and we all agree it's the best thing since sliced bread, until next week, when we do it again.

But none of those charts answer the only question that matters. Can the model actually run your daily briefing or send a quote to a customer? At Hyperagent, we give you access to all of these models on top of a best-in-class agent. So we ran them ourselves. Fourteen models, forty-two runs, on the real knowledge work our customers hand to agents every day. Three things stood out.

  • The bill ranged 50x on identical work. The same test battery cost $1.48 on Qwen 3.7 Plus and $75.16 on Fable 5.
  • The benchmark winner wasn't the work winner. GPT 5.6 Sol topped our baseline tests but didn't crack the top three once it was running real jobs inside Hyperagent. Grok 4.5 was middle-of-the-pack on baseline and came out first on finished work, speed, and cost.
  • And no model won everything. The right choice moved with the job.

The takeaway isn't a new favorite model. It's a way to decide, job by job, which one earns its cost.

The AI model map for knowledge work

Every job you'd hand an agent comes down to two questions: how much judgment does it need, and how often does it run? Plot those two axes and knowledge work falls into four territories, each of which wants a different kind of model.

Andrew Busse - inline image

Volume work (top left). Triage, monitoring, routing, data sync. Clear procedure, runs constantly, and a miss is cheap to fix. This work wants something fast and inexpensive.

The daily craft (top right). Drafting, reporting, research, client comms. The recurring knowledge work that fills most of the week - it needs real synthesis and a good writing voice, done reliably and often.

High stakes (bottom right). Contract review, pricing calls, deep dives. Rare, but a wrong answer costs a client, a contract, or a quarter's margin. This is where the downside dwarfs the model bill, and where the priciest models earn their rate.

Don't optimize (bottom left). The occasional, simple ask where any capable model clears the bar and the price gap is too small to deserve a decision. It's the one square on the map we tell people not to think about.

The three model classes

Three working classes emerge from the map. They describe the economics and judgment profile of the job rather than grading a model as good or bad.

Sprinters are built for volume work - fast, cheap, and consistent when the procedure is clear and a miss is easy to fix. Haiku 4.5, Gemini 3.5 Flash, and DeepSeek V4 Pro.

Middleweights run the daily craft. They synthesize sources, hold a multi-step plan, and write well without charging top-tier rates for every report. Sonnet 5, GPT 5.6 Terra, Grok 4.5, and GLM 5.2 - where most knowledge work belongs.

Heavyweights take the high-stakes calls, bringing stronger judgment and more reasoning time to work where a wrong answer costs far more than the model bill. Opus 4.8, GPT 5.6 Sol, and Fable 5.

The instinct most people have is to start high and work down - run everything on a Fable 5 or an Opus 4.8, get the output you want, then try cheaper models until quality drops off. It works, but we've found most jobs don't need it. Once you can name the job and place it on the map, you can usually start in the right class from the beginning. For the daily craft, which is most of what you do, that means starting on a solid middleweight like Sonnet 5. You only reach for the top-down search when a job genuinely sits at the high-stakes end.

The field is wider than Claude, GPT, and Gemini

Most teams default to the handful of model names they already know. That's a reasonable place to start, but it leaves a lot of the roster untouched - and some of the models doing the most useful work in Hyperagent aren't the household names. These are our operating impressions from the battery and from daily use, plus the work we've watched customers run.

Grok 4.5 is fast, real-time research. It moves quickly through tool-heavy research and performed especially well paired with Hyperagent's native Exa Search. It also has a native line into X, so a question about live sentiment can come back with the actual posts behind the answer. It led our finished-work score at $9.71 for the full run.

Muse Spark 1.1 writes clean and checks itself. It finished second in the harness, and our early read is concise, well-structured prose with a useful habit of auditing its own work before it hands it back. It's still new, so we'd start it on supervised everyday jobs before the long unattended ones.

Kimi K2.6 is deep research at an open-model price. It runs searches in parallel and pulls the evidence back into a single answer without charging flagship rates, which makes it a strong research specialist. Its one real limit is document size, with a context window that tops out at 256,000 tokens.

GLM 5.2 is the value pick. It came surprisingly close to the closed middleweights on routine research, drafting, and long-document work at roughly a third of the price, and its prose is among the strongest outside Sonnet and Opus. It's become a daily driver for a lot of the Hyperagent team.

Qwen 3.7 Plus is the screen worker. It's particularly good at finding buttons, fields, and menus on rendered pages, and it's inexpensive for structured extraction - it produced the lowest-cost complete run in our battery. Reach for it when the job is working an interface or moving data, not when voice matters.

DeepSeek V4 Pro is structured analysis at very low cost. It's strongest on reconciliation, data cleanup, scheduled number work, and similar structured tasks. Its writing is literal and terse, which serves internal explanations well and client-facing prose poorly. A specialist, and a very economical one.

Staffing your always-on agents at the right price

Steve Juba runs a six-person travel company. He built a briefing agent connected to eight live data sources that runs several times a day, and it cut his morning context-gathering from more than three hours across eight tabs to fifteen minutes.

An agent like this reads far more than it writes and does the same job hundreds of times a year, so a small per-run price difference stops being a rounding error and becomes a real budget line. Consider a briefing that reads 100,000 tokens and writes 2,000 every day. At July 2026 list prices, that shape of work runs roughly:

  • $16 a year on Qwen 3.7 Plus (sprinter)
  • $38 a year on Kimi K2.6 (middleweight)
  • $40 a year on Haiku 4.5 (sprinter)
  • $80 a year on Sonnet 5 (middleweight)
Andrew Busse - inline image

The open-weight options change the math on read-heavy work like this. Kimi K2.6 does middleweight-quality research and synthesis for $38 on this job - less than half of Sonnet 5, and about the same as running the Haiku sprinter. You're not trading judgment for cost; you're getting middleweight work at sprinter money. If the job were pure extraction with no judgment, Qwen would run it for $16 - but the moment it needs real synthesis, Kimi gives you that for close to the same price.

Swap the model, keep the agent

None of this matters if switching models means rebuilding the agent. Inside Hyperagent it doesn't, because the model is just one setting and the role sits above it. Change the model and the agent keeps everything that makes it useful:

  • Its instructions and scope
  • Its skills, scripts, and working procedures
  • Its memory of corrections
  • Its reference files and company context
  • Its connections to the systems where the work lives
  • Its rubrics for evaluating quality
Andrew Busse - inline image

That turns model choice into a test rather than a migration. The same agent can run the same job on two models without being rebuilt, so you can find the cheapest one that clears your bar instead of guessing.

It also lets the expensive part of a hard workflow pay for itself once. When a job genuinely needs it, use a heavier model to design and debug the workflow - then codify the procedure as a skill so a middleweight or sprinter can run it every day at a fraction of the cost.

There's a hidden cost to watch, though. A cheaper model with a low bill isn't cheap if every output needs correcting - the human review time and the cost of a miss are part of the real price. Call it the review tax. It sets the floor on how light you can go. What's different in Hyperagent is that the review doesn't just get paid, it compounds: as the agent learns from your corrections through memory, that review effort is captured by the harness and produces a better agent on the next run.

Customers already staff agents this way. Ankit Das built a marketing operation with a Growth Orchestrator on Sonnet 5 making the judgment calls and eight channel specialists on GLM 5.2 producing the recurring channel work, with separate QA agents checking brand compliance and writing quality - all kicked off from a single thread. Eli Weiss describes his split more simply: roughly 85% Sonnet and 15% Opus across his skills and routines. Most work runs on the daily driver; Opus gets the smaller set of jobs where the extra judgment earns its cost.

That's the practical value of model choice inside Hyperagent. Spend deliberately across the work, and reserve the extra dollars for the jobs where extra judgment changes the outcome.

Choose the model after you choose the job

Use this sequence on an agent you already run:

  1. Name the job. “Prepare the Monday client report” is more useful than “help with reporting.”
  2. Place it on the map. Decide how much judgment it needs and how often it will run.
  3. Start where the map puts it. For most work that's a capable middleweight - use it until you understand the job's edge cases and codify the process into skills.
  4. Test one class lighter for five real runs. Keep the instructions, skills, memory, sources, and rubric unchanged.
  5. Compare the whole cost. Track finished quality, consistency, time, factual misses, model bill, and human review time.
  6. Escalate on purpose. Bring in a heavyweight when the review burden or cost of a wrong answer exceeds the model premium.
  7. Use a different model to review important work. The model that made the mistake is often the least likely to see it.

Keep the least expensive model that reliably clears your standard. Then spend the difference on the calls that deserve it.

YouMind’da yeniden üret

Turn one viral article into a full content workflow

Collect the source, decode the pattern, create assets, draft the story, and distribute from one AI workspace.

Explore YouMind
Üreticiler için

Markdown'ınızı temiz bir 𝕏 makalesine dönüştürün

Kendi uzun yazılarınızı yayımlarken görselleri, tabloları ve kod bloklarını 𝕏 için biçimlendirmek zahmetlidir. YouMind, eksiksiz bir Markdown taslağını temiz ve hemen paylaşılabilir bir 𝕏 makalesine dönüştürür.

Markdown'dan 𝕏'e deneyin

Çözülecek daha fazla kalıp

Son viral makaleler

Daha fazla viral makale keşfet