Have you ever encountered this situation?
The same Claude, the same GPT-4o—one person uses it to write 1 million lines of code in 5 months, while another can't even get it to run stably for two hours.
The models are identical, but the results are worlds apart.
Where is the problem?
I recently read a bunch of articles from OpenAI, Anthropic, Martin Fowler, and Phil Schmid, and I found they are all talking about the same thing.
They call it Harness Engineering.
Simply put, it's building an "operating system" for your Agent.
First, Understand What a Harness Is

Phil Schmid made a great analogy in a HuggingFace blog post.
Think of an Agent system like a computer.
The model is the CPU, providing raw computing power. The context window is the RAM, storing things temporarily. The Agent is the application running on top.
So, what's the operating system?
The Harness is the operating system.
Without an OS, even the most powerful CPU is just a chip. You can't type on a chip.
Similarly, without a Harness, even the smartest model is just a chat box. If you let it run a complex task for an hour, what if it forgets the context? Who stops it from writing garbage code? What if it makes a mistake and doesn't even know?
These aren't problems you solve by "switching to a smarter model."
Martin Fowler said something that stuck with me: Harnesses might become "service templates" in the future. Just as you start a new project today with a service template, you'll start a new Agent with a Harness template.
I think this prediction is likely to come true.
Why is it Suddenly Exploding in 2026?

Because models are strong enough now.
In 2024, everyone was competing on whose model was smarter. By 2026, the gap between top-tier models has become very small. If you give Claude and GPT the same problem, their scores are only a few points apart.
But if you let them work for 8 hours straight, the gap appears.
This gap isn't in the model itself; it's in the "harness" surrounding it.
OpenAI's Codex team has a staggering statistic. They used Codex to build a complete product—5 months, 1 million lines of code, zero lines handwritten. Throughout the process, they found the bottleneck was no longer "can the model write code."
The bottleneck was whether humans could review the code fast enough.
Model output speed has surpassed human review speed. At this point, what's the use of optimizing the model? You should optimize the review process, quality control, and architectural constraints.
That's what the Harness does.
The Three Pillars

So, what does a Harness actually contain?
After reading these articles, I found that while terms vary, there are three core pillars.
1. Evaluation Closed-Loop
This is what Anthropic emphasizes most.
The core idea is simple: An Agent cannot grade itself.
Think about it: if an intern finishes a report and you ask them how they did, they'll say "it's okay." You need an independent person to evaluate.
Anthropic calls this "Evaluation-Driven Development." First define what "doing well" looks like, then let the Agent do it, and finally have an independent evaluator score it.
Evaluation-Driven Development is the Agent version of TDD. Write the tests first, then the code. Except here, the "tests" are for the Agent.
The evaluator doesn't just look at the code. It actually operates the product—using Playwright to click buttons, fill forms, and run tests—then judges based on clear standards.
There's a fascinating case here.
Anthropic's Opus 4.5 found a loophole in a booking policy during a flight reservation test, finding a solution better than the standard answer.
But the evaluator marked it as a "failure."
Why? Because the evaluator didn't expect such a creative solution. There was only one standard answer, and because the Agent found a better one, it was penalized.
This story shows two things: first, Agents are smart enough to find solutions humans haven't thought of. Second, the evaluation loop isn't just checking the Agent; it's also checking the evaluation itself. If your evaluator is too rigid, it becomes the bottleneck.
Another data point: Opus 4.5 initially scored 42% on CORE-Bench. After they fixed scoring bugs and relaxed scaffold constraints, the score jumped to 95%.
Often, it's not that the model isn't good enough; it's that your Harness has issues.
Using this method, Anthropic had an Agent build a complete game in 6 hours for $200.
2. Architectural Constraints
This is the OpenAI Codex team's specialty.
You tell an intern "the code needs to be layered," they nod, then immediately write UI logic into the database layer.
Talking is useless.
OpenAI's approach is to enforce it mechanically via linters and CI. Code that violates architectural rules is rejected immediately, without even getting a review.
Their code layering looks like this: Types → Config → Service → UI. Each layer can only depend on the layer above it, never the other way around. This rule isn't just written in a document; it's written in a linter for automatic checking.
Even better, these linters themselves are generated by Codex.
The Agent writes its own rules and then follows them.
Martin Fowler said after reading OpenAI's article:
"Increasing trust and reliability requires constraining the solution space. This means giving up some of the flexibility to 'generate anything.'"
The more constraints, the more reliable.
It sounds counterintuitive, but the data speaks. LangChain did an experiment: without changing the model, they only changed the Harness, and the Terminal Bench 2.0 pass rate jumped from 52.8% to 66.5%. Vercel went further, deleting 80% of Agent tools, resulting in fewer steps, faster speeds, and better results.
Fewer tools often lead to better performance—this conclusion has been repeatedly verified in the Agent field.
3. Memory Governance
This pillar is discussed less, but I think it's the most important in the long run.
PrismerCloud has done deep work in this direction.
The problem is: when multiple Agents share a knowledge base, Agent A writes an experience, and Agent B reads it as truth. But what if Agent A was wrong?
One Agent's hallucination can pollute all Agents through the shared knowledge base.
PrismerCloud's approach is to build an "Evolution Engine." Every Agent experience is first recorded as a "signal." Once verified, signals are distilled into "genes," which are continuously optimized based on actual results.
Simply put, genes are verified, effective knowledge. If it's not verified, it doesn't count.
There's an interesting stat: 3 lines of prompt plus a memory system performs roughly as well as 200 lines of carefully crafted expert prompts. Moreover, the former evolves, while the latter is static.
This means if your memory system is good, you don't need complex prompts. The Agent will naturally improve over time.
Bonus: Entropy Resistance
This isn't a standalone pillar but is worth mentioning.
Agent systems naturally decay over time. Documents expire, architectures are bypassed, and knowledge bases fill with outdated info.
OpenAI's approach is to periodically run a "Refactoring Agent" to scan for document inconsistencies and architectural violations. They said it best:
"When an Agent struggles, we treat it as a signal: find out what's missing, feed it back into the codebase, and always let Codex write the fix."
When an Agent has problems, don't just fix the Agent—fix the Harness. This mindset is key.
Who is Doing This?

The field is split into two paths: open-source projects you can use today, and internal practices of commercial companies where you can only learn the methodology.
Open Source Projects: Ready to Use
LangChain DeepAgents: Likely the closest open-source project to a "universal Claude Code." Planning, file operations, sub-agent delegation, automatic context compression—ready out of the box. 115k stars on GitHub.
DeerFlow 2.0: From ByteDance. Open-sourced in March, it hit 39k stars in a month. It calls itself a "SuperAgent Harness." It's a complete rewrite from v1 with sandbox execution, persistent memory, and skill systems based on LangGraph.
OpenHands: Specialized for coding Agents. It hit 77.6% on SWE-bench Verified. It's model-agnostic and uses Laminar for observability, tracing every Agent action.
SWE-agent: From Princeton and Stanford. It focuses on perfecting "evaluation-driven" development.
Goose: Open-sourced by Block (Square/Cash App). A general on-machine Agent that can install dependencies, run tests, and manage files.
PrismerCloud: Focuses on memory governance and the evolution engine. It's the most mature solution for preventing hallucination pollution in multi-agent systems.
Cognee: A knowledge-graph-driven memory engine for Agents that helps establish semantic connections between data.
Commercial Practices: Learn the Methodology
Claude Code + Agent SDK: Anthropic's benchmark for a general Harness. It's not just for coding; they use it for research, video creation, and note-taking.
OpenAI Codex: The ultimate practice in architectural constraints. 1 million lines of code with zero handwriting, relying on auto-generated linters and Agent peer reviews.
A Lesson That Stuck With Me

Rich Sutton wrote a classic paper called "The Bitter Lesson." The gist is that general methods leveraging computation always beat human-designed specific methods in the long run.
This lesson is being proven again in the Agent field.
Manus refactored its Harness 5 times in 6 months. LangChain re-architected 3 times in a year. Vercel deleted 80% of its tools.
Build to Delete.
The "clever logic" you write today might be obsolete tomorrow when the model upgrades. Your architecture must be modular and ready to be scrapped.
Phil Schmid said something worth remembering:
"Competitive advantage is no longer the prompt; it's the trajectories captured by your Harness. Every success and failure is data for training the next generation."
The longer your Harness runs and the more trajectories it accumulates, the stronger your Agent becomes. You can't catch up just by switching models.
The Three Stages

Think of the Harness's place in AI engineering like this.
Prompt Engineering solves "what to say." A single interaction.
Context Engineering solves "what to know." Providing references and history.
Harness Engineering solves "how to work continuously, stably, and at scale." Evaluation loops ensure quality, architectural constraints ensure rules, and memory governance ensures experience accumulation.
Without a Harness, an Agent might remember things but has no oversight, leading to chaos. When all three layers are in place, you have a character that can truly work long-term.
OpenAI, Anthropic, and LangChain are already doing this.
Sources: OpenAI Harness Engineering, Anthropic Demystifying Evals, Phil Schmid (HuggingFace) The Importance of Agent Harness in 2026, Martin Fowler Harness Engineering, LangChain Agent Frameworks.





