Most people are trying to fix AI agents at the wrong layer.
When an agent fails, they rewrite the prompt. When it fails again, they add more instructions, switch models, increase the context window, or connect another tool.
Then the same problems come back.
The agent forgets an important decision. It uses the wrong tool. It loses track of what happened three steps earlier. It claims the task is finished without checking the result. It retries the same failed action until the budget is gone.
The problem is not always the model.
The problem is the environment around it.
That environment is the harness.
Harness Engineering is the practice of building the system around a model that decides what it can see, what it can do, what it remembers, what counts as success, and what happens when something fails.
A better prompt can improve one response.
A better harness improves every run.
Follow my Substack for more practical breakdowns of AI agents, automation and production systems:
1. The Model Is Not the Agent
A model can reason, generate, compare and choose.
That does not make it a reliable agent.
A real agent also needs to find the right context, use tools, preserve state, respect permissions, verify its own work and recover when the environment behaves differently than expected.
The model is only the reasoning engine.
The harness is everything that turns that reasoning into real execution.
1USER REQUEST2 |3 v4+-----------------------------+5| HARNESS |6| |7| contract context |8| tools state |9| policy verification |10| traces recovery |11+-----------------------------+12 |13 v14 MODEL15 |16 v17REAL ENVIRONMENT
Put the same model inside a chat box and it answers questions.

Put it inside a repository with terminal access, tests, browser tools, project memory, controlled permissions and a review loop and it can complete real work.
The model did not change.
The harness did.
2. Turn Every Request Into a Contract
Natural language is flexible.
Autonomous execution should not be.
A request like:
Improve the onboarding flow.
is fine when a human is sitting next to the model.
It is terrible as a production instruction.
Before the agent acts, turn the request into a bounded task contract.

1objective: reduce onboarding drop-off23inputs:4 - product brief5 - analytics data6 - repository78constraints:9 - preserve authentication10 - do not change database schema11 - preserve current mobile behavior1213deliverable:14 - reviewable pull request1516done_when:17 - tests pass18 - analytics event fires correctly19 - desktop flow passes review20 - mobile flow passes review2122approval_required:23 - production deployment
The important part is done_when.
Without it, the agent can solve a slightly easier version of the problem and still confidently say the task is complete.
With it, completion becomes measurable.
The agent should not ask:
What should I do next?
It should ask:
What action moves the current environment closer to the contracted outcome?
That is a much stronger loop.
3. Give the Agent a Map, Not a Giant Context Window
A common reaction to agent mistakes is to give the model more context.
More documentation.
More conversation history.
More files.
More tool output.
Eventually the agent receives everything and understands less.
Context is not storage.
It is an attention budget.

Instead of dumping the entire project into every run, give the agent a small map of where useful information lives.
1PROJECT MAP23product rules -> docs/product/4architecture -> docs/architecture.md5frontend -> apps/web/6backend -> services/api/7tests -> tests/8commands -> docs/commands.md9security -> docs/security.md
Then expand only when needed.
1TASK2 |3 v4PROJECT MAP5 |6 v7RELEVANT SYSTEM8 |9 v10EXACT FILES11 |12 v13LOCAL INSTRUCTIONS
The source material describes this as progressive disclosure: the harness should load more information because the task needs it, not simply because the information exists.
The goal is not maximum context.
The goal is maximum useful signal.
4. Put a Gateway Between the Model and Its Tools
A model with twenty tools is not automatically twenty times more capable.
It may just have twenty more ways to fail.
Every tool should have a contract.
1TOOL: edit_file23INPUTS4path5patch67PRECONDITIONS8path exists9path is inside workspace1011SUCCESS12patch applied13diff returned1415FAILURE16structured error17no partial overwrite1819RISK20reversible
Then the execution path becomes:
1MODEL PROPOSES2 |3 v4GATEWAY VALIDATES5 |6 v7POLICY AUTHORIZES8 |9 v10TOOL EXECUTES11 |12 v13HARNESS RECORDS RESULT
The model decides what action it wants.
The harness decides whether that action is valid, permitted and safe.

That distinction becomes critical when tools can send messages, modify production, spend money or delete data.
A good tool gateway can also add timeouts, validate arguments, restrict file paths, normalize errors and make retries safe.
Good tools reduce the number of things the model has to guess.
5. Move Memory Outside the Conversation
The conversation should not be the system of record.
Long-running agents eventually hit context limits, crash, restart, or hand work to another session.
If every important decision exists only inside the transcript, the workflow is fragile.
Store durable state separately.

1{2 "task_id": "feature_042",3 "status": "verifying",4 "current_step": "mobile_check",56 "completed": [7 "implementation",8 "unit_tests",9 "desktop_check"10 ],1112 "decisions": [13 "reuse existing export endpoint",14 "preserve current date format"15 ],1617 "artifacts": [18 "export.csv",19 "desktop-after.png"20 ],2122 "open_risks": [23 "mobile toolbar may overflow"24 ],2526 "next_action": "render mobile viewport"27}
A useful system separates memory into four categories:
1FACTS2stable knowledge34DECISIONS5what was chosen and why67STATE8where the current run is910LESSONS11failures that should affect future runs
The next agent session should inherit the state of the work, not a compressed story about the previous conversation.
6. Make Evidence the Gate to Completion
An agent saying "done" is not proof that the job is done.

It is another model output.
The harness needs observable evidence.
1CLAIM EVIDENCE23"bug is fixed" failing test now passes45"page works" browser flow completed67"data is correct" values match source89"migration is safe" dry run + rollback pass1011"task is complete" every acceptance check passes
Use deterministic checks first.
1syntax2 |3 v4types5 |6 v7focused tests8 |9 v10integration tests11 |12 v13visual / semantic review14 |15 v16human approval
Do not ask another model to answer something a compiler, test, schema or database query can prove.
Use models for judgment.
Use deterministic systems for facts.
The model creates the artifact.
The environment creates evidence about the artifact.
The harness decides whether the evidence is sufficient.
7. Separate the Builder From the Verifier
There is another problem with self-review.
The agent that created the mistake often carries the same assumptions into the review.

A stronger architecture separates the worker and the verifier.
1BUILDER2 |3 v4creates candidate5 |6 v7VERIFIER8 |9 +-- checks contract10 +-- searches for missing cases11 +-- tests unsupported claims12 +-- tries to break result13 |14 +------ PASS ------> ACCEPT15 |16 +------ FAIL ------> RETURN EVIDENCE
The verifier should not ask:
Does this look good?
It should ask:
What would make this unacceptable?
That changes review from confirmation into attempted disproof.
The source material explicitly recommends giving verification its own rejection criteria and enough independence to challenge the assumptions that produced the first result.
8. Move Permissions Outside the Model
Some rules should never depend on the model remembering them.
1never publish without approval2never expose secrets3never exceed the spend limit4never write outside the workspace5never claim tests passed unless they ran
These are not prompt suggestions.

They are policy.
A simple permission ladder:
1LOW RISK23read4search5inspect67-> automatic89REVERSIBLE1011edit workspace12run tests13create draft1415-> automatic + trace1617EXTERNAL EFFECT1819send20deploy21purchase2223-> approval required2425IRREVERSIBLE / SENSITIVE2627delete data28rotate credentials29publish globally3031-> hard gate or prohibited
The stronger the consequence, the stronger the control.
The model can recommend the action.
The harness authorizes it.
The tool executes it.
Autonomy is not the absence of control.
It is freedom inside an enforced boundary.
9. Stop Retrying Blindly
One of the worst recovery policies is:
Something failed. Try again.
If nothing changes, the system is simply paying to reproduce the same failure.
Failures should be classified first.

1TOOL TIMEOUT2-> retry with backoff34INVALID ARGUMENTS5-> repair tool call67MISSING CONTEXT8-> retrieve missing source910FAILED TEST11-> inspect failing behavior1213PERMISSION DENIED14-> request approval1516CONFLICTING REQUIREMENTS17-> escalate1819UNCHANGED REPEATED FAILURE20-> stop
A useful agent loop looks like this:
1OBSERVE2 |3 v4DECIDE5 |6 v7ACT8 |9 v10MEASURE11 |12 +---- ACCEPT13 |14 +---- REPAIR15 |16 +---- ESCALATE17 |18 +---- STOP
Every loop should have limits on attempts, time, spend and destructive scope.
A reliable agent needs to know how to continue.
It also needs to know when another attempt is no longer worth it.
10. Turn Repeated Instructions Into Infrastructure
Suppose the prompt contains:
Always run the formatter.
That rule is stronger if the formatter runs automatically.
Suppose the instructions say:
UI code may not directly access the database.
That is stronger as an architecture test that fails when the rule is broken.
The progression looks like this:
1EXPLANATION2 |3 v4CHECKLIST5 |6 v7TEMPLATE8 |9 v10AUTOMATED CHECK11 |12 v13ENFORCED POLICY
The prompt should explain judgment.
The harness should enforce invariants.
Every recurring mistake should move a little further down this ladder.
Eventually the model no longer needs to remember the lesson.
The environment remembers it for the model.
11. Record the Run
A perfect final artifact can hide a terrible execution path.
Maybe the agent accessed the wrong source.
Maybe it ignored a failed command.
Maybe it repeated an external action twice.
Maybe it spent ten times the expected budget.
Maybe it got the right answer for the wrong reason.
Record enough information to reconstruct what happened.
109:14 task contract created209:15 architecture.md loaded309:17 checkout.ts edited409:18 focused test failed509:21 implementation repaired609:22 focused test passed709:24 integration test passed809:25 deployment blocked: approval required
Useful traces include context sources, tool calls, state changes, verification results, retry reasons, approval decisions, cost and latency.
The point is not to collect logs for fun.
The point is to make failure local.
If step 18 breaks, you should be able to repair step 18.
You should not need to replay the entire run.
12. Give Every Run a Receipt
Do not force the human to review a forty-message transcript.
Compile the result into a small receipt.
1OBJECTIVE23Fix duplicate coupon application.45CHANGED67checkout validation8regression test910VERIFIED1112lint passed13unit tests passed14integration test passed1516NOT VERIFIED1718production payment provider1920RISKS2122legacy mobile client unavailable2324APPROVAL NEEDED2526deploy to staging
This is not a summary of what the model claims happened.
It is a summary of what the harness can prove happened.
That distinction makes the receipt useful for review, handoffs and future agent sessions.
13. Make Every Failure Improve the Harness
Most teams fix the failed output.
The better approach is to fix the system that allowed the failure.
1MISSING CONTEXT2-> improve project map34WRONG TOOL5-> improve routing or tool contract67BAD OUTPUT8-> add validator910REPEATED LOOP11-> add retry cap1213UNSAFE ACTION14-> add permission gate1516LOST DECISION17-> persist state1819UNKNOWN FAILURE20-> improve tracing
This is where harness engineering starts to compound.
One repaired output helps one run.
One repaired harness improves every run after it.
The best agent systems get more reliable because mistakes leave infrastructure behind.
14. Start With the Smallest Useful Harness
You do not need a huge orchestration platform to start.
Build in layers.
1LEVEL 023prompt4model56LEVEL 178task contract9project map10tools1112LEVEL 21314structured state15verification16bounded loop1718LEVEL 31920permissions21traces22recovery23human gates
A short research task may only need a prompt and one review.
A six-hour coding task with file access, network access and deployment capability needs much more.
Add complexity when the failure surface earns it.
Not because agent architecture looks cool.
The Harness Engineering Checklist
Before giving an agent meaningful autonomy, ask:
1[ ] Is success defined before execution?23[ ] Can the agent find the right context4 without loading everything?56[ ] Does every tool have a clear purpose,7 schema and failure state?89[ ] Are important decisions stored10 outside the conversation?1112[ ] Does completion require evidence?1314[ ] Are risky actions protected by policy?1516[ ] Does every loop have a retry limit?1718[ ] Can the run resume after interruption?1920[ ] Can you reconstruct every important action?2122[ ] Does failure improve a rule, tool,23 test, map or permission?2425[ ] Can the final change be rolled back?
If several answers are no, a stronger model will not automatically make the agent reliable.
It may simply make the failure faster and more expensive.
The Real Shift
Prompt engineering asks:
What should I tell the model?
Context engineering asks:
What should the model know right now?
Harness engineering asks:
What system lets the model act, verify its work, recover from failure and operate safely?
1PROMPT2-> instruction34CONTEXT5-> working view67HARNESS8-> operating environment910LOOP11-> local correction1213GRAPH14-> coordination
Models will keep changing.
The durable advantage lives around them.
Your contracts get better.
Your tools get better.
Your tests get better.
Your state gets cleaner.
Your permissions get safer.
Your recovery logic gets smarter.
Your failures turn into infrastructure.
That is how capable models become reliable agents.
That is Harness Engineering.
If You Made It This Far
Bookmark this guide.
Follow me on X: x.com/0xjmori
Subscribe to my Substack: substack.com/@lunarresearcher
Send this article to someone who is still trying to fix every agent failure with a longer prompt.



![[Apology] I No Longer Recommend Freelancing for Independence.](/cdn-cgi/image/width=1920,quality=90,format=auto,metadata=none/https%3A%2F%2Fcms-assets.youmind.com%2Fmedia%2F1790615558140_ndvona_HTQ9s6laYAAX4J7.jpg)

