YouMind
تسجيل الدخول

Harness Engineering: How to Build AI Agents That Actually Work

@0xjmori
الإنجليزية27 سبتمبر 2026
328K
196
24
17
638

ليرة تركية؛ د

This article introduces 'Harness Engineering,' arguing that AI agent reliability depends on the surrounding system (contracts, tools, state, verification) rather than just the model or prompt. It provides a comprehensive guide to building robust agent infrastructures.

Most people are trying to fix AI agents at the wrong layer.

When an agent fails, they rewrite the prompt. When it fails again, they add more instructions, switch models, increase the context window, or connect another tool.

Then the same problems come back.

The agent forgets an important decision. It uses the wrong tool. It loses track of what happened three steps earlier. It claims the task is finished without checking the result. It retries the same failed action until the budget is gone.

The problem is not always the model.

The problem is the environment around it.

That environment is the harness.

Harness Engineering is the practice of building the system around a model that decides what it can see, what it can do, what it remembers, what counts as success, and what happens when something fails.

A better prompt can improve one response.

A better harness improves every run.

Follow my Substack for more practical breakdowns of AI agents, automation and production systems:

substack.com/@lunarresearcher

1. The Model Is Not the Agent

A model can reason, generate, compare and choose.

That does not make it a reliable agent.

A real agent also needs to find the right context, use tools, preserve state, respect permissions, verify its own work and recover when the environment behaves differently than expected.

The model is only the reasoning engine.

The harness is everything that turns that reasoning into real execution.

text
1USER REQUEST
2 |
3 v
4+-----------------------------+
5| HARNESS |
6| |
7| contract context |
8| tools state |
9| policy verification |
10| traces recovery |
11+-----------------------------+
12 |
13 v
14 MODEL
15 |
16 v
17REAL ENVIRONMENT

Put the same model inside a chat box and it answers questions.

Mori - inline image

Put it inside a repository with terminal access, tests, browser tools, project memory, controlled permissions and a review loop and it can complete real work.

The model did not change.

The harness did.

2. Turn Every Request Into a Contract

Natural language is flexible.

Autonomous execution should not be.

A request like:

Improve the onboarding flow.

is fine when a human is sitting next to the model.

It is terrible as a production instruction.

Before the agent acts, turn the request into a bounded task contract.

Mori - inline image
yaml
1objective: reduce onboarding drop-off
2
3inputs:
4 - product brief
5 - analytics data
6 - repository
7
8constraints:
9 - preserve authentication
10 - do not change database schema
11 - preserve current mobile behavior
12
13deliverable:
14 - reviewable pull request
15
16done_when:
17 - tests pass
18 - analytics event fires correctly
19 - desktop flow passes review
20 - mobile flow passes review
21
22approval_required:
23 - production deployment

The important part is done_when.

Without it, the agent can solve a slightly easier version of the problem and still confidently say the task is complete.

With it, completion becomes measurable.

The agent should not ask:

What should I do next?

It should ask:

What action moves the current environment closer to the contracted outcome?

That is a much stronger loop.

3. Give the Agent a Map, Not a Giant Context Window

A common reaction to agent mistakes is to give the model more context.

More documentation.

More conversation history.

More files.

More tool output.

Eventually the agent receives everything and understands less.

Context is not storage.

It is an attention budget.

Mori - inline image

Instead of dumping the entire project into every run, give the agent a small map of where useful information lives.

text
1PROJECT MAP
2
3product rules -> docs/product/
4architecture -> docs/architecture.md
5frontend -> apps/web/
6backend -> services/api/
7tests -> tests/
8commands -> docs/commands.md
9security -> docs/security.md

Then expand only when needed.

text
1TASK
2 |
3 v
4PROJECT MAP
5 |
6 v
7RELEVANT SYSTEM
8 |
9 v
10EXACT FILES
11 |
12 v
13LOCAL INSTRUCTIONS

The source material describes this as progressive disclosure: the harness should load more information because the task needs it, not simply because the information exists.

The goal is not maximum context.

The goal is maximum useful signal.

4. Put a Gateway Between the Model and Its Tools

A model with twenty tools is not automatically twenty times more capable.

It may just have twenty more ways to fail.

Every tool should have a contract.

text
1TOOL: edit_file
2
3INPUTS
4path
5patch
6
7PRECONDITIONS
8path exists
9path is inside workspace
10
11SUCCESS
12patch applied
13diff returned
14
15FAILURE
16structured error
17no partial overwrite
18
19RISK
20reversible

Then the execution path becomes:

text
1MODEL PROPOSES
2 |
3 v
4GATEWAY VALIDATES
5 |
6 v
7POLICY AUTHORIZES
8 |
9 v
10TOOL EXECUTES
11 |
12 v
13HARNESS RECORDS RESULT

The model decides what action it wants.

The harness decides whether that action is valid, permitted and safe.

Mori - inline image

That distinction becomes critical when tools can send messages, modify production, spend money or delete data.

A good tool gateway can also add timeouts, validate arguments, restrict file paths, normalize errors and make retries safe.

Good tools reduce the number of things the model has to guess.

5. Move Memory Outside the Conversation

The conversation should not be the system of record.

Long-running agents eventually hit context limits, crash, restart, or hand work to another session.

If every important decision exists only inside the transcript, the workflow is fragile.

Store durable state separately.

Mori - inline image
json
1{
2 "task_id": "feature_042",
3 "status": "verifying",
4 "current_step": "mobile_check",
5
6 "completed": [
7 "implementation",
8 "unit_tests",
9 "desktop_check"
10 ],
11
12 "decisions": [
13 "reuse existing export endpoint",
14 "preserve current date format"
15 ],
16
17 "artifacts": [
18 "export.csv",
19 "desktop-after.png"
20 ],
21
22 "open_risks": [
23 "mobile toolbar may overflow"
24 ],
25
26 "next_action": "render mobile viewport"
27}

A useful system separates memory into four categories:

text
1FACTS
2stable knowledge
3
4DECISIONS
5what was chosen and why
6
7STATE
8where the current run is
9
10LESSONS
11failures that should affect future runs

The next agent session should inherit the state of the work, not a compressed story about the previous conversation.

6. Make Evidence the Gate to Completion

An agent saying "done" is not proof that the job is done.

Mori - inline image

It is another model output.

The harness needs observable evidence.

text
1CLAIM EVIDENCE
2
3"bug is fixed" failing test now passes
4
5"page works" browser flow completed
6
7"data is correct" values match source
8
9"migration is safe" dry run + rollback pass
10
11"task is complete" every acceptance check passes

Use deterministic checks first.

text
1syntax
2 |
3 v
4types
5 |
6 v
7focused tests
8 |
9 v
10integration tests
11 |
12 v
13visual / semantic review
14 |
15 v
16human approval

Do not ask another model to answer something a compiler, test, schema or database query can prove.

Use models for judgment.

Use deterministic systems for facts.

The model creates the artifact.

The environment creates evidence about the artifact.

The harness decides whether the evidence is sufficient.

7. Separate the Builder From the Verifier

There is another problem with self-review.

The agent that created the mistake often carries the same assumptions into the review.

Mori - inline image

A stronger architecture separates the worker and the verifier.

text
1BUILDER
2 |
3 v
4creates candidate
5 |
6 v
7VERIFIER
8 |
9 +-- checks contract
10 +-- searches for missing cases
11 +-- tests unsupported claims
12 +-- tries to break result
13 |
14 +------ PASS ------> ACCEPT
15 |
16 +------ FAIL ------> RETURN EVIDENCE

The verifier should not ask:

Does this look good?

It should ask:

What would make this unacceptable?

That changes review from confirmation into attempted disproof.

The source material explicitly recommends giving verification its own rejection criteria and enough independence to challenge the assumptions that produced the first result.

8. Move Permissions Outside the Model

Some rules should never depend on the model remembering them.

text
1never publish without approval
2never expose secrets
3never exceed the spend limit
4never write outside the workspace
5never claim tests passed unless they ran

These are not prompt suggestions.

Mori - inline image

They are policy.

A simple permission ladder:

text
1LOW RISK
2
3read
4search
5inspect
6
7-> automatic
8
9REVERSIBLE
10
11edit workspace
12run tests
13create draft
14
15-> automatic + trace
16
17EXTERNAL EFFECT
18
19send
20deploy
21purchase
22
23-> approval required
24
25IRREVERSIBLE / SENSITIVE
26
27delete data
28rotate credentials
29publish globally
30
31-> hard gate or prohibited

The stronger the consequence, the stronger the control.

The model can recommend the action.

The harness authorizes it.

The tool executes it.

Autonomy is not the absence of control.

It is freedom inside an enforced boundary.

9. Stop Retrying Blindly

One of the worst recovery policies is:

Something failed. Try again.

If nothing changes, the system is simply paying to reproduce the same failure.

Failures should be classified first.

Mori - inline image
text
1TOOL TIMEOUT
2-> retry with backoff
3
4INVALID ARGUMENTS
5-> repair tool call
6
7MISSING CONTEXT
8-> retrieve missing source
9
10FAILED TEST
11-> inspect failing behavior
12
13PERMISSION DENIED
14-> request approval
15
16CONFLICTING REQUIREMENTS
17-> escalate
18
19UNCHANGED REPEATED FAILURE
20-> stop

A useful agent loop looks like this:

text
1OBSERVE
2 |
3 v
4DECIDE
5 |
6 v
7ACT
8 |
9 v
10MEASURE
11 |
12 +---- ACCEPT
13 |
14 +---- REPAIR
15 |
16 +---- ESCALATE
17 |
18 +---- STOP

Every loop should have limits on attempts, time, spend and destructive scope.

A reliable agent needs to know how to continue.

It also needs to know when another attempt is no longer worth it.

10. Turn Repeated Instructions Into Infrastructure

Suppose the prompt contains:

Always run the formatter.

That rule is stronger if the formatter runs automatically.

Suppose the instructions say:

UI code may not directly access the database.

That is stronger as an architecture test that fails when the rule is broken.

The progression looks like this:

text
1EXPLANATION
2 |
3 v
4CHECKLIST
5 |
6 v
7TEMPLATE
8 |
9 v
10AUTOMATED CHECK
11 |
12 v
13ENFORCED POLICY

The prompt should explain judgment.

The harness should enforce invariants.

Every recurring mistake should move a little further down this ladder.

Eventually the model no longer needs to remember the lesson.

The environment remembers it for the model.

11. Record the Run

A perfect final artifact can hide a terrible execution path.

Maybe the agent accessed the wrong source.

Maybe it ignored a failed command.

Maybe it repeated an external action twice.

Maybe it spent ten times the expected budget.

Maybe it got the right answer for the wrong reason.

Record enough information to reconstruct what happened.

text
109:14 task contract created
209:15 architecture.md loaded
309:17 checkout.ts edited
409:18 focused test failed
509:21 implementation repaired
609:22 focused test passed
709:24 integration test passed
809:25 deployment blocked: approval required

Useful traces include context sources, tool calls, state changes, verification results, retry reasons, approval decisions, cost and latency.

The point is not to collect logs for fun.

The point is to make failure local.

If step 18 breaks, you should be able to repair step 18.

You should not need to replay the entire run.

12. Give Every Run a Receipt

Do not force the human to review a forty-message transcript.

Compile the result into a small receipt.

text
1OBJECTIVE
2
3Fix duplicate coupon application.
4
5CHANGED
6
7checkout validation
8regression test
9
10VERIFIED
11
12lint passed
13unit tests passed
14integration test passed
15
16NOT VERIFIED
17
18production payment provider
19
20RISKS
21
22legacy mobile client unavailable
23
24APPROVAL NEEDED
25
26deploy to staging

This is not a summary of what the model claims happened.

It is a summary of what the harness can prove happened.

That distinction makes the receipt useful for review, handoffs and future agent sessions.

13. Make Every Failure Improve the Harness

Most teams fix the failed output.

The better approach is to fix the system that allowed the failure.

text
1MISSING CONTEXT
2-> improve project map
3
4WRONG TOOL
5-> improve routing or tool contract
6
7BAD OUTPUT
8-> add validator
9
10REPEATED LOOP
11-> add retry cap
12
13UNSAFE ACTION
14-> add permission gate
15
16LOST DECISION
17-> persist state
18
19UNKNOWN FAILURE
20-> improve tracing

This is where harness engineering starts to compound.

One repaired output helps one run.

One repaired harness improves every run after it.

The best agent systems get more reliable because mistakes leave infrastructure behind.

14. Start With the Smallest Useful Harness

You do not need a huge orchestration platform to start.

Build in layers.

text
1LEVEL 0
2
3prompt
4model
5
6LEVEL 1
7
8task contract
9project map
10tools
11
12LEVEL 2
13
14structured state
15verification
16bounded loop
17
18LEVEL 3
19
20permissions
21traces
22recovery
23human gates

A short research task may only need a prompt and one review.

A six-hour coding task with file access, network access and deployment capability needs much more.

Add complexity when the failure surface earns it.

Not because agent architecture looks cool.

The Harness Engineering Checklist

Before giving an agent meaningful autonomy, ask:

text
1[ ] Is success defined before execution?
2
3[ ] Can the agent find the right context
4 without loading everything?
5
6[ ] Does every tool have a clear purpose,
7 schema and failure state?
8
9[ ] Are important decisions stored
10 outside the conversation?
11
12[ ] Does completion require evidence?
13
14[ ] Are risky actions protected by policy?
15
16[ ] Does every loop have a retry limit?
17
18[ ] Can the run resume after interruption?
19
20[ ] Can you reconstruct every important action?
21
22[ ] Does failure improve a rule, tool,
23 test, map or permission?
24
25[ ] Can the final change be rolled back?

If several answers are no, a stronger model will not automatically make the agent reliable.

It may simply make the failure faster and more expensive.

The Real Shift

Prompt engineering asks:

What should I tell the model?

Context engineering asks:

What should the model know right now?

Harness engineering asks:

What system lets the model act, verify its work, recover from failure and operate safely?

text
1PROMPT
2-> instruction
3
4CONTEXT
5-> working view
6
7HARNESS
8-> operating environment
9
10LOOP
11-> local correction
12
13GRAPH
14-> coordination

Models will keep changing.

The durable advantage lives around them.

Your contracts get better.

Your tools get better.

Your tests get better.

Your state gets cleaner.

Your permissions get safer.

Your recovery logic gets smarter.

Your failures turn into infrastructure.

That is how capable models become reliable agents.

That is Harness Engineering.

If You Made It This Far

Bookmark this guide.

Follow me on X: x.com/0xjmori

Subscribe to my Substack: substack.com/@lunarresearcher

Send this article to someone who is still trying to fix every agent failure with a longer prompt.

بنقرة واحدة حفظ

استخدم YouMind للقراءة العميقة للمقالات سريعة الانتشار بتقنية الذكاء الاصطناعي

احفظ المصدر، واطرح أسئلة مركزة، ولخص الحجة، وحوّل المقالة واسعة الانتشار إلى ملاحظات قابلة لإعادة الاستخدام في مساحة عمل واحدة تعمل بالذكاء الاصطناعي.

اكتشف YouMind
للمبدعين

حول Markdown إلى مقالة 𝕏 نظيفة

عندما تنشر كتاباتك الطويلة، فإن الصور والجداول وكتل التعليمات البرمجية تجعل تنسيق 𝕏 مؤلمًا. YouMind يحول مسودة Markdown كاملة إلى مقالة نظيفة وجاهزة للنشر 𝕏.

حاول Markdown إلى 𝕏

المزيد من الأنماط لفك التشفير

المقالات الفيروسية الأخيرة

استكشاف المزيد من المقالات الفيروسية