YouMind
تسجيل الدخول

The 5 Levels of AI Agents: From Prompt to Production

@undefinedKi
الإنجليزية28 سبتمبر 2026
357K
246
24
12
637

ليرة تركية؛ د

This article breaks down AI agent development into five essential layers: context, loop, Jev, harness, and evals. It explains how to structure these components for production reliability and offers practical implementation tips and shortcuts using tools like Viktor.

There are five words people throw around about agents right now. Context engineering, loop engineering, Jev engineering, harness engineering, eval engineering.

They sound like five competing approaches. They're not. They're five layers of one system, and each one answers one question.

The easiest way to see them is to imagine you just hired a new employee:

  • Context - what's on his desk when you ask him something
  • Loop - whether he follows your checklist or figures out the next step himself
  • Jev - the front desk that sorts the mail, so he only sees what matters
  • Harness - his office. The tools, the keys, and who checks his work
  • Evals - the same test every month, so you know he actually got better

An AI agent is that employee. The model is the person, and these five layers are everything around him. Get these five layers right and a small team can take on work that used to mean hiring. The repeat work goes to an agent that holds up in production, and your people keep the decisions. That's what scaling without adding headcount looks like in practice.

For each layer I'll break down what it is, how it works, where to use it and how to build it.

And for each layer I've also prepared a shortcut. A simpler way to get the same result without building it from scratch, made for beginners. My friends at Viktor helped me put these shortcuts together.

Viktor is an AI employee who lives in your Slack or Microsoft Teams. He sits in the same channels as everyone else and works like a teammate, one you add to the team without opening a new role.

Working with him looks like working with a person:

  • You mention him in a channel or a thread and describe the task.
  • He figures out the steps, does the work across your tools, and posts the result back in the same thread.

Setup is simple. Just add Viktor to your workspace, and he'll be listed as a participant just like any other member of your team. If you want to try him while reading, use code YARCHI100.

Yarchi - inline image

pic1. The five layers of an agent

1. Context engineering

Definition

Before each question, you put papers on the employee's desk. Put the right three pages there and he answers in seconds. Put three hundred and the answer is buried in the middle of the pile. He misses it.

The model is the employee. The desk is the context window: everything the model sees before it answers. Your instructions, the chat so far, documents pulled from a database, the output of every tool it ran.

Context engineering is deciding what goes on the desk, and where.

How it works

More context doesn't mean better answers. Past a certain point it means worse ones.

Stanford tested this directly. Give a model 20 to 30 documents and its accuracy on the ones in the middle drops to around 50 to 57 percent. With no documents at all, the same model scored 56 percent. The answer was in the window and the model did worse than with nothing.

Two things drive this:

  • Models pay the most attention to the start and the end of the window, the least to the middle.
  • Every token costs money and time, whether it helps or not.

The one mechanic worth knowing is prompt caching. The provider stores the start of your prompt and reuses it on the next call, and a cached token costs about ten times less than a fresh one.

The catch: the cache only works if that start is identical, byte for byte. Change one character near the top and everything after it is billed at full price. So the order is always stable stuff first, changing stuff last.

How to build it

Every piece of context goes into one of four places:

  1. System prompt. Only what's true on every call: role, constraints, output format. Keep it byte for byte identical. No timestamps at the top, no JSON keys in random order.
  2. Tools. Don't add or remove them mid conversation. It breaks the cache and leaves the model calling tools that no longer exist. To restrict a tool at some step, block the call and keep the definition.
  3. Disk. Anything large or long lived goes into a file, and only the path stays in the window. Same with a web page: keep the URL, drop the body. Drop the content, keep the key that gets it back.
  4. Tail. Every few steps, restate the current goal near the end of the context. You can't fix the middle, so keep what matters out of it.

Tools have a limit too. Anthropic measured 58 tool definitions eating about 55,000 tokens before the user typed a word. Letting the model search for tools instead of loading all of them took Opus 4 from 49 to 74 percent on their benchmark.

Under about 20 tools, keep them loaded. Over that, switch to search.

One more trick for big jobs: send a sub-agent. It reads the 50 files in its own window and hands back a one page summary. Your main context only ever sees the summary.

Yarchi - inline image

pic2. What goes into one model call

Shortcut

Viktor takes most of this off your hands, because his context lives at the company level.

  • Memory. He keeps a persistent memory of your business across the whole team. What he learned from your cofounder last week doesn't need to be pasted into your request today.
  • Connected sources. With Notion, Google Drive or HubSpot connected, you stop pasting files into chat. You name the doc or the record and he reads the source. That's the disk rule, done for you.
  • Skills. You record your screen doing a task once. He turns the recording into a written procedure, you correct it and approve it. From then on that instruction is fixed and reviewed, like a good system prompt.

What's left on your side:

  • Write a short company brief once: what you sell, who buys, which numbers matter, what he should never do.
  • One task per thread, with what done means in the first message.
  • If he learned something wrong, wipe his memory from settings instead of correcting him in every thread.

2. Loop engineering

Definition

You can give the employee a checklist: open the file, change line 12, save. Or you can give him a goal: make the test pass.

With a goal, he tries something, looks at what happened, and decides what to do next. The checklist is a workflow. The goal is a loop.

How it works

A loop is four moves on repeat: think, act, observe, decide. Fixing a bug looks like this:

  1. Run the tests. Three fail.
  2. Read the first error. It's a missing import.
  3. Add the import, run again.
  4. One test still fails. Read that error, fix it, run again.
  5. Everything passes. Stop.

Nobody wrote those steps in advance. The model chose each one after seeing the last result.

That's the whole difference. A hundred steps you wrote is still a workflow. Three steps the model chose is a loop.

Use a loop only when you can't write the steps in advance. If you can write them, write them. A workflow is cheaper, runs in parallel, and when step four fails you rerun step four, not everything.

Loops are also expensive. An agent uses roughly four times the tokens of a single call. Multi agent setups run around fifteen times.

How to build it

A loop needs four parts. Take away any one and it stops working:

  1. A goal with a clear done. Not "fix the bug." Instead: "the failing test in auth_test.py passes and nothing else broke."
  2. A checker. Something outside the model that says pass or fail: a test suite, a compiler, a linter. Research on self correction is consistent here. It works with real external feedback and fails when the model just reviews itself. No checker means no loop, just open ended spend.
  3. A stop rule. The checker passes, or you hit the turn limit, or the last two attempts gave the same output.
  4. A budget. Turns and dollars, both.

In code, the whole thing fits in a few lines:

python
1for turn in range(MAX_TURNS):
2 action = model.next_step(goal, history)
3 result = run(action)
4 history.append(result)
5
6 if checker(result): break # done
7 if repeated(history, 2): break # stuck
8 if spent() > BUDGET: break # too expensive
9else:
10 fallback_workflow(goal)

The best production setup I've seen published is a hybrid. Atlan runs a deterministic filter first, and only about 14 percent of incoming alerts even reach the agent.

The loop then gets three cycles at most. If confidence is still under 50 percent after three, a fixed Python workflow takes over.

Filter first, loop briefly, fall back.

Yarchi - inline image

pic3. How to decide between a workflow and a loop

Shortcut

Every task you give Viktor in a thread is a loop. You write the goal, he picks the steps, works through your tools and comes back to the thread with the result.

The four parts map like this:

  • Goal. He pushes back on incomplete briefs and asks instead of guessing. Still, write the done. "Revenue report for last week, totals match Stripe, posted in #finance by 9am Monday" beats "make the report."
  • Checker. He flags numbers that look off before posting. Make it stronger by naming the source to check against: the Stripe total, the row count, the test suite.
  • Stop rule. Sensitive actions pause for a human, and the result always comes back to you.
  • Budget. Credits. The reasoning tier sets the price of every step, and on recurring jobs frequency matters. An hourly report costs far more than a weekly one.

The Atlan hybrid works without code too. Work that repeats becomes a scheduled task: he proposes it, and it stays paused until you approve it. That's your workflow.

Anything open ended goes into a thread, and that's your loop. You're the fallback.

3. Jev engineering

Definition

The expert in an office doesn't open every envelope. Someone at the front desk sorts the mail: bills in one pile, spam in the bin, contracts to the lawyer.

Your agent does two kinds of work: writing something, and deciding something. Right now one big model does both, so you're paying the lawyer to sort the mail.

Jev is the front desk. It never writes. It only picks.

How it works

You give Jev a question and the possible answers in advance. It returns one of three things, plus a confidence score:

  • yes or no
  • one option from a set
  • a number on a scale

Because it only picks, it's fast and cheap. The claimed numbers are 70 to 500 milliseconds against 3 to 329 seconds, and $0.042 per million input tokens with output free.

Now the honest part. Jev is two weeks old and independent tests are only starting to come in.

  • On email classification, plain logistic regression scored 98.9 percent against Jev's 98.6.
  • On phishing detection, Jev got 62.6 percent while Claude Haiku 4.5 got 81.3.
  • The "zero hallucinations" figure comes with the authors' own footnote: it's not empirical, it only means the output always matches the schema.

The confidence score isn't a real probability out of the box either. Treat it as a ranking and set thresholds on your own labeled data.

How to build it

The sensible setup is a gate in front of the expensive model:

  1. Everything comes in.
  2. Jev answers one narrow question about each item.
  3. Confident and routine: handled cheaply. Labeled, routed or dropped.
  4. Unsure or unusual: goes to the main model.
python
1d = jev.choose(item, options=["spam", "order_status", "refund", "other"])
2
3if d.confidence >= 0.7 and d.option != "other":
4 handle_cheap(d.option, item)
5else:
6 main_model(item) # fail closed: unsure goes to the expensive path

Before any of that, try the dumbest thing that could work. Label a few hundred real examples, train a basic classifier, and move up only if it's not good enough.

That's thirty minutes of work, and it's the baseline any vendor claim has to beat.

Yarchi - inline image

pic4. The three kinds of questions Jev answers

Shortcut

Jev isn't part of Viktor's published stack, but the idea applies in two places you control:

  • The reasoning tier. He runs on three: Smart on Claude Opus, Balanced on Claude Sonnet at about half the cost, Ultra on Claude Fable at about twice. Sorting requests or pulling one number doesn't need the top tier.
  • The gate in front of him. If you want him to handle a stream like support emails or alerts, don't hand him the whole stream. Put a classifier or a Jev call first, so only the items that need judgment become tasks for him.

This is one of the two layers where your own engineering still matters, even with Viktor.

4. Harness engineering

Definition

Same employee, two offices. In the first he has the right tools, a manual on the wall, a colleague who checks his work, and no key to the safe. In the second he has a laptop and your admin password.

Same skills, very different results. The employee is the model. The office is the harness.

Agent = model + harness.

How it works

The harness is everything that isn't the model: tools, permissions, the sandbox, files that explain the project, checks on the output.

It became its own discipline in 2026 because people started measuring, and the model explained less than expected:

  • Anthropic changed only the container resources and moved a benchmark score by 6 points.
  • LangChain froze the model and moved the same benchmark 13.7 points by changing only the harness.
  • Then they tuned the harness around an open model ten times cheaper until it scored 0.86 against Opus 4.8's 0.87.

You're not buying a model anymore. You're buying a model and a harness together.

You need one the moment the agent touches something real: a repo, an inbox, a payment, a production database.

How to build it

Build from the outside in:

  1. Containment. What the agent physically can't reach. A container, a separate branch, a read only database user, no network except an allowlist. Do this before the first prompt.
  2. Guides. What steers it before it acts. A file in the repo like AGENTS.md, tool descriptions clear enough that the model picks the right one, a few examples of good output.
  3. Sensors. What checks it after it acts. Linter, type checker, test suite: fast and deterministic, so run them on everything. Slower checks, like a second model reviewing a diff, only on what matters.
  4. Permissions. When agents ask for approval, people approve 93 percent of the time. The approval prompt protects almost nothing. Real protection is the actions that simply aren't available. Save approval for the few things that are truly irreversible.

A guide file doesn't need to be long. Four lines already change behavior:

markdown
1# AGENTS.md
2- Monorepo: /api (FastAPI), /web (Next.js), /jobs (cron)
3- Run tests with `make test`. They must pass before any commit
4- Never edit /migrations by hand, use `make migration`
5- DB access is read only. Ask before any schema change

Hooks are the enforced version of the same idea. In Claude Code, a hook is a small script that runs before a tool call and can block it. "Never push to main" works as a hook, not as a line in the prompt.

One warning. Every piece of your harness is a bet that the model can't do something, and those bets expire. Anthropic deleted a whole scaffolding component after a model upgrade made it unnecessary.

Reread your harness every few months and delete what the model has outgrown.

Yarchi - inline image

pic5. The four rings of a harness

Shortcut

If agent = model + harness, most of what you get with Viktor is harness. The model underneath is Claude. Everything around the model comes ready:

  • Tools. 3,200+ integrations: GitHub, Linear, HubSpot, Stripe, Notion, Google Drive and more. For a tool with no ready connection, he can build one himself.
  • Guides. Skills. Instead of writing AGENTS.md yourself, you record your screen, he drafts the procedure, you edit it.
  • Sensors. He flags data that doesn't add up and questions briefs with missing pieces. And every step lands in a Slack thread your team can read.
  • Permissions. Customer emails and financial changes pause for approval. New scheduled automations stay paused until someone turns them on.

The part no product can build for you is containment, because it's made of the credentials you hand over. Remember the 93 percent, and connect him with the least access the job needs:

  • a read only database user
  • a restricted Stripe key
  • a shared support inbox, not your personal one
  • a GitHub token scoped to specific repos

What he can't reach, he can't break.

Yarchi - inline image

pic6. Viktor's harness

5. Evals engineering

Definition

How do you know a new hire got better? You give him the same test you gave him last month and compare.

Without the same test, every change you make is a guess. Evals are that test for your agent: a set of tasks where you already know the right answer, run every time you change something.

How it works

There are two kinds of checks:

  • End to end. Did the final answer come out right? It tells you the score moved, not why.
  • Behavioral. Did one specific thing happen? Did it call search before answering. Did it ask a clarifying question when the request was vague. Did it verify before saying done.

Behavioral checks run on the trace, the log of everything the agent did, not only on the final answer. Google's rule is that this suite should finish in under five seconds, so it can run on every change.

If a model does the grading, that's a judge, and the judge needs checking too. Airbnb found that about three quarters of their model generated reference answers came out different on repeated runs of the same input. Their eval was measuring its own noise.

How to build it

  1. Start from real failures. Go through actual runs, find what went wrong, turn each into a test case. Airbnb's golden sets run 50 to 100 examples, and failures are required.
  2. Write one behavioral check per case. One case, one thing to verify.
  3. Check your judge. Grade a sample by hand and see how often you and the model agree before you trust it.
  4. Sample production. Airbnb pulls 5 percent of live traffic every day. Your eval set decays as real usage drifts away from it.

A single case can be this small:

yaml
1input: "Refund order #1042, customer says it arrived broken"
2expect:
3 - looks up the order before replying
4 - asks for approval before issuing the refund
5check:
6 - trace has get_order before send_reply
7 - trace has approval_request before refund

Don't watch the overall pass rate. It can go up while one specific behavior quietly breaks. Watch the individual checks.

Yarchi - inline image

pic7. Where good eval cases come from

Shortcut

No product can ship this layer for you, because only you know what the right answer looks like for your business. The method carries over directly though:

  • After two weeks, pick 20 of his threads where you know the correct answer. Include every one where you had to correct him.
  • Write one plain check per case. Did he link a source for every number, did he ask when the brief was vague, did he pause before anything left the company.
  • Rerun the set after every change: a new skill, an edited brief, a different tier. Moving from Smart to Balanced saves about half the credits, so test it before you switch for good.
  • Once a week, grade five random threads by hand.

Putting it together

Five layers, five questions. Context is what it sees. Loop is who decides. Jev handles the cheap decisions. Harness is what it can reach and who checks it. Evals are how you know.

With Viktor, three of them come mostly built in: context, loop and harness. Jev, evals and the credentials you hand over stay your engineering.

If you're starting today, the cheapest useful move on each:

  • Context: move your stable instructions into the system prompt and stop changing them.
  • Loop: put a turn limit and a dollar limit on it before you leave it running.
  • Jev: label 200 examples of whatever your agent decides most often.
  • Harness: run your agent as a user that can't delete anything.
  • Evals: write down your last five failures as test cases.

On Viktor, the same day looks like this:

  • Context: write the company brief and record your first skill.
  • Loop: put a definition of done in every task and pick the tier on purpose.
  • Jev: filter any stream before it turns into tasks for him.
  • Harness: connect him with read only, scoped credentials.
  • Evals: save 20 real threads as your first test set.

That's a day of work and it covers all five.

My take: the model is no longer the hard part. The five layers around it are. Get them right and you add capacity without adding headcount. Most teams should take the first three off the shelf and spend their own time on the two that decide quality, the cheap decisions and the evals.

Thanks Viktor for sponsoring this article.

Try free at @viktor_com. $100 in credits, no card. Full link in my first reply.

Use code YARCHI100 when you sign up.

Paid Partnership

بنقرة واحدة حفظ

استخدم YouMind للقراءة العميقة للمقالات سريعة الانتشار بتقنية الذكاء الاصطناعي

احفظ المصدر، واطرح أسئلة مركزة، ولخص الحجة، وحوّل المقالة واسعة الانتشار إلى ملاحظات قابلة لإعادة الاستخدام في مساحة عمل واحدة تعمل بالذكاء الاصطناعي.

اكتشف YouMind
للمبدعين

حول Markdown إلى مقالة 𝕏 نظيفة

عندما تنشر كتاباتك الطويلة، فإن الصور والجداول وكتل التعليمات البرمجية تجعل تنسيق 𝕏 مؤلمًا. YouMind يحول مسودة Markdown كاملة إلى مقالة نظيفة وجاهزة للنشر 𝕏.

حاول Markdown إلى 𝕏

المزيد من الأنماط لفك التشفير

المقالات الفيروسية الأخيرة

استكشاف المزيد من المقالات الفيروسية