YouMind
로그인

ORO Bench: Continuous Eval as an Incentive Mechanism for Agentic Commerce

@oroagents
영어2026년 9월 22일
138K
36
10
2
3

TL;DR

ORO introduces ORO Bench, a dynamic evaluation framework for agentic commerce that generates fresh environments daily to prevent benchmark saturation and gaming.

The ORO model: Eval as an Incentive Mechanism

Subnet owners that are testing agent capabilities should be thinking about their validation as EVAL as an incentive mechanism (EAAIM). If you build a robust, high quality eval, you lay the groundwork necessary to have a high quality incentive mechanism. After this, It's the system design of your subnet, and the scaffolding in which you choose to serve your IM that helps you get the most out of your subnet. We've got some more details about that part too, check out the video below. We did this with ShoppingBench. Today, we’re excited to launch and announce how we’re doing this with our new benchmark, ORO Bench.

Why we needed to move away from ShoppingBench

Previously, our EAAIM was ShoppingBench. As model capability increased and agent capability followed, we found that ShoppingBench got saturated. It's a static benchmark with a lot to offer in complexity for multi-step reasoning, however, static benchmarks will nevertheless still get saturated. This is the industry expectation. Unless you treat your benchmark as continuous software, as Ryan Marten, links in "Continuous Benchmarks", this is the expectation. It's not the first time, as ShoppingReasoningBench got saturated within 2 weeks of launch. Launched on June 2026, Gemini 3.5 Pro had a 77% pass rate on it within 2 weeks of launch. In fact, a systematic ICML 2026 study finds nearly half of benchmarks exhibit saturation, with rates increasing with age.

Our next Benchmark

As a team, our next step was to, well, build the benchmark that's going to become the defining consumer shopping agent benchmark. One that will take time to saturate (80%+ performance).

Before we can dive into this new benchmark, it’s important to establish why it is necessary for us:

  • Lack of Open Source work in agentic commerce. How do I know the best agent that's going to shop for me will do a good job? Smaller, less performative agents perform worse in the market. We talk about this in a previous article, where Opus 4.5 agents paid $2.45 less as buyers and extracted $2.68 more as sellers than Haiku 4.5 agents for the same items, and the people represented by the weaker model did not notice.
  • Eval hardening by hosting it on a Bittensor subnet. Bittensor is the perfect playground for academia to harden their evals, because if a miner will exploit whatever avenue possible. Recently, Epoch AI published the SWE-bench Verified review, Sep 2026 and found that over half of the commonly unsolved tasks had tests that rejected functionally correct submissions, and every frontier model had seen some of the problems in training. These "flaws" in benchmarks would have, with certainty, been uncovered by miners on Bittensor.
  • Every submission on our arena produces a trajectory. The full sequence of actions, observations, tool calls and outcomes from a run. This interaction data is how vertical companies have etched out a market in their niche (Cursor in coding, OpenEvidence in clinical medicine, Sierra in customer service, Harvey & Legora in legal). We realized these trajectories can be valuable post-training data (read our trajectory paper: distilling a shopping agent from subnet traces).

So as we're thinking about the next benchmark, we asked ourselves, how do we build something that doesn't immediately get saturated? How do we build something where we're not continually trying to play catch up with? Well, the answer is, we don't build a benchmark. We build an environment generator.

Introducing, ORO Bench

If we build a generalization engine for agentic commerce, we can continuously generate new, verifiable commerce environments and tasks on a daily basis, rewarding agents for how efficiently they reason through environments they have never seen, instead of a benchmark they can memorize.

This is a machine that produces a constant stream of fresh environments for agentic reasoning. We ground them in the methods linked here to ensure verifiability and scalability to feed both miner incentives and a future product.

What does an environment consist of?

Every environment starts from a product catalog: a snapshot of a real e-commerce index with 11 million source SKUs. The generator that creates this environment makes no model calls. It slices the catalog by category, attribute and price band, and each of seven task families draws its tasks from those slices, so a task is derived from what the catalog can support rather than fitted to a template. One task gives the agent:

  • A shopper goal in plain language with a budget and the constraints the shopper states up front. Since the shopper is a user-sim, the agent has the opportunity to ask further clarifying questions in in order to reveal soft preferences, which is rewarded in the validation.
  • A starting set of ten tools: search, filter, view, compare, stock and cart inspection, cart edits, a message channel to a simulated shopper, and an order call.
  • For recovery tasks, a sealed event: the chosen item goes out of stock, or is repriced above budget, at a fixed point in the episode.
  • A family-specific verifier that inspects the final state and pays a reward in two stages: hard gates first (the ordered product is in the accepted set, in stock, within budget, no illegal side effects), then a family metric (binary for intent, constraint and justification; continuous for retrieval, preference, ranking and recovery). A failed gate pays zero.

Tasks are compiled into a release (an EnvPack): a public qualifying roster that miners can iterate against locally, and a hidden race roster used for the daily race. The specifics are less important, since how tasks are generated will keep changing as miners find the edges, which is the point. What stays fixed is the contract in the ORO Bench docs.

The eval is inherently resistant to gaming, as it continuously updates. If we follow from our approach of “Eval as an Incentive Mechanism”, this also helps us lay the groundwork to make the incentive inherently resistant to gaming.

Working with Jarrod Barnes from Dynamical Systems

We’ve been working with Jarrod Barnes from Dynamical Systems to develop this new incentive mechanism. He noted that his inspiration here is from ARC-AGI-3, applied to commerce. Measuring intelligence by dropping agents into novel environments they cannot have memorized. The idea with ORO Bench is to do the same. The shopper's goal is given, so scoring stays verifiable, and only the environment is new each cycle. The agent has to explore and reason over real preferences, constraints, and tradeoffs instead of replaying a hardcoded harness.

We're excited to introduce our not-a-benchmark, ORO Bench today. We've been running this on our subnet with promising results (a released race episode from a top agent, recovering from a sold-out item).

What this means for our subnet

The only way to keep up with intelligent miners is to create a new incentive mechanism every day.

We're excited to see where miners take this. If you're a builder and you’re not submitting agents on ORO, well, you're missing out. Go check out oroagents.com.

By @shardiban, @ironseth_s, @JarrodBarnes

원클릭 저장

YouMind로 바이럴 글을 AI 심층 읽기

소스를 저장하고, 핵심 질문을 던지고, 주장을 요약해 바이럴 글을 다시 활용할 수 있는 노트로 바꾸세요. 하나의 AI 워크스페이스에서 모두 할 수 있습니다.

YouMind 둘러보기
크리에이터를 위해

당신의 Markdown을 깔끔한 𝕏 글로

직접 쓴 장문을 올릴 때 이미지, 표, 코드 블록을 𝕏에 맞게 정리하는 일은 번거롭습니다. YouMind는 전체 Markdown 초안을 깔끔하고 바로 게시할 수 있는 𝕏 글로 바꿔 줍니다.

Markdown → 𝕏 사용해 보기

분석할 패턴 더 보기

최근 바이럴 아티클

더 많은 바이럴 아티클 보기