YouMind
Sign in

ORO Bench: Continuous Eval as an Incentive Mechanism for Agentic Commerce

@oroagents
ENGLISHSep 22, 2026
138K
36
10
2
3

TL;DR

ORO introduces ORO Bench, a dynamic evaluation framework for agentic commerce that generates fresh environments daily to prevent benchmark saturation and gaming.

The ORO model: Eval as an Incentive Mechanism

Subnet owners that are testing agent capabilities should be thinking about their validation as EVAL as an incentive mechanism (EAAIM). If you build a robust, high quality eval, you lay the groundwork necessary to have a high quality incentive mechanism. After this, It's the system design of your subnet, and the scaffolding in which you choose to serve your IM that helps you get the most out of your subnet. We've got some more details about that part too, check out the video below. We did this with ShoppingBench. Today, we’re excited to launch and announce how we’re doing this with our new benchmark, ORO Bench.

Why we needed to move away from ShoppingBench

Previously, our EAAIM was ShoppingBench. As model capability increased and agent capability followed, we found that ShoppingBench got saturated. It's a static benchmark with a lot to offer in complexity for multi-step reasoning, however, static benchmarks will nevertheless still get saturated. This is the industry expectation. Unless you treat your benchmark as continuous software, as Ryan Marten, links in "Continuous Benchmarks", this is the expectation. It's not the first time, as ShoppingReasoningBench got saturated within 2 weeks of launch. Launched on June 2026, Gemini 3.5 Pro had a 77% pass rate on it within 2 weeks of launch. In fact, a systematic ICML 2026 study finds nearly half of benchmarks exhibit saturation, with rates increasing with age.

Our next Benchmark

As a team, our next step was to, well, build the benchmark that's going to become the defining consumer shopping agent benchmark. One that will take time to saturate (80%+ performance).

Before we can dive into this new benchmark, it’s important to establish why it is necessary for us:

  • Lack of Open Source work in agentic commerce. How do I know the best agent that's going to shop for me will do a good job? Smaller, less performative agents perform worse in the market. We talk about this in a previous article, where Opus 4.5 agents paid $2.45 less as buyers and extracted $2.68 more as sellers than Haiku 4.5 agents for the same items, and the people represented by the weaker model did not notice.
  • Eval hardening by hosting it on a Bittensor subnet. Bittensor is the perfect playground for academia to harden their evals, because if a miner will exploit whatever avenue possible. Recently, Epoch AI published the SWE-bench Verified review, Sep 2026 and found that over half of the commonly unsolved tasks had tests that rejected functionally correct submissions, and every frontier model had seen some of the problems in training. These "flaws" in benchmarks would have, with certainty, been uncovered by miners on Bittensor.
  • Every submission on our arena produces a trajectory. The full sequence of actions, observations, tool calls and outcomes from a run. This interaction data is how vertical companies have etched out a market in their niche (Cursor in coding, OpenEvidence in clinical medicine, Sierra in customer service, Harvey & Legora in legal). We realized these trajectories can be valuable post-training data (read our trajectory paper: distilling a shopping agent from subnet traces).

So as we're thinking about the next benchmark, we asked ourselves, how do we build something that doesn't immediately get saturated? How do we build something where we're not continually trying to play catch up with? Well, the answer is, we don't build a benchmark. We build an environment generator.

Introducing, ORO Bench

If we build a generalization engine for agentic commerce, we can continuously generate new, verifiable commerce environments and tasks on a daily basis, rewarding agents for how efficiently they reason through environments they have never seen, instead of a benchmark they can memorize.

This is a machine that produces a constant stream of fresh environments for agentic reasoning. We ground them in the methods linked here to ensure verifiability and scalability to feed both miner incentives and a future product.

What does an environment consist of?

Every environment starts from a product catalog: a snapshot of a real e-commerce index with 11 million source SKUs. The generator that creates this environment makes no model calls. It slices the catalog by category, attribute and price band, and each of seven task families draws its tasks from those slices, so a task is derived from what the catalog can support rather than fitted to a template. One task gives the agent:

  • A shopper goal in plain language with a budget and the constraints the shopper states up front. Since the shopper is a user-sim, the agent has the opportunity to ask further clarifying questions in in order to reveal soft preferences, which is rewarded in the validation.
  • A starting set of ten tools: search, filter, view, compare, stock and cart inspection, cart edits, a message channel to a simulated shopper, and an order call.
  • For recovery tasks, a sealed event: the chosen item goes out of stock, or is repriced above budget, at a fixed point in the episode.
  • A family-specific verifier that inspects the final state and pays a reward in two stages: hard gates first (the ordered product is in the accepted set, in stock, within budget, no illegal side effects), then a family metric (binary for intent, constraint and justification; continuous for retrieval, preference, ranking and recovery). A failed gate pays zero.

Tasks are compiled into a release (an EnvPack): a public qualifying roster that miners can iterate against locally, and a hidden race roster used for the daily race. The specifics are less important, since how tasks are generated will keep changing as miners find the edges, which is the point. What stays fixed is the contract in the ORO Bench docs.

The eval is inherently resistant to gaming, as it continuously updates. If we follow from our approach of “Eval as an Incentive Mechanism”, this also helps us lay the groundwork to make the incentive inherently resistant to gaming.

Working with Jarrod Barnes from Dynamical Systems

We’ve been working with Jarrod Barnes from Dynamical Systems to develop this new incentive mechanism. He noted that his inspiration here is from ARC-AGI-3, applied to commerce. Measuring intelligence by dropping agents into novel environments they cannot have memorized. The idea with ORO Bench is to do the same. The shopper's goal is given, so scoring stays verifiable, and only the environment is new each cycle. The agent has to explore and reason over real preferences, constraints, and tradeoffs instead of replaying a hardcoded harness.

We're excited to introduce our not-a-benchmark, ORO Bench today. We've been running this on our subnet with promising results (a released race episode from a top agent, recovering from a sold-out item).

What this means for our subnet

The only way to keep up with intelligent miners is to create a new incentive mechanism every day.

We're excited to see where miners take this. If you're a builder and you’re not submitting agents on ORO, well, you're missing out. Go check out oroagents.com.

By @shardiban, @ironseth_s, @JarrodBarnes

One-click save

Use YouMind for AI deep reading of viral articles

Save the source, ask focused questions, summarize the argument, and turn a viral article into reusable notes in one AI workspace.

Explore YouMind
For creators

Turn your Markdown into a clean 𝕏 article

When you publish your own long-form writing, images, tables, and code blocks make 𝕏 formatting painful. YouMind turns a full Markdown draft into a clean, ready-to-post 𝕏 article.

Try Markdown to 𝕏

More patterns to decode

Recent viral articles

Explore more viral articles