YouMind
تسجيل الدخول

Benchmarking agentic reasoning: Gemini 3 Pro vs. Gemini 2.5 Pro in Pokémon Crystal

@GoogleAIStudio
الإنجليزية15 ديسمبر 2025
455K
1.0K
144
38
360

ليرة تركية؛ د

A head-to-head benchmark reveals Gemini 3 Pro is up to 8x faster than Gemini 2.5 Pro at completing Pokémon Crystal, showcasing superior tool creation and visual reasoning.

Gemini 3 Pro was the winner of the race, obtaining 16 badges, defeating the Elite Four and Champion, and defeating the hidden boss Red in roughly half the tokens and turns it took Gemini 2.5 Pro to acquire only four badges. This fully autonomous head-to-head race in Pokémon Crystal was conducted by Joel Zhang (@TheCodeOfJoel) of the ARISE Foundation (and streamed on Twitch). His detailed blog post comparing the models revealed multiple fascinating distinctions in their behavior; overall, Gemini 3 Pro is at least 2x faster than Gemini 2.5 Pro at completing Crystal, and if we extrapolate, a more accurate estimate suggests the older model is around 8x slower.

Google AI Studio - inline image

The rate of completion of Gemini 3 Pro compared to Gemini 2.5 Pro. Credit: Joel Zhang

This culminated in the final battle against Red. Facing a level disadvantage, the 3.0 agent devised a complex, multi-stage strategy it termed "Operation Zombie Phoenix," combining passive recovery, stat reduction, resource exhaustion, and a "revive loop" to secure victory in a marathon 7-hour battle.

Google AI Studio - inline image

0:52

The victory over Red. Credit: Joel Zhang

An AI scientist prompt

The harness setup for this race was identical across the two agents to ensure a fair comparison. Notably, the agents were not prompted to “complete the game as fast as possible,” but rather to employ the scientific method and not assume that their prior knowledge about the game was correct. The unstructured notepad function allowed the agents to record hypotheses and test ideas, while keeping track of their gameplay.

This philosophy is in line with the harness’ flexibility that allowed the agents to design their own code tools and sub-agents within the harness. In some sense, this race also tested how fast the agents could adapt to their environment and build a working setup to succeed in the world of Pokémon Crystal.

Discarding the "training wheels"

Gemini 3 Pro exhibits a higher likelihood of trusting its tools. When an action fails, it re-evaluates the environment rather than the codebase. This awareness led to a fascinating behavior regarding harness restrictions.

The harness enforces strict input handling, prohibiting "mixed button inputs" (e.g., pressing A and Up in sequence) to keep 2.5 Pro stable and prevent emulator desyncs. When Gemini 3 Pro encountered a situation requiring complex input sequences—specifically nicknaming a Pokémon—it found the single-press restriction inefficient.

Rather than accepting the constraint, it utilized the define_tool capability to write a custom tool called press_sequence, as custom tools lack the mixed input restriction for button presses.

This script allowed it to batch input sequences locally, effectively writing its own driver to bypass the harness restrictions to improve its efficiency via the clever intended loophole. The 3.0 agent treated the harness constraints as engineering problems to be solved, not immutable laws.

Multi-modal advantage

In the 8th Gym, the solution requires dropping boulders from a floor above to chart a path across a floor made of lava. The state change for the bottom floor is difficult to track based only on the RAM data from the harness since there are no mentions of fallen boulders in the data.

Gemini 3 Pro utilized the visual feed to identify the fallen boulders to unstick itself from a loop it had fallen into, assuming the puzzle was not yet solved (a fact exacerbated by the decoy boulders remaining on the second level). It ignored the potentially confusing state data and relied on the screenshot to identify the boulder positions, correcting its strategy based on visual evidence. This ability to switch data modalities—from RAM inspection to raw vision—helped the 3.0 agent escape a "stuck" state that had left it looping for hours.

Also remarkable was the 3.0 agent’s ability to “read” the health bar of opponents. This information, incredibly significant to understanding the optimal move to make in a battle, is not provided by the RAM state, and must be deduced by the agent from the screen. The 3.0 agent was able to quite accurately estimate the fraction of health remaining during the battle with Red, a fact which likely contributed to its success.

Battle efficiency and state management

The efficiency gap and improved battle reasoning performance was extremely significant in Gemini 3 Pro’s victory. Gemini 2.5 Pro lost twice to the 3rd Gym leader (Whitney) due to poorer strategizing capabilities and as a result spent an excessive amount of time grinding levels far beyond what was necessary to obtain the 3rd badge.

Gemini 3 Pro completed the entire game, including the final hidden boss battle with Red, without a single loss.

It demonstrated superior tactical reasoning, performing live damage calculations to optimize move selection. For example, it correctly chose Swift over Flamethrower after recognizing that the opponent's Snorlax had boosted its Special Defense, and also factored in calculations based on weather (rain reduces fire damage). During the Elite Four gauntlet, it managed hit point conservation proactively, using items to top off health between rounds—behavior that 2.5 Pro historically struggles to prioritize over immediate combat moves.

Current limitations

Despite the performance leap, Gemini 3 Pro is not without flaws.

  • Assumptions without verification: The biggest failure mode observed was forming a hypothesis and refusing to test it. In one instance, the 3.0 agent assumed the radio interface worked like a standard menu (Left/Right) rather than a visual dial (Up/Down), ignoring visual cues and wasting hours in a loop. In another case, the 3.0 agent spent a long time testing increasingly complicated theories about a locked door puzzle, failing to talk to the hint-giving NPCs nearby.
  • Proactive planning: While reactive tactics are strong, proactive goal management remains inconsistent. The 3.0 agent often identifies a strategic need (e.g., "switch Pokémon order") but fails to execute it until the battle has already started.
  • Dry runs: There are many instances where the 3.0 agent called a tool but made a mistake with the tool call parameter, resulting in a dry run. However, unlike the 2.5 agent, it typically recognizes this mistake and self-corrects in the subsequent turn.
  • Parallel planning: The 3.0 agent has a tough time planning to execute multiple large goals in parallel for efficiency gains, instead preferring to solve tasks one at a time, even if it would be possible to make progress on multiple goals simultaneously.

The takeaway

In this race, Gemini 3 Pro moved beyond simple instruction following and demonstrated genuine spatial reasoning, improvised tool creation, and a "scientific" approach to hypothesis testing.

This reasoning capability translated directly to efficiency. Gemini 3 Pro completed the run in 17 days using 1.88 billion tokens. Based on the Mineral Badge milestone, Gemini 2.5 Pro is projected to require 69 days and over 15 billion tokens to achieve the same outcome.

To start building your own autonomous agents, check out the Gemini 3 documentation for technical implementation details.

بنقرة واحدة حفظ

استخدم YouMind للقراءة العميقة للمقالات سريعة الانتشار بتقنية الذكاء الاصطناعي

احفظ المصدر، واطرح أسئلة مركزة، ولخص الحجة، وحوّل المقالة واسعة الانتشار إلى ملاحظات قابلة لإعادة الاستخدام في مساحة عمل واحدة تعمل بالذكاء الاصطناعي.

اكتشف YouMind
للمبدعين

حول Markdown إلى مقالة 𝕏 نظيفة

عندما تنشر كتاباتك الطويلة، فإن الصور والجداول وكتل التعليمات البرمجية تجعل تنسيق 𝕏 مؤلمًا. YouMind يحول مسودة Markdown كاملة إلى مقالة نظيفة وجاهزة للنشر 𝕏.

حاول Markdown إلى 𝕏

المزيد من الأنماط لفك التشفير

المقالات الفيروسية الأخيرة

استكشاف المزيد من المقالات الفيروسية