The model most similar to Fable isn’t Opus. It’s not Sonnet either.

@arena
АНГЛИЙСКИЙ11 сент. 2026 г.
122K
197
10
10
109

Суть

An analysis of over 30,000 AI model battles reveals that LLMs are converging conceptually across national boundaries, meaning higher token prices rarely buy unique perspectives for creative brainstorming.

by @DawidGalarowicz and @petergostev

Arena's Battle Mode lets us compare two models answering the same user prompt. That makes it possible to measure not only preference, but also how much conceptual ground they share.

We reviewed 30,086 pairs of AI answers to the same real-world Text Arena prompts submitted between May 1 and September 2. Here is what we found about how similar the models are.

Key findings

  • Different LLMs cover different ground. Models shared just 43.1% of their ideas.
  • Adjacent generations of the same model family echo each other. Grok 4.5 and Grok 4.6 shared 59.7% of their ideas, GPT 5.5 and GPT 5.6 shared 59.2%.
  • US and Chinese models draw on the same ideas. Fable 5 and GLM 5.3 shared 55.9% of the points, with GLM, Kimi, Opus and Sonnet forming a higher-overlap neighborhood.
  • Paying a higher token price does not buy a different perspective. Fable 5 and DeepSeek V4 Pro were separated by a 12× price difference but shared 59.2% of their ideas.
  • Models are converging in how they think. Most recent releases appear to cover the same ideas more frequently than their prior versions.

Models output vary more than you might expect

Arena.ai - inline image

Conceptual similarity between each pair of models, along with the similarity between a single model and all other LLMs.

To capture how similar the content of the model responses was, we defined a metric called conceptual similarity. This approach boiled down the user-AI interactions into five key points two separate LLMs made in response to the same query, and estimated the overlap between the resulting ten.

On average, about four in ten distinct points identified within a battle appeared in both responses. In spite of frontier models following similar training recipes, we can note they regularly select different subsets of the available ideas.

For instance, two models might cover the same steps when asked for a pancake recipe, but the first one might decide to recommend toppings, while the other one can include mistakes to avoid.

This difference is narrowed down among adjacent model generations - they often lead to limited second opinions. Grok 4.5 and Grok 4.6 shared 59.7% of their ideas, while GPT 5.5 and GPT 5.6 shared 59.2%. These pairings may be useful when the goal is validation, but they provide comparatively less conceptual range.

The highest observed overlap was between Kimi K3 and Muse Spark 1.2, at 63.5%. At the other extreme, Bytedance’s Dola Seed 2.0 Pro was especially differentiated, averaging about 31.1% across its observed pairings.

US and Chinese models think alike

Arena.ai - inline image

Cross-model relationships presented spatially - nearby models selected similar ideas in response to the same user prompts; axes and orientation have no intrinsic meaning, it’s their relative distance and neighbourhoods that indicate conceptual closeness. Areas surrounding Anthropic models in focus.

Given there is a closer relationship between models released by the same lab, we might also expect to see a stronger link between releases from the US compared to those outside of the country.

However, it turns out that model similarity cuts across national boundaries. GLM, Kimi, Opus and Sonnet occupy the same higher-overlap neighbourhood rather than forming isolated American and Chinese blocs. This is especially clear when the overlap with the three Claude models is overlayed.

While the specific reasons for this are hard to pin down, the conceptual similarity at the frontier could simply stem from the models sharing their core training materials and the recent cross-lab focus.

Model overlap is growing

Arena.ai - inline image

Conceptual similarity for each model considering similarity to releases that preceded it.

This type of similarity is also on the rise. Comparing each release against previous models that appeared in Battle Arena shows the conceptual overlap has grown. For example, Muse Spark 1.2 (xHigh) showed greater conceptual overlap with earlier models than Muse Spark 1.1 - 48% versus 45%.

The broad picture also suggests that the likelihood of wildcard responses may be coming down, and the next generation of low-cost models is likely to raise the bar of robustness even higher.

On the other hand, there are scenarios where “reverting to the mean” might be a problem. For example, brainstorming a book chapter with a set of agents, or looking for innovative ways to market a brand of cereals in a new geographical market.

Creative work gains most from a second perspective

Arena.ai - inline image

Conceptual similarity across models considering battles which fell into one of the categories above - mean scores along with 95% confidence intervals.

Luckily, task type changes model overlap substantially, and in a direction that accommodates both exploration and exploitation of ideas.

Medicine and healthcare produced the most similar answers, at 51.1%, followed by life and social sciences at 49.1% and legal and government at 48.9%. At the other end, creative writing produced the least similar answers, at 34.9%. Entertainment and media followed at 36.8%, with coding at 39.7%.

As constrained tasks tended to contain a narrower set of facts or procedures, models converged. Conversely, open-ended tasks gave models more room to choose their own framing and examples, and the diversity between answers was much higher.

Paying more does not lead to better brainstorming

Arena.ai - inline image

Number of new ideas generated by using another LLM in addition to Fable and contrasted with costs of doing so - mean and 95% confidence intervals for the ideas count.

What if exploration, rather than exploitation was key? The data suggests that once a frontier model has provided a strong foundation, using another SOTA model to expand on it leads to a poor payoff.

For example, where Fable 5 answers were available, using MiniMax M3 resulted in more ideas than spinning up Opus 5. Similarly, accompanying it with Mistral Medium 3.5 answers would have provided more ideas per dollar than using GPT 5.6 Sol.

One limitation of this is that the ideas might not be of equal quality. However, once reviewed by a human expert, their diversity is likely to spur more lateral thinking, that ultimately is required in creative tasks. Optimally, a frontier model would be paired with a smaller, yet still robust LLM.

Methodology

Data

The analysis covers 30,086 Text Arena comparisons recorded from 1 May to 2 September 2026. It focuses on English-language prompts and includes Battle and Direct Battle interactions.

Approach

LLM-as-a-judge extracted up to five supported key ideas from each model's answer. A second judge pass reviewed both original answers and the combined ideas, grouping every idea as shared by both models, present only in A or present only in B. The complete judgement was repeated with the model order reversed to reduce ordering effects. Mean values from those judgements form the basis of this analysis.

For each comparison:

conceptual similarity = shared ideas / (shared + A-only + B-only ideas)

A score of 1 means every supported idea was shared; a score of 0 means none was.

Two real-world examples showing what overlap means

  • Writing an online article

A user asked for an SEO-optimised article about unzipping folders on a MacBook. DeepSeek V4 Pro High and Fable 5 High produced closely aligned responses, including similar structures. Their few differences were narrow: DeepSeek noted that macOS lacks native support for some archive formats, while Fable highlighted failure cases such as file corruption.

The conceptual similarity score was 0.833. The second answer mostly validated the first.

  • Drafting an HR policy email

A user asked for an email communicating a change to an HR medical policy. Opus 5 High discussed registered family members, PTO protections and a checklist of issues to resolve before sending the message. GLM 5.3 Max instead raised how the policy might affect employees already on PTO.

The conceptual similarity score was 0.106. Neither answer simply repeated the other. Taken together they expose a broader set of policy risks.

Переделать в YouMind

Превратите одну вирусную статью в полноценный рабочий процесс создания контента

Собирайте источники, расшифровывайте паттерны, создавайте активы, пишите черновики и публикуйте контент из одного рабочего пространства ИИ.

Исследовать YouMind
Для авторов

Превратите ваш Markdown в аккуратную статью для 𝕏

Когда вы публикуете длинные тексты, изображения, таблицы и блоки кода, форматирование в 𝕏 становится мучением. YouMind превращает полный черновик в Markdown в чистую статью, готовую к публикации в 𝕏.

Попробовать Markdown для 𝕏

Другие паттерны для анализа

Недавние виральные статьи

Смотреть другие виральные статьи