YouMind
Увійти

The Real Gap in Chinese AI May Be Widening: Insights from a Former ByteDance Researcher

@Gorden_Sun
СПРОЩЕНА КИТАЙСЬКА26 квіт. 2026 р.
217K
465
92
43
625

Коротко

Peking University professor Zhang Chi explains how 'bench-maxing' and a lack of high-quality user feedback are crippling Chinese AI development, forcing companies to rely on 'distilling' US models instead of innovating.

The release of DeepSeek V4 did not replicate last year's frenzy. In fact, compared to Claude Sonnet 4.5 released six months ago, their capabilities are roughly in the same tier, but the gap is much larger than six months because Sonnet 4.5 was only considered second-tier half a year ago. However, in social media articles, we often see Chinese large models producing increasingly beautiful benchmark data, with claims of "only six months behind" or "basically caught up" being heard everywhere.

What is the actual situation regarding the AI gap between China and the US?

On April 22, in the "Into Asia" podcast, Zhang Chi, an assistant professor of AI at Peking University, told the truth as he sees it. Zhang Chi is currently an assistant professor at Peking University and recently resigned from ByteDance's core large model team (Seed LLM).

As an R&D professional who has truly worked on the front lines of a major tech company, his judgment of current domestic AI is quite stinging:

"I do not agree with the view that Chinese models are catching up. I believe we are still far behind, and this gap may be widening."

▸ False Prosperity: Everyone is "Teaching to the Test," but Real Combat is Lacking

To the outside world, models from various tech giants are engaged in a fierce battle on various benchmarks, with scores repeatedly hitting new highs. But internally, this is just a massive "exam-oriented education" for large models.

Zhang Chi revealed in the interview that inside ByteDance (and he suspects other big tech firms are similar), the working atmosphere is actually relatively "chill" (with a two-hour lunch break and about 9 actual working hours a day), but everyone faces an implicit KPI pressure—Bench-maxing.

Leaders pay close attention to model scores on specific leaderboards. If the module you are responsible for does not match the scores of leading US models, your performance review will look very bad.

Result: The data on paper is extremely gorgeous, but once it hits complex real-world applications, the experience is frustrating.

▸ The Chasm in Compute and Infrastructure: Three Months for Others, Maybe Half a Year for Us

Hardware bottlenecks are an old story, but the chain reaction they cause is deeper than we imagine.

Currently, a large part of what domestic giants use to train their core models is still NVIDIA chips stockpiled before the ban, or the compliant H20 special editions. Fortunately, starting with DeepSeek V4, there is a full transition to Huawei Ascend graphics cards, which is expected to improve the domestic training ecosystem.

But the gap in computing power is already directly reflected in "iteration speed."

Zhang Chi mentioned an industry rumor: Google might now only need 3 months to complete a full round of pre-training and post-training for a large language model. For domestic giants, limited by the scale of computing power and infrastructure, this cycle could be as long as half a year.

More hidden is the gap in infrastructure (Infra). Zhang Chi, who interned at Google, lamented that the underlying infrastructure there is so well-done that researchers only need to write code on a smooth graphical interface without worrying about the underlying architecture. In domestic tech giants, training frequently freezes or throws errors; these friction costs are invisibly slowing down the pace of catching up.

▸ "Users are all using US models; where will we get the data to improve?"

If computing power is the first sword hanging over Chinese AI, then in Zhang Chi's view, the second sword—and currently the most unsolvable one—is the rupture of the "data flywheel."

He offered a very sharp insight in the interview: Leading US models have established a positive cycle that is extremely difficult to overcome. GPT and Claude have massive global user bases. These users use the models in actual work and "like" or "dislike" the results. This high-quality feedback constitutes the most precious training data for real-world scenarios.

In contrast, due to the objective gap in basic capabilities, high-value users who need AI assistance the most—such as programmers and hardcore researchers—are "defecting" en masse.

"I now mainly use Claude Code and Cursor for programming," Zhang Chi said bluntly. "I even feel I don't need to recruit so many PhD students to help me; I can completely treat Claude Code and Cursor as my students. I can mentor them and give them instructions to do what I want. But I am also conflicted: if my generation doesn't train new people, who will continue the research when I'm old?"

This daily choice by a top Chinese AI scientist reflects the cold reality: When the top Chinese developers who should be contributing feedback data to domestic models are all using US models to increase efficiency, where will Chinese large model companies get the high-quality interaction data to optimize programming and reasoning capabilities?

▸ The Price of Taking Shortcuts: "Distilled" Intelligence Has No Soul

If there is no time to polish infrastructure and one faces the urgent pressure of catching up with KPIs, what do domestic giants do?

The answer is one word: Distillation.

If you want to train a high-intelligence model, the most hardcore way is to hire extremely professional industry experts to write high-quality reasoning data stroke by stroke, which is both expensive and time-consuming.

But there is a shortcut: Ask GPT, Claude, or Gemini directly. After getting the correct answer and reasoning process, copy it over and feed it to your own model. This is known as "distillation" in the AI circle—essentially copying the top student's homework.

Zhang Chi admitted that we might already be world-class in "distillation" technology, but this may not translate into a true advantage in the long run. Copying homework can help you quickly go from failing to passing, or even to a score of 80, but you can never become a true top scholar by copying.

Because you lack your own deep data pipeline. When foreign models begin to evolve autonomously, "shortcuts" instead become shackles that bind our original capabilities.

▸ The Only Remaining Confidence: Hardware and the Dream of "Embodied AI"

Despite his strong pessimism about the prospects of catching up in pure large language models, Zhang Chi still pointed out a few structural advantages in China's AI ecosystem.

In his view, the advantage lies in manufacturing. He mentioned Unitree, which recently sparked public discussion, believing that China has global competitiveness in hardware bodies and motor motion control. Regarding the currently hot "Embodied AI," Zhang Chi's view is that if your language model is only used to perform relatively simple tasks (like grabbing objects), then the capabilities of existing Chinese large models are "good enough."

But he also poured cold water: currently, the vast majority of robot manufacturers are still stuck in the "motion control" stage and haven't truly put intelligence into the robot's brain. Once complex reasoning and generalized "dexterous manipulation" are involved, we are likely to hit the same ceiling that large language models currently face.

▸ Future?

Limited chips, weak data pipelines, lagging infrastructure, lack of user feedback loops, and over-reliance on distillation—these problems combined cannot be solved by a single technical breakthrough. Fortunately, DeepSeek V4 is fully adapted to domestic graphics cards. Although the overall capability is somewhat behind, there is still hope to catch up once the ecosystem is perfected, and without relying on distillation.

Original Podcast Link: [https://www.buzzsprout.com/2546300/episodes/19057945-a-year-inside-bytedance-s-ai-lab](https://www.buzzsprout.com/2546300/episodes/19057945-a-year-inside-bytedance-s-ai-lab)

Збереження в один клік

Використовуйте YouMind для AI-глибокого читання віральних статей

Зберігайте джерела, ставте цілеспрямовані запитання, підсумовуйте аргументи та перетворюйте віральні статті на корисні нотатки в одному AI-робочому просторі.

Дослідити YouMind
Для авторів

Перетворіть свій Markdown на охайну статтю для 𝕏

Коли ви публікуєте власні лонгріди, зображення, таблиці та блоки коду роблять форматування в 𝕏 складним. YouMind перетворює повну чернетку в Markdown на чисту статтю для 𝕏, готову до публікації.

Спробувати Markdown для 𝕏

Більше патернів для аналізу

Останні віральні статті

Переглянути більше віральних статей