Reviewing Kimi K3: Small-Town Youths vs. AI Capitalists

@0xcherry
TIẾNG TRUNG2 ngày trước · 19 thg 7, 2026
159K
308
49
23
289

TL;DR

Kimi K3 marks a shift toward massive scaling (2.8T parameters) to compete with top-tier US models. By releasing open weights, it creates an asymmetric competition that undermines the high-margin business models of closed-source labs.

**

车厘子 - inline image

The Three Chariots of Chinese AI

For a long time, Chinese AI models have roughly followed three paths.

The first path is intentionally going small. Examples include minimax-m3, Kimi K2.5, and hy3, which actively limit activated parameters to around 20B. Whether 20B is truly "small" is a subtle question, but in engineering terms, we generally consider it so.

Smaller activated parameters are usually a hard economic constraint. Maximizing performance within these constraints is a typical "small model" mindset.

The second path is DeepSeek, and the manufacturers with similar technical stacks and sizes that emerged after it. The hallmark of the DeepSeek route is not just being fully open-source, but more importantly, that DeepSeek-V3/R1 were large-scale models from the start, benchmarking against OpenAI's top-tier models.

As R1 exploded in popularity, DeepSeek added a vast technical stack focused on economic efficiency. Thus, the DeepSeek route is characterized by "high performance while ensuring inference economy." GLM-5 is a systematic successor and practitioner of this path, with GLM-5.2 being one of its well-known results.

The third path is Kimi. Kimi's technical stack has always been independent and is distantly related to DeepSeek. While the K2.5 to K2.7 series significantly absorbed DeepSeek's tech stack, the new K3 series has returned to its proprietary technology.

K3 has pushed total parameters to 2.8T, with activated parameters estimated at 50B, using entirely proprietary tech. The scaling law of size has taken effect, but it has also caused Kimi K3's economic efficiency to drop significantly, entering the price range of Claude Sonnet.

Although still very cheap compared to Claude, Kimi's inference clusters are located within China. Given the limitations on GPU counts, Kimi likely did not prioritize economy when building K3.

In other words, K3 is the product of an independent tech stack, massive size, and an all-out pursuit of high performance.

Unlike models on the DeepSeek path that balance performance and economy, K3's goal from the start was not to make a better R1, but to benchmark against the cutting-edge, ultra-large models from OpenAI and Anthropic.

Why is K3 So Big?

Describing Kimi as an independent technical route never related to DeepSeek is not entirely accurate.

The Kimi K2 paper made it clear: K2 used an ultra-sparse MoE and MLA architecture similar to DeepSeek-V3. Both have 61 layers and shared experts, with each token activating 8 routed experts. K2 improved on this base by increasing the number of experts to 384, halving attention heads, and using self-developed MuonClip, data recipes, agentic post-training, and independent infrastructure to create its own fork.

Therefore, K2 was more like a Kimi branch on a DeepSeek-V3-style chassis rather than a different species from scratch.

What is truly noteworthy is that the base size from K2.5 to K2.7 did not continue to expand: total parameters remained around 1T, activated parameters around 32B, still with 8 experts selected from 384.

In other words, Kimi proved it could squeeze many generations out of the same chassis using post-training and agent systems without expanding the base.

But K3 is different. K3 is not just a routine version upgrade. It marks Kimi's active decision to end the phase of "continuing post-training on an existing base" and return to pre-training scale expansion.

车厘子 - inline image

K3 heavily utilizes previously dispersed proprietary technologies, such as its own training base, attention mechanism, and expert routing structure. At this stage, K3's kinship with DeepSeek has become quite distant.

K3 aims to prove two things: first, that Kimi can define its own technical paradigm; second, that Kimi is qualified to use new technology to train brand-new, ultra-large models that benchmark against OpenAI and Anthropic.

As it turns out, K3 succeeded.

From Public Utility to Privileged Monopoly

Before DeepSeek, China did not lack usable models, but they were only barely functional.

The emergence of DeepSeek brought a high-performance model available in China for the first time, becoming a public utility for a generation of Chinese AI. The core mission of a public utility is Adoption: how to promote "good" capabilities on a large scale—especially on an inference base that isn't exactly wealthy?

Therefore, DeepSeek's role was not just a performance breakthrough, but also a heavy responsibility for economic optimization.

From V3 to V4, DeepSeek has maintained near-charitable API pricing, backed by a mature technical stack oriented toward inference optimization.

In this stage, the problem facing Chinese AI was: first meet the performance needs of the public utility, then reduce the computing costs of that utility.

For Chinese manufacturers other than DeepSeek, the task suddenly became twofold: first, performance must catch up with R1, or the model has no competitive standing; second, price, throughput, and deployment efficiency must be as close to DeepSeek as possible, or adoption simply won't happen.

GLM's phased stack transition is the clearest microcosm of this industrial pressure.

In this sense, DeepSeek brought Chinese models into the Adoption phase. The core question of this phase is: how to let more people use sufficiently strong intelligence, and how to handle the resulting explosion in call volume.

But a frontier lab must eventually answer another question: if you only ever pursue being a cheaper version of others, where will higher-performance models come from?

All major internet companies made a choice: find ways to bypass bans and import from the US. For example, Alibaba's famous Singapore entity allowed users to run Claude Code via remote login to Singapore machines.

Then they got blocked. If Uncle Sam hadn't quickly delivered the 5.6-Sol blow, Dario might have published a few more articles stepping on the heads of the Chinese.

Through extensive engineering optimization, China has gradually realized the foundation for the public utility of AI. But without larger, higher-performance frontier models, it will always be a step behind, falling into a long-term disadvantageous "privileged competition."

Xiang Zhuang's Sword Dance: Aiming Beyond the Surface

K3's 2.8T is Moonshot's answer to this problem.

K2.5 to K2.7 can still be understood as improving cost-performance on a 1T-A32B base through continued pre-training, RL, tool environments, and agent systems.

K3, however, pushes total capacity directly to 2.8T, accepting deployment requirements of over 64 acceleration cards and raising API prices from K2.7 Code's $0.95/$4 to $3/$15. Although this price is only comparable to Sonnet, it is still very expensive.

The key is the 64 acceleration cards. Where are you going to find so many 64-card nodes in China's inference clusters?

K3 might still hope to maintain economy, but that is nearly impossible for K3. A 2.8T giant model is a major test even for American inference companies. Therefore, K3's intent is clearly not for widespread adoption, but to open-source a model standard that can be widely adopted in the most cutting-edge battlefields.

Compared to most PR pieces claiming to crush American models, Kimi's official wording is more moderate: K3 overall still lags behind Fable 5 and GPT-5.6 Sol.

Of course, from another perspective, it's very aggressive: it's only slightly behind the top-tier models from OpenAI and Anthropic.

Once it enters this bracket, K3 first destabilizes Anthropic's most important commercial narrative of the past year.

Anthropic's Perpetual Motion Story

Anthropic's methodology has always been clear: we create an irreplaceable model, and once users find it can complete tasks other models can't, they will be willing to pay a premium for higher performance.

Fable 5 is the most extreme product of this methodology. It targets the most difficult knowledge work, coding, and complex asynchronous tasks that can last for days, with API prices reaching $10 per million input tokens and $50 per million output tokens—double the regular price of Opus 4.8.

Anthropic isn't selling more expensive tokens; it's selling a new task interval: projects that couldn't be stably completed in the past can now be handled by Fable. In these scenarios, Fable has irreplaceable value.

As long as this irreplaceability exists, the price premium is perfectly reasonable. After all, high-performance models replace high-knowledge labor, and the latter's salary is much more expensive than the model.

After Fable's release, the market briefly formed an illusion that OpenAI could no longer catch up with Anthropic. SemiAnalysis's super-bullish projection even estimated that Anthropic's ARR could reach $300 billion by the end of 2027, with a $6 trillion enterprise value at a 20x ARR multiple.

This prediction relied on a set of assumptions: API accounting for 75%—85% of revenue, high API gross margins, extremely high net revenue retention, and continuous penetration of Claude Code.

Standing at the time of July 18, it seemed as if SemiAnalysis was hallucinating. But on July 8, many people found it quite plausible.

Reality arranged a highly dramatic timeline for this narrative. The $6 trillion valuation projection was widely circulated on July 8, and GPT-5.6 Sol was officially released on July 9.

Anthropic's perpetual motion story lasted less than a month before being overturned. During that time, a few articles criticizing the distillation of Chinese models were also thrown in.

Who Killed Cock Robin?

After the release of GLM-5.2, I wrote a systematic analysis of its significance, which was hailed by the public as a "masterpiece."

Considering Zhipu has never paid me for promotion, I remain an independent researcher (laughs).

However, GLM-5.2 did indeed shake Anthropic's foundation; people just didn't realize the severity of the situation at the time.

GLM-5.2 performs better than Opus 4.8 on most tasks, second only to Opus 4.7. You read that right: better than 4.8 while being second only to 4.7, which makes one wonder how much water Anthropic diluted into 4.8.

Meanwhile, GLM-5.2's price is only one-fifth of Opus, and it doesn't suffer from the rate-limiting issues of Anthropic's subscription plans. The emergence of GLM-5.2 has already significantly threatened the price anchor of Anthropic's main model segment.

As a result, the market hoped to price Anthropic for "more cutting-edge models," which led to SemiAnalysis's $6 trillion valuation prediction.

Immediately after, Sol collided head-on with Fable in the highest capability tier. Anthropic subsequently extended Fable's subscription limits multiple times and eventually announced it would become a permanent benefit for Max and Team Premium.

车厘子 - inline image

Then K3 came out. Its comprehensive capability lags behind Fable and Sol—and only behind Fable and Sol. Meanwhile, K3 is open-source, and any inference cluster can deploy it.

Children, advanced AI models really do grow in the fields.

Asymmetric Competition

In the past, the logic of US advanced computing chip controls on China was straightforward: slow down China's training of frontier models and the construction of an independent semiconductor ecosystem by limiting upstream computing power. National security, military, and surveillance uses were the public legal reasons, while maintaining US leadership in key technologies was a concurrent industrial goal.

Of course, the bigger problem was that the US's own advanced chips weren't enough. The US was not only worried about China obtaining computing power but also didn't want Chinese companies participating in bidding, crowding out the card-buying capacity of US labs and data centers.

These controls certainly increased the difficulty for China to train frontier models, but they also created a subtle counter-effect: they deprived Chinese AI companies of the ability to earn excess profits, while simultaneously turning this disadvantage into a weapon in the game.

Since you won't let me earn excess profits, and I don't want my technology blocked, I'll just give it away for free once I make it.

And open-source models aren't picky about graphics cards. Chinese cards can deploy them, and US computing clusters can too. Fireworks, TogetherAI, and a bunch of NeoClouds all have plenty of GPUs.

Thus, a highly abstract combination has emerged in the US-China AI competition:

Chinese Models + US and Global Inference Clusters VS US Closed-Source AI Labs

Chinese labs undertake expensive model R&D, trading open weights for adoption and frontier status; US cloud providers and inference platforms use local GPUs to handle calls, pricing based on cost and making a killing.

The ones truly losing profit are the US frontier labs that rely on closed weights and capability scarcity to charge premium prices.

For Chinese labs, sacrificing open weights means losing potential API rent that was already limited by their own inference capacity.

For US closed labs, what is sacrificed is the actual excess profit built on abundant local computing power and global capability scarcity.

Therefore, the US faces not a simple choice of "continue sanctions or open sales," but an impossible trinity:

  • Sanction China, and China will continue to release open-source models, so no one makes money.
  • Let China buy cards, and collective bidding might leave local labs short on cards.
  • Give China some cards appropriately to achieve a compromise, but how many is appropriate? No one knows, and no one dares to speak up.

This is typical asymmetric competition. Washington loves to talk about asymmetric competition; now the term has come back to haunt Washington.

Loyalty to Micron!

Pushing this chain one step further upstream touches the entire AI CapEx narrative.

In the past, the high inference gross margins of closed-source frontier models were the most important cushion for the AI CapEx chain. As long as model capabilities were scarce, APIs and subscriptions could maintain high prices; labs and cloud providers obtained high gross margins, allowing them to tolerate more expensive GPUs, faster depreciation, and more aggressive data center investments; upstream would then continue to expand production and maintain high pricing.

This chain can be simply written as:

Model Capability Scarcity → High API Prices → Downstream High Gross Margins → Ability to Withstand GPU Premiums → CapEx Expansion → Upstream Continued Prosperity.

When multiple platforms can deploy the same weight, inference prices shift from "how much the frontier lab is willing to charge" back to "how much it costs an efficient cluster to run a successful task."

Open weights don't change the fact that computing power costs money; they remove the "model tax" attached to computing consumption. Fireworks doesn't have to bear K3's pre-training costs for Moonshot, nor does it have to pay full model rent to a closed-source model upstream per token.

It still has to bear the costs of GPUs, inference engines, quantization, caching, networking, financing, SLAs, and enterprise sales, but it only needs to quote based on actual computing costs, utilization, service costs, and competitive profit. This is typical cost-based pricing logic, and for an industry dependent on prosperity, cost-based pricing has a disastrous impact on the upstream.

This will inversely change the capital return calculation for every GPU.

In the past, service providers could put expensive hardware costs into high-priced model APIs and pass them on to end customers.

When inference platforms compete around the same open weights, it becomes harder for any of them to continue passing on this premium. Thus, GPU procurement requires higher utilization, shorter payback periods, lower financing costs, and more certain customer contracts.

In the short term, more dispersed downstream buyers will not weaken the bargaining power of computing producers (especially NVDA). But in the long term, the systemic decline in downstream gross margins will inevitably lead to worse bargaining elasticity for the upstream, which is an uncontrollable and huge risk for the entire AI CapEx narrative.

Too Big to Fail

At this point, K3's 2.8T has revealed two seemingly contradictory meanings.

On one hand, K3 is not a very economical model, pushing models into the Sonnet price range and the 64-card+ deployment bracket.

On the other hand, it is precisely this "not prioritizing economy" that qualifies K3 to enter the highest capability group, destabilize Anthropic's model rent, restructure US-China inference profits, and influence the AI CapEx narrative.

In the past, many companies liked to talk about edge-side, small sizes, and "good enough." These routes are certainly important; they determine how far existing intelligence can spread to devices, users, and industries.

But the term "good enough" is very boring; it only applies to markets where capability boundaries are already fixed.

Frontier AI never faces a fixed set of tasks: a model that just learned single-file programming today will have to handle entire codebases tomorrow; one that just learned tool calling will have to work continuously for hours or even days in the next generation.

Moderately sized models answer "how to be used by more people," while larger models compete for "who has the right to deliver the oracle."

A model that is already strong but expensive can later be made faster and cheaper through sparsification, quantization, caching, distillation, and inference engineering; it can also serve as a teacher model for an entire line of small models.

A small model that didn't learn a certain capability from the start, however, is hard to supplement out of thin air through extended thinking or deployment optimization.

This asymmetry makes "going big so you won't lose" a crude but often effective frontier methodology.

In 2024, OpenAI released o1, shifting the focus of scaling from larger pre-training bases to training-time reinforcement learning and test-time thinking computation.

o1 itself is a smaller inference model. It proved that "smaller base + more thinking" can create amazing progress in math, code, and formal reasoning, thereby starting a cycle of making parameters smaller.

The smaller the base, the more it relies on test-time compute compensation; the more successful the inference tricks, the less motivation the organization has to expand the base. Consequently, OpenAI's models began to deteriorate rapidly.

In 2025, Anthropic went the opposite way. Anthropic insisted on making larger, more expensive models, then used Claude Code to turn capabilities into products.

User loyalty to Claude isn't because it's cheap, but because on the most difficult real-world software engineering tasks, they feel more secure giving the work to Opus. By the end of 2025, Anthropic claimed Claude Code occupied over half of the AI coding market, with its enterprise market share rising from 24% to 40%. Revenue also surpassed OpenAI's.

OpenAI subsequently realized that inference efficiency and routing cannot replace the highest capability base. GPT-5 re-established a complete ladder of main, thinking, mini, and Pro; by GPT-5.6, it clearly set three model sizes—Sol, Terra, and Luna—and put the large-scale Sol back in the flagship position for the most difficult work.

As a result, Codex users quickly surpassed 10 million, forcing Anthropic to add Fable back to its subscription plans.

Now, K3 is released, second only to Sol and Fable, even sparking discussions about whether US frontier AI is still up to the task.

When you are truly willing to compete for first place, the result is usually not too bad. This is Too Big to Fail.

Conclusion: Small-Town Youths Beating Up Capitalists

Chinese AI model companies are often described as small-town youths: limited resources, poor location, even worse international reputation, and lucky if their GPUs can sustain inference, let alone training—just as long as they don't fall behind.

Meanwhile, advanced US AI labs have a capitalist flair: discussing anthropology and safety, discussing benevolent machines, discussing whether the future world belongs to freedom or unfreedom, whether AI belongs to the earth or the sky.

Under such a dramatic gap, a play where small-town youths beat up capitalists can actually be staged. No matter how you spin it, it's enough to make one laugh their head off.

—Actually, no need to rush. The best is yet to come.

Viết lại trong YouMind

Turn one viral article into a full content workflow

Collect the source, decode the pattern, create assets, draft the story, and distribute from one AI workspace.

Explore YouMind
Dành cho nhà sáng tạo

Biến Markdown của bạn thành bài viết 𝕏 gọn gàng

Khi bạn đăng bài viết dài của riêng mình, việc định dạng hình ảnh, bảng và khối mã cho 𝕏 rất mệt mỏi. YouMind biến cả bản nháp Markdown thành một bài viết 𝕏 gọn gàng, sẵn sàng để đăng.

Thử Markdown sang 𝕏

Thêm pattern để giải mã

Bài viết viral gần đây

Khám phá thêm bài viết viral