A Conversation with Yang Zhilin of Kimi: Advancing Toward the Endless, Unknown Snow Mountains

@zhang_benita
УПРОЩЁННЫЙ КИТАЙСКИЙ3 дня назад · 19 июл. 2026 г.
176K
267
28
13
274

Суть

Moonshot AI founder Yang Zhilin discusses the technical philosophy behind Kimi, emphasizing long context as a path to AGI and the future of unified world models.

This was my first interview with Yang Zhilin, conducted in early 2024 and published on March 1, 2024—exactly the first anniversary of Kimi's founding. At the time, Kimi had only 80 people, working out of their first, somewhat run-down office. There was no logo at the entrance. Only a white piano standing guard by the door.

This article generated quite a stir in China's tech community at the time.

Back then, I was still a print journalist, so this interview exists only as text and an audio podcast.

As we can see, many of the views expressed in this article have since been borne out.

Just rereading these words from two years ago, you really can't help but marvel at how dramatically the world has changed!

(This translation was generated by Kimi K3.)

Yang Zhilin: “If everyone thinks you are normal—if your dream is one that anyone could have—it adds nothing to the sum total of humanity’s dreams.”

By Zhang Xiaojun

Just one year ago, AI scientist Yang Zhilin did a precise calculation in Silicon Valley. He realized that if he decided to launch a foundation-model startup aimed at AGI, he would need to raise more than $100 million within the next few months.

Yet that was merely a ticket to the game. A year later, that figure had multiplied thirteenfold.

For foundation-model companies, competition is less a scientific contest than, first and foremost, a brutal contest of money. With investors holding their purse strings tight, you have to stay ahead of your rivals in raising more money, buying more GPUs, and grabbing more talent.

“It requires a concentration of talent and a concentration of capital,” says Yang Zhilin, founder and CEO of Moonshot AI, the foundation-model company established on March 1, 2023.

Over the past year, Chinese foundation-model companies have seemed to live on a tense, constricted edge of survival. On the surface, each of them holds hefty sums of cash. But on one hand, they must immediately pour freshly raised money into extremely costly research to chase OpenAI—first catching up to GPT-3.5, and before GPT-4 is even reached, along comes Sora. On the other hand, they must race nonstop to find viable real-world use cases, to validate for themselves that they are companies, not research institutes that only devour capital. And that is not enough: for every one of these ventures, whether the exit is an IPO or an acquisition, the way out remains anything but clear.

Among the founders of China’s foundation-model companies, Yang Zhilin is the youngest, born in 1992. The industry describes him as a staunch AGI believer and a founder with rare technical charisma. Much of his academic and professional record is tied to general-purpose AI, and his papers have been cited more than 22,000 times.

In mid-2023, China’s tech community turned abruptly from euphoria to chill on foundation models, and accelerating real-world deployment became the pragmatic mainstream melody. This inevitably left foundation-model CEOs torn violently between ideal and reality. In a Chinese AI ecosystem where everyone chants PMF (product/market fit) and everyone chants commercialization, this founder—an AI researcher by training—is in no particular hurry.

With 80 people, Moonshot AI has the smallest headcount among China’s leading foundation-model companies. Unlike his rivals, Yang did not opt for the safer B2B business or seek deployment in verticals such as healthcare or gaming. He built one—and only one—consumer product: the AI assistant Kimi, which accepts inputs of up to 200,000 Chinese characters. Kimi is also Yang Zhilin’s English name.

Yang prefers to see his company as a system that combines science, engineering, and business. You might picture it this way: above the human world, he is erecting an AI laboratory bench—with one hand he runs experiments, and with the other he brings cutting-edge technology down into the real world, discovering applications through interaction with people and delivering those applications into consumers’ hands. Ideally, the former burns through billions and tens of billions of dollars of capital; the latter earns that money back hundreds or thousands of times over. However you hear it, it sounds as thrilling—and as perilous—as walking a tightrope.

“AI is not about what PMF I can find in the next year or two; it’s about how to change the world over the next ten to twenty years,” he says.

Such abstract, idealistic thinking makes one sweat nervously on his behalf: can a young AI scientist carve out room to survive in a realist China?

In February 2024, Moonshot AI closed a large funding round against the market tide. It is understood that the company raised a Series B of more than $1 billion at a $1.5 billion pre-money valuation, led by Alibaba with follow-on participation from Monolith Management, Xiaohongshu, and others. Upon completion of the deal, Moonshot AI’s post-money valuation stood at roughly $2.5 billion—making it, at this stage, the highest-valued unicorn in China’s foundation-model race. (The company declined to respond to or comment on the matter.)

In the midst of this third funding round, we sat down with Yang Zhilin to talk about his first year of entrepreneurship—a cross-section, in miniature, of a year in which Chinese foundation-model companies raced ahead from the starting line.

His company did not set up in Sohu Network Plaza in Beijing, the hub where foundation-model companies cluster. For a company with total funding of about RMB 9 billion, this office in the Liangzi Xinzuo building looks crude and run-down. There is not even a company logo at the entrance—only a white piano standing guard by the door.

The meeting room sits in a corner; with its small windows it is dark inside, and the heater hums as it blows warm air against the winter cold. In the dim light, Yang describes how the past year has felt to him: “It’s a bit like driving down a road with a range of snow mountains stretching out ahead. You don’t know what’s inside them. You just keep walking forward, one step at a time.”

Below is the full interview with Yang Zhilin. (For readability, the author has made some textual edits.)

张小珺 Xiaojun Zhang - inline image

This photo was taken in early 2024 at Kimi’s first office in Beijing. They’ve moved out of this location now. No logos lined the entrance — only a white piano stood quietly by the door.

**

Part 1

Standing at the Beginning

“You Have to Ride the Wave”

**

Zhang Xiaojun: How have you been lately?

Yang Zhilin: Busy—there’s a lot going on. But I’m still excited. We’re standing at the very beginning of an industry, and there is enormous room for imagination.

Zhang Xiaojun: When I came in just now, I saw a pure white piano at your company’s entrance.

Yang Zhilin: There’s a Pink Floyd album sitting on it, too. I have no idea who put them there—I suddenly noticed them a couple of days ago and haven’t had a chance to ask. (Pink Floyd is the British rock band that released the album The Dark Side of the Moon.)

Zhang Xiaojun: On the day ChatGPT was released in November 2022, what were you doing?

Yang Zhilin: I was already preparing for this—recruiting people, building a team, exchanging new ideas. Seeing ChatGPT was thrilling. Three to five years earlier—even in 2021—it would have been inconceivable. That kind of higher-order reasoning had been very hard to achieve.

I sensed that many variables were about to shift in the market: capital on one side, talent on the other—the core factors of production for AI. If those variables fell into place, it would become possible to build a proper company to do this—an organization built for AGI could go from 0 to 1. That was a major epiphany. An independent company made more sense, but it wasn’t something you could do the moment you wanted to; ChatGPT jolted the variables and brought the factors of production together. You have to ride the wave.

Zhang Xiaojun: After you decided to found an AGI company, what preparations did you make? How did you assemble the two factors of production—capital and talent?

Yang Zhilin: It was a winding process. ChatGPT took time to diffuse. Some people learned of it early, some late; some doubted at first, then were shocked, then became believers. Finding people and finding money were tightly bound to timing.

We began focusing on our first funding round in February 2023. Had we delayed to April, we basically would have had no chance. But doing it in December 2022 or January 2023 wouldn’t have worked either—the pandemic was still on, and people hadn’t processed it yet. So the real window was just one month.

One night in the United States, I did a precise calculation. When I finished, I concluded we needed to raise at least $100 million within a few months. Many in the market hadn’t started fundraising yet, and many didn’t believe you could necessarily raise that much. But it turned out to be possible—even more than that.

The talent market started moving, too. Inspired by ChatGPT, many people had this realization in March or April 2023: this is the only thing worth doing in the next decade. You have to reach out to the right people at the right time. A year or two earlier, talent would not have clustered to this degree. Back then, more people were doing traditional AI or AI-adjacent businesses—none of it was general-purpose AI.

Zhang Xiaojun: To sum up: February was the window for fundraising, and March and April were the window for hiring?

Yang Zhilin: More or less.

Zhang Xiaojun: That night in the U.S.—where were you when you did this math? How exactly did you calculate it?

Yang Zhilin: From late 2022 into early 2023, I spent a month or two in the U.S., talking to people. I did it where I was living. You work out how many FLOPs you need, the training cost, inference, and the user numbers.

Zhang Xiaojun: At that moment, what was the prevailing mood in Silicon Valley?

Yang Zhilin: The product began picking up many early adopters, concentrated in the tech circle. We were in that circle ourselves, so we felt it more keenly. At the big Silicon Valley companies, people have to write performance reviews every six months, and many started writing them with ChatGPT. Some people whose writing was usually not that professional turned in reviews written with ChatGPT, and everyone sounded dead serious.

Undercurrents were stirring. Many people were thinking about their next job or about starting a company. Quite a few friends who talked with us later went off to found startups. And there was intense FOMO—fear of missing out. Nobody could sleep. Whether it was midnight, 1 a.m., or 2 a.m., if you reached out, people were always there. A bit anxious, a bit FOMO, and very excited.

Zhang Xiaojun: The night you calculated you needed to raise $100 million—how late did you stay up?

Yang Zhilin: It was fine—the calculation itself didn’t take long. But afterward, I couldn’t tell too many people. If I had, no one would have believed it could be done.

Part 2

Technical Lineage

“Free Yourself from Endless Carving”

Zhang Xiaojun: When the venture capital world talks about you, they say, “The founder is brilliant, has technical charisma, and the team is full of technical stars.” So before we discuss your foundation-model venture, I’d like to start with your academic background. You studied computer science at Tsinghua as an undergraduate and earned your PhD at Carnegie Mellon’s School of Computer Science. Has AI always been your focus?

Yang Zhilin: I was born in 1992 and started my undergraduate degree in 2011. From my sophomore year to now—more than a decade—I’ve been in this field. At first I explored more divergently, looking around everywhere; I did some work related to graphs and to multimodality. In 2017, I converged on language models. At the time I felt language models were a relatively important problem; later I came to feel it was the only important problem.

Zhang Xiaojun: In 2017, how did the AI industry generally understand language models, and how did that understanding evolve?

Yang Zhilin: Back then it was a model used to rank speech-recognition results. (Laughs.) After a segment of speech was recognized, you’d get many candidate results, and you’d use the language model to see which one had the highest probability and output the most likely one. Its applications were very limited.

But you come to realize it’s a fundamental problem, because you are modeling the probabilities of the world. Language is limited, but it’s a projection of the world; in theory, if you make the token space—the space of all possible tokens—large enough, you can build a general world model. How everything in the world arises and develops can be assigned a probability. Every problem can be reduced to how to estimate probabilities.

Zhang Xiaojun: Your academic mentors are very prominent: your PhD advisors were Ruslan Salakhutdinov, head of AI at Apple, and William W. Cohen, chief scientist of Google AI. Both straddle industry and academia.

Yang Zhilin: In previous years, industry and academia came together more, but the trend is now shifting: more valuable breakthroughs will happen in industry. That’s an inevitable law of development. It starts with exploratory research and gradually shifts into a more mature industrialization process. That doesn’t mean research is unnecessary during industrialization—only that pure research will struggle to produce valuable breakthroughs.

Zhang Xiaojun: What did you learn from these renowned mentors?

Yang Zhilin: I learned the most at Google, where I interned for a long time. I began working on Transformer-based language models in late 2018. My biggest learning was freeing myself from endless “carving”—the obsessive refinement of surface details. That was crucial.

You should look at what the big direction is, the big gradient. When ten roads lie before you, the average person worries about how to brake for a pedestrian ahead on this one road—short-term details. But which of the ten roads to take is what matters most.

This field previously had exactly that problem. For example, on a dataset of only one or two million tokens, you’d look at how to push perplexity lower, how to push loss lower, how to improve accuracy—and you’d fall into endless carving. People invented many bizarre architectures; these were carving tricks. After carving, you might do better on that kind of dataset, but you miss the essence of the problem.

The essence is analyzing what the field is missing. What is the first principle? Why can the scaling law serve as a first principle? You only need to find a structure that satisfies two conditions: first, it is sufficiently general; second, it is scalable. General means you can model all problems within this framework; scalable means that as long as you pour in enough compute, it keeps getting better.

This is the thinking I learned at Google: if something can be explained by something more fundamental, you shouldn’t over-carve at the upper layers. There’s an important line I strongly agree with: if you can solve a problem with scale, don’t solve it with a new algorithm. The greatest value of a new algorithm is in how it lets you scale better. When you free yourself from carving, you can see much more.

Zhang Xiaojun: Was Google also a follower of the scaling law back then? How did it implement first-principles thinking?

Yang Zhilin: Many such ideas already existed there, but Google didn’t implement them especially well. It had this way of thinking, but it couldn’t organize itself into a true moonshot. It was more like: here are five people pursuing my first principles, and over there five people pursuing theirs. There was nothing top-down.

Zhang Xiaojun: During your PhD, you published papers in collaboration with Turing Award winners Yann LeCun and Yoshua Bengio—and you were first author on those papers. How did those collaborations come about? What I mean is: they’re Turing Award laureates and they weren’t your advisors—what did you rely on to attract them?

Yang Zhilin: Academia is very open. As long as you have a good idea and a meaningful problem, it’s fine. What two brains—or n brains—produce is more than one brain alone. This applies when developing AGI, too. An important strategy in AI is called “ensembling”—using multiple different models or methods and combining their predictions for better performance. It’s essentially doing the same thing: when you have diverse viewpoints, you can spark many new things. Collaboration is hugely beneficial.

Zhang Xiaojun: Would you first have an idea and then ask them whether they were interested?

Yang Zhilin: That’s roughly how it went.

Zhang Xiaojun: Which is harder: winning over academic heavyweights in research, or winning over capital heavyweights in fundraising? What are the similarities?

Yang Zhilin: “Winning over” isn’t a good phrase—the essence behind it is cooperation. Cooperation means both sides win, because mutual benefit is the precondition for cooperation. So there’s really no difference: you need to offer others unique value.

Zhang Xiaojun: How do you earn their trust? What do you think your gift is?

Yang Zhilin: There’s no particular gift—just working hard.

Part 3

The Old System No Longer Works

“AGI Needs a New Kind of Organization”

**

Zhang Xiaojun: You just said “more valuable breakthroughs will happen in industry”—does that include startups and the giants’ AI labs?

Yang Zhilin: Labs are history. Google Brain used to be the biggest AI lab in industry, but it was a research organization embedded inside a big company. That kind of organization can explore new ideas, but it’s very hard for it to produce a great system—it could produce the Transformer, but it couldn't produce ChatGPT.

The way development now evolves is that you’re building an enormous system, which requires new algorithms, solid engineering, and even a lot of product and commercialization work. It’s like the early 2000s: you couldn’t research information retrieval in a lab; it had to live in the real world, as a huge system, a product with users—like Google. So research and education systems will shift their function toward primarily cultivating talent.

Zhang Xiaojun: How would you describe this new form of system? Is OpenAI its prototype?

Yang Zhilin: It’s the most mature organization of this kind today, and it’s still gradually evolving.

Zhang Xiaojun: So it can be understood as an organization established for humanity’s grand scientific goals?

Yang Zhilin: I want to emphasize: it is not pure science—it’s a combination of science, engineering, and business. It has to be a commercial organization, a company, not a research institute. But this company is built from zero to one, because AGI needs a new kind of organization. First, the mode of production differs from the internet era; second, it shifts from pure research to a combination of research, engineering, product, and business.

At its core, it should be a moonshot program, with a great deal of top-down planning—yet within that planning there is room for innovation, because not all the technology is predetermined. Bottom-up elements exist within a top-down framework. Such an organization didn’t exist before, but the organization must adapt to the technology, because technology determines the mode of production; if they don’t match, you can’t produce effectively. We believe it will very likely need to be redesigned from scratch.

Zhang Xiaojun: During last year’s OpenAI boardroom coup, one option for Sam Altman was to join Microsoft and lead a new Microsoft AI team. What is the essential difference between that and being CEO of OpenAI?

Yang Zhilin: You’d have to grow a new organization inside an old culture—and that is extremely difficult.

Zhang Xiaojun: You want to build “China’s OpenAI”—can we put it that way?

Yang Zhilin: Not quite accurate. We don’t want to be China’s anything, and we don’t necessarily want to be OpenAI.

First, real AGI will definitely be global. There is no such thing—at least not long-term—as an AGI company confined to some regional market because of market-protection mechanisms. Globalization, AGI, and having a product with a very large user base: these three are ultimately necessary conditions.

Second, should it be OpenAI? If you look at 2017–2018, OpenAI had a terrible reputation. When people in our circle looked for jobs, they generally considered places like Google. Many people who talked with Ilya Sutskever, OpenAI’s chief scientist, came away thinking the man was crazy and far too full of himself—OpenAI was either madmen or scammers. But they committed very early, found the non-consensus, and found what is now the only first principle that works in AI: scaling through next-token prediction.

I believe there will be a company greater than OpenAI. A truly great company can combine technological idealism with a great product, co-creating with its users—AGI will ultimately be something produced by co-working with all of its users. So it’s not only about technology; it also requires pragmatism and real-world pursuits—ultimately, a perfect combination of the two.

Still, we should learn from OpenAI’s technological idealism. If everyone thinks you’re normal—if your dream is one that anyone could have—it adds nothing to the sum total of humanity’s dreams.

Part 4

The Moonshot’s First Step Is “Long Context”—What’s the Second?

“Two Big Milestones Are Coming Next”

**

Zhang Xiaojun: Back to the moment you decided to start the company—did you launch the first funding round immediately after returning to China?

Yang Zhilin: It began in the U.S. in February (last year), some of it remotely. In the end, domestic investors made up the majority.

Zhang Xiaojun: Did the first round raise $100 million?

Yang Zhilin: The first round wasn’t that much; later rounds exceeded that figure. We completed two rounds in 2023, totaling nearly RMB 2 billion.

This is now the third round. We haven’t formally announced the financing, so I can’t comment at this time.

Zhang Xiaojun: Some people say that since the second half of 2023, no one has been willing to invest in foundation-model companies anymore. Are they wrong?

Yang Zhilin: There still are. You can indeed see the shift in sentiment, but it’s not that no one is investing—at least for now, there’s quite a lot of investment interest in the market.

Zhang Xiaojun: Besides capital and people, what other key decisions did you make in 2023?

Yang Zhilin: Deciding what to do. That’s the advantage of companies like ours—having a technical vision for decisions at the highest level.

We do long context. That requires judgment about the future: you need to know what is fundamental and where things are heading next. Again, it’s first principles—the process of “de-carving.” If you focus on carving, you can only look at what OpenAI has already done and figure out how to reproduce it.

You’ll find that doing lossless long-text compression in Kimi gives the product a unique experience. When you read English-language papers, it helps you understand them remarkably well. Using Claude or GPT-4 today, you won’t necessarily do as well; this required laying the groundwork in advance. We worked on it for over half a year. That’s very different from spotting a long-context trend today, hastily assembling two teams, and developing it at maximum speed.

Of course, the marathon has only just begun; more differentiation will come, and that requires you to anticipate in advance what counts as “a non-consensus that holds true.”

Zhang Xiaojun: In what month was this decision made?

Yang Zhilin: February or March—it was decided as soon as the company was founded.

Zhang Xiaojun: Why is long context the first step of the moonshot?

Yang Zhilin: Because it’s fundamental. It is the new computer’s memory.

The old computer’s memory grew by several orders of magnitude over the past few decades, and the same thing will happen with the new computer. It can solve many of today’s problems. For example, current multimodal architectures still need a tokenizer, but with a losslessly compressed long context, you don’t need one—you can put the raw input in directly. Taken further, it’s the foundation for making the new computing paradigm more general.

The old computer could represent everything with 0s and 1s; everything could be digitized. But today’s new computer can’t yet—there isn’t enough context, so it isn't that general. To become a general world model, you need long context.

Second, it enables personalization. AI’s core value is personalized interaction; the value ultimately lands on personalization, and AGI will be more personalized than the previous generation of recommendation engines.

But personalization isn’t achieved through fine-tuning—it’s achieved by supporting very long context. Your entire history with the machine is context, and that context defines the personalization process. It cannot be replicated, and it makes for more direct dialogue—dialogue that generates information.

Zhang Xiaojun: How much room is there to scale this up?

Yang Zhilin: Enormous. On one hand, expanding the window itself still has a long way to go—several orders of magnitude.

On the other hand, you can’t only expand the window, and you can’t just look at the number; whether the window is a few million tokens or several billion today is meaningless in itself. You have to look at the reasoning ability it enables within that window, the faithfulness—fidelity to the original information—and the instruction-following ability. You shouldn’t chase a single metric; you have to combine metrics with capabilities.

If these two dimensions keep improving, you can do a great deal. It could follow an instruction tens of thousands of words long, and the instruction itself could define many agents—highly personalized.

Zhang Xiaojun: Are the technologies behind long context and catching up with GPT-4 reusable for each other? Are they the same thing?

Yang Zhilin: I don’t think so. It’s more about adding a new dimension—a dimension GPT-4 doesn’t have.

Zhang Xiaojun: Many people say the leading Chinese foundation-model companies are all doing roughly the same thing—chasing GPT-3.5 in 2023, chasing GPT-4 in 2024. Do you agree?

Yang Zhilin: Improving general capabilities certainly has key milestones, so that statement is right to a degree—as a latecomer, you inevitably go through a catching-up process. But it’s also one-sided. Beyond general capabilities, there is a lot of space to develop distinctive capabilities and reach state-of-the-art in certain directions. Long context is one. DALL-E 3’s image generation is thoroughly outclassed by Midjourney V6. So you have to work on both fronts.

Zhang Xiaojun: What proportion of time and resources goes to general capabilities versus new dimensions?

Yang Zhilin: They have to be combined. A new dimension can’t exist apart from general capabilities, so it’s hard to give a direct ratio. But sufficient investment is required to do the new dimension well.

Zhang Xiaojun: Will all these new dimensions be carried by Kimi?

Yang Zhilin: Kimi is certainly a very important product for us, and we’ll have some other attempts as well.

Zhang Xiaojun: What do you make of the comment by Li Guangmi, founder of Shixiang, that the technological distinctiveness of Chinese foundation-model companies is still not very high today?

Yang Zhilin: I think it’s fine—we’ve already produced quite a lot of differentiation today. It’s a matter of time; this year you should see more dimensions. Last year, everyone was just putting up the scaffolding and getting things running first.

Zhang Xiaojun: If the moonshot’s first step is long context, what’s the second?

Yang Zhilin: There will be two big milestones ahead. First, a truly unified world model—one that unifies all the different modalities, a truly scalable and general architecture.

Second, enabling AI to keep evolving without human data input.

Zhang Xiaojun: How long will it take to reach these two milestones?

Yang Zhilin: Two to three years—possibly faster.

Zhang Xiaojun: So three years from now, we’ll already be looking at a world completely different from today’s.

Yang Zhilin: At the current pace of development, yes. The technology is now in a budding, fast-growing stage.

Zhang Xiaojun: Can you imagine what will exist three years from now?

Yang Zhilin: There will be a certain degree of AGI. Many of the things we do today, AI will also be able to do—even better than us. But the key is how we use it.

Zhang Xiaojun: And for you—for Moonshot AI as a company—what’s the second step?

Yang Zhilin: We will go do those two things. All the remaining problems are derived from these two factors. The reasoning and agents people talk about today are byproducts of solving these two problems. Some additional carving is needed, but there’s no fundamental blocker.

Zhang Xiaojun: Will you go all in on catching up with GPT-4?

Yang Zhilin: GPT-4 is a necessary stop on the road to AGI. The key is not to be satisfied with merely matching GPT-4. First, you have to ask what the real non-consensus is now: beyond GPT-4, what’s next? What should GPT-5 and GPT-6 look like? Second, you have to see which distinctive capabilities you have within that—and that matters more.

Zhang Xiaojun: Other foundation-model companies publish their

Переделать в YouMind

Превратите одну вирусную статью в полноценный рабочий процесс создания контента

Собирайте источники, расшифровывайте паттерны, создавайте активы, пишите черновики и публикуйте контент из одного рабочего пространства ИИ.

Исследовать YouMind
Для авторов

Превратите ваш Markdown в аккуратную статью для 𝕏

Когда вы публикуете длинные тексты, изображения, таблицы и блоки кода, форматирование в 𝕏 становится мучением. YouMind превращает полный черновик в Markdown в чистую статью, готовую к публикации в 𝕏.

Попробовать Markdown для 𝕏

Другие паттерны для анализа

Недавние виральные статьи

Смотреть другие виральные статьи