Fireside Weekly: The Eight Conversations That Mattered This Week

@firesidealpha
АНГЛІЙСЬКА1 день тому · 25 лип. 2026 р.
148K
32
3
0
71

Коротко

This weekly digest synthesizes eight major tech discussions, covering NVIDIA's growth projections, the shift toward custom AI silicon, and how agents like Devin are transforming software engineering.

For the time-strapped builders and investors who want to learn from the best tech and business conversations without giving up the hours it takes to follow them all.

The eight best episodes published this week, recapped in one place, with each full piece a click away. Make sure to subscribe at Fireside Alpha to receive daily breakdowns and analysis on new episodes.

This week’s biggest:

  • Jensen Huang says the chip industry still has to get five to ten times bigger before the AI build-out is done, and dates the bubble to “someday, just not today.” ▶ Watch · 00:32:01
  • Charlie Kawwas (President, Broadcom) has one customer scaling from 1.5 GW of AI silicon this year to a contracted 6.5 GW next year to more than 10 the year after. ▶ Watch · 00:16:40
  • Sassine Ghazi (CEO, Synopsys) grades AI chip design at L4 today, the front-end already running on six to eight agents, after two years of being wrong that it was five to six years away. ▶ Watch · 00:12:56
  • Andy Hock (Chief Strategy Officer, Cerebras) says OpenAI agreed to buy 750 megawatts of Cerebras compute, and the faster Codex option in its coding product runs on Cerebras underneath. ▶ Watch · 00:49:58
  • Pat Gelsinger and the HBM panel including SK hynix note the memory prints close to 90% margins for a business that used to get “one good year out of four,” and still call it a bad memory. ▶ Watch · 0:03:11
  • The SemiAnalysis desk makes the case that Anthropic is already profitable, with more than 80% of ARR on the high-margin API side and Q3 operating profit possibly clearing $1 billion.
  • Scott Wu shows 90%+ of Cognition’s code is now committed by Devin, and the agent already starts the majority of its own sessions. ▶ Watch · 5:31
  • Mansour Karam of Aria Networks says the same million tokens can cost $0.20 or $20, a 100x spread, and the layer that decides is the network. ▶ Watch · 0:03:38

Jensen Huang · Axios

Fireside Alpha - inline image

Full Article: Jensen Huang on the AI Doomers: The Bubble Comes Someday, Just Not in the Next Five Years

Jensen Huang, co-founder and CEO of NVIDIA, on Axios “Behind the Curtain” with Mike Allen from a new Fort Worth chip plant, on why the AI bubble is real but years off and what has to get built first.

The chip industry has to get five to ten times bigger over the next ten years. Huang’s estimate rests on shortages across memory, storage, optical interconnects, packaging, and TSMC capacity all at once.

“I believe we probably need our semiconductor industry to be somewhere between five to 10 times larger than it is.”

Watch · 00:32:01

The bubble comes “someday,” not in the next five years. Allen pressed him on the oldest question in semiconductors, whether a boom this big is due for a bust.

“The bubble will come someday. It’s just not today. You know, this is in the very beginning of the build-out.”

Watch · 00:36:21

NVIDIA’s China sales are “approximately zero today.” Huang told investors to expect none while the Chinese market stays closed.

“Our China sales is approximately zero today. And I’ve told all of our investors not to expect any China sales.”

Watch · 00:06:07

Manufacturing jobs are up some 50% in the last several years. Huang ties the rise to the physical build of the AI data centers that generate tokens.

“The number of manufacturing jobs has increased some 50% in the last several years because these computers have to go into these AI data centers, which generate the tokens, the intelligence.”

Watch · 00:17:27

Someone who uses AI takes your job, not AI itself. Huang’s anti-doom line splits automating a task from eliminating a whole job.

“AI is not going to destroy all of our jobs. Someone who uses AI is going to take our jobs.”

Watch · 00:23:41

A hundred billion to a trillion agents running all the time. Against roughly 100 million people using a computer at any given moment today.

“We’re going to have a hundred billion, a trillion agents that are running all of the time. Smart agents, lesser smart agents, specialized agents, super agents, all kinds of different types of agents running all the time.”

Watch · 01:04:54

Charlie Kawwas · RAISE Summit

Fireside Alpha - inline image

Full Article: Broadcom (Charlie Kawwas): How Custom Silicon Went From Zero to $100bn

Charlie Kawwas, president of Broadcom’s semiconductor solutions group, on the RAISE Summit stage, walking BlackRock’s Tony Kim through how custom silicon went from zero to $100bn and why the biggest AI labs are leaving Nvidia’s general-purpose rack.

Nvidia’s general-purpose rack carries two taxes, an efficiency tax and a margin tax. Kawwas is laying out the cost structure a frontier lab inherits when it buys Nvidia’s pre-built compute platform.

“This general compute comes with a tax. The tax is first the efficiency of that compute platform, meaning it’s built for everybody. If you have a specific workload in your AI frontier model, guess what, you’re paying that extra tax. And then the other tax is the high margins that come with it.”

Watch · 00:01:56

Each chip generation triples memory and runs almost 3x the compute of the one before. Kawwas said this while handing the host actual chips from consecutive years, two memory cubes against six.

“Last year to this year, you are now getting three times more memory. There you have two cubes of memory. Here you have six cubes of memory. And the compute that’s in the middle is almost 3x stronger.”

Watch · 00:04:09

Four of the big five AI labs are building custom platforms with Broadcom. The five are OpenAI, Anthropic, Google, Meta and xAI, roughly 90% of world AI spend by Kawwas’s count.

“The big five in the US, which today are spending maybe 90% of the spend in the world, are OpenAI, Anthropic, Google with Gemini, Meta and xAI... Amongst these five is literally 80, 90 plus percent of the market. And four out of the five have realized the Nvidia technology is so good, but they can’t be buying this general compute platform that’s custom at the rack level.”

Watch · 00:07:56

Co-design runs at two levels, the XPU and the whole rack. Kawwas is answering why a lab can’t buy this custom silicon off a shelf, and where Broadcom pushes for an open rack.

“There’s two portions to that haute couture. One is exactly the deep co-design for the XPU you described. The second is at the rack level. So not only we have to do it at the XPU level for their LLMs and their specific use cases, but we actually engage heavily into how do we make sure that the entire rack is open.”

Watch · 00:12:18

One customer goes from 1.5 GW this year to a contracted 6.5 GW next year to over 10 the year after. Asked where AI infrastructure goes next, Kawwas answered in power for a single large customer, not chip counts.

“For one of the large five, this year we’re putting in about one and a half gigawatts with one of the chips I showed you. Next year, we’re going to probably put in over five. We already contracted five... The year after that, I’m pretty sure it’ll be over 10.”

Watch · 00:16:40

Sassine Ghazi · RAISE Summit

Fireside Alpha - inline image

Full Article: Synopsys (Sassine Ghazi): Chip Design Has Reached "L4," and Wafer Capacity Is the Ceiling

Sassine Ghazi, CEO of Synopsys, on the RAISE Summit stage, on how fast AI is now designing the chips themselves and where the buildout runs out of wafers.

Every chip on earth runs through Synopsys before it exists, per its CEO. Synopsys sells the electronic-design-automation tools chip engineers use to lay out silicon, now extended into physics by the Ansys acquisition.

“There is no chip today that is designed without Synopsys, so we are essential and pervasive in terms of our use at the silicon level.”

Watch · 00:01:08

The move from training to inference forces customization up the whole stack. Ghazi’s account of why one general-purpose AI chip is fragmenting into GPUs, Broadcom-style ASICs, and hyperscaler customer-owned tooling.

“In order to get to that inference efficiency, you need to look at the whole stack... And that’s where you need both customization of the silicon, looking all the way up along the stack, in order to achieve the right efficiency, performance, cost.”

Watch · 00:03:38

Data-center margins are cost, and hyperscalers are co-designing to cut them. The Ansys acquisition, one year old next week, is what lets Synopsys model thermal, fluid, and structural physics alongside the electronics.

“When you’re designing a data center, typically you’re designing with a lot of margins, and margins equal cost... We’re seeing many of the hyperscalers and advanced customers trying to reduce this margin by co-designing across multiple engineering domains.”

Watch · 00:07:40

Ghazi called autonomous chip design five to six years out, and has been wrong two years running. Synopsys grades design autonomy L1 to L5 like self-driving and puts the industry at L4 today, with the front-end already on six to eight agents.

“If you were to ask me that question two years ago, I’d say this is about five, six years out, and I’ve been wrong for the last two years because it’s coming so fast.”

Watch · 00:12:56

Wafer capacity is the ceiling in both logic and memory, and Tesla sees a silicon shortfall by 2030. The binding constraint on the buildout is physical fab output, not demand, which Ghazi expects to keep pushing against it.

“The current constraints are limited by wafer capacity, and that’s a logic wafer capacity, as well as memory... When you hear companies like Tesla say that by 2030 they will not have, given their ambition from EV to robotics, there will not be enough capacity to support the silicon demand they envision.”

Watch · 00:14:45

Andy Hock · Cerebras

Fireside Alpha - inline image

Full Article: How Cerebras Plans to Kill Nvidia (Andy Hock, Chief Strategy Officer)

Andy Hock, chief strategy officer at Cerebras, on why the industry keeps buying a chip built for graphics, and how one wafer-size processor left whole ended up powering OpenAI’s 750-megawatt compute deal.

Cerebras builds one chip and leaves the wafer whole instead of slicing it into hundreds. The industry prints a wafer, cuts it into small chips, and wires them back together across boards and data centers.

“What if we just put as much compute as we could together on a chip, on a big chip, and then allowed the compute elements to talk to each other over silicon, so we wouldn’t have to wire a bunch of small chips together. [...] What if we just didn’t do that last part? What if we built one big chip and then left it whole?”

Watch · 00:15:13

The wafer scale engine holds nearly a million cores, 900,000 in production machines. Each core carries its own bank of fast on-chip memory, all connected over silicon rather than copper and fiber.

“This is the wafer scale engine. [...] It has nearly one million cores all on one device, and 900,000 in our production machines. They’re in a 2D grid across the entire surface of the wafer, and all those cores are directly connected to each other. [...] This design gives us what you might think of as a cluster worth of AI compute on a single device, and literally three orders of magnitude more communication bandwidth and memory bandwidth than is possible with traditional chip designs.”

Watch · 00:15:45

Training is a cost center, inference is where you ring the cash register. Hock divides the AI economy into building the model and running it for end users.

“Training is a cost center. That’s where you build your model, that’s where you teach it how to do what you want it to do. Inference is a value center. Once the model does the thing you need it to do, or the thing that’s valuable for you or your end users, inference is where you ring the cash register.”

Watch · 00:21:45

Cerebras serves Gemma 4 at 1,500 tokens per second, 10x any GPU implementation. The company put the Google model into preview on the day of the interview and quoted its speed against the field.

“We’re serving Gemma 4 at 1,500 tokens per second. It’s 10 times faster than any GPU implementation and 15 times faster than something like Claude’s Haiku model. [...] Fast coding agents, where speed equals software engineer productivity.”

Watch · 00:23:16

OpenAI agreed to buy 750 megawatts of Cerebras compute, and Codex Spark runs on it. The faster, higher-priced coding option in OpenAI’s product is Cerebras underneath.

“The public terms of the deal are they’re going to buy 750 megawatts of compute by Cerebras to power their applications. The coding model that uses Cerebras now is Codex Spark, and we’re working with the OpenAI team to bring up additional models in the future, from their GPT frontier models to next generation coding models.”

Watch · 00:49:58

Cerebras keeps all its memory on the chip, so the HBM shortage misses it. The high-bandwidth memory GPUs pull weights from is fab-limited and can’t scale on command.

“That memory is subject to fabrication limits, manufacturing limits. [...] You can’t just turn a knob and produce 10x more. And so that memory is a bottleneck right now for those types of chips. For us, it’s not. We have all of our memory on the chip. It’s right here. It’s not part of that supply chain.”

Watch · 00:30:41

HBM Panel · RAISE Summit

Fireside Alpha - inline image

Full Article: HBM Prints 90% Margins and Everyone on Stage Still Calls It a Bad Memory: Inside AI’s Memory Bottleneck

At RAISE Summit in Paris, Pat Gelsinger, SK hynix Fellow Hoshik Kim, and investors Andrew Homan and Vijay Shilpiekandula worked through why HBM prints roughly 90% margins while everyone on stage still calls it a bad memory.

HBM earns close to 90% margins for a business that used to get one good year in four. Gelsinger ran Intel through multiple memory cycles and now invests in the space at Playground Global.

“Memory has never had it this good. We used to joke in the semiconductor industry that if you were a memory company, you had one good year out of four. The other three sucked.” Pat Gelsinger, Playground Global

Watch · 0:03:11

The former Intel CEO calls HBM “a lousy memory” with the SK hynix Fellow one seat away. Gelsinger’s objection is physical, not commercial: stacking DRAM builds a thermal sandwich and pins bandwidth to the shoreline where memory meets the GPU.

“HBM is a lousy memory. I know my memory friends are about to come on stage, and calling their baby ugly is not maybe the best way to make friends and influence people.” Pat Gelsinger, Playground Global

Watch · 0:04:35

SK hynix’s own Fellow says HBM will not solve the memory wall. Kim runs system-architecture research at SK hynix, the HBM market leader.

“It’s a good memory, but it’s not the final answer. HBM will not solve the memory wall problem. Memory wall problem is an inherent AI problem, which you cannot avoid.” Hoshik Kim, SK hynix

Watch · 0:12:49

HBM drops a fab’s throughput from 12 waffles an hour to four. Homan reached for the kitchen to explain why manufacturing throughput keeps the shortage from clearing fast.

“SK hynix is a waffle maker. So you have a waffle machine that usually can cook a waffle in five minutes. So in that scenario, you can make 12 waffles per hour. With HBM, it now takes 15 minutes to make that waffle. So now you can only make four waffles per hour.” Andrew Homan, Maverick Silicon

Watch · 0:19:41

Apple raised prices across its line, and it is the most sophisticated memory buyer on the planet. HBM sits at roughly a 7% supply deficit in 2026, and Homan pointed to Cupertino for the clearest read on how long the tightness lasts.

“Apple raised prices across its product line. They’re the most sophisticated memory procurement machines on the planet, and clearly they expect this to last for a while, otherwise they would not have done that.” Andrew Homan, Maverick Silicon

Watch · 0:27:50

Several more years of memory shortage still ahead, by the ex-Intel CEO’s read. Gelsinger set his timeline against capacity that was already locked by capital decisions made three years ago.

“The years of plenty are many years away. The years of memory shortages, we have several more years in front of us.” Pat Gelsinger, Playground Global

Watch · 0:29:26

SemiAnalysis Tokenomics · Ep020

Fireside Alpha - inline image

Full Article: SemiAnalysis Tokenomics: Anthropic Is Already Profitable ($1B+ in Q3), and OpenAI Just Clawed Back to Even

The SemiAnalysis Tokenomics desk (Crystal Huang, Max Kan, Joey Brookhart, host Jordan Nanos) spends the hour on model-lab margins: why Anthropic is already profitable, how subscriptions subsidize the API, and revenue per megawatt doubling every three quarters.

Anthropic’s API is 80%-plus of ARR, and Q3 operating profit could top $1 billion. Joey Brookhart puts the profit in the P&L on a non-GAAP, ex-stock-comp basis.

“Operating profit was profitable in Q2, could be profitable to the tune of 1 billion plus in Q3.” Joey Brookhart

85%-plus API gross margin for Anthropic, and the 20x subscription plans may run negative. Max Kan on where the subscription-versus-API margin actually lands.

“Forget being worse margin than API, which we think is likely 85% plus for Anthropic at least. It might just be negative margin in general for these 20x plans. And so of course these businesses want to move as much volume as possible to API.” Max Kan

The 99th-percentile company spends $100,000 per employee a year on AI. Joey Brookhart reads Ramp’s card data as the ROI check on heavy coding spend.

“The Ramp data is like the 99th percentile of companies, they spend a hundred thousand per year on AI per employee. And if you’ve gotten to that point where you’re spending that much, you’re obviously getting pretty good ROI.” Joey Brookhart

Enterprise plans block subscriptions and force per-token billing, where the margin is. Max Kan on why the labs meter their heaviest users onto the API.

“Anybody who can use a subscription definitely should use a subscription just because it’s so subsidized. OpenAI and Anthropic are very wise to this. When they have their enterprise plans, they don’t let you use a subscription. They force you to pay per token because they know that’s where the margin is.” Max Kan

The desk’s read: Elon rents most of Colossus to Anthropic at 3x to 4x market rate, with a 90-day clawback clause. Max Kan narrates the rent-the-frontier logic from the compute owner’s seat.

“I’m going to leave you guys with enough compute to prove that you’re capable of reaching the frontier. In the meantime, I’m going to monetize all the additional compute by renting it out on these extremely high margin 3x, maybe 4x market rate rental deals. I’m going to make sure there’s this 90-day clawback clause included in every single one of them.” Max Kan

Frontier-lab data budgets cross $10 billion this year, roughly 10x last year. Max Kan on what actually gates model capability, where the bottleneck is RL environments rather than raw compute.

“Reinforcement learning is probably the most important scaling law for improving model capabilities today. There are a lot of people who believe that the only thing stopping the models from being able to do literally anything a human can do on a computer is having sufficient RL environments.” Max Kan

Scott Wu · RAISE Summit

Fireside Alpha - inline image

Full Article (link below): Cognition (Scott Wu): The Golden Age of Software Engineering

Scott Wu, CEO of Cognition and maker of the Devin AI software engineer, on the RAISE Summit stage, arguing the productivity jump inside his own company is a reason to grow the engineering team, not shrink it.

90%+ of Cognition’s code is now committed by Devin. Wu put up his internal chart and noted the agent has started triggering its own sessions off performance regressions and DataDog errors, no human in the loop.

“You see the percent of code committed by Devin going up to 90-plus percent. We basically use Devin for pretty much everything right now. The thing that’s even crazier is Devin actually also does the majority of starting Devins.”

Watch · 5:31

Cognition shipped 10x more software in six months on ~40% more engineers. Wu ran November to end of April as an experiment on his own company, growing the team from roughly 45 to 60, and credits models and agents rather than headcount.

“What changed exactly over the course of six months that allowed us to ship 10 times more and do 10 times more?”

Watch · 2:36

GPT-3.5 could run 30 seconds of unassisted work, the latest models run 16-hour jobs. Wu pointed to the METR study, which measures how long a task an AI can complete before it has to come back to a human.

“You take a look at GPT-3.5, right in the middle of that chart, 30 seconds of work, which is not bad, by the way. And that was the ChatGPT moment.”

Watch · 3:09

Back in December 2025, Devin was an extra 130 engineers of bandwidth per thousand. Wu framed the added agent capacity as headcount you could hand an existing team, before noting it has since climbed toward 89 to 90% of the work.

“Think about a team of a thousand software engineers and imagine you could tell them, okay, you are getting an extra 130 software engineers worth of bandwidth to go and build things. You can do a lot with that.”

Watch · 8:55

Every engineer is 10x more productive, and there’s 10x more to build. Wu’s answer to the layoffs fear, backed by deliberately dull examples of unbuilt software such as a DMV that works and pullable medical records.

“Every human software engineer is ten times more productive, but guess what? We have ten times more things to go build.”

Watch · 10:57

Wu bookended the talk with one question: what you would do with an infinite army of software engineers. His customers include Goldman Sachs, Mercedes-Benz, and Lowe’s, each running tens of thousands of engineers against a backlog that never clears.

“What would you do with an infinite army of software engineers? I would ask all of you guys to spend a second and just think about that.”

Watch · 1:26

Mansour Karam · RAISE Summit

Fireside Alpha - inline image

Mansour Karam, founder and CEO of Aria Networks, on the RAISE Summit stage with SemiAnalysis's Dylan Patel, on why the "one general-purpose cluster for everything" era is ending and why the network decides inference economics.

Training is a cost center, serving looks more like renting an apartment complex. Karam splits an AI cluster into two businesses that share the hardware but almost none of the economics.

"Training is essentially building a model. That's like building a product. But then serving the model, the economics look a lot more like when you're renting an apartment complex in terms of the financing."

Watch · 0:01:34

Specializing a cluster for one workload is worth a factor of 10 or 100. Karam sees room for new companies to form around single use cases, each tuned to one metric.

"There is a lot of opportunity that I see around new companies being formed to optimize for a given use case, and the difference can be a factor of 10 or 100, depending on what metric you're trying to optimize for."

Watch · 0:02:17

A million tokens can cost $0.20 or $20, a 100x spread set by how you build. Karam ties the swing to whether the infrastructure was built for cheap high-throughput inference or something else.

"You can go anywhere from, what, 20 cents to like $20 or something like that for a million tokens. I mean, this is like a 100x difference. If you really want to optimize for high-throughput simple inference, you can serve it really cheap, but you have to build your infrastructure with that in mind."

Watch · 0:03:38

Dylan Patel's line: "speed is the moat." Patel supplies the demand-side reason the fast corner is suddenly where the money goes: agents run in sequential try-verify loops.

"Agents are very sequential, task-oriented. Trying something, verifying it, checking if it's right, and you do that in a loop. It ends up being very slow. And speed is the moat. The faster you go, speed is the moat." Dylan Patel, SemiAnalysis

Watch · 0:07:44

One transceiver with a grain of dust stalls every user on the cluster. A single inference query crosses several networks in sequence, and the slowest link sets the pace for everyone.

"If you have packet loss, or you have one transceiver that has a grain of dust... then everybody else has to wait, and it's all sequential."

Watch · 0:11:00

One second inside an inference cluster "is like a century at human scale." Karam's pitch for Aria is microsecond-resolution telemetry to catch the stalls standard one-second monitoring misses.

"One second, in the context of an AI cluster running inference, is like a century at human scale. A lot happens in a second. And so what you really need is to collect this telemetry at microsecond resolution, so you can capture those events that could be causing a stall."

Watch · 0:14:52

Збереження в один клік

Використовуйте YouMind для AI-глибокого читання віральних статей

Зберігайте джерела, ставте цілеспрямовані запитання, підсумовуйте аргументи та перетворюйте віральні статті на корисні нотатки в одному AI-робочому просторі.

Дослідити YouMind
Для авторів

Перетворіть свій Markdown на охайну статтю для 𝕏

Коли ви публікуєте власні лонгріди, зображення, таблиці та блоки коду роблять форматування в 𝕏 складним. YouMind перетворює повну чернетку в Markdown на чисту статтю для 𝕏, готову до публікації.

Спробувати Markdown для 𝕏

Більше патернів для аналізу

Останні віральні статті

Переглянути більше віральних статей