Read best config against best config, SemiAnalysis's own dashboard gives NVIDIA 9.7x outperformance on the TPU it just crowned.
On September 7 SemiAnalysis published a headline that says Google's Ironwood TPU beats NVIDIA's Blackwell by up to 50% on performance per dollar. SemiAnalysis's own dashboard, with each side running the way its maker actually ships it, says NVIDIA wins by 8 to 12x. Same shop, same benchmark, one day apart. We pulled the dashboard's own numbers, and this piece is what they say.
The article is "TPU Inference Externalization Full Steam Ahead," and the InferenceX site carries a free copy. Its section title is "TPU is King on Performance per Dollar." The comparison under it puts Ironwood, which the dashboard labels TPU7x, against a single box of eight NVIDIA chips, running the slower of NVIDIA's two number formats, writing one token at a time, on one pool of chips that reads and writes together instead of the split setup the rack runs. SemiAnalysis calls that apples to apples. It calls NVIDIA's full 72-chip rack against the TPU "apples to bananas," and keeps the rack out of the headline.
https://x.com/SemiAnalysis_/status/2097054644258103448
That is the part worth sitting with. The dashboard is one click from the headline. The data that contradicts the headline is SemiAnalysis's own, run by SemiAnalysis's own methodology, priced on SemiAnalysis's own cost model, and in the default table the rack's curve is sorted to the top. A benchmark shop is the referee in this fight. This time the referee published the halftime score as the final.
The question we kept asking this week: what would you do to find that common denominator for both, to make it a fair comparison?
The answer a buyer uses is simple to say. Tokens are the word-pieces a model reads and writes. How many of them do you get for a dollar, at the speed a user sees the answer come back, with each machine set up the way its maker recommends. The bill starts in that number. Read that way, on the same model and the same workload the article uses, SemiAnalysis's dashboard has GB300 NVL72 delivering 9.7x the tokens per dollar of TPU7x when each user gets 100 tokens a second, and 11.9x at the fastest speed the TPU reaches. The discount is on the chip. The bill is per token.
What the Headline Switched Off
Four things are off on the NVIDIA side of the "apples-to-apples" chart, and all four are visible in SemiAnalysis's own text and in the dashboard's configuration labels. The test runs NVIDIA in eight-bit math, FP8, instead of the four-bit math, NVFP4, that Blackwell was built around, which does twice the arithmetic per second on Blackwell and three times on Blackwell Ultra, the chip inside GB300; SemiAnalysis notes Ironwood cannot run four-bit natively and the next TPU will. It turns off multi-token prediction, the trick where the chip drafts several tokens ahead and checks them in one pass, the way a good autocomplete does. It uses a box of eight chips instead of the 72-chip rack NVIDIA sells as one computer. And it serves the model aggregated, reading the question and writing the answer on the same chips, instead of disaggregated, which splits those two jobs across different chips the way a kitchen splits prep cooks from the line. That split is a serving choice, not a software brand: vLLM and SGLang can run it too, and the article's setup does not.
Each switch has a price on the same dashboard. The multi-token trick alone is worth about 2.3x on the rack when each user gets 100 tokens a second. Going from a single box serving aggregated to the rack serving disaggregated, same number format and the same engine, is worth about 2.7x. The engine itself, TensorRT-LLM against SGLang on the same rack with the same split, about 3x. Four-bit over eight-bit on the same box, about 1.9x. The ratios do not multiply cleanly, because the curves were run on different dates, and the table says so. Together they are where the 8 to 12x comes from. None of this is hidden. NVIDIA has said in public for two years that the rack, the number format and the serving software are the product, SemiAnalysis's own methodology says which of them it switched off, and if you turn a lot of that stuff off it doesn't work out.
Then there is what the benchmark leaves out. It runs one model, Qwen3.5 397B, released in February; on Artificial Analysis's intelligence ranking it now sits 52nd of 113 with a score of 19, while Qwen3.8 Flash Next, released August 26, scores 40, ranks fifth, and was not run. Neither were DeepSeek V4 Pro, at 1.6 trillion parameters, or Kimi K3, at 2.8 trillion, both of which are already on the dashboard for NVIDIA chips. It runs one workload of three: a long question, 8,000 tokens in, with a short answer, 1,000 out. It runs nothing from AgentX, the multi-step agent benchmark, which SemiAnalysis says comes for TPU "later in the year." It carries zero accuracy results for any TPU on a page that held 2,746 of them as of September 11, so nobody has checked that the TPU's answers are as good as the GPU's on this test. And it carries a caption under the TPU curve: "TPU7x results are an official preview and may change as validation and publication continue." The GB300 NVL72 curve carries no such label. It is a published result, and every point on it links to the public run that produced it.
The objection here is about the reader, not the authors. Why come out with such a report if things are so skewed, when the average person doesn't really understand all these things that are missing?

Eight switches. The four configuration switches, all on the NVIDIA side, are each worth 1.9x to 3x on SemiAnalysis's own dashboard at 100 tokens per second per user. Flip those four back and "up to 50% better" becomes 8 to 12x the other way.
We are calling this the Peeled Apple: an apples-to-apples comparison that only comes out even after one apple has been peeled. SemiAnalysis's label for rack against TPU is "apples to bananas." The buyer does not get a bananas column. The buyer pays for tokens at the speed the user wants them, on whatever each vendor ships.
https://x.com/SemiAnalysis_/status/2097309231082807795
Apples to apples is the whole point of a benchmark, and it cuts the other way from how the article uses it. Strip the rack, the number format, the read-write split and the guessing trick off one side and the comparison is no fairer; the other side just looks like a contender. Fair is each chip at its best, on the same model, at the same speed, on the same accuracy test, and the dashboard can show exactly that. The headline is what you get when you stop one dropdown short of it.
Best Against Best Is 8 to 12x
We ran a commercial lighting company from 2005 to 2019, and every fixture we sold came with an LM-79 report. LM-79 is the lighting industry's test for a product as sold: the whole fixture, LEDs, driver and optics together, sent to a third-party lab that measures the light coming out in lumens and the power going in at the wall, and reports the number the buyer actually cares about, lumens per watt. Nobody bought on the LED chip's datasheet. The chip's number on a test bench at room temperature is not what the building pays for; the fixture's system wattage is, every hour, for a decade. So buyers compared products on the tested number, because the electricity bill, the utility rebate and the payback were all computed on it. Efficiency matters. These are decisions that buyers are making, and the ROI is determined by the actual performance. Headlines cannot be misleading, because in that business a misleading headline sat on the customer's power bill for ten years.
The chip-hour is the fixture price. The token is the lumen. Interactivity, tokens per second per user, is the light level the design has to hit, and a serving fleet is tuned to a target speed the way a lighting layout is tuned to a target brightness. So a buyer counts tokens per dollar at the speed the user expects, with each machine running the way its maker ships it, the way we counted lumens per watt on the fixture as tested. Nobody priced a retrofit on the sticker of the fixture, and nobody should price a fleet on the chip-hour. We said the same thing to AMD in July: move the wattage and you move the ratio, so publish the number you can defend.
On SemiAnalysis's cost model, the TPU chip-hour costs $1.21 against $2.31 for GB300 NVL72, or $1.03 if you are Google and pay the fab yourself. The chip-hour is about half the price. Now count lumens. When each user gets 100 tokens a second, GB300 NVL72 delivers 9.7x the tokens per dollar, 8.2x if you give Google its own lower cost. At 50 tokens a second the gap is about 6x. At 130, the fastest the TPU goes on this chart, it is about 12x. Turn it sideways and the same money buys speed: at the tokens per dollar the TPU delivers at 100 tokens a second, the NVIDIA rack delivers about 430, 4.3x faster. The TPU curve ends at about 130 tokens per second per user. The rack's runs to about 600.
Two things about the chart's accounting belong in print. The dashboard charges every chip, the ones reading the question and the ones writing the answer, so the rack's per-chip number is honest. But it is a rack: the winning NVIDIA points run 17 to 42 chips working together, while every TPU preview point runs on four. A buyer who wants the rack's economics buys the rack. And the metric counts tokens read and tokens written together, which for a long question and a short answer is nine read for every one written. Count only the written ones and the ratio barely moves: 9.6x at 100 tokens per second per user, 11.9x at 130.

Best configuration against best configuration on SemiAnalysis's own dashboard: GB300 NVL72 delivers 9.7x the tokens per dollar of TPU7x at 100 tokens per second per user and 11.9x at 130, or 4.3x the speed for the same money. The chip-hour is half the price. The token is not.
Anyone can rebuild that chart in two minutes. Open the dashboard, pick Qwen3.5 397B, the 8K / 1K scenario, FP4 and FP8, Total Tokens per $1 TCO on the y-axis and Interactivity on the x-axis, then flip the TCO basis between External and Internal and switch on Optimal Only. The table view gives the numbers behind every ratio in this piece, and every point links to the public GitHub run that produced it. Play with it yourself. That is the whole point of a public benchmark, and it is why the dashboard is worth more than the headline.
Tokens per watt is the number the trap was written in, and the preview publishes no power figure for TPU7x. Using SemiAnalysis's own power assumption for each chip, 2.12 kilowatts for GB300 and 1.207 for TPU7x, the gap is about 10.5x at 100 tokens per second per user and 13x at 130. That is a model, not a measurement.
As we wrote in The Cheap Chip Trap: "The trap is not the chip price. It is the gigawatt that has to feed the chip for the next five years." That piece was a spec-sheet call built on a Morgan Stanley FLOPs-per-watt chart, and we said so at the time. This is its first mark against a delivered TPU curve, on a public dashboard, on SemiAnalysis's own cost model.
Jensen said it at Stanford in March 2024, and it read as bravado then: NVIDIA's total cost of ownership is "so good that even when the competitor's chips are free, it's not cheap enough." Run it on this chart. Google's own $1.03 an hour only moves the ratio from 9.7x to 8.2x. Halve the TPU's entire chip-hour, silicon and power and building together, and the gap is still 4x. Quarter it and the gap is still 2x. No chip price on this chart closes it.
The Apples-to-Apples Numbers Moved the Next Day
Even inside SemiAnalysis's own framing, the dashboard moved the day after the article. The article's B200 number at the operating point it headlines was 8,903 tokens per second per chip, and $0.222 per million tokens when each user gets 100 tokens a second, with the TPU "approximately 19% lower cost than B200." The dashboard's September 8 run of the same B200 setup, on a newer release of the open-source software, shows 10,389 tokens per second per chip and $0.149 per million.
On that run, eight-bit against eight-bit, a single B200 box beats the TPU by 24 to 44% across the speeds a user would notice, 50 to 130 tokens a second. The TPU leads only at the slow end, under about 24 tokens per second per user, which is the point the article headlines, and there its lead is 29% on SemiAnalysis's cost model and 51% on Google's own. The B300 curve in the comparison dates from May 21, a four-month-old software build, and matches the article's number exactly. The authors used the runs they had. One of them moved the next day, the other had not moved since May, and on September 10 the B200's four-bit curve moved again, which is less a knock on SemiAnalysis than the point of the piece: on NVIDIA hardware the software moves every week, and a benchmark that runs continuously moves with it.
The point where the TPU wins is also the point where the user waits. Time to first token is how long you stare at a blank screen before the answer starts. At the article's headline operating point the TPU's median is 3.6 seconds and its average 5.4, which the article reports. B200 is 1.7 and 3.0. B300 is 0.8 and 2.4.
Read the other way, the wait favors the TPU. At its own operating points, 48 to 132 tokens per second per user, the TPU's first token arrives in about 0.3 seconds. The rack's cheapest points at the same speeds make the user wait 1 to 3 seconds, because they are serving hundreds to thousands of people at once. A buyer who needs a snappy first token at small scale is reading a different chart.

Same benchmark, one day apart. The B200 number in the headline moved 17% the next day; on the current run the TPU's lead at concurrency 256 is 29%, and at 100 tokens per second per user it is gone.
One thing cuts the other way, and it should be in print. NVIDIA's rack in eight-bit math without the multi-token trick, a curve dated July 13, trails the TPU at 75 to 130 tokens per second per user, by 10 to 30%. That is the "apples to bananas" comparison SemiAnalysis did make, where it found "about a 30% perf per dollar advantage for GB300 against TPUv7 agg" mid-curve and expects TPU software to close it "within a few months." On a July build at the wrong number format, the rack is beatable. That is a statement about a July build: the same curve sits below a single B200 box on the September 8 run, which is how stale it is.
The article also asserts that "When serving models using FP4 on NVIDIA GPUs, there is quality loss versus FP8." That is a testable claim, and SemiAnalysis's own site tests it. Its accuracy page runs the model on a set of grade-school math problems, GSM8K, the only test listed for this model, and scores the four-bit and eight-bit versions the same: 0.9685 against 0.9700 on B300, 0.9682 against 0.9694 on B200. That test is easy and we would not hang a precision verdict on it. The claim is asserted; the measurement on the same site does not show it.
One more mark against ourselves. In All Roads and Rockets Lead to NVIDIA in August we wrote that "the buyers stopped paying for tokens some time ago, and the two numbers do not rank the same," meaning a buyer ranks silicon on what a finished task costs, not on tokens. This dashboard has no task number, and for the TPU it has no accuracy number either, which is the other reason the empty evals column matters. Tokens per dollar is where the bill starts. A finished task is where the buyer's comparison ends.
Where the Data Cuts the Other Way
The first risk is that this is a software snapshot, not a silicon verdict. The TPU curve is a preview: open-source software, one token at a time, reading and writing on the same chips, one run. The two software features that would help it most, the read-write split and the multi-token trick, are not on the public dashboard yet, and SemiAnalysis expects them within months. Our gate: on InferenceX's Qwen3.5 chart, if Google's best public TPU7x setup, with both features on, comes within 2x of GB300 NVL72's best at 100 tokens per second per user by the end of March 2027, the 12x was a software-maturity snapshot and the Peeled Apple was a snapshot too. Anyone with a browser can grade it.
The second is one model, one workload. One model on one long-question, short-answer workload is one point. No long-answer workload, no agent workload, no accuracy submission. A TPU result on agentic workloads could land anywhere, and until it lands the 8 to 12x is a statement about one chart. The second half of the gate: a TPU7x accuracy-eval submission and an AgentX TPU result on the public dashboard. Both, or the preview stays a preview.
The third is the one we have carried since May. The live bear on the Cheap Chip Trap is that models keep learning to do the same work with far fewer tokens, faster than the world adds work. A 12x gap in tokens per dollar matters less in a world that needs a tenth of the tokens. The volume leg of the custom-silicon bear is also untouched. In The Neocloud Hypothesis in March we called that bear "more credible today than it was six months ago," because Claude and Gemini run mostly on TPUs and Trainium. This chart says nothing about where Anthropic's tokens run, only what they cost when they run there. The cost leg now has measured data against it. The volume leg does not.
The first two change the slope of the gap. Only the third changes the direction of the call.
So What?
SemiAnalysis's own text says why a public dashboard matters this year: "Ironwood (TPUv7) is the first generation in which Google is competing for others' inference workloads with chips that can be purchased outright or rented through its own cloud." It also dovetails with the Google TPU Blackstone effort to basically sell this to other companies as well, the WSJ-reported joint venture we covered in May with $5 billion of Blackstone equity behind it. The buyer of an externalized TPU is a stranger with a spreadsheet. The stranger can read the dashboard.
That is also the answer to the "Nvidia tax." VentureBeat's Sam Witteveen put the take plainly in April: "Google pays fab, packaging and engineering costs on its TPUs. It does not pay that margin." True, and it shows up on the dashboard as a chip-hour half the price. The margin the buyer pays is per token, and on this chart the tax runs the other way.
There is a second reason the gap on this chart matters, and Jensen gave it yesterday at Goldman's Communacopia conference. General-purpose compute, he said, "gives you fungibility, durability, versatility, rentability, and very importantly today, because of capital constraints, everybody's balance sheets, investability. The NVIDIA Compute is finally a computer that could be asset backed. We can use it to secure loans." The plain version: a GPU can be used by different customers. The tenant who leaves gets replaced by the next one, running a different model, on the same software everyone already uses. A custom ASIC or TPU cannot be ported as easily. Its workloads are tuned to it, its software is young outside its owner's walls, and if the one customer who runs it walks away there is no line of second customers waiting. That is why the GPU is the easier asset to finance: an asset anyone can use is an asset a lender can take back and re-rent. We wrote in August, off the $500 billion of financing platforms NVIDIA announced with six firms, that compute is the new collateral. This chart is the other half of that story. Financeability rests on the asset being usable by anyone, and usable means the buyer can see what it delivers, on public data, before the money moves. You need the data to show that, so customers can make the right decision. The dashboard is the first place that data exists for a TPU sold to outsiders, and it says what it says.
In April, after Google Cloud Next, we wrote that Camp Two's efficiency advantage "does not come from having a better accelerator than NVIDIA. It comes from having a better system". That still reads right, and it cuts both ways: the system on the dashboard is a preview running open-source software. Google's own claim for its next inference chip, TPU 8i, is about 80% better performance per dollar than Ironwood. Eighty percent does not close 8 to 12x. Google's published figure for TPU 8i is 10.1 petaflops of four-bit math per chip against Ironwood's 4.6 at eight-bit, a little over 2x per generation; NVIDIA's published rack figure for Vera Rubin is about 3.3x GB300 NVL72 on the same four-bit math. The comparison to watch is the next TPU against NVIDIA's next rack, four-bit against four-bit, on a dashboard neither company runs.
The question we keep landing on is whether "TPUs are just best served horizontally for Google type workloads," and whether a buyer with a choice would spend the money on NVIDIA GPUs and be better off. The dashboard is the first public place that question has a number.
For NVDA, which we have been publicly long since 2016, the Nvidia-tax leg of the bear just lost its data, and the share-loss story needs a new number. It is one data point for the sixth gate in The Only Company That Can Afford Free, open inference staying on NVIDIA iron; that gate grades on serving mix over two quarters, not on one chart.
For Google, no position, the dual-fleet path from Cloud Next, Vera Rubin A5X next to TPU, looks like the right hedge, because the TPU pitch to outsiders now lands on a public chart. For Broadcom and Marvell, no position, the custom-chip programs they are paid to build through 2027 and 2028 are not the question; how much of the accelerator market those chips end up with is what a measured 8 to 12x squeezes.
We wrote this down before pulling a single number, and it still holds. Either way SemiAnalysis stays the reference. One shop runs these numbers, so its assumptions become the market's assumptions, holes and all. The fix is not a better rebuttal. It is a second, independent set of numbers.
Coming Up
- We re-pull the InferenceX dashboard the morning this publishes and any morning a TPU configuration changes. If Google's disaggregated or MTP curves land, the update runs in the subscriber chat the same day with the gate regraded.
- The next MLPerf inference round is due; we will grade it against this chart when the results are public.
- We will be at the All In Summit this week in Los Angeles
Related BEP Research
- The Cheap Chip Trap Jensen Just Confirmed At Dell World (May 18): the spec-sheet call this chart marks
- The Third Announcement Most People Missed at Google Cloud Next (April 25): two camps, one dual-fleet cloud
- The NeoCloud Hypothesis (March 11): where the custom-silicon bear was named
- The Only Company That Can Afford Free (September 3): the sixth gate this chart feeds
- The Fourth Piece Ships (March 17): the first InferenceX number we printed, off a vendor slide
- AMD Showed a 15% Lead. Its Own Slide Showed Why It May Not Last. (July 27): reading a benchmark against its footnotes
Disclosure: BEP Research's principal is long NVDA, LITE, CRDO, TSEM, ALAB, WOLF, NOW, SMCI, BE, NBIS, and ORCL (2027 LEAPS), and has been publicly long NVDA since 2016. This is investment research, not investment advice. Do your own work.





