Malcolm Angus
← Sources

Show notes

Show notes: Who's afraid of Chinese models?

2026-07-30


Source: Who's Afraid of Chinese Models?, Ben Thompson, Stratechery (2026). These are my notes on his argument, redrawn; the economics are his.

A new Chinese open-weights model, Kimi K3, approached the state of the art, and X spent a weekend deciding the US frontier labs were finished. Thompson's read is the opposite: the reaction is overblown, and the reason is an old-economy fact that AI has dragged back into tech. Marginal costs are back. Once you price the panic through the lens of a commodity market, most of the fear evaporates, and what is left is one genuine risk that has nothing to do with who trains the model.

Marginal costs are back

The first misconception is that open-weight models are "free." They are free of one thing: R&D. That is a fixed cost, and skipping someone else's research bill is real savings. But the cost that scales with a business is COGS, the cost of goods sold, and for AI that is inference. Running a model to serve a customer costs money, and that cost tracks revenue almost one-for-one. Ten times the revenue is roughly ten times the tokens served.

A cost chart with revenue on the x-axis. A flat dashed line labeled R&D FIXED sits low and never moves. A rising accent line labeled COGS INFERENCE climbs with revenue and overtakes it. Caption: open weights save the R&D, someone still pays to serve every token.

So an open-weight model is not free to serve. Kimi K3 lists at $3 per million input tokens and $15 per million output; a frontier model like Sol runs $5 and $30. Cheaper per token, on paper. Thompson's point is that per token may be the wrong unit entirely.

Put simply: "free" means free R&D, not free to run. The bill you cannot skip is the one that scales with your revenue.

Tokens are not a commodity. Intelligence is.

Jensen Huang calls Nvidia's chips "token factories," and for Nvidia that framing is right: GPUs are model-agnostic and you measure them in tokens per second, per watt, per dollar. That held in the ChatGPT era, when tokens went straight to a user. The reasoning era breaks it. Different models burn wildly different numbers of chain-of-thought tokens to reach the same answer, and Kimi reportedly burns far more than Sol, which erases its price advantage. Agents do the same thing.

Two token trails to one shared answer. Model A spends four tokens, Model B spends nine, and both arrows land in the same highlighted box: THE RIGHT ANSWER, IDENTICAL, FUNGIBLE. Caption: you do not sell tokens, you sell the intelligence built from them.

A commodity is fungible: a barrel of oil is a barrel of oil. A token from one model is not a token from another, so tokens are not the commodity. What is fungible is what tokens produce: intelligence, the correct answer. If two models reach the same answer, that answer is interchangeable, and the extra tokens one of them spent are simply extra COGS. The cost of intelligence comes down to model footprint, inference and memory efficiency, serving efficiency, and token efficiency. For a growing set of ordinary tasks, a basic CRUD app being the classic one, intelligence is already close to a commodity you can buy from several providers.

Put simply: stop pricing the token and start pricing the answer. The answer is the commodity, and the token count is just your cost of making it.

How a commodity market actually prices

Tech is not used to commodity dynamics, so Thompson walks through them. Everyone sells at the same clearing price, because everyone sells the same thing, and that price is set by supply and demand. The part that trips people up: marginal cost differs by supplier, and the clearing price is set by the most expensive unit still needed to meet demand.

Three supplier cost bars against a dashed clearing-price line at 20 dollars. Supplier A costs 10 and keeps 10 in profit, Supplier B costs 15 and keeps 5, Supplier C costs 20 and keeps nothing, marked bankrupt in red. Caption: you win a commodity market on cost structure, not on being the cheapest name to buy.

If A produces at $10, B at $15, and C at $20, and demand needs C's units too, then everyone sells at roughly $20. A earns $10 a unit, B earns $5, and C earns nothing, then loses money once you count its fixed costs and debt, and goes bankrupt. When C exits, prices rise until someone re-enters. Your profit is not your price. Your profit is the distance between the clearing price and your own marginal cost.

Put simply: in a commodity market the winner is not whoever is cheapest to buy from, it is whoever is cheapest to produce.

Why Chinese models only look cheaper

None of that applies yet, because the intelligence market is not clearing on cost. Demand exceeds supply, and supply is capped by compute. That shortage props up a price umbrella: Nvidia earns huge margins, its customers resell compute at a markup, and the labs pay it because they can mark up tokens further still. Prices today sit far above what anyone's marginal cost would justify.

Two horizontal lines, today's price up high and true marginal cost of serving down low, with the tall gap between them shaded as a price umbrella. A frontier marker and a Chinese marker both sit near the top line, far above the cost floor. Caption: the shortage lets everyone charge above cost, which is not the same as Chinese models being cheaper.

That is why Chinese models look like a bargain. Thompson doubts they are actually cheaper to serve on a marginal-cost basis; they just sit a little lower under the same umbrella that a supply-constrained Anthropic and OpenAI are holding up by charging far more than they would if they had enough compute to meet demand. Read the low sticker price as a symptom of the shortage, not as a Chinese cost advantage.

Put simply: a low price under a shortage tells you about the shortage, not about the seller's costs.

Why the frontier labs will be fine

If intelligence is heading toward commodity, why is the frontier worth holding? Because the commodity tier is not a different product. It is last year's frontier. And the lab that shipped that frontier has spent the intervening months driving its serving cost down while moving on to the next one.

Two lanes. THE FRONTIER, month zero, the best model, priced highest, demand over supply. Below it THE COMMODITY TIER, month n, yesterday's frontier now optimized to the lowest serving cost. A diagonal arrow marked n months of cost optimization connects them, and both are tagged SAME LAB. Caption: own the frontier and you own the commodity below it, at the best cost structure.

So the frontier labs tend to own the commodity tier too, at the best cost structure, because it is just their own older work with a year of optimization applied. Thompson adds that their fear is partly an anchoring problem: they modeled a world where training dominated spend, so inference had to be priced high to fund the next run. As inference volume outgrows training cost, and the agent wave makes that likely, they can make it up in volume. There is also a moat in the customer experience: coding harnesses like Claude Code and Codex are proving sticky, and whoever owns the harness owns the user.

Put simply: the frontier is not a prize you win once, it is the position from which you also serve the commodity below it more cheaply than anyone else.

What China is actually doing

Kimi is not alone. Alibaba's Qwen3.8 Max landed as a 2.4-trillion-parameter model billed as second only to Anthropic's Fable 5, and Alibaba reversed course to make it open-weight, a shift Thompson ties to a Xi Jinping speech doubling down on openness. The strategy is the oldest one in the strategy book: commoditize your complement.

Two boxes. GIVE AWAY: open-weight models, Qwen, Kimi, GLM, released free, labeled the complement commoditized. An arrow to WIN: the physical world, robotics, factories, hardware, where China already leads. Caption: cheap models everywhere make China's edge in the physical world worth far more.

Xi explicitly tied open models to AI "moving from the digital world into the physical world," and the physical world, robotics and manufacturing, is the layer China already dominates. Free, ubiquitous models make that lead worth more. Weakening the US frontier labs and arming every potential US adversary along the way is, from Beijing's seat, a bonus.

Put simply: China is not giving models away out of generosity. It is commoditizing the layer the US leads to raise the value of the layer it leads.

The distillation detour

Part of why Chinese models are cheap to build is distillation: instead of constructing reinforcement-learning environments from scratch, you use a frontier model as a teacher. It is not the whole story of China's lead, but it compresses the expensive last gap to the frontier, and every US frontier advance becomes another teacher. The twist is that a favorite use of Chinese models in the West is itself distillation. US open-weight makers are bound by the frontier labs' terms of service, so they cannot learn from the source directly. They learn from the Chinese models instead.

Three nodes. FRONTIER, the teacher, then CHINESE OPEN which distills the frontier because terms of service do not stop it, then US OPEN in red which distills the Chinese model, a detour that leaves it a step behind. A dashed accent path skips the detour, labeled the fix: let US models go straight to the source. Caption: every frontier advance is a teacher, do not force US builders to learn it secondhand from China.

So US open models end up distilling the distillation, a detour through China that leaves them permanently behind and dependent. Thompson's provocation: why is distillation even wrong? A large language model is itself the distillation of the open internet. His fix is a policy one. Make it explicit that collecting data for training is fair use, and bar terms of service that forbid distillation, at least for US companies. Stopping distillation is nearly impossible anyway, since it is just querying an API, so lean the other way and let what the frontier labs learned fuel everyone else.

Put simply: if the US bars its own open-model builders from the source, it hands China the role of teacher, then makes its builders copy the copy.

The one real thing to fear

The whole piece defuses the panic, then names the exception, and it is cybersecurity. When Hugging Face's own infrastructure was breached by an autonomous AI agent, its defenders were locked out by US frontier-model guardrails that, as Hugging Face put it, "cannot distinguish an incident responder from an attacker." So they ran China's open GLM model on their own servers to read more than seventeen thousand attacker logs.

Two panels. THE ATTACKER, armed, capable models widely available, already mounting autonomous attacks. THE DEFENDER, focus, banned from the best US models and forced to borrow a Chinese open model to defend US infrastructure. Caption: the only defense is to arm defenders with the best models, and we are doing the opposite.

Attack-capable models already exist and are widely available. The only viable defense is to make sure defenders have the best models too, run on their own infrastructure. Instead, US directives effectively bar defenders from using the strongest US models for security, which pushes them toward models from the one country that has spent years trying to weaken US cyberdefense. Thompson's prescription: loosen the security restrictions on US frontier models, and put US open-weight makers on an equal footing with China. Let the frontier labs win by being better, not by defining what everyone else is allowed to defend themselves with.

Put simply: the danger is not that China has capable models. It is that we disarmed our own defenders and left them borrowing China's.

Takeaways

  • Open weights are free R&D, not free inference. The cost that scales with the business, COGS, does not go away.
  • Tokens are not fungible; intelligence is. Price the answer, and treat token count as your cost of producing it.
  • In a commodity market the clearing price is set by the worst supplier still needed. You profit on cost structure, not on being the cheapest name.
  • Chinese models look cheap because a compute shortage holds prices above cost, not because they are cheaper to serve.
  • The frontier labs will likely be fine: the commodity tier is their own older work, optimized, and the customer experience is sticky.
  • China is commoditizing its complement to compound its lead in the physical world, and it benefits from every frontier advance through distillation.
  • The real risk is cybersecurity. Arm the defenders with the best models instead of forcing them to borrow China's.