DeepSeek Killed the Commodity Tier, Not the Frontier
Cheap open-weight models killed single-shot chat pricing. Multi-step agentic reliability is a different market, priced on different logic.
Originally published elsewhere.

The “AI is a cheap utility now” argument is right about one market and confidently wrong about the one where the money is.
There is a sentence doing the rounds on every podcast and timeline right now. DeepSeek is almost as good and a fraction of the price, so the American frontier labs are cooked. Their economics do not survive contact with an open-weight model that costs a tenth as much.
Half of that is correct. I want to give the correct half its due before I take the other half apart, because the people saying it are not fools and the strawman version helps nobody.
DeepSeek V4 is real. At its April 2026 release it landed within a fifth of a point of the best closed model on SWE-Bench Verified, ran a million-token context, and priced output around 8.6x below the closed frontier. It is open-weight. For a large and growing class of work, the right call is now to route to the cheapest model that clears the bar, and that model is increasingly Chinese and increasingly free to self-host. If your workload lives there, the utility thesis is not a prediction. It already happened.
So why is the conclusion wrong? Because “AI” in that sentence is one word standing in for two markets, and they price on opposite logic.
Two markets wearing one name
Call them the commodity tier and the frontier-agentic tier. Do not define them by which model you use. Define them by what binds.
The commodity tier is work where capability has saturated. Classification, extraction, summarisation, single-turn chat, retrieval answers. Good-enough is the whole bar, and most models cleared it eighteen months ago. When capability saturates, the only axis left to compete on is price, and price falls to the marginal cost of inference. There is no moat here and there was never going to be one. DeepSeek did not kill this tier. It euthanised something that was already terminal.
The frontier-agentic tier is different in kind. Long-horizon work. Many dependent steps, tool calls that touch real systems, an agent that has to stay coherent across the chain and recover when a step fails. Here the binding constraint is not price. It is reliability across the chain. And reliability is exactly where the cheap-and-almost-as-good story falls apart.
Why “almost as good” stops being almost good enough
This is the part the utility argument never models, because it reasons from single-shot leaderboards.
A three-point benchmark gap on a one-turn eval is nothing. You would not feel it in a demo. Now put that same model in a forty-step agent loop where each step depends on the last. Per-step reliability compounds.
A model that is right 99% of the time per step finishes the chain intact about two times in three. Drop to 97% per step and you finish intact roughly one time in three. Same three points. Completely different system. One ships. One pages you at 2am.

The commodity tier is single-shot, so the gap is cosmetic. The agentic tier is multiplicative, so the gap is structural.

That is the whole distinction, and collapsing it is how a true observation about cheap chat becomes a false conclusion about frontier economics.
The token bill is not where you think it is
There is a pricing version of the same error. Cheap tokens, the story goes, are what push enterprise spend through the roof. The Jevons-paradox framing: make the unit cheaper and total consumption explodes, so the discount is the cause of the bill.
It does not hold up. The thing inflating the bill is not unit price. It is intensity and breadth. An agentic task consumes vastly more tokens than autocomplete did, the rough order being a hundred to a thousand times more, and that figure is an illustration rather than a measured constant. The heavy spend then sits at the frontier tier, where price has fallen the least. Per-token cost for a fixed capability level is dropping fast at the commodity end and slowly at the frontier. So the place generating the giant invoices is precisely the place the discount has not reached. The cheap-tokens framing points you at the wrong end of the curve.

“But enterprises are clearly paying for it”
The standard rejoinder is to point at the spend. Some firm burned an enormous sum on frontier credits in a single month, so obviously the frontier is worth paying for.
The spend is real. It just does not prove what people want it to prove. Heavy adoption is evidence of perceived value, not of attributable return. Large enterprises (Google, Uber etc.) ran an internal leaderboard ranking teams by how much AI tooling they used, which means high usage is partly an artefact of the incentive rather than a clean signal of worth. Its own operating chief said publicly that he could not draw a line from the coding-agent usage to shipped consumer features. Controlled-study productivity gains of around 55% on scoped tasks thin out sharply in the field. None of that says the tools failed. It says “they spent a fortune” and “it was worth it” are two different claims, and only the first one is in evidence. The spend settles neither side of the utility debate. It is a separate question that only looks like this one.
What follows, and what to actually watch
If your work is commodity-tier, believe the utility thesis and act on it. Route to the cheapest adequate open-weight model and stop paying frontier prices for saturated capability.
If your work is agentic, invert the conclusion. The model is the cheapest part of your system. The cost and the risk live in the orchestration, the eval harness, the observability, the guardrails that catch a bad step before it acts. Model commoditisation is not an argument against that layer. It is the argument for it. The cheaper and more interchangeable the model gets, the more your durable advantage sits in the structure you wrap around it, and that structure does not commoditise on the same clock.
Here is what would make me wrong, because a claim with no failure condition is not worth posting. The two tiers collapse into one only if two things happen together. Per-task token consumption at the agentic tier falls materially through caching, routing, and handing sub-tasks to smaller models. And open-weight models close the reliability-across-many-steps gap, not just the single-shot benchmark gap. If both land, the frontier tier loses its moat and the utility crowd were simply early. So watch step-level reliability under long horizons, and watch per-task economics. Do not watch the headline leaderboard. The leaderboard is measuring the tier that was already lost.