I run a small fleet of AI agents. They write, they publish, they watch markets, they file reports while I sleep. Which means that every month I do something most people arguing about AI on the internet never do: I look at the bill.
So when CNBC reported this month that Chinese AI models have been carrying at least 30 percent of OpenRouter’s token traffic every single week since early February, peaking at 46 percent, I wasn’t shocked. Mostly I was surprised anyone was surprised.
The AI race didn’t flip because of a breakthrough. It flipped because of an invoice.
The week nobody noticed
OpenRouter is the closest thing the AI industry has to a public traffic camera. It’s a routing layer that sits between developers and every major model, which means its token volume is one of the few honest windows into what developers actually use, as opposed to what they say they use at conferences.

Here’s what that traffic camera recorded. In the first half of 2025, Chinese-origin models carried about 4.5 percent of OpenRouter’s volume. Over the following twelve months the average climbed to 11 percent. Then, during the week of February 9, 2026, something happened that would have been front-page news in any other industry: Chinese models passed American models on an American-built platform serving largely American developers.
Not for one anomalous week. Every week since. By mid-2026 the weekly peak hit 46 percent. In the most recent provider breakdown, by one common measure, Chinese-origin models held 46.4 percent of routed tokens against 35.7 percent for US-origin models, on a platform now moving roughly 29 trillion tokens a week. DeepSeek alone accounts for 17.6 percent, the single largest vendor on the platform. Alibaba’s Qwen holds another 13.9 percent.
Read that again. The largest single AI vendor by traffic on one of America’s most important AI infrastructure platforms is DeepSeek.
While this was happening, the public conversation about AI was consumed by other things. GPT-5.6 launching through a government review gate. Grok pricing stunts. Humanoid robots running half-marathons in Beijing. All real stories. I wrote about one of them. But the structural story, the one about who actually processes the developer world’s AI workloads, slipped past almost everyone.
It’s the prices, stupid
The explanation isn’t complicated, and it isn’t ideological. Justin Summerville at OpenRouter put it plainly: Chinese open-source models run 60 to 90 percent cheaper than the leading American offerings.
The specific numbers are more brutal than the range suggests. As of June, DeepSeek V4 Flash costs 14 cents per million input tokens. GPT-5.5, the model that still anchors OpenAI’s public price list while its gated successor rolls out, costs five dollars. That’s not a discount so much as a different product category, the difference between bottled water and a municipal supply.
And for a growing share of real workloads, the quality gap no longer justifies the price gap. Kyle Chan at the Brookings Institution estimates Chinese models trail the American frontier by six to nine months. The US government’s own Center for AI Standards and Innovation put the lag at roughly eight months in a May report that tested cybersecurity, software development, math, science, and abstract reasoning.

Eight months behind, at three percent of the price.
If you’re a startup burning venture money on inference, that isn’t a hard decision. It’s barely a decision at all. The AI startup Lindy moved traffic from Anthropic’s Claude to DeepSeek, and its CEO Flo Crivello said the switch saves millions. He isn’t an outlier, just one of the few willing to say it on the record, because there’s still a mild social cost in Silicon Valley to admitting your American AI product runs on Chinese weights.
I understand the economics viscerally because I live a small version of them. My agent fleet burns tokens around the clock. When a routine task costs 35 times more on one model than another and both produce an acceptable result, loyalty isn’t a line item. Nobody’s loyalty survives a 97 percent discount forever. Mine included.
The part everyone gets wrong
The standard reaction to this story, when people notice it at all, splits into two takes.
The first take says this proves America is losing the AI race. It doesn’t. The frontier is still American. The most capable models in the world are still built in San Francisco and Seattle, and the eight-month lag is real. If you need the absolute best reasoning available on Earth, you’re still buying American, and the labs know it. That’s exactly why they price the way they do.
The second take is smarter, and it deserves to be stated at full strength. OpenRouter is a router for cost-conscious developers, not a census of the AI economy. Tokens aren’t dollars: the American labs sell most of their capacity through direct contracts and cloud partnerships that never touch a routing layer, and measured in revenue rather than volume, the race hasn’t flipped at all. On this view, the cheap tier has simply found its natural home in Chinese open weights, while the market that pays, the enterprise contracts, the premium subscriptions, the frontier API deals, remains solidly American. Nothing to see here.
Every sentence of that is true, and it still shouldn’t comfort anyone, because it describes exactly how incumbents get displaced. In technology, the cheap layer is where the ecosystem grows. The developers building on 14-cent tokens today are the companies setting architecture standards in three years. Every workload that becomes a solved problem migrates down the price curve, and the price curve now has a Chinese floor. Revenue tells you who’s winning this quarter. Volume tells you where the next generation is being built.

We’ve run a version of this experiment before. American companies invented the memory chip, then decided commodity DRAM was a low-margin business not worth defending. Japanese firms took the volume, and by the mid-1980s Intel had exited memory entirely, retreating up the stack to processors. It worked, once. But note what’s different this time, because it makes the situation stranger, not safer. The Japanese needed DRAM revenue to fund their next generation of fabs. Chinese labs give their models away; open weights make revenue almost beside the point. This flywheel doesn’t run on money, it runs on adoption. Every developer who builds on cheap Chinese weights is a developer whose tools, habits, and default architectures form around them, and defaults are what standards are made of. Intel could climb back up the stack. Nobody has yet demonstrated how you buy back a generation of developers’ defaults.
You cannot export-control a torrent
There’s also a quieter irony here. Washington spent the first half of this year building approval gates around frontier American models, the export-control theater I covered when GPT-5.6 became the first model to launch with what amounted to a permission slip. Whatever you think of that policy, note what it can’t touch: open weights. DeepSeek and Qwen ship their models to the world as downloadable files. You can’t export-control a torrent. The tighter the gate around American closed models, the better the open Chinese alternative looks to everyone standing outside it.
What the token flow is telling you
Traffic data is honest in a way that benchmarks and keynotes aren’t. Nobody routes five trillion tokens a week through DeepSeek to make a geopolitical statement. They do it because the math works, and the math is the message.
The message is that AI is splitting into two markets. There’s a frontier market, still American, where capability is the product and price is almost irrelevant. And there’s a volume market, increasingly Chinese, where AI is an input cost like electricity, and the only questions are reliability and price per kilowatt. The frontier market gets the headlines. The volume market gets the world.
American labs are betting that the frontier premium lasts, that there will always be enough customers who need the best badly enough to pay 35 times more for it. Maybe. But every month, more workloads discover they don’t need the best. They need good enough, always on, at a price that disappears into the budget. That migration only runs in one direction.
Six months ago the story was that Chinese AI was a copycat sideshow. Today it carries nearly half the traffic on an American routing platform, in plain sight, documented weekly, while the discourse argues about model names and demo videos.
The race didn’t flip when someone announced it. It flipped on the invoices, one procurement decision at a time, which is how these races always actually flip. The memo went out in February.
Check your own AI bill. You may find you already got it.
Chris Meredith writes about AI, technology, and what it actually means for real people. Follow along on Substack: monkeyattack.substack.com