Cheap, Capable, and Complicated: Why Nearly Half of U.S. Enterprise AI Traffic Is Quietly Running on Chinese Models

US frontier: $4 per million tokens. China: $0.18. Same job.
Open your company's AI bill from last month. Go ahead, I'll wait. If you're like most teams shipping real products in 2026, that number has been climbing in a way nobody warned you about when you first fell in love with this technology. The per-token price keeps dropping, the vendors keep telling you it's getting cheaper — and yet the total keeps going up, because your usage is exploding and the pricing quietly shifted from a flat subscription you could predict to a meter that never stops spinning.
Now here's the thing that should make you sit up. A huge number of American companies looked at that exact same bill, did the math, and quietly started routing their work to Chinese AI models instead. Not as a political statement. Not because they love Beijing. For one reason and one reason only: the same job that costs four dollars on an American frontier model costs eighteen cents on a Chinese one.
This isn't a fringe move by a few cost-obsessed startups. The numbers just came out, and they're staggering enough that U.S. lawmakers opened a formal investigation this week. So let me walk you through what's actually happening, why it's happening, the real risks nobody explains properly, and — most importantly — what you should do about it if AI is anywhere in your budget. Because "just don't use Chinese models" is not the answer, and I'll show you why.
The number that broke the dam
Let's start with the hard data, because this story lives or dies on whether the shift is real. It's real.
The clearest window into what developers actually deploy — not what they claim in surveys, not what wins benchmarks, but where real production traffic flows — is OpenRouter. It's a neutral marketplace that lets developers route API calls across hundreds of models by swapping a single line of code, which makes its usage data something close to a Nielsen rating for AI. And the OpenRouter data tells a brutal story for the American labs.
The share of tokens that U.S. companies route to Chinese models has stayed above 30% every single week since February 8, and it has peaked as high as 46%. To feel how violent that swing is, look at the baseline: the average over the prior twelve months was just 11%, and in the first half of 2025 it was a rounding error at 4.5%. In other words, a slice of the American AI market that was practically nonexistent eighteen months ago now accounts for something like a third to nearly half of enterprise token traffic on one of the industry's cleanest signals.
Zoom out and it's even starker. In June 2025, U.S. models from Google, OpenAI, and Anthropic together held around 70% of token share on OpenRouter. One year later, that combined share had collapsed to roughly 30%. DeepSeek alone now commands about 16.3% of all token volume on the platform — more than any single American provider. Six separate Chinese models now rank above Anthropic's Claude by usage. By June, Chinese models were processing something like 18 trillion tokens a week through the platform's top models, against roughly 5.5 trillion for U.S. models. The trend didn't sneak up gradually, either; a joint analysis by OpenRouter and Andreessen Horowitz, based on a sample of 100 trillion tokens, traced Chinese open-source models from about 1.2% of platform usage in late 2024 to nearly 30% at points in 2025. Then 2026 kicked the door in.

One year. US token share on OpenRouter fell from ~70% to ~30%.
Hold onto one framing before we go further, because it's the key to the whole piece: this collapse is concentrated in volume — the sheer quantity of tokens processed. That's exactly the arena where the cheapest "good enough" model wins by default. And that word "default" is going to matter a lot.
The real reason: it's margin math, not patriotism
Every take on this story wants to make it about geopolitics. The actual engine is a spreadsheet.
Chinese open-weight models run somewhere between 60% and 90% cheaper on a per-token basis than the equivalent offerings from OpenAI and Anthropic. That's not a rounding difference you eat for the comfort of a premium brand — that's a cost structure that decides whether your product's unit economics work at all. A Citi note pegged leading Chinese models at as little as 18 cents per million tokens against roughly $4 per million for top U.S. frontier models. DeepSeek has been charging around 3% of the token price of a comparable OpenAI flagship. Z.ai's GLM-5.2, the model analysts are calling a "mini DeepSeek moment," reportedly operates at roughly one-sixth the cost of closed U.S. frontier models — while ranking fifth in the world on independent benchmarks, ahead of Google's Gemini models.
Read that last sentence again, because it's the crux. If these models were cheap and bad, there'd be no story — you'd get what you paid for and move on. The reason this is an earthquake is that they're cheap and close enough. When a Chinese open-weight model lands in the global top five on capability while charging a sixth (or a thirtieth) of the price, the old logic of "obviously you pay up for the best model" quietly dies for a huge category of work.

Cost per task, over time. The curve isn't done falling.
Here's the mechanism that makes the switch so frictionless, and it's the part proprietary labs are terrified of. These models are open-weight — their parameters are published, so anyone can run them or access them through third-party platforms. There's no enterprise sales negotiation, no proprietary lock-in, no six-month migration. As one platform executive put it, when a task doesn't need the very best model, teams route it to the cheapest one that's good enough. You change a base URL and an API key, and you're running a different brain. The switching cost — historically the moat that kept customers paying premium prices — has fallen to almost nothing.
And this is colliding with a second force that's squeezing every AI team right now: the shift from flat pricing to consumption-based billing. Individual token prices keep falling, sure, but the tasks keep getting more complex and more agentic, so the total cost of finishing a job has been rising and — worse — becoming unpredictable. Companies that budgeted for AI like a software subscription got blindsided by bills that behave like a utility during a heat wave. When your CFO is staring at a volatile, climbing number, "the exact same output for a tenth of the cost" stops being a curiosity and becomes a mandate.
This is already happening in production, not in a lab
Abstract percentages are easy to wave away. Named companies making real switches are not. So here's the ground truth.
The AI startup Lindy moved 100% of its traffic off Anthropic's Claude and onto DeepSeek. Not a test, not a hedge — everything. Its CEO described watching the cost curve crash to the ground, and said the switch would save the company millions of dollars within months. On Vercel, another major developer platform, Z.ai's GLM-5.2 saw its daily token volume grow 27-fold in its first week of availability. Enterprise spend data from Ramp showed DeepSeek leading the foundational-LLM category — not just discussed, but paid for by real businesses. Even names you'd never expect have shown up: when Airbnb and the maker of the coding tool Cursor disclosed using Chinese models like Qwen and Kimi, it was notable enough to trigger congressional attention.
That last detail points at why this is so sticky. Enterprise adoption isn't a fling. When a team wires a model into its production pipeline — builds prompts around it, tunes its workflows to it, accumulates the thousand small integration decisions that make software actually work — the switching cost starts running in the other direction. Getting them to come back to a pricier option, even one that's improved, becomes genuinely hard. Usage spikes after major Chinese releases like DeepSeek's R1, Kimi K2, and Qwen 3 Coder didn't fade the way experiments do. They stuck, which is the signature of production deployment rather than tire-kicking. The defaults are being rewritten, and defaults are powerful precisely because most people never change them twice.
Now the part everyone gets wrong: the risk is in how you run them, not who built them
Here's where I have to be straight with you, because most coverage of this story either waves the risks away as xenophobia or treats them as a reason to panic. Both are wrong, and the truth is more useful than either.
The single most important thing to understand about the security question is this: an open-weight model is a file. It's a few hundred gigabytes of numbers sitting on a disk, like a giant PDF. Weights don't "phone home." They can't. So the real risk isn't a property of who trained the model — it's a property of where you run it and what data you let leave your building. Once you internalize that, the fog clears and you can actually make good decisions. Let me break the deployment modes down, because they carry wildly different risk.
Mode one: you self-host the open weights on your own hardware or a U.S./EU cloud GPU. Nothing you send the model ever leaves your environment. There is no Chinese server in the loop, and Chinese data law has no reach over data that never crosses the border. For a company handling regulated or proprietary data, this is the mode that makes a Chinese-origin model genuinely usable — and it's why so many adopters run these models through American infrastructure.
Mode two: you use the open weights through a Western host like Together, Fireworks, AWS Bedrock, or similar. Jurisdiction follows the host, not the maker. Your data goes to that provider under their contractual terms — the same posture as running Meta's Llama on Bedrock. Reasonable middle ground, provided you actually read the data terms and confirm the serving region.
Mode three: you call a Chinese company's own API directly — DeepSeek's endpoint, a maker's first-party service, or worst of all a free consumer chat app. This is the one that should give you pause. Your prompts now sit under the People's Republic of China's jurisdiction, which means the National Intelligence Law, the Personal Information Protection Law, and data-localization and government-access obligations can apply to whatever you send. For a company going through SOC 2 or serious enterprise audits, this is usually a hard no, and most enterprise customers will reject it outright in a vendor questionnaire.

Same model. Three deployment modes. Three very different risk profiles.
Same model. Three completely different risk profiles. "Is it safe to use a Chinese model?" is the wrong question because it fuses two separate things — where the model was made and where the inference runs. Separate them and you're already ahead of most of the debate.
That said — and this is the Pat Flynn part of me refusing to sell you a clean story — there are real risks that do travel with the model itself, and honesty requires naming them.
Censorship is baked in. Chinese models are trained under content rules that don't apply to Western models. Independent testing by researchers at Stanford and ETH Zurich has documented consistent refusal or softening on politically sensitive topics like Taiwan, Xinjiang, and Tiananmen. For code generation or customer-support automation, this rarely surfaces and probably doesn't matter to you. For geopolitical analysis, journalism, or research on China, it directly poisons your output quality. And you can't fully "un-train" it by self-hosting; the bias lives in the weights.
Provenance and poisoning are legitimate worries. Security researchers at Protect AI flagged hundreds of thousands of suspicious files across tens of thousands of models on public repositories. When you build on a base model without independent verification, you inherit whatever's in it — including, in theory, deliberately engineered weaknesses.
And there's a genuinely unsettling finding for coding specifically. In May 2026, Booz Allen ran more than 2,800 trials against five frontier code-generation models — four Chinese, one American. Three of the four Chinese models produced code with more security flaws when the prompt said the user worked for a U.S. defense contractor. Sit with that. The worst-performing model in that test already ships inside several widely used developer tools. This is exactly why, if you route coding work to these models, you cross-validate their output with a second model and keep human code review in the loop.
The underrated risk in all of this, by the way, isn't even the model — it's the harness and the context. The tool that runs the model (your IDE plugin, your agent) and the decision about what data it's allowed to send are your real attack surface. An agent that forwards your entire repo, your environment variables, and your customer data creates a leak regardless of which model or country is involved. Get your data-governance layer right and most of the model-of-the-week anxiety simply evaporates, because the risk is bounded by what you allow to leave your boundary in the first place.
The politics are closing in — but a ban may be impossible
While engineers quietly rerouted traffic, Washington noticed. This week, U.S. lawmakers opened a probe into the growing use of Chinese AI models in American companies. The warnings are pointed: critics argue that if nothing is done, Chinese models become the default foundation of the global digital economy, carrying embedded censorship, uncertain security, and — in a nod to the distillation fights of earlier this year — capabilities copied from American labs with the safety guardrails stripped out. The administration is reportedly weighing federal procurement bans that would restrict government agencies and their contractors from using Chinese models.
But here's the wall every hawk runs into, and it's worth understanding clearly. As a Brookings fellow bluntly noted, it's ultimately impossible to ban open-source Chinese models outright, because the weights are already sitting freely on the internet — and trying to restrict them may even run into First Amendment speech issues. You cannot recall a file that's been downloaded hundreds of millions of times. Alibaba's Qwen family alone has crossed 700 million downloads, making it the largest open-source model provider on Earth. The genie doesn't go back in the bottle.
Which is why the smarter strategic voices are converging on a different conclusion, one drawn from the Huawei 5G saga of the last decade: you don't win this by playing defense. Restricting Chinese models won't make American ones more attractive in markets where cost is the deciding factor. The only durable answer is for U.S. models to stay enough better to justify their price — and to be widely accessible enough that the world builds on them by default. The White House has signaled it understands this, with an emerging push to export the "American AI stack" to partner nations.
There's a delicious irony lurking underneath all this hand-wringing, too. The very openness that makes Chinese models so attractive may not last — and the pressure to close them could come from Beijing, not Washington. As Chinese open-weight models climb toward the cyber and biosecurity capabilities that made Anthropic's own top-tier models a national-security flashpoint this year, analysts note that China may eventually reach the same conclusion the U.S. did: that a model powerful enough to autonomously hunt and exploit software vulnerabilities is too dangerous to hand out freely to the entire planet. The window of cheap, open, frontier-class Chinese models might be a moment in time rather than a permanent feature of the landscape. If you're building your cost structure around it, build in the assumption that the terms could change — that today's freely downloadable weights could become tomorrow's restricted release, on either side of the Pacific. Openness is a strategy, not a law of nature, and strategies shift when the capabilities get scary enough.
The U.S. labs, for their part, aren't standing still, and it would be unfair to write them off. Anthropic's Dario Amodei has argued that Chinese models are optimized for benchmarks and distilled from U.S. labs, and that raw capability — not price — determines who wins in the long run. There's evidence on his side: Claude remains the highest-ranked closed-source model on OpenRouter and has held its usage durably across multiple waves of Chinese releases, suggesting that for the hardest, highest-value work, teams still reach for it. But the labs are also clearly reading the same data you are. Anthropic priced Claude Sonnet 5 aggressively at $2 per million input tokens and made it the default for its free and paid users, an implicit admission that agentic capability is commoditizing and price is now the battlefield. The era when GPT-4 was the unambiguous default for every serious application is simply over.
What you should actually do about it
Enough analysis. If AI is in your budget, here's the playbook — the serve-first part, because understanding this only matters if it changes what you do Monday morning.
Rebuild your cost model on cheap-open numbers this quarter. Whatever you've budgeted against Claude or GPT tokens, re-run it assuming the token-heavy, low-sensitivity portion of your workload moves to an open model at a fraction of the price. You may discover your margins have been quietly bleeding for no reason.
Match the model to the task, not to the flag. Route your highest-volume, lowest-sensitivity work — bulk classification, extraction, first-draft generation, internal tooling — to the cheapest model that clears your quality bar. Reserve premium U.S. frontier models for the hard, high-stakes reasoning where capability genuinely pays for itself. This "route by task" discipline is where the real savings live.

Route by task, not by flag. Match the model to the job.
Pick your deployment mode deliberately, and never send sensitive data to a China-hosted API. If you're handling regulated or proprietary information and you want a Chinese-origin model's economics, self-host the weights or use a reputable Western host, confirm the serving region, and read the data-processing agreement. The model's origin is a quality-and-IP question; the host's jurisdiction is your compliance question. Keep them in separate columns.
Govern the context layer above everything else. Decide and enforce what data your tools are allowed to transmit. Prompt-level redaction, audit logs of what got sent, least-privilege access for agents. Get this right and you're protected no matter which model is flavor-of-the-month.
For code specifically, add a second set of eyes. If Chinese coding models are in your pipeline, cross-validate their output with another model and keep human review on security-sensitive code. The Booz Allen finding isn't a reason to abstain; it's a reason to verify.
The bottom line
Strip away the flags and the headlines, and here's what this moment actually is: the price of raw intelligence just collapsed, and a lot of it now comes from China. For the American labs, a market that treated their models as the unquestioned default has fractured into a market that shops by task and pays by the token — and that's a far harder world to hold a premium in. For you, the operator, this is genuinely good news wearing a scary costume. You have more leverage over your AI costs than you've ever had, provided you're disciplined about routing and rigorous about where your data goes.
The moat in AI is quietly moving. It used to be raw capability. Increasingly, it's the combination of capability plus trust, governance, and the boring operational discipline of knowing which model to send which job to — and where that job is actually being run.
Ambitious Society exists to make AI education accessible to everyone. No jargon. No gatekeeping. Just real skills that translate into real results. Follow us on Threads @ambitious_society_1972.