Export controls were supposed to make Chinese AI slower. Instead they made it cheaper — and cheapness turned out to be the more defensible asset.
Key takeaways
- Token economics replaced capability races as the industry's central variable. Chinese commentary describes 2026 as the year the sector moved from a single contest over capability to a dual contest of "capability first, cost decisive."
- Open source crossed majority: on OpenRouter, open models went from 34% of processed tokens in January 2026 to 65% in June 2026, with 500+ organisations switching away from proprietary models.
- Price compression is extreme: Chinese open models have been available at roughly $0.18 per million output tokens against $30 for a leading US frontier model.
- Cost is now a national competitiveness question, not just a company one: cheap, efficient model supply is treated in Chinese analysis as industrial infrastructure.
- Efficiency is being pushed down the stack: models are being co-designed with domestic chips, with reported Day-0 adaptation across nine domestic AI chips and per-token serving costs on domestic clusters approaching parity with mainstream NVIDIA GPUs.
How constraint became capability
When you cannot assume unlimited high-end accelerators, engineering culture changes. Mixture-of-experts architectures, aggressive quantisation, distillation, long-context efficiency work and inference-time optimisations stop being nice-to-haves and become the main event.
The result is a set of models that deliver frontier-adjacent capability at a fraction of the inference cost. DeepSeek V4, released open-source in April 2026, became the flagship example: a million-token-class context window with weights and low-level code released globally, priced at levels that made entire categories of application economically viable for the first time.
Chinese executives have described the principle plainly — that strong models can be built "not by stacking computing power, but by relying on efficiency and innovation."
What cheapness unlocks
The important consequence is not that existing workloads get cheaper. It is that previously impossible workloads become possible.
When a million tokens cost cents rather than dollars, you can afford to run an agent that reads an entire codebase. You can run a hundred parallel evaluations. You can serve a national education system. You can put a model in a robot, in a hospital triage tool, in a farmer's phone.
This is why Chinese analysts describe cheap model supply as having risen "from a company competitive advantage to a national industrial competitiveness question." It is also why the country reports daily token consumption above 140 trillion and more than 600 million generative AI users — the scale only makes sense if inference is genuinely inexpensive.
The stack-level version
Efficiency work in China is no longer confined to model architecture. It now runs through the whole stack:
- CANN, Huawei's compute architecture for Ascend, has been fully open-sourced, with an ecosystem of over 3,000 partners.
- DeepSeek V4 achieved Day-0 adaptation across nine domestic AI chips, meaning Chinese models run on Chinese silicon from the day they ship.
- Zhipu's GLM-5.3-Flash has been reported serving large-scale production traffic on domestic chip clusters at per-token cost comparable to mainstream NVIDIA GPUs — the threshold at which domestic compute stops being a subsidy and starts being competitive.
That last point is the one to watch. Cost parity on domestic silicon converts a geopolitical risk into an engineering fact.
The counter-pressure
None of this means capability stopped mattering. Chinese labs are still shipping flagship models — MiniMax M3, GLM-5.2, Qwen3-Max-Thinking — each chasing differentiated strengths in long context, code and inference. A model that is cheap but weak has no market; the entire argument rests on the gap having narrowed to roughly 2.7% on competitive coding leaderboards.
There is also a real risk in the cost story: price competition can compress margins industry-wide. Analysts have noted the sector's shift toward "looking at revenue, calculating costs and comparing efficiency" as investors move away from pricing on expectation. Chinese AI companies now have to win the same way everyone else does — by being both good and cheap, sustainably.
Pricing and usage figures as reported by OpenRouter and cited in Chinese and international media in 2026; efficiency claims based on company disclosures.
