The AI conversation in wealthy markets is about frontier capability. The AI conversation everywhere else is about cost, connectivity and control — and on those axes, Chinese open models have a structural advantage that has very little to do with geopolitics.
Key takeaways
- Size fits reality: 83% of all-time model downloads are under 1 billion parameters; models above 100 billion account for about 1%. Only 3% of 2026 download volume came from anything above 70B.
- Licence fits reality: about 59% of Chinese releases above 20B parameters in 2026 are Apache 2.0 and 22% MIT, versus roughly 29% Apache-or-MIT among comparable US releases, with 41% custom terms and 30% no declared licence.
- Offline fits reality: Qwen leads GGUF — the quantised format used by llama.cpp, Ollama and LM Studio — at roughly 39.6 million pulls a month, against Gemma's 20.8 million and Llama's 7.5 million.
- Reach: Kling reached users across 224 countries and regions; Alibaba pushes Qwen through its cloud into Southeast Asia and Africa.
The three constraints that decide model choice outside the rich world
1. Hardware. A team building a local-language assistant does not have an H100 cluster. It has a workstation, maybe a single GPU, possibly a laptop. A model family that ships usable variants from 0.5B upward, in quantised formats, with a permissive licence, is not one option among many — it is the only option that runs.
2. Connectivity. API-dependent products are unavailable or unaffordable where bandwidth is expensive and latency is high. Local deployment is not a preference; it is a requirement. The GGUF pull numbers are the clearest evidence of this: that format is associated with someone actually running a model on their own machine, not evaluating it in a notebook.
3. Legal certainty. A startup in Nairobi cannot retain counsel to interpret a custom licence with field-of-use restrictions. Apache 2.0 is three paragraphs that any founder can read. This is why the licensing asymmetry in Hugging Face's data is the most decision-relevant fact in the report, even though it gets a fraction of the attention that benchmark rankings receive.
What "default workflow" actually means
Hugging Face's conclusion — that Qwen has become part of the default workflow for developers deciding what to fine-tune and deploy — sounds abstract. In practice it looks like this:
A developer in Southeast Asia building a customer-service bot for a regional language starts from a Qwen base because the tooling examples are in Qwen, the quantised versions exist, the fine-tuning scripts are published, and the licence permits commercial use. Having done it once, they do it again next time.
Defaults compound through tooling, tutorials and team habits. That is a deeper moat than a benchmark lead that changes every quarter.
The honest caveats
Three things need saying alongside this.
Downloads are not deployments. Hugging Face measures activity on its own hub and does not capture API usage or private deployment. Alibaba's 3 billion figure and Hugging Face's 2.045 billion use different bases; cite one and name it.
This is about distribution, not capability. Winning download share is not the same as winning the frontier. Chinese labs lead in reach; the frontier is contested.
The reach is policy-contingent. Beijing has weighed restrictions on overseas access to its strongest models. The global distribution that produced these numbers is a strategic asset that could be withdrawn — which would be, from the perspective of a developer in Lagos, a supply shock rather than a political event.
Why it matters beyond the numbers
There is a version of the AI future in which capability concentrates in a handful of proprietary systems, priced in dollars, accessible via API, and shaped by the priorities of the markets that can pay. There is another in which the substrate is open, local, and adaptable to languages and contexts that no frontier lab will prioritise.
Chinese open-weight labs are, whatever their motivation, currently the strongest force pushing toward the second version. That is worth stating plainly, because for developers in most of the world, it is the difference between building and renting.
Sources: Hugging Face, State of Open Models, 14 August 2026; Alibaba disclosures; Bloomberg and Tech in Asia reporting, August 2026.
