The most consequential number in Chinese AI infrastructure is not a FLOPS figure. It is a procurement decision: how much of the stack can be bought domestically.
In 2026 the answer moved decisively. Henan province landed what is described as the first all-domestic 100,000-card AI supercluster — a cluster built entirely from Chinese accelerators, not a mixed deployment with imported hardware as the primary tier.
Key takeaways
- All-domestic at scale: Henan's 100,000-card cluster is the first reported AI supercluster built entirely from domestic accelerators.
- Supernode architecture: Huawei's CloudMatrix 384 links 384 Ascend 910C NPUs and 192 Kunpeng CPUs across 16 racks into a single schedulable machine, at roughly 300 PFLOPS dense BF16.
- Deployment volume: Huawei reported at WAIC 2026 that the 384-chip supernode had been used in more than 750 commercial projects.
- Regional specialisation: Anhui is building AI chip capability; Hubei is leveraging its optoelectronics industry for AI compute demand; Henan is repositioning from a rail hub to a compute hub.
- Power is being treated as part of the problem: state enterprise AI programmes explicitly include "deepening compute-electricity coordination" (算电协同) as a workstream.
Why "all-domestic" is the milestone that matters
A mixed cluster is a hedge. An all-domestic cluster is a commitment: it means the operators have concluded the domestic stack is good enough to bet a nine-figure deployment on, including the software.
That last word is the hard part. A cluster of 100,000 accelerators is useless if you cannot port your training framework, your operators and your debugging tools. Which is why the software open-sourcing of 2025 was a prerequisite, not a nicety: CANN open-sourced, operators published, MindSpore available under Apache 2.0 since 2020, and a commitment of 1,500 PFLOPS of compute and 30,000 development boards a year to ecosystem work.
By WAIC 2026, Huawei reported the CANN community at 67 projects and more than 3,500 monthly active developers. Small in absolute terms — but it is a real number attached to a real ecosystem, and it is growing.
The engineering logic: bigger machine, weaker chip
When you cannot buy the fastest chip, you compensate in three ways, and China is doing all three simultaneously:
- Scale up harder. 384 chips in one coherence domain, with 3,168 optical fibres and 6,912 400G LPO modules, and cross-node bandwidth degradation under 3%. The result, per independent analysis, is close to double the dense BF16 compute of NVIDIA's GB200 NVL72.
- Pay in power. The same analysis estimates about 4.1× the power draw, roughly 2.5× worse per watt. This is the honest cost of the approach, and Chinese engineers discuss it openly — hence "compute-electricity coordination" appearing in national programmes.
- Optimise the software. Serving DeepSeek-R1 with 320-way expert parallelism and INT8 quantisation, Huawei and SiliconFlow documented prefill throughput of 6,688 tokens/second/NPU and decode of 1,943 tokens/second/NPU within a 50 ms per-output-token budget.
The regional pattern
One underappreciated feature of this build-out is that it is distributed by design. Rather than concentrating all compute in coastal tech hubs, provinces are competing on differentiated advantages: Anhui on chips, Hubei on optoelectronics and photonics for interconnect, Henan on scale and power availability.
This mirrors how Chinese manufacturing capability developed — regional clusters with complementary specialisations, competing for investment and talent. Whether it produces efficient allocation or duplicated capacity is a genuine open question, but the pattern is deliberate.
What remains unresolved
Stated plainly, from China's own policy analysis: high-end chips remain a chokepoint, high-end compute supply and demand are mismatched, and basic software still lags.
The 100,000-card all-domestic cluster does not eliminate those gaps. What it does is establish that the domestic stack can be deployed at commercially meaningful scale — which is the precondition for the investment needed to close them.
Infrastructure built under constraint has a particular property: it is optimised for the constraint, and when the constraint eases, the optimisation does not disappear. China is building for the world it expects to operate in.
Sources: Xinhua and China Youth Daily reporting on regional compute build-out, 2026; Huawei disclosures at WAIC 2026; SemiAnalysis, April 2025.
