In May 2026, Nvidia's CEO told CNBC the company had "largely conceded" China's AI-chip market to Huawei. The shock was not the words — it was that the leader of the world's most valuable chipmaker said them on earnings day, with revenue up 85% to US$81.62 billion and China's data-center contribution effectively zero.
That single quote reframes the whole China-chip story. It is no longer "when will Huawei catch Nvidia per chip?" It is "what happens when a buyer is pushed onto a different stack entirely?"
The sanction that redrew the map
US export controls have tightened since 2022, first cutting the A100/H100 and later requiring a license for H100/H200 sales to China (CNBC, May 2026). Reuters, via CNBC, reported that some Chinese firms — Alibaba, Tencent, ByteDance, JD.com — received US approvals to buy H200 chips, but Beijing discouraged uptake, pushing domestic developers onto homegrown silicon.
Huang told investors to "expect nothing" on China approvals. The result, in his own words: "We've really largely conceded that market to them."
The chip: 昇腾 (Ascend) 910C on SMIC
Huawei's workhorse is the 昇腾 (Ascend) 910C, a dual-chiplet accelerator built on SMIC's 7-nanometre-class DUV process. Analyst and secondary reporting puts it at roughly 800 TFLOPS of FP16 compute, 128 GB of HBM, and 3.2 TB/s of bandwidth — with each 910C delivering about one-third the BF16 throughput of Nvidia's B200.
Volume is the strategic bet. Multiple analyst notes cited by several outlets estimate Huawei is targeting roughly 600,000 Ascend 910C units in 2026 — about double 2025 output — across a total Ascend line that could reach ~1.6 million dies. Treat those as estimates, not Huawei-disclosed shipments.
The real move: a rack, not a die
The more interesting answer is the CloudMatrix 384. Huawei packs 384 Ascend 910C NPUs plus 192 Kunpeng CPUs into 16 racks, linked by an all-to-all "UnifiedBus" mesh. Huawei claims over 300 petaFLOPS of BF16 compute; SemiAnalysis' independent April-2025 calculation put it at roughly 300 PFLOPS dense BF16 — about double Nvidia's GB200 NVL72 (~180 PFLOPS), with 3.6× the memory capacity and 2.1× the bandwidth, at roughly 4× the power.
By the World AI Conference in July 2026, Huawei said the 384-chip supernode had been deployed in 750+ commercial projects across internet, telecom and finance. The logic is simple: if one chip is a third as fast, use five times as many and schedule the rack as one machine.
The harder problem: software, not silicon
Hardware was never the only wall. Nvidia's moat is CUDA. Huawei's answer is CANN (Compute Architecture for Neural Networks), which it open-sourced in August 2025, alongside the MindSpore framework (open-sourced in 2020). CANN bridges to PyTorch and TensorFlow, but coverage for multimodal and custom operators is thinner, and the 910C lacks confirmed FP8 hardware — so production inference often falls back to INT8 or FP16.
The honest read: horizontal scaling works for inference at production scale. Frontier model pre-training is where the gap bites — some labs have reportedly reverted to Nvidia H800s for key training runs after stability and throughput issues on Ascend.
The bottleneck nobody talks about: yields
The quiet constraint is SMIC's 7nm-class yield. Analyst estimates put it at roughly 50–60%, against 80–90% at TSMC. That is the whole story: China can design and package the chip, but throws away a large share of every wafer and cannot buy the EUV machine that would fix it.
Honest limitations
- Per-chip specs (800 TFLOPS FP16, 128 GB HBM) and the 2026 volume targets (~600k 910C, ~1.6M dies) come from analyst notes and secondary reporting, not Huawei's own disclosed shipment figures.
- The "Nvidia China share fell to ~8% in 2026, Huawei ~50%" figures are Bernstein Research estimates reported by trade media, not official market data.
- We did not benchmark CloudMatrix 384; the ~300 PFLOPS and efficiency comparisons are Huawei's claims, with SemiAnalysis' calculation broadly consistent.
- Frontier-training maturity on Ascend remains contested (some labs reverted to Nvidia for key runs) — verify per workload, not per headline.
What readers can do now
- Plan for CANN, not CUDA: if you run China-facing inference, budget engineering time for CANN porting and design for INT8/FP16 rather than FP8.
- Watch SMIC yields, not launch events: the 50–60% vs 80–90% gap is the real supply constraint — track it before believing any "mass production" claim.
- Separate hype from procurement: any "domestic chip beats Nvidia" line must be checked against per-watt efficiency and software-stack maturity, not peak TFLOPS.
