A single AI accelerator card has not been the bottleneck for a while. The bottleneck is the distance between them.
At Huawei Connect 2026 in Shanghai this month, Huawei stopped talking about one more chip and started talking about 4,096 of them behaving as one machine.
The shift is a strategic admission: when you cannot buy the fastest silicon on the market, you make the system wrapped around it smarter.
What Huawei actually shipped
At Huawei Connect 2026, which opened on September 17 in Shanghai, Huawei unveiled the Ascend 960 (昇腾960) supernode — what it calls the world's first supernode built on NPO (near-package optics, 近封装光学). Rotating chairman Wang Tao (汪涛) framed the company's AI strategy around computing power and "hardware monetization," built through a "supernode + cluster" approach rather than chasing a single flagship chip.
The numbers that matter:
- A single Ascend 960 supernode scales to 4,096 accelerator cards, up from 1,024 in the prior-generation Ascend 950 supernode.
- It delivers 8 EFLOPS of FP8 compute and 16 EFLOPS of FP4 compute.
- The NPO engine, branded Hi-ONE, carries 7.2T of total transmission capacity. Huawei says it replaces 48,000 conventional 800G optical modules with 5,500 Hi-ONE engines, cutting system power by more than 550 kW and lifting availability to 99.8%.
String several Ascend 960 supernodes together and you get clusters of up to 512,000 cards, or — with a multi-track topology — as many as 1 million cards.
Why a supernode, not a faster chip
The physics Huawei is attacking is familiar to anyone training a large model (大模型): once a cluster crosses roughly 100,000 cards, server-to-server communication can consume more than 40% of training time, gutting the model floating-point utilization (MFU) that decides whether a cluster is actually useful.
Huawei's answer is to bind hundreds or thousands of cards behind one memory-addressing fabric — its Lingqu (灵衢) interconnect — so a weaker-per-card domestic chip can still win on system throughput. The Ascend 950 supernode already offers 1 EFLOPS FP8 / 2 EFLOPS FP4 across 1,024 cards with a 256 TB unified memory space. The 960 doubles the card count and swaps copper for optical links to break past copper's distance and bandwidth ceilings.
The wider domestic compute base
Huawei is not beginning from scratch. Wang said more than 1,000 Ascend 910C supernodes have already been deployed, and the Ascend 950 supernode is in commercial use across internet, telecom, finance, education, healthcare, transport and manufacturing customers.
The company also reiterated a "one generation per year" rhythm. The Ascend 960DT is specified at 2 PFLOPS FP8 / 4 PFLOPS FP4, up to 288 GB of HBM and 9.6 TB/s memory bandwidth, and is now expected as early as Q1 2027 — three quarters ahead of plan. The 960PR follows in Q3 2027, with the 970 (2028) and 980 (2029) already mapped. The CANN compiler stack has been fully open-sourced.
The other domestic pillar is Cambricon (寒武纪). In its first-half 2026 report (released August 7), the company posted revenue of RMB 5.996 billion (about US$840 million), up 108.13% year on year, and net profit of RMB 2.311 billion (roughly US$320 million), up 122.61%. But momentum cooled within the half: Q2 revenue was RMB 3.111 billion (about US$435 million), only 7.8% above Q1, and inventory reached RMB 8.248 billion (about US$1.15 billion) — 45.32% of total assets — a reminder that order books and shipped silicon are not the same thing when high-bandwidth memory is scarce.
Underlying demand is real. China's National Data Administration reported daily AI token calls of nearly 175 trillion in June 2026, the highest of any country. That is the pressure cooker pushing both Huawei's supernodes and Cambricon's delivery ramp.
External researchers put Huawei's 2025 Ascend output near 800,000 units and project roughly 1.5 million for 2026 — still a fraction of Nvidia's expected volume. The point is not parity; it is a credible, growing second source.
Honest limitations
- Every hard specification above — card counts, EFLOPS ratings, HBM capacity, roadmap dates — comes from Huawei's own keynote and product disclosures at Huawei Connect 2026. We have not independently benchmarked the hardware, and vendor-stated throughput has historically diverged from real-world MFU.
- The Cambricon H1 figures are from the company's official semi-annual filing. The "delivery bottleneck" reading is analyst interpretation (Morgan Stanley, via secondary coverage), not a Cambricon statement.
- The ~1.5-million-unit 2026 estimate is from research group Epoch AI (which cites the Wall Street Journal) and is a modeling projection, not a shipment tally.
- Export controls, HBM allocation and SMIC yield remain live variables; a roadmap date that slipped by quarters once can slip again.
What readers can do now
- If you evaluate AI infrastructure in China, stop benchmarking single cards and start asking about supernode density, unified memory and interconnect bandwidth — that is where usable throughput lives.
- Track CANN and Lingqu (灵衢) ecosystem adoption (partner count, model ports) rather than chip-spec sheets; software maturity is the harder moat to climb.
- Treat "domestic alternative" as a portfolio question, not a single-vendor bet — Huawei and Cambricon (寒武纪) cover different procurement and performance profiles.
