NeuroAI NEUROAINEUROAI.SITE
ESC

Can Huawei's Supernode (超节点) Outsmart the Chip Ban—or Just Defer the Problem?

At Connect 2026 Huawei said it has deployed over 1,000 Ascend 910C supernodes and unveiled the NPO-based Ascend 960 supernode. The strategy is system-level networking over a single miracle chip.

2026-09-20 · 922 words · NeuroAI
Can Huawei's Supernode (超节点) Outsmart the Chip Ban—or Just Defer the Problem?

On a Shanghai stage on 17 September 2026, Huawei's rotating chairman pulled back the curtain on a server cabinet big enough to rewire China's AI plans. Inside sits the Ascend 960 supernode (超节点) — a system that links 4,096 chips into something the company claims behaves like one computer. The bet is simple and desperate at once: if you cannot buy the world's best single accelerator, connect thousands of decent ones so tightly that the gap stops mattering.

Huawei is not whispering this. At its Connect 2026 keynote it laid out deployment numbers, a chip roadmap and a networking architecture in one breath. Whether the claims survive independent benchmarking is a separate question, addressed below.

What Huawei actually announced

Rotating chairman Wang Tao (汪涛) presented the following at Huawei Connect 2026 in Shanghai on 17 September 2026:

  • Ascend 910C supernodes (超节点) have been deployed in more than 1,000 sets.
  • Ascend 950 supernodes have entered scale commercial use.
  • The Ascend 960 chip is ahead of schedule, with performance doubled versus the prior plan.
  • The 960DT is targeted for Q1 2027, three quarters early; the 960PR for Q3 2027, one quarter early.
  • The roadmap continues at one generation per year: Ascend 970 in 2028 and Ascend 980 in 2029, each doubling compute规格.

The throughline is cadence. Huawei is promising a predictable annual beat of new silicon, which matters more than any single spec because it signals the supply chain behind the chip is stabilising.

The NPO supernode (超节点) explained

The headline hardware is the Ascend 960 supernode, described as the industry's first to use NPO (Near-Packaged Optics, 近封装光学). The claimed specs:

  • Scales to 4,096 accelerators in a single node.
  • Delivers up to 8 EFLOPS of FP8 compute and 1 PB of HBM capacity.
  • Designed to train and serve ten-trillion-parameter large models (大模型).
  • Uses 5,500 Hi-ONE optical engines to replace 48,000 800G optical modules, cutting over 550 kW of power, doubling fault-free runtime, and reaching 99.8% availability.

Huawei says multiple such supernodes can be networked into far larger systems — up to 512,000 cards, and, with multi-rail topology, a theoretical 1,000,000-card supernode cluster. Those top-end numbers are architectural ceilings, not deployed systems, and should be read as such.

The logic is sound where it counts: in large-model training, the bottleneck is rarely a single chip's math. It is the communication between chips and the memory each can reach. Wiring thousands of chips into one addressable pool attacks exactly that wall.

The money and the ecosystem behind it

None of this is cheap, and Huawei's own books show where the cash goes. In its 2026 semi-annual report (released 31 August via domestic clearing houses), Huawei posted:

  • Revenue of 467.819 billion yuan (about US$65.9 billion / HK$510 billion) for H1 2026, up 9.55% year on year.
  • R&D spend of 121.382 billion yuan (about US$17.1 billion), a record for any first half, or 25.94% of revenue.

That R&D ratio is the real story. Huawei is spending more than a quarter of revenue on engineering — and a large share is feeding exactly the compute and networking stack described above.

The software side got quieter but arguably more important updates:

  • openEuler has passed 20 million installations and is the No. 1 server operating system in China by share.
  • Ascend became directly installable from the official PyTorch site — the first China-built compute platform to be offered that way.
  • The CANN compute framework is fully open-sourced, a direct answer to the ecosystem gap that long dogged domestic accelerators.

Why this matters outside China

The strategic wager is that system-level innovation offsets a process-node gap. The United States restricted the most advanced manufacturing equipment; Huawei's response is to stop depending on a single breakthrough die and instead stitch many available ones into a coherent whole. If the networking holds, a million loosely-coupled chips can approximate the work of a smaller number of cutting-edge ones.

It is also a bet on a parallel stack. A compute base that runs PyTorch natively, ships its own OS, and open-sources its compiler is one that foreign labs can, in principle, adopt without touching a blacklisted part.

What readers can do now

  1. Track deployments, not keynote specs. The "1,000+ Ascend 910C supernodes" figure is the one to watch — it is a real shipment count, not a projection.
  2. Watch the software metrics. PyTorch-installability and CANN's open-source community are better leading indicators of adoption than any TFLOPS claim.
  3. Discount the ceiling numbers. The million-card cluster is an architectural target; treat it as a design claim until a customer references a live system.

Honest limitations

  • Every performance and deployment figure above originates from Huawei's Connect 2026 keynote on 17 September 2026, as reported by established Chinese outlets including China Youth Daily, Securities Times and Southern+. They are vendor-stated claims, not independently benchmarked results.
  • I could not locate a direct Reuters, Bloomberg or AP primary report verifying the Ascend 960 specs at publication; international coverage of the event exists but frames it as a high-profile, uncertain bid rather than confirming the numbers.
  • The 467.819 billion yuan revenue and 121.382 billion yuan R&D figures come from Huawei's own 2026 semi-annual report, published through domestic financial clearing platforms and carried by China Daily and Securities Times.
  • The 512,000- and 1,000,000-card cluster sizes are Huawei's stated theoretical maximums, not confirmed live deployments.
  • Currency conversions use approximately 7.1 RMB/US$ and 1.09 RMB/HK$.

More in “Chips & Compute” → · Back to home · Markdown version