NeuroAI NEUROAINEUROAI.SITE
ESC

Huawei's New Supernode Puts 4,096 Chips in One "Computer" — But the Win Is in the Glass, Not the Silicon

At Huawei Connect 2026 (17 September) Huawei launched the Ascend 960 supernode (昇腾960超节点), billed as the world's first to use NPO optical interconnect, scaling to 4,096 accelerators. With Cambricon's trillion-yuan market cap, China's domestic compute build-out is moving from chip races to system races.

2026-09-16 · 855 words · NeuroAI
Huawei's New Supernode Puts 4,096 Chips in One "Computer" — But the Win Is in the Glass, Not the Silicon

Someone figured out that the bottleneck in AI is no longer the chip — it is the wire between the chips. Huawei's answer, unveiled this month, puts the optical transceiver next to the silicon and calls the whole thing one computer.

On 17 September 2026, at Huawei Connect 2026 in Shanghai, Rotating Chairman Wang Tao (汪涛) launched the Ascend 960 supernode (昇腾960超节点) — described as the world's first supernode to use NPO (Near-Packaged Optics, 近封装光学). The pitch is simple to state and hard to build: connect up to 4,096 accelerators so they behave like a single machine, and do it with light instead of copper at the edges.

What "supernode" actually means

A supernode is a middle layer between one server and a whole data centre. Huawei's design uses the LingQu (灵衢) high-speed interconnect plus a new optical engine called Hi-ONE to give every component equal access to memory across physical servers. The result:

  • Scale: up to 4,096 cards in one supernode, up from 1,024 in the previous Ascend 950 generation.
  • Compute: 8 EFLOPS FP8 and 16 EFLOPS FP4 per supernode.
  • Optics: 5,500 Hi-ONE modules replace what would have been 48,000 800G optical modules, cutting power by more than 550 kW and lifting availability to 99.8%.
  • Cluster: multiple 960s link via LingQu or RoCE into clusters of 512,000 cards, and with multi-rail topology up to 1,000,000 cards.

The headline claim is that NPO — turning electricity to light centimetres from the chip — breaks the distance and bandwidth wall that copper hits. Hi-ONE is billed as the first mass-produced NPO product, at 7.2T capacity, and the only one with a built-in light source. Huawei says it has opened the LingQu 2.0 protocol to the industry and is joining a MIIT effort to standardise domestic interconnect — a tell that the real prize is an ecosystem, not a single box.

For context, Nvidia's GB200 NVL72 packs 72 GPUs; Huawei's argument is not that one 960 beats one NVL72 on efficiency — it is that 4,096 domestic accelerators wired as one machine clear a scale that Western export rules block. The win is in the system, not any single chip.

The context: China's compute build-out

The 960 is the front edge of a stack already in the field.

  • The Ascend 910C (CloudMatrix 384) supernode, 384 NPUs plus 192 Kunpeng CPUs, has been deployed in more than 1,000 sets across internet, telecom, finance and other sectors.
  • The Ascend 950 supernode (Atlas 950 SuperPoD) — 8,192 Ascend 950DT chips, 1 EFLOPS FP8, 256 TB unified memory — is slated for commercial launch in Q4 2026.
  • Huawei says the air-cooled and liquid-cooled 960 supernodes will ship in Q2 and Q3 2027.

The strategic logic is explicit: if you cannot buy the fastest chip, build the largest, best-wired machine from domestic parts.

The domestic rival: Cambricon's quiet quarter

Huawei is not the only one. Cambricon (寒武纪), the other flagship of Chinese AI silicon, posted a breakout 2026 first quarter: revenue of 2.885 billion yuan (≈ US$406M / HK$3.14B), up 159.6%, and net profit of 1.013 billion yuan (≈ US$143M / HK$1.10B), up 185%. Its MLU590 chip — 7 nm, about 256 TFLOPS FP16 and 512 TOPS INT8 — is in volume delivery and pitched at roughly 80% of Nvidia's A100. On 30 June 2026 Cambricon's market cap crossed 1 trillion yuan (≈ US$140B / HK$1.09T), the first STAR Market company to do so.

The synergy is the part worth watching: DeepSeek's V4-Pro achieved Day-0 adaptation to the Ascend platform in 2026, and several reports say DeepSeek chose Ascend 950PR as a flagship compute base. Model and silicon, both domestic, closing the loop.

The honest limits

The numbers are impressive and almost entirely Huawei's or Cambricon's own.

  • Benchmarks are vendor-stated. FP8/FP4 throughput, the 550 kW saving and 99.8% availability come from Huawei's launch, not an independent lab.
  • "World's first NPO supernode" is Huawei's framing. NPO is an industry-wide research area; the claim is about mass production and integration, not the concept.
  • Scale needs yield. Building 5,500 optical engines per supernode at volume is a manufacturing bet, not yet a track record.
  • Export controls are the backdrop, not the cure. Domestic parts solve availability, not the software-ecosystem gap versus CUDA.

What readers can do now

  • If you buy AI compute in China, ask for cluster-level numbers, not chip-level peaks — MFU, latency, and real token throughput per supernode.
  • Track Cambricon's MLU590 adoption as the canary for non-Huawei domestic silicon.
  • Watch the 950/960 commercial deployments in Q4 2026–2027 before treating "million-card clusters" as deployed reality.

Honest limitations

Ascend 960 figures come from Chinese state and technical media (Xinhua, Science and Technology Daily, CGTN, Tencent/Universal coverage of Huawei Connect 2026) and Huawei's own disclosures. Cambricon financials are from Eastmoney and the company's filings. Deployment counts (1,000+ 910C supernodes, 750+ 384 systems) vary slightly by source and date and are Huawei-reported. The "world's first NPO supernode" and all performance figures are vendor claims not independently audited. Exchange rates use roughly 7.1 RMB/USD and 1.09 RMB/HKD. Reporting is current to late September 2026.

More in “Chips & Compute” → · Back to home · Markdown version