NeuroAI NEUROAINEUROAI.SITE
ESC

Huawei's next Ascend chips — a 2028 roadmap built on system scale, not single-chip speed

Huawei's Ascend (昇腾) 910C is now shipping at volume, but the more telling move came in September 2025 when the company laid out a multi-year roadmap: the Ascend 950 in 2026, 960 in 2027 and 970 in 2028. The strategy is explicit — win through cluster scale and fast iteration rather than any one chip matching Nvidia's best.

2026-09-28 · 1005 words · NeuroAI
Huawei's next Ascend chips — a 2028 roadmap built on system scale, not single-chip speed

The most important thing about Huawei's AI chips is not the chip you can buy today. It is the plan the company wrote down for the next three years. In a market where everyone compares dies, Huawei is arguing that the die is the wrong unit of comparison.

That argument starts with a part already in the field — and gets interesting with the parts that are only on a slide.

The chip that's shipping: Ascend 910C

The Ascend 910C (昇腾910C) began mass shipments in May 2025. It is, in essence, two 910B-class dies packaged together using chiplet technology, built on SMIC's second-generation 7nm-class (N+2) process, with around 530 billion transistors and an estimated 55% domestic-content rate.

On paper, the widely cited figures are roughly 640 TFLOPS at FP16 (some estimates run higher, toward 800) and about 3.2 TB/s of memory bandwidth, at a typical ~310 W draw. In inference, DeepSeek's own testing — cited by Reuters — put the 910C at about 60% of an Nvidia H100's performance.

One caveat up front: Huawei has not published a full datasheet for the 910C. Almost every per-chip number above is an analyst or supply-chain estimate, not a vendor-confirmed spec. Treat them as directional.

The number that matters is the cluster, not the die

Huawei's real answer to Nvidia lives one level up. The Ascend 384 supernode (昇腾384超节点) — the Atlas 900 A3 SuperPoD, also marketed as CloudMatrix 384 — links 384 Ascend 910C chips into a single high-bandwidth fabric, announced in May 2025, with reportedly over 500 such units deployed.

The logic is blunt: a single 910C is well behind an H100, but 384 of them, wired together with all-optical interconnect, can present an aggregate system that competes at the cluster level for training and serving large models. Analysts such as SemiAnalysis note the trade-off plainly — having several times as many Ascends offsets each chip being only a fraction as fast, at the cost of far higher power draw (the CloudMatrix 384 draws roughly 4 times the watts of Nvidia's GB200 NVL72).

That rack-scale system has a price tag too: a CloudMatrix 384 set has been reported at around RMB 60 million (≈ US$8.3M / HK$65M). The strategy is to spend more chips and more electricity to close the per-chip gap.

The roadmap: Ascend 950, 960, 970

At its Connect conference in September 2025, Huawei laid out a forward roadmap that puts the 910C in context:

  • Ascend 950 — arriving in 2026
  • Ascend 960 — planned for 2027
  • Ascend 970 — planned for 2028

Each step is described as raising low-precision throughput and shipping alongside larger Atlas SuperPoD and SuperCluster systems scaling toward hundreds of thousands, then over a million accelerators. Huawei framed this as a multi-year plan to compete through system scale and frequent iteration — explicitly, not through any single chip matching Nvidia's best.

What the 950 actually promises

The most concrete next step is the Ascend 950 series. Reported specifications for the 950PR variant (targeted at inference prefill, Q1 2026) include about 1.56 PFLOPS at FP4, 112 GB of in-house HBM, and a 600 W envelope; the 950DT variant (Q4 2026, for decode and training) is described with up to 4 TB/s of bandwidth. The 950 generation also brings native FP8 and FP4 low-precision formats, the formats that matter for running giant models cheaply.

These are vendor roadmap claims about products not yet broadly shipped, so they describe intentions and design targets, not measured field performance. The in-house HBM in particular is a notable bet — it signals Huawei trying to control the memory supply that constrains every Chinese accelerator.

The rumored 910D — and why caution is warranted

A further part, the Ascend 910D, has circulated in reporting (The Wall Street Journal reported in April 2025 that Huawei had approached Chinese firms about testing it, with first samples expected around May 2025). It is described in leaks as a multi-chiplet design possibly on a refined process, aimed at closing more of the gap with high-end Nvidia parts.

But the 910D is not an officially announced or spec-confirmed product. Anything stated about it should be treated as rumor until Huawei publishes details. For now, the 950 series is the credible "next generation" that is actually on the company's own roadmap.

What readers can do now

  • If you build on Chinese AI infrastructure, evaluate Ascend at the supernode level, not the chip level. Ask about CloudMatrix 384 deployment, the CANN software stack, and model-porting effort — those decide real-world usability more than TFLOPS.
  • If you follow the hardware race, track the Ascend 950's 2026 delivery and in-house HBM yield. A domestic HBM win would matter more than any single FLOPS number.
  • If you report on the sector, separate confirmed shipments (910C) from roadmap claims (950/960/970) and rumors (910D). Mixing them is how the picture gets distorted.

Honest limitations

Core facts (Ascend 910C mass shipments beginning May 2025; chiplet dual-910B design on SMIC N+2 7nm, ~530B transistors, ~55% domestic content; ~640 TFLOPS FP16 and ~3.2 TB/s bandwidth as widely cited estimates; ~60% of H100 inference per DeepSeek testing cited by Reuters; the Ascend 384 / Atlas 900 A3 SuperPoD / CloudMatrix 384 with 384 chips announced May 2025 and >500 deployed; Huawei's September 2025 Connect roadmap of Ascend 950/960/970 for 2026/2027/2028; Ascend 950PR ~1.56 PFLOPS FP4, 112GB in-house HBM, 600W, and 950DT ~4 TB/s; CloudMatrix 384 ~RMB 60 million per SemiAnalysis; the Ascend 910D as an unconfirmed rumor per WSJ reporting) are drawn from Reuters (via Wikiwand and aiwiki.ai), SemiAnalysis commentary, Baidu Baike (citing Huawei's disclosures), and Huawei's own September 2025 roadmap announcement. Huawei has not published a full 910C datasheet; per-chip figures are analyst/supply-chain estimates. The RMB 60 million figure is converted at roughly 7.2 RMB/USD and 1.09 HKD/RMB. The 950-series specs are vendor roadmap claims for not-yet-broadly-shipped products, and the 910D is unconfirmed. Facts are current to 28 September 2026 and this article is not investment advice.

Related coverage

More in “Chips & Compute” → · Back to home · Markdown version