---
title: "Alibaba's new AI chip is 3× last year's — and it's a bet on the whole stack"
date: 2026-09-28
category: Chips & Compute
site: NeuroAI
canonical: https://neuroai.site/a/na-chips-zhenwu-v900-ai-chip
language: en
---

# Alibaba's new AI chip is 3× last year's — and it's a bet on the whole stack

> At the 2026 Apsara Conference, Alibaba's chip unit T-Head unveiled Zhenwu V900, a training-inference AI chip it says is three times as fast as its previous generation. The real story is not the speed number — it is how China's largest cloud builder is stitching chips, models and data centres into one sovereign stack.

A chip you cannot buy yet just became the clearest signal of where Chinese AI infrastructure is heading. At a conference in Hangzhou, Alibaba showed off the part that sits inside its own data centres — and the number it led with was "three times."

For anyone outside the semiconductor world, the interesting part is not the transistor math. It is the fact that China's biggest cloud company is no longer content to buy its brains from somewhere else.

## The chip that actually showed up

On 22 September 2026, during Alibaba's Apsara Conference (云栖大会), the company's chip arm **T-Head (平头哥)** introduced **Zhenwu V900 (真武V900)**, a training-and-inference AI chip. Alibaba calls it the most powerful domestically designed AI chip in China. The headline claims:

- **3× the performance** of its previous-generation Zhenwu M890.

- **216 GB of on-chip HBM** memory, with **1,200 GB/s** inter-chip bandwidth.

- **Native FP8 and FP4** low-precision compute, the formats that matter for running giant models cheaply.

- **Mass production planned for the first quarter of 2027.**

The previous chip, M890, had already done something concrete: it ran models like Qwen3.8 and Kimi K3 at over **two trillion parameters**, and Alibaba says the Zhenwu family had served more than **650 enterprise customers** by June 2026. So this is not a slide-deck prototype. It is the next step in a line that is already inside production clusters.

## Why "3×" is the wrong headline

Speed beats are easy to announce and hard to verify. The more useful question is *why* a cloud company is building its own silicon at all.

Alibaba's answer is integration. The V900 was shown as part of a "compute–storage–network" full-stack story: the chip plus a self-designed interconnect switch (ICN Switch), a smart network card, and a storage controller, all working together so that **thousands of chips act like one**. Alibaba says a single cluster built this way can scale toward **500,000 cards**.

That matters because the bottleneck for large models is rarely one chip. It is getting thousands of them to cooperate without the system falling over. Owning the interconnect and the scheduler — not just the die — is what lets a hyperscaler squeeze cost out of the whole machine.

## The stack play: chip, model, cloud

The V900 does not live alone. It is the hardware half of a loop Alibaba is deliberately closing:

- **Model pulls chip.** Alibaba says future Qwen models will scale toward **5 to 10 trillion parameters**, and it is already training the next-generation Qwen4.

- **Chip serves model.** V900 and its supernodes are the substrate those models run on.

- **Cloud sells both.** Alibaba Cloud repackages the compute as a service, and the company has pledged to run **more than 20 gigawatts** of global data-centre capacity by 2032.

The logic is simple and a little intimidating: control the layer underneath the model, and you control who can run the model affordably. For Chinese developers worried about export controls and overseas-API outages, a domestic chip that already runs a two-trillion-parameter model is a sovereignty argument as much as an engineering one.

## What it cannot do yet

Two caveats are worth saying out loud.

First, **it is not shipping**. Mass production is targeted for early 2027. Today the V900 is an announcement with a spec sheet, not a product you can order. The M890 it replaces is the one actually in clusters now.

Second, **the "3×" is Alibaba's own claim**, against its own previous part, on its own workloads. Independent benchmarks of the V900 do not yet exist, and real-world efficiency depends on software, not just silicon. A chip is only as good as the compiler and runtime that drive it — and that is the part outsiders cannot test yet.

## What readers can do now

- **If you build on Chinese AI infrastructure**, treat "domestic chip + domestic model" as a fast-maturing default, not a fallback. Test inference latency on an M890-backed endpoint today; the V900 generation is the upgrade path.

- **If you invest or report on AI hardware**, watch *cluster scale* and *interconnect*, not peak teraflops. The 500,000-card claim is where the real moat lives.

- **If you follow policy**, read this as a marker of the "full stack" doctrine: chips, models and clouds are now planned as one strategic unit, not separate markets.

## Honest limitations

Core facts (release date 22 September 2026 at Alibaba's Apsara Conference; T-Head as maker; 3× vs M890; 216 GB HBM; 1,200 GB/s; native FP8/FP4; Q1 2027 mass production; 650+ enterprise clients by June 2026; 500,000-card cluster target; 20 GW data-centre goal by 2032; Qwen 5–10T parameter plan) come from Alibaba's own conference disclosures, reported by the Associated Press (22 September 2026), 广州日报 (Dayoo), 科创板日报 (STAR Market Daily), 新浪财经 (Sina Finance) and a CITIC Construction Research note (27 September 2026). The "3× performance" and benchmark framing are Alibaba's vendor claims and have not been independently benchmarked by a third party. The chip is not yet in mass production; all performance and cluster-scale figures describe announced targets or current M890-generation deployments, not V900 in the field. This article contains no RMB amounts, so no currency conversion applies. Analysis is current to 28 September 2026 and is not investment advice.

---

Published by NeuroAI (https://neuroai.site/) — https://neuroai.site/a/na-chips-zhenwu-v900-ai-chip
Free to quote with attribution and a link to the original.
