NeuroAI NEUROAINEUROAI.SITE
ESC

Zhipu's GLM-5 Ships 744B Open Weights — Trained Entirely on Huawei Chips

Released 11 February 2026, Zhipu's GLM-5 is a 744-billion-parameter open-weight large model (大模型) trained wholly on Huawei Ascend chips, scoring 77.8% on SWE-bench Verified — the strongest open coding result at launch.

2026-10-03 · 955 words · NeuroAI
Zhipu's GLM-5 Ships 744B Open Weights — Trained Entirely on Huawei Chips

A lab that went public in Hong Kong this year just handed the world a model anyone can download. The catch is where it was built: not a single NVIDIA GPU touched the training run. That single fact may matter more than any benchmark.

What actually shipped

On 11 February 2026, Zhipu AI — now branded Z.ai internationally — released GLM-5, the newest flagship in its GLM large model (大模型) series. It is a mixture-of-experts model with roughly 744 billion total parameters and about 40 billion active per token, up from the 355B-class GLM-4.5 and GLM-4.6 lines. The weights are published under the MIT license on Hugging Face and ModelScope, so anyone may download, fine-tune, and ship them commercially.

GLM-5 is explicitly built for "agentic engineering" — not chatting, but planning, editing across many files, running tests, and fixing what breaks. It carries a 200,000-token context window, can emit up to 128,000 tokens in one response, and ships a dedicated thinking mode, function calling, and structured output.

Why "open weight" still means something

Plenty of vendors say "open" and mean "open a tab in our chatbot." GLM-5 is not that. The checkpoint lands in BF16 and FP8, and the MIT licence removes the commercial restrictions that fence in many rivals. A clinic in São Paulo, a startup in Lagos, or a research group in Berlin can run it without sending a prompt to a foreign server.

The honest footnote is size. The BF16 weights run to roughly 1.5 terabytes, so "open" does not mean "runs on your laptop." It means open to anyone with a server — which is most of the enterprises that would actually deploy it. Community FP8 quantizations shrink that footprint, at some cost to precision on code-critical tasks.

The Huawei Ascend bet

The detail that separates GLM-5 from almost every other frontier release is the silicon. Zhipu trained the model entirely on Huawei Ascend 910B chips — about 100,000 of them — using the MindSpore framework. No NVIDIA GPUs were used in pretraining.

That is a direct response to U.S. export controls that limit China's access to the most advanced NVIDIA hardware. A model of this size proving it can be trained end-to-end on domestic chips is, for Chinese AI strategy, a stronger signal than another percentage point on a leaderboard. It also widens the field of buyers: GLM-5 has been shown running inference on chips from Huawei, Moore Threads, Cambricon, Kunlunxin and others.

Coding is the battleground

Zhipu aimed GLM-5 straight at software engineering. At release it scored 77.8% on SWE-bench Verified, the highest result among open-weight models at the time and within roughly three points of Claude Opus 4.5 (80.9%). On Terminal-Bench 2.0 it posted 56.2%, and it led open models on agentic tests including BrowseComp and τ²-Bench. Zhipu also released Z Code, a programming tool that decomposes tasks, writes code, and debugs in a loop.

The efficiency story matters as much as the score. GLM-5 integrates DeepSeek Sparse Attention (DSA), which keeps long-context serving affordable, and was pretrained on about 28.5 trillion tokens using an asynchronous reinforcement-learning framework (Slime) for long-horizon alignment.

The public-company effect

In January 2026 Zhipu became the first Chinese AI software company to list, and the first globally focused specifically on AGI foundation models, when it floated on the Hong Kong Stock Exchange. The IPO was heavily oversubscribed. Shares rose 34% on the GLM-5 launch day and, by mid-February, had gained more than 500% from the offer price, lifting the market capitalization past US$40 billion.

In its first annual results as a listed company (fiscal 2025, released late March 2026), Zhipu reported revenue of RMB 724 million (≈ US$100M / HK$780M), up 132% year over year, alongside a widened net loss of RMB 4.72 billion (≈ US$650M / HK$5.1B) — figures from the company's filing as summarised by secondary trackers, not independently re-audited. The loss is the cost of a frontier lab spending on compute and talent; the growth is the open-weight strategy pulling in enterprise users.

How it stacks up

GLM-5 is not the biggest model in China — Baidu's ERNIE 5.0 and Alibaba's Qwen3-Max both scale larger on paper. But it is the one a competitor can legally download and run. That positions Zhipu less as a model vendor and more as the default open-weight backbone for enterprises that want sovereignty over their stack.

Honest limitations

Release date, 744B/40B parameter split, MIT licence, Hugging Face/ModelScope weights, 200K context, 128K max output, DSA, and the Huawei Ascend training claim come from Zhipu's official GLM-5 materials and corroborating coverage (AIbase, AIWiki, AICost). The 77.8% SWE-bench Verified, 56.2% Terminal-Bench 2.0, and Artificial Analysis Intelligence Index figures are Zhipu-reported or aggregator-reported scores; we found no fully independent third-party audit at writing and present them as company/publisher claims. The "100,000 Ascend 910B" and "28.5T tokens" figures appear in secondary write-ups of the launch and should be treated as vendor-flavoured. The RMB 724M revenue / RMB 4.72B loss are from the company's fiscal-2025 filing as summarised by trackers; currency conversions use ≈7.24 RMB/US$ and ≈7.8 HKD/US$. Stock-performance figures (34% day gain, 500% by mid-Feb, >US$40B cap) are from secondary market summaries and move with the tape.

What readers can do now

  1. If you need sovereign, self-hostable coding: pull the GLM-5 FP8 weights and benchmark agentic SWE tasks against your current model — the MIT licence lets you ship it commercially.
  2. If you run long-context agents: test the 200K window with DSA against your GPU memory budget before committing; cache and quantize to fit.
  3. If you track China's AI industry: read GLM-5 as a supply-chain signal, not just a model — a 744B training run on Huawei silicon is the more durable story.

Related coverage

More in “Foundation Models” → · Back to home · Markdown version