NeuroAI NEUROAINEUROAI.SITE
ESC

GLM-4.5: Zhipu open-sources a 355B model built for agents, not just chat

Zhipu AI released GLM-4.5, a 355-billion-parameter open-weight (开源权重) model tuned for intelligent agents (智能体), under an MIT licence on Hugging Face and ModelScope.

2026-10-01 · 754 words · NeuroAI
GLM-4.5: Zhipu open-sources a 355B model built for agents, not just chat

The GPU cluster at Zhipu's lab reportedly "smoked" the night GLM-4.5 launched — not from a fire, but from a stampede of users. When a Chinese lab drops a 355-billion-parameter model with an MIT licence, developers show up.

On the evening of 28 July 2025, Zhipu AI (智谱AI) released GLM-4.5, its new flagship foundation model (基础模型). Unlike a typical chatbot upgrade, the model was positioned from the start as "born for agents (为智能体而生)". The weights landed on Hugging Face and ModelScope the same night, free for commercial use and modification.

What GLM-4.5 actually is

GLM-4.5 uses a Mixture-of-Experts (混合专家, MoE) architecture. The headline version carries 355 billion total parameters with 32 billion active per token; a lighter GLM-4.5-Air ships 106 billion total with 12 billion active. Because only a slice of the network fires per request, the smaller active count keeps inference costs down while the large total pool preserves capability.

Both versions are hybrid reasoning models. They expose two modes:

  • Thinking mode — for complex reasoning and tool use (工具调用)
  • Non-thinking mode — for instant responses

That single-model, two-personality design means one deployment serves both quick Q&A and deep agent workflows, instead of running a separate reasoning model alongside a chat model.

Why "agent" matters here

Zhipu's pitch is that modern value is shifting from "answer my question" to "do my multi-step task." GLM-4.5 was tuned for exactly that:

  • Tool-use and function calling are first-class, not bolted on.
  • The model is built to chain searches, code execution, and API calls into a plan.
  • It ships with a parser compatible with popular serving stacks (vLLM, SGLang, Transformers).

Zhipu also made it drop-in compatible with the Claude Code framework through its BigModel.cn platform, so developers used to that tooling could point it at GLM-4.5 with minimal changes.

The numbers Zhipu published

On its own 12-benchmark evaluation, Zhipu reports GLM-4.5 scoring 63.2, which it places third among all proprietary and open models at release, with GLM-4.5-Air at 59.8. These are the lab's internal figures, not an independent leaderboard, so treat them as directional.

Pricing on the API was set low: input ¥0.8 per million tokens (≈ US$0.11) and output ¥2 per million tokens (≈ US$0.28), with a high-speed variant reaching up to 100 tokens per second. Free tiers are available on z.ai and chatglm.cn.

The bigger bet: agents as the product surface

The interesting move is not the size — China already has several hundred-billion-parameter open models. It is the positioning. Zhipu is arguing that the next competitive layer is not "who has the biggest model" but "whose model actually completes a workflow end to end." A model that can call a calculator, run a script, query a database, and report back is more useful to an enterprise than a smarter parrot.

GLM-4.5-Air exists for that exact reason: a 12-billion-active model that is cheap enough to embed in a product, yet capable enough to orchestrate tools. Zhipu is effectively selling two price points of the same agent brain.

Where it fits in China's open-model wave

GLM-4.5 joined a crowded 2025 field of Chinese open-weight (开源权重) releases — DeepSeek, Qwen, MiniMax, and others had already published large models under permissive licences. Zhipu's differentiation is the agent framing plus the Air variant for cost-sensitive deployments, and a 100-tokens-per-second high-speed tier for latency-sensitive products.

Honest limitations

A few caveats readers should hold in mind:

  • Performance scores (63.2, "third place") come from Zhipu's own benchmark suite and methodology; independent reproduction on public leaderboards was not verified for this article.
  • The "GPU smoked" line is a colourful quote from Chinese tech coverage, not a measured technical claim.
  • Exact training-data composition, training-compute spend, and safety-eval results were not disclosed in the sources reviewed.
  • This article does not compare GLM-4.5 head-to-head against models released after July 2025, so "state of the open art" is a moving target.
  • The MIT licence covers the published weights; enterprise support, larger context tiers, and any hosted extras are separate commercial terms.

What readers can do now

  1. Try it directly — pull GLM-4.5 or GLM-4.5-Air from Hugging Face (zai-org) or ModelScope and run a small agent that calls a search or calculator tool, to feel the think/non-think toggle.
  2. Benchmark on your own tasks — if you run agentic code-gen or tool-use workflows, run your private eval suite rather than trusting vendor numbers.
  3. Model the cost curve — at ≈ US$0.11–0.28 per million tokens, open-weight self-hosting may already beat closed APIs for high-volume internal agents; calculate your own total cost before committing.

Related coverage

More in “Foundation Models” → · Back to home · Markdown version