---
title: "Zhipu open-sourced GLM-4.5 with 355B parameters — and gave agents a thinking mode"
date: 2026-09-28
category: Foundation Models
site: NeuroAI
canonical: https://neuroai.site/a/na-model-zhipu-glm
language: en
---

# Zhipu open-sourced GLM-4.5 with 355B parameters — and gave agents a thinking mode

> On 28 July 2025, Beijing-based Zhipu AI (智谱AI) released GLM-4.5 as a fully open-weights model under the permissive MIT license, available on Hugging Face and ModelScope. The flagship carries 355 billion total parameters (32 billion active) in a mixture-of-experts design and is built specifically for agentic use — with a switchable "thinking" mode for hard tasks and a fast mode for instant replies.

A Chinese lab quietly did something the big American closed models usually avoid: it handed the weights of a frontier-scale model to the public, for free, with almost no strings attached. You can download GLM-4.5, run it, fork it, or build a product on it — and the license will not come after you.

For developers outside China, that is the part worth understanding. The model is not a trimmed "community" edition; it is the real thing.

## What actually shipped

Zhipu AI (智谱AI), the Beijing company behind the GLM (General Language Model / 大模型) family, released **GLM-4.5** on **28 July 2025**. Two variants:

- **GLM-4.5** — **355 billion total parameters, 32 billion active**, a mixture-of-experts (MoE) design.

- **GLM-4.5-Air** — **106 billion total, 12 billion active**, a lighter sibling that still lands near far larger proprietary models on reasoning benchmarks.

Both are **open-weights under the MIT License** and published on **Hugging Face and ModelScope**. MIT is about as permissive as open-source licenses get — commercial use included, attribution aside.

## The "agent" framing is the point

Zhipu did not market this as a chatbot. It positioned GLM-4.5 as a **foundation model for intelligent agents (智能体)** — software that plans, calls tools, and acts over many steps rather than answering one prompt. The design choices follow that goal:

- **Hybrid reasoning.** A single set of weights switches between a **"thinking" mode** for complex reasoning and tool use, and a **"non-thinking" mode** for instant responses. You do not need two models; you flip a switch.

- **128k context window** and **native function calling**, the plumbing agents need to hold long state and invoke external tools.

- A reported **100 tokens/second** on the high-speed API tier.

Zhipu says the model was trained on roughly **15 trillion tokens** of general pre-training plus **8 trillion tokens** of targeted training in code, reasoning, and agentic tasks, then sharpened with reinforcement learning.

## Where it sits on the leaderboard

On a 12-benchmark composite covering reasoning, coding, and agentic tasks, Zhipu reported GLM-4.5 at **3rd place globally** — behind only the very top proprietary models — and **first among open-source models** at the time of release. On agentic benchmarks (τ-bench, Berkeley Function Calling Leaderboard v3) it claimed parity with Claude 4 Sonnet, and on the web-browsing BrowseComp test it reported **26.4%**, ahead of Claude-4-Opus (18.8%).

Those are the company's own benchmark numbers, and they should be read as marketing-adjacent until independently reproduced. But the architecture claim is concrete and verifiable: a 355B MoE that activates only 32B per token is a genuine efficiency story, letting a capable model run on far less hardware than its parameter count suggests.

## The price, converted

Zhipu priced the API aggressively: **input at 0.8 RMB per million tokens, output at 2 RMB per million tokens** — roughly **US$0.11 / HK$0.87** to send a million input tokens and **US$0.28 / HK$2.18** for a million output tokens. The company also made it compatible with the **Claude Code** agent framework on its BigModel.cn platform, a deliberate move to let developers port agent workflows with minimal rewiring.

## Why an open-weight Chinese model matters

Most of the world's most capable models are behind APIs you cannot inspect. Zhipu's choice to release GLM-4.5 with MIT terms puts a large-model (大模型) capability in the hands of anyone who can host it — researchers in countries without API access, companies wary of depending on a single foreign provider, and regulators who want to audit what they deploy.

It also reflects a broader Chinese strategy: open-weight releases as a way to build developer ecosystems and standards, not just products. Zhipu, spun out of Tsinghua University, is one of the labs pursuing that path most aggressively.

## What readers can do now

- **If you build agents**, pull GLM-4.5 from Hugging Face or ModelScope and test the thinking/non-thinking switch on a real multi-step task; the 32B-active MoE is light enough to self-host on modest clusters.

- **If you are locked to a US API**, treat GLM-4.5 as a drop-in option for cost-sensitive or compliance-sensitive workloads, especially given Claude Code compatibility.

- **If you follow the open-model race**, watch Zhipu's benchmark claims against independent reruns — the "3rd globally / #1 open" line is the vendor's, and the real test is community reproduction.

## Honest limitations

Verified facts: Zhipu AI (智谱AI) released GLM-4.5 on 28 July 2025, open-weights on Hugging Face and ModelScope under the MIT License, GLM-4.5 at 355B total / 32B active and GLM-4.5-Air at 106B total / 12B active, hybrid thinking/non-thinking modes, 128k context, native function calling, and API pricing of 0.8 RMB / 2 RMB per million tokens — confirmed by Zhipu's official z.ai blog and corroborated by Shanghai Securities News (上证报), Beijing News (新京报), and the Beijing Mentougou government portal (30 July 2025). The training-data scale (15T general + 8T targeted) is from the Beijing Mentougou and Southern Metropolis Daily coverage. The "3rd globally / #1 open-source / parity with Claude 4 Sonnet / 26.4% BrowseComp" performance claims are Zhipu's own benchmark results reported by those outlets and are not independently re-verified in this article; treat them as vendor claims pending community reproduction. Currency converted at ~7.1 RMB/USD and ~1.09 HKD/RMB: input ≈ US$0.11 / HK$0.87 per million tokens, output ≈ US$0.28 / HK$2.18. No political figures or stock-touting asserted. Analysis is current to 28 September 2026 and is not investment advice.

---

Published by NeuroAI (https://neuroai.site/) — https://neuroai.site/a/na-model-zhipu-glm
Free to quote with attribution and a link to the original.
