NeuroAI NEUROAINEUROAI.SITE
ESC

Alibaba open-sourced its biggest model ever — a 2.4-trillion-parameter (大模型) you can download

On 12 August 2026, Alibaba's Qwen team (通义千问) published Qwen3.8-2.4T-A95B as open weights on Hugging Face — a 2.4-trillion-parameter mixture-of-experts large model (大模型) and the first time Alibaba has open-sourced a Max-class flagship.

2026-09-30 · 744 words · NeuroAI
Alibaba open-sourced its biggest model ever — a 2.4-trillion-parameter (大模型) you can download

The biggest model a company makes is usually the one it keeps behind a paid API. Alibaba just did the opposite, and the download is larger than most data centers can stomach.

On 12 August 2026, Alibaba's Qwen team (通义千问) published Qwen3.8-2.4T-A95B as open weights on Hugging Face (and mirrored on ModelScope). The model card states plainly that this is "the first time Qwen3.8 brings a Qwen-Max-class model to open release." The hosted Qwen3.8-Max API had gone live on 3 August; the weights followed nine days later.

How big, and how it works

The numbers are the story:

  • 2.4 trillion total parameters, with 95 billion activated per token — a sparse mixture-of-experts design.
  • 512 experts, of which 11 fire per token (10 routed + 1 shared).
  • Native context of 262,144 tokens, extensible to about 1.01 million.
  • Hybrid architecture mixing full attention and linear attention (Gated DeltaNet) to keep compute and memory bounded as context grows.

Multiple model trackers describe it as the largest open-weight language model published to date. The release matters because "open weights" at this scale changes who can study, audit, and modify a frontier-class checkpoint — not just rent it.

The license is the catch

This is not a standard Apache release. The open weights ship under a custom Qwen3.8-Max License, not the Apache 2.0 terms Alibaba has used for smaller Qwen models. The thresholds matter for businesses:

  • Above 100 million monthly active users or US$20 million in monthly revenue, you must prominently display the model name.
  • Model-as-a-Service or "AI work assistant" businesses above US$50 million in trailing-twelve-month revenue need a separate commercial license from Alibaba.

Alibaba did pair the Max-class drop with a smaller sibling, Qwen3.8-27B, released under Apache 2.0 — a single-GPU-friendly model for teams that want open terms and local deployment without the custom-license strings.

What you actually get — and what you don't

The open checkpoint is text-only. The hosted Qwen3.8-Max adds vision input, optional non-thinking mode, a 1-million-token context by default, and built-in tools; the download strips those. The released weights also require "thinking" to be on for every response, so it reasons before it answers.

Running it is the other catch. At full BF16 precision the repository is about 4.89 TB; community quantizations bring that down, with Unsloth's 1-bit build reported near 397 GB. Even the practical floor still calls for well-provisioned servers or workstations — essentially no consumer GPU can serve it at full precision. Hosted API pricing, for comparison, runs about US$2 per million input tokens and US$6 per million output tokens.

Why open a flagship at all

Alibaba's move fits a wider Chinese pattern: the largest labs are using open weights as a distribution and ecosystem strategy, not just a research gesture. Releasing the Max-class checkpoint invites universities, enterprises, and tooling developers to build on top of Qwen rather than a competitor's stack — and it pressures closed rivals on the "can I see inside?" question.

For global readers, the practical upside is optionality: audit the model, avoid vendor lock-in, deploy privately, and modify the weights. The downside is the hardware and license friction that keeps the "open" in open weights from meaning "free to run anywhere."

Honest limitations

  • Parameter counts, expert counts, and license terms are drawn from the official Hugging Face model card and Alibaba materials, which are strong primary sources — but "largest open-weight model to date" depends on how one defines "open-weight," and the claim should be re-checked as the field moves.
  • The open release is a subset of the hosted product: text-only, thinking-required, smaller native context than the API's default 1M. Benchmark comparisons to closed models are vendor-reported.
  • Full-precision serving needs data-center hardware; reported quantized sizes (397 GB and up) carry accuracy trade-offs that independent evaluation had not fully measured at the time of this review.
  • The custom license adds commercial restrictions absent from Alibaba's Apache-released Qwen models — read it before embedding the model in a paid product.

What readers can do now

  1. Pull Qwen3.8-2.4T-A95B from Hugging Face or ModelScope if you have data-center GPUs and want to study or fine-tune a frontier-scale checkpoint.
  2. Use the Apache-2.0 Qwen3.8-27B for single-GPU local work when you need open terms and a model that fits on accessible hardware.
  3. If you are a business, check the custom-license thresholds (100M MAU / US$20M monthly revenue / US$50M MaaS revenue) before shipping a product built on the Max-class weights.

Related coverage

More in “Foundation Models” → · Back to home · Markdown version