NeuroAI NEUROAINEUROAI.SITE
ESC

Alibaba's Qwen3-Max Crosses One Trillion Parameters — but Stays Behind a Paywall

Unveiled at Alibaba's Apsara Conference in September 2025, Qwen3-Max is the company's largest large model (大模型) at over one trillion parameters and 36T training tokens — but it ships as a hosted API, not open weights.

2026-10-03 · 782 words · NeuroAI
Alibaba's Qwen3-Max Crosses One Trillion Parameters — but Stays Behind a Paywall

A model bigger than almost anything China has shipped was unveiled to a packed conference hall — then handed to developers as a service, not a download. The same company gives away hundreds of smaller models for free. This one, it keeps.

What actually shipped

At Alibaba's Apsara Conference in late September 2025, the Tongyi (通义) team unveiled Qwen3-Max, which it calls its largest and most capable large model (大模型) to date. The Base version carries more than one trillion parameters and was pretrained on 36 trillion tokens, following the Qwen3 MoE design with a global-batch load-balancing loss. Alibaba describes a notably stable training run — no loss spikes, no rollbacks — and a 30% relative gain in training efficiency (MFU) over its prior Qwen2.5-Max base.

Qwen3-Max ships in two flavours: Qwen3-Max-Instruct for general use, and Qwen3-Max-Thinking for explicit reasoning. The Instruct model is available through Qwen Chat and Alibaba Cloud's Model Studio API; the Thinking variant was still in training at the time of the technical write-up but already showing extreme results.

The closed-flag contrast

Here is the part easy to miss: unlike the hundreds of Qwen models Alibaba has open-sourced, Qwen3-Max does not ship open weights. The official materials point developers to the hosted API and the Qwen Chat web interface. That is a deliberate split. Alibaba has built the most-downloaded open model family on earth — over 300 models, 600 million downloads, 170,000 derivatives — yet it is keeping the crown jewel behind a paywall.

At the same conference, Alibaba also showed Qwen3-VL and Qwen3-Omni, and reaffirmed a three-year, RMB 380 billion (≈ US$52.5B / HK$410B) commitment to AI and cloud infrastructure. The strategy is clear: open the ecosystem, charge for the apex.

What the benchmarks claim

Qwen3-Max-Instruct's preview ranked third on the LMArena text leaderboard, ahead of GPT-5-Chat. The formal release pushed coding and agentic scores higher: 69.6 on SWE-bench Verified and 74.8 on Tau2-Bench, the latter described by Alibaba as surpassing Claude Opus 4 and DeepSeek V3.1 on tool-calling.

The Thinking variant is the headline-grabber. Augmented with a code interpreter and parallel test-time compute, Qwen3-Max-Thinking scored 100% on AIME 25 and HMMT — two gruelling maths benchmarks — according to Alibaba. Those are vendor-reported numbers on tasks where tool use and extra compute swing the result, so read them as Alibaba's claim, not a verdict.

Why "just scale it" is the thesis

Alibaba's own blog title for the release is "Just Scale it." The argument is unfashionable in a year when everyone preaches efficiency: Qwen3-Max's gains came from simply making the model and its data bigger. The company leans on long-context training tricks — a ChunkFlow strategy said to deliver 3× throughput over context parallelism, enabling a 1-million-token training context — and on fault-tolerance fixes that cut hardware-failure downtime to a fifth of the prior generation.

Whether scaling alone still buys frontier performance is the live debate. Qwen3-Max's answer is a confident yes; the rest of the field is betting on reasoning and agents instead.

Pricing and access

Alibaba lists tiered API pricing for Qwen3-Max that rises with context length. For the shortest tier (≤32K tokens) reported rates run about US$1.20 per million input tokens and US$6.00 per million output tokens, scaling to US$3.00 / US$15.00 for 128K+ prompts. Pricing differs by region and deployment (international vs mainland China), so a production budget needs the console figure for your zone.

Honest limitations

Parameter count (">1 trillion"), 36T training tokens, MoE design, LMArena ranking, SWE-bench 69.6, Tau2-Bench 74.8, and the AIME 25 / HMMT 100% Thinking scores all come from Alibaba's official Qwen blog and the Apsara Conference coverage (Xinhua, Alibaba News). The "2.4 trillion parameters" figure that appears in some secondary outlets is not supported by Alibaba's primary materials, which state "over 1 trillion"; we used the primary figure. Qwen3-Max weights are not open per the official release — only API/Qwen Chat access is documented. The US$1.20 / US$6.00 pricing is the ≤32K international tier as reported by model aggregators; longer-context and domestic tiers differ. The RMB 380 billion infrastructure commitment is from Alibaba's official newsroom. All benchmark numbers are vendor-reported and lack independent third-party audit at writing.

What readers can do now

  1. If you want the strongest Qwen: call qwen3-max through Alibaba Cloud Model Studio or Qwen Chat — there is no weight download, so plan for API dependency.
  2. If you need open weights: stay on Qwen's open releases (Qwen3-235B and smaller) and treat Max as a capability ceiling to compare against, not a model you can self-host.
  3. If you benchmark: wait for independent SWE-bench / Tau2 numbers before trusting the 69.6 / 74.8 claims in procurement decisions.

Related coverage

More in “Foundation Models” → · Back to home · Markdown version