Most of the AI world spent 2024 arguing about who had the biggest model. A Shanghai lab quietly asked a different question: what if the model could just read everything at once? On 15 January 2025, MiniMax open-sourced an answer.
The release was not a bigger parameter count for its own sake. It was a structural bet that the next useful leap is not raw size but how much a model can hold in its head in a single pass.
The bet: context over raw size
MiniMax-01 ships as two models: MiniMax-Text-01, a foundation language model, and MiniMax-VL-01, a vision–language model built by continued training on top of it. Both are open source under the permissive Apache 2.0 license, with full weights published on GitHub and Hugging Face.
The headline number is the context window. Where most leading models in early 2025 topped out at 128,000 to 200,000 tokens, MiniMax-Text-01 was built to handle up to 4 million tokens at inference (a 1-million-token window during training). In the company's framing, that is 20 to 32 times the effective context of models like GPT-4o.
The point is not a benchmark party trick. MiniMax's explicit thesis is that 2025 would be the year of the AI agent (智能体), and agents need sustained memory — both to remember a single long task and to coordinate across many agents talking to each other. A long context is the substrate that makes that possible.
What Lightning Attention actually changes
The reason this matters technically is the architecture underneath. Standard Transformers use softmax attention, whose cost grows quadratically with sequence length. Double the text and you more than double the compute. That is why long context has historically been so expensive.
MiniMax's fix is Lightning Attention (闪电注意力), a linear-attention mechanism. In the MiniMax-01 design, 7 of every 8 layers use Lightning Attention and the 8th keeps a traditional softmax layer. The company says this is the first time linear attention has been scaled to a commercial-grade, hundreds-of-billions-parameter model rather than a research demo.
The payoff is near-linear cost as input grows. On a long "needle in a haystack" retrieval test, MiniMax reports 100% accuracy at 4 million tokens, and on its own real-world assistant test set the model degrades far less than comparable models as the input stretches.
The numbers that matter
- 456 billion total parameters, of which 45.9 billion activate per token — a Mixture-of-Experts (混合专家, MoE) design with 32 experts.
- 4 million tokens max context at inference; 1 million during training.
- Released open source under Apache 2.0 on 15 January 2025.
- API priced in US dollars: US$0.20 per million input tokens and US$1.10 per million output tokens — among the cheapest frontier-tier APIs at launch.
The model already powers MiniMax's own consumer product, Hailuo AI (海螺AI), which is live globally.
The sequel came fast
MiniMax did not stop at the foundation release. On 16 June 2025 it open-sourced MiniMax M1, a reasoning model built on the same hybrid architecture with a 1-million-token window and an 80,000-token "thinking budget." M1 was positioned as the first open-weight, large-scale hybrid-attention reasoning model.
One detail from that release is unusually concrete. MiniMax said its entire reinforcement-learning phase ran on 512 H800 GPUs for three weeks at a rental cost of US$534,700 — an order of magnitude cheaper than the company initially expected, thanks to a custom RL algorithm called CISPO. For builders, that is a real data point on how cheap frontier-grade training can get when the architecture is efficient.
Where it still falls short
Two caveats are worth stating plainly.
First, the context-length claim is MiniMax's own, measured on its own tests. Independent, third-party long-context benchmarks of the open weights exist, but the 4-million-token figure describes the model's designed capacity, not a universally agreed external measurement. The weights are public, so anyone can verify — but most users will take the vendor's word.
Second, linear attention is still a young bet. The software and ecosystem around it are less battle-tested than the standard Transformer stack, and running 4-million-token prompts at scale means real infrastructure and cost questions that a spec sheet does not answer.
What readers can do now
- If you build agents, treat a 1M–4M-token window as a design option, not a curiosity. Test MiniMax-01 or M1 on a task that currently forces you to chunk documents — full-codebase reasoning, long contracts, multi-document research — and measure whether the unchunked version actually performs better.
- If you self-host open weights, the Apache 2.0 license is permissive; start from the Hugging Face checkpoints and benchmark inference cost at your target context length before committing.
- If you follow the architecture race, watch linear attention closely. If MiniMax's scaling holds up, the "quadratic attention tax" that has shaped model design for years may finally have a credible alternative.
Honest limitations
Core facts (open-source release of MiniMax-01 on 15 January 2025; two models MiniMax-Text-01 and MiniMax-VL-01; 456B total / 45.9B activated parameters; 32-expert MoE; Lightning Attention with 7-of-8 linear layers; 4M-token inference / 1M-token training context; Apache 2.0 license; API pricing US$0.20 in / US$1.10 out per million tokens; Hailuo AI as the product; MiniMax M1 open-sourced 16 June 2025 with 1M context and 80K reasoning budget; RL training cost US$534,700 on 512 H800s for three weeks; CISPO algorithm) come from MiniMax's own disclosures on minimax.io and the MiniMax-01 technical materials, corroborated by Shanghai Securities News (中国证券网) and the model's public GitHub/Hugging Face repositories. MiniMax is a Shanghai-based company; the 4-million-token context and benchmark results are vendor-reported and have not been independently re-measured by a neutral lab at full length. All prices are stated by MiniMax in US dollars, so no RMB conversion applies. Figures are current to 28 September 2026 and this article is not investment advice.
