Most long-context models quietly extrapolate a 128K or 200K window and hope it holds. MiniMax (稀宇科技) took a different route: it rebuilt the attention mechanism so a 4-million-token context is the design, not the stretch goal.
In January 2025, MiniMax released and open-sourced the MiniMax-01 series — MiniMax-Text-01 for language and MiniMax-VL-01 for vision. The language model is a mixture-of-experts (MoE) system with 456 billion total parameters and 45.9 billion activated per token, putting it in the same weight class as the largest open models of its moment.
The attention trick
The headline is the architecture. Inside every 8 layers, 7 use Lightning Attention — a linear-attention mechanism — and 1 uses traditional softmax attention. This is described as the first time linear attention was scaled to a commercial-grade model. The payoff is an approach to linear complexity on long inputs, which is what makes the extreme context window affordable.
The numbers:
- Training context: 1 million tokens.
- Inference context: up to 4 million tokens — roughly 20 to 32 times longer than leading peers at the time.
- Architecture: 80 layers, 32 experts, top-2 routing, rotary position encoding.
In MiniMax's own "Needle-In-A-Haystack" test at 4 million tokens, the model reported 100 percent retrieval accuracy. That is a vendor benchmark, but it points at the real goal: a model that does not collapse as the document grows.
Why long context is the agent thesis
MiniMax's explicit motivation is the AI agent (智能体) era. Whether it is sustained memory inside a single agent or the back-and-forth communication between many agents, increasingly long contexts are treated as foundational. A model that degrades slowly as input length grows is, in this view, the substrate agents will be built on — which is why MiniMax open-sourced the weights on GitHub and Hugging Face rather than keeping them behind an API wall.
Performance positioning
On standard academic benchmarks, MiniMax reported MiniMax-Text-01 as matching top-tier models such as GPT-4o and Claude-3.5-Sonnet, while leading on long-context evaluations. Those comparisons are the company's own evaluations and should be read as positioning, not as independently audited superiority. The accompanying API was priced at US$0.2 per million input tokens and US$1.1 per million output tokens — a deliberately aggressive number tied to the efficiency the architecture claims.
Where it sits among Chinese open models
MiniMax-Text-01 is a useful data point in the broader open-model story: it is not chasing raw parameter count, but architectural efficiency for a specific workload. That distinguishes it from reasoning-first releases and from the agentic, tool-using models that optimize for action over context.
The engineering trade-off most coverage skipped
Linear attention is not free. Replacing softmax attention with a linear variant changes what the model can express per layer, and MiniMax's own architecture is a hedge rather than a conversion: seven linear layers to every one softmax layer, keeping some full-attention capacity in the mix. That ratio is the honest summary of the whole field's current judgment — linear attention wins on cost at extreme lengths but is not yet trusted alone for the hardest reasoning steps. For builders, the practical implication is that "supports 4M tokens" and "reasons well over 4M tokens" are different claims, and only the second one matters for real workloads. Test with documents your users actually bring, at lengths your users actually hit.
Honest limitations
Facts (January 2025 release and open-sourcing of MiniMax-01 series; 456B total / 45.9B active parameters; hybrid Lightning + softmax attention with 7-of-8 linear layers; 1M trained / 4M inference context; 80 layers, 32 experts, top-2 routing; GitHub and Hugging Face availability; API pricing US$0.2 / US$1.1 per million tokens; 100 percent 4M Needle-In-A-Haystack per the company) come from MiniMax's official news post, its arXiv technical report (2501.08313) and the Hugging Face model card, cross-checked against the GitHub README. Performance claims versus GPT-4o and Claude-3.5-Sonnet are MiniMax's own evaluations and are not independently verified here. The "first commercial-scale linear attention" claim is MiniMax's and is presented as the company's assertion. Real-world latency and cost at 4M tokens depend heavily on deployment and are not benchmarked in this article. No investment advice is given.
What readers can do now
- If you build long-document or multi-agent workflows, test MiniMax-Text-01 where context length is the bottleneck — legal, scientific or code-repository tasks that break other models.
- If you choose open weights, pull it from GitHub or Hugging Face and validate the 4M-token behavior on your own retrieval tasks rather than trusting the vendor benchmark.
- If you follow architecture trends, treat Lightning Attention as a serious alternative to pure softmax transformers — the 4M context is the proof point that linear attention has left the research lab.
