A model that can write code is useful. A model that can rewrite the system that trains it is something else — and a Chinese lab just put that claim on the table.
On 18 March 2026, Shanghai-based MiniMax (稀宇科技) released M2.7, which it describes as the first public demonstration of a "model self-evolution" path. The pitch is not raw benchmark scores. It is that the model can construct its own agent (智能体) scaffolding — what MiniMax calls an Agent Harness — and then use that scaffold to improve itself.
What "self-evolution" actually means here
MiniMax is careful to define the term narrowly. Instead of a model that autonomously rewrites its own weights in the open, M2.7 is positioned as a research agent that participates in MiniMax's own development pipeline: data preparation, experiment design, training tuning, and evaluation feedback.
The concrete mechanism is an Agent Harness the model assembles itself. In MiniMax's telling, earlier M2-series models were steered into a research-agent role that collaborates with human teams across data pipelines, training environments, and evaluation systems. For reinforcement-learning work, the agent can propose an experiment, debate it with researchers, then autonomously do log analysis, bug tracing, metric optimization, and code fixes — lowering how often a human has to step in.
The numbers MiniMax reports
The company's own figures, carried by Xinhua and Xinhua Finance, are specific:
- In some R&D workflows, M2.7 can take on 30%–50% of the workload.
- In internal tests, it ran more than 100 "analyze – improve – verify" cycles, autonomously adjusting sampling parameters and workflow strategy.
- On an internal evaluation set, that loop produced roughly a 30% improvement.
- On software-engineering benchmarks MiniMax cites, M2.7 scored 56.22% on SWE-bench Pro, 55.6% on VIBE-Pro, and 57.0% on Terminal Bench 2.
In production-engineering scenarios, MiniMax says, the model does more than generate code: it reads monitoring metrics and deployment timelines, performs causal analysis, connects to a database to test a hypothesis, and proposes a fix. The company claims some online incident recoveries were cut to under three minutes.
Where M2.7 sits in MiniMax's agent line
M2.7 is the latest step in a deliberate agent (智能体) strategy, not a pivot. MiniMax's M2.5, released on 12–13 February 2026, was positioned as a "native agent production-grade model" — strong at coding, with tool-calling and search built in, and open-sourced for local deployment a day after launch. MiniMax reported M2.5 at 80.2% on SWE-Bench Verified and said API call volume crossed 3 trillion tokens within a week of release.
The throughline: M2.5 was an "executing agent." M2.7 is meant to be an "evolving agent" — one that does not just run tasks but helps build the next version of itself. MiniMax has framed the broader shift as AI competition moving from "model capability" to "execution-system capability."
Why this matters for buyers
For teams evaluating Chinese models, the practical question is whether agent-native design changes the cost of getting work done. MiniMax's argument is that when the model can plan architecture, call tools, and self-correct inside a long workflow, the unit cost of a completed task drops faster than raw intelligence alone would suggest.
The company has also stressed multi-agent collaboration (Agent Teams), where the model plays several roles for adversarial reasoning, and complex Skills with a reported 97% instruction-following rate on office tasks. In finance, MiniMax says M2.7 can read annual reports, synthesize research, build a revenue model, and produce a presentation — behaving, in its words, like a junior analyst that corrects itself across turns.
Honest limitations
- The "self-evolution," 30%–50% workload, 100-cycle loop, and ~30% improvement figures are MiniMax's own reported numbers from internal evaluations, not independently audited results. Treat them as directional evidence of a capability, not as measured productivity gains.
- "Self-evolving" here means the model assists its own R&D pipeline via agent scaffolding — not open-ended autonomous self-improvement. The language is easy to over-read.
- Benchmark scores (SWE-bench Pro, VIBE-Pro, Terminal Bench 2) are vendor-reported and use mixed harnesses; independent reproduction was not available at the time of this review.
- Like all agent models, real-world reliability depends heavily on the tools, permissions, and guardrails you give it.
What readers can do now
- Test M2.7 on MiniMax Agent and the open platform on a real coding or office workflow before trusting any benchmark claim.
- Compare agent reliability head-to-head with Claude, GPT, and other Chinese agents on your own tasks — tool-use consistency over long runs is what separates demos from deployments.
- Watch the cost curve, not just the score — the thesis is that self-improvement lowers MiniMax's own training cost over time; track whether that shows up in API pricing.
