Did a $0.30-Per-Million-Token Chinese Model Just Catch Up to Claude Opus in Coding?
MiniMax M2.5 scores within 0.6 points of Claude Opus 4.6 on SWE-Bench Verified and beats it outright on multi-step agent tasks — at roughly a tenth of the price. It also arrives wrapped in a distillation controversy Anthropic says is real.
- What M2.5 Actually Is
- The Number That Flips the Table: 80.2% on SWE-Bench
- The Bigger Surprise: Agentic Task Dominance
- A Price That Changes Agent Economics
- Strengths at a Glance
- The Dark Side: Hallucinations, Math, and a Distillation Fight
- When to Use M2.5 — and When to Avoid It
- Expert Editorial Opinion
- Final Verdict
- FAQ
MiniMax M2.5 review 2026: until early this year, the pecking order in frontier coding models felt settled — Claude, GPT, and Gemini traded the top spots, and Chinese labs competed mostly on price and "good enough" performance a tier below. On February 12, 2026, Shanghai-based MiniMax released M2.5 and that assumption got harder to defend, at least in one specific lane.
This review breaks down what M2.5 actually scores on the benchmarks that matter for real coding and agent work, where it genuinely leads, where it clearly still trails the Western frontier, and the distillation controversy that Anthropic has publicly attached to MiniMax's name. It's built from MiniMax's own model documentation, independent benchmark trackers, and verified reporting — not a claim of hands-on testing of M2.5 itself unless stated otherwise.
What M2.5 Actually Is
Third Generation, Real Funding Behind It
MiniMax was founded in Shanghai in 2021 by researchers with SenseTime backgrounds, raised $619 million in a Hong Kong IPO in January 2026, and is backed by Alibaba and Tencent. M2.5 is the third model in its line, following M2 and M2.1.
230B Mixture-of-Experts, 10B Active
M2.5 uses a Mixture-of-Experts architecture with 230 billion total parameters and roughly 10 billion active at inference — a design choice that keeps inference cost down while retaining a large total knowledge base.
~200K Token Context, Text-Only
The context window sits around 200,000 tokens depending on the provider. There is no image, screenshot, or multimodal input support — this is a text-only model, which matters for visual debugging workflows.
Open Weights, Commercial License
Weights are published on Hugging Face under a permissive MIT-family license that allows commercial use and self-hosting — a real point of differentiation from Claude, GPT, or Gemini, none of which publish weights at all.
The Number That Flips the Table: 80.2% on SWE-Bench
SWE-Bench Verified is the benchmark most worth trusting for real coding ability: it grades a model on genuine GitHub issues across production repositories, with terminal access and the ability to actually edit files — not a curated toy problem set. M2.5 scored 80.2%, putting it just 0.6 points behind Claude Opus 4.6 and ahead of both GPT-5.2 and Gemini 3 Pro on this specific measure. Independent trackers, including Artificial Analysis and multiple benchmark write-ups published after launch, corroborate the figure. It's also the highest score recorded by any open-weight model on this benchmark as of its release.
| Benchmark | M2.5 | Opus 4.6 | GPT-5.2 | Gemini 3 Pro |
|---|---|---|---|---|
| SWE-Bench Verified | 80.2% | 80.8% | 80% | 78% |
| Multi-SWE-Bench | 51.3% | 50.3% | — | 42.7% |
| BFCL Multi-Turn (agent) | 76.8% | 63.3% | — | 61% |
| BrowseComp (w/ context mgmt) | 76.3% | 84% | 65.8% | 59.2% |
| AIME 2025 (math) | 86.3 | 95.6 | 98.0 | 96.0 |
| GPQA Diamond | 85.2 | 90.0 | — | — |
| AA Intelligence Index | 42 | 53 | — | 57 |
Read the whole table, not just the headline row, and the real shape of the story appears: M2.5 is essentially tied at the top for coding and multilingual coding, but it trails by 9-12 points on math and general reasoning. This isn't a model that beats the Western frontier across the board — it's one that closed a very specific, very commercially valuable gap.
The Bigger Surprise: Agentic Task Dominance
The coding parity is the headline, but the more consequential number for anyone building AI agents is BFCL Multi-Turn — a benchmark for multi-step function calling across complex API schemas, which is exactly where real agent pipelines tend to break. M2.5 scored 76.8% here, a full 13.5 points ahead of Opus 4.6's 63.3% and 15.8 points ahead of Gemini 3 Pro. That's not a marginal edge; it's the kind of gap that shows up directly as fewer failed tool calls and fewer wasted retries in a long-running agent loop.
— OpenHands' independent evaluation, describing M2.5's behavior on extended agent runs.
MiniMax backs this up with two supporting figures worth noting: M2.5 completes agentic tasks using roughly 20% fewer rounds than competing models, and it finishes SWE-Bench tasks 37% faster than its own predecessor, M2.1. Together, that points to a model specifically trained — via MiniMax's Forge reinforcement-learning framework across more than 200,000 simulated environments — for the exact failure mode that makes production agents expensive: losing context and burning tokens on retries.
A Price That Changes Agent Economics
This is the number that turns M2.5 from "an interesting open-weight model" into something enterprises actually have to evaluate. The Lightning variant runs at $2.40 per million output tokens, against roughly $75 for Opus 4.6, $60 for GPT-5.2, and $20 for Gemini 3 Pro — a 10 to 20x gap, not a marginal discount.
| Model | Input $/M tokens | Output $/M tokens | Est. daily cost* |
|---|---|---|---|
| M2.5 Standard | $0.15 | $1.20 | ~$4.70/day |
| M2.5 Lightning | $0.30 | $2.40 | ~$1/hour continuous |
| Claude Opus 4.6 | $5.00 | $25.00 | ~$100/day |
| Gemini 3.1 Pro | $2.00 | $8.00 | ~$36/day |
*Estimated for 1M input + 2M output tokens/day. Rates vary slightly across hosting providers (OpenRouter, Bedrock, MiniMax's own API); figures above reflect MiniMax's published first-party pricing.
At roughly $1 per hour of continuous 24/7 agent operation, M2.5 makes always-on agentic workflows economically viable in a way the frontier Western models simply don't at their current pricing — that's the practical, bottom-line consequence of the benchmark table above.
Strengths at a Glance
The Dark Side: Hallucinations, Math, and a Distillation Fight
Hallucination rate spiked. Artificial Analysis measured an 88% hallucination rate for M2.5, up sharply from 67% for M2.1, with the AA-Omniscience reliability score dropping from -30 to -41. That's a meaningful step backward in factual grounding — less damaging for pure code generation, where output either compiles and passes tests or it doesn't, but a real liability for any task involving factual claims or information retrieval.
The math and reasoning gap is real and MiniMax doesn't hide it. On AIME 2025, M2.5 scored 86.3 against 95.6 for Opus, 98.0 for GPT-5.2, and 96.0 for Gemini 3 Pro — a 9-12 point gap that repeats on the Artificial Analysis Intelligence Index (42 vs. 53-57). This isn't a model optimized for hard math, and it isn't marketed as one; anyone benchmarking it against frontier reasoning specifically will come away disappointed.
The distillation controversy. In February 2026, Anthropic published findings accusing DeepSeek, Moonshot AI, and MiniMax of running a coordinated "distillation attack" against Claude — using roughly 24,000 fraudulent accounts to generate more than 16 million exchanges with Claude, with MiniMax alone responsible for over 13 million of those exchanges. Anthropic further alleged that MiniMax redirected close to half its traffic within 24 hours of any new Anthropic model release, specifically to harvest capabilities from the newest system. MiniMax has not publicly confirmed or denied the specifics. This doesn't invalidate M2.5 as a product — the benchmarks above are independently reproducible regardless of how the model was trained — but it's a legitimate reason to read direct comparisons against Claude specifically with some added skepticism, and it's part of a wider pattern Anthropic has since documented across multiple China-based labs.
Practical limits worth knowing before you deploy it. Beyond text-only input and the verbosity issue above, the raw model weights are roughly 457GB uncompressed — self-hosting requires a serious multi-GPU setup, not a hobbyist machine.
When to Use M2.5 — and When to Avoid It
✅ Use it when...
you're running high-volume coding agents — multi-file refactors, SWE-Bench-style tasks, long tool-calling loops — where cost is a real constraint or you want genuinely affordable 24/7 agent operation, or you specifically want open weights for self-hosting and fine-tuning.
❌ Avoid it when...
you need high factual reliability or information retrieval accuracy (88% hallucination rate is disqualifying for that), the task centers on hard math or deep scientific reasoning, or you need multimodal input like screenshots, diagrams, or visual debugging.
Expert Editorial Opinion
The honest framing for M2.5 isn't "China caught up to the West in AI" — that headline is both too broad and too generous. The more accurate claim is narrower and, in some ways, more interesting: a Chinese open-weight lab matched the closed Western frontier in one specific, commercially critical lane — real-world coding and agentic tool use — while still trailing clearly in deep reasoning and factual reliability. Precision matters here because the broad version of the claim gets disproven the moment someone runs an AIME math benchmark, and the narrow version is the one actually worth taking seriously.
What makes the agentic number more convincing than the coding number, to me, is that it's harder to game. SWE-Bench has been optimized against by nearly every major lab for over a year now; a 0.6-point gap at the top is close enough that scaffold and harness differences could easily explain it. A 13.5-point lead on multi-turn tool calling is a bigger, structurally different kind of result, and it lines up with MiniMax's own stated training approach — reinforcement learning across 200,000+ simulated environments specifically targeting long-horizon tool use. That's a testable mechanism behind the number, not just a benchmark that happened to land favorably.
The hallucination jump deserves more attention than it's gotten in most coverage I've seen. Going from 67% to 88% between M2.1 and M2.5 is not a rounding error — it suggests MiniMax traded some factual grounding for agentic and coding gains, which is a real, specific tradeoff, not a vague "still improving" caveat. For code generation, where correctness is externally verifiable by running the code, that tradeoff is survivable. For anything involving research, fact-checking, or information synthesis, it's a serious limitation that should rule the model out entirely.
The distillation allegations are the part every review has to handle carefully, because the two things people want to conflate — "is the model good" and "how was the model built" — are genuinely separate questions with genuinely separate answers. Anthropic's findings are specific and dated (24,000 fraudulent accounts, 16+ million exchanges, MiniMax accounting for the majority of it), and MiniMax hasn't offered a public rebuttal. That's a legitimate shadow over any head-to-head marketing MiniMax does against Claude specifically. It does not, however, change what SWE-Bench Verified or BFCL actually measure — those are external, reproducible evaluations, not MiniMax's own claims.
Put practically: M2.5 is a genuinely strong, genuinely cheap coding and agent model that earns real consideration for high-volume production agent work — used as a fast, low-cost drafting layer with a more reliable model verifying anything that touches facts or hard reasoning, exactly as MiniMax's own positioning suggests. Treating it as a like-for-like Claude replacement across the board would be a mistake; treating it as proof that the coding/agent gap between Chinese and Western frontier models has meaningfully narrowed is a defensible, evidence-backed conclusion.
Final Verdict
Score band: 7.0-7.9 — Competent but compromised: a strong fit for a specific, well-defined use case (high-volume coding agents), not a general-purpose replacement for the Western frontier.
| Dimension | Weight | Score /10 | Why |
|---|---|---|---|
| Technical quality | 30% | 8.0/10 | Near-frontier coding and agentic tool-calling, offset by a real hallucination-rate regression and text-only limits |
| Price-to-value | 25% | 9.5/10 | 10-20x cheaper than Western frontier models at comparable coding performance — the clearest win in this review |
| Maturity & documentation | 20% | 6.5/10 | Brand-new (Feb 2026) release with a published model card, but rising hallucination rate and unresolved distillation questions cut into confidence |
| Ceiling & flexibility | 15% | 7.5/10 | Open weights and self-hosting are real advantages, capped by the lack of multimodal input |
| Honesty of positioning | 10% | 6.0/10 | MiniMax markets coding parity accurately, but doesn't foreground the hallucination jump, and hasn't addressed Anthropic's distillation allegations publicly |
Weighted calculation: (8.0×0.30) + (9.5×0.25) + (6.5×0.20) + (7.5×0.15) + (6.0×0.10) = 7.80.

Comments
Post a Comment