↑
Press ESC or click to close
Latest
Loading latest reviews…

Did a $0.30 Chinese AI Model Just Catch Claude Opus in Coding?

Mahmoud Salamoun · September 22, 2026 · 5 min read
Did a $0.30 Chinese AI Model Just Catch Claude Opus in Coding?
AI Models Updated Sep 2026

Did a $0.30-Per-Million-Token Chinese Model Just Catch Up to Claude Opus in Coding?

MiniMax M2.5 scores within 0.6 points of Claude Opus 4.6 on SWE-Bench Verified and beats it outright on multi-step agent tasks — at roughly a tenth of the price. It also arrives wrapped in a distillation controversy Anthropic says is real.

7.8/ 10
📋 Technical Desk Review — built from official documentation, benchmark publishers, and independently verified third-party coverage. No hands-on testing claimed.
Last verified: September 19, 2026
📋 Table of Contents
  1. What M2.5 Actually Is
  2. The Number That Flips the Table: 80.2% on SWE-Bench
  3. The Bigger Surprise: Agentic Task Dominance
  4. A Price That Changes Agent Economics
  5. Strengths at a Glance
  6. The Dark Side: Hallucinations, Math, and a Distillation Fight
  7. When to Use M2.5 — and When to Avoid It
  8. Expert Editorial Opinion
  9. Final Verdict
  10. FAQ

MiniMax M2.5 review 2026: until early this year, the pecking order in frontier coding models felt settled — Claude, GPT, and Gemini traded the top spots, and Chinese labs competed mostly on price and "good enough" performance a tier below. On February 12, 2026, Shanghai-based MiniMax released M2.5 and that assumption got harder to defend, at least in one specific lane.

This review breaks down what M2.5 actually scores on the benchmarks that matter for real coding and agent work, where it genuinely leads, where it clearly still trails the Western frontier, and the distillation controversy that Anthropic has publicly attached to MiniMax's name. It's built from MiniMax's own model documentation, independent benchmark trackers, and verified reporting — not a claim of hands-on testing of M2.5 itself unless stated otherwise.

MS
Mahmoud Salamoun
Founder, ToolRadar · Reviewed Sep 2026
Independent AI tools reviewer with a background in marketing and content, and hands-on daily experience directing AI tools like Gemini and ChatGPT for real work. This review is based on official documentation, published pricing, and verified third-party coverage — not a claim of hands-on testing of M2.5 itself unless stated otherwise.

What M2.5 Actually Is

01

Third Generation, Real Funding Behind It

MiniMax was founded in Shanghai in 2021 by researchers with SenseTime backgrounds, raised $619 million in a Hong Kong IPO in January 2026, and is backed by Alibaba and Tencent. M2.5 is the third model in its line, following M2 and M2.1.

02

230B Mixture-of-Experts, 10B Active

M2.5 uses a Mixture-of-Experts architecture with 230 billion total parameters and roughly 10 billion active at inference — a design choice that keeps inference cost down while retaining a large total knowledge base.

03

~200K Token Context, Text-Only

The context window sits around 200,000 tokens depending on the provider. There is no image, screenshot, or multimodal input support — this is a text-only model, which matters for visual debugging workflows.

04

Open Weights, Commercial License

Weights are published on Hugging Face under a permissive MIT-family license that allows commercial use and self-hosting — a real point of differentiation from Claude, GPT, or Gemini, none of which publish weights at all.

The Number That Flips the Table: 80.2% on SWE-Bench

SWE-Bench Verified is the benchmark most worth trusting for real coding ability: it grades a model on genuine GitHub issues across production repositories, with terminal access and the ability to actually edit files — not a curated toy problem set. M2.5 scored 80.2%, putting it just 0.6 points behind Claude Opus 4.6 and ahead of both GPT-5.2 and Gemini 3 Pro on this specific measure. Independent trackers, including Artificial Analysis and multiple benchmark write-ups published after launch, corroborate the figure. It's also the highest score recorded by any open-weight model on this benchmark as of its release.

BenchmarkM2.5Opus 4.6GPT-5.2Gemini 3 Pro
SWE-Bench Verified80.2%80.8%80%78%
Multi-SWE-Bench51.3%50.3%—42.7%
BFCL Multi-Turn (agent)76.8%63.3%—61%
BrowseComp (w/ context mgmt)76.3%84%65.8%59.2%
AIME 2025 (math)86.395.698.096.0
GPQA Diamond85.290.0——
AA Intelligence Index4253—57

Read the whole table, not just the headline row, and the real shape of the story appears: M2.5 is essentially tied at the top for coding and multilingual coding, but it trails by 9-12 points on math and general reasoning. This isn't a model that beats the Western frontier across the board — it's one that closed a very specific, very commercially valuable gap.

The Bigger Surprise: Agentic Task Dominance

The coding parity is the headline, but the more consequential number for anyone building AI agents is BFCL Multi-Turn — a benchmark for multi-step function calling across complex API schemas, which is exactly where real agent pipelines tend to break. M2.5 scored 76.8% here, a full 13.5 points ahead of Opus 4.6's 63.3% and 15.8 points ahead of Gemini 3 Pro. That's not a marginal edge; it's the kind of gap that shows up directly as fewer failed tool calls and fewer wasted retries in a long-running agent loop.

"A closed-loop system that misses nothing, and proactively adds value during long-horizon tasks."

— OpenHands' independent evaluation, describing M2.5's behavior on extended agent runs.

MiniMax backs this up with two supporting figures worth noting: M2.5 completes agentic tasks using roughly 20% fewer rounds than competing models, and it finishes SWE-Bench tasks 37% faster than its own predecessor, M2.1. Together, that points to a model specifically trained — via MiniMax's Forge reinforcement-learning framework across more than 200,000 simulated environments — for the exact failure mode that makes production agents expensive: losing context and burning tokens on retries.

A Price That Changes Agent Economics

This is the number that turns M2.5 from "an interesting open-weight model" into something enterprises actually have to evaluate. The Lightning variant runs at $2.40 per million output tokens, against roughly $75 for Opus 4.6, $60 for GPT-5.2, and $20 for Gemini 3 Pro — a 10 to 20x gap, not a marginal discount.

ModelInput $/M tokensOutput $/M tokensEst. daily cost*
M2.5 Standard$0.15$1.20~$4.70/day
M2.5 Lightning$0.30$2.40~$1/hour continuous
Claude Opus 4.6$5.00$25.00~$100/day
Gemini 3.1 Pro$2.00$8.00~$36/day

*Estimated for 1M input + 2M output tokens/day. Rates vary slightly across hosting providers (OpenRouter, Bedrock, MiniMax's own API); figures above reflect MiniMax's published first-party pricing.

At roughly $1 per hour of continuous 24/7 agent operation, M2.5 makes always-on agentic workflows economically viable in a way the frontier Western models simply don't at their current pricing — that's the practical, bottom-line consequence of the benchmark table above.

Strengths at a Glance

✓Coding performance within a point of Claude Opus 4.6 on SWE-Bench Verified — the highest score any open-weight model has posted on this benchmark.
✓Clear lead on multi-turn agentic tool calling (BFCL), the benchmark that most directly predicts real agent-pipeline reliability.
✓10-20x cheaper than Western frontier alternatives, with genuinely usable 24/7-agent economics.
✓Open weights under a commercially permissive license, enabling self-hosting and fine-tuning that closed models don't allow.
✕Hallucination rate jumped to 88%, up from 67% on the previous M2.1 generation, per Artificial Analysis — a real regression in factual reliability.
✕A clear 9-12 point gap versus the Western frontier on hard math and deep reasoning benchmarks (AIME, GPQA).
✕Text-only — no image input, no visual debugging, no multimodal support of any kind.
✕Unusually verbose: it generated roughly 56 million tokens during one intelligence evaluation, versus an average of about 16 million for comparable models — a real cost multiplier the headline per-token price doesn't show.

The Dark Side: Hallucinations, Math, and a Distillation Fight

Hallucination rate spiked. Artificial Analysis measured an 88% hallucination rate for M2.5, up sharply from 67% for M2.1, with the AA-Omniscience reliability score dropping from -30 to -41. That's a meaningful step backward in factual grounding — less damaging for pure code generation, where output either compiles and passes tests or it doesn't, but a real liability for any task involving factual claims or information retrieval.

The math and reasoning gap is real and MiniMax doesn't hide it. On AIME 2025, M2.5 scored 86.3 against 95.6 for Opus, 98.0 for GPT-5.2, and 96.0 for Gemini 3 Pro — a 9-12 point gap that repeats on the Artificial Analysis Intelligence Index (42 vs. 53-57). This isn't a model optimized for hard math, and it isn't marketed as one; anyone benchmarking it against frontier reasoning specifically will come away disappointed.

The distillation controversy. In February 2026, Anthropic published findings accusing DeepSeek, Moonshot AI, and MiniMax of running a coordinated "distillation attack" against Claude — using roughly 24,000 fraudulent accounts to generate more than 16 million exchanges with Claude, with MiniMax alone responsible for over 13 million of those exchanges. Anthropic further alleged that MiniMax redirected close to half its traffic within 24 hours of any new Anthropic model release, specifically to harvest capabilities from the newest system. MiniMax has not publicly confirmed or denied the specifics. This doesn't invalidate M2.5 as a product — the benchmarks above are independently reproducible regardless of how the model was trained — but it's a legitimate reason to read direct comparisons against Claude specifically with some added skepticism, and it's part of a wider pattern Anthropic has since documented across multiple China-based labs.

Practical limits worth knowing before you deploy it. Beyond text-only input and the verbosity issue above, the raw model weights are roughly 457GB uncompressed — self-hosting requires a serious multi-GPU setup, not a hobbyist machine.

When to Use M2.5 — and When to Avoid It

✅ Use it when...

you're running high-volume coding agents — multi-file refactors, SWE-Bench-style tasks, long tool-calling loops — where cost is a real constraint or you want genuinely affordable 24/7 agent operation, or you specifically want open weights for self-hosting and fine-tuning.

❌ Avoid it when...

you need high factual reliability or information retrieval accuracy (88% hallucination rate is disqualifying for that), the task centers on hard math or deep scientific reasoning, or you need multimodal input like screenshots, diagrams, or visual debugging.

Expert Editorial Opinion

The honest framing for M2.5 isn't "China caught up to the West in AI" — that headline is both too broad and too generous. The more accurate claim is narrower and, in some ways, more interesting: a Chinese open-weight lab matched the closed Western frontier in one specific, commercially critical lane — real-world coding and agentic tool use — while still trailing clearly in deep reasoning and factual reliability. Precision matters here because the broad version of the claim gets disproven the moment someone runs an AIME math benchmark, and the narrow version is the one actually worth taking seriously.

What makes the agentic number more convincing than the coding number, to me, is that it's harder to game. SWE-Bench has been optimized against by nearly every major lab for over a year now; a 0.6-point gap at the top is close enough that scaffold and harness differences could easily explain it. A 13.5-point lead on multi-turn tool calling is a bigger, structurally different kind of result, and it lines up with MiniMax's own stated training approach — reinforcement learning across 200,000+ simulated environments specifically targeting long-horizon tool use. That's a testable mechanism behind the number, not just a benchmark that happened to land favorably.

The hallucination jump deserves more attention than it's gotten in most coverage I've seen. Going from 67% to 88% between M2.1 and M2.5 is not a rounding error — it suggests MiniMax traded some factual grounding for agentic and coding gains, which is a real, specific tradeoff, not a vague "still improving" caveat. For code generation, where correctness is externally verifiable by running the code, that tradeoff is survivable. For anything involving research, fact-checking, or information synthesis, it's a serious limitation that should rule the model out entirely.

The distillation allegations are the part every review has to handle carefully, because the two things people want to conflate — "is the model good" and "how was the model built" — are genuinely separate questions with genuinely separate answers. Anthropic's findings are specific and dated (24,000 fraudulent accounts, 16+ million exchanges, MiniMax accounting for the majority of it), and MiniMax hasn't offered a public rebuttal. That's a legitimate shadow over any head-to-head marketing MiniMax does against Claude specifically. It does not, however, change what SWE-Bench Verified or BFCL actually measure — those are external, reproducible evaluations, not MiniMax's own claims.

Put practically: M2.5 is a genuinely strong, genuinely cheap coding and agent model that earns real consideration for high-volume production agent work — used as a fast, low-cost drafting layer with a more reliable model verifying anything that touches facts or hard reasoning, exactly as MiniMax's own positioning suggests. Treating it as a like-for-like Claude replacement across the board would be a mistake; treating it as proof that the coding/agent gap between Chinese and Western frontier models has meaningfully narrowed is a defensible, evidence-backed conclusion.

Final Verdict

ToolRadar Performance Score
7.8 / 10

Score band: 7.0-7.9 — Competent but compromised: a strong fit for a specific, well-defined use case (high-volume coding agents), not a general-purpose replacement for the Western frontier.

DimensionWeightScore /10Why
Technical quality30%8.0/10Near-frontier coding and agentic tool-calling, offset by a real hallucination-rate regression and text-only limits
Price-to-value25%9.5/1010-20x cheaper than Western frontier models at comparable coding performance — the clearest win in this review
Maturity & documentation20%6.5/10Brand-new (Feb 2026) release with a published model card, but rising hallucination rate and unresolved distillation questions cut into confidence
Ceiling & flexibility15%7.5/10Open weights and self-hosting are real advantages, capped by the lack of multimodal input
Honesty of positioning10%6.0/10MiniMax markets coding parity accurately, but doesn't foreground the hallucination jump, and hasn't addressed Anthropic's distillation allegations publicly

Weighted calculation: (8.0×0.30) + (9.5×0.25) + (6.5×0.20) + (7.5×0.15) + (6.0×0.10) = 7.80.

❓ Frequently Asked Questions

On SWE-Bench Verified specifically, yes — M2.5 scores 80.2% against Opus 4.6's 80.8%, a 0.6-point gap. It even leads Opus on multi-turn agentic tool calling. But Opus remains clearly ahead on hard math, deep reasoning, and factual reliability, so "as good" only holds for coding and agent tasks specifically, not general capability.
MiniMax hasn't published a detailed explanation. Artificial Analysis measured an increase from 67% (M2.1) to 88% (M2.5), alongside a drop in the AA-Omniscience reliability score. It appears to coincide with the model's heavier optimization toward coding and agentic tool use, though that connection isn't officially confirmed by MiniMax.
Anthropic published specific, dated findings in February 2026 accusing MiniMax, DeepSeek, and Moonshot AI of coordinated distillation attacks against Claude, including account and exchange counts. MiniMax has not publicly confirmed or denied the allegations. It remains an accusation from one party rather than an independently adjudicated fact, but it is a credible, specific, and sourced claim worth factoring into how you read MiniMax's own comparisons to Claude.
MiniMax's published API pricing is around $0.15 per million input tokens and $1.20 per million output tokens for the Standard tier, or $0.30/$2.40 for the faster Lightning variant — roughly 10-20x cheaper than Claude Opus 4.6 at comparable coding performance. Note that M2.5 is also unusually verbose, which can partly offset the low per-token price on real workloads.
Yes — the weights are published on Hugging Face under a commercially permissive, MIT-family license. In practice this requires serious infrastructure: the uncompressed model is roughly 457GB, which means a multi-GPU setup rather than a single consumer machine.
Share this review
Mahmoud Salamoun
Written by
Mahmoud Salamoun
Independent AI tools reviewer based in the Middle East. I test and rate AI tools so you don't have to — no sponsorships, no bias, just honest analysis.
Rate this review
★ ★ ★ ★ ★
(-/5)

Comments