Model comparison

Claude Opus 4.1 vs MiMo-V2.5

MiMo-V2.5 is the stronger model overall, scoring 43.4 to 41.0 on the Noometry Index.

Last verified . 19 shared benchmarks.

Claude Opus 4.1 Anthropic

41.0

Rank #142 Confirmed

MiMo-V2.5 Xiaomi

43.4

Rank #93 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Claude Opus 4.1 scores higher in 7 categories and MiMo-V2.5 in 2 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where MiMo-V2.5 leads 36.8 to 22.3.
  • MiMo-V2.5 is cheaper at $0.14 / $0.28 per million input/output tokens, against $15 / $75 for Claude Opus 4.1.
  • MiMo-V2.5 accepts more context: 1.05M tokens versus 200K.
  • MiMo-V2.5 has downloadable open weights; the other is API-only.

Side by side

Claude Opus 4.1 and MiMo-V2.5 specifications
Claude Opus 4.1MiMo-V2.5
ProviderAnthropicXiaomi
Noometry Index41.043.4
Released2025-08-052026-04-22
WeightsProprietaryOpen
Context window200K1.05M
Max output32K131K
Input $ / M tokens$15$0.14
Output $ / M tokens$75$0.28
Results tracked4823

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude Opus 4.1: 44.4 (#73), MiMo-V2.5: 43.9 (#81)

Coding benchmarks
BenchmarkClaude Opus 4.1MiMo-V2.5
LMArena WebDev13901438
LMArena Coding14791469
ALE-Bench674.77513.95
SWE-bench Verified73.3%—
SciCode—43.1%
WeirdML45.9%—
AlgoTune1.34—

Agentic & Tool Use Not comparable

Claude Opus 4.1: 35.0 (#41), MiMo-V2.5: —

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4.1MiMo-V2.5
Terminal-Bench38%—
GDPval43.6%—
Cybench42%—
DeepResearch Bench48.3%—
LMArena Search1148—
METR Time Horizons66.8%—

Reasoning Claude Opus 4.1 leads

Claude Opus 4.1: 32.2 (#76), MiMo-V2.5: 28.6 (#101)

Reasoning benchmarks
BenchmarkClaude Opus 4.1MiMo-V2.5
LMArena Hard Prompts14431450
SimpleBench60%—
CritPt—3.7%
Chess Puzzles7%—
EnigmaEval7.2%—
EBR-Bench7.9%—
Mystery Game Puzzles21%—
DTBench80%—
LMCA37.1%—
Epoch Capabilities Index144.12—
ForecastBench62—

Math MiMo-V2.5 leads

Claude Opus 4.1: 22.3 (#277), MiMo-V2.5: 36.8 (#163)

Math benchmarks
BenchmarkClaude Opus 4.1MiMo-V2.5
LMArena Math14311436
FrontierMath (Tiers 1-3)12.6%—
FrontierMath Tier 42.4%—
OTIS Mock AIME 2024-202568.9%—
ProofBench—16%
FrontierMath (Feb 2025 set)7.2%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge Claude Opus 4.1 leads

Claude Opus 4.1: 42.0 (#101), MiMo-V2.5: 40.8 (#115)

Knowledge benchmarks
BenchmarkClaude Opus 4.1MiMo-V2.5
LMArena Expert14391460
GPQA Diamond77.3%—
Humanity's Last Exam11.5%—
Confabulations17.1%—
Vectara Hallucination Rate11.8%—

Multimodal MiMo-V2.5 leads

Claude Opus 4.1: 26.8 (#119), MiMo-V2.5: 39.8 (#54)

Multimodal benchmarks
BenchmarkClaude Opus 4.1MiMo-V2.5
LMArena Vision—1247
VPCT35%—

Multilingual Too close to call

Claude Opus 4.1: 52.0 (#95), MiMo-V2.5: 51.9 (#99)

Multilingual benchmarks
BenchmarkClaude Opus 4.1MiMo-V2.5
LMArena Non-English14051404
LMArena Chinese14271468
LMArena French14311447
LMArena German14131421
LMArena Japanese13781306
LMArena Korean13801363
LMArena Russian14221395
LMArena Spanish14481416

Instruction Following Too close to call

Claude Opus 4.1: 75.6 (#58), MiMo-V2.5: 75.5 (#60)

Instruction Following benchmarks
BenchmarkClaude Opus 4.1MiMo-V2.5
LMArena Instruction Following14351434

Long Context Too close to call

Claude Opus 4.1: 44.5 (#63), MiMo-V2.5: 44.2 (#73)

Long Context benchmarks
BenchmarkClaude Opus 4.1MiMo-V2.5
LMArena Longer Query14551445

Writing & Preference Too close to call

Claude Opus 4.1: 62.4 (#74), MiMo-V2.5: 61.6 (#86)

Writing & Preference benchmarks
BenchmarkClaude Opus 4.1MiMo-V2.5
LMArena Text14191428
LMArena Creative Writing14121393
LMArena Multi-Turn14441445
Short-Story Creative Writing84.7%—

Frequently asked questions

Is Claude Opus 4.1 better than MiMo-V2.5?

MiMo-V2.5 is the stronger model overall, scoring 43.4 to 41.0 on the Noometry Index.

Which is cheaper, Claude Opus 4.1 or MiMo-V2.5?

MiMo-V2.5 is cheaper. It lists at $0.14 per million input tokens and $0.28 per million output tokens; Claude Opus 4.1 lists at $15 and $75.

Is Claude Opus 4.1 or MiMo-V2.5 better for coding?

They score almost the same on coding (44.4 vs 43.9); test both on your own repository before choosing.

Which has the bigger context window?

MiMo-V2.5 does, with 1.05M tokens against 200K.

How many benchmarks do Claude Opus 4.1 and MiMo-V2.5 share?

19 benchmarks have published results for both models. Claude Opus 4.1 has 48 scored results on Noometry and MiMo-V2.5 has 23.

Related comparisons

Go deeper