Model comparison

Claude 3 Sonnet vs MiMo-V2-Flash

MiMo-V2-Flash is the stronger model overall, scoring 41.3 to 29.0 on the Noometry Index.

Last verified . 17 shared benchmarks.

Claude 3 Sonnet Anthropic

29.0

Rank #319 Confirmed

MiMo-V2-Flash Xiaomi

41.3

Rank #138 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Claude 3 Sonnet scores higher in 0 categories and MiMo-V2-Flash in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where MiMo-V2-Flash leads 38.3 to 10.7.
  • MiMo-V2-Flash has downloadable open weights; the other is API-only.

Side by side

Claude 3 Sonnet and MiMo-V2-Flash specifications
Claude 3 SonnetMiMo-V2-Flash
ProviderAnthropicXiaomi
Noometry Index29.041.3
Released2024-02-292025-12-16
WeightsProprietaryOpen
Context window—262K
Max output—66K
Input $ / M tokens—$0.14
Output $ / M tokens—$0.28
Results tracked3021

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiMo-V2-Flash leads

Claude 3 Sonnet: 29.6 (#302), MiMo-V2-Flash: 36.1 (#211)

Coding benchmarks
BenchmarkClaude 3 SonnetMiMo-V2-Flash
LMArena Coding12231443
LMArena WebDev—1330
SciCode—25.9%
WeirdML10.2%—
BigCodeBench Instruct42.7%—
BigCodeBench Complete53.8%—
ALE-Bench—737.95
HumanEval+64%—
MBPP+69.3%—

Reasoning MiMo-V2-Flash leads

Claude 3 Sonnet: 20.5 (#237), MiMo-V2-Flash: 24.9 (#157)

Reasoning benchmarks
BenchmarkClaude 3 SonnetMiMo-V2-Flash
LMArena Hard Prompts11971420
CritPt—0%
DTBench53.6%—
Epoch Capabilities Index120.7—
WinoGrande75.1%—

Math MiMo-V2-Flash leads

Claude 3 Sonnet: 10.7 (#310), MiMo-V2-Flash: 38.3 (#139)

Math benchmarks
BenchmarkClaude 3 SonnetMiMo-V2-Flash
LMArena Math12131396
OTIS Mock AIME 2024-20252.5%—
MATH Level 518.2%—

Knowledge MiMo-V2-Flash leads

Claude 3 Sonnet: 21.1 (#276), MiMo-V2-Flash: 39.7 (#131)

Knowledge benchmarks
BenchmarkClaude 3 SonnetMiMo-V2-Flash
LMArena Expert11731425
GPQA Diamond40.6%—
MMLU75.9%—

Multimodal Not comparable

Claude 3 Sonnet: 25.2 (#125), MiMo-V2-Flash: —

Multimodal benchmarks
BenchmarkClaude 3 SonnetMiMo-V2-Flash
LMArena Vision984—

Multilingual MiMo-V2-Flash leads

Claude 3 Sonnet: 37.8 (#234), MiMo-V2-Flash: 51.0 (#113)

Multilingual benchmarks
BenchmarkClaude 3 SonnetMiMo-V2-Flash
LMArena Non-English12051392
LMArena Chinese11891462
LMArena French12291429
LMArena German12041395
LMArena Japanese11311325
LMArena Korean11281358
LMArena Russian12271387
LMArena Spanish12041420

Instruction Following MiMo-V2-Flash leads

Claude 3 Sonnet: 62.8 (#235), MiMo-V2-Flash: 73.5 (#120)

Instruction Following benchmarks
BenchmarkClaude 3 SonnetMiMo-V2-Flash
LMArena Instruction Following11991392

Long Context MiMo-V2-Flash leads

Claude 3 Sonnet: 36.7 (#228), MiMo-V2-Flash: 43.0 (#110)

Long Context benchmarks
BenchmarkClaude 3 SonnetMiMo-V2-Flash
LMArena Longer Query12111409

Writing & Preference MiMo-V2-Flash leads

Claude 3 Sonnet: 42.1 (#238), MiMo-V2-Flash: 59.7 (#106)

Writing & Preference benchmarks
BenchmarkClaude 3 SonnetMiMo-V2-Flash
LMArena Text12181411
LMArena Creative Writing11861375
LMArena Multi-Turn12271404

Frequently asked questions

Is Claude 3 Sonnet better than MiMo-V2-Flash?

MiMo-V2-Flash is the stronger model overall, scoring 41.3 to 29.0 on the Noometry Index.

Is Claude 3 Sonnet or MiMo-V2-Flash better for coding?

MiMo-V2-Flash scores higher on coding benchmarks: 36.1 versus 29.6 in the Noometry coding category.

How many benchmarks do Claude 3 Sonnet and MiMo-V2-Flash share?

17 benchmarks have published results for both models. Claude 3 Sonnet has 30 scored results on Noometry and MiMo-V2-Flash has 21.

Related comparisons

Go deeper