Model comparison

Grok 4.3 vs MiMo-V2-Pro

Grok 4.3 and MiMo-V2-Pro score almost the same on the Noometry Index (43.8 vs 43.0), so choose on price, context window or the category you care about most.

Last verified . 20 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

MiMo-V2-Pro Xiaomi

43.0

Rank #103 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Grok 4.3 scores higher in 4 categories and MiMo-V2-Pro in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.3 leads 35.9 to 22.1.
  • The biggest single-benchmark swing is NYT Connections (extended): 55.2% for Grok 4.3 and 25.8% for MiMo-V2-Pro.
  • MiMo-V2-Pro is cheaper at $0.43 / $0.87 per million input/output tokens, against $1.25 / $2.50 for Grok 4.3.
  • MiMo-V2-Pro accepts more context: 1.05M tokens versus 1M.

Side by side

Grok 4.3 and MiMo-V2-Pro specifications
Grok 4.3MiMo-V2-Pro
ProviderxAIXiaomi
Noometry Index43.843.0
Released2026-04-172026-03-18
WeightsProprietaryProprietary
Context window1M1.05M
Max output30K131K
Input $ / M tokens$1.25$0.43
Output $ / M tokens$2.50$0.87
Results tracked4023

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiMo-V2-Pro leads

Grok 4.3: 41.6 (#121), MiMo-V2-Pro: 43.8 (#83)

Coding benchmarks
BenchmarkGrok 4.3MiMo-V2-Pro
LMArena WebDev13571433
LMArena Coding14151476
ALE-Bench944.17785.17
SciCode47.3%—
WeirdML49.9%—

Agentic & Tool Use Not comparable

Grok 4.3: 27.7 (#99), MiMo-V2-Pro: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.3MiMo-V2-Pro
GDP.pdf8%—
LMArena Search1165—
Vending-Bench 235.26—

Reasoning Grok 4.3 leads

Grok 4.3: 35.9 (#68), MiMo-V2-Pro: 22.1 (#206)

Reasoning benchmarks
BenchmarkGrok 4.3MiMo-V2-Pro
NYT Connections (extended)55.2%25.8%
LMArena Hard Prompts13961457
CritPt8%—
Chess Puzzles25%—
Thematic Generalization—45.9%
DTBench90.7%—
LMCA38.3%—
Epoch Capabilities Index149.16—
ForecastBench60.3—

Math Grok 4.3 leads

Grok 4.3: 46.0 (#74), MiMo-V2-Pro: 39.5 (#102)

Math benchmarks
BenchmarkGrok 4.3MiMo-V2-Pro
LMArena Math13881447
FrontierMath (Tiers 1-3)42.8%—
FrontierMath Tier 414.6%—
OTIS Mock AIME 2024-202593.3%—
ProofBench11%—

Knowledge Grok 4.3 leads

Grok 4.3: 52.5 (#62), MiMo-V2-Pro: 41.4 (#111)

Knowledge benchmarks
BenchmarkGrok 4.3MiMo-V2-Pro
LMArena Expert13851478
GPQA Diamond88.8%—
SimpleQA Verified33.2%—

Multimodal Not comparable

Grok 4.3: 31.6 (#104), MiMo-V2-Pro: —

Multimodal benchmarks
BenchmarkGrok 4.3MiMo-V2-Pro
LMArena Vision1229—
Blueprint-Bench 20%—

Multilingual MiMo-V2-Pro leads

Grok 4.3: 50.5 (#120), MiMo-V2-Pro: 52.7 (#81)

Multilingual benchmarks
BenchmarkGrok 4.3MiMo-V2-Pro
LMArena Non-English13851416
LMArena Chinese14221456
LMArena French14121469
LMArena German13951417
LMArena Japanese13791366
LMArena Korean13561400
LMArena Russian13991427
LMArena Spanish13981457

Instruction Following MiMo-V2-Pro leads

Grok 4.3: 72.1 (#140), MiMo-V2-Pro: 76.0 (#49)

Instruction Following benchmarks
BenchmarkGrok 4.3MiMo-V2-Pro
LMArena Instruction Following13661445

Long Context Too close to call

Grok 4.3: 42.5 (#123), MiMo-V2-Pro: 41.5 (#138)

Long Context benchmarks
BenchmarkGrok 4.3MiMo-V2-Pro
LMArena Longer Query13931455
CL-bench—15.7%
CL-bench Life—6.9%

Writing & Preference MiMo-V2-Pro leads

Grok 4.3: 58.5 (#118), MiMo-V2-Pro: 62.8 (#70)

Writing & Preference benchmarks
BenchmarkGrok 4.3MiMo-V2-Pro
LMArena Text13971436
LMArena Creative Writing13801415
LMArena Multi-Turn14061456
EQ-Bench 41075—

Frequently asked questions

Is Grok 4.3 better than MiMo-V2-Pro?

Grok 4.3 and MiMo-V2-Pro score almost the same on the Noometry Index (43.8 vs 43.0), so choose on price, context window or the category you care about most.

Which is cheaper, Grok 4.3 or MiMo-V2-Pro?

MiMo-V2-Pro is cheaper. It lists at $0.43 per million input tokens and $0.87 per million output tokens; Grok 4.3 lists at $1.25 and $2.50.

Is Grok 4.3 or MiMo-V2-Pro better for coding?

MiMo-V2-Pro scores higher on coding benchmarks: 43.8 versus 41.6 in the Noometry coding category.

Which has the bigger context window?

MiMo-V2-Pro does, with 1.05M tokens against 1M.

How many benchmarks do Grok 4.3 and MiMo-V2-Pro share?

20 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and MiMo-V2-Pro has 23.

Related comparisons

Go deeper