Model comparison

MiMo-V2-Flash vs Qwen3.8 27B

Qwen3.8 27B is the stronger model overall, scoring 46.0 to 41.3 on the Noometry Index. MiMo-V2-Flash costs 6.4× less per token, which makes it the better buy when Qwen3.8 27B's lead doesn't matter for your workload.

Last verified . 20 shared benchmarks.

MiMo-V2-Flash Xiaomi

41.3

Rank #138 Confirmed

Qwen3.8 27B Alibaba (Qwen)

46.0

Rank #68 Confirmed

Summary

  • They share 20 benchmarks with published results for both. MiMo-V2-Flash scores higher in 1 category and Qwen3.8 27B in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.8 27B leads 41.0 to 24.9.
  • The biggest single-benchmark swing is SciCode: 25.9% for MiMo-V2-Flash and 46.6% for Qwen3.8 27B.
  • MiMo-V2-Flash is cheaper at $0.14 / $0.28 per million input/output tokens, against $0.99 / $1.49 for Qwen3.8 27B.

Side by side

MiMo-V2-Flash and Qwen3.8 27B specifications
MiMo-V2-FlashQwen3.8 27B
ProviderXiaomiAlibaba (Qwen)
Noometry Index41.346.0
Released2025-12-162026-08-14
WeightsOpenOpen
Context window262K262K
Max output66K33K
Input $ / M tokens$0.14$0.99
Output $ / M tokens$0.28$1.49
Results tracked2131

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.8 27B leads

MiMo-V2-Flash: 36.1 (#211), Qwen3.8 27B: 50.5 (#44)

Coding benchmarks
BenchmarkMiMo-V2-FlashQwen3.8 27B
LMArena WebDev13301593
SciCode25.9%46.6%
LMArena Coding14431482
ALE-Bench737.95—

Agentic & Tool Use Not comparable

MiMo-V2-Flash: —, Qwen3.8 27B: 32.9 (#57)

Agentic & Tool Use benchmarks
BenchmarkMiMo-V2-FlashQwen3.8 27B
APEX-Agents—47.5%

Reasoning Qwen3.8 27B leads

MiMo-V2-Flash: 24.9 (#157), Qwen3.8 27B: 41.0 (#54)

Reasoning benchmarks
BenchmarkMiMo-V2-FlashQwen3.8 27B
CritPt0%5.4%
LMArena Hard Prompts14201460
ARC-AGI-2—42.4%
NYT Connections (extended)—54.5%
ARC-AGI-1—87.5%
DTBench—88%
LMCA—41.4%
Surface Evolver Bench—45%
Epoch Capabilities Index—149.38

Math MiMo-V2-Flash leads

MiMo-V2-Flash: 38.3 (#139), Qwen3.8 27B: 37.1 (#161)

Math benchmarks
BenchmarkMiMo-V2-FlashQwen3.8 27B
LMArena Math13961456
ProofBench—16%

Knowledge Qwen3.8 27B leads

MiMo-V2-Flash: 39.7 (#131), Qwen3.8 27B: 41.6 (#109)

Knowledge benchmarks
BenchmarkMiMo-V2-FlashQwen3.8 27B
LMArena Expert14251482

Multimodal Not comparable

MiMo-V2-Flash: —, Qwen3.8 27B: 41.3 (#37)

Multimodal benchmarks
BenchmarkMiMo-V2-FlashQwen3.8 27B
LMArena Vision—1271

Multilingual Qwen3.8 27B leads

MiMo-V2-Flash: 51.0 (#113), Qwen3.8 27B: 53.7 (#60)

Multilingual benchmarks
BenchmarkMiMo-V2-FlashQwen3.8 27B
LMArena Non-English13921430
LMArena Chinese14621504
LMArena French14291465
LMArena German13951438
LMArena Japanese13251384
LMArena Korean13581393
LMArena Russian13871415
LMArena Spanish14201448

Instruction Following Qwen3.8 27B leads

MiMo-V2-Flash: 73.5 (#120), Qwen3.8 27B: 75.8 (#53)

Instruction Following benchmarks
BenchmarkMiMo-V2-FlashQwen3.8 27B
LMArena Instruction Following13921439

Long Context Qwen3.8 27B leads

MiMo-V2-Flash: 43.0 (#110), Qwen3.8 27B: 44.3 (#70)

Long Context benchmarks
BenchmarkMiMo-V2-FlashQwen3.8 27B
LMArena Longer Query14091450

Writing & Preference Qwen3.8 27B leads

MiMo-V2-Flash: 59.7 (#106), Qwen3.8 27B: 65.8 (#43)

Writing & Preference benchmarks
BenchmarkMiMo-V2-FlashQwen3.8 27B
LMArena Text14111441
LMArena Creative Writing13751384
LMArena Multi-Turn14041441
EQ-Bench Creative Writing—1671

Frequently asked questions

Is MiMo-V2-Flash better than Qwen3.8 27B?

Qwen3.8 27B is the stronger model overall, scoring 46.0 to 41.3 on the Noometry Index. MiMo-V2-Flash costs 6.4× less per token, which makes it the better buy when Qwen3.8 27B's lead doesn't matter for your workload.

Which is cheaper, MiMo-V2-Flash or Qwen3.8 27B?

MiMo-V2-Flash is cheaper. It lists at $0.14 per million input tokens and $0.28 per million output tokens; Qwen3.8 27B lists at $0.99 and $1.49.

Is MiMo-V2-Flash or Qwen3.8 27B better for coding?

Qwen3.8 27B scores higher on coding benchmarks: 50.5 versus 36.1 in the Noometry coding category.

Which has the bigger context window?

Both accept 262K tokens.

How many benchmarks do MiMo-V2-Flash and Qwen3.8 27B share?

20 benchmarks have published results for both models. MiMo-V2-Flash has 21 scored results on Noometry and Qwen3.8 27B has 31.

Related comparisons

Go deeper