Model comparison

MiMo-V2.6-Flash vs Qwen3.5 397B-A17B

MiMo-V2.6-Flash is the stronger model overall, scoring 48.5 to 46.0 on the Noometry Index.

Last verified . 16 shared benchmarks.

MiMo-V2.6-Flash Xiaomi

48.5

Rank #55 Confirmed

Qwen3.5 397B-A17B Alibaba (Qwen)

46.0

Rank #67 Confirmed

Summary

  • They share 16 benchmarks with published results for both. MiMo-V2.6-Flash scores higher in 7 categories and Qwen3.5 397B-A17B in 2 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in coding, where MiMo-V2.6-Flash leads 53.4 to 42.0.
  • MiMo-V2.6-Flash is cheaper at $0.14 / $0.28 per million input/output tokens, against $0.60 / $3.60 for Qwen3.5 397B-A17B.
  • MiMo-V2.6-Flash accepts more context: 1.05M tokens versus 262K.

Side by side

MiMo-V2.6-Flash and Qwen3.5 397B-A17B specifications
MiMo-V2.6-FlashQwen3.5 397B-A17B
ProviderXiaomiAlibaba (Qwen)
Noometry Index48.546.0
Released2026-09-212026-02-01
WeightsOpenOpen
Context window1.05M262K
Max output131K66K
Input $ / M tokens$0.14$0.60
Output $ / M tokens$0.28$3.60
Results tracked1936

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiMo-V2.6-Flash leads

MiMo-V2.6-Flash: 53.4 (#30), Qwen3.5 397B-A17B: 42.0 (#114)

Coding benchmarks
BenchmarkMiMo-V2.6-FlashQwen3.5 397B-A17B
LMArena WebDev16371400
LMArena Coding15041465
SciCode51.3%—

Agentic & Tool Use Not comparable

MiMo-V2.6-Flash: —, Qwen3.5 397B-A17B: 33.3 (#53)

Agentic & Tool Use benchmarks
BenchmarkMiMo-V2.6-FlashQwen3.5 397B-A17B
APEX-Agents—24.9%
τ²-bench Airline—81.5%
τ²-bench Banking—9.8%
τ²-bench Retail—84.4%
τ²-bench Telecom—97.8%

Reasoning MiMo-V2.6-Flash leads

MiMo-V2.6-Flash: 36.5 (#66), Qwen3.5 397B-A17B: 34.5 (#70)

Reasoning benchmarks
BenchmarkMiMo-V2.6-FlashQwen3.5 397B-A17B
LMArena Hard Prompts14821448
Kagi LLM Benchmark—73.7%
NYT Connections (extended)—58.9%
CritPt12%—
Chess Puzzles—13%
Thematic Generalization—65.1%
Mystery Game Puzzles—18%
DTBench—87.5%
LMCA—37.9%
Epoch Capabilities Index—146.65

Math MiMo-V2.6-Flash leads

MiMo-V2.6-Flash: 51.9 (#52), Qwen3.5 397B-A17B: 46.1 (#73)

Math benchmarks
BenchmarkMiMo-V2.6-FlashQwen3.5 397B-A17B
LMArena Math14681454
FrontierMath (Tiers 1-3)—31.2%
OTIS Mock AIME 2024-2025—88.9%
ProofBench63%—

Knowledge Qwen3.5 397B-A17B leads

MiMo-V2.6-Flash: 42.2 (#99), Qwen3.5 397B-A17B: 53.3 (#58)

Knowledge benchmarks
BenchmarkMiMo-V2.6-FlashQwen3.5 397B-A17B
LMArena Expert15011462
GPQA Diamond—86.4%

Multimodal Too close to call

MiMo-V2.6-Flash: 40.5 (#47), Qwen3.5 397B-A17B: 40.7 (#44)

Multimodal benchmarks
BenchmarkMiMo-V2.6-FlashQwen3.5 397B-A17B
LMArena Vision12591263

Multilingual Too close to call

MiMo-V2.6-Flash: 54.0 (#51), Qwen3.5 397B-A17B: 53.7 (#59)

Multilingual benchmarks
BenchmarkMiMo-V2.6-FlashQwen3.5 397B-A17B
LMArena Non-English14341430
LMArena Chinese15111500
LMArena French14751461
LMArena Russian14091429
LMArena Spanish14561441
LMArena German—1447
LMArena Japanese—1426
LMArena Korean—1384

Instruction Following MiMo-V2.6-Flash leads

MiMo-V2.6-Flash: 76.8 (#35), Qwen3.5 397B-A17B: 75.0 (#77)

Instruction Following benchmarks
BenchmarkMiMo-V2.6-FlashQwen3.5 397B-A17B
LMArena Instruction Following14631424

Long Context Too close to call

MiMo-V2.6-Flash: 44.8 (#57), Qwen3.5 397B-A17B: 44.1 (#74)

Long Context benchmarks
BenchmarkMiMo-V2.6-FlashQwen3.5 397B-A17B
LMArena Longer Query14631442

Writing & Preference Too close to call

MiMo-V2.6-Flash: 63.1 (#67), Qwen3.5 397B-A17B: 62.3 (#79)

Writing & Preference benchmarks
BenchmarkMiMo-V2.6-FlashQwen3.5 397B-A17B
LMArena Text14551438
LMArena Creative Writing14001401
LMArena Multi-Turn14511446
EQ-Bench Creative Writing—1478

Frequently asked questions

Is MiMo-V2.6-Flash better than Qwen3.5 397B-A17B?

MiMo-V2.6-Flash is the stronger model overall, scoring 48.5 to 46.0 on the Noometry Index.

Which is cheaper, MiMo-V2.6-Flash or Qwen3.5 397B-A17B?

MiMo-V2.6-Flash is cheaper. It lists at $0.14 per million input tokens and $0.28 per million output tokens; Qwen3.5 397B-A17B lists at $0.60 and $3.60.

Is MiMo-V2.6-Flash or Qwen3.5 397B-A17B better for coding?

MiMo-V2.6-Flash scores higher on coding benchmarks: 53.4 versus 42.0 in the Noometry coding category.

Which has the bigger context window?

MiMo-V2.6-Flash does, with 1.05M tokens against 262K.

How many benchmarks do MiMo-V2.6-Flash and Qwen3.5 397B-A17B share?

16 benchmarks have published results for both models. MiMo-V2.6-Flash has 19 scored results on Noometry and Qwen3.5 397B-A17B has 36.

Related comparisons

Go deeper