Model comparison

MiMo-V2-Flash vs Trinity Large Thinking

MiMo-V2-Flash is the stronger model overall, scoring 41.3 to 38.6 on the Noometry Index.

Last verified . 20 shared benchmarks.

MiMo-V2-Flash Xiaomi

41.3

Rank #138 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 20 benchmarks with published results for both. MiMo-V2-Flash scores higher in 7 categories and Trinity Large Thinking in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where MiMo-V2-Flash leads 24.9 to 16.9.
  • The biggest single-benchmark swing is SciCode: 25.9% for MiMo-V2-Flash and 36.1% for Trinity Large Thinking.
  • MiMo-V2-Flash is cheaper at $0.14 / $0.28 per million input/output tokens, against $0.25 / $0.80 for Trinity Large Thinking.

Side by side

MiMo-V2-Flash and Trinity Large Thinking specifications
MiMo-V2-FlashTrinity Large Thinking
ProviderXiaomiArcee AI
Noometry Index41.338.6
Released2025-12-162026-04-01
WeightsOpenOpen
Context window262K262K
Max output66K80K
Input $ / M tokens$0.14$0.25
Output $ / M tokens$0.28$0.80
Results tracked2124

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiMo-V2-Flash leads

MiMo-V2-Flash: 36.1 (#211), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkMiMo-V2-FlashTrinity Large Thinking
LMArena WebDev13301238
SciCode25.9%36.1%
LMArena Coding14431381
ALE-Bench737.95—

Reasoning MiMo-V2-Flash leads

MiMo-V2-Flash: 24.9 (#157), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkMiMo-V2-FlashTrinity Large Thinking
CritPt0%0.9%
LMArena Hard Prompts14201350
NYT Connections (extended)—16.5%
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%

Math Too close to call

MiMo-V2-Flash: 38.3 (#139), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkMiMo-V2-FlashTrinity Large Thinking
LMArena Math13961366

Knowledge Trinity Large Thinking leads

MiMo-V2-Flash: 39.7 (#131), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkMiMo-V2-FlashTrinity Large Thinking
LMArena Expert14251360
Vectara Hallucination Rate—6.9%

Multilingual MiMo-V2-Flash leads

MiMo-V2-Flash: 51.0 (#113), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkMiMo-V2-FlashTrinity Large Thinking
LMArena Non-English13921325
LMArena Chinese14621373
LMArena French14291374
LMArena German13951356
LMArena Japanese13251311
LMArena Korean13581306
LMArena Russian13871337
LMArena Spanish14201357

Instruction Following MiMo-V2-Flash leads

MiMo-V2-Flash: 73.5 (#120), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkMiMo-V2-FlashTrinity Large Thinking
LMArena Instruction Following13921334

Long Context MiMo-V2-Flash leads

MiMo-V2-Flash: 43.0 (#110), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkMiMo-V2-FlashTrinity Large Thinking
LMArena Longer Query14091355

Writing & Preference MiMo-V2-Flash leads

MiMo-V2-Flash: 59.7 (#106), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkMiMo-V2-FlashTrinity Large Thinking
LMArena Text14111340
LMArena Creative Writing13751320
LMArena Multi-Turn14041342

Frequently asked questions

Is MiMo-V2-Flash better than Trinity Large Thinking?

MiMo-V2-Flash is the stronger model overall, scoring 41.3 to 38.6 on the Noometry Index.

Which is cheaper, MiMo-V2-Flash or Trinity Large Thinking?

MiMo-V2-Flash is cheaper. It lists at $0.14 per million input tokens and $0.28 per million output tokens; Trinity Large Thinking lists at $0.25 and $0.80.

Is MiMo-V2-Flash or Trinity Large Thinking better for coding?

MiMo-V2-Flash scores higher on coding benchmarks: 36.1 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Both accept 262K tokens.

How many benchmarks do MiMo-V2-Flash and Trinity Large Thinking share?

20 benchmarks have published results for both models. MiMo-V2-Flash has 21 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper