Model comparison

DeepSeek-R1-Distill-Llama-70B vs MiMo-V2.5

MiMo-V2.5 is the stronger model overall, scoring 43.4 to 37.8 on the Noometry Index.

Last verified . 0 shared benchmarks.

MiMo-V2.5 Xiaomi

43.4

Rank #93 Confirmed

Summary

  • The widest gap is in writing & preference, where MiMo-V2.5 leads 61.6 to 49.0.

Side by side

DeepSeek-R1-Distill-Llama-70B and MiMo-V2.5 specifications
DeepSeek-R1-Distill-Llama-70BMiMo-V2.5
ProviderDeepSeekXiaomi
Noometry Index37.843.4
Released2025-01-202026-04-22
WeightsOpenOpen
Context window—1.05M
Max output—131K
Input $ / M tokens—$0.14
Output $ / M tokens—$0.28
Results tracked1323

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiMo-V2.5 leads

DeepSeek-R1-Distill-Llama-70B: 36.8 (#202), MiMo-V2.5: 43.9 (#81)

Coding benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BMiMo-V2.5
LMArena WebDev—1438
SciCode—43.1%
BigCodeBench Instruct35.3%—
LiveBench Coding51.6%—
LMArena Coding—1469
BigCodeBench Complete49.9%—
ALE-Bench—513.95

Reasoning MiMo-V2.5 leads

DeepSeek-R1-Distill-Llama-70B: 24.9 (#156), MiMo-V2.5: 28.6 (#101)

Reasoning benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BMiMo-V2.5
Kagi LLM Benchmark52.3%—
CritPt—3.7%
LiveBench Reasoning67.6%—
LMArena Hard Prompts—1450
LiveBench Data Analysis55.9%—
LiveBench54.5%—

Math Too close to call

DeepSeek-R1-Distill-Llama-70B: 36.0 (#176), MiMo-V2.5: 36.8 (#163)

Math benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BMiMo-V2.5
OTIS Mock AIME 2024-202551.4%—
ProofBench—16%
LiveBench Math58.1%—
LMArena Math—1436
MATH Level 589.9%—

Knowledge MiMo-V2.5 leads

DeepSeek-R1-Distill-Llama-70B: 30.7 (#225), MiMo-V2.5: 40.8 (#115)

Knowledge benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BMiMo-V2.5
GPQA Diamond55.7%—
LMArena Expert—1460

Multimodal Not comparable

DeepSeek-R1-Distill-Llama-70B: —, MiMo-V2.5: 39.8 (#54)

Multimodal benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BMiMo-V2.5
LMArena Vision—1247

Multilingual Not comparable

DeepSeek-R1-Distill-Llama-70B: —, MiMo-V2.5: 51.9 (#99)

Multilingual benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BMiMo-V2.5
LMArena Non-English—1404
LMArena Chinese—1468
LMArena French—1447
LMArena German—1421
LMArena Japanese—1306
LMArena Korean—1363
LMArena Russian—1395
LMArena Spanish—1416

Instruction Following MiMo-V2.5 leads

DeepSeek-R1-Distill-Llama-70B: 68.2 (#190), MiMo-V2.5: 75.5 (#60)

Instruction Following benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BMiMo-V2.5
LiveBench Instruction Following69.9%—
LMArena Instruction Following—1434

Long Context Not comparable

DeepSeek-R1-Distill-Llama-70B: —, MiMo-V2.5: 44.2 (#73)

Long Context benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BMiMo-V2.5
LMArena Longer Query—1445

Writing & Preference MiMo-V2.5 leads

DeepSeek-R1-Distill-Llama-70B: 49.0 (#194), MiMo-V2.5: 61.6 (#86)

Writing & Preference benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BMiMo-V2.5
LMArena Text—1428
LMArena Creative Writing—1393
LMArena Multi-Turn—1445
LiveBench Language23.8%—

Frequently asked questions

Is DeepSeek-R1-Distill-Llama-70B better than MiMo-V2.5?

MiMo-V2.5 is the stronger model overall, scoring 43.4 to 37.8 on the Noometry Index.

Is DeepSeek-R1-Distill-Llama-70B or MiMo-V2.5 better for coding?

MiMo-V2.5 scores higher on coding benchmarks: 43.9 versus 36.8 in the Noometry coding category.

How many benchmarks do DeepSeek-R1-Distill-Llama-70B and MiMo-V2.5 share?

0 benchmarks have published results for both models. DeepSeek-R1-Distill-Llama-70B has 13 scored results on Noometry and MiMo-V2.5 has 23.

Related comparisons

Go deeper