Model comparison

Mistral Nemo vs Qwen2.5 Plus 1127

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 26.4 on the Noometry Index.

Last verified . 0 shared benchmarks.

Mistral Nemo Mistral AI

26.4

Rank #337 Confirmed

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Summary

  • The widest gap is in knowledge, where Qwen2.5 Plus 1127 leads 35.5 to 12.3.
  • Mistral Nemo has downloadable open weights; the other is API-only.

Side by side

Mistral Nemo and Qwen2.5 Plus 1127 specifications
Mistral NemoQwen2.5 Plus 1127
ProviderMistral AIAlibaba (Qwen)
Noometry Index26.438.8
Released2024-07-01—
WeightsOpenProprietary
Context window128K—
Max output128K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.15—
Results tracked1014

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral Nemo: —, Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkMistral NemoQwen2.5 Plus 1127
LMArena Coding—1314

Agentic & Tool Use Not comparable

Mistral Nemo: 23.5 (#125), Qwen2.5 Plus 1127: —

Agentic & Tool Use benchmarks
BenchmarkMistral NemoQwen2.5 Plus 1127
Berkeley Function Calling Leaderboard27.6%—
BALROG17.6%—

Reasoning Qwen2.5 Plus 1127 leads

Mistral Nemo: 20.7 (#232), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkMistral NemoQwen2.5 Plus 1127
LMArena Hard Prompts—1299
DTBench48.6%—
Epoch Capabilities Index118.68—
PIQA83.5%—

Math Qwen2.5 Plus 1127 leads

Mistral Nemo: 25.5 (#268), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkMistral NemoQwen2.5 Plus 1127
LMArena Math—1298
MATH Level 510.8%—
GSM8K84.2%—

Knowledge Qwen2.5 Plus 1127 leads

Mistral Nemo: 12.3 (#298), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkMistral NemoQwen2.5 Plus 1127
GPQA Diamond29.9%—
LMArena Expert—1289
BoolQ82.5%—

Multilingual Not comparable

Mistral Nemo: —, Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkMistral NemoQwen2.5 Plus 1127
LMArena Non-English—1265
LMArena Chinese—1314
LMArena German—1231
LMArena Japanese—1207
LMArena Russian—1271

Instruction Following Not comparable

Mistral Nemo: —, Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkMistral NemoQwen2.5 Plus 1127
LMArena Instruction Following—1275

Long Context Not comparable

Mistral Nemo: —, Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkMistral NemoQwen2.5 Plus 1127
LMArena Longer Query—1292

Writing & Preference Qwen2.5 Plus 1127 leads

Mistral Nemo: 28.5 (#296), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkMistral NemoQwen2.5 Plus 1127
LMArena Text—1299
LMArena Creative Writing—1262
EQ-Bench Creative Writing881—
LMArena Multi-Turn—1299

Frequently asked questions

Is Mistral Nemo better than Qwen2.5 Plus 1127?

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 26.4 on the Noometry Index.

How many benchmarks do Mistral Nemo and Qwen2.5 Plus 1127 share?

0 benchmarks have published results for both models. Mistral Nemo has 10 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper