Model comparison

Mistral Small 3.1 vs Qwen2.5 Plus 1127

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 31.7 on the Noometry Index.

Last verified . 14 shared benchmarks.

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Mistral Small 3.1 scores higher in 1 category and Qwen2.5 Plus 1127 in 7 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen2.5 Plus 1127 leads 36.1 to 14.7.
  • Mistral Small 3.1 has downloadable open weights; the other is API-only.

Side by side

Mistral Small 3.1 and Qwen2.5 Plus 1127 specifications
Mistral Small 3.1Qwen2.5 Plus 1127
ProviderMistral AIAlibaba (Qwen)
Noometry Index31.738.8
Released2025-03-17—
WeightsOpenProprietary
Context window128K—
Max output102K—
Input $ / M tokens$0.35—
Output $ / M tokens$0.56—
Results tracked2814

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mistral Small 3.1: 38.3 (#179), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkMistral Small 3.1Qwen2.5 Plus 1127
LMArena Coding13091314

Reasoning Qwen2.5 Plus 1127 leads

Mistral Small 3.1: 19.7 (#254), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkMistral Small 3.1Qwen2.5 Plus 1127
LMArena Hard Prompts12781299
Chess Puzzles1%—
Epoch Capabilities Index127.48—

Math Qwen2.5 Plus 1127 leads

Mistral Small 3.1: 14.7 (#301), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkMistral Small 3.1Qwen2.5 Plus 1127
LMArena Math12621298
OTIS Mock AIME 2024-20253.9%—
Omni-MATH24.8%—

Knowledge Qwen2.5 Plus 1127 leads

Mistral Small 3.1: 22.6 (#271), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkMistral Small 3.1Qwen2.5 Plus 1127
LMArena Expert12571289
GPQA Diamond41.9%—
MMLU-Pro61%—
GPQA (HELM)39.2%—

Multimodal Not comparable

Mistral Small 3.1: 33.2 (#99), Qwen2.5 Plus 1127: —

Multimodal benchmarks
BenchmarkMistral Small 3.1Qwen2.5 Plus 1127
LMArena Vision1136—

Multilingual Too close to call

Mistral Small 3.1: 41.2 (#209), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkMistral Small 3.1Qwen2.5 Plus 1127
LMArena Non-English12551265
LMArena Chinese12531314
LMArena German12661231
LMArena Japanese12081207
LMArena Russian12631271
LMArena French1273—
LMArena Korean1206—
LMArena Spanish1283—

Instruction Following Qwen2.5 Plus 1127 leads

Mistral Small 3.1: 63.6 (#230), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkMistral Small 3.1Qwen2.5 Plus 1127
LMArena Instruction Following12641275
IFEval75%—

Long Context Too close to call

Mistral Small 3.1: 39.5 (#178), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkMistral Small 3.1Qwen2.5 Plus 1127
LMArena Longer Query12991292

Writing & Preference Qwen2.5 Plus 1127 leads

Mistral Small 3.1: 37.0 (#259), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkMistral Small 3.1Qwen2.5 Plus 1127
LMArena Text12771299
LMArena Creative Writing12531262
LMArena Multi-Turn12701299
EQ-Bench Creative Writing761—
WildBench78.8%—

Frequently asked questions

Is Mistral Small 3.1 better than Qwen2.5 Plus 1127?

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 31.7 on the Noometry Index.

Is Mistral Small 3.1 or Qwen2.5 Plus 1127 better for coding?

They score almost the same on coding (38.3 vs 38.5); test both on your own repository before choosing.

How many benchmarks do Mistral Small 3.1 and Qwen2.5 Plus 1127 share?

14 benchmarks have published results for both models. Mistral Small 3.1 has 28 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper