Model comparison

Mistral Medium 3.5 vs Qwen2.5 Plus 1127

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 38.8 on the Noometry Index.

Last verified . 13 shared benchmarks.

Mistral Medium 3.5 Mistral AI

40.2

Rank #152 Confirmed

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Mistral Medium 3.5 scores higher in 6 categories and Qwen2.5 Plus 1127 in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Mistral Medium 3.5 leads 51.9 to 41.9.
  • Mistral Medium 3.5 has downloadable open weights; the other is API-only.

Side by side

Mistral Medium 3.5 and Qwen2.5 Plus 1127 specifications
Mistral Medium 3.5Qwen2.5 Plus 1127
ProviderMistral AIAlibaba (Qwen)
Noometry Index40.238.8
Released——
WeightsOpenProprietary
Context window262K—
Max output210K—
Input $ / M tokens$1.50—
Output $ / M tokens$7.50—
Results tracked2214

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 Plus 1127 leads

Mistral Medium 3.5: 36.0 (#213), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkMistral Medium 3.5Qwen2.5 Plus 1127
LMArena Coding14611314
LMArena WebDev1264—

Reasoning Qwen2.5 Plus 1127 leads

Mistral Medium 3.5: 17.3 (#295), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkMistral Medium 3.5Qwen2.5 Plus 1127
LMArena Hard Prompts14361299
Kagi LLM Benchmark41.4%—
NYT Connections (extended)12.9%—
Epoch Capabilities Index141.35—

Math Mistral Medium 3.5 leads

Mistral Medium 3.5: 39.1 (#113), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkMistral Medium 3.5Qwen2.5 Plus 1127
LMArena Math14311298

Knowledge Mistral Medium 3.5 leads

Mistral Medium 3.5: 40.0 (#126), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkMistral Medium 3.5Qwen2.5 Plus 1127
LMArena Expert14321289

Multimodal Not comparable

Mistral Medium 3.5: 38.3 (#65), Qwen2.5 Plus 1127: —

Multimodal benchmarks
BenchmarkMistral Medium 3.5Qwen2.5 Plus 1127
LMArena Vision1223—

Multilingual Mistral Medium 3.5 leads

Mistral Medium 3.5: 51.9 (#100), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkMistral Medium 3.5Qwen2.5 Plus 1127
LMArena Non-English14041265
LMArena Chinese14421314
LMArena German14511231
LMArena Russian13951271
LMArena French1448—
LMArena Japanese—1207
LMArena Korean1385—
LMArena Spanish1409—

Instruction Following Mistral Medium 3.5 leads

Mistral Medium 3.5: 74.6 (#90), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkMistral Medium 3.5Qwen2.5 Plus 1127
LMArena Instruction Following14151275

Long Context Mistral Medium 3.5 leads

Mistral Medium 3.5: 43.2 (#103), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkMistral Medium 3.5Qwen2.5 Plus 1127
LMArena Longer Query14151292

Writing & Preference Mistral Medium 3.5 leads

Mistral Medium 3.5: 58.5 (#117), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkMistral Medium 3.5Qwen2.5 Plus 1127
LMArena Text14211299
LMArena Creative Writing13741262
LMArena Multi-Turn14231299
EQ-Bench 4993—

Frequently asked questions

Is Mistral Medium 3.5 better than Qwen2.5 Plus 1127?

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 38.8 on the Noometry Index.

Is Mistral Medium 3.5 or Qwen2.5 Plus 1127 better for coding?

Qwen2.5 Plus 1127 scores higher on coding benchmarks: 38.5 versus 36.0 in the Noometry coding category.

How many benchmarks do Mistral Medium 3.5 and Qwen2.5 Plus 1127 share?

13 benchmarks have published results for both models. Mistral Medium 3.5 has 22 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper