Model comparison

Mistral vs Qwen2.5 Plus 1127

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 29.9 on the Noometry Index.

Last verified . 14 shared benchmarks.

Mistral Mistral AI

29.9

Rank #303 Confirmed

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Mistral scores higher in 0 categories and Qwen2.5 Plus 1127 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen2.5 Plus 1127 leads 35.5 to 16.6.

Side by side

Mistral and Qwen2.5 Plus 1127 specifications
MistralQwen2.5 Plus 1127
ProviderMistral AIAlibaba (Qwen)
Noometry Index29.938.8
Released——
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2214

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 Plus 1127 leads

Mistral: 33.8 (#250), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkMistralQwen2.5 Plus 1127
LMArena Coding11621314

Reasoning Qwen2.5 Plus 1127 leads

Mistral: 22.2 (#200), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkMistralQwen2.5 Plus 1127
LMArena Hard Prompts11491299

Math Qwen2.5 Plus 1127 leads

Mistral: 22.3 (#278), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkMistralQwen2.5 Plus 1127
LMArena Math11801298
Omni-MATH7.2%—

Knowledge Qwen2.5 Plus 1127 leads

Mistral: 16.6 (#288), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkMistralQwen2.5 Plus 1127
LMArena Expert11251289
MMLU-Pro27.7%—
GPQA (HELM)30.3%—

Multilingual Qwen2.5 Plus 1127 leads

Mistral: 32.8 (#254), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkMistralQwen2.5 Plus 1127
LMArena Non-English11291265
LMArena Chinese11091314
LMArena German11551231
LMArena Japanese10131207
LMArena Russian11681271
LMArena French1180—
LMArena Korean1032—
LMArena Spanish1143—

Instruction Following Qwen2.5 Plus 1127 leads

Mistral: 52.6 (#288), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkMistralQwen2.5 Plus 1127
LMArena Instruction Following11521275
IFEval56.8%—

Long Context Qwen2.5 Plus 1127 leads

Mistral: 35.0 (#245), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkMistralQwen2.5 Plus 1127
LMArena Longer Query11531292

Writing & Preference Qwen2.5 Plus 1127 leads

Mistral: 37.0 (#260), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkMistralQwen2.5 Plus 1127
LMArena Text11651299
LMArena Creative Writing11581262
LMArena Multi-Turn11471299
WildBench66%—

Frequently asked questions

Is Mistral better than Qwen2.5 Plus 1127?

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 29.9 on the Noometry Index.

Is Mistral or Qwen2.5 Plus 1127 better for coding?

Qwen2.5 Plus 1127 scores higher on coding benchmarks: 38.5 versus 33.8 in the Noometry coding category.

How many benchmarks do Mistral and Qwen2.5 Plus 1127 share?

14 benchmarks have published results for both models. Mistral has 22 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper