Model comparison

Mistral 7B vs Qwen2.5 Plus 1127

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 23.0 on the Noometry Index.

Last verified . 14 shared benchmarks.

Mistral 7B Mistral AI

23.0

Rank #351 Confirmed

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Mistral 7B scores higher in 0 categories and Qwen2.5 Plus 1127 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen2.5 Plus 1127 leads 35.5 to 7.4.
  • Mistral 7B has downloadable open weights; the other is API-only.

Side by side

Mistral 7B and Qwen2.5 Plus 1127 specifications
Mistral 7BQwen2.5 Plus 1127
ProviderMistral AIAlibaba (Qwen)
Noometry Index23.038.8
Released2023-09-27—
WeightsOpenProprietary
Context window8K—
Max output8K—
Input $ / M tokens$0.25—
Output $ / M tokens$0.25—
Results tracked3714

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 Plus 1127 leads

Mistral 7B: 26.4 (#326), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkMistral 7BQwen2.5 Plus 1127
LMArena Coding10821314
BigCodeBench Instruct19.5%—
BigCodeBench Complete27.3%—
HumanEval+36%—
MBPP+42.1%—

Reasoning Qwen2.5 Plus 1127 leads

Mistral 7B: 13.1 (#336), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkMistral 7BQwen2.5 Plus 1127
LMArena Hard Prompts10671299
Chess Puzzles0%—
DTBench42.5%—
Adversarial NLI47.1%—
BIG-Bench Hard56.1%—
Epoch Capabilities Index112.21—
HellaSwag81%—
PIQA83%—
WinoGrande75.3%—

Math Qwen2.5 Plus 1127 leads

Mistral 7B: 8.1 (#325), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkMistral 7BQwen2.5 Plus 1127
LMArena Math10851298
OTIS Mock AIME 2024-20250.3%—
MATH Level 53.7%—
GSM8K54.4%—

Knowledge Qwen2.5 Plus 1127 leads

Mistral 7B: 7.4 (#311), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkMistral 7BQwen2.5 Plus 1127
LMArena Expert10361289
GPQA Diamond15.2%—
ARC (AI2) Challenge78.6%—
BoolQ87.4%—
MMLU62.5%—
OpenBookQA79.8%—
TriviaQA75.2%—

Multilingual Qwen2.5 Plus 1127 leads

Mistral 7B: 25.8 (#283), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkMistral 7BQwen2.5 Plus 1127
LMArena Non-English10121265
LMArena Chinese10091314
LMArena German9871231
LMArena Japanese8781207
LMArena Russian10181271
LMArena French1037—
LMArena Spanish1026—

Instruction Following Qwen2.5 Plus 1127 leads

Mistral 7B: 54.2 (#280), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkMistral 7BQwen2.5 Plus 1127
LMArena Instruction Following10601275

Long Context Qwen2.5 Plus 1127 leads

Mistral 7B: 32.2 (#271), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkMistral 7BQwen2.5 Plus 1127
LMArena Longer Query10601292

Writing & Preference Qwen2.5 Plus 1127 leads

Mistral 7B: 30.7 (#286), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkMistral 7BQwen2.5 Plus 1127
LMArena Text10901299
LMArena Creative Writing10681262
LMArena Multi-Turn10621299

Frequently asked questions

Is Mistral 7B better than Qwen2.5 Plus 1127?

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 23.0 on the Noometry Index.

Is Mistral 7B or Qwen2.5 Plus 1127 better for coding?

Qwen2.5 Plus 1127 scores higher on coding benchmarks: 38.5 versus 26.4 in the Noometry coding category.

How many benchmarks do Mistral 7B and Qwen2.5 Plus 1127 share?

14 benchmarks have published results for both models. Mistral 7B has 37 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper