Model comparison

Mixtral 8x7B vs Qwen2.5 Plus 1127

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 27.1 on the Noometry Index.

Last verified . 14 shared benchmarks.

Mixtral 8x7B Mistral AI

27.1

Rank #334 Confirmed

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Mixtral 8x7B scores higher in 0 categories and Qwen2.5 Plus 1127 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen2.5 Plus 1127 leads 35.5 to 11.0.
  • Mixtral 8x7B has downloadable open weights; the other is API-only.

Side by side

Mixtral 8x7B and Qwen2.5 Plus 1127 specifications
Mixtral 8x7BQwen2.5 Plus 1127
ProviderMistral AIAlibaba (Qwen)
Noometry Index27.138.8
Released2023-12-11—
WeightsOpenProprietary
Context window32K—
Max output32K—
Input $ / M tokens$0.70—
Output $ / M tokens$0.70—
Results tracked3814

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 Plus 1127 leads

Mixtral 8x7B: 32.8 (#269), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkMixtral 8x7BQwen2.5 Plus 1127
LMArena Coding11261314
HumanEval+39.6%—
MBPP+49.7%—

Reasoning Qwen2.5 Plus 1127 leads

Mixtral 8x7B: 18.2 (#285), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkMixtral 8x7BQwen2.5 Plus 1127
LMArena Hard Prompts11151299
DTBench49.6%—
Adversarial NLI55.2%—
Epoch Capabilities Index118.47—
ForecastBench56.3—
HellaSwag86.7%—
PIQA83.6%—
WinoGrande77.2%—

Math Qwen2.5 Plus 1127 leads

Mixtral 8x7B: 18.8 (#289), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkMixtral 8x7BQwen2.5 Plus 1127
LMArena Math11471298
Omni-MATH10.5%—
MATH Level 510%—
GSM8K74.4%—

Knowledge Qwen2.5 Plus 1127 leads

Mixtral 8x7B: 11.0 (#301), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkMixtral 8x7BQwen2.5 Plus 1127
LMArena Expert10881289
GPQA Diamond30.6%—
MMLU-Pro33.5%—
GPQA (HELM)29.6%—
ARC (AI2) Challenge87.3%—
MMLU70.6%—
OpenBookQA85.8%—
TriviaQA82.2%—

Multilingual Qwen2.5 Plus 1127 leads

Mixtral 8x7B: 29.6 (#266), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkMixtral 8x7BQwen2.5 Plus 1127
LMArena Non-English10771265
LMArena Chinese10551314
LMArena German11141231
LMArena Japanese9311207
LMArena Russian10901271
LMArena French1166—
LMArena Korean968—
LMArena Spanish1111—

Instruction Following Qwen2.5 Plus 1127 leads

Mixtral 8x7B: 51.0 (#297), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkMixtral 8x7BQwen2.5 Plus 1127
LMArena Instruction Following11091275
IFEval57.5%—

Long Context Qwen2.5 Plus 1127 leads

Mixtral 8x7B: 33.4 (#260), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkMixtral 8x7BQwen2.5 Plus 1127
LMArena Longer Query11031292

Writing & Preference Qwen2.5 Plus 1127 leads

Mixtral 8x7B: 34.2 (#270), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkMixtral 8x7BQwen2.5 Plus 1127
LMArena Text11321299
LMArena Creative Writing11091262
LMArena Multi-Turn11151299
WildBench67.3%—

Frequently asked questions

Is Mixtral 8x7B better than Qwen2.5 Plus 1127?

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 27.1 on the Noometry Index.

Is Mixtral 8x7B or Qwen2.5 Plus 1127 better for coding?

Qwen2.5 Plus 1127 scores higher on coding benchmarks: 38.5 versus 32.8 in the Noometry coding category.

How many benchmarks do Mixtral 8x7B and Qwen2.5 Plus 1127 share?

14 benchmarks have published results for both models. Mixtral 8x7B has 38 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper