Model comparison

Mistral Large vs Qwen Turbo

Mistral Large is the stronger model overall, scoring 31.9 to 27.1 on the Noometry Index. Qwen Turbo costs 34× less per token, which makes it the better buy when Mistral Large's lead doesn't matter for your workload.

Last verified . 3 shared benchmarks.

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Qwen Turbo Alibaba (Qwen)

27.1

Rank #335 Reported

Summary

  • They share 3 benchmarks with published results for both. Mistral Large scores higher in 2 categories and Qwen Turbo in 0 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Mistral Large leads 30.1 to 22.2.
  • The biggest single-benchmark swing is GPQA Diamond: 51.3% for Mistral Large and 41.8% for Qwen Turbo.
  • Qwen Turbo is cheaper at $0.05 / $0.20 per million input/output tokens, against $2 / $6 for Mistral Large.
  • Qwen Turbo accepts more context: 1M tokens versus 131K.
  • Mistral Large has downloadable open weights; the other is API-only.

Side by side

Mistral Large and Qwen Turbo specifications
Mistral LargeQwen Turbo
ProviderMistral AIAlibaba (Qwen)
Noometry Index31.927.1
Released2024-02-262024-11-01
WeightsOpenProprietary
Context window131K1M
Max output16K16K
Input $ / M tokens$2$0.05
Output $ / M tokens$6$0.20
Results tracked513

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral Large: 34.3 (#240), Qwen Turbo: —

Coding benchmarks
BenchmarkMistral LargeQwen Turbo
SciCode36.2%—
BigCodeBench Instruct30%—
LiveBench Coding47.1%—
LMArena Coding1277—
BigCodeBench Complete38.3%—
ALE-Bench264.7—
HumanEval+62.2%—
MBPP+59.5%—

Agentic & Tool Use Not comparable

Mistral Large: 28.6 (#89), Qwen Turbo: —

Agentic & Tool Use benchmarks
BenchmarkMistral LargeQwen Turbo
Berkeley Function Calling Leaderboard38.4%—

Reasoning Not comparable

Mistral Large: 15.8 (#310), Qwen Turbo: —

Reasoning benchmarks
BenchmarkMistral LargeQwen Turbo
SimpleBench22.5%—
CritPt0%—
LiveBench Reasoning43.5%—
LMArena Hard Prompts1257—
DTBench65.1%—
LiveBench Data Analysis50.1%—
LMCA16.7%—
Epoch Capabilities Index128.52—
ForecastBench57.1—
LiveBench48.4%—

Math Mistral Large leads

Mistral Large: 18.2 (#291), Qwen Turbo: 15.3 (#297)

Math benchmarks
BenchmarkMistral LargeQwen Turbo
OTIS Mock AIME 2024-20258.5%6.1%
MATH Level 550.3%56.2%
Omni-MATH28.1%—
LiveBench Math42.5%—
LMArena Math1262—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Mistral Large leads

Mistral Large: 30.1 (#230), Qwen Turbo: 22.2 (#272)

Knowledge benchmarks
BenchmarkMistral LargeQwen Turbo
GPQA Diamond51.3%41.8%
MMLU-Pro59.9%—
Confabulations21.4%—
Vectara Hallucination Rate4.5%—
GPQA (HELM)43.5%—
LMArena Expert1232—
MMLU80%—

Multilingual Not comparable

Mistral Large: 40.0 (#219), Qwen Turbo: —

Multilingual benchmarks
BenchmarkMistral LargeQwen Turbo
LMArena Non-English1237—
LMArena Chinese1240—
LMArena French1325—
LMArena German1254—
LMArena Japanese1188—
LMArena Korean1202—
LMArena Russian1257—
LMArena Spanish1268—

Instruction Following Not comparable

Mistral Large: 67.9 (#191), Qwen Turbo: —

Instruction Following benchmarks
BenchmarkMistral LargeQwen Turbo
LiveBench Instruction Following67.9%—
IFEval87.7%—
LMArena Instruction Following1249—

Long Context Not comparable

Mistral Large: 38.3 (#199), Qwen Turbo: —

Long Context benchmarks
BenchmarkMistral LargeQwen Turbo
LMArena Longer Query1261—

Writing & Preference Not comparable

Mistral Large: 40.7 (#242), Qwen Turbo: —

Writing & Preference benchmarks
BenchmarkMistral LargeQwen Turbo
LMArena Text1266—
LMArena Creative Writing1243—
Short-Story Creative Writing69%—
EQ-Bench Creative Writing985—
WildBench80.1%—
LMArena Multi-Turn1260—
LiveBench Language39.4%—

Frequently asked questions

Is Mistral Large better than Qwen Turbo?

Mistral Large is the stronger model overall, scoring 31.9 to 27.1 on the Noometry Index. Qwen Turbo costs 34× less per token, which makes it the better buy when Mistral Large's lead doesn't matter for your workload.

Which is cheaper, Mistral Large or Qwen Turbo?

Qwen Turbo is cheaper. It lists at $0.05 per million input tokens and $0.20 per million output tokens; Mistral Large lists at $2 and $6.

Which has the bigger context window?

Qwen Turbo does, with 1M tokens against 131K.

How many benchmarks do Mistral Large and Qwen Turbo share?

3 benchmarks have published results for both models. Mistral Large has 51 scored results on Noometry and Qwen Turbo has 3.

Related comparisons

Go deeper