Model comparison

Mistral Large 4 vs Qwen2-72B

Mistral Large 4 is the stronger model overall, scoring 43.1 to 30.0 on the Noometry Index.

Last verified . 12 shared benchmarks.

Mistral Large 4 Mistral AI

43.1

Rank #99 Confirmed

Qwen2-72B Alibaba (Qwen)

30.0

Rank #300 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Mistral Large 4 scores higher in 7 categories and Qwen2-72B in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Large 4 leads 60.4 to 40.8.
  • Qwen2-72B has downloadable open weights; the other is API-only.

Side by side

Mistral Large 4 and Qwen2-72B specifications
Mistral Large 4Qwen2-72B
ProviderMistral AIAlibaba (Qwen)
Noometry Index43.130.0
Released2026-10-062024-06-07
WeightsProprietaryOpen
Context window1.05M—
Max output262K—
Input $ / M tokens$0.68—
Output $ / M tokens$2.09—
Results tracked1526

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large 4 leads

Mistral Large 4: 48.6 (#57), Qwen2-72B: 29.1 (#310)

Coding benchmarks
BenchmarkMistral Large 4Qwen2-72B
LMArena Coding14751196
LMArena WebDev1541—
WeirdML—11.3%
BigCodeBench Instruct—38.5%
BigCodeBench Complete—54%

Agentic & Tool Use Not comparable

Mistral Large 4: —, Qwen2-72B: 17.0 (#146)

Agentic & Tool Use benchmarks
BenchmarkMistral Large 4Qwen2-72B
TheAgentCompany—1.1%
METR Time Horizons—29.9%

Reasoning Too close to call

Mistral Large 4: 22.5 (#192), Qwen2-72B: 23.2 (#181)

Reasoning benchmarks
BenchmarkMistral Large 4Qwen2-72B
LMArena Hard Prompts14441191
NYT Connections (extended)27.4%—
Epoch Capabilities Index—125.28

Math Mistral Large 4 leads

Mistral Large 4: 40.4 (#91), Qwen2-72B: 30.2 (#236)

Math benchmarks
BenchmarkMistral Large 4Qwen2-72B
LMArena Math14881235
MATH Level 5—39.1%

Knowledge Mistral Large 4 leads

Mistral Large 4: 36.6 (#166), Qwen2-72B: 21.2 (#275)

Knowledge benchmarks
BenchmarkMistral Large 4Qwen2-72B
LMArena Expert14471171
GPQA Diamond—40.8%
SimpleQA Verified20%—
MMLU—82.4%

Multilingual Mistral Large 4 leads

Mistral Large 4: 52.6 (#82), Qwen2-72B: 35.9 (#244)

Multilingual benchmarks
BenchmarkMistral Large 4Qwen2-72B
LMArena Non-English14151176
LMArena Chinese14911240
LMArena Russian14141169
LMArena French—1170
LMArena German—1151
LMArena Japanese—1111
LMArena Korean—1083
LMArena Spanish—1169

Instruction Following Mistral Large 4 leads

Mistral Large 4: 75.0 (#76), Qwen2-72B: 61.7 (#241)

Instruction Following benchmarks
BenchmarkMistral Large 4Qwen2-72B
LMArena Instruction Following14241181

Long Context Mistral Large 4 leads

Mistral Large 4: 43.6 (#89), Qwen2-72B: 36.1 (#235)

Long Context benchmarks
BenchmarkMistral Large 4Qwen2-72B
LMArena Longer Query14291192

Writing & Preference Mistral Large 4 leads

Mistral Large 4: 60.4 (#97), Qwen2-72B: 40.8 (#241)

Writing & Preference benchmarks
BenchmarkMistral Large 4Qwen2-72B
LMArena Text14271203
LMArena Creative Writing13611181
LMArena Multi-Turn14241196

Frequently asked questions

Is Mistral Large 4 better than Qwen2-72B?

Mistral Large 4 is the stronger model overall, scoring 43.1 to 30.0 on the Noometry Index.

Is Mistral Large 4 or Qwen2-72B better for coding?

Mistral Large 4 scores higher on coding benchmarks: 48.6 versus 29.1 in the Noometry coding category.

How many benchmarks do Mistral Large 4 and Qwen2-72B share?

12 benchmarks have published results for both models. Mistral Large 4 has 15 scored results on Noometry and Qwen2-72B has 26.

Related comparisons

Go deeper