Model comparison

Codestral vs Qwen Max

Qwen Max is the stronger model overall, scoring 34.7 to 30.6 on the Noometry Index. Codestral costs 6.2× less per token, which makes it the better buy when Qwen Max's lead doesn't matter for your workload.

Last verified . 1 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 1 benchmark with published results for both. Codestral scores higher in 0 categories and Qwen Max in 2 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen Max leads 25.1 to 19.8.
  • The biggest single-benchmark swing is Aider Polyglot: 11.1% for Codestral and 21.8% for Qwen Max.
  • Codestral is cheaper at $0.30 / $0.90 per million input/output tokens, against $1.60 / $6.40 for Qwen Max.
  • Codestral accepts more context: 256K tokens versus 33K.

Side by side

Codestral and Qwen Max specifications
CodestralQwen Max
ProviderMistral AIAlibaba (Qwen)
Noometry Index30.634.7
Released2024-05-292024-04-03
WeightsProprietaryProprietary
Context window256K33K
Max output8K8K
Input $ / M tokens$0.30$1.60
Output $ / M tokens$0.90$6.40
Results tracked723

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen Max leads

Codestral: 27.3 (#321), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkCodestralQwen Max
Aider Polyglot11.1%21.8%
BigCodeBench Instruct41.8%—
LMArena Coding—1288
BigCodeBench Complete52.5%—
ALE-Bench137.78—
HumanEval+73.8%—
MBPP+61.9%—

Reasoning Qwen Max leads

Codestral: 19.8 (#251), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkCodestralQwen Max
Kagi LLM Benchmark32.5%—
LMArena Hard Prompts—1269

Math Not comparable

Codestral: —, Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkCodestralQwen Max
OTIS Mock AIME 2024-2025—16.1%
LMArena Math—1275
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Not comparable

Codestral: —, Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkCodestralQwen Max
GPQA Diamond—56.1%
LMArena Expert—1248

Multilingual Not comparable

Codestral: —, Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkCodestralQwen Max
LMArena Non-English—1263
LMArena Chinese—1254
LMArena French—1330
LMArena German—1254
LMArena Japanese—1205
LMArena Korean—1142
LMArena Russian—1274
LMArena Spanish—1290

Instruction Following Not comparable

Codestral: —, Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkCodestralQwen Max
LMArena Instruction Following—1262

Long Context Not comparable

Codestral: —, Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkCodestralQwen Max
Fiction.LiveBench—66.7%
LMArena Longer Query—1288

Writing & Preference Not comparable

Codestral: —, Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkCodestralQwen Max
LMArena Text—1282
LMArena Creative Writing—1248
LMArena Multi-Turn—1277

Frequently asked questions

Is Codestral better than Qwen Max?

Qwen Max is the stronger model overall, scoring 34.7 to 30.6 on the Noometry Index. Codestral costs 6.2× less per token, which makes it the better buy when Qwen Max's lead doesn't matter for your workload.

Which is cheaper, Codestral or Qwen Max?

Codestral is cheaper. It lists at $0.30 per million input tokens and $0.90 per million output tokens; Qwen Max lists at $1.60 and $6.40.

Is Codestral or Qwen Max better for coding?

Qwen Max scores higher on coding benchmarks: 30.7 versus 27.3 in the Noometry coding category.

Which has the bigger context window?

Codestral does, with 256K tokens against 33K.

How many benchmarks do Codestral and Qwen Max share?

1 benchmark has published results for both models. Codestral has 7 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper