Model comparison

Codestral vs Qwen3.5 Max Preview

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 30.6 on the Noometry Index.

Last verified . 0 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

Qwen3.5 Max Preview Alibaba (Qwen)

45.3

Rank #71 Confirmed

Summary

  • The widest gap is in coding, where Qwen3.5 Max Preview leads 44.0 to 27.3.

Side by side

Codestral and Qwen3.5 Max Preview specifications
CodestralQwen3.5 Max Preview
ProviderMistral AIAlibaba (Qwen)
Noometry Index30.645.3
Released2024-05-29—
WeightsProprietaryProprietary
Context window256K—
Max output8K—
Input $ / M tokens$0.30—
Output $ / M tokens$0.90—
Results tracked717

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5 Max Preview leads

Codestral: 27.3 (#321), Qwen3.5 Max Preview: 44.0 (#77)

Coding benchmarks
BenchmarkCodestralQwen3.5 Max Preview
Aider Polyglot11.1%—
BigCodeBench Instruct41.8%—
LMArena Coding—1487
BigCodeBench Complete52.5%—
ALE-Bench137.78—
HumanEval+73.8%—
MBPP+61.9%—

Reasoning Qwen3.5 Max Preview leads

Codestral: 19.8 (#251), Qwen3.5 Max Preview: 30.8 (#84)

Reasoning benchmarks
BenchmarkCodestralQwen3.5 Max Preview
Kagi LLM Benchmark32.5%—
LMArena Hard Prompts—1483

Math Not comparable

Codestral: —, Qwen3.5 Max Preview: 40.1 (#94)

Math benchmarks
BenchmarkCodestralQwen3.5 Max Preview
LMArena Math—1474

Knowledge Not comparable

Codestral: —, Qwen3.5 Max Preview: 41.8 (#107)

Knowledge benchmarks
BenchmarkCodestralQwen3.5 Max Preview
LMArena Expert—1489

Multilingual Not comparable

Codestral: —, Qwen3.5 Max Preview: 56.2 (#22)

Multilingual benchmarks
BenchmarkCodestralQwen3.5 Max Preview
LMArena Non-English—1465
LMArena Chinese—1534
LMArena French—1484
LMArena German—1487
LMArena Japanese—1495
LMArena Korean—1438
LMArena Russian—1471
LMArena Spanish—1470

Instruction Following Not comparable

Codestral: —, Qwen3.5 Max Preview: 77.0 (#31)

Instruction Following benchmarks
BenchmarkCodestralQwen3.5 Max Preview
LMArena Instruction Following—1467

Long Context Not comparable

Codestral: —, Qwen3.5 Max Preview: 45.2 (#45)

Long Context benchmarks
BenchmarkCodestralQwen3.5 Max Preview
LMArena Longer Query—1476

Writing & Preference Not comparable

Codestral: —, Qwen3.5 Max Preview: 66.0 (#41)

Writing & Preference benchmarks
BenchmarkCodestralQwen3.5 Max Preview
LMArena Text—1470
LMArena Creative Writing—1464
LMArena Multi-Turn—1478

Frequently asked questions

Is Codestral better than Qwen3.5 Max Preview?

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 30.6 on the Noometry Index.

Is Codestral or Qwen3.5 Max Preview better for coding?

Qwen3.5 Max Preview scores higher on coding benchmarks: 44.0 versus 27.3 in the Noometry coding category.

How many benchmarks do Codestral and Qwen3.5 Max Preview share?

0 benchmarks have published results for both models. Codestral has 7 scored results on Noometry and Qwen3.5 Max Preview has 17.

Related comparisons

Go deeper