Model comparison

Codestral vs Qwen3 32B

Qwen3 32B is the stronger model overall, scoring 39.2 to 30.6 on the Noometry Index. Codestral costs 2.7× less per token, which makes it the better buy when Qwen3 32B's lead doesn't matter for your workload.

Last verified . 2 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

Qwen3 32B Alibaba (Qwen)

39.2

Rank #172 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Codestral scores higher in 0 categories and Qwen3 32B in 2 categories; one gap is clear of the uncertainty.
  • The widest gap is in coding, where Qwen3 32B leads 37.7 to 27.3.
  • The biggest single-benchmark swing is Aider Polyglot: 11.1% for Codestral and 40% for Qwen3 32B.
  • Codestral is cheaper at $0.30 / $0.90 per million input/output tokens, against $0.70 / $2.80 for Qwen3 32B.
  • Codestral accepts more context: 256K tokens versus 131K.
  • Qwen3 32B has downloadable open weights; the other is API-only.

Side by side

Codestral and Qwen3 32B specifications
CodestralQwen3 32B
ProviderMistral AIAlibaba (Qwen)
Noometry Index30.639.2
Released2024-05-292025-04
WeightsProprietaryOpen
Context window256K131K
Max output8K16K
Input $ / M tokens$0.30$0.70
Output $ / M tokens$0.90$2.80
Results tracked726

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 32B leads

Codestral: 27.3 (#321), Qwen3 32B: 37.7 (#190)

Coding benchmarks
BenchmarkCodestralQwen3 32B
Aider Polyglot11.1%40%
SciCode—35.4%
BigCodeBench Instruct41.8%—
LMArena Coding—1358
BigCodeBench Complete52.5%—
ALE-Bench137.78—
HumanEval+73.8%—
MBPP+61.9%—

Agentic & Tool Use Not comparable

Codestral: —, Qwen3 32B: 32.6 (#62)

Agentic & Tool Use benchmarks
BenchmarkCodestralQwen3 32B
Berkeley Function Calling Leaderboard—48.7%

Reasoning Too close to call

Codestral: 19.8 (#251), Qwen3 32B: 20.2 (#241)

Reasoning benchmarks
BenchmarkCodestralQwen3 32B
Kagi LLM Benchmark32.5%54.9%
CritPt—0.3%
Chess Puzzles—5%
LMArena Hard Prompts—1334
DTBench—67.5%
LMCA—17.3%
Epoch Capabilities Index—138.51

Math Not comparable

Codestral: —, Qwen3 32B: 39.7 (#99)

Math benchmarks
BenchmarkCodestralQwen3 32B
OTIS Mock AIME 2024-2025—66.9%
LMArena Math—1399

Knowledge Not comparable

Codestral: —, Qwen3 32B: 40.0 (#125)

Knowledge benchmarks
BenchmarkCodestralQwen3 32B
GPQA Diamond—65.7%
Vectara Hallucination Rate—5.9%
LMArena Expert—1362

Multilingual Not comparable

Codestral: —, Qwen3 32B: 45.6 (#167)

Multilingual benchmarks
BenchmarkCodestralQwen3 32B
LMArena Non-English—1317
LMArena Chinese—1357
LMArena German—1341
LMArena Russian—1311

Instruction Following Not comparable

Codestral: —, Qwen3 32B: 68.9 (#179)

Instruction Following benchmarks
BenchmarkCodestralQwen3 32B
LMArena Instruction Following—1305

Long Context Not comparable

Codestral: —, Qwen3 32B: 43.8 (#87)

Long Context benchmarks
BenchmarkCodestralQwen3 32B
Fiction.LiveBench—74.2%
LMArena Longer Query—1327

Writing & Preference Not comparable

Codestral: —, Qwen3 32B: 52.9 (#163)

Writing & Preference benchmarks
BenchmarkCodestralQwen3 32B
LMArena Text—1340
LMArena Creative Writing—1297
LMArena Multi-Turn—1331

Frequently asked questions

Is Codestral better than Qwen3 32B?

Qwen3 32B is the stronger model overall, scoring 39.2 to 30.6 on the Noometry Index. Codestral costs 2.7× less per token, which makes it the better buy when Qwen3 32B's lead doesn't matter for your workload.

Which is cheaper, Codestral or Qwen3 32B?

Codestral is cheaper. It lists at $0.30 per million input tokens and $0.90 per million output tokens; Qwen3 32B lists at $0.70 and $2.80.

Is Codestral or Qwen3 32B better for coding?

Qwen3 32B scores higher on coding benchmarks: 37.7 versus 27.3 in the Noometry coding category.

Which has the bigger context window?

Codestral does, with 256K tokens against 131K.

How many benchmarks do Codestral and Qwen3 32B share?

2 benchmarks have published results for both models. Codestral has 7 scored results on Noometry and Qwen3 32B has 26.

Related comparisons

Go deeper