Model comparison

Codestral vs Qwen3.7 Plus

Qwen3.7 Plus is the stronger model overall, scoring 45.3 to 30.6 on the Noometry Index. Codestral costs 1.6× less per token, which makes it the better buy when Qwen3.7 Plus's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

Qwen3.7 Plus Alibaba (Qwen)

45.3

Rank #72 Confirmed

Summary

  • The widest gap is in reasoning, where Qwen3.7 Plus leads 39.3 to 19.8.
  • Codestral is cheaper at $0.30 / $0.90 per million input/output tokens, against $0.40 / $1.60 for Qwen3.7 Plus.
  • Qwen3.7 Plus accepts more context: 1M tokens versus 256K.

Side by side

Codestral and Qwen3.7 Plus specifications
CodestralQwen3.7 Plus
ProviderMistral AIAlibaba (Qwen)
Noometry Index30.645.3
Released2024-05-292026-06-02
WeightsProprietaryProprietary
Context window256K1M
Max output8K131K
Input $ / M tokens$0.30$0.40
Output $ / M tokens$0.90$1.60
Results tracked732

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.7 Plus leads

Codestral: 27.3 (#321), Qwen3.7 Plus: 36.6 (#206)

Coding benchmarks
BenchmarkCodestralQwen3.7 Plus
FrontierCode—10.2%
Aider Polyglot11.1%—
SciCode—45.5%
BigCodeBench Instruct41.8%—
LMArena Coding—1473
BigCodeBench Complete52.5%—
ALE-Bench137.78—
HumanEval+73.8%—
MBPP+61.9%—

Agentic & Tool Use Not comparable

Codestral: —, Qwen3.7 Plus: 21.4 (#138)

Agentic & Tool Use benchmarks
BenchmarkCodestralQwen3.7 Plus
OSWorld 2.0—2.8%

Reasoning Qwen3.7 Plus leads

Codestral: 19.8 (#251), Qwen3.7 Plus: 39.3 (#59)

Reasoning benchmarks
BenchmarkCodestralQwen3.7 Plus
Kagi LLM Benchmark32.5%—
NYT Connections (extended)—74.8%
CritPt—9.1%
Chess Puzzles—24%
LMArena Hard Prompts—1460
Mystery Game Puzzles—17%
DTBench—84%
LMCA—37.6%
Epoch Capabilities Index—147.37

Math Not comparable

Codestral: —, Qwen3.7 Plus: 50.5 (#56)

Math benchmarks
BenchmarkCodestralQwen3.7 Plus
FrontierMath (Tiers 1-3)—34.4%
OTIS Mock AIME 2024-2025—93.3%
LMArena Math—1466

Knowledge Not comparable

Codestral: —, Qwen3.7 Plus: 54.9 (#51)

Knowledge benchmarks
BenchmarkCodestralQwen3.7 Plus
GPQA Diamond—87.9%
LMArena Expert—1467

Multimodal Not comparable

Codestral: —, Qwen3.7 Plus: 41.8 (#33)

Multimodal benchmarks
BenchmarkCodestralQwen3.7 Plus
LMArena Vision—1279
LMArena Document—1444

Multilingual Not comparable

Codestral: —, Qwen3.7 Plus: 54.8 (#38)

Multilingual benchmarks
BenchmarkCodestralQwen3.7 Plus
LMArena Non-English—1445
LMArena Chinese—1510
LMArena French—1473
LMArena German—1471
LMArena Japanese—1413
LMArena Korean—1415
LMArena Russian—1457
LMArena Spanish—1457

Instruction Following Not comparable

Codestral: —, Qwen3.7 Plus: 75.8 (#52)

Instruction Following benchmarks
BenchmarkCodestralQwen3.7 Plus
LMArena Instruction Following—1440

Long Context Not comparable

Codestral: —, Qwen3.7 Plus: 44.5 (#65)

Long Context benchmarks
BenchmarkCodestralQwen3.7 Plus
LMArena Longer Query—1455

Writing & Preference Not comparable

Codestral: —, Qwen3.7 Plus: 64.3 (#56)

Writing & Preference benchmarks
BenchmarkCodestralQwen3.7 Plus
LMArena Text—1455
LMArena Creative Writing—1439
LMArena Multi-Turn—1460

Frequently asked questions

Is Codestral better than Qwen3.7 Plus?

Qwen3.7 Plus is the stronger model overall, scoring 45.3 to 30.6 on the Noometry Index. Codestral costs 1.6× less per token, which makes it the better buy when Qwen3.7 Plus's lead doesn't matter for your workload.

Which is cheaper, Codestral or Qwen3.7 Plus?

Codestral is cheaper. It lists at $0.30 per million input tokens and $0.90 per million output tokens; Qwen3.7 Plus lists at $0.40 and $1.60.

Is Codestral or Qwen3.7 Plus better for coding?

Qwen3.7 Plus scores higher on coding benchmarks: 36.6 versus 27.3 in the Noometry coding category.

Which has the bigger context window?

Qwen3.7 Plus does, with 1M tokens against 256K.

How many benchmarks do Codestral and Qwen3.7 Plus share?

0 benchmarks have published results for both models. Codestral has 7 scored results on Noometry and Qwen3.7 Plus has 32.

Related comparisons

Go deeper