Model comparison

Devstral Small 2505 vs Mixtral 8x22B

Devstral Small 2505 is the stronger model overall, scoring 34.3 to 27.1 on the Noometry Index.

Last verified . 0 shared benchmarks.

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Summary

  • The widest gap is in coding, where Devstral Small 2505 leads 38.9 to 24.2.
  • Devstral Small 2505 is cheaper at $0.10 / $0.30 per million input/output tokens, against $2 / $6 for Mixtral 8x22B.
  • Devstral Small 2505 accepts more context: 128K tokens versus 64K.

Side by side

Devstral Small 2505 and Mixtral 8x22B specifications
Devstral Small 2505Mixtral 8x22B
ProviderMistral AIMistral AI
Noometry Index34.327.1
Released2025-05-072024-04-17
WeightsOpenOpen
Context window128K64K
Max output128K64K
Input $ / M tokens$0.10$2
Output $ / M tokens$0.30$6
Results tracked434

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Devstral Small 2505 leads

Devstral Small 2505: 38.9 (#166), Mixtral 8x22B: 24.2 (#329)

Coding benchmarks
BenchmarkDevstral Small 2505Mixtral 8x22B
SWE-bench Verified (bash only)56.4%—
SciCode28.8%—
WeirdML—3.2%
BigCodeBench Instruct—40.6%
LMArena Coding—1166
BigCodeBench Complete—50.2%
HumanEval+—72%
MBPP+—64.3%

Agentic & Tool Use Not comparable

Devstral Small 2505: —, Mixtral 8x22B: 23.1 (#127)

Agentic & Tool Use benchmarks
BenchmarkDevstral Small 2505Mixtral 8x22B
Cybench—7.5%

Reasoning Too close to call

Devstral Small 2505: 19.7 (#252), Mixtral 8x22B: 19.9 (#248)

Reasoning benchmarks
BenchmarkDevstral Small 2505Mixtral 8x22B
Kagi LLM Benchmark37.7%—
CritPt0%—
LMArena Hard Prompts—1150
DTBench—55.1%
Epoch Capabilities Index—122.03
ForecastBench—56.3

Math Not comparable

Devstral Small 2505: —, Mixtral 8x22B: 22.9 (#275)

Math benchmarks
BenchmarkDevstral Small 2505Mixtral 8x22B
Omni-MATH—16.3%
LMArena Math—1184
MATH Level 5—24.2%

Knowledge Not comparable

Devstral Small 2505: —, Mixtral 8x22B: 15.1 (#293)

Knowledge benchmarks
BenchmarkDevstral Small 2505Mixtral 8x22B
GPQA Diamond—34.1%
MMLU-Pro—46%
GPQA (HELM)—33.4%
LMArena Expert—1113
MMLU—77.8%

Multilingual Not comparable

Devstral Small 2505: —, Mixtral 8x22B: 32.8 (#255)

Multilingual benchmarks
BenchmarkDevstral Small 2505Mixtral 8x22B
LMArena Non-English—1128
LMArena Chinese—1116
LMArena French—1166
LMArena German—1141
LMArena Japanese—1037
LMArena Korean—1057
LMArena Russian—1158
LMArena Spanish—1151

Instruction Following Not comparable

Devstral Small 2505: —, Mixtral 8x22B: 57.7 (#266)

Instruction Following benchmarks
BenchmarkDevstral Small 2505Mixtral 8x22B
IFEval—72.4%
LMArena Instruction Following—1147

Long Context Not comparable

Devstral Small 2505: —, Mixtral 8x22B: 34.7 (#247)

Long Context benchmarks
BenchmarkDevstral Small 2505Mixtral 8x22B
LMArena Longer Query—1144

Writing & Preference Not comparable

Devstral Small 2505: —, Mixtral 8x22B: 36.9 (#262)

Writing & Preference benchmarks
BenchmarkDevstral Small 2505Mixtral 8x22B
LMArena Text—1162
LMArena Creative Writing—1141
WildBench—71.1%
LMArena Multi-Turn—1130

Frequently asked questions

Is Devstral Small 2505 better than Mixtral 8x22B?

Devstral Small 2505 is the stronger model overall, scoring 34.3 to 27.1 on the Noometry Index.

Which is cheaper, Devstral Small 2505 or Mixtral 8x22B?

Devstral Small 2505 is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; Mixtral 8x22B lists at $2 and $6.

Is Devstral Small 2505 or Mixtral 8x22B better for coding?

Devstral Small 2505 scores higher on coding benchmarks: 38.9 versus 24.2 in the Noometry coding category.

Which has the bigger context window?

Devstral Small 2505 does, with 128K tokens against 64K.

How many benchmarks do Devstral Small 2505 and Mixtral 8x22B share?

0 benchmarks have published results for both models. Devstral Small 2505 has 4 scored results on Noometry and Mixtral 8x22B has 34.

Related comparisons

Go deeper