Model comparison

Claude 3.5 Haiku vs Pixtral Large

Pixtral Large is the stronger model overall, scoring 32.2 to 29.2 on the Noometry Index.

Last verified . 2 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Pixtral Large Mistral AI

32.2

Rank #259 Reported

Summary

  • They share 2 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 1 category and Pixtral Large in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Claude 3.5 Haiku leads 42.7 to 32.9.
  • Pixtral Large has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Haiku and Pixtral Large specifications
Claude 3.5 HaikuPixtral Large
ProviderAnthropicMistral AI
Noometry Index29.232.2
Released2024-10-222024-11-01
WeightsProprietaryOpen
Context window—128K
Max output—128K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked493

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Claude 3.5 Haiku: 32.9 (#265), Pixtral Large: —

Coding benchmarks
BenchmarkClaude 3.5 HaikuPixtral Large
Aider Polyglot28%—
SciCode27.4%—
WeirdML30.7%—
BigCodeBench Instruct46.1%—
LiveBench Coding51.4%—
LMArena Coding1286—
BigCodeBench Complete59%—
CadEval32%—

Agentic & Tool Use Not comparable

Claude 3.5 Haiku: 28.0 (#95), Pixtral Large: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuPixtral Large
BALROG19.3%—

Reasoning Pixtral Large leads

Claude 3.5 Haiku: 17.7 (#290), Pixtral Large: 21.7 (#218)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuPixtral Large
CritPt0%—
EnigmaEval—0.8%
LiveBench Reasoning28.1%—
LMArena Hard Prompts1251—
DTBench56.7%—
LiveBench Data Analysis48.5%—
Epoch Capabilities Index127.15—
LiveBench43.5%—

Math Not comparable

Claude 3.5 Haiku: 14.7 (#300), Pixtral Large: —

Math benchmarks
BenchmarkClaude 3.5 HaikuPixtral Large
OTIS Mock AIME 2024-20254.3%—
Omni-MATH22.4%—
LiveBench Math35.5%—
LMArena Math1244—
MATH Level 546.4%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Not comparable

Claude 3.5 Haiku: 18.7 (#281), Pixtral Large: —

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuPixtral Large
GPQA Diamond38.1%—
MMLU-Pro60.5%—
Confabulations36.7%—
GPQA (HELM)36.3%—
LMArena Expert1208—
MMLU74.3%—

Multimodal Pixtral Large leads

Claude 3.5 Haiku: 26.8 (#117), Pixtral Large: 30.6 (#111)

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuPixtral Large
LMArena Vision10921089
GeoBench34%—

Multilingual Not comparable

Claude 3.5 Haiku: 40.0 (#218), Pixtral Large: —

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuPixtral Large
LMArena Non-English1238—
LMArena Chinese1229—
LMArena French1264—
LMArena German1237—
LMArena Japanese1175—
LMArena Korean1173—
LMArena Russian1253—
LMArena Spanish1261—

Instruction Following Not comparable

Claude 3.5 Haiku: 62.9 (#234), Pixtral Large: —

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuPixtral Large
LiveBench Instruction Following61.9%—
IFEval79.2%—
LMArena Instruction Following1241—

Long Context Not comparable

Claude 3.5 Haiku: 38.3 (#200), Pixtral Large: —

Long Context benchmarks
BenchmarkClaude 3.5 HaikuPixtral Large
LMArena Longer Query1261—

Writing & Preference Claude 3.5 Haiku leads

Claude 3.5 Haiku: 42.7 (#234), Pixtral Large: 32.9 (#278)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuPixtral Large
EQ-Bench Creative Writing1146988
LMArena Text1255—
LMArena Creative Writing1233—
Short-Story Creative Writing73.5%—
WildBench76%—
LMArena Multi-Turn1265—
LiveBench Language35.4%—

Frequently asked questions

Is Claude 3.5 Haiku better than Pixtral Large?

Pixtral Large is the stronger model overall, scoring 32.2 to 29.2 on the Noometry Index.

How many benchmarks do Claude 3.5 Haiku and Pixtral Large share?

2 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Pixtral Large has 3.

Related comparisons

Go deeper