Model comparison

Claude 3.5 Haiku vs Grok 4.1

Grok 4.1 is the stronger model overall, scoring 41.5 to 29.2 on the Noometry Index.

Last verified . 17 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 0 categories and Grok 4.1 in 9 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok 4.1 leads 38.9 to 14.7.

Side by side

Claude 3.5 Haiku and Grok 4.1 specifications
Claude 3.5 HaikuGrok 4.1
ProviderAnthropicxAI
Noometry Index29.241.5
Released2024-10-222025-11-17
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4919

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude 3.5 Haiku: 32.9 (#265), Grok 4.1: 33.7 (#253)

Coding benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.1
LMArena Coding12861445
Aider Polyglot28%—
LMArena WebDev—1214
SciCode27.4%—
WeirdML30.7%—
BigCodeBench Instruct46.1%—
LiveBench Coding51.4%—
BigCodeBench Complete59%—
CadEval32%—

Agentic & Tool Use Grok 4.1 leads

Claude 3.5 Haiku: 28.0 (#95), Grok 4.1: 34.1 (#49)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.1
Cybench—39%
BALROG19.3%—

Reasoning Grok 4.1 leads

Claude 3.5 Haiku: 17.7 (#290), Grok 4.1: 29.5 (#91)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.1
LMArena Hard Prompts12511435
CritPt0%—
LiveBench Reasoning28.1%—
DTBench56.7%—
LiveBench Data Analysis48.5%—
Epoch Capabilities Index127.15—
LiveBench43.5%—

Math Grok 4.1 leads

Claude 3.5 Haiku: 14.7 (#300), Grok 4.1: 38.9 (#120)

Math benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.1
LMArena Math12441422
OTIS Mock AIME 2024-20254.3%—
Omni-MATH22.4%—
LiveBench Math35.5%—
MATH Level 546.4%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Grok 4.1 leads

Claude 3.5 Haiku: 18.7 (#281), Grok 4.1: 39.5 (#133)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.1
LMArena Expert12081417
GPQA Diamond38.1%—
MMLU-Pro60.5%—
Confabulations36.7%—
GPQA (HELM)36.3%—
MMLU74.3%—

Multimodal Not comparable

Claude 3.5 Haiku: 26.8 (#117), Grok 4.1: —

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.1
LMArena Vision1092—
GeoBench34%—

Multilingual Grok 4.1 leads

Claude 3.5 Haiku: 40.0 (#218), Grok 4.1: 53.4 (#68)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.1
LMArena Non-English12381425
LMArena Chinese12291465
LMArena French12641448
LMArena German12371446
LMArena Japanese11751397
LMArena Korean11731407
LMArena Russian12531434
LMArena Spanish12611438

Instruction Following Grok 4.1 leads

Claude 3.5 Haiku: 62.9 (#234), Grok 4.1: 73.8 (#111)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.1
LMArena Instruction Following12411400
LiveBench Instruction Following61.9%—
IFEval79.2%—

Long Context Grok 4.1 leads

Claude 3.5 Haiku: 38.3 (#200), Grok 4.1: 43.2 (#100)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.1
LMArena Longer Query12611416

Writing & Preference Grok 4.1 leads

Claude 3.5 Haiku: 42.7 (#234), Grok 4.1: 62.4 (#75)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.1
LMArena Text12551437
LMArena Creative Writing12331411
LMArena Multi-Turn12651437
Short-Story Creative Writing73.5%—
EQ-Bench Creative Writing1146—
WildBench76%—
LiveBench Language35.4%—

Frequently asked questions

Is Claude 3.5 Haiku better than Grok 4.1?

Grok 4.1 is the stronger model overall, scoring 41.5 to 29.2 on the Noometry Index.

Is Claude 3.5 Haiku or Grok 4.1 better for coding?

They score almost the same on coding (32.9 vs 33.7); test both on your own repository before choosing.

How many benchmarks do Claude 3.5 Haiku and Grok 4.1 share?

17 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Grok 4.1 has 19.

Related comparisons

Go deeper