Model comparison

Claude 3.5 Haiku vs Grok-2 (Dec 2024)

Grok-2 (Dec 2024) is the stronger model overall, scoring 33.7 to 29.2 on the Noometry Index.

Last verified . 33 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Grok-2 (Dec 2024) xAI

33.7

Rank #239 Confirmed

Summary

  • They share 33 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 1 category and Grok-2 (Dec 2024) in 7 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok-2 (Dec 2024) leads 29.8 to 18.7.
  • The biggest single-benchmark swing is LiveBench Reasoning: 28.1% for Claude 3.5 Haiku and 54.8% for Grok-2 (Dec 2024).

Side by side

Claude 3.5 Haiku and Grok-2 (Dec 2024) specifications
Claude 3.5 HaikuGrok-2 (Dec 2024)
ProviderAnthropicxAI
Noometry Index29.233.7
Released2024-10-222024-08-13
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4934

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude 3.5 Haiku: 32.9 (#265), Grok-2 (Dec 2024): 33.3 (#258)

Coding benchmarks
BenchmarkClaude 3.5 HaikuGrok-2 (Dec 2024)
WeirdML30.7%22.2%
LiveBench Coding51.4%46.4%
LMArena Coding12861287
Aider Polyglot28%—
SciCode27.4%—
BigCodeBench Instruct46.1%—
BigCodeBench Complete59%—
CadEval32%—

Agentic & Tool Use Not comparable

Claude 3.5 Haiku: 28.0 (#95), Grok-2 (Dec 2024): —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuGrok-2 (Dec 2024)
BALROG19.3%—

Reasoning Too close to call

Claude 3.5 Haiku: 17.7 (#290), Grok-2 (Dec 2024): 16.9 (#299)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuGrok-2 (Dec 2024)
LiveBench Reasoning28.1%54.8%
LMArena Hard Prompts12511272
DTBench56.7%65.2%
LiveBench Data Analysis48.5%54.5%
Epoch Capabilities Index127.15130.48
LiveBench43.5%54.3%
SimpleBench—22.7%
CritPt0%—

Math Grok-2 (Dec 2024) leads

Claude 3.5 Haiku: 14.7 (#300), Grok-2 (Dec 2024): 20.8 (#284)

Math benchmarks
BenchmarkClaude 3.5 HaikuGrok-2 (Dec 2024)
OTIS Mock AIME 2024-20254.3%11.5%
LiveBench Math35.5%54.9%
LMArena Math12441283
MATH Level 546.4%63.5%
FrontierMath (Feb 2025 set)0.3%0.7%
Omni-MATH22.4%—

Knowledge Grok-2 (Dec 2024) leads

Claude 3.5 Haiku: 18.7 (#281), Grok-2 (Dec 2024): 29.8 (#233)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuGrok-2 (Dec 2024)
GPQA Diamond38.1%53.8%
Confabulations36.7%20.1%
LMArena Expert12081254
MMLU-Pro60.5%—
GPQA (HELM)36.3%—
MMLU74.3%—

Multimodal Not comparable

Claude 3.5 Haiku: 26.8 (#117), Grok-2 (Dec 2024): —

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuGrok-2 (Dec 2024)
LMArena Vision1092—
GeoBench34%—

Multilingual Grok-2 (Dec 2024) leads

Claude 3.5 Haiku: 40.0 (#218), Grok-2 (Dec 2024): 43.1 (#188)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuGrok-2 (Dec 2024)
LMArena Non-English12381282
LMArena Chinese12291289
LMArena French12641318
LMArena German12371287
LMArena Japanese11751244
LMArena Korean11731237
LMArena Russian12531286
LMArena Spanish12611281

Instruction Following Grok-2 (Dec 2024) leads

Claude 3.5 Haiku: 62.9 (#234), Grok-2 (Dec 2024): 66.9 (#202)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuGrok-2 (Dec 2024)
LiveBench Instruction Following61.9%69.6%
LMArena Instruction Following12411270
IFEval79.2%—

Long Context Too close to call

Claude 3.5 Haiku: 38.3 (#200), Grok-2 (Dec 2024): 38.8 (#190)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuGrok-2 (Dec 2024)
LMArena Longer Query12611276

Writing & Preference Grok-2 (Dec 2024) leads

Claude 3.5 Haiku: 42.7 (#234), Grok-2 (Dec 2024): 48.6 (#198)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuGrok-2 (Dec 2024)
LMArena Text12551305
LMArena Creative Writing12331284
Short-Story Creative Writing73.5%63.6%
LMArena Multi-Turn12651290
LiveBench Language35.4%45.6%
EQ-Bench Creative Writing1146—
WildBench76%—

Frequently asked questions

Is Claude 3.5 Haiku better than Grok-2 (Dec 2024)?

Grok-2 (Dec 2024) is the stronger model overall, scoring 33.7 to 29.2 on the Noometry Index.

Is Claude 3.5 Haiku or Grok-2 (Dec 2024) better for coding?

They score almost the same on coding (32.9 vs 33.3); test both on your own repository before choosing.

How many benchmarks do Claude 3.5 Haiku and Grok-2 (Dec 2024) share?

33 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Grok-2 (Dec 2024) has 34.

Related comparisons

Go deeper