Model comparison

Claude 3.5 Haiku vs Grok 4.3

Grok 4.3 is the stronger model overall, scoring 43.8 to 29.2 on the Noometry Index.

Last verified . 25 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 1 category and Grok 4.3 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.3 leads 52.5 to 18.7.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.3% for Claude 3.5 Haiku and 93.3% for Grok 4.3.

Side by side

Claude 3.5 Haiku and Grok 4.3 specifications
Claude 3.5 HaikuGrok 4.3
ProviderAnthropicxAI
Noometry Index29.243.8
Released2024-10-222026-04-17
WeightsProprietaryProprietary
Context window—1M
Max output—30K
Input $ / M tokens—$1.25
Output $ / M tokens—$2.50
Results tracked4940

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.3 leads

Claude 3.5 Haiku: 32.9 (#265), Grok 4.3: 41.6 (#121)

Coding benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.3
SciCode27.4%47.3%
WeirdML30.7%49.9%
LMArena Coding12861415
Aider Polyglot28%—
LMArena WebDev—1357
BigCodeBench Instruct46.1%—
LiveBench Coding51.4%—
BigCodeBench Complete59%—
CadEval32%—
ALE-Bench—944.17

Agentic & Tool Use Too close to call

Claude 3.5 Haiku: 28.0 (#95), Grok 4.3: 27.7 (#99)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.3
BALROG19.3%—
GDP.pdf—8%
LMArena Search—1165
Vending-Bench 2—35.26

Reasoning Grok 4.3 leads

Claude 3.5 Haiku: 17.7 (#290), Grok 4.3: 35.9 (#68)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.3
CritPt0%8%
LMArena Hard Prompts12511396
DTBench56.7%90.7%
Epoch Capabilities Index127.15149.16
NYT Connections (extended)—55.2%
Chess Puzzles—25%
LiveBench Reasoning28.1%—
LiveBench Data Analysis48.5%—
LMCA—38.3%
ForecastBench—60.3
LiveBench43.5%—

Math Grok 4.3 leads

Claude 3.5 Haiku: 14.7 (#300), Grok 4.3: 46.0 (#74)

Math benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.3
OTIS Mock AIME 2024-20254.3%93.3%
LMArena Math12441388
FrontierMath (Tiers 1-3)—42.8%
FrontierMath Tier 4—14.6%
ProofBench—11%
Omni-MATH22.4%—
LiveBench Math35.5%—
MATH Level 546.4%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Grok 4.3 leads

Claude 3.5 Haiku: 18.7 (#281), Grok 4.3: 52.5 (#62)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.3
GPQA Diamond38.1%88.8%
LMArena Expert12081385
SimpleQA Verified—33.2%
MMLU-Pro60.5%—
Confabulations36.7%—
GPQA (HELM)36.3%—
MMLU74.3%—

Multimodal Grok 4.3 leads

Claude 3.5 Haiku: 26.8 (#117), Grok 4.3: 31.6 (#104)

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.3
LMArena Vision10921229
GeoBench34%—
Blueprint-Bench 2—0%

Multilingual Grok 4.3 leads

Claude 3.5 Haiku: 40.0 (#218), Grok 4.3: 50.5 (#120)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.3
LMArena Non-English12381385
LMArena Chinese12291422
LMArena French12641412
LMArena German12371395
LMArena Japanese11751379
LMArena Korean11731356
LMArena Russian12531399
LMArena Spanish12611398

Instruction Following Grok 4.3 leads

Claude 3.5 Haiku: 62.9 (#234), Grok 4.3: 72.1 (#140)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.3
LMArena Instruction Following12411366
LiveBench Instruction Following61.9%—
IFEval79.2%—

Long Context Grok 4.3 leads

Claude 3.5 Haiku: 38.3 (#200), Grok 4.3: 42.5 (#123)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.3
LMArena Longer Query12611393

Writing & Preference Grok 4.3 leads

Claude 3.5 Haiku: 42.7 (#234), Grok 4.3: 58.5 (#118)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.3
LMArena Text12551397
LMArena Creative Writing12331380
LMArena Multi-Turn12651406
Short-Story Creative Writing73.5%—
EQ-Bench Creative Writing1146—
WildBench76%—
EQ-Bench 4—1075
LiveBench Language35.4%—

Frequently asked questions

Is Claude 3.5 Haiku better than Grok 4.3?

Grok 4.3 is the stronger model overall, scoring 43.8 to 29.2 on the Noometry Index.

Is Claude 3.5 Haiku or Grok 4.3 better for coding?

Grok 4.3 scores higher on coding benchmarks: 41.6 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and Grok 4.3 share?

25 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Grok 4.3 has 40.

Related comparisons

Go deeper