Model comparison

Claude Haiku 5.5 vs Grok 4

Claude Haiku 5.5 is the stronger model overall, scoring 49.5 to 48.1 on the Noometry Index.

Last verified . 2 shared benchmarks.

Claude Haiku 5.5 Anthropic

49.5

Rank #52 Confirmed

Grok 4 xAI

48.1

Rank #56 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Claude Haiku 5.5 scores higher in 2 categories and Grok 4 in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Haiku 5.5 leads 73.6 to 48.4.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 98.9% for Claude Haiku 5.5 and 84% for Grok 4.

Side by side

Claude Haiku 5.5 and Grok 4 specifications
Claude Haiku 5.5Grok 4
ProviderAnthropicxAI
Noometry Index49.548.1
Released2026-10-072025-07-09
WeightsProprietaryProprietary
Context window1M—
Max output128K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.50—
Results tracked948

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4 leads

Claude Haiku 5.5: 49.2 (#50), Grok 4: 50.3 (#46)

Coding benchmarks
BenchmarkClaude Haiku 5.5Grok 4
Aider Polyglot—79.6%
LMArena WebDev1587—
WeirdML—45.7%
LMArena Coding—1408

Agentic & Tool Use Not comparable

Claude Haiku 5.5: —, Grok 4: 32.3 (#68)

Agentic & Tool Use benchmarks
BenchmarkClaude Haiku 5.5Grok 4
Terminal-Bench—27.2%
Berkeley Function Calling Leaderboard—63%
GDPval—21.1%
Cybench—43%
DeepResearch Bench—47.3%
BALROG—43.6%
LMArena Search—1142
METR Time Horizons—66.6%

Reasoning Grok 4 leads

Claude Haiku 5.5: 35.2 (#69), Grok 4: 36.7 (#65)

Reasoning benchmarks
BenchmarkClaude Haiku 5.5Grok 4
ARC-AGI-2—16%
SimpleBench—60.5%
Kagi LLM Benchmark—73.6%
NYT Connections (extended)65.7%—
ARC-AGI-1—66.7%
Chess Puzzles—28%
LMArena Hard Prompts—1409
Mystery Game Puzzles30%—
Epoch Capabilities Index—146.44
ForecastBench—60.9

Math Claude Haiku 5.5 leads

Claude Haiku 5.5: 73.6 (#18), Grok 4: 48.4 (#64)

Math benchmarks
BenchmarkClaude Haiku 5.5Grok 4
OTIS Mock AIME 2024-202598.9%84%
FrontierMath (Tiers 1-3)75.1%—
FrontierMath Tier 446.3%—
Omni-MATH—60.3%
LMArena Math—1422
FrontierMath (Feb 2025 set)—19.7%
FrontierMath Tier 4 (v1)—2.1%

Knowledge Grok 4 leads

Claude Haiku 5.5: 50.8 (#70), Grok 4: 53.8 (#55)

Knowledge benchmarks
BenchmarkClaude Haiku 5.5Grok 4
GPQA Diamond89.6%87%
SimpleQA Verified23.8%—
MMLU-Pro—85.1%
Confabulations—12.4%
GPQA (HELM)—72.7%
LMArena Expert—1415

Multimodal Claude Haiku 5.5 leads

Claude Haiku 5.5: 43.7 (#21), Grok 4: 33.7 (#94)

Multimodal benchmarks
BenchmarkClaude Haiku 5.5Grok 4
LMArena Vision—1210
GeoBench—45%
Furniture Assembly47.5%—

Multilingual Not comparable

Claude Haiku 5.5: —, Grok 4: 51.8 (#103)

Multilingual benchmarks
BenchmarkClaude Haiku 5.5Grok 4
LMArena Non-English—1403
LMArena Chinese—1427
LMArena French—1418
LMArena German—1429
LMArena Japanese—1394
LMArena Korean—1377
LMArena Russian—1410
LMArena Spanish—1420

Instruction Following Not comparable

Claude Haiku 5.5: —, Grok 4: 79.2 (#5)

Instruction Following benchmarks
BenchmarkClaude Haiku 5.5Grok 4
IFEval—94.9%
LMArena Instruction Following—1387

Long Context Not comparable

Claude Haiku 5.5: —, Grok 4: 63.1 (#4)

Long Context benchmarks
BenchmarkClaude Haiku 5.5Grok 4
Fiction.LiveBench—94.4%
LMArena Longer Query—1409

Writing & Preference Not comparable

Claude Haiku 5.5: —, Grok 4: 58.5 (#116)

Writing & Preference benchmarks
BenchmarkClaude Haiku 5.5Grok 4
LMArena Text—1411
LMArena Creative Writing—1397
Short-Story Creative Writing—76.9%
WildBench—79.7%
LMArena Multi-Turn—1416

Frequently asked questions

Is Claude Haiku 5.5 better than Grok 4?

Claude Haiku 5.5 is the stronger model overall, scoring 49.5 to 48.1 on the Noometry Index.

Is Claude Haiku 5.5 or Grok 4 better for coding?

Grok 4 scores higher on coding benchmarks: 50.3 versus 49.2 in the Noometry coding category.

How many benchmarks do Claude Haiku 5.5 and Grok 4 share?

2 benchmarks have published results for both models. Claude Haiku 5.5 has 9 scored results on Noometry and Grok 4 has 48.

Related comparisons

Go deeper