Model comparison

Claude 3 Haiku vs Claude Sonnet 4.6

Claude Sonnet 4.6 is the stronger model overall, scoring 50.3 to 25.9 on the Noometry Index.

Last verified . 26 shared benchmarks.

Claude 3 Haiku Anthropic

25.9

Rank #340 Confirmed

Claude Sonnet 4.6 Anthropic

50.3

Rank #50 Confirmed

Summary

  • They share 26 benchmarks with published results for both. Claude 3 Haiku scores higher in 0 categories and Claude Sonnet 4.6 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Sonnet 4.6 leads 52.9 to 9.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 1.8% for Claude 3 Haiku and 85.8% for Claude Sonnet 4.6.

Side by side

Claude 3 Haiku and Claude Sonnet 4.6 specifications
Claude 3 HaikuClaude Sonnet 4.6
ProviderAnthropicAnthropic
Noometry Index25.950.3
Released2024-03-072026-02-17
WeightsProprietaryProprietary
Context window—1M
Max output—128K
Input $ / M tokens—$3
Output $ / M tokens—$15
Results tracked3757

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Sonnet 4.6 leads

Claude 3 Haiku: 26.4 (#325), Claude Sonnet 4.6: 46.3 (#67)

Coding benchmarks
BenchmarkClaude 3 HaikuClaude Sonnet 4.6
WeirdML9.8%66.1%
LMArena Coding11991504
SWE-bench Verified—75.2%
DeepSWE—29.9%
FrontierCode—24.3%
LMArena WebDev—1522
SciCode—46.8%
BigCodeBench Instruct39.4%—
BigCodeBench Complete50.1%—
CadEval12%—
ALE-Bench—1,327
HumanEval+68.9%—
MBPP+68.8%—

Agentic & Tool Use Not comparable

Claude 3 Haiku: —, Claude Sonnet 4.6: 39.1 (#28)

Agentic & Tool Use benchmarks
BenchmarkClaude 3 HaikuClaude Sonnet 4.6
Terminal-Bench—53.4%
APEX-Agents—43%
OSWorld 2.0—9.3%
DeepResearch Bench—54.9%
OSWorld—72.1%
ExploitBench—23.6%
GBAEval—48.8%
GDP.pdf—18%
LMArena Search—1221
Vending-Bench 2—7,204

Reasoning Claude Sonnet 4.6 leads

Claude 3 Haiku: 16.3 (#307), Claude Sonnet 4.6: 46.1 (#45)

Reasoning benchmarks
BenchmarkClaude 3 HaikuClaude Sonnet 4.6
LMArena Hard Prompts11741484
DTBench50.1%89.9%
LMCA8.8%46.5%
Epoch Capabilities Index118.35152.24
ForecastBench53.262
ARC-AGI-2—60.4%
Kagi LLM Benchmark34.2%—
NYT Connections (extended)—80.9%
ARC-AGI-1—86.5%
CritPt—3.1%
Chess Puzzles—13%
Thematic Generalization—76.3%
Mystery Game Puzzles—16%
WinoGrande74.2%—

Math Claude Sonnet 4.6 leads

Claude 3 Haiku: 9.8 (#319), Claude Sonnet 4.6: 52.9 (#49)

Math benchmarks
BenchmarkClaude 3 HaikuClaude Sonnet 4.6
OTIS Mock AIME 2024-20251.8%85.8%
LMArena Math11881462
ProofBench—45%
MATH Level 514.9%—
FrontierMath (Feb 2025 set)—32.4%
FrontierMath Tier 4 (v1)—8.3%

Knowledge Claude Sonnet 4.6 leads

Claude 3 Haiku: 17.3 (#285), Claude Sonnet 4.6: 51.7 (#65)

Knowledge benchmarks
BenchmarkClaude 3 HaikuClaude Sonnet 4.6
GPQA Diamond36.3%87.4%
LMArena Expert11481500
SimpleQA Verified—35.5%
Confabulations34.2%—
Vectara Hallucination Rate—10.6%
MMLU73.8%—

Multimodal Claude Sonnet 4.6 leads

Claude 3 Haiku: 23.6 (#128), Claude Sonnet 4.6: 38.0 (#68)

Multimodal benchmarks
BenchmarkClaude 3 HaikuClaude Sonnet 4.6
LMArena Vision9501283
Blueprint-Bench 2—6.7%
LMArena Document—1482
ScienceQA72%—

Multilingual Claude Sonnet 4.6 leads

Claude 3 Haiku: 36.0 (#243), Claude Sonnet 4.6: 54.4 (#41)

Multilingual benchmarks
BenchmarkClaude 3 HaikuClaude Sonnet 4.6
LMArena Non-English11781440
LMArena Chinese11551491
LMArena French11951465
LMArena German11741428
LMArena Japanese11021420
LMArena Korean11091411
LMArena Russian12041440
LMArena Spanish11661464

Instruction Following Claude Sonnet 4.6 leads

Claude 3 Haiku: 61.3 (#247), Claude Sonnet 4.6: 77.4 (#25)

Instruction Following benchmarks
BenchmarkClaude 3 HaikuClaude Sonnet 4.6
LMArena Instruction Following11731475

Long Context Claude Sonnet 4.6 leads

Claude 3 Haiku: 36.1 (#237), Claude Sonnet 4.6: 45.3 (#44)

Long Context benchmarks
BenchmarkClaude 3 HaikuClaude Sonnet 4.6
LMArena Longer Query11901479

Writing & Preference Claude Sonnet 4.6 leads

Claude 3 Haiku: 29.7 (#291), Claude Sonnet 4.6: 70.2 (#22)

Writing & Preference benchmarks
BenchmarkClaude 3 HaikuClaude Sonnet 4.6
LMArena Text11951458
LMArena Creative Writing11571435
EQ-Bench Creative Writing7171810
LMArena Multi-Turn11901464
EQ-Bench 4—1207

Frequently asked questions

Is Claude 3 Haiku better than Claude Sonnet 4.6?

Claude Sonnet 4.6 is the stronger model overall, scoring 50.3 to 25.9 on the Noometry Index.

Is Claude 3 Haiku or Claude Sonnet 4.6 better for coding?

Claude Sonnet 4.6 scores higher on coding benchmarks: 46.3 versus 26.4 in the Noometry coding category.

How many benchmarks do Claude 3 Haiku and Claude Sonnet 4.6 share?

26 benchmarks have published results for both models. Claude 3 Haiku has 37 scored results on Noometry and Claude Sonnet 4.6 has 57.

Related comparisons

Go deeper