Anthropic, proprietary
Claude 3.5 Haiku
Claude 3.5 Haiku by Anthropic ranks 315th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.2. Its strongest category is agentic & tool use, where it ranks 95th.
Last verified
Specifications
- Noometry rank
- #315 of 354
- Index score
- 29.2
- Evidence
- Confirmed 49 results
- Provider
- Anthropic
- Released
- October 22, 2024
- Weights
- Proprietary
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- Not measured
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Coding 32.9
- Agentic & Tool Use 28.0
- Reasoning 17.7
- Math 14.7
- Knowledge 18.7
- Multimodal 26.8
- Multilingual 40.0
- Instruction Following 62.9
- Long Context 38.3
- Writing & Preference 42.7
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 32.9 | #265 | 8 |
| Agentic & Tool Use | 28.0 | #95 | 1 |
| Reasoning | 17.7 | #290 | 5 |
| Math | 14.7 | #300 | 5 |
| Knowledge | 18.7 | #281 | 5 |
| Multimodal | 26.8 | #117 | 2 |
| Multilingual | 40.0 | #218 | 1 |
| Instruction Following | 62.9 | #234 | 3 |
| Long Context | 38.3 | #200 | 1 |
| Writing & Preference | 42.7 | #234 | 7 |
Strengths and weaknesses
Categories where Claude 3.5 Haiku places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Agentic & Tool Use | 28.0 | −2.3 | #95 of 154, top 62% |
| Long Context | 38.3 | −2.7 | #200 of 296, top 68% |
| Multilingual | 40.0 | −7.4 | #218 of 297, top 74% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Math | 14.7 | −21.9 | #300 of 327, top 92% |
| Multimodal | 26.8 | −11.7 | #117 of 128, top 92% |
| Knowledge | 18.7 | −18.6 | #281 of 314, top 90% |
Closest competitors
The models ranked just above and below Claude 3.5 Haiku. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| DBRX | #311 | 29.4 | — | — | Compare |
| Gemma 2 27B | #312 | 29.4 | $0.65 | — | Compare |
| Gemma 1.1 2b IT | #313 | 29.3 | — | — | Compare |
| Phi 3 Small 8k Instruct | #314 | 29.3 | — | — | Compare |
| GPT-4 | #316 | 29.1 | $37.50 | — | Compare |
| Llama 2-7B | #317 | 29.1 | — | — | Compare |
| Granite 4.0 Micro | #318 | 29.0 | $0.0408 | — | Compare |
| Claude 3 Sonnet | #319 | 29.0 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Aider Polyglot | 28% | #33 of 44, top 75% | Epoch AI | ||
| SciCode | 27.4% | #109 of 121, top 91% | Epoch AI | ||
| WeirdML | 30.7% | #96 of 119, top 81% | Epoch AI | ||
| BigCodeBench Instruct | 46.1% | #12 of 64, top 19% | BigCodeBench | 2024-10-22 | |
| LiveBench Coding | 51.4% | #17 of 39, top 44% | Epoch AI | ||
| LMArena Coding | 1286 | #208 of 294, top 71% | LMArena | 2026-10-08 | |
| BigCodeBench Complete | 59% | #7 of 66, top 11% | BigCodeBench | 2024-10-22 | |
| CadEval | 32% | #10 of 14, top 72% | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| BALROG | 19.3% | #25 of 35, top 72% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| CritPt | 0% | #101 of 134, top 76% | Epoch AI | ||
| LiveBench Reasoning | 28.1% | #31 of 39, top 80% | Epoch AI | ||
| LMArena Hard Prompts | 1251 | #220 of 297, top 75% | LMArena | 2026-10-08 | |
| DTBench | 56.7% | #121 of 151, top 81% | Epoch AI | ||
| LiveBench Data Analysis | 48.5% | #26 of 39, top 67% | Epoch AI | ||
| Epoch Capabilities Index | 127.15 | #150 of 213, top 71% | Epoch AI | 2024-10-22 | |
| LiveBench | 43.5% | #28 of 39, top 72% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 4.3% | #151 of 173, top 88% | Epoch AI | 2025-02-25 | |
| Omni-MATH | 22.4% | #48 of 57, top 85% | HELM Capabilities | ||
| LiveBench Math | 35.5% | #31 of 39, top 80% | Epoch AI | ||
| LMArena Math | 1244 | #217 of 285, top 77% | LMArena | 2026-10-08 | |
| MATH Level 5 | 46.4% | #49 of 79, top 63% | Epoch AI | 2025-03-12 | |
| FrontierMath (Feb 2025 set) | 0.3% | #64 of 68, top 95% | Epoch AI | 2025-03-07 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 38.1% | #152 of 186, top 82% | Epoch AI | 2025-03-12 | |
| MMLU-Pro | 60.5% | #42 of 58, top 73% | HELM Capabilities | ||
| Confabulations (lower is better) | 36.7% | #49 of 51, top 97% | Lech Mazur benchmarks | ||
| GPQA (HELM) | 36.3% | #48 of 57, top 85% | HELM Capabilities | ||
| LMArena Expert | 1208 | #219 of 273, top 81% | LMArena | 2026-10-08 | |
| MMLU | 74.3% | #35 of 81, top 44% | Epoch AI |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1092 | #108 of 122, top 89% | LMArena | 2026-10-09 | |
| GeoBench | 34% | #25 of 25, top 100% | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1238 | #218 of 297, top 74% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1229 | #218 of 285, top 77% | LMArena | 2026-10-08 | |
| LMArena French | 1264 | #178 of 223, top 80% | LMArena | 2026-10-08 | |
| LMArena German | 1237 | #181 of 231, top 79% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1175 | #168 of 211, top 80% | LMArena | 2026-10-08 | |
| LMArena Korean | 1173 | #173 of 213, top 82% | LMArena | 2026-10-08 | |
| LMArena Russian | 1253 | #211 of 283, top 75% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1261 | #176 of 226, top 78% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LiveBench Instruction Following | 61.9% | #26 of 39, top 67% | Epoch AI | ||
| IFEval | 79.2% | #44 of 57, top 78% | HELM Capabilities | ||
| LMArena Instruction Following | 1241 | #220 of 298, top 74% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1261 | #215 of 291, top 74% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1255 | #225 of 297, top 76% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1233 | #219 of 295, top 75% | LMArena | 2026-10-08 | |
| Short-Story Creative Writing | 73.5% | #27 of 39, top 70% | Epoch AI | ||
| EQ-Bench Creative Writing | 1146 | #88 of 115, top 77% | EQ-Bench | ||
| WildBench | 76% | #43 of 57, top 76% | HELM Capabilities | ||
| LMArena Multi-Turn | 1265 | #216 of 295, top 74% | LMArena | 2026-10-08 | |
| LiveBench Language | 35.4% | #22 of 39, top 57% | Epoch AI |
Compare Claude 3.5 Haiku
- Claude 3.5 Haiku vs Claude 3 Haiku
- Claude 3.5 Haiku vs Phi 3 Small 8k Instruct
- Claude 3.5 Haiku vs GPT-4
- Claude 3.5 Haiku vs Gemma 1.1 2b IT
- Claude 3.5 Haiku vs Llama 2-7B
- Claude 3.5 Haiku vs Gemma 2 27B
- Claude 3.5 Haiku vs Granite 4.0 Micro
- Claude 3.5 Haiku vs GPT-6 Astra
- Claude 3.5 Haiku vs Gemini 3.8 Flash
- Claude 3.5 Haiku vs Kimi K3
- Claude 3.5 Haiku vs Grok 4.6
- Claude 3.5 Haiku vs Qwen3.8 Max
- Claude 3.5 Haiku vs GLM-5.3
- Claude 3.5 Haiku vs Muse Spark 1.3
Other Anthropic models
- Claude Fable 5.169.0
- Claude Opus 5.568.6
- Claude Opus 567.8
- Claude Fable 566.8
- Claude Sonnet 5.561.9
- Claude Opus 4.860.7
- Claude Opus 4.758.3
- Claude Opus 4.658.2
Frequently asked questions
How good is Claude 3.5 Haiku?
Claude 3.5 Haiku by Anthropic ranks 315th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.2. Its strongest category is agentic & tool use, where it ranks 95th.
Is Claude 3.5 Haiku open source?
No. Claude 3.5 Haiku is proprietary and available only through Anthropic's API and partner platforms.
What are Claude 3.5 Haiku's strengths and weaknesses?
Relative to other ranked models, Claude 3.5 Haiku places best in agentic & tool use, long context, multilingual and lowest in math, multimodal, knowledge.
What is Claude 3.5 Haiku best at?
Its best category is agentic & tool use, where it ranks 95th on Noometry.