Anthropic, proprietary

Claude Opus 4

Claude Opus 4 by Anthropic ranks 100th of 354 ranked models on the Noometry Index as of October 2026, with a score of 43.1. Its strongest category is instruction following, where it ranks 28th. API pricing starts at $15 per million input tokens and $75 per million output tokens, with a 200K-token context window.

Last verified

Specifications

Noometry rank
#100 of 354
Index score
43.1
Evidence
Confirmed 56 results
Provider
Anthropic
Released
May 22, 2025
Weights
Proprietary
Reasoning
Yes
Context window
200K
Max output
32K
Input price
$15 / M
Output price
$75 / M
Blended price
$30 / M
Output speed
29 tokens/s Kagi
Value
#211 of 219
Knowledge cutoff
March 2025
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

Claude Opus 4 category scores
  1. Coding 47.2
  2. Agentic & Tool Use 34.8
  3. Reasoning 27.3
  4. Math 42.0
  5. Knowledge 44.0
  6. Multimodal 31.5
  7. Multilingual 48.8
  8. Instruction Following 77.1
  9. Long Context 39.6
  10. Writing & Preference 61.2
Claude Opus 4 category ranks
CategoryScoreRankResults
Coding47.2#626
Agentic & Tool Use34.8#422
Reasoning27.3#1219
Math42.0#864
Knowledge44.0#887
Multimodal31.5#1063
Multilingual48.8#1381
Instruction Following77.1#282
Long Context39.6#1722
Writing & Preference61.2#896

Strengths and weaknesses

Categories where Claude Opus 4 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Claude Opus 4: strongest categories
CategoryScorevs medianRank
Instruction Following77.1+5.8#28 of 305, top 10%
Coding47.2+8.5#62 of 340, top 19%
Math42.0+5.4#86 of 327, top 27%

Weakest categories

Claude Opus 4: weakest categories
CategoryScorevs medianRank
Multimodal31.5−7.1#106 of 128, top 83%
Long Context39.6−1.4#172 of 296, top 59%
Multilingual48.8+1.4#138 of 297, top 47%

Closest competitors

The models ranked just above and below Claude Opus 4. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Claude Opus 4
ModelRankScoreBlended $/MSpeed
Seed 2.0 Pro#9643.2$1.13—Compare
DeepSeek-V3.1-Terminus#9743.1$0.4530Compare
Hunyuan Vision 1.5#9843.1——Compare
Mistral Large 4#9943.1$1.03—Compare
Amazon Nova Experimental Chat 11 10#10143.0——Compare
Qwen3-Next 80B-A3B Instruct#10243.0$0.88111Compare
MiMo-V2-Pro#10343.0$0.54—Compare
Amazon Nova Experimental Chat 12 10#10442.9——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Claude Opus 4 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified70.7%#24 of 32, top 75%Epoch AI2026-02-06
SWE-bench Verified (bash only)67.6%#12 of 39, top 31%SWE-bench2025-08-02
Aider Polyglot70.7%Epoch AI
Aider Polyglot72%#7 of 44, top 16%32KEpoch AI
GSO6.9%#19 of 31, top 62%Epoch AI
WeirdML43.7%#62 of 119, top 53%16KEpoch AI
LMArena Coding1442#96 of 294, top 33%thinking-16kLMArena2026-10-08
AlgoTune1.33#17 of 18, top 95%Epoch AI

Agentic & Tool Use

Claude Opus 4 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Cybench38%#7 of 21, top 34%Epoch AI
DeepResearch Bench46.8%#12 of 24, top 50%2KEpoch AI
LMArena Search1127#31 of 32, top 97%LMArena2026-08-24
METR Time Horizons61.5%Epoch AI
METR Time Horizons63.9%#15 of 32, top 47%16KEpoch AI

Reasoning

Claude Opus 4 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-21.3%Epoch AI
ARC-AGI-28.6%#50 of 83, top 61%16KEpoch AI
ARC-AGI-20%1KEpoch AI
ARC-AGI-24.5%8KEpoch AI
SimpleBench58.8%#27 of 77, top 36%Epoch AI
SimpleBench58.8%12KEpoch AI
Kagi LLM Benchmark74.3%#14 of 99, top 15%Kagi LLM Benchmark
Kagi LLM Benchmark59.6%Kagi LLM Benchmark
ARC-AGI-122.5%Epoch AI
ARC-AGI-135.7%#62 of 83, top 75%16KEpoch AI
ARC-AGI-127%1KEpoch AI
ARC-AGI-130.7%8KEpoch AI
CritPt0.3%#91 of 134, top 68%Epoch AI
EnigmaEval5.6%#22 of 38, top 58%Epoch AI
LMArena Hard Prompts1399#126 of 297, top 43%thinking-16kLMArena2026-10-08
DTBench81.6%#68 of 151, top 46%Epoch AI
LMCA37.4%#56 of 125, top 45%Epoch AI
Epoch Capabilities Index142.67#95 of 213, top 45%Epoch AI2025-05-22
ForecastBench61.1#15 of 72, top 21%Epoch AI

Math

Claude Opus 4 Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-202542.2%Epoch AI2025-05-22
OTIS Mock AIME 2024-202560%16KEpoch AI2025-05-22
OTIS Mock AIME 2024-202564.4%#101 of 173, top 59%27KEpoch AI2025-05-28
Omni-MATH61.6%#8 of 57, top 15%HELM Capabilities
Omni-MATH51.1%HELM Capabilities
LMArena Math1390#136 of 285, top 48%thinking-16kLMArena2026-10-08
MATH Level 585%#20 of 79, top 26%Epoch AI2025-05-22
FrontierMath (Feb 2025 set)4.5%#47 of 68, top 70%Epoch AI2025-07-04
FrontierMath (Feb 2025 set)4.1%27KEpoch AI2025-08-05
FrontierMath Tier 4 (v1)0%Epoch AI2025-07-01
FrontierMath Tier 4 (v1)4.2%#27 of 55, top 50%27KEpoch AI2025-07-01

Knowledge

Claude Opus 4 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond69.2%Epoch AI2025-05-22
GPQA Diamond76.3%#89 of 186, top 48%16KEpoch AI2025-05-22
Humanity's Last Exam10.7%#24 of 41, top 59%Epoch AI
MMLU-Pro85.9%HELM Capabilities
MMLU-Pro87.5%#2 of 58, top 4%HELM Capabilities
Confabulations (lower is better)15.9%#23 of 51, top 46%Lech Mazur benchmarks
Confabulations (lower is better)17.1%no reasoningLech Mazur benchmarks
Vectara Hallucination Rate (lower is better)12%#72 of 96, top 75%Vectara Hallucination Leaderboard
GPQA (HELM)70.8%#9 of 57, top 16%HELM Capabilities
GPQA (HELM)66.6%HELM Capabilities
LMArena Expert1386#134 of 273, top 50%thinking-16kLMArena2026-10-08

Multimodal

Claude Opus 4 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1192#83 of 122, top 69%thinking-16kLMArena2026-10-09
GeoBench49%#21 of 25, top 84%32KEpoch AI
VPCT33%Epoch AI
VPCT38%#16 of 24, top 67%16KEpoch AI

Multilingual

Claude Opus 4 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1362#138 of 297, top 47%thinking-16kLMArena2026-10-08
LMArena Chinese1386#145 of 285, top 51%thinking-16kLMArena2026-10-08
LMArena French1372#134 of 223, top 61%thinking-16kLMArena2026-10-08
LMArena German1391#101 of 231, top 44%thinking-16kLMArena2026-10-08
LMArena Japanese1331#110 of 211, top 53%thinking-16kLMArena2026-10-08
LMArena Korean1321#118 of 213, top 56%thinking-16kLMArena2026-10-08
LMArena Russian1392#114 of 283, top 41%thinking-16kLMArena2026-10-08
LMArena Spanish1389#120 of 226, top 54%thinking-16kLMArena2026-10-08

Instruction Following

Claude Opus 4 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval91.8%#7 of 57, top 13%HELM Capabilities
IFEval84.9%HELM Capabilities
LMArena Instruction Following1406#91 of 298, top 31%thinking-16kLMArena2026-10-08

Long Context

Claude Opus 4 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench61.1%#28 of 47, top 60%Epoch AI
LMArena Longer Query1422#86 of 291, top 30%thinking-16kLMArena2026-10-08

Writing & Preference

Claude Opus 4 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1377#139 of 297, top 47%thinking-16kLMArena2026-10-08
LMArena Creative Writing1387#99 of 295, top 34%thinking-16kLMArena2026-10-08
Short-Story Creative Writing83.1%Epoch AI
Short-Story Creative Writing83.6%#7 of 39, top 18%16KEpoch AI
EQ-Bench Creative Writing1580#44 of 115, top 39%EQ-Bench
WildBench83.3%HELM Capabilities
WildBench85.2%#12 of 57, top 22%HELM Capabilities
LMArena Multi-Turn1396#123 of 295, top 42%thinking-16kLMArena2026-10-08

API pricing by provider

Claude Opus 4 API prices
RouteInput $/MOutput $/MCached input $/MChecked
vertex$15$75$1.502026-10-10

Compare Claude Opus 4

Other Anthropic models

Frequently asked questions

How good is Claude Opus 4?

Claude Opus 4 by Anthropic ranks 100th of 354 ranked models on the Noometry Index as of October 2026, with a score of 43.1. Its strongest category is instruction following, where it ranks 28th. API pricing starts at $15 per million input tokens and $75 per million output tokens, with a 200K-token context window.

How much does Claude Opus 4 cost?

Claude Opus 4 costs $15 per million input tokens and $75 per million output tokens on vertex, with cached input at $1.50.

What is Claude Opus 4's context window?

Claude Opus 4 accepts up to 200K tokens of input and can write up to 32K tokens in one response.

Is Claude Opus 4 open source?

No. Claude Opus 4 is proprietary and available only through Anthropic's API and partner platforms.

How fast is Claude Opus 4?

Claude Opus 4 generated about 29 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Claude Opus 4's strengths and weaknesses?

Relative to other ranked models, Claude Opus 4 places best in instruction following, coding, math and lowest in multimodal, long context, multilingual.

What is Claude Opus 4 best at?

Its best category is instruction following, where it ranks 28th on Noometry.