Anthropic, proprietary

Claude Opus 4.6

Claude Opus 4.6 by Anthropic ranks 20th of 354 ranked models on the Noometry Index as of October 2026, with a score of 58.2. Its strongest category is agentic & tool use, where it ranks 4th. API pricing starts at $5 per million input tokens and $25 per million output tokens, with a 1M-token context window.

Last verified

Specifications

Noometry rank
#20 of 354
Index score
58.2
Evidence
Confirmed 68 results
Provider
Anthropic
Released
February 4, 2026
Weights
Proprietary
Reasoning
Yes
Context window
1M
Max output
128K
Input price
$5 / M
Output price
$25 / M
Blended price
$10 / M
Output speed
19 tokens/s Kagi
Value
#203 of 219
Knowledge cutoff
May 2025
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

Claude Opus 4.6 category scores
  1. Coding 57.2
  2. Agentic & Tool Use 51.1
  3. Reasoning 57.8
  4. Math 63.0
  5. Knowledge 61.9
  6. Multimodal 37.3
  7. Multilingual 57.9
  8. Instruction Following 79.5
  9. Long Context 48.1
  10. Writing & Preference 73.5
Claude Opus 4.6 category ranks
CategoryScoreRankResults
Coding57.2#208
Agentic & Tool Use51.1#47
Reasoning57.8#2313
Math63.0#316
Knowledge61.9#265
Multimodal37.3#742
Multilingual57.9#61
Instruction Following79.5#41
Long Context48.1#133
Writing & Preference73.5#105

Strengths and weaknesses

Categories where Claude Opus 4.6 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Claude Opus 4.6: strongest categories
CategoryScorevs medianRank
Instruction Following79.5+8.2#4 of 305, top 2%
Multilingual57.9+10.5#6 of 297, top 3%
Agentic & Tool Use51.1+20.7#4 of 154, top 3%

Weakest categories

Claude Opus 4.6: weakest categories
CategoryScorevs medianRank
Multimodal37.3−1.3#74 of 128, top 58%
Math63.0+26.4#31 of 327, top 10%
Knowledge61.9+24.5#26 of 314, top 9%

Closest competitors

The models ranked just above and below Claude Opus 4.6. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Claude Opus 4.6
ModelRankScoreBlended $/MSpeed
GPT-5.4#1659.4$5.6312Compare
GPT-5.6 Terra#1759.2$4.5011Compare
GPT-5.4 Pro#1858.9$67.50—Compare
Claude Opus 4.7#1958.3$1033Compare
Grok 4.6#2156.9$3—Compare
Qwen3.8 Max#2256.8$3—Compare
Gemini 3.1 Pro Preview#2356.7$4.50—Compare
Gemini 4 Argon#2456.5——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Claude Opus 4.6 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified78.7%#4 of 32, top 13%Epoch AI2026-02-18
FrontierCode26.6%#29 of 37, top 79%Epoch AI
SWE-bench Verified (bash only)75.6%#4 of 39, top 11%SWE-bench2026-02-17
LMArena WebDev1547#34 of 113, top 31%highLMArena2026-10-08
SWE-bench Multilingual72%#2 of 13, top 16%SWE-bench2026-02-13
GSO33.3%Epoch AI
GSO41.2%#7 of 31, top 23%highEpoch AI
WeirdML77.9%Epoch AI
WeirdML78%#12 of 119, top 11%highEpoch AI
LMArena Coding1536#3 of 294, top 2%LMArena2026-10-08
ALE-Bench996.5#43 of 105, top 41%Epoch AI
AlgoTune1.47#12 of 18, top 67%highEpoch AI

Agentic & Tool Use

Claude Opus 4.6 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Terminal-Bench79.8%#5 of 41, top 13%Epoch AI
APEX-Agents46.3%#31 of 49, top 64%maxEpoch AI
Remote Labor Index4.2%#8 of 14, top 58%Epoch AI
τ²-bench Banking27.3%#14 of 26, top 54%maxτ²-bench2026-05-05
Cybench93%Best of 21Epoch AI
DeepResearch Bench55.3%Best of 24highEpoch AI
DeepResearch Bench51.4%lowEpoch AI
DeepResearch Bench53.2%mediumEpoch AI
GBAEval44.1%#11 of 23, top 48%Epoch AI
LMArena Search1253#2 of 32, top 7%LMArena2026-08-24
METR Time Horizons78.9%#2 of 32, top 7%Epoch AI
Vending-Bench 28,018#12 of 60, top 20%Epoch AI

Reasoning

Claude Opus 4.6 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-269.2%#21 of 83, top 26%120KEpoch AI
SimpleBench67.6%#14 of 77, top 19%Epoch AI
Kagi LLM Benchmark83.6%#4 of 99, top 5%Kagi LLM Benchmark
Kagi LLM Benchmark72.4%Kagi LLM Benchmark
NYT Connections (extended)76.4%Lech Mazur benchmarks
NYT Connections (extended)92.1%#13 of 91, top 15%high reasoningLech Mazur benchmarks
ARC-AGI-194%#18 of 83, top 22%120KEpoch AI
Chess Puzzles13%120KEpoch AI2026-02-20
Chess Puzzles17%#65 of 129, top 51%32KEpoch AI2026-02-06
Chess Puzzles10%64KEpoch AI2026-02-06
Chess Puzzles14%maxEpoch AI2026-08-06
EnigmaEval6.8%Epoch AI
EnigmaEval7.6%#17 of 38, top 45%maxEpoch AI
Thematic Generalization80.6%Best of 23high reasoningLech Mazur benchmarks
EBR-Bench12.7%#17 of 24, top 71%maxEpoch AI2026-06-29
LMArena Hard Prompts1527#3 of 297, top 2%highLMArena2026-10-08
Mystery Game Puzzles15%Epoch AI2026-08-06
Mystery Game Puzzles7%lowEpoch AI2026-08-06
Mystery Game Puzzles25%#33 of 74, top 45%maxEpoch AI2026-07-25
DTBench91.2%#30 of 151, top 20%maxEpoch AI
LMCA55.8%#9 of 125, top 8%maxEpoch AI
Epoch Capabilities Index155.24#30 of 213, top 15%Epoch AI2026-02-05
ForecastBench60#32 of 72, top 45%Epoch AI

Math

Claude Opus 4.6 Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)66%#29 of 81, top 36%maxEpoch AI2026-06-11
FrontierMath Tier 426.8%#32 of 63, top 51%maxEpoch AI2026-06-11
MathArena Final-Answer Competitions78.5%#6 of 29, top 21%highMathArena
OTIS Mock AIME 2024-202593.1%32KEpoch AI2026-02-06
OTIS Mock AIME 2024-202594.4%#37 of 173, top 22%64KEpoch AI2026-02-06
OTIS Mock AIME 2024-202591.1%maxEpoch AI2026-08-06
ProofBench50%#28 of 77, top 37%maxEpoch AI
LMArena Math1519#6 of 285, top 3%highLMArena2026-10-08
FrontierMath (Feb 2025 set)38.3%Epoch AI2026-02-05
FrontierMath (Feb 2025 set)40%32KEpoch AI2026-02-06
FrontierMath (Feb 2025 set)39.7%64KEpoch AI2026-02-06
FrontierMath (Feb 2025 set)40.7%#7 of 68, top 11%maxEpoch AI2026-02-12
FrontierMath Tier 4 (v1)14.6%Epoch AI2026-02-05
FrontierMath Tier 4 (v1)20.8%32KEpoch AI2026-02-06
FrontierMath Tier 4 (v1)20.8%64KEpoch AI2026-02-06
FrontierMath Tier 4 (v1)22.9%#9 of 55, top 17%maxEpoch AI2026-02-12

Knowledge

Claude Opus 4.6 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond90.5%#34 of 186, top 19%32KEpoch AI2026-02-06
GPQA Diamond88.8%64KEpoch AI2026-02-06
GPQA Diamond88.4%maxEpoch AI2026-08-06
Humanity's Last Exam19%Epoch AI
Humanity's Last Exam34.4%#10 of 41, top 25%maxEpoch AI
SimpleQA Verified47%#32 of 77, top 42%maxEpoch AI2026-08-27
Vectara Hallucination Rate (lower is better)12.2%#76 of 96, top 80%Vectara Hallucination Leaderboard
LMArena Expert1546#4 of 273, top 2%highLMArena2026-10-08

Multimodal

Claude Opus 4.6 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1316#7 of 122, top 6%highLMArena2026-10-09
Furniture Assembly28.3%#24 of 31, top 78%maxEpoch AI2026-09-10
LMArena Document1507#3 of 38, top 8%highLMArena2026-09-13

Multilingual

Claude Opus 4.6 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1489#6 of 297, top 3%highLMArena2026-10-08
LMArena Chinese1551#6 of 285, top 3%highLMArena2026-10-08
LMArena French1513#5 of 223, top 3%highLMArena2026-10-08
LMArena German1502#4 of 231, top 2%highLMArena2026-10-08
LMArena Japanese1484#14 of 211, top 7%highLMArena2026-10-08
LMArena Korean1464#7 of 213, top 4%LMArena2026-10-08
LMArena Russian1497#9 of 283, top 4%highLMArena2026-10-08
LMArena Spanish1510#3 of 226, top 2%LMArena2026-10-08

Instruction Following

Claude Opus 4.6 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1523#3 of 298, top 2%highLMArena2026-10-08

Long Context

Claude Opus 4.6 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
CL-bench20.7%#6 of 19, top 32%Epoch AI
CL-bench Life13.6%Epoch AI
CL-bench Life17%#4 of 13, top 31%highEpoch AI
LMArena Longer Query1520#4 of 291, top 2%highLMArena2026-10-08

Writing & Preference

Claude Opus 4.6 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1503#5 of 297, top 2%highLMArena2026-10-08
LMArena Creative Writing1505#4 of 295, top 2%highLMArena2026-10-08
EQ-Bench Creative Writing1809#23 of 115, top 20%EQ-Bench
EQ-Bench 41223#13 of 28, top 47%EQ-Bench
LMArena Multi-Turn1513#2 of 295, top 1%highLMArena2026-10-08

API pricing by provider

Claude Opus 4.6 API prices
RouteInput $/MOutput $/MCached input $/MChecked
anthropic$5$25$0.502026-10-10
azure$5$25$0.502026-10-10
bedrock$5.50$27.50$0.552026-10-10
openrouter$5$25$0.502026-10-10
vertex$5$25$0.502026-10-10

Compare Claude Opus 4.6

Other Anthropic models

Frequently asked questions

How good is Claude Opus 4.6?

Claude Opus 4.6 by Anthropic ranks 20th of 354 ranked models on the Noometry Index as of October 2026, with a score of 58.2. Its strongest category is agentic & tool use, where it ranks 4th. API pricing starts at $5 per million input tokens and $25 per million output tokens, with a 1M-token context window.

How much does Claude Opus 4.6 cost?

Claude Opus 4.6 costs $5 per million input tokens and $25 per million output tokens on Anthropic's own API, with cached input at $0.50.

What is Claude Opus 4.6's context window?

Claude Opus 4.6 accepts up to 1M tokens of input and can write up to 128K tokens in one response.

Is Claude Opus 4.6 open source?

No. Claude Opus 4.6 is proprietary and available only through Anthropic's API and partner platforms.

How fast is Claude Opus 4.6?

Claude Opus 4.6 generated about 19 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Claude Opus 4.6's strengths and weaknesses?

Relative to other ranked models, Claude Opus 4.6 places best in instruction following, multilingual, agentic & tool use and lowest in multimodal, math, knowledge.

What is Claude Opus 4.6 best at?

Its best category is agentic & tool use, where it ranks 4th on Noometry.