Anthropic, proprietary
Claude Opus 4.6
Claude Opus 4.6 by Anthropic ranks 20th of 354 ranked models on the Noometry Index as of October 2026, with a score of 58.2. Its strongest category is agentic & tool use, where it ranks 4th. API pricing starts at $5 per million input tokens and $25 per million output tokens, with a 1M-token context window.
Last verified
Specifications
- Noometry rank
- #20 of 354
- Index score
- 58.2
- Evidence
- Confirmed 68 results
- Provider
- Anthropic
- Released
- February 4, 2026
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 1M
- Max output
- 128K
- Input price
- $5 / M
- Output price
- $25 / M
- Blended price
- $10 / M
- Output speed
- 19 tokens/s Kagi
- Value
- #203 of 219
- Knowledge cutoff
- May 2025
- Input
- text, image, pdf
Category scores
Each category score combines every public result we have in that category.
- Coding 57.2
- Agentic & Tool Use 51.1
- Reasoning 57.8
- Math 63.0
- Knowledge 61.9
- Multimodal 37.3
- Multilingual 57.9
- Instruction Following 79.5
- Long Context 48.1
- Writing & Preference 73.5
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 57.2 | #20 | 8 |
| Agentic & Tool Use | 51.1 | #4 | 7 |
| Reasoning | 57.8 | #23 | 13 |
| Math | 63.0 | #31 | 6 |
| Knowledge | 61.9 | #26 | 5 |
| Multimodal | 37.3 | #74 | 2 |
| Multilingual | 57.9 | #6 | 1 |
| Instruction Following | 79.5 | #4 | 1 |
| Long Context | 48.1 | #13 | 3 |
| Writing & Preference | 73.5 | #10 | 5 |
Strengths and weaknesses
Categories where Claude Opus 4.6 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Instruction Following | 79.5 | +8.2 | #4 of 305, top 2% |
| Multilingual | 57.9 | +10.5 | #6 of 297, top 3% |
| Agentic & Tool Use | 51.1 | +20.7 | #4 of 154, top 3% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multimodal | 37.3 | −1.3 | #74 of 128, top 58% |
| Math | 63.0 | +26.4 | #31 of 327, top 10% |
| Knowledge | 61.9 | +24.5 | #26 of 314, top 9% |
Closest competitors
The models ranked just above and below Claude Opus 4.6. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| GPT-5.4 | #16 | 59.4 | $5.63 | 12 | Compare |
| GPT-5.6 Terra | #17 | 59.2 | $4.50 | 11 | Compare |
| GPT-5.4 Pro | #18 | 58.9 | $67.50 | — | Compare |
| Claude Opus 4.7 | #19 | 58.3 | $10 | 33 | Compare |
| Grok 4.6 | #21 | 56.9 | $3 | — | Compare |
| Qwen3.8 Max | #22 | 56.8 | $3 | — | Compare |
| Gemini 3.1 Pro Preview | #23 | 56.7 | $4.50 | — | Compare |
| Gemini 4 Argon | #24 | 56.5 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified | 78.7% | #4 of 32, top 13% | Epoch AI | 2026-02-18 | |
| FrontierCode | 26.6% | #29 of 37, top 79% | Epoch AI | ||
| SWE-bench Verified (bash only) | 75.6% | #4 of 39, top 11% | SWE-bench | 2026-02-17 | |
| LMArena WebDev | 1547 | #34 of 113, top 31% | high | LMArena | 2026-10-08 |
| SWE-bench Multilingual | 72% | #2 of 13, top 16% | SWE-bench | 2026-02-13 | |
| GSO | 33.3% | Epoch AI | |||
| GSO | 41.2% | #7 of 31, top 23% | high | Epoch AI | |
| WeirdML | 77.9% | Epoch AI | |||
| WeirdML | 78% | #12 of 119, top 11% | high | Epoch AI | |
| LMArena Coding | 1536 | #3 of 294, top 2% | LMArena | 2026-10-08 | |
| ALE-Bench | 996.5 | #43 of 105, top 41% | Epoch AI | ||
| AlgoTune | 1.47 | #12 of 18, top 67% | high | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Terminal-Bench | 79.8% | #5 of 41, top 13% | Epoch AI | ||
| APEX-Agents | 46.3% | #31 of 49, top 64% | max | Epoch AI | |
| Remote Labor Index | 4.2% | #8 of 14, top 58% | Epoch AI | ||
| τ²-bench Banking | 27.3% | #14 of 26, top 54% | max | τ²-bench | 2026-05-05 |
| Cybench | 93% | Best of 21 | Epoch AI | ||
| DeepResearch Bench | 55.3% | Best of 24 | high | Epoch AI | |
| DeepResearch Bench | 51.4% | low | Epoch AI | ||
| DeepResearch Bench | 53.2% | medium | Epoch AI | ||
| GBAEval | 44.1% | #11 of 23, top 48% | Epoch AI | ||
| LMArena Search | 1253 | #2 of 32, top 7% | LMArena | 2026-08-24 | |
| METR Time Horizons | 78.9% | #2 of 32, top 7% | Epoch AI | ||
| Vending-Bench 2 | 8,018 | #12 of 60, top 20% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 69.2% | #21 of 83, top 26% | 120K | Epoch AI | |
| SimpleBench | 67.6% | #14 of 77, top 19% | Epoch AI | ||
| Kagi LLM Benchmark | 83.6% | #4 of 99, top 5% | Kagi LLM Benchmark | ||
| Kagi LLM Benchmark | 72.4% | Kagi LLM Benchmark | |||
| NYT Connections (extended) | 76.4% | Lech Mazur benchmarks | |||
| NYT Connections (extended) | 92.1% | #13 of 91, top 15% | high reasoning | Lech Mazur benchmarks | |
| ARC-AGI-1 | 94% | #18 of 83, top 22% | 120K | Epoch AI | |
| Chess Puzzles | 13% | 120K | Epoch AI | 2026-02-20 | |
| Chess Puzzles | 17% | #65 of 129, top 51% | 32K | Epoch AI | 2026-02-06 |
| Chess Puzzles | 10% | 64K | Epoch AI | 2026-02-06 | |
| Chess Puzzles | 14% | max | Epoch AI | 2026-08-06 | |
| EnigmaEval | 6.8% | Epoch AI | |||
| EnigmaEval | 7.6% | #17 of 38, top 45% | max | Epoch AI | |
| Thematic Generalization | 80.6% | Best of 23 | high reasoning | Lech Mazur benchmarks | |
| EBR-Bench | 12.7% | #17 of 24, top 71% | max | Epoch AI | 2026-06-29 |
| LMArena Hard Prompts | 1527 | #3 of 297, top 2% | high | LMArena | 2026-10-08 |
| Mystery Game Puzzles | 15% | Epoch AI | 2026-08-06 | ||
| Mystery Game Puzzles | 7% | low | Epoch AI | 2026-08-06 | |
| Mystery Game Puzzles | 25% | #33 of 74, top 45% | max | Epoch AI | 2026-07-25 |
| DTBench | 91.2% | #30 of 151, top 20% | max | Epoch AI | |
| LMCA | 55.8% | #9 of 125, top 8% | max | Epoch AI | |
| Epoch Capabilities Index | 155.24 | #30 of 213, top 15% | Epoch AI | 2026-02-05 | |
| ForecastBench | 60 | #32 of 72, top 45% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 66% | #29 of 81, top 36% | max | Epoch AI | 2026-06-11 |
| FrontierMath Tier 4 | 26.8% | #32 of 63, top 51% | max | Epoch AI | 2026-06-11 |
| MathArena Final-Answer Competitions | 78.5% | #6 of 29, top 21% | high | MathArena | |
| OTIS Mock AIME 2024-2025 | 93.1% | 32K | Epoch AI | 2026-02-06 | |
| OTIS Mock AIME 2024-2025 | 94.4% | #37 of 173, top 22% | 64K | Epoch AI | 2026-02-06 |
| OTIS Mock AIME 2024-2025 | 91.1% | max | Epoch AI | 2026-08-06 | |
| ProofBench | 50% | #28 of 77, top 37% | max | Epoch AI | |
| LMArena Math | 1519 | #6 of 285, top 3% | high | LMArena | 2026-10-08 |
| FrontierMath (Feb 2025 set) | 38.3% | Epoch AI | 2026-02-05 | ||
| FrontierMath (Feb 2025 set) | 40% | 32K | Epoch AI | 2026-02-06 | |
| FrontierMath (Feb 2025 set) | 39.7% | 64K | Epoch AI | 2026-02-06 | |
| FrontierMath (Feb 2025 set) | 40.7% | #7 of 68, top 11% | max | Epoch AI | 2026-02-12 |
| FrontierMath Tier 4 (v1) | 14.6% | Epoch AI | 2026-02-05 | ||
| FrontierMath Tier 4 (v1) | 20.8% | 32K | Epoch AI | 2026-02-06 | |
| FrontierMath Tier 4 (v1) | 20.8% | 64K | Epoch AI | 2026-02-06 | |
| FrontierMath Tier 4 (v1) | 22.9% | #9 of 55, top 17% | max | Epoch AI | 2026-02-12 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 90.5% | #34 of 186, top 19% | 32K | Epoch AI | 2026-02-06 |
| GPQA Diamond | 88.8% | 64K | Epoch AI | 2026-02-06 | |
| GPQA Diamond | 88.4% | max | Epoch AI | 2026-08-06 | |
| Humanity's Last Exam | 19% | Epoch AI | |||
| Humanity's Last Exam | 34.4% | #10 of 41, top 25% | max | Epoch AI | |
| SimpleQA Verified | 47% | #32 of 77, top 42% | max | Epoch AI | 2026-08-27 |
| Vectara Hallucination Rate (lower is better) | 12.2% | #76 of 96, top 80% | Vectara Hallucination Leaderboard | ||
| LMArena Expert | 1546 | #4 of 273, top 2% | high | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1316 | #7 of 122, top 6% | high | LMArena | 2026-10-09 |
| Furniture Assembly | 28.3% | #24 of 31, top 78% | max | Epoch AI | 2026-09-10 |
| LMArena Document | 1507 | #3 of 38, top 8% | high | LMArena | 2026-09-13 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1489 | #6 of 297, top 3% | high | LMArena | 2026-10-08 |
| LMArena Chinese | 1551 | #6 of 285, top 3% | high | LMArena | 2026-10-08 |
| LMArena French | 1513 | #5 of 223, top 3% | high | LMArena | 2026-10-08 |
| LMArena German | 1502 | #4 of 231, top 2% | high | LMArena | 2026-10-08 |
| LMArena Japanese | 1484 | #14 of 211, top 7% | high | LMArena | 2026-10-08 |
| LMArena Korean | 1464 | #7 of 213, top 4% | LMArena | 2026-10-08 | |
| LMArena Russian | 1497 | #9 of 283, top 4% | high | LMArena | 2026-10-08 |
| LMArena Spanish | 1510 | #3 of 226, top 2% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1523 | #3 of 298, top 2% | high | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| CL-bench | 20.7% | #6 of 19, top 32% | Epoch AI | ||
| CL-bench Life | 13.6% | Epoch AI | |||
| CL-bench Life | 17% | #4 of 13, top 31% | high | Epoch AI | |
| LMArena Longer Query | 1520 | #4 of 291, top 2% | high | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1503 | #5 of 297, top 2% | high | LMArena | 2026-10-08 |
| LMArena Creative Writing | 1505 | #4 of 295, top 2% | high | LMArena | 2026-10-08 |
| EQ-Bench Creative Writing | 1809 | #23 of 115, top 20% | EQ-Bench | ||
| EQ-Bench 4 | 1223 | #13 of 28, top 47% | EQ-Bench | ||
| LMArena Multi-Turn | 1513 | #2 of 295, top 1% | high | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| anthropic | $5 | $25 | $0.50 | 2026-10-10 |
| azure | $5 | $25 | $0.50 | 2026-10-10 |
| bedrock | $5.50 | $27.50 | $0.55 | 2026-10-10 |
| openrouter | $5 | $25 | $0.50 | 2026-10-10 |
| vertex | $5 | $25 | $0.50 | 2026-10-10 |
Compare Claude Opus 4.6
- Claude Opus 4.6 vs Claude Opus 4.5
- Claude Opus 4.6 vs Claude Opus 4.7
- Claude Opus 4.6 vs Grok 4.6
- Claude Opus 4.6 vs GPT-5.4 Pro
- Claude Opus 4.6 vs Qwen3.8 Max
- Claude Opus 4.6 vs GPT-5.6 Terra
- Claude Opus 4.6 vs Gemini 3.1 Pro Preview
- Claude Opus 4.6 vs GPT-6 Astra
- Claude Opus 4.6 vs Gemini 3.8 Flash
- Claude Opus 4.6 vs Kimi K3
- Claude Opus 4.6 vs GLM-5.3
- Claude Opus 4.6 vs Muse Spark 1.3
- Claude Opus 4.6 vs DeepSeek V4 Pro
- Claude Opus 4.6 vs MiMo-V2.6-Pro
Other Anthropic models
- Claude Fable 5.169.0
- Claude Opus 5.568.6
- Claude Opus 567.8
- Claude Fable 566.8
- Claude Sonnet 5.561.9
- Claude Opus 4.860.7
- Claude Opus 4.758.3
- Claude Sonnet 554.6
Frequently asked questions
How good is Claude Opus 4.6?
Claude Opus 4.6 by Anthropic ranks 20th of 354 ranked models on the Noometry Index as of October 2026, with a score of 58.2. Its strongest category is agentic & tool use, where it ranks 4th. API pricing starts at $5 per million input tokens and $25 per million output tokens, with a 1M-token context window.
How much does Claude Opus 4.6 cost?
Claude Opus 4.6 costs $5 per million input tokens and $25 per million output tokens on Anthropic's own API, with cached input at $0.50.
What is Claude Opus 4.6's context window?
Claude Opus 4.6 accepts up to 1M tokens of input and can write up to 128K tokens in one response.
Is Claude Opus 4.6 open source?
No. Claude Opus 4.6 is proprietary and available only through Anthropic's API and partner platforms.
How fast is Claude Opus 4.6?
Claude Opus 4.6 generated about 19 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Claude Opus 4.6's strengths and weaknesses?
Relative to other ranked models, Claude Opus 4.6 places best in instruction following, multilingual, agentic & tool use and lowest in multimodal, math, knowledge.
What is Claude Opus 4.6 best at?
Its best category is agentic & tool use, where it ranks 4th on Noometry.