Anthropic, proprietary
Claude Opus 4.1
Claude Opus 4.1 by Anthropic ranks 142nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 41.0. Its strongest category is agentic & tool use, where it ranks 41st. API pricing starts at $15 per million input tokens and $75 per million output tokens, with a 200K-token context window.
Last verified
Specifications
- Noometry rank
- #142 of 354
- Index score
- 41.0
- Evidence
- Confirmed 48 results
- Provider
- Anthropic
- Released
- August 5, 2025
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 200K
- Max output
- 32K
- Input price
- $15 / M
- Output price
- $75 / M
- Blended price
- $30 / M
- Output speed
- Not measured
- Value
- #212 of 219
- Knowledge cutoff
- March 2025
- Input
- text, image, pdf
Category scores
Each category score combines every public result we have in that category.
- Coding 44.4
- Agentic & Tool Use 35.0
- Reasoning 32.2
- Math 22.3
- Knowledge 42.0
- Multimodal 26.8
- Multilingual 52.0
- Instruction Following 75.6
- Long Context 44.5
- Writing & Preference 62.4
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 44.4 | #73 | 4 |
| Agentic & Tool Use | 35.0 | #41 | 4 |
| Reasoning | 32.2 | #76 | 8 |
| Math | 22.3 | #277 | 4 |
| Knowledge | 42.0 | #101 | 5 |
| Multimodal | 26.8 | #119 | 1 |
| Multilingual | 52.0 | #95 | 1 |
| Instruction Following | 75.6 | #58 | 1 |
| Long Context | 44.5 | #63 | 1 |
| Writing & Preference | 62.4 | #74 | 4 |
Strengths and weaknesses
Categories where Claude Opus 4.1 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Instruction Following | 75.6 | +4.3 | #58 of 305, top 20% |
| Long Context | 44.5 | +3.6 | #63 of 296, top 22% |
| Coding | 44.4 | +5.7 | #73 of 340, top 22% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multimodal | 26.8 | −11.7 | #119 of 128, top 93% |
| Math | 22.3 | −14.3 | #277 of 327, top 85% |
| Knowledge | 42.0 | +4.7 | #101 of 314, top 33% |
Closest competitors
The models ranked just above and below Claude Opus 4.1. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| MiMo-V2-Flash | #138 | 41.3 | $0.18 | — | Compare |
| Hunyuan Turbos 20250226 | #139 | 41.3 | — | — | Compare |
| Kimi K2 (Jul 2025) | #140 | 41.2 | $1 | 201 | Compare |
| Grok-3 mini | #141 | 41.2 | — | 10 | Compare |
| o1 | #143 | 40.9 | $26.25 | — | Compare |
| Gemini 3.1 Flash Lite | #144 | 40.8 | $0.56 | 10 | Compare |
| Claude Sonnet 4 | #145 | 40.8 | $6 | 31 | Compare |
| Qwen2.5-Max | #146 | 40.7 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified | 73.3% | #20 of 32, top 63% | Epoch AI | 2026-02-11 | |
| LMArena WebDev | 1390 | #77 of 113, top 69% | LMArena | 2026-10-08 | |
| WeirdML | 45.9% | #58 of 119, top 49% | 16K | Epoch AI | |
| LMArena Coding | 1479 | #48 of 294, top 17% | thinking-16k | LMArena | 2026-10-08 |
| ALE-Bench | 674.77 | #71 of 105, top 68% | 16K | Epoch AI | |
| AlgoTune | 1.34 | #16 of 18, top 89% | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Terminal-Bench | 38% | #25 of 41, top 61% | Epoch AI | ||
| GDPval | 43.6% | #3 of 11, top 28% | Epoch AI | ||
| Cybench | 42% | #5 of 21, top 24% | Epoch AI | ||
| DeepResearch Bench | 48.3% | #9 of 24, top 38% | Epoch AI | ||
| LMArena Search | 1148 | #24 of 32, top 75% | LMArena | 2026-08-24 | |
| METR Time Horizons | 61.6% | Epoch AI | |||
| METR Time Horizons | 66.8% | #12 of 32, top 38% | 16K | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SimpleBench | 60% | #26 of 77, top 34% | Epoch AI | ||
| Chess Puzzles | 7% | #86 of 129, top 67% | Epoch AI | 2026-07-20 | |
| EnigmaEval | 7.2% | #18 of 38, top 48% | Epoch AI | ||
| EBR-Bench | 7.9% | #21 of 24, top 88% | Epoch AI | 2026-06-25 | |
| LMArena Hard Prompts | 1443 | #79 of 297, top 27% | thinking-16k | LMArena | 2026-10-08 |
| Mystery Game Puzzles | 21% | #39 of 74, top 53% | 24K | Epoch AI | 2026-07-25 |
| DTBench | 80% | #75 of 151, top 50% | Epoch AI | ||
| LMCA | 37.1% | #58 of 125, top 47% | Epoch AI | ||
| Epoch Capabilities Index | 144.12 | #87 of 213, top 41% | Epoch AI | 2025-08-05 | |
| ForecastBench | 62 | #2 of 72, top 3% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 12.6% | #75 of 81, top 93% | 32K | Epoch AI | 2026-06-11 |
| FrontierMath Tier 4 | 2.4% | #57 of 63, top 91% | 32K | Epoch AI | 2026-06-11 |
| OTIS Mock AIME 2024-2025 | 40% | Epoch AI | 2025-08-05 | ||
| OTIS Mock AIME 2024-2025 | 64.4% | 16K | Epoch AI | 2025-08-05 | |
| OTIS Mock AIME 2024-2025 | 68.9% | #94 of 173, top 55% | 27K | Epoch AI | 2025-08-05 |
| LMArena Math | 1431 | #84 of 285, top 30% | thinking-16k | LMArena | 2026-10-08 |
| FrontierMath (Feb 2025 set) | 5.9% | Epoch AI | 2025-08-05 | ||
| FrontierMath (Feb 2025 set) | 7.2% | #41 of 68, top 61% | 27K | Epoch AI | 2025-08-05 |
| FrontierMath Tier 4 (v1) | 4.2% | #28 of 55, top 51% | 27K | Epoch AI | 2025-08-05 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 73.2% | Epoch AI | 2025-08-05 | ||
| GPQA Diamond | 77.3% | #85 of 186, top 46% | 16K | Epoch AI | 2025-08-05 |
| GPQA Diamond | 76.8% | 27K | Epoch AI | 2025-08-05 | |
| Humanity's Last Exam | 11.5% | #23 of 41, top 57% | Epoch AI | ||
| Confabulations (lower is better) | 17.1% | #26 of 51, top 51% | Lech Mazur benchmarks | ||
| Confabulations (lower is better) | 18.5% | no reasoning | Lech Mazur benchmarks | ||
| Vectara Hallucination Rate (lower is better) | 11.8% | #70 of 96, top 73% | Vectara Hallucination Leaderboard | ||
| LMArena Expert | 1439 | #85 of 273, top 32% | thinking-16k | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| VPCT | 35% | #19 of 24, top 80% | Epoch AI | ||
| VPCT | 33% | 16K | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1405 | #95 of 297, top 32% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1427 | #121 of 285, top 43% | LMArena | 2026-10-08 | |
| LMArena French | 1431 | #91 of 223, top 41% | thinking-16k | LMArena | 2026-10-08 |
| LMArena German | 1413 | #84 of 231, top 37% | thinking-16k | LMArena | 2026-10-08 |
| LMArena Japanese | 1378 | #82 of 211, top 39% | thinking-16k | LMArena | 2026-10-08 |
| LMArena Korean | 1380 | #79 of 213, top 38% | thinking-16k | LMArena | 2026-10-08 |
| LMArena Russian | 1422 | #80 of 283, top 29% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1448 | #52 of 226, top 24% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1435 | #55 of 298, top 19% | thinking-16k | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1455 | #46 of 291, top 16% | thinking-16k | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1419 | #95 of 297, top 32% | thinking-16k | LMArena | 2026-10-08 |
| LMArena Creative Writing | 1412 | #61 of 295, top 21% | thinking-16k | LMArena | 2026-10-08 |
| Short-Story Creative Writing | 84.7% | #3 of 39, top 8% | Epoch AI | ||
| Short-Story Creative Writing | 84.5% | 16K | Epoch AI | ||
| LMArena Multi-Turn | 1444 | #66 of 295, top 23% | thinking-16k | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $15 | $75 | $1.50 | 2026-10-10 |
| bedrock | $15 | $75 | $1.50 | 2026-10-10 |
| openrouter | $15 | $75 | $1.50 | 2026-10-10 |
| vertex | $15 | $75 | $1.50 | 2026-10-10 |
Compare Claude Opus 4.1
- Claude Opus 4.1 vs Claude Opus 4
- Claude Opus 4.1 vs Grok-3 mini
- Claude Opus 4.1 vs o1
- Claude Opus 4.1 vs Kimi K2 (Jul 2025)
- Claude Opus 4.1 vs Gemini 3.1 Flash Lite
- Claude Opus 4.1 vs Hunyuan Turbos 20250226
- Claude Opus 4.1 vs Claude Sonnet 4
- Claude Opus 4.1 vs GPT-6 Astra
- Claude Opus 4.1 vs Gemini 3.8 Flash
- Claude Opus 4.1 vs Kimi K3
- Claude Opus 4.1 vs Grok 4.6
- Claude Opus 4.1 vs Qwen3.8 Max
- Claude Opus 4.1 vs GLM-5.3
- Claude Opus 4.1 vs Muse Spark 1.3
Other Anthropic models
- Claude Fable 5.169.0
- Claude Opus 5.568.6
- Claude Opus 567.8
- Claude Fable 566.8
- Claude Sonnet 5.561.9
- Claude Opus 4.860.7
- Claude Opus 4.758.3
- Claude Opus 4.658.2
Frequently asked questions
How good is Claude Opus 4.1?
Claude Opus 4.1 by Anthropic ranks 142nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 41.0. Its strongest category is agentic & tool use, where it ranks 41st. API pricing starts at $15 per million input tokens and $75 per million output tokens, with a 200K-token context window.
How much does Claude Opus 4.1 cost?
Claude Opus 4.1 costs $15 per million input tokens and $75 per million output tokens on azure, with cached input at $1.50.
What is Claude Opus 4.1's context window?
Claude Opus 4.1 accepts up to 200K tokens of input and can write up to 32K tokens in one response.
Is Claude Opus 4.1 open source?
No. Claude Opus 4.1 is proprietary and available only through Anthropic's API and partner platforms.
What are Claude Opus 4.1's strengths and weaknesses?
Relative to other ranked models, Claude Opus 4.1 places best in instruction following, long context, coding and lowest in multimodal, math, knowledge.
What is Claude Opus 4.1 best at?
Its best category is agentic & tool use, where it ranks 41st on Noometry.