Anthropic, proprietary
Claude Opus 4.5
Claude Opus 4.5 by Anthropic ranks 47th of 354 ranked models on the Noometry Index as of October 2026, with a score of 50.5. Its strongest category is agentic & tool use, where it ranks 12th. API pricing starts at $5 per million input tokens and $25 per million output tokens, with a 200K-token context window.
Last verified
Specifications
- Noometry rank
- #47 of 354
- Index score
- 50.5
- Evidence
- Confirmed 69 results
- Provider
- Anthropic
- Released
- November 1, 2025
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 200K
- Max output
- 64K
- Input price
- $5 / M
- Output price
- $25 / M
- Blended price
- $10 / M
- Output speed
- 13 tokens/s Kagi
- Value
- #205 of 219
- Knowledge cutoff
- May 2025
- Input
- text, image, pdf
Category scores
Each category score combines every public result we have in that category.
- Coding 54.8
- Agentic & Tool Use 47.3
- Reasoning 42.6
- Math 38.6
- Knowledge 56.5
- Multimodal 31.4
- Multilingual 54.3
- Instruction Following 77.5
- Long Context 46.5
- Writing & Preference 68.1
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 54.8 | #27 | 7 |
| Agentic & Tool Use | 47.3 | #12 | 12 |
| Reasoning | 42.6 | #51 | 12 |
| Math | 38.6 | #132 | 5 |
| Knowledge | 56.5 | #44 | 5 |
| Multimodal | 31.4 | #107 | 3 |
| Multilingual | 54.3 | #47 | 1 |
| Instruction Following | 77.5 | #19 | 1 |
| Long Context | 46.5 | #22 | 2 |
| Writing & Preference | 68.1 | #28 | 4 |
Strengths and weaknesses
Categories where Claude Opus 4.5 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Instruction Following | 77.5 | +6.3 | #19 of 305, top 7% |
| Long Context | 46.5 | +5.5 | #22 of 296, top 8% |
| Agentic & Tool Use | 47.3 | +17.0 | #12 of 154, top 8% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multimodal | 31.4 | −7.1 | #107 of 128, top 84% |
| Math | 38.6 | +2.1 | #132 of 327, top 41% |
| Multilingual | 54.3 | +6.9 | #47 of 297, top 16% |
Closest competitors
The models ranked just above and below Claude Opus 4.5. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Qwen3.6 Max Preview | #43 | 51.5 | $2.92 | — | Compare |
| GLM-5.2 | #44 | 51.1 | $2.15 | 23 | Compare |
| GPT-5 | #45 | 50.9 | $3.44 | 2 | Compare |
| Muse Spark | #46 | 50.6 | — | — | Compare |
| Muse Spark 1.2 | #48 | 50.3 | $2 | — | Compare |
| MiMo-V2.6-Pro | #49 | 50.3 | $0.54 | — | Compare |
| Claude Sonnet 4.6 | #50 | 50.3 | $6 | — | Compare |
| Muse Spark 1.1 | #51 | 49.9 | $2 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified | 76.7% | #9 of 32, top 29% | Epoch AI | 2026-02-05 | |
| SWE-bench Verified (bash only) | 76.8% | Best of 39 | high | SWE-bench | 2026-02-17 |
| SWE-bench Verified (bash only) | 74.4% | medium | SWE-bench | 2025-11-24 | |
| LMArena WebDev | 1494 | #49 of 113, top 44% | LMArena | 2026-10-08 | |
| LMArena WebDev | 1469 | LMArena | 2026-10-08 | ||
| SWE-bench Multilingual | 70.7% | #3 of 13, top 24% | SWE-bench | 2026-02-13 | |
| GSO | 26.5% | #12 of 31, top 39% | Epoch AI | ||
| WeirdML | 63.7% | #24 of 119, top 21% | 16K | Epoch AI | |
| LMArena Coding | 1499 | LMArena | 2026-10-08 | ||
| LMArena Coding | 1504 | #16 of 294, top 6% | LMArena | 2026-10-08 | |
| ALE-Bench | 1,025 | #41 of 105, top 40% | 16K | Epoch AI | |
| AlgoTune | 1.77 | #5 of 18, top 28% | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Terminal-Bench | 63.1% | #11 of 41, top 27% | Epoch AI | ||
| Terminal-Bench | 59.1% | 128K | Epoch AI | ||
| Berkeley Function Calling Leaderboard | 77.5% | Best of 49 | fc | Berkeley Function Calling Leaderboard | |
| GDPval | 45.5% | #2 of 11, top 19% | Epoch AI | ||
| Remote Labor Index | 3.8% | #9 of 14, top 65% | Epoch AI | ||
| τ²-bench Airline | 84% | Best of 7 | high | τ²-bench | 2026-02-26 |
| τ²-bench Banking | 24.7% | #19 of 26, top 74% | high | τ²-bench | 2026-02-26 |
| τ²-bench Retail | 79.6% | #3 of 7, top 43% | high | τ²-bench | 2026-02-26 |
| τ²-bench Telecom | 92.3% | #2 of 7, top 29% | high | τ²-bench | 2026-02-26 |
| Cybench | 82% | #2 of 21, top 10% | Epoch AI | ||
| DeepResearch Bench | 54.8% | #3 of 24, top 13% | high | Epoch AI | |
| DeepResearch Bench | 53.7% | low | Epoch AI | ||
| OSWorld | 66.3% | #2 of 8, top 25% | Epoch AI | ||
| BALROG | 43.5% | #10 of 35, top 29% | Epoch AI | ||
| BALROG | 43% | 64K | Epoch AI | ||
| LMArena Search | 1180 | #19 of 32, top 60% | LMArena | 2026-08-24 | |
| METR Time Horizons | 73% | Epoch AI | |||
| METR Time Horizons | 75% | #5 of 32, top 16% | 16K | Epoch AI | |
| Vending-Bench 2 | 4,967 | #31 of 60, top 52% | Epoch AI |
Reasoning
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 34.4% | #57 of 81, top 71% | 32K | Epoch AI | 2026-06-11 |
| FrontierMath Tier 4 | 4.9% | #54 of 63, top 86% | 32K | Epoch AI | 2026-06-11 |
| OTIS Mock AIME 2024-2025 | 48.1% | Epoch AI | 2025-11-24 | ||
| OTIS Mock AIME 2024-2025 | 81.7% | 16K | Epoch AI | 2025-11-24 | |
| OTIS Mock AIME 2024-2025 | 86.1% | #68 of 173, top 40% | 32K | Epoch AI | 2025-11-24 |
| ProofBench | 36% | #37 of 77, top 49% | Epoch AI | ||
| LMArena Math | 1458 | LMArena | 2026-10-08 | ||
| LMArena Math | 1463 | #50 of 285, top 18% | LMArena | 2026-10-08 | |
| FrontierMath (Feb 2025 set) | 20.7% | #30 of 68, top 45% | Epoch AI | 2025-11-25 | |
| FrontierMath (Feb 2025 set) | 20.3% | 16K | Epoch AI | 2025-11-25 | |
| FrontierMath (Feb 2025 set) | 20.7% | 32K | Epoch AI | 2025-11-25 | |
| FrontierMath Tier 4 (v1) | 4.2% | #29 of 55, top 53% | Epoch AI | 2025-11-25 | |
| FrontierMath Tier 4 (v1) | 2.1% | 16K | Epoch AI | 2025-11-25 | |
| FrontierMath Tier 4 (v1) | 4.2% | 32K | Epoch AI | 2025-11-25 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 80.7% | Epoch AI | 2025-11-24 | ||
| GPQA Diamond | 85.5% | 16K | Epoch AI | 2025-11-25 | |
| GPQA Diamond | 86% | #60 of 186, top 33% | 32K | Epoch AI | 2025-11-24 |
| Humanity's Last Exam | 25.2% | #14 of 41, top 35% | Epoch AI | ||
| SimpleQA Verified | 45.7% | #35 of 77, top 46% | 32K | Epoch AI | 2026-08-27 |
| Vectara Hallucination Rate (lower is better) | 10.9% | #65 of 96, top 68% | Vectara Hallucination Leaderboard | ||
| LMArena Expert | 1487 | #35 of 273, top 13% | LMArena | 2026-10-08 | |
| LMArena Expert | 1481 | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GeoBench | 75% | #9 of 25, top 36% | Epoch AI | ||
| VPCT | 40% | #12 of 24, top 50% | 32K | Epoch AI | |
| Furniture Assembly | 28.3% | #23 of 31, top 75% | 64K | Epoch AI | 2026-09-10 |
| LMArena Document | 1462 | #17 of 38, top 45% | LMArena | 2026-09-13 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1438 | #47 of 297, top 16% | LMArena | 2026-10-08 | |
| LMArena Non-English | 1435 | LMArena | 2026-10-08 | ||
| LMArena Chinese | 1470 | #72 of 285, top 26% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1466 | LMArena | 2026-10-08 | ||
| LMArena French | 1457 | LMArena | 2026-10-08 | ||
| LMArena French | 1471 | #38 of 223, top 18% | LMArena | 2026-10-08 | |
| LMArena German | 1449 | #45 of 231, top 20% | LMArena | 2026-10-08 | |
| LMArena German | 1440 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1413 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1416 | #42 of 211, top 20% | LMArena | 2026-10-08 | |
| LMArena Korean | 1424 | #34 of 213, top 16% | LMArena | 2026-10-08 | |
| LMArena Korean | 1374 | LMArena | 2026-10-08 | ||
| LMArena Russian | 1447 | #47 of 283, top 17% | LMArena | 2026-10-08 | |
| LMArena Russian | 1437 | LMArena | 2026-10-08 | ||
| LMArena Spanish | 1458 | #37 of 226, top 17% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1453 | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1478 | #16 of 298, top 6% | LMArena | 2026-10-08 | |
| LMArena Instruction Following | 1473 | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| CL-bench | 21.1% | #4 of 19, top 22% | Epoch AI | ||
| LMArena Longer Query | 1478 | LMArena | 2026-10-08 | ||
| LMArena Longer Query | 1480 | #24 of 291, top 9% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1448 | LMArena | 2026-10-08 | ||
| LMArena Text | 1451 | #42 of 297, top 15% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1445 | #32 of 295, top 11% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1443 | LMArena | 2026-10-08 | ||
| EQ-Bench Creative Writing | 1687 | #33 of 115, top 29% | EQ-Bench | ||
| LMArena Multi-Turn | 1461 | LMArena | 2026-10-08 | ||
| LMArena Multi-Turn | 1466 | #34 of 295, top 12% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| anthropic | $5 | $25 | $0.50 | 2026-10-10 |
| azure | $5 | $25 | $0.50 | 2026-10-10 |
| bedrock | $5 | $25 | $0.50 | 2026-10-10 |
| openrouter | $5 | $25 | $0.50 | 2026-10-10 |
| vertex | $5 | $25 | $0.50 | 2026-10-10 |
Compare Claude Opus 4.5
- Claude Opus 4.5 vs Claude Opus 4.1
- Claude Opus 4.5 vs Muse Spark
- Claude Opus 4.5 vs Muse Spark 1.2
- Claude Opus 4.5 vs GPT-5
- Claude Opus 4.5 vs MiMo-V2.6-Pro
- Claude Opus 4.5 vs GLM-5.2
- Claude Opus 4.5 vs Claude Sonnet 4.6
- Claude Opus 4.5 vs GPT-6 Astra
- Claude Opus 4.5 vs Gemini 3.8 Flash
- Claude Opus 4.5 vs Kimi K3
- Claude Opus 4.5 vs Grok 4.6
- Claude Opus 4.5 vs Qwen3.8 Max
- Claude Opus 4.5 vs GLM-5.3
- Claude Opus 4.5 vs Muse Spark 1.3
Other Anthropic models
- Claude Fable 5.169.0
- Claude Opus 5.568.6
- Claude Opus 567.8
- Claude Fable 566.8
- Claude Sonnet 5.561.9
- Claude Opus 4.860.7
- Claude Opus 4.758.3
- Claude Opus 4.658.2
Frequently asked questions
How good is Claude Opus 4.5?
Claude Opus 4.5 by Anthropic ranks 47th of 354 ranked models on the Noometry Index as of October 2026, with a score of 50.5. Its strongest category is agentic & tool use, where it ranks 12th. API pricing starts at $5 per million input tokens and $25 per million output tokens, with a 200K-token context window.
How much does Claude Opus 4.5 cost?
Claude Opus 4.5 costs $5 per million input tokens and $25 per million output tokens on Anthropic's own API, with cached input at $0.50.
What is Claude Opus 4.5's context window?
Claude Opus 4.5 accepts up to 200K tokens of input and can write up to 64K tokens in one response.
Is Claude Opus 4.5 open source?
No. Claude Opus 4.5 is proprietary and available only through Anthropic's API and partner platforms.
How fast is Claude Opus 4.5?
Claude Opus 4.5 generated about 13 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Claude Opus 4.5's strengths and weaknesses?
Relative to other ranked models, Claude Opus 4.5 places best in instruction following, long context, agentic & tool use and lowest in multimodal, math, multilingual.
What is Claude Opus 4.5 best at?
Its best category is agentic & tool use, where it ranks 12th on Noometry.