Coding benchmark

# CursorBench leaderboard

> CursorBench results for 14 AI models, led by Claude Opus 5.5 at 57.8%. What the benchmark measures, who runs it, and a source for every score.
- Canonical page: https://noometry.com/benchmarks/cursorbench
- Last updated: 2026-10-10
- Title: CursorBench Leaderboard (October 2026): Scores by Model

As of October 2026, Claude Opus 5.5 has the highest published CursorBench score on Noometry at 57.8%, out of 14 models with results.

Last verified October 10, 2026

## About CursorBench

Coding-agent tasks drawn from real Cursor sessions, published by Cursor.

- **Category:** [Coding](https://noometry.com/best/coding)
- **Introduced:** 2026
- **Unit:** Percent (random guessing ≈ 0%)
- **Official site:** [cursor.com](https://cursor.com)

## Top 14 models

Top models on CursorBench

1.  Claude Opus 5.5 57.8%
2.  Claude Sonnet 5.5 55.5%
3.  Claude Fable 5.1 51.8%
4.  Claude Opus 5 46.6%
5.  Grok 4.7 46.3%
6.  GLM-5.3 42.6%
7.  GPT-5.6 Sol 41.7%
8.  Muse Spark 1.3 41.6%
9.  Grok 4.6 41.4%
10.  GPT-5.6 Terra 41.3%
11.  Gemini 3.8 Flash 39.6%
12.  GLM-5.3-Flash 36.8%
13.  GPT-5.6 Luna 35.9%
14.  Claude Sonnet 5 34.1%
15.  2030405060

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## All results

CursorBench results by model
| # | Model | Provider | Score | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | [Claude Opus 5.5](https://noometry.com/models/claude-opus-5-5) | [Anthropic](https://noometry.com/providers/anthropic) | 57.8% | max | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 2 | [Claude Sonnet 5.5](https://noometry.com/models/claude-sonnet-5-5) | [Anthropic](https://noometry.com/providers/anthropic) | 55.5% | max | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 3 | [Claude Fable 5.1](https://noometry.com/models/claude-fable-5-1) | [Anthropic](https://noometry.com/providers/anthropic) | 51.8% | max | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 4 | [Claude Opus 5](https://noometry.com/models/claude-opus-5) | [Anthropic](https://noometry.com/providers/anthropic) | 46.6% | max | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 5 | [Grok 4.7](https://noometry.com/models/grok-4-7) | [xAI](https://noometry.com/providers/xai) | 46.3% | xhigh | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 6 | [GLM-5.3](https://noometry.com/models/glm-5-3) | [Z.ai (Zhipu)](https://noometry.com/providers/zai) | 42.6% | max | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 7 | [GPT-5.6 Sol](https://noometry.com/models/gpt-5-6-sol) | [OpenAI](https://noometry.com/providers/openai) | 41.7% | max | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 8 | [Muse Spark 1.3](https://noometry.com/models/muse-spark-1-3) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 41.6% | max | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 9 | [Grok 4.6](https://noometry.com/models/grok-4-6) | [xAI](https://noometry.com/providers/xai) | 41.4% | xhigh | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 10 | [GPT-5.6 Terra](https://noometry.com/models/gpt-5-6-terra) | [OpenAI](https://noometry.com/providers/openai) | 41.3% | max | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 11 | [Gemini 3.8 Flash](https://noometry.com/models/gemini-3-8-flash) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 39.6% | high | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 12 | [GLM-5.3-Flash](https://noometry.com/models/glm-5-3-flash) | [Z.ai (Zhipu)](https://noometry.com/providers/zai) | 36.8% | max | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 13 | [GPT-5.6 Luna](https://noometry.com/models/gpt-5-6-luna) | [OpenAI](https://noometry.com/providers/openai) | 35.9% | max | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 14 | [Claude Sonnet 5](https://noometry.com/models/claude-sonnet-5) | [Anthropic](https://noometry.com/providers/anthropic) | 34.1% | max | [Epoch AI](https://epoch.ai/benchmarks) |  |

## Compare the leaders

-   [Claude Opus 5.5 vs Claude Sonnet 5.5](https://noometry.com/compare/claude-opus-5-5-vs-claude-sonnet-5-5)
-   [Claude Opus 5.5 vs Claude Fable 5.1](https://noometry.com/compare/claude-fable-5-1-vs-claude-opus-5-5)
-   [Claude Opus 5.5 vs Claude Opus 5](https://noometry.com/compare/claude-opus-5-vs-claude-opus-5-5)
-   [Claude Opus 5.5 vs Grok 4.7](https://noometry.com/compare/claude-opus-5-5-vs-grok-4-7)
-   [Claude Sonnet 5.5 vs Claude Fable 5.1](https://noometry.com/compare/claude-fable-5-1-vs-claude-sonnet-5-5)
-   [Claude Sonnet 5.5 vs Claude Opus 5](https://noometry.com/compare/claude-opus-5-vs-claude-sonnet-5-5)

## Other coding benchmarks

-   [SWE-bench Verified](https://noometry.com/benchmarks/swe-bench-verified)
-   [DeepSWE](https://noometry.com/benchmarks/deepswe)
-   [FrontierCode](https://noometry.com/benchmarks/frontiercode)
-   [SWE-bench Verified (bash only)](https://noometry.com/benchmarks/swe-bench-bash-only)
-   [Aider Polyglot](https://noometry.com/benchmarks/aider-polyglot)
-   [LMArena WebDev](https://noometry.com/benchmarks/arena-webdev)
-   [SWE-bench Multilingual](https://noometry.com/benchmarks/swe-bench-multilingual)
-   [FrontierSWE](https://noometry.com/benchmarks/frontierswe)
-   [SciCode](https://noometry.com/benchmarks/scicode)
-   [GSO](https://noometry.com/benchmarks/gso-bench)
-   [WeirdML](https://noometry.com/benchmarks/weirdml)
-   [LMArena Coding](https://noometry.com/benchmarks/arena-coding)

## Frequently asked questions

### What does CursorBench measure?

Coding-agent tasks drawn from real Cursor sessions, published by Cursor.

### Which model has the highest CursorBench score?

As of October 2026, Claude Opus 5.5 has the highest published CursorBench score on Noometry at 57.8%, out of 14 models with results.

### What is the best open-weight model on CursorBench?

GLM-5.3 has the highest CursorBench accuracy among open-weight models at 42.6%, ranking 6 of 14 overall.

### Cite this page

Noometry. (2026). CursorBench leaderboard. Retrieved October 10, 2026, from https://noometry.com/benchmarks/cursorbench

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/benchmarks/cursorbench.md).
