Benchmark directory

# LLM benchmarks

> Every AI benchmark Noometry tracks, grouped into 10 categories: what each one measures, which model leads it, and where each score comes from.
- Canonical page: https://noometry.com/benchmarks
- Last updated: 2026-10-10
- Title: LLM Benchmarks Explained (October 2026): 131 Tracked

Noometry tracks 131 AI benchmarks across 10 categories, from coding and agentic tool use to math, knowledge and long context. Each one has a page explaining what it measures, which model leads it and where every score comes from.

Last verified October 10, 2026

## Coding

Writing, editing and repairing real code: repository-level bug fixing, multi-language exercises and code generation. [Category ranking →](https://noometry.com/best/coding)

Coding benchmarks
| Benchmark | Models | Leader | Top score |
| --- | --- | --- | --- |
| [SWE-bench Verified](https://noometry.com/benchmarks/swe-bench-verified) | 32 | [Claude Opus 4.7](https://noometry.com/models/claude-opus-4-7) | 83.5% |
| [DeepSWE](https://noometry.com/benchmarks/deepswe) | 29 | [GPT-6.1 Sol](https://noometry.com/models/gpt-6-1-sol) | 75.2% |
| [FrontierCode](https://noometry.com/benchmarks/frontiercode) | 37 | [Claude Opus 5.5](https://noometry.com/models/claude-opus-5-5) | 54.6% |
| [SWE-bench Verified (bash only)](https://noometry.com/benchmarks/swe-bench-bash-only) | 39 | [Claude Opus 4.5](https://noometry.com/models/claude-opus-4-5) | 76.8% |
| [Aider Polyglot](https://noometry.com/benchmarks/aider-polyglot) | 44 | [GPT-5](https://noometry.com/models/gpt-5) | 88% |
| [LMArena WebDev](https://noometry.com/benchmarks/arena-webdev) | 113 | [Claude Opus 5.5](https://noometry.com/models/claude-opus-5-5) | 1813 |
| [CursorBench](https://noometry.com/benchmarks/cursorbench) | 14 | [Claude Opus 5.5](https://noometry.com/models/claude-opus-5-5) | 57.8% |
| [SWE-bench Multilingual](https://noometry.com/benchmarks/swe-bench-multilingual) | 13 | [Gemini 3 Flash Preview](https://noometry.com/models/gemini-3-flash-preview) | 72.7% |
| [FrontierSWE](https://noometry.com/benchmarks/frontierswe) | 18 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | 65.5% |
| [SciCode](https://noometry.com/benchmarks/scicode) | 121 | [Claude Opus 5.5](https://noometry.com/models/claude-opus-5-5) | 66.9% |
| [GSO](https://noometry.com/benchmarks/gso-bench) | 31 | [Claude Fable 5.1](https://noometry.com/models/claude-fable-5-1) | 88.2% |
| [WeirdML](https://noometry.com/benchmarks/weirdml) | 119 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | 93.6% |
| [LMArena Coding](https://noometry.com/benchmarks/arena-coding) | 294 | [Gemini 4 Argon](https://noometry.com/models/gemini-4-argon) | 1550 |
| [BigCodeBench Instruct](https://noometry.com/benchmarks/bigcodebench-instruct) | 64 | [GPT-4o](https://noometry.com/models/gpt-4o) | 51.1% |
| [LiveBench Coding](https://noometry.com/benchmarks/livebench-coding) | 39 | [Gemini 2.5 Pro](https://noometry.com/models/gemini-2-5-pro) | 85.9% |
| [MirrorCode](https://noometry.com/benchmarks/mirrorcode) | 9 | [Claude Opus 5.5](https://noometry.com/models/claude-opus-5-5) | 77.4% |
| [BigCodeBench Complete](https://noometry.com/benchmarks/bigcodebench-complete) | 66 | [DeepSeek-V3](https://noometry.com/models/deepseek-v3) | 62.2% |
| [CadEval](https://noometry.com/benchmarks/cadeval) | 14 | [o3](https://noometry.com/models/o3) | 74% |
| [ALE-Bench](https://noometry.com/benchmarks/ale-bench) | 105 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | 2,951 |
| [AlgoTune](https://noometry.com/benchmarks/algotune) | 18 | [GPT-5.2](https://noometry.com/models/gpt-5-2) | 2.05 |
| [HumanEval+](https://noometry.com/benchmarks/humaneval-plus) | 45 | [o1](https://noometry.com/models/o1) | 89% |
| [MBPP+](https://noometry.com/benchmarks/mbpp-plus) | 38 | [o1](https://noometry.com/models/o1) | 80.2% |

## Agentic & Tool Use

Completing multi-step tasks with tools, terminals, browsers and computers, where the model plans and acts without step-by-step help. [Category ranking →](https://noometry.com/best/agentic)

Agentic & Tool Use benchmarks
| Benchmark | Models | Leader | Top score |
| --- | --- | --- | --- |
| [Terminal-Bench](https://noometry.com/benchmarks/terminal-bench) | 41 | [GPT-5.5](https://noometry.com/models/gpt-5-5) | 84.7% |
| [APEX-Agents](https://noometry.com/benchmarks/apex-agents) | 49 | [Gemini 4 Argon](https://noometry.com/models/gemini-4-argon) | 82.2% |
| [Berkeley Function Calling Leaderboard](https://noometry.com/benchmarks/bfcl) | 49 | [Claude Opus 4.5](https://noometry.com/models/claude-opus-4-5) | 77.5% |
| [OSWorld 2.0](https://noometry.com/benchmarks/osworld-2) | 9 | [Claude Opus 5](https://noometry.com/models/claude-opus-5) | 31.4% |
| [GDPval](https://noometry.com/benchmarks/gdpval) | 11 | [GPT-5.2](https://noometry.com/models/gpt-5-2) | 49.7% |
| [Remote Labor Index](https://noometry.com/benchmarks/remote-labor-index) | 14 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | 20.8% |
| [TheAgentCompany](https://noometry.com/benchmarks/the-agent-company) | 14 | [DeepSeek-V3.2-Exp](https://noometry.com/models/deepseek-v3-2-exp) | 42.9% |
| [τ²-bench Airline](https://noometry.com/benchmarks/tau2-airline) | 7 | [Claude Opus 4.5](https://noometry.com/models/claude-opus-4-5) | 84% |
| [τ²-bench Banking](https://noometry.com/benchmarks/tau2-banking) | 26 | [Qwen3.8 Max](https://noometry.com/models/qwen3-8-max) | 55.1% |
| [τ²-bench Retail](https://noometry.com/benchmarks/tau2-retail) | 7 | [Qwen3.5 397B-A17B](https://noometry.com/models/qwen3-5-397b-a17b) | 84.4% |
| [τ²-bench Telecom](https://noometry.com/benchmarks/tau2-telecom) | 7 | [Qwen3.5 397B-A17B](https://noometry.com/models/qwen3-5-397b-a17b) | 97.8% |
| [Cybench](https://noometry.com/benchmarks/cybench) | 21 | [Claude Opus 4.6](https://noometry.com/models/claude-opus-4-6) | 93% |
| [DeepResearch Bench](https://noometry.com/benchmarks/deepresearch-bench) | 24 | [Claude Opus 4.6](https://noometry.com/models/claude-opus-4-6) | 55.3% |
| [OSWorld](https://noometry.com/benchmarks/osworld) | 8 | [Claude Sonnet 4.6](https://noometry.com/models/claude-sonnet-4-6) | 72.1% |
| [PostTrainBench](https://noometry.com/benchmarks/posttrainbench) | 11 | [Claude Fable 5](https://noometry.com/models/claude-fable-5) | 41.8% |
| [BALROG](https://noometry.com/benchmarks/balrog) | 35 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | 68.3% |
| [ExploitBench](https://noometry.com/benchmarks/exploitbench) | 9 | [Claude Mythos Preview](https://noometry.com/models/claude-mythos-preview) | 73.8% |
| [GBAEval](https://noometry.com/benchmarks/gbaeval) | 23 | [Claude Opus 5](https://noometry.com/models/claude-opus-5) | 79.6% |
| [GDP.pdf](https://noometry.com/benchmarks/gdp-pdf) | 36 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | 34.2% |
| [LMArena Search](https://noometry.com/benchmarks/arena-search) | 32 | [GPT-5.6 Sol](https://noometry.com/models/gpt-5-6-sol) | 1257 |
| [METR Time Horizons](https://noometry.com/benchmarks/metr-time-horizons) | 32 | [Claude Mythos Preview](https://noometry.com/models/claude-mythos-preview) | 85.2% |
| [Vending-Bench 2](https://noometry.com/benchmarks/vending-bench-2) | 60 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | 15,515 |

## Reasoning

Novel problem solving that cannot be answered from memory: abstraction puzzles, trick questions and multi-step logic. [Category ranking →](https://noometry.com/best/reasoning)

Reasoning benchmarks
| Benchmark | Models | Leader | Top score |
| --- | --- | --- | --- |
| [ARC-AGI-2](https://noometry.com/benchmarks/arc-agi-2) | 83 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | 95% |
| [SimpleBench](https://noometry.com/benchmarks/simplebench) | 77 | [Claude Fable 5](https://noometry.com/models/claude-fable-5) | 81.9% |
| [Kagi LLM Benchmark](https://noometry.com/benchmarks/kagi-reasoning) | 99 | [Claude Fable 5](https://noometry.com/models/claude-fable-5) | 91.4% |
| [NYT Connections (extended)](https://noometry.com/benchmarks/nyt-connections) | 91 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | 98.1% |
| [ARC-AGI-1](https://noometry.com/benchmarks/arc-agi-1) | 83 | [Claude Fable 5](https://noometry.com/models/claude-fable-5) | 98.5% |
| [CritPt](https://noometry.com/benchmarks/critpt) | 134 | [GPT-5.6 Sol](https://noometry.com/models/gpt-5-6-sol) | 32.3% |
| [Chess Puzzles](https://noometry.com/benchmarks/chess-puzzles) | 129 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | 72% |
| [EnigmaEval](https://noometry.com/benchmarks/enigmaeval) | 38 | [Claude Fable 5](https://noometry.com/models/claude-fable-5) | 39.3% |
| [Thematic Generalization](https://noometry.com/benchmarks/thematic-generalization) | 23 | [Claude Opus 4.6](https://noometry.com/models/claude-opus-4-6) | 80.6% |
| [LMArena Hard Prompts](https://noometry.com/benchmarks/arena-hard-prompts) | 297 | [Gemini 4 Argon](https://noometry.com/models/gemini-4-argon) | 1548 |
| [EBR-Bench](https://noometry.com/benchmarks/ebr-bench) | 24 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | 76.2% |
| [LiveBench Reasoning](https://noometry.com/benchmarks/livebench-reasoning) | 39 | [GPT-5.1](https://noometry.com/models/gpt-5-1) | 95.8% |
| [Mystery Game Puzzles](https://noometry.com/benchmarks/mystery-game-puzzles) | 74 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | 84% |
| [DTBench](https://noometry.com/benchmarks/dtbench) | 151 | [Claude Opus 5.5](https://noometry.com/models/claude-opus-5-5) | 98.9% |
| [LiveBench Data Analysis](https://noometry.com/benchmarks/livebench-data-analysis) | 39 | [Gemini 2.5 Pro](https://noometry.com/models/gemini-2-5-pro) | 79.9% |
| [LMCA](https://noometry.com/benchmarks/lmca) | 125 | [Claude Opus 5.5](https://noometry.com/models/claude-opus-5-5) | 68.2% |
| [Surface Evolver Bench](https://noometry.com/benchmarks/surface-evolver-bench) | 25 | [Claude Fable 5](https://noometry.com/models/claude-fable-5) | 95% |
| [Adversarial NLI](https://noometry.com/benchmarks/anli) | 9 | [GPT-3.5-turbo](https://noometry.com/models/gpt-3-5-turbo) | 58.1% |
| [BIG-Bench Hard](https://noometry.com/benchmarks/bbh) | 27 | [Gemini 1.5 Pro (May 2024)](https://noometry.com/models/gemini-1-5-pro) | 89.2% |
| [Bench to the Future 3](https://noometry.com/benchmarks/btf-3) | 10 | [GLM-5.3](https://noometry.com/models/glm-5-3) | 0.15 |
| [CommonsenseQA 2.0](https://noometry.com/benchmarks/csqa2) | 2 | [GPT-3.5-turbo](https://noometry.com/models/gpt-3-5-turbo) | 57% |
| [Epoch Capabilities Index](https://noometry.com/benchmarks/epoch-capabilities-index) | 213 | [Claude Opus 5.5](https://noometry.com/models/claude-opus-5-5) | 167.33 |
| [ForecastBench](https://noometry.com/benchmarks/forecastbench) | 72 | [o3](https://noometry.com/models/o3) | 62.5 |
| [HellaSwag](https://noometry.com/benchmarks/hellaswag) | 29 | [GPT-4](https://noometry.com/models/gpt-4) | 95.3% |
| [LAMBADA](https://noometry.com/benchmarks/lambada) | 9 | [Falcon-180B](https://noometry.com/models/falcon-180b) | 79.8% |
| [LiveBench](https://noometry.com/benchmarks/livebench) | 39 | [Gemini 2.5 Pro](https://noometry.com/models/gemini-2-5-pro) | 82.3% |
| [PIQA](https://noometry.com/benchmarks/piqa) | 27 | [GPT-4o mini](https://noometry.com/models/gpt-4o-mini) | 88.7% |
| [WinoGrande](https://noometry.com/benchmarks/winogrande) | 43 | [Llama 3.1-405B](https://noometry.com/models/llama-3-1-405b) | 89.2% |

## Math

Competition and research mathematics, from AIME-style problems to unpublished research-level questions. [Category ranking →](https://noometry.com/best/math)

Math benchmarks
| Benchmark | Models | Leader | Top score |
| --- | --- | --- | --- |
| [FrontierMath (Tiers 1-3)](https://noometry.com/benchmarks/frontiermath) | 81 | [GPT-6.1 Sol](https://noometry.com/models/gpt-6-1-sol) | 93.7% |
| [FrontierMath Tier 4](https://noometry.com/benchmarks/frontiermath-tier-4) | 63 | [GPT-6.1 Sol](https://noometry.com/models/gpt-6-1-sol) | 100% |
| [MathArena Final-Answer Competitions](https://noometry.com/benchmarks/matharena) | 29 | [GPT-5.5](https://noometry.com/models/gpt-5-5) | 94.3% |
| [OTIS Mock AIME 2024-2025](https://noometry.com/benchmarks/otis-mock-aime) | 173 | [Claude Fable 5](https://noometry.com/models/claude-fable-5) | 100% |
| [ProofBench](https://noometry.com/benchmarks/proofbench) | 77 | [Claude Fable 5.1](https://noometry.com/models/claude-fable-5-1) | 100% |
| [Omni-MATH](https://noometry.com/benchmarks/omni-math) | 57 | [GPT-5 Mini](https://noometry.com/models/gpt-5-mini) | 72.2% |
| [LMArena Math](https://noometry.com/benchmarks/arena-math) | 285 | [Gemini 4 Argon](https://noometry.com/models/gemini-4-argon) | 1536 |
| [LiveBench Math](https://noometry.com/benchmarks/livebench-math) | 39 | [GPT-5.1](https://noometry.com/models/gpt-5-1) | 94.5% |
| [MATH Level 5](https://noometry.com/benchmarks/math-level-5) | 79 | [GPT-5](https://noometry.com/models/gpt-5) | 98.1% |
| [FrontierMath (Feb 2025 set)](https://noometry.com/benchmarks/frontiermath-2025-02) | 68 | [GPT-5.5 Pro](https://noometry.com/models/gpt-5-5-pro) | 52.4% |
| [FrontierMath Erdős](https://noometry.com/benchmarks/frontiermath-erdos) | 7 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | 2.9% |
| [FrontierMath Tier 4 (v1)](https://noometry.com/benchmarks/frontiermath-tier-4-v1) | 55 | [AI Co-Mathematician](https://noometry.com/models/ai-co-mathematician) | 47.9% |
| [GSM8K](https://noometry.com/benchmarks/gsm8k) | 38 | [Deepseek Coder v2](https://noometry.com/models/deepseek-coder-v2) | 94.5% |

## Knowledge

Expert-level factual and scientific knowledge, including graduate-level science questions and short-form factual accuracy. [Category ranking →](https://noometry.com/best/knowledge)

Knowledge benchmarks
| Benchmark | Models | Leader | Top score |
| --- | --- | --- | --- |
| [GPQA Diamond](https://noometry.com/benchmarks/gpqa-diamond) | 186 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | 95.8% |
| [Humanity's Last Exam](https://noometry.com/benchmarks/hle) | 41 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | 54.8% |
| [SimpleQA Verified](https://noometry.com/benchmarks/simpleqa-verified) | 77 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | 75.6% |
| [MMLU-Pro](https://noometry.com/benchmarks/mmlu-pro) | 58 | [Gemini 3 Pro](https://noometry.com/models/gemini-3-pro) | 90.3% |
| [Confabulations](https://noometry.com/benchmarks/confabulations) | 51 | [GPT-5](https://noometry.com/models/gpt-5) | 10.3% |
| [Vectara Hallucination Rate](https://noometry.com/benchmarks/vectara-hallucination) | 96 | [GPT-5.4 nano](https://noometry.com/models/gpt-5-4-nano) | 3.1% |
| [LMArena Expert](https://noometry.com/benchmarks/arena-expert) | 273 | [Claude Opus 5](https://noometry.com/models/claude-opus-5) | 1557 |
| [GPQA (HELM)](https://noometry.com/benchmarks/helm-gpqa) | 57 | [Gemini 3 Pro](https://noometry.com/models/gemini-3-pro) | 80.3% |
| [ARC (AI2) Challenge](https://noometry.com/benchmarks/arc-challenge) | 39 | [DeepSeek-V3](https://noometry.com/models/deepseek-v3) | 95.3% |
| [BoolQ](https://noometry.com/benchmarks/boolq) | 23 | [Falcon-180B](https://noometry.com/models/falcon-180b) | 89% |
| [MMLU](https://noometry.com/benchmarks/mmlu) | 81 | [GPT-4o](https://noometry.com/models/gpt-4o) | 88.1% |
| [OpenBookQA](https://noometry.com/benchmarks/openbookqa) | 19 | [Phi 3 Mini 4k Instruct](https://noometry.com/models/phi-3-mini-4k-instruct) | 88% |
| [TriviaQA](https://noometry.com/benchmarks/triviaqa) | 25 | [Llama 2-70B](https://noometry.com/models/llama-2-70b) | 87.6% |

## Multimodal

Understanding images, charts, documents and video alongside text. [Category ranking →](https://noometry.com/best/multimodal)

Multimodal benchmarks
| Benchmark | Models | Leader | Top score |
| --- | --- | --- | --- |
| [LMArena Vision](https://noometry.com/benchmarks/arena-vision) | 122 | [Claude Fable 5](https://noometry.com/models/claude-fable-5) | 1324 |
| [Video-MME](https://noometry.com/benchmarks/video-mme) | 15 | [video-SALMONN 2+](https://noometry.com/models/video-salmonn-2-plus) | 79.7% |
| [GeoBench](https://noometry.com/benchmarks/geobench) | 25 | [Gemini 3 Flash Preview](https://noometry.com/models/gemini-3-flash-preview) | 88% |
| [VPCT](https://noometry.com/benchmarks/vpct) | 24 | [Gemini 3 Pro](https://noometry.com/models/gemini-3-pro) | 91% |
| [Blueprint-Bench 2](https://noometry.com/benchmarks/blueprint-bench-2) | 31 | [Gemini 4 Argon](https://noometry.com/models/gemini-4-argon) | 54.4% |
| [Furniture Assembly](https://noometry.com/benchmarks/furniture-assembly) | 31 | [Claude Opus 5.5](https://noometry.com/models/claude-opus-5-5) | 83.3% |
| [LMArena Document](https://noometry.com/benchmarks/arena-document) | 38 | [Claude Opus 5](https://noometry.com/models/claude-opus-5) | 1516 |
| [MindCube](https://noometry.com/benchmarks/mindcube) | 2 | [Gemma 3 12B](https://noometry.com/models/gemma-3-12b) | 46.7% |
| [ScienceQA](https://noometry.com/benchmarks/scienceqa) | 6 | [GPT-4o](https://noometry.com/models/gpt-4o) | 88.5% |
| [SpatialViz-Bench](https://noometry.com/benchmarks/spatialviz-bench) | 8 | [Gemini 2.5 Pro](https://noometry.com/models/gemini-2-5-pro) | 44.7% |

## Multilingual

Quality in languages other than English. [Category ranking →](https://noometry.com/best/multilingual)

Multilingual benchmarks
| Benchmark | Models | Leader | Top score |
| --- | --- | --- | --- |
| [LMArena Non-English](https://noometry.com/benchmarks/arena-non-english) | 297 | [Gemini 4 Argon](https://noometry.com/models/gemini-4-argon) | 1523 |
| [LMArena Chinese](https://noometry.com/benchmarks/arena-chinese) | 285 | [Gemini 4 Argon](https://noometry.com/models/gemini-4-argon) | 1610 |
| [LMArena French](https://noometry.com/benchmarks/arena-french) | 223 | [Claude Fable 5.1](https://noometry.com/models/claude-fable-5-1) | 1525 |
| [LMArena German](https://noometry.com/benchmarks/arena-german) | 231 | [Claude Opus 5](https://noometry.com/models/claude-opus-5) | 1524 |
| [LMArena Japanese](https://noometry.com/benchmarks/arena-japanese) | 211 | [Claude Fable 5.1](https://noometry.com/models/claude-fable-5-1) | 1543 |
| [LMArena Korean](https://noometry.com/benchmarks/arena-korean) | 213 | [Claude Fable 5.1](https://noometry.com/models/claude-fable-5-1) | 1534 |
| [LMArena Russian](https://noometry.com/benchmarks/arena-russian) | 283 | [Gemini 4 Argon](https://noometry.com/models/gemini-4-argon) | 1537 |
| [LMArena Spanish](https://noometry.com/benchmarks/arena-spanish) | 226 | [Claude Opus 5](https://noometry.com/models/claude-opus-5) | 1519 |

## Instruction Following

Following explicit formatting, length and content constraints exactly. [Category ranking →](https://noometry.com/best/instruction-following)

Instruction Following benchmarks
| Benchmark | Models | Leader | Top score |
| --- | --- | --- | --- |
| [LiveBench Instruction Following](https://noometry.com/benchmarks/livebench-if) | 39 | [GPT-5.1](https://noometry.com/models/gpt-5-1) | 93.3% |
| [LMArena Instruction Following](https://noometry.com/benchmarks/arena-instruction-following) | 298 | [Gemini 4 Argon](https://noometry.com/models/gemini-4-argon) | 1538 |
| [IFEval](https://noometry.com/benchmarks/ifeval) | 57 | [Grok-3 mini](https://noometry.com/models/grok-3-mini) | 95.1% |

## Long Context

Retrieving and reasoning over information spread across very long inputs. [Category ranking →](https://noometry.com/best/long-context)

Long Context benchmarks
| Benchmark | Models | Leader | Top score |
| --- | --- | --- | --- |
| [Fiction.LiveBench](https://noometry.com/benchmarks/fiction-livebench) | 47 | [GPT-5](https://noometry.com/models/gpt-5) | 97.2% |
| [CL-bench](https://noometry.com/benchmarks/cl-bench) | 19 | [GPT-5.4](https://noometry.com/models/gpt-5-4) | 27.9% |
| [LMArena Longer Query](https://noometry.com/benchmarks/arena-longer-query) | 291 | [Gemini 4 Argon](https://noometry.com/models/gemini-4-argon) | 1549 |
| [CL-bench Life](https://noometry.com/benchmarks/cl-bench-life) | 13 | [GPT-5.5](https://noometry.com/models/gpt-5-5) | 22.2% |

## Writing & Preference

How the model's answers are rated in blind side-by-side comparisons, by people and by LLM judges, including creative writing. [Category ranking →](https://noometry.com/best/writing)

Writing & Preference benchmarks
| Benchmark | Models | Leader | Top score |
| --- | --- | --- | --- |
| [LMArena Text](https://noometry.com/benchmarks/arena-text) | 297 | [Gemini 4 Argon](https://noometry.com/models/gemini-4-argon) | 1534 |
| [LMArena Creative Writing](https://noometry.com/benchmarks/arena-creative-writing) | 295 | [Claude Opus 5.5](https://noometry.com/models/claude-opus-5-5) | 1533 |
| [Short-Story Creative Writing](https://noometry.com/benchmarks/lech-mazur-writing) | 39 | [GPT-5](https://noometry.com/models/gpt-5) | 86% |
| [EQ-Bench Creative Writing](https://noometry.com/benchmarks/eqbench-creative-writing) | 115 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | 2173 |
| [WildBench](https://noometry.com/benchmarks/wildbench) | 57 | [Qwen3 235B-A22B](https://noometry.com/models/qwen3-235b-a22b) | 86.6% |
| [LMArena Multi-Turn](https://noometry.com/benchmarks/arena-multi-turn) | 295 | [Gemini 4 Argon](https://noometry.com/models/gemini-4-argon) | 1557 |
| [EQ-Bench 4](https://noometry.com/benchmarks/eqbench-4) | 28 | [Claude Opus 5](https://noometry.com/models/claude-opus-5) | 1385 |
| [LiveBench Language](https://noometry.com/benchmarks/livebench-language) | 39 | [GPT-5.1](https://noometry.com/models/gpt-5-1) | 80.2% |
