Benchmark directory

LLM benchmarks

Noometry tracks 131 AI benchmarks across 10 categories, from coding and agentic tool use to math, knowledge and long context. Each one has a page explaining what it measures, which model leads it and where every score comes from.

Last verified

Coding

Writing, editing and repairing real code: repository-level bug fixing, multi-language exercises and code generation. Category ranking →

Agentic & Tool Use

Completing multi-step tasks with tools, terminals, browsers and computers, where the model plans and acts without step-by-step help. Category ranking →

Reasoning

Novel problem solving that cannot be answered from memory: abstraction puzzles, trick questions and multi-step logic. Category ranking →

Math

Competition and research mathematics, from AIME-style problems to unpublished research-level questions. Category ranking →

Knowledge

Expert-level factual and scientific knowledge, including graduate-level science questions and short-form factual accuracy. Category ranking →

Multimodal

Understanding images, charts, documents and video alongside text. Category ranking →

Multilingual

Quality in languages other than English. Category ranking →

Instruction Following

Following explicit formatting, length and content constraints exactly. Category ranking →

Instruction Following benchmarks
BenchmarkModelsLeaderTop score
LiveBench Instruction Following39GPT-5.193.3%
LMArena Instruction Following298Gemini 4 Argon1538
IFEval57Grok-3 mini95.1%

Long Context

Retrieving and reasoning over information spread across very long inputs. Category ranking →

Long Context benchmarks
BenchmarkModelsLeaderTop score
Fiction.LiveBench47GPT-597.2%
CL-bench19GPT-5.427.9%
LMArena Longer Query291Gemini 4 Argon1549
CL-bench Life13GPT-5.522.2%

Writing & Preference

How the model's answers are rated in blind side-by-side comparisons, by people and by LLM judges, including creative writing. Category ranking →