# Noometry > Noometry ranks 354 AI language models (of 514 tracked) on 93 public benchmarks grouped into 10 categories, with API prices, context windows and a source for every number. Rankings use the Noometry Index v1.0; missing results are never counted as zero. Data as of October 10, 2026. ## Main pages - [LLM leaderboard](https://noometry.com/): Overall ranking of every ranked model with score, price and context - [Models](https://noometry.com/models): Directory of all 514 tracked models - [Compare](https://noometry.com/compare): Head-to-head comparisons; 51,689 pairs have enough shared results to judge - [Benchmarks](https://noometry.com/benchmarks): What each of 131 benchmarks measures and who leads it - [Providers](https://noometry.com/providers): AI labs ranked by their best model - [Best-of rankings](https://noometry.com/best): Rankings by skill, type, provider, region and value - [Best model for your job](https://noometry.com/best/for): Shortlists for coding, writing, research, agents and more - [LLM API pricing](https://noometry.com/llm-pricing): Price per million tokens for 335 models - [Price vs performance](https://noometry.com/llm-value): Which models give the most score per dollar - [Rankings over time](https://noometry.com/llm-rankings-over-time): How the leaderboard has moved since 2023 - [Methodology](https://noometry.com/methodology): How the Noometry Index is calculated - [Sources](https://noometry.com/sources): Every data source, license and attribution ## Citable statistics All statistics are regenerated from the live dataset (as of October 10, 2026). Cite as "Noometry" with a link. - [Best AI model right now](https://noometry.com/best/overall): As of October 2026, GPT-6 Astra (OpenAI) leads the Noometry Index at 70.8, ahead of Claude Fable 5.1 at 69.0. - [Best open-weight model](https://noometry.com/best/open-source): Kimi K3 is the highest-ranked model with downloadable weights, #15 overall with 59.5. - [Open vs closed](https://noometry.com/best/open-source): 8 of the top 50 ranked models (16%) have open weights. - [Largest context window](https://noometry.com/best/large-context-window): GPT-6 Astra accepts 1.05M tokens, the most among tracked models. - [Cheapest top-20 model](https://noometry.com/llm-pricing): Gemini 3.8 Flash is the cheapest model in the top 20 at $0.75 per million input tokens. - [Coverage](https://noometry.com/sources): Noometry holds 12,766 benchmark results for 514 models from 41 providers. ## Rankings - [Best LLMs for Coding](https://noometry.com/best/coding): Claude Fable 5.1 leads - [Best Agentic AI Models](https://noometry.com/best/agentic): Claude Opus 5 leads - [Best Reasoning LLMs](https://noometry.com/best/reasoning): GPT-6 Astra leads - [Best LLMs for Math](https://noometry.com/best/math): GPT-6.1 Sol leads - [Most Knowledgeable LLMs](https://noometry.com/best/knowledge): GPT-6 Astra leads - [Best Multimodal AI Models](https://noometry.com/best/multimodal): Claude Opus 5.5 leads - [Best Multilingual LLMs](https://noometry.com/best/multilingual): Gemini 4 Argon leads - [Best LLMs for Instruction Following](https://noometry.com/best/instruction-following): GPT-5.1 leads - [Best Long-Context LLMs](https://noometry.com/best/long-context): o3-pro leads - [Best LLMs for Writing](https://noometry.com/best/writing): Claude Opus 5 leads - [Best AI Models Overall](https://noometry.com/best/overall): GPT-6 Astra leads - [Best Open-Source LLMs](https://noometry.com/best/open-source): Kimi K3 leads - [Best Proprietary LLMs](https://noometry.com/best/proprietary): GPT-6 Astra leads - [Best Reasoning Models](https://noometry.com/best/reasoning-models): GPT-6 Astra leads - [Best Non-Reasoning LLMs](https://noometry.com/best/non-reasoning-models): Grok 4.20 (Non-Reasoning) leads - [Best New AI Models](https://noometry.com/best/newest): GPT-6 Astra leads - [Best LLMs With a 1M+ Token Context Window](https://noometry.com/best/large-context-window): GPT-6 Astra leads - [Best Cheap LLMs (Under $1 per Million Tokens)](https://noometry.com/best/cheap): GPT-5.6 Luna leads - [Best Value LLMs](https://noometry.com/best/best-value-overall): Qwen3.7 Flash leads - [Best Value LLMs for Coding](https://noometry.com/best/best-value-coding): Mercury 2.5 leads - [Best Value LLMs for Agentic & Tool Use](https://noometry.com/best/best-value-agentic): GPT-6 Luna leads - [Best Value LLMs for Reasoning](https://noometry.com/best/best-value-reasoning): Qwen3.7 Flash leads - [Best Value LLMs for Math](https://noometry.com/best/best-value-math): gpt-oss-20b leads - [Best Value LLMs for Knowledge](https://noometry.com/best/best-value-knowledge): Qwen3.7 Flash leads - [Best Value LLMs for Multimodal](https://noometry.com/best/best-value-multimodal): Gemma 4 26B A4B IT leads - [Best OpenAI Models](https://noometry.com/best/openai-models): GPT-6 Astra leads - [Best Anthropic Models](https://noometry.com/best/anthropic-models): Claude Fable 5.1 leads - [Best Google Models](https://noometry.com/best/google-models): Gemini 3.8 Flash leads - [Best Meta Models](https://noometry.com/best/meta-models): Muse Spark 1.3 leads - [Best DeepSeek Models](https://noometry.com/best/deepseek-models): DeepSeek V4 Pro leads - [Best Mistral Models](https://noometry.com/best/mistral-models): Mistral Large 4 leads - [Best xAI Grok Models](https://noometry.com/best/xai-models): Grok 4.6 leads - [Best Alibaba Qwen Models](https://noometry.com/best/alibaba-models): Qwen3.8 Max leads - [Best Moonshot Kimi Models](https://noometry.com/best/moonshot-models): Kimi K3 leads - [Best Z.ai GLM Models](https://noometry.com/best/zai-models): GLM-5.3 leads - [Best MiniMax Models](https://noometry.com/best/minimax-models): MiniMax-M3 leads - [Best Microsoft Models](https://noometry.com/best/microsoft-models): Wizardlm 70b leads - [Best NVIDIA Models](https://noometry.com/best/nvidia-models): Nemotron 3 Ultra leads - [Best Amazon Nova Models](https://noometry.com/best/amazon-models): Amazon Nova Experimental Chat 26 02 10 leads - [Best Cohere Models](https://noometry.com/best/cohere-models): Command A leads - [Best Chinese AI Models](https://noometry.com/best/chinese-models): Kimi K3 leads - [Best European AI Models](https://noometry.com/best/european-models): Mistral Large 4 leads - [Best American AI Models](https://noometry.com/best/american-models): GPT-6 Astra leads ## Best model for your job - [Best AI model for coding](https://noometry.com/best/for/coding): Claude Fable 5.1 scores highest on current data - [Best AI model for debugging](https://noometry.com/best/for/debugging): GPT-6 Astra scores highest on current data - [Best AI model for building websites](https://noometry.com/best/for/frontend): Claude Fable 5.1 scores highest on current data - [Best local LLM for coding](https://noometry.com/best/for/local-code): Kimi K3 scores highest on current data - [Best AI model for agents](https://noometry.com/best/for/agents): Claude Opus 5 scores highest on current data - [Best AI model for RAG](https://noometry.com/best/for/rag): GPT-5 scores highest on current data - [Best AI model for research](https://noometry.com/best/for/research): GPT-6 Astra scores highest on current data - [Best AI model for studying](https://noometry.com/best/for/studying): GPT-6 Astra scores highest on current data - [Best AI model for math](https://noometry.com/best/for/math): GPT-6 Astra scores highest on current data - [Best AI model for hard reasoning](https://noometry.com/best/for/reasoning): GPT-6 Astra scores highest on current data - [Best AI model for data analysis](https://noometry.com/best/for/data-analysis): GPT-6 Astra scores highest on current data - [Best AI model for Excel](https://noometry.com/best/for/spreadsheets): GPT-6 Astra scores highest on current data - [Best AI model for data extraction](https://noometry.com/best/for/extraction): GPT-5 scores highest on current data - [Best AI model for writing](https://noometry.com/best/for/writing): Claude Opus 5 scores highest on current data - [Best AI model for emails](https://noometry.com/best/for/emails): Claude Opus 5 scores highest on current data - [Best AI model for copywriting](https://noometry.com/best/for/marketing): Claude Opus 5 scores highest on current data - [Best AI model for creative writing](https://noometry.com/best/for/fiction): Claude Opus 5 scores highest on current data - [Best AI model for translation](https://noometry.com/best/for/translation): Claude Fable 5.1 scores highest on current data - [Best AI model for language learning](https://noometry.com/best/for/language-learning): Claude Fable 5.1 scores highest on current data - [Best AI model for customer service](https://noometry.com/best/for/support): Claude Opus 5 scores highest on current data - [Best AI model for presentations](https://noometry.com/best/for/presentations): Claude Opus 5.5 scores highest on current data - [Best AI model for PDFs](https://noometry.com/best/for/documents): GPT-5 scores highest on current data - [Best AI assistant for everyday use](https://noometry.com/best/for/general): GPT-6 Astra scores highest on current data ## Top models - [GPT-6 Astra](https://noometry.com/models/gpt-6-astra): #1, OpenAI, 70.8/100, proprietary, 1.05M context, $10/$50 per M tokens - [Claude Fable 5.1](https://noometry.com/models/claude-fable-5-1): #2, Anthropic, 69.0/100, proprietary, 1M context, $10/$50 per M tokens - [Claude Opus 5.5](https://noometry.com/models/claude-opus-5-5): #3, Anthropic, 68.6/100, proprietary, 1M context, $4/$20 per M tokens - [Claude Opus 5](https://noometry.com/models/claude-opus-5): #4, Anthropic, 67.8/100, proprietary, 1M context, $5/$25 per M tokens - [Claude Fable 5](https://noometry.com/models/claude-fable-5): #5, Anthropic, 66.8/100, proprietary, 1M context, $10/$50 per M tokens - [GPT-6.1 Sol](https://noometry.com/models/gpt-6-1-sol): #6, OpenAI, 65.6/100, proprietary, 1.05M context, $2/$10 per M tokens - [GPT-5.6 Sol](https://noometry.com/models/gpt-5-6-sol): #7, OpenAI, 65.0/100, proprietary, 1.05M context, $4/$20 per M tokens - [GPT-5.5 Pro](https://noometry.com/models/gpt-5-5-pro): #8, OpenAI, 64.3/100, proprietary, 1.05M context, $30/$180 per M tokens - [GPT-5.5](https://noometry.com/models/gpt-5-5): #9, OpenAI, 63.4/100, proprietary, 1.05M context, $5/$30 per M tokens - [Claude Sonnet 5.5](https://noometry.com/models/claude-sonnet-5-5): #10, Anthropic, 61.9/100, proprietary, 1M context, $2/$10 per M tokens - [Gemini 3.8 Flash](https://noometry.com/models/gemini-3-8-flash): #11, Google, 61.8/100, proprietary, 1.05M context, $0.75/$3.75 per M tokens - [GPT-6 Sol](https://noometry.com/models/gpt-6-sol): #12, OpenAI, 61.8/100, proprietary, 1.05M context, $2/$10 per M tokens - [Claude Opus 4.8](https://noometry.com/models/claude-opus-4-8): #13, Anthropic, 60.7/100, proprietary, 1M context, $5/$25 per M tokens - [Gemini 3.7 Flash](https://noometry.com/models/gemini-3-7-flash): #14, Google, 59.8/100, proprietary, 1.05M context, $0.75/$3.75 per M tokens - [Kimi K3](https://noometry.com/models/kimi-k3): #15, Moonshot AI, 59.5/100, open weights, 1.05M context, $3/$15 per M tokens - [GPT-5.4](https://noometry.com/models/gpt-5-4): #16, OpenAI, 59.4/100, proprietary, 1.05M context, $2.50/$15 per M tokens - [GPT-5.6 Terra](https://noometry.com/models/gpt-5-6-terra): #17, OpenAI, 59.2/100, proprietary, 1.05M context, $2/$12 per M tokens - [GPT-5.4 Pro](https://noometry.com/models/gpt-5-4-pro): #18, OpenAI, 58.9/100, proprietary, 1.05M context, $30/$180 per M tokens - [Claude Opus 4.7](https://noometry.com/models/claude-opus-4-7): #19, Anthropic, 58.3/100, proprietary, 1M context, $5/$25 per M tokens - [Claude Opus 4.6](https://noometry.com/models/claude-opus-4-6): #20, Anthropic, 58.2/100, proprietary, 1M context, $5/$25 per M tokens - [Grok 4.6](https://noometry.com/models/grok-4-6): #21, xAI, 56.9/100, proprietary, 500K context, $2/$6 per M tokens - [Qwen3.8 Max](https://noometry.com/models/qwen3-8-max): #22, Alibaba (Qwen), 56.8/100, proprietary, 1M context, $2/$6 per M tokens - [Gemini 3.1 Pro Preview](https://noometry.com/models/gemini-3-1-pro-preview): #23, Google, 56.7/100, proprietary, 1.05M context, $2/$12 per M tokens - [Gemini 4 Argon](https://noometry.com/models/gemini-4-argon): #24, Google, 56.5/100, proprietary - [Grok 4.5](https://noometry.com/models/grok-4-5): #25, xAI, 55.0/100, proprietary, 500K context, $2/$6 per M tokens - [GLM-5.3](https://noometry.com/models/glm-5-3): #26, Z.ai (Zhipu), 54.8/100, open weights, 1M context, $1.40/$4.40 per M tokens - [Muse Spark 1.3](https://noometry.com/models/muse-spark-1-3): #27, Meta, 54.8/100, proprietary, 1.05M context, $1.25/$4.25 per M tokens - [Gemini 3 Pro](https://noometry.com/models/gemini-3-pro): #28, Google, 54.8/100, proprietary - [Claude Sonnet 5](https://noometry.com/models/claude-sonnet-5): #29, Anthropic, 54.6/100, proprietary, 1M context, $2/$10 per M tokens - [GPT-5.6 Luna](https://noometry.com/models/gpt-5-6-luna): #30, OpenAI, 54.6/100, proprietary, 1.05M context, $0.20/$1.20 per M tokens ## Benchmarks - [Terminal-Bench](https://noometry.com/benchmarks/terminal-bench): 41 models; GPT-5.5 leads - [APEX-Agents](https://noometry.com/benchmarks/apex-agents): 49 models; Gemini 4 Argon leads - [Berkeley Function Calling Leaderboard](https://noometry.com/benchmarks/bfcl): 49 models; Claude Opus 4.5 leads - [OSWorld 2.0](https://noometry.com/benchmarks/osworld-2): 9 models; Claude Opus 5 leads - [GDPval](https://noometry.com/benchmarks/gdpval): 11 models; GPT-5.2 leads - [Remote Labor Index](https://noometry.com/benchmarks/remote-labor-index): 14 models; GPT-6 Astra leads - [TheAgentCompany](https://noometry.com/benchmarks/the-agent-company): 14 models; DeepSeek-V3.2-Exp leads - [τ²-bench Airline](https://noometry.com/benchmarks/tau2-airline): 7 models; Claude Opus 4.5 leads - [τ²-bench Banking](https://noometry.com/benchmarks/tau2-banking): 26 models; Qwen3.8 Max leads - [τ²-bench Retail](https://noometry.com/benchmarks/tau2-retail): 7 models; Qwen3.5 397B-A17B leads - [τ²-bench Telecom](https://noometry.com/benchmarks/tau2-telecom): 7 models; Qwen3.5 397B-A17B leads - [Cybench](https://noometry.com/benchmarks/cybench): 21 models; Claude Opus 4.6 leads - [DeepResearch Bench](https://noometry.com/benchmarks/deepresearch-bench): 24 models; Claude Opus 4.6 leads - [OSWorld](https://noometry.com/benchmarks/osworld): 8 models; Claude Sonnet 4.6 leads - [PostTrainBench](https://noometry.com/benchmarks/posttrainbench): 11 models; Claude Fable 5 leads - [BALROG](https://noometry.com/benchmarks/balrog): 35 models; GPT-6 Astra leads - [ExploitBench](https://noometry.com/benchmarks/exploitbench): 9 models; Claude Mythos Preview leads - [GBAEval](https://noometry.com/benchmarks/gbaeval): 23 models; Claude Opus 5 leads - [GDP.pdf](https://noometry.com/benchmarks/gdp-pdf): 36 models; GPT-6 Astra leads - [SWE-bench Verified](https://noometry.com/benchmarks/swe-bench-verified): 32 models; Claude Opus 4.7 leads - [DeepSWE](https://noometry.com/benchmarks/deepswe): 29 models; GPT-6.1 Sol leads - [FrontierCode](https://noometry.com/benchmarks/frontiercode): 37 models; Claude Opus 5.5 leads - [SWE-bench Verified (bash only)](https://noometry.com/benchmarks/swe-bench-bash-only): 39 models; Claude Opus 4.5 leads - [Aider Polyglot](https://noometry.com/benchmarks/aider-polyglot): 44 models; GPT-5 leads - [LMArena WebDev](https://noometry.com/benchmarks/arena-webdev): 113 models; Claude Opus 5.5 leads - [CursorBench](https://noometry.com/benchmarks/cursorbench): 14 models; Claude Opus 5.5 leads - [SWE-bench Multilingual](https://noometry.com/benchmarks/swe-bench-multilingual): 13 models; Gemini 3 Flash Preview leads - [FrontierSWE](https://noometry.com/benchmarks/frontierswe): 18 models; GPT-6 Astra leads - [SciCode](https://noometry.com/benchmarks/scicode): 121 models; Claude Opus 5.5 leads - [GSO](https://noometry.com/benchmarks/gso-bench): 31 models; Claude Fable 5.1 leads - [WeirdML](https://noometry.com/benchmarks/weirdml): 119 models; GPT-6 Astra leads - [LMArena Coding](https://noometry.com/benchmarks/arena-coding): 294 models; Gemini 4 Argon leads - [BigCodeBench Instruct](https://noometry.com/benchmarks/bigcodebench-instruct): 64 models; GPT-4o leads - [LiveBench Coding](https://noometry.com/benchmarks/livebench-coding): 39 models; Gemini 2.5 Pro leads - [MirrorCode](https://noometry.com/benchmarks/mirrorcode): 9 models; Claude Opus 5.5 leads - [BigCodeBench Complete](https://noometry.com/benchmarks/bigcodebench-complete): 66 models; DeepSeek-V3 leads - [CadEval](https://noometry.com/benchmarks/cadeval): 14 models; o3 leads - [LiveBench Instruction Following](https://noometry.com/benchmarks/livebench-if): 39 models; GPT-5.1 leads - [LMArena Instruction Following](https://noometry.com/benchmarks/arena-instruction-following): 298 models; Gemini 4 Argon leads - [IFEval](https://noometry.com/benchmarks/ifeval): 57 models; Grok-3 mini leads ## Markdown mirrors Every page has a Markdown version at `https://noometry.com/md/.md` (the home page is `/md/index.md`), for example https://noometry.com/md/best/coding.md. Comparison pages also return Markdown at their canonical URL when requested with `Accept: text/markdown`, and at `https://noometry.com/md/compare/-vs-.md`. - [Full text of key pages](https://noometry.com/llms-full.txt) ## Notes - Data last updated: October 10, 2026 - Method: Noometry Index v1.0 - Sources: models.dev, Epoch AI Benchmarking Hub, OpenRouter, LMArena, Kagi LLM Benchmark, HELM Capabilities, Berkeley Function Calling Leaderboard, Lech Mazur benchmarks, EvalPlus, MathArena, EQ-Bench, BigCodeBench, SWE-bench, τ²-bench, Vectara Hallucination Leaderboard, and model cards (self-reported, labelled as such) - Sitemap: https://noometry.com/sitemap.xml