Noometry Index v1.0

# How the Noometry Index works

> How Noometry turns 12,766 public benchmark results into one score per AI model: the sources it uses, how it keeps comparisons fair and what the ten categories cover.
- Canonical page: https://noometry.com/methodology
- Last updated: 2026-10-10
- Title: How the Noometry Index Works

The Noometry Index gives every AI model one score from 0 to 100, built from public benchmark results by independent evaluators. It compares models fairly even when they have been tested on different benchmarks, and it favours independent results over what labs report about themselves.

Last verified October 10, 2026

By [Noometry Editorial](https://noometry.com/authors/editorial)

## What goes in

We collect 12,766 results from 15 public sources, most of them independent evaluators and leaderboards, plus the numbers labs publish in their model cards. Every result keeps a link to where it came from, the setting it was run at and the date. The [sources page](https://noometry.com/sources) lists them all.

## How we keep it fair

-   **Independent results come first.** When a model has been tested by an outside evaluator, that result is preferred over the lab's own number. Self-reported results are labelled and count for less.
-   **Harder tests count for more.** A strong result on a difficult benchmark says more than the same result on an easy one, so every benchmark is adjusted for how hard it is.
-   **Models are compared on what they actually ran.** Skipping a hard test is neither rewarded nor punished, and a missing result is never treated as a zero.
-   **One lucky result can't top the table.** A model needs a body of evidence to rank highly. Models with only a few results stay conservative until more are published.

## Ten categories, one score

Each model gets a score in each of ten categories, and the overall score combines them, with the skills most people rely on day to day counting for more.

| Category | What it covers |
| --- | --- |
| [Coding](https://noometry.com/best/coding) | Writing, editing and repairing real code: repository-level bug fixing, multi-language exercises and code generation. |
| [Agentic & Tool Use](https://noometry.com/best/agentic) | Completing multi-step tasks with tools, terminals, browsers and computers, where the model plans and acts without step-by-step help. |
| [Reasoning](https://noometry.com/best/reasoning) | Novel problem solving that cannot be answered from memory: abstraction puzzles, trick questions and multi-step logic. |
| [Math](https://noometry.com/best/math) | Competition and research mathematics, from AIME-style problems to unpublished research-level questions. |
| [Knowledge](https://noometry.com/best/knowledge) | Expert-level factual and scientific knowledge, including graduate-level science questions and short-form factual accuracy. |
| [Multimodal](https://noometry.com/best/multimodal) | Understanding images, charts, documents and video alongside text. |
| [Multilingual](https://noometry.com/best/multilingual) | Quality in languages other than English. |
| [Instruction Following](https://noometry.com/best/instruction-following) | Following explicit formatting, length and content constraints exactly. |
| [Long Context](https://noometry.com/best/long-context) | Retrieving and reasoning over information spread across very long inputs. |
| [Writing & Preference](https://noometry.com/best/writing) | How the model's answers are rated in blind side-by-side comparisons, by people and by LLM judges, including creative writing. |

## How much evidence is behind a score

-   **Confirmed**: backed by plenty of independent results across several categories.
-   **Reported**: ranked, but with fewer results or mostly self-reported ones.
-   **Sparse**: not enough to rank yet. The model still has a page with every result we have.

## What the index is not

It measures what public benchmarks measure. It says nothing about uptime, safety policies or how a model handles your particular data. Use it to build a shortlist, then test the shortlist on your own work.

## Changelog

-   **v1.0** (September 2026): first public version, with ten categories and evidence tiers.

## Frequently asked questions

### What data does the Noometry Index use?

Results from independent evaluators and public leaderboards, including Epoch AI, LMArena, SWE-bench, HELM, MathArena, τ²-bench, the Berkeley Function Calling Leaderboard, Kagi, Vectara and EQ-Bench, plus model cards published by the labs. The full list is on the sources page. Self-reported numbers are labelled and count for less than independent results.

### Why is a model with fewer benchmarks ranked lower?

It is not penalised for missing tests. Gaps simply cannot lift a model, so a strong model with thin coverage can sit a little lower until more independent results are published, then it moves up.

### Can a lab pay to change its score?

No. Scores come only from published results and the same method applies to every model. Advertising is labelled and kept out of the rankings.

### When do scores change?

Scores move when a source publishes new results, when a new model is added, or when the method version changes. Method changes are logged on this page.
