Reasoning benchmark

# LAMBADA leaderboard

> LAMBADA results for 9 AI models, led by Falcon-180B at 79.8%. What the benchmark measures, who runs it, and a source for every score.
- Canonical page: https://noometry.com/benchmarks/lambada
- Last updated: 2026-10-10
- Title: LAMBADA Leaderboard (October 2026): Scores by Model

As of October 2026, Falcon-180B has the highest published LAMBADA score on Noometry at 79.8%, out of 9 models with results.

Last verified October 10, 2026

## About LAMBADA

Predicting the last word of a passage, which needs the broader context.

- **Category:** [Reasoning](https://noometry.com/best/reasoning)
- **Introduced:** 2016
- **Format:** Word prediction
- **Unit:** Percent (random guessing ≈ 0%)
- **Official site:** [zenodo.org](https://zenodo.org/record/2630551)

## Top 9 models

Top models on LAMBADA

1.  Falcon-180B 79.8%
2.  Llama 2-70B 78.9%
3.  Falcon-40B 77.3%
4.  Llama 2-13B 76.5%
5.  Llama 13b 75.2%
6.  Falcon-7B 74.9%
7.  Llama 2-7B 73.3%
8.  Qwen-14B 71.1%
9.  Qwen-7B 67.9%
10.  6065707580

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## All results

LAMBADA results by model
| # | Model | Provider | Score | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | [Falcon-180B](https://noometry.com/models/falcon-180b) |  [![](/logos/tii.svg) Technology Innovation Institute](https://noometry.com/providers/tii) | 79.8% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 2 | [Llama 2-70B](https://noometry.com/models/llama-2-70b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 78.9% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 3 | [Falcon-40B](https://noometry.com/models/falcon-40b) |  [![](/logos/tii.svg) Technology Innovation Institute](https://noometry.com/providers/tii) | 77.3% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 4 | [Llama 2-13B](https://noometry.com/models/llama-2-13b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 76.5% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 5 | [Llama 13b](https://noometry.com/models/llama-13b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 75.2% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 6 | [Falcon-7B](https://noometry.com/models/falcon-7b) |  [![](/logos/tii.svg) Technology Innovation Institute](https://noometry.com/providers/tii) | 74.9% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 7 | [Llama 2-7B](https://noometry.com/models/llama-2-7b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 73.3% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 8 | [Qwen-14B](https://noometry.com/models/qwen-14b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 71.1% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 9 | [Qwen-7B](https://noometry.com/models/qwen-7b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 67.9% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

## Compare the leaders

-   [Falcon-180B vs Llama 2-70B](https://noometry.com/compare/falcon-180b-vs-llama-2-70b)
-   [Falcon-180B vs Falcon-40B](https://noometry.com/compare/falcon-180b-vs-falcon-40b)
-   [Falcon-180B vs Llama 2-13B](https://noometry.com/compare/falcon-180b-vs-llama-2-13b)
-   [Falcon-180B vs Llama 13b](https://noometry.com/compare/falcon-180b-vs-llama-13b)
-   [Llama 2-70B vs Falcon-40B](https://noometry.com/compare/falcon-40b-vs-llama-2-70b)
-   [Llama 2-70B vs Llama 2-13B](https://noometry.com/compare/llama-2-13b-vs-llama-2-70b)

## Other reasoning benchmarks

-   [ARC-AGI-2](https://noometry.com/benchmarks/arc-agi-2)
-   [SimpleBench](https://noometry.com/benchmarks/simplebench)
-   [Kagi LLM Benchmark](https://noometry.com/benchmarks/kagi-reasoning)
-   [NYT Connections (extended)](https://noometry.com/benchmarks/nyt-connections)
-   [ARC-AGI-1](https://noometry.com/benchmarks/arc-agi-1)
-   [CritPt](https://noometry.com/benchmarks/critpt)
-   [Chess Puzzles](https://noometry.com/benchmarks/chess-puzzles)
-   [EnigmaEval](https://noometry.com/benchmarks/enigmaeval)
-   [Thematic Generalization](https://noometry.com/benchmarks/thematic-generalization)
-   [LMArena Hard Prompts](https://noometry.com/benchmarks/arena-hard-prompts)
-   [EBR-Bench](https://noometry.com/benchmarks/ebr-bench)
-   [LiveBench Reasoning](https://noometry.com/benchmarks/livebench-reasoning)

## Frequently asked questions

### What does LAMBADA measure?

Predicting the last word of a passage, which needs the broader context.

### Which model has the highest LAMBADA score?

As of October 2026, Falcon-180B has the highest published LAMBADA score on Noometry at 79.8%, out of 9 models with results.

### What is the best open-weight model on LAMBADA?

Falcon-180B has the highest LAMBADA accuracy among open-weight models at 79.8%, ranking 1 of 9 overall.

### Cite this page

Noometry. (2026). LAMBADA leaderboard. Retrieved October 10, 2026, from https://noometry.com/benchmarks/lambada

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/benchmarks/lambada.md).
