Reasoning benchmark

# CommonsenseQA 2.0 leaderboard

> CommonsenseQA 2.0 results for 2 AI models, led by GPT-3.5-turbo at 57%. What the benchmark measures, who runs it, and a source for every score.
- Canonical page: https://noometry.com/benchmarks/csqa2
- Last updated: 2026-10-10
- Title: CommonsenseQA 2.0 Leaderboard (October 2026): Scores by Model

As of October 2026, GPT-3.5-turbo has the highest published CommonsenseQA 2.0 score on Noometry at 57%, out of 2 models with results.

Last verified October 10, 2026

## About CommonsenseQA 2.0

Yes/no commonsense questions collected through a model-in-the-loop game.

- **Category:** [Reasoning](https://noometry.com/best/reasoning)
- **Introduced:** 2022
- **Format:** Yes/no
- **Unit:** Percent (random guessing ≈ 50%)
- **Official site:** [allenai.github.io](https://allenai.github.io/csqa2/)

## Top 2 models

Top models on CommonsenseQA 2.0

1.  GPT-3.5-turbo 57%
2.  Llama 2-70B 50%
3.  485052545658

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## All results

CommonsenseQA 2.0 results by model
| # | Model | Provider | Score | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | [GPT-3.5-turbo](https://noometry.com/models/gpt-3-5-turbo) | [OpenAI](https://noometry.com/providers/openai) | 57% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 2 | [Llama 2-70B](https://noometry.com/models/llama-2-70b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 50% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

## Compare the leaders

-   [GPT-3.5-turbo vs Llama 2-70B](https://noometry.com/compare/gpt-3-5-turbo-vs-llama-2-70b)

## Other reasoning benchmarks

-   [ARC-AGI-2](https://noometry.com/benchmarks/arc-agi-2)
-   [SimpleBench](https://noometry.com/benchmarks/simplebench)
-   [Kagi LLM Benchmark](https://noometry.com/benchmarks/kagi-reasoning)
-   [NYT Connections (extended)](https://noometry.com/benchmarks/nyt-connections)
-   [ARC-AGI-1](https://noometry.com/benchmarks/arc-agi-1)
-   [CritPt](https://noometry.com/benchmarks/critpt)
-   [Chess Puzzles](https://noometry.com/benchmarks/chess-puzzles)
-   [EnigmaEval](https://noometry.com/benchmarks/enigmaeval)
-   [Thematic Generalization](https://noometry.com/benchmarks/thematic-generalization)
-   [LMArena Hard Prompts](https://noometry.com/benchmarks/arena-hard-prompts)
-   [EBR-Bench](https://noometry.com/benchmarks/ebr-bench)
-   [LiveBench Reasoning](https://noometry.com/benchmarks/livebench-reasoning)

## Frequently asked questions

### What does CommonsenseQA 2.0 measure?

Yes/no commonsense questions collected through a model-in-the-loop game.

### Which model has the highest CommonsenseQA 2.0 score?

As of October 2026, GPT-3.5-turbo has the highest published CommonsenseQA 2.0 score on Noometry at 57%, out of 2 models with results.

### What is the best open-weight model on CommonsenseQA 2.0?

Llama 2-70B has the highest CommonsenseQA 2.0 accuracy among open-weight models at 50%, ranking 2 of 2 overall.

### Cite this page

Noometry. (2026). CommonsenseQA 2.0 leaderboard. Retrieved October 10, 2026, from https://noometry.com/benchmarks/csqa2

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/benchmarks/csqa2.md).
