Math benchmark

# MathArena Final-Answer Competitions leaderboard

> MathArena Final-Answer Competitions results for 29 AI models, led by GPT-5.5 at 94.3%. What the benchmark measures, who runs it, and a source for every score.
- Canonical page: https://noometry.com/benchmarks/matharena
- Last updated: 2026-10-10
- Title: MathArena Final-Answer Competitions Leaderboard (October 2026): Scores by Model

As of October 2026, GPT-5.5 has the highest published MathArena Final-Answer Competitions score on Noometry at 94.3%, out of 29 models with results.

Last verified October 10, 2026

## About MathArena Final-Answer Competitions

Problems from recent math competitions such as AIME and HMMT, run shortly after each contest so the answers are unlikely to be in training data, averaged across competitions.

- **Category:** [Math](https://noometry.com/best/math)
- **Introduced:** 2025
- **Format:** Final answer
- **Unit:** Percent (random guessing ≈ 0%)
- **Official site:** [matharena.ai](https://matharena.ai/)

## Top 15 models

Top models on MathArena Final-Answer Competitions

1.  GPT-5.5 94.3%
2.  Claude Opus 4.8 91.8%
3.  Kimi K3 87.8%
4.  Gemini 3.1 Pro Preview 86.5%
5.  GPT-5.4 83.1%
6.  Claude Opus 4.6 78.5%
7.  DeepSeek V4 Pro 76.6%
8.  DeepSeek V4 Flash 76.5%
9.  Gemini 3.5 Flash 76.3%
10.  Claude Opus 4.7 73.6%
11.  Kimi K2.6 72.9%
12.  GPT-5.2 72%
13.  Gemini 3.6 Flash 70.8%
14.  Step 3.7 Flash 68.5%
15.  Gemini 3 Flash Preview 67.6%
16.  60708090100

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## All results

MathArena Final-Answer Competitions results by model
| # | Model | Provider | Score | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | [GPT-5.5](https://noometry.com/models/gpt-5-5) | [OpenAI](https://noometry.com/providers/openai) | 94.3% | xhigh | [MathArena](https://matharena.ai/) |  |
| 2 | [Claude Opus 4.8](https://noometry.com/models/claude-opus-4-8) | [Anthropic](https://noometry.com/providers/anthropic) | 91.8% | max | [MathArena](https://matharena.ai/) |  |
| 3 | [Kimi K3](https://noometry.com/models/kimi-k3) | [Moonshot AI](https://noometry.com/providers/moonshot) | 87.8% | think | [MathArena](https://matharena.ai/) |  |
| 4 | [Gemini 3.1 Pro Preview](https://noometry.com/models/gemini-3-1-pro-preview) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 86.5% |  | [MathArena](https://matharena.ai/) |  |
| 5 | [GPT-5.4](https://noometry.com/models/gpt-5-4) | [OpenAI](https://noometry.com/providers/openai) | 83.1% | xhigh | [MathArena](https://matharena.ai/) |  |
| 6 | [Claude Opus 4.6](https://noometry.com/models/claude-opus-4-6) | [Anthropic](https://noometry.com/providers/anthropic) | 78.5% | high | [MathArena](https://matharena.ai/) |  |
| 7 | [DeepSeek V4 Pro](https://noometry.com/models/deepseek-v4-pro) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 76.6% | max | [MathArena](https://matharena.ai/) |  |
| 8 | [DeepSeek V4 Flash](https://noometry.com/models/deepseek-v4-flash) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 76.5% | max | [MathArena](https://matharena.ai/) |  |
| 9 | [Gemini 3.5 Flash](https://noometry.com/models/gemini-3-5-flash) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 76.3% |  | [MathArena](https://matharena.ai/) |  |
| 10 | [Claude Opus 4.7](https://noometry.com/models/claude-opus-4-7) | [Anthropic](https://noometry.com/providers/anthropic) | 73.6% | xhigh | [MathArena](https://matharena.ai/) |  |
| 11 | [Kimi K2.6](https://noometry.com/models/kimi-k2-6) | [Moonshot AI](https://noometry.com/providers/moonshot) | 72.9% | think | [MathArena](https://matharena.ai/) |  |
| 12 | [GPT-5.2](https://noometry.com/models/gpt-5-2) | [OpenAI](https://noometry.com/providers/openai) | 72% | high | [MathArena](https://matharena.ai/) |  |
| 13 | [Gemini 3.6 Flash](https://noometry.com/models/gemini-3-6-flash) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 70.8% |  | [MathArena](https://matharena.ai/) |  |
| 14 | [Step 3.7 Flash](https://noometry.com/models/step-3-7-flash) |  [![](/logos/stepfun.svg) StepFun](https://noometry.com/providers/stepfun) | 68.5% |  | [MathArena](https://matharena.ai/) |  |
| 15 | [Gemini 3 Flash Preview](https://noometry.com/models/gemini-3-flash-preview) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 67.6% |  | [MathArena](https://matharena.ai/) |  |
| 16 | [GLM-5.2](https://noometry.com/models/glm-5-2) | [Z.ai (Zhipu)](https://noometry.com/providers/zai) | 67.6% |  | [MathArena](https://matharena.ai/) |  |
| 17 | [GLM-5.1](https://noometry.com/models/glm-5-1) | [Z.ai (Zhipu)](https://noometry.com/providers/zai) | 67.1% |  | [MathArena](https://matharena.ai/) |  |
| 18 | [Gemini 3 Pro](https://noometry.com/models/gemini-3-pro) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 67% | preview | [MathArena](https://matharena.ai/) |  |
| 19 | [Step 3.5 Flash](https://noometry.com/models/step-3-5-flash) |  [![](/logos/stepfun.svg) StepFun](https://noometry.com/providers/stepfun) | 66.8% |  | [MathArena](https://matharena.ai/) |  |
| 20 | [GLM-5](https://noometry.com/models/glm-5) | [Z.ai (Zhipu)](https://noometry.com/providers/zai) | 65.7% |  | [MathArena](https://matharena.ai/) |  |
| 21 | [Kimi K2.5](https://noometry.com/models/kimi-k2-5) | [Moonshot AI](https://noometry.com/providers/moonshot) | 62.3% | think | [MathArena](https://matharena.ai/) |  |
| 22 | [Grok 4.1 Fast](https://noometry.com/models/grok-4-1-fast) | [xAI](https://noometry.com/providers/xai) | 60.9% | reasoning | [MathArena](https://matharena.ai/) |  |
| 23 | [Nemotron 3 Super](https://noometry.com/models/nemotron-3-super) |  [![](/logos/nvidia.svg) NVIDIA](https://noometry.com/providers/nvidia) | 60.4% |  | [MathArena](https://matharena.ai/) |  |
| 24 | [DeepSeek-V3.2-Exp](https://noometry.com/models/deepseek-v3-2-exp) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 57.7% | think | [MathArena](https://matharena.ai/) |  |
| 25 | [Qwen3.5 27B](https://noometry.com/models/qwen3-5-27b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 56.7% |  | [MathArena](https://matharena.ai/) |  |
| 26 | [Qwen3.5 35B-A3B](https://noometry.com/models/qwen3-5-35b-a3b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 56% |  | [MathArena](https://matharena.ai/) |  |
| 27 | [Qwen3.5-9B](https://noometry.com/models/qwen3-5-9b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 48.5% |  | [MathArena](https://matharena.ai/) |  |
| 28 | [Qwen3-30B-A3B](https://noometry.com/models/qwen3-30b-a3b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 47.8% |  | [MathArena](https://matharena.ai/) |  |
| 29 | [Qwen3-4B](https://noometry.com/models/qwen3-4b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 38.5% |  | [MathArena](https://matharena.ai/) |  |

## Compare the leaders

-   [GPT-5.5 vs Claude Opus 4.8](https://noometry.com/compare/claude-opus-4-8-vs-gpt-5-5)
-   [GPT-5.5 vs Kimi K3](https://noometry.com/compare/gpt-5-5-vs-kimi-k3)
-   [GPT-5.5 vs Gemini 3.1 Pro Preview](https://noometry.com/compare/gemini-3-1-pro-preview-vs-gpt-5-5)
-   [GPT-5.5 vs GPT-5.4](https://noometry.com/compare/gpt-5-4-vs-gpt-5-5)
-   [Claude Opus 4.8 vs Kimi K3](https://noometry.com/compare/claude-opus-4-8-vs-kimi-k3)
-   [Claude Opus 4.8 vs Gemini 3.1 Pro Preview](https://noometry.com/compare/claude-opus-4-8-vs-gemini-3-1-pro-preview)

## Other math benchmarks

-   [FrontierMath (Tiers 1-3)](https://noometry.com/benchmarks/frontiermath)
-   [FrontierMath Tier 4](https://noometry.com/benchmarks/frontiermath-tier-4)
-   [OTIS Mock AIME 2024-2025](https://noometry.com/benchmarks/otis-mock-aime)
-   [ProofBench](https://noometry.com/benchmarks/proofbench)
-   [Omni-MATH](https://noometry.com/benchmarks/omni-math)
-   [LMArena Math](https://noometry.com/benchmarks/arena-math)
-   [LiveBench Math](https://noometry.com/benchmarks/livebench-math)
-   [MATH Level 5](https://noometry.com/benchmarks/math-level-5)
-   [FrontierMath (Feb 2025 set)](https://noometry.com/benchmarks/frontiermath-2025-02) (reference)
-   [FrontierMath Erdős](https://noometry.com/benchmarks/frontiermath-erdos) (reference)
-   [FrontierMath Tier 4 (v1)](https://noometry.com/benchmarks/frontiermath-tier-4-v1) (reference)
-   [GSM8K](https://noometry.com/benchmarks/gsm8k) (reference)

## Frequently asked questions

### What does MathArena Final-Answer Competitions measure?

Problems from recent math competitions such as AIME and HMMT, run shortly after each contest so the answers are unlikely to be in training data, averaged across competitions.

### Which model has the highest MathArena Final-Answer Competitions score?

As of October 2026, GPT-5.5 has the highest published MathArena Final-Answer Competitions score on Noometry at 94.3%, out of 29 models with results.

### What is the best open-weight model on MathArena Final-Answer Competitions?

Kimi K3 has the highest MathArena Final-Answer Competitions accuracy among open-weight models at 87.8%, ranking 3 of 29 overall.

### Cite this page

Noometry. (2026). MathArena Final-Answer Competitions leaderboard. Retrieved October 10, 2026, from https://noometry.com/benchmarks/matharena

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/benchmarks/matharena.md).
