Multimodal benchmark

# Video-MME leaderboard

> Video-MME results for 15 AI models, led by video-SALMONN 2+ at 79.7%. What the benchmark measures, who runs it, and a source for every score.
- Canonical page: https://noometry.com/benchmarks/video-mme
- Last updated: 2026-10-10
- Title: Video-MME Leaderboard (October 2026): Scores by Model

As of October 2026, video-SALMONN 2+ has the highest published Video-MME score on Noometry at 79.7%, out of 15 models with results.

Last verified October 10, 2026

## About Video-MME

Multiple-choice questions about short, medium and long videos. Scores here are without subtitles.

- **Category:** [Multimodal](https://noometry.com/best/multimodal)
- **Introduced:** 2024
- **Size:** 900 videos
- **Format:** Multiple choice
- **Unit:** Percent (random guessing ≈ 25%)
- **Official site:** [video-mme.github.io](https://video-mme.github.io)

## Top 15 models

Top models on Video-MME

1.  video-SALMONN 2+ 79.7%
2.  Gemini 1.5 Pro (May 2024) 75%
3.  Gemini 1.5 Pro (May 2024) 75%
4.  Qwen2.5-VL 72B Instruct 73.5%
5.  InternVL2\_5-78B 72.1%
6.  GPT-4o 71.9%
7.  GPT-4o 71.9%
8.  Gemini 1.5 Flash (May 2024) 70.3%
9.  GPT-4o mini 64.8%
10.  NVILA 8B 64.2%
11.  InternVL2-40B 61.2%
12.  Claude 3.5 Sonnet 60%
13.  Claude 3.5 Sonnet 60%
14.  GPT-4V 59.9%
15.  Qwen-VL Max 51.3%
16.  4050607080

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## All results

Video-MME results by model
| # | Model | Provider | Score | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | [video-SALMONN 2+](https://noometry.com/models/video-salmonn-2-plus) |  [![](/logos/bytedance.svg) ByteDance Seed](https://noometry.com/providers/bytedance) | 79.7% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 2 | [Gemini 1.5 Pro (May 2024)](https://noometry.com/models/gemini-1-5-pro) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 75% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 3 | [Gemini 1.5 Pro (May 2024)](https://noometry.com/models/gemini-1-5-pro) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 75% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 4 | [Qwen2.5-VL 72B Instruct](https://noometry.com/models/qwen2-5-vl-72b-instruct) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 73.5% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 5 | [InternVL2\_5-78B](https://noometry.com/models/internvl2-5-78b) |  [![](/logos/shanghai-ai-lab.svg) Shanghai AI Lab](https://noometry.com/providers/shanghai-ai-lab) | 72.1% | 5-78B | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 6 | [GPT-4o](https://noometry.com/models/gpt-4o) | [OpenAI](https://noometry.com/providers/openai) | 71.9% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 7 | [GPT-4o](https://noometry.com/models/gpt-4o) | [OpenAI](https://noometry.com/providers/openai) | 71.9% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 8 | [Gemini 1.5 Flash (May 2024)](https://noometry.com/models/gemini-1-5-flash) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 70.3% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 9 | [GPT-4o mini](https://noometry.com/models/gpt-4o-mini) | [OpenAI](https://noometry.com/providers/openai) | 64.8% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 10 | [NVILA 8B](https://noometry.com/models/nvila-8b) |  [![](/logos/nvidia.svg) NVIDIA](https://noometry.com/providers/nvidia) | 64.2% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 11 | [InternVL2-40B](https://noometry.com/models/internvl2-40b) |  [![](/logos/shanghai-ai-lab.svg) Shanghai AI Lab](https://noometry.com/providers/shanghai-ai-lab) | 61.2% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 12 | [Claude 3.5 Sonnet](https://noometry.com/models/claude-3-5-sonnet) | [Anthropic](https://noometry.com/providers/anthropic) | 60% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 13 | [Claude 3.5 Sonnet](https://noometry.com/models/claude-3-5-sonnet) | [Anthropic](https://noometry.com/providers/anthropic) | 60% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 14 | [GPT-4V](https://noometry.com/models/gpt-4v) | [OpenAI](https://noometry.com/providers/openai) | 59.9% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 15 | [Qwen-VL Max](https://noometry.com/models/qwen-vl-max) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 51.3% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

## Compare the leaders

-   [video-SALMONN 2+ vs Gemini 1.5 Pro (May 2024)](https://noometry.com/compare/gemini-1-5-pro-vs-video-salmonn-2-plus)
-   [video-SALMONN 2+ vs Gemini 1.5 Pro (May 2024)](https://noometry.com/compare/gemini-1-5-pro-vs-video-salmonn-2-plus)
-   [video-SALMONN 2+ vs Qwen2.5-VL 72B Instruct](https://noometry.com/compare/qwen2-5-vl-72b-instruct-vs-video-salmonn-2-plus)
-   [video-SALMONN 2+ vs InternVL2\_5-78B](https://noometry.com/compare/internvl2-5-78b-vs-video-salmonn-2-plus)
-   [Gemini 1.5 Pro (May 2024) vs Gemini 1.5 Pro (May 2024)](https://noometry.com/compare/gemini-1-5-pro-vs-gemini-1-5-pro)
-   [Gemini 1.5 Pro (May 2024) vs Qwen2.5-VL 72B Instruct](https://noometry.com/compare/gemini-1-5-pro-vs-qwen2-5-vl-72b-instruct)

## Other multimodal benchmarks

-   [LMArena Vision](https://noometry.com/benchmarks/arena-vision)
-   [GeoBench](https://noometry.com/benchmarks/geobench)
-   [VPCT](https://noometry.com/benchmarks/vpct)
-   [Blueprint-Bench 2](https://noometry.com/benchmarks/blueprint-bench-2)
-   [Furniture Assembly](https://noometry.com/benchmarks/furniture-assembly)
-   [LMArena Document](https://noometry.com/benchmarks/arena-document) (reference)
-   [MindCube](https://noometry.com/benchmarks/mindcube) (reference)
-   [ScienceQA](https://noometry.com/benchmarks/scienceqa) (reference)
-   [SpatialViz-Bench](https://noometry.com/benchmarks/spatialviz-bench) (reference)

## Frequently asked questions

### What does Video-MME measure?

Multiple-choice questions about short, medium and long videos. Scores here are without subtitles.

### Which model has the highest Video-MME score?

As of October 2026, video-SALMONN 2+ has the highest published Video-MME score on Noometry at 79.7%, out of 15 models with results.

### What is the best open-weight model on Video-MME?

Qwen2.5-VL 72B Instruct has the highest Video-MME accuracy among open-weight models at 73.5%, ranking 4 of 15 overall.

### Cite this page

Noometry. (2026). Video-MME leaderboard. Retrieved October 10, 2026, from https://noometry.com/benchmarks/video-mme

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/benchmarks/video-mme.md).
