Google, proprietary

Gemini 2.5 Pro

Gemini 2.5 Pro by Google ranks 75th of 354 ranked models on the Noometry Index as of October 2026, with a score of 45.0. Its strongest category is long context, where it ranks 5th. API pricing starts at $1.25 per million input tokens and $10 per million output tokens, with a 1.05M-token context window.

Last verified

Specifications

Noometry rank
#75 of 354
Index score
45.0
Evidence
Confirmed 78 results
Provider
Google
Released
March 25, 2025
Weights
Proprietary
Reasoning
Yes
Context window
1.05M
Max output
66K
Input price
$1.25 / M
Output price
$10 / M
Blended price
$3.44 / M
Output speed
5 tokens/s Kagi
Value
#171 of 219
Knowledge cutoff
January 2025
Input
text, image, audio, video, pdf

Category scores

Each category score combines every public result we have in that category.

Gemini 2.5 Pro category scores
  1. Coding 42.4
  2. Agentic & Tool Use 29.2
  3. Reasoning 28.8
  4. Math 32.5
  5. Knowledge 56.0
  6. Multimodal 45.2
  7. Multilingual 55.3
  8. Instruction Following 75.0
  9. Long Context 59.8
  10. Writing & Preference 63.7
Gemini 2.5 Pro category ranks
CategoryScoreRankResults
Coding42.4#10110
Agentic & Tool Use29.2#887
Reasoning28.8#9912
Math32.5#2137
Knowledge56.0#467
Multimodal45.2#183
Multilingual55.3#311
Instruction Following75.0#753
Long Context59.8#52
Writing & Preference63.7#627

Strengths and weaknesses

Categories where Gemini 2.5 Pro places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Gemini 2.5 Pro: strongest categories
CategoryScorevs medianRank
Long Context59.8+18.8#5 of 296, top 2%
Multilingual55.3+7.9#31 of 297, top 11%
Multimodal45.2+6.6#18 of 128, top 15%

Weakest categories

Gemini 2.5 Pro: weakest categories
CategoryScorevs medianRank
Math32.5−4.1#213 of 327, top 66%
Agentic & Tool Use29.2−1.2#88 of 154, top 58%
Coding42.4+3.7#101 of 340, top 30%

Closest competitors

The models ranked just above and below Gemini 2.5 Pro. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Gemini 2.5 Pro
ModelRankScoreBlended $/MSpeed
Qwen3.5 Max Preview#7145.3——Compare
Qwen3.7 Plus#7245.3$0.70—Compare
Hy4 preview#7345.3$1.13—Compare
MiMo-V2.5-Pro#7445.2$0.54—Compare
GPT-5.4 mini#7645.0$1.6910Compare
Amazon Nova Experimental Chat 26 02 10#7744.5——Compare
DeepSeek-V3.2-Exp#7844.3$0.2916Compare
Hy3#7944.2$0.14—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Gemini 2.5 Pro Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified57.6%#30 of 32, top 94%Epoch AI2026-02-13
SWE-bench Verified (bash only)53.6%#27 of 39, top 70%SWE-bench2025-07-26
Aider Polyglot72.9%Epoch AI
Aider Polyglot76.9%Epoch AI
Aider Polyglot79.1%Epoch AI
Aider Polyglot72.9%Epoch AI
Aider Polyglot83.1%#3 of 44, top 7%32KEpoch AI
LMArena WebDev1227#107 of 113, top 95%LMArena2026-10-08
SciCode42.8%#68 of 121, top 57%Epoch AI
GSO3.9%#25 of 31, top 81%Epoch AI
GSO3.9%#25 of 31, top 81%Epoch AI
WeirdML54%#41 of 119, top 35%16KEpoch AI
LiveBench Coding85.9%Best of 39Epoch AI
LMArena Coding1452#85 of 294, top 29%LMArena2026-10-08
CadEval64%#2 of 14, top 15%Epoch AI
ALE-Bench785.52#61 of 105, top 59%32KEpoch AI
AlgoTune1.51#11 of 18, top 62%Epoch AI

Agentic & Tool Use

Gemini 2.5 Pro Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Terminal-Bench32.6%#31 of 41, top 76%Epoch AI
GDPval23.3%#9 of 11, top 82%Epoch AI
Remote Labor Index0.8%#14 of 14, top 100%Epoch AI
TheAgentCompany30.3%#5 of 14, top 36%Epoch AI
τ²-bench Banking13.7%#23 of 26, top 89%highτ²-bench2026-05-05
DeepResearch Bench41.5%Epoch AI
DeepResearch Bench42.8%#18 of 24, top 75%Epoch AI
BALROG43.3%#11 of 35, top 32%Epoch AI
LMArena Search1142#26 of 32, top 82%LMArena2026-08-24
METR Time Horizons55.4%#21 of 32, top 66%Epoch AI
Vending-Bench 2573.64#48 of 60, top 80%Epoch AI

Reasoning

Gemini 2.5 Pro Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-24%16KEpoch AI
ARC-AGI-20%1KEpoch AI
ARC-AGI-24.9%#56 of 83, top 68%32KEpoch AI
ARC-AGI-22.9%8KEpoch AI
SimpleBench62.4%#17 of 77, top 23%Epoch AI
SimpleBench51.6%Epoch AI
SimpleBench51.6%Epoch AI
Kagi LLM Benchmark70.3%#22 of 99, top 23%Kagi LLM Benchmark
ARC-AGI-133%Epoch AI
ARC-AGI-141%#60 of 83, top 73%16KEpoch AI
ARC-AGI-131.3%1KEpoch AI
ARC-AGI-137%32KEpoch AI
ARC-AGI-129.5%8KEpoch AI
CritPt2%#72 of 134, top 54%Epoch AI
Chess Puzzles20%#57 of 129, top 45%Epoch AI2025-12-08
EnigmaEval2.4%Epoch AI
EnigmaEval4.1%Epoch AI
EnigmaEval5.6%#23 of 38, top 61%Epoch AI
LiveBench Reasoning89.8%#3 of 39, top 8%Epoch AI
LMArena Hard Prompts1455#54 of 297, top 19%LMArena2026-10-08
DTBench82.4%#64 of 151, top 43%Epoch AI
LiveBench Data Analysis79.9%Best of 39Epoch AI
LMCA34.8%#63 of 125, top 51%Epoch AI
Epoch Capabilities Index145.32#82 of 213, top 39%Epoch AI2025-06-05
Epoch Capabilities Index144.16Epoch AI2025-03-31
Epoch Capabilities Index142.47Epoch AI2025-05-06
ForecastBench61.3#13 of 72, top 19%Epoch AI
ForecastBench60.4Epoch AI
LiveBench82.3%Best of 39Epoch AI

Math

Gemini 2.5 Pro Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)24.6%#65 of 81, top 81%Epoch AI2026-06-11
FrontierMath Tier 40%#61 of 63, top 97%Epoch AI2026-06-11
OTIS Mock AIME 2024-202584.7%#70 of 173, top 41%Epoch AI2025-11-16
Omni-MATH41.6%#25 of 57, top 44%HELM Capabilities
LiveBench Math90.2%#2 of 39, top 6%Epoch AI
LMArena Math1450#62 of 285, top 22%LMArena2026-10-08
MATH Level 595.9%#10 of 79, top 13%Epoch AI2025-05-08
MATH Level 595.6%Epoch AI2025-05-07
FrontierMath (Feb 2025 set)10.3%Epoch AI2025-07-03
FrontierMath (Feb 2025 set)14.1%#35 of 68, top 52%Epoch AI2025-11-24
FrontierMath Tier 4 (v1)4.2%#32 of 55, top 59%Epoch AI2025-07-03
FrontierMath Tier 4 (v1)2.1%Epoch AI2025-07-03

Knowledge

Gemini 2.5 Pro Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond85.3%#64 of 186, top 35%Epoch AI2025-11-16
GPQA Diamond83.8%Epoch AI2025-03-31
GPQA Diamond66.7%Epoch AI2025-06-03
GPQA Diamond84.8%Epoch AI2025-06-05
Humanity's Last Exam18.2%Epoch AI
Humanity's Last Exam17.8%Epoch AI
Humanity's Last Exam21.6%#17 of 41, top 42%Epoch AI
MMLU-Pro86.3%#4 of 58, top 7%HELM Capabilities
Confabulations (lower is better)10.6%#2 of 51, top 4%Lech Mazur benchmarks
Confabulations (lower is better)10.8%Lech Mazur benchmarks
Confabulations (lower is better)12.4%Lech Mazur benchmarks
Vectara Hallucination Rate (lower is better)7%#28 of 96, top 30%Vectara Hallucination Leaderboard
GPQA (HELM)74.9%#5 of 57, top 9%HELM Capabilities
LMArena Expert1452#69 of 273, top 26%LMArena2026-10-08

Multimodal

Gemini 2.5 Pro Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1263#43 of 122, top 36%LMArena2026-10-09
GeoBench86%#2 of 25, top 8%Epoch AI
GeoBench81%Epoch AI
VPCT48%#8 of 24, top 34%Epoch AI
VPCT46.4%Epoch AI
VPCT40.5%Epoch AI
LMArena Document1421#31 of 38, top 82%LMArena2026-09-13
SpatialViz-Bench44.7%Best of 8Epoch AI

Multilingual

Gemini 2.5 Pro Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1451#32 of 297, top 11%LMArena2026-10-08
LMArena Chinese1507#40 of 285, top 15%LMArena2026-10-08
LMArena French1472#36 of 223, top 17%LMArena2026-10-08
LMArena German1487#16 of 231, top 7%LMArena2026-10-08
LMArena Japanese1461#20 of 211, top 10%LMArena2026-10-08
LMArena Korean1434#27 of 213, top 13%LMArena2026-10-08
LMArena Russian1461#32 of 283, top 12%LMArena2026-10-08
LMArena Spanish1473#17 of 226, top 8%LMArena2026-10-08

Instruction Following

Gemini 2.5 Pro Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following80.6%#10 of 39, top 26%Epoch AI
IFEval84%#24 of 57, top 43%HELM Capabilities
LMArena Instruction Following1437#54 of 298, top 19%LMArena2026-10-08

Long Context

Gemini 2.5 Pro Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench66.7%Epoch AI
Fiction.LiveBench91.7%#5 of 47, top 11%Epoch AI
Fiction.LiveBench66.7%Epoch AI
Fiction.LiveBench66.7%Epoch AI
LMArena Longer Query1449#54 of 291, top 19%LMArena2026-10-08

Writing & Preference

Gemini 2.5 Pro Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1458#36 of 297, top 13%LMArena2026-10-08
LMArena Creative Writing1454#26 of 295, top 9%LMArena2026-10-08
Short-Story Creative Writing80.9%Epoch AI
Short-Story Creative Writing80.5%Epoch AI
Short-Story Creative Writing83.8%#6 of 39, top 16%Epoch AI
EQ-Bench Creative Writing1421#65 of 115, top 57%EQ-Bench
EQ-Bench Creative Writing1396EQ-Bench
WildBench85.7%#7 of 57, top 13%HELM Capabilities
LMArena Multi-Turn1453#50 of 295, top 17%LMArena2026-10-08
LiveBench Language67.8%#2 of 39, top 6%Epoch AI

API pricing by provider

Gemini 2.5 Pro API prices
RouteInput $/MOutput $/MCached input $/MChecked
google$1.25$10$0.132026-10-10
openrouter$1.25$10$0.132026-10-10
vertex$1.25$10$0.132026-10-10

Compare Gemini 2.5 Pro

Other Google models

Frequently asked questions

How good is Gemini 2.5 Pro?

Gemini 2.5 Pro by Google ranks 75th of 354 ranked models on the Noometry Index as of October 2026, with a score of 45.0. Its strongest category is long context, where it ranks 5th. API pricing starts at $1.25 per million input tokens and $10 per million output tokens, with a 1.05M-token context window.

How much does Gemini 2.5 Pro cost?

Gemini 2.5 Pro costs $1.25 per million input tokens and $10 per million output tokens on Google's own API, with cached input at $0.13.

What is Gemini 2.5 Pro's context window?

Gemini 2.5 Pro accepts up to 1.05M tokens of input and can write up to 66K tokens in one response.

Is Gemini 2.5 Pro open source?

No. Gemini 2.5 Pro is proprietary and available only through Google's API and partner platforms.

How fast is Gemini 2.5 Pro?

Gemini 2.5 Pro generated about 5 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Gemini 2.5 Pro's strengths and weaknesses?

Relative to other ranked models, Gemini 2.5 Pro places best in long context, multilingual, multimodal and lowest in math, agentic & tool use, coding.

What is Gemini 2.5 Pro best at?

Its best category is long context, where it ranks 5th on Noometry.