Google, proprietary
Gemini 2.5 Flash
Gemini 2.5 Flash by Google ranks 170th of 354 ranked models on the Noometry Index as of October 2026, with a score of 39.3. Its strongest category is long context, where it ranks 17th. API pricing starts at $0.30 per million input tokens and $2.50 per million output tokens, with a 1.05M-token context window.
Last verified
Specifications
- Noometry rank
- #170 of 354
- Index score
- 39.3
- Evidence
- Confirmed 54 results
- Provider
Google
- Released
- April 17, 2025
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 1.05M
- Max output
- 66K
- Input price
- $0.30 / M
- Output price
- $2.50 / M
- Blended price
- $0.85 / M
- Output speed
- 152 tokens/s Kagi
- Value
- #103 of 219
- Knowledge cutoff
- January 2025
- Input
- text, image, audio, video, pdf
Category scores
Each category score combines every public result we have in that category.
- Coding 35.8
- Agentic & Tool Use 30.8
- Reasoning 18.1
- Math 39.9
- Knowledge 36.4
- Multimodal 41.8
- Multilingual 52.3
- Instruction Following 75.7
- Long Context 47.5
- Writing & Preference 53.8
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 35.8 | #220 | 4 |
| Agentic & Tool Use | 30.8 | #74 | 4 |
| Reasoning | 18.1 | #286 | 9 |
| Math | 39.9 | #98 | 3 |
| Knowledge | 36.4 | #168 | 6 |
| Multimodal | 41.8 | #32 | 3 |
| Multilingual | 52.3 | #88 | 1 |
| Instruction Following | 75.7 | #54 | 2 |
| Long Context | 47.5 | #17 | 2 |
| Writing & Preference | 53.8 | #157 | 6 |
Strengths and weaknesses
Categories where Gemini 2.5 Flash places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 47.5 | +6.6 | #17 of 296, top 6% |
| Instruction Following | 75.7 | +4.5 | #54 of 305, top 18% |
| Multimodal | 41.8 | +3.3 | #32 of 128, top 25% |
Closest competitors
The models ranked just above and below Gemini 2.5 Flash. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| DeepSeek-V3 | #166 | 39.5 | $0.41 | 73 | Compare |
| Grok 4 Fast | #167 | 39.4 | — | 577 | Compare |
| Olmo 3.1 32b Instruct | #168 | 39.4 | — | — | Compare |
| Granite 4.2 3b | #169 | 39.4 | — | — | Compare |
| Step 2 16k Exp 202412 | #171 | 39.2 | — | — | Compare |
| Qwen3 32B | #172 | 39.2 | $1.22 | 86 | Compare |
| Gemini 2.0 Pro | #173 | 39.1 | — | — | Compare |
| Molmo 2 8b | #174 | 39.1 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified (bash only) | 28.7% | #32 of 39, top 83% | SWE-bench | 2025-07-26 | |
| Aider Polyglot | 44% | Epoch AI | |||
| Aider Polyglot | 47.1% | Epoch AI | |||
| Aider Polyglot | 55.1% | #19 of 44, top 44% | 23K | Epoch AI | |
| WeirdML | 41% | Epoch AI | |||
| WeirdML | 41% | 16K | Epoch AI | ||
| WeirdML | 41.9% | #71 of 119, top 60% | 16k | Epoch AI | |
| LMArena Coding | 1424 | #116 of 294, top 40% | LMArena | 2026-10-08 | |
| LMArena Coding | 1400 | LMArena | 2026-10-08 | ||
| ALE-Bench | 661.88 | #73 of 105, top 70% | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Terminal-Bench | 17.1% | Epoch AI | |||
| Terminal-Bench | 17.1% | #39 of 41, top 96% | Epoch AI | ||
| Berkeley Function Calling Leaderboard | 56.2% | #12 of 49, top 25% | fc | Berkeley Function Calling Leaderboard | |
| TheAgentCompany | 41.1% | #2 of 14, top 15% | Epoch AI | ||
| BALROG | 33.5% | #13 of 35, top 38% | Epoch AI | ||
| Vending-Bench 2 | 548.84 | #49 of 60, top 82% | Epoch AI |
Reasoning
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 73.1% | #88 of 173, top 51% | Epoch AI | 2025-04-22 | |
| OTIS Mock AIME 2024-2025 | 70.8% | Epoch AI | 2025-05-20 | ||
| Omni-MATH | 38.5% | #28 of 57, top 50% | HELM Capabilities | ||
| LMArena Math | 1415 | #109 of 285, top 39% | LMArena | 2026-10-08 | |
| LMArena Math | 1411 | LMArena | 2026-10-08 | ||
| FrontierMath (Feb 2025 set) | 4.8% | #46 of 68, top 68% | Epoch AI | 2025-12-18 | |
| FrontierMath Tier 4 (v1) | 4.2% | #31 of 55, top 57% | Epoch AI | 2025-12-18 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Humanity's Last Exam | 12.1% | #22 of 41, top 54% | Epoch AI | ||
| Humanity's Last Exam | 11% | Epoch AI | |||
| MMLU-Pro | 63.9% | #39 of 58, top 68% | HELM Capabilities | ||
| Confabulations (lower is better) | 16.8% | #25 of 51, top 50% | Lech Mazur benchmarks | ||
| Vectara Hallucination Rate (lower is better) | 7.8% | #35 of 96, top 37% | Vectara Hallucination Leaderboard | ||
| GPQA (HELM) | 39% | #44 of 57, top 78% | HELM Capabilities | ||
| LMArena Expert | 1418 | LMArena | 2026-10-08 | ||
| LMArena Expert | 1426 | #101 of 273, top 37% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1235 | LMArena | 2026-10-09 | ||
| LMArena Vision | 1253 | #53 of 122, top 44% | LMArena | 2026-10-09 | |
| GeoBench | 73% | Epoch AI | |||
| GeoBench | 76% | #8 of 25, top 32% | Epoch AI | ||
| VPCT | 38% | Epoch AI | |||
| VPCT | 46.2% | #9 of 24, top 38% | Epoch AI | ||
| SpatialViz-Bench | 36.9% | #3 of 8, top 38% | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1396 | LMArena | 2026-10-08 | ||
| LMArena Non-English | 1409 | #89 of 297, top 30% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1449 | LMArena | 2026-10-08 | ||
| LMArena Chinese | 1450 | #101 of 285, top 36% | LMArena | 2026-10-08 | |
| LMArena French | 1433 | #88 of 223, top 40% | LMArena | 2026-10-08 | |
| LMArena French | 1432 | LMArena | 2026-10-08 | ||
| LMArena German | 1417 | LMArena | 2026-10-08 | ||
| LMArena German | 1418 | #79 of 231, top 35% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1396 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1405 | #58 of 211, top 28% | LMArena | 2026-10-08 | |
| LMArena Korean | 1385 | LMArena | 2026-10-08 | ||
| LMArena Korean | 1385 | #70 of 213, top 33% | LMArena | 2026-10-08 | |
| LMArena Russian | 1415 | #87 of 283, top 31% | LMArena | 2026-10-08 | |
| LMArena Russian | 1393 | LMArena | 2026-10-08 | ||
| LMArena Spanish | 1421 | #88 of 226, top 39% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1405 | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 89.8% | #10 of 57, top 18% | HELM Capabilities | ||
| LMArena Instruction Following | 1405 | #95 of 298, top 32% | LMArena | 2026-10-08 | |
| LMArena Instruction Following | 1394 | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 47.2% | Epoch AI | |||
| Fiction.LiveBench | 77.8% | #12 of 47, top 26% | Epoch AI | ||
| LMArena Longer Query | 1419 | #93 of 291, top 32% | LMArena | 2026-10-08 | |
| LMArena Longer Query | 1406 | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1417 | #102 of 297, top 35% | LMArena | 2026-10-08 | |
| LMArena Text | 1406 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1387 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1400 | #82 of 295, top 28% | LMArena | 2026-10-08 | |
| Short-Story Creative Writing | 76.5% | #21 of 39, top 54% | Epoch AI | ||
| EQ-Bench Creative Writing | 1137 | #90 of 115, top 79% | EQ-Bench | ||
| WildBench | 81.7% | #22 of 57, top 39% | HELM Capabilities | ||
| LMArena Multi-Turn | 1397 | LMArena | 2026-10-08 | ||
| LMArena Multi-Turn | 1408 | #113 of 295, top 39% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| $0.30 | $2.50 | $0.03 | 2026-10-10 | |
| openrouter | $0.30 | $2.50 | $0.03 | 2026-10-10 |
| vertex | $0.30 | $2.50 | $0.03 | 2026-10-10 |
Compare Gemini 2.5 Flash
- Gemini 2.5 Flash vs Gemini 1.5 Flash 8B
- Gemini 2.5 Flash vs Granite 4.2 3b
- Gemini 2.5 Flash vs Step 2 16k Exp 202412
- Gemini 2.5 Flash vs Olmo 3.1 32b Instruct
- Gemini 2.5 Flash vs Qwen3 32B
- Gemini 2.5 Flash vs Grok 4 Fast
- Gemini 2.5 Flash vs Gemini 2.0 Pro
- Gemini 2.5 Flash vs GPT-6 Astra
- Gemini 2.5 Flash vs Claude Fable 5.1
- Gemini 2.5 Flash vs Kimi K3
- Gemini 2.5 Flash vs Grok 4.6
- Gemini 2.5 Flash vs Qwen3.8 Max
- Gemini 2.5 Flash vs GLM-5.3
- Gemini 2.5 Flash vs Muse Spark 1.3
Other Google models
Frequently asked questions
How good is Gemini 2.5 Flash?
Gemini 2.5 Flash by Google ranks 170th of 354 ranked models on the Noometry Index as of October 2026, with a score of 39.3. Its strongest category is long context, where it ranks 17th. API pricing starts at $0.30 per million input tokens and $2.50 per million output tokens, with a 1.05M-token context window.
How much does Gemini 2.5 Flash cost?
Gemini 2.5 Flash costs $0.30 per million input tokens and $2.50 per million output tokens on Google's own API, with cached input at $0.03.
What is Gemini 2.5 Flash's context window?
Gemini 2.5 Flash accepts up to 1.05M tokens of input and can write up to 66K tokens in one response.
Is Gemini 2.5 Flash open source?
No. Gemini 2.5 Flash is proprietary and available only through Google's API and partner platforms.
How fast is Gemini 2.5 Flash?
Gemini 2.5 Flash generated about 152 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Gemini 2.5 Flash's strengths and weaknesses?
Relative to other ranked models, Gemini 2.5 Flash places best in long context, instruction following, multimodal and lowest in reasoning, coding, knowledge.
What is Gemini 2.5 Flash best at?
Its best category is long context, where it ranks 17th on Noometry.