Google, proprietary
Gemini 1.5 Flash (May 2024)
Gemini 1.5 Flash (May 2024) by Google ranks 246th of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.2. Its strongest category is multimodal, where it ranks 81st.
Last verified
Specifications
- Noometry rank
- #246 of 354
- Index score
- 33.2
- Evidence
- Confirmed 42 results
- Provider
Google
- Released
- May 14, 2024
- Weights
- Proprietary
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- Not measured
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Coding 34.4
- Agentic & Tool Use 26.6
- Reasoning 21.7
- Math 22.1
- Knowledge 26.2
- Multimodal 36.0
- Multilingual 42.9
- Instruction Following 66.8
- Long Context 39.0
- Writing & Preference 48.7
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 34.4 | #236 | 4 |
| Agentic & Tool Use | 26.6 | #102 | 1 |
| Reasoning | 21.7 | #215 | 2 |
| Math | 22.1 | #281 | 4 |
| Knowledge | 26.2 | #260 | 4 |
| Multimodal | 36.0 | #81 | 3 |
| Multilingual | 42.9 | #189 | 1 |
| Instruction Following | 66.8 | #205 | 2 |
| Long Context | 39.0 | #187 | 1 |
| Writing & Preference | 48.7 | #196 | 4 |
Strengths and weaknesses
Categories where Gemini 1.5 Flash (May 2024) places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Reasoning | 21.7 | −1.9 | #215 of 350, top 62% |
| Writing & Preference | 48.7 | −5.1 | #196 of 312, top 63% |
| Long Context | 39.0 | −1.9 | #187 of 296, top 64% |
Closest competitors
The models ranked just above and below Gemini 1.5 Flash (May 2024). When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Mercury 2.5 | #242 | 33.5 | $0.0675 | — | Compare |
| Mistral Small | #243 | 33.4 | $0.26 | 120 | Compare |
| Nova 2.0 Pro Preview | #244 | 33.4 | — | — | Compare |
| Qwen2.5-Coder-32B | #245 | 33.4 | $0.74 | — | Compare |
| Granite 3.1 2b Instruct | #247 | 33.2 | — | — | Compare |
| Gemma 2 2b IT | #248 | 33.1 | — | — | Compare |
| Wizardlm 70b | #249 | 33.0 | — | — | Compare |
| Phi 3 Medium 4k Instruct | #250 | 33.0 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| WeirdML | 24.9% | #100 of 119, top 85% | Epoch AI | ||
| BigCodeBench Instruct | 43.5% | #26 of 64, top 41% | BigCodeBench | 2024-05-14 | |
| LMArena Coding | 1236 | LMArena | 2026-10-08 | ||
| LMArena Coding | 1261 | #221 of 294, top 76% | LMArena | 2026-10-08 | |
| BigCodeBench Complete | 55.1% | #18 of 66, top 28% | BigCodeBench | 2024-05-14 | |
| HumanEval+ | 75.6% | #14 of 45, top 32% | EvalPlus | ||
| MBPP+ | 67.5% | #18 of 38, top 48% | EvalPlus |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| BALROG | 14.6% | #31 of 35, top 89% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Hard Prompts | 1220 | LMArena | 2026-10-08 | ||
| LMArena Hard Prompts | 1257 | #213 of 297, top 72% | LMArena | 2026-10-08 | |
| DTBench | 53.8% | #126 of 151, top 84% | Epoch AI | ||
| DTBench | 51.4% | Epoch AI | |||
| Epoch Capabilities Index | 122.59 | Epoch AI | 2024-05-23 | ||
| Epoch Capabilities Index | 129.36 | #141 of 213, top 67% | Epoch AI | 2024-09-24 | |
| ForecastBench | 53.9 | #69 of 72, top 96% | Epoch AI | ||
| PIQA | 87.5% | #3 of 27, top 12% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 3.9% | Epoch AI | 2025-02-25 | ||
| OTIS Mock AIME 2024-2025 | 16.3% | #129 of 173, top 75% | Epoch AI | 2025-02-25 | |
| Omni-MATH | 30.4% | #37 of 57, top 65% | HELM Capabilities | ||
| LMArena Math | 1229 | LMArena | 2026-10-08 | ||
| LMArena Math | 1269 | #203 of 285, top 72% | LMArena | 2026-10-08 | |
| MATH Level 5 | 61.9% | #39 of 79, top 50% | Epoch AI | 2025-01-27 | |
| MATH Level 5 | 25.1% | Epoch AI | 2025-01-27 | ||
| FrontierMath (Feb 2025 set) | 0% | #67 of 68, top 99% | Epoch AI | 2025-03-07 | |
| GSM8K | 82.4% | #11 of 38, top 29% | Epoch AI |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 40.4% | Epoch AI | 2025-01-27 | ||
| GPQA Diamond | 47.3% | #136 of 186, top 74% | Epoch AI | 2025-01-27 | |
| MMLU-Pro | 67.8% | #36 of 58, top 63% | HELM Capabilities | ||
| GPQA (HELM) | 43.7% | #38 of 57, top 67% | HELM Capabilities | ||
| LMArena Expert | 1199 | LMArena | 2026-10-08 | ||
| LMArena Expert | 1233 | #207 of 273, top 76% | LMArena | 2026-10-08 | |
| BoolQ | 85.8% | #7 of 23, top 31% | Epoch AI | ||
| MMLU | 77.9% | #26 of 81, top 33% | Epoch AI | ||
| MMLU | 73.9% | Epoch AI | |||
| MMLU | 77.8% | Epoch AI |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1141 | #101 of 122, top 83% | LMArena | 2026-10-09 | |
| LMArena Vision | 1008 | LMArena | 2026-10-09 | ||
| Video-MME | 70.3% | #8 of 15, top 54% | Epoch AI | ||
| GeoBench | 76% | #7 of 25, top 29% | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1231 | LMArena | 2026-10-08 | ||
| LMArena Non-English | 1278 | #189 of 297, top 64% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1295 | #190 of 285, top 67% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1234 | LMArena | 2026-10-08 | ||
| LMArena French | 1241 | LMArena | 2026-10-08 | ||
| LMArena French | 1258 | #182 of 223, top 82% | LMArena | 2026-10-08 | |
| LMArena German | 1262 | #169 of 231, top 74% | LMArena | 2026-10-08 | |
| LMArena German | 1216 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1186 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1252 | #139 of 211, top 66% | LMArena | 2026-10-08 | |
| LMArena Korean | 1201 | LMArena | 2026-10-08 | ||
| LMArena Korean | 1221 | #154 of 213, top 73% | LMArena | 2026-10-08 | |
| LMArena Russian | 1239 | LMArena | 2026-10-08 | ||
| LMArena Russian | 1288 | #183 of 283, top 65% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1243 | #185 of 226, top 82% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1224 | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 83.1% | #31 of 57, top 55% | HELM Capabilities | ||
| LMArena Instruction Following | 1212 | LMArena | 2026-10-08 | ||
| LMArena Instruction Following | 1258 | #205 of 298, top 69% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1284 | #200 of 291, top 69% | LMArena | 2026-10-08 | |
| LMArena Longer Query | 1252 | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1287 | #202 of 297, top 69% | LMArena | 2026-10-08 | |
| LMArena Text | 1239 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1222 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1285 | #182 of 295, top 62% | LMArena | 2026-10-08 | |
| WildBench | 79.2% | #34 of 57, top 60% | HELM Capabilities | ||
| LMArena Multi-Turn | 1253 | #222 of 295, top 76% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1228 | LMArena | 2026-10-08 |
Compare Gemini 1.5 Flash (May 2024)
- Gemini 1.5 Flash (May 2024) vs Qwen2.5-Coder-32B
- Gemini 1.5 Flash (May 2024) vs Granite 3.1 2b Instruct
- Gemini 1.5 Flash (May 2024) vs Nova 2.0 Pro Preview
- Gemini 1.5 Flash (May 2024) vs Gemma 2 2b IT
- Gemini 1.5 Flash (May 2024) vs Mistral Small
- Gemini 1.5 Flash (May 2024) vs Wizardlm 70b
- Gemini 1.5 Flash (May 2024) vs GPT-6 Astra
- Gemini 1.5 Flash (May 2024) vs Claude Fable 5.1
- Gemini 1.5 Flash (May 2024) vs Kimi K3
- Gemini 1.5 Flash (May 2024) vs Grok 4.6
- Gemini 1.5 Flash (May 2024) vs Qwen3.8 Max
- Gemini 1.5 Flash (May 2024) vs GLM-5.3
- Gemini 1.5 Flash (May 2024) vs Muse Spark 1.3
- Gemini 1.5 Flash (May 2024) vs DeepSeek V4 Pro
Other Google models
Frequently asked questions
How good is Gemini 1.5 Flash (May 2024)?
Gemini 1.5 Flash (May 2024) by Google ranks 246th of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.2. Its strongest category is multimodal, where it ranks 81st.
Is Gemini 1.5 Flash (May 2024) open source?
No. Gemini 1.5 Flash (May 2024) is proprietary and available only through Google's API and partner platforms.
What are Gemini 1.5 Flash (May 2024)'s strengths and weaknesses?
Relative to other ranked models, Gemini 1.5 Flash (May 2024) places best in reasoning, writing & preference, long context and lowest in math, knowledge, coding.
What is Gemini 1.5 Flash (May 2024) best at?
Its best category is multimodal, where it ranks 81st on Noometry.