OpenAI, proprietary

GPT-4o mini

GPT-4o mini by OpenAI ranks 343rd of 354 ranked models on the Noometry Index as of October 2026, with a score of 25.5. Its strongest category is agentic & tool use, where it ranks 101st. API pricing starts at $0.15 per million input tokens and $0.60 per million output tokens, with a 128K-token context window.

Last verified

Specifications

Noometry rank
#343 of 354
Index score
25.5
Evidence
Confirmed 60 results
Provider
OpenAI
Released
July 18, 2024
Weights
Proprietary
Reasoning
No
Context window
128K
Max output
16K
Input price
$0.15 / M
Output price
$0.60 / M
Blended price
$0.26 / M
Output speed
120 tokens/s Kagi
Value
#63 of 219
Knowledge cutoff
September 2023
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

GPT-4o mini category scores
  1. Coding 22.0
  2. Agentic & Tool Use 27.5
  3. Reasoning 8.7
  4. Math 10.4
  5. Knowledge 17.7
  6. Multimodal 25.9
  7. Multilingual 42.0
  8. Instruction Following 61.9
  9. Long Context 39.1
  10. Writing & Preference 39.5
GPT-4o mini category ranks
CategoryScoreRankResults
Coding22.0#3356
Agentic & Tool Use27.5#1011
Reasoning8.7#34710
Math10.4#3146
Knowledge17.7#2846
Multimodal25.9#1224
Multilingual42.0#1991
Instruction Following61.9#2393
Long Context39.1#1861
Writing & Preference39.5#2487

Strengths and weaknesses

Categories where GPT-4o mini places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

GPT-4o mini: strongest categories
CategoryScorevs medianRank
Long Context39.1−1.8#186 of 296, top 63%
Agentic & Tool Use27.5−2.9#101 of 154, top 66%
Multilingual42.0−5.4#199 of 297, top 68%

Weakest categories

GPT-4o mini: weakest categories
CategoryScorevs medianRank
Reasoning8.7−14.9#347 of 350, top 100%
Coding22.0−16.8#335 of 340, top 99%
Math10.4−26.2#314 of 327, top 97%

Closest competitors

The models ranked just above and below GPT-4o mini. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to GPT-4o mini
ModelRankScoreBlended $/MSpeed
DeepSeek-R1-Distill-Qwen-1.5B#33926.1——Compare
Claude 3 Haiku#34025.9—41Compare
Gemma 2 9B#34125.9——Compare
Dolly 2.0-12b#34225.5——Compare
Llama 3-8B#34425.5——Compare
Claude 2.1#34525.2——Compare
Claude 2#34625.0——Compare
DeepSeek LLM 67B#34724.9——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

GPT-4o mini Coding benchmark results
BenchmarkScorePositionSettingSourceDate
Aider Polyglot3.6%#44 of 44, top 100%Epoch AI
WeirdML11.8%#111 of 119, top 94%Epoch AI
BigCodeBench Instruct46.1%#13 of 64, top 21%BigCodeBench2024-07-18
LiveBench Coding43.1%#22 of 39, top 57%Epoch AI
LMArena Coding1290#205 of 294, top 70%LMArena2026-10-08
BigCodeBench Complete57.4%#14 of 66, top 22%BigCodeBench2024-07-18
HumanEval+83.5%#8 of 45, top 18%july 2024EvalPlus
MBPP+72.2%#12 of 38, top 32%july 2024EvalPlus

Agentic & Tool Use

GPT-4o mini Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
BALROG17.4%#27 of 35, top 78%Epoch AI

Reasoning

GPT-4o mini Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-20%#78 of 83, top 94%Epoch AI
SimpleBench10.7%#77 of 77, top 100%Epoch AI
Kagi LLM Benchmark28.8%#94 of 99, top 95%Kagi LLM Benchmark
Chess Puzzles0%#116 of 129, top 90%Epoch AI2026-07-15
LiveBench Reasoning32.8%#29 of 39, top 75%Epoch AI
LMArena Hard Prompts1267#208 of 297, top 71%LMArena2026-10-08
Mystery Game Puzzles12%#57 of 74, top 78%Epoch AI2026-08-27
DTBench54.4%#124 of 151, top 83%Epoch AI
LiveBench Data Analysis50%#23 of 39, top 59%Epoch AI
LMCA10.4%#110 of 125, top 88%Epoch AI
Epoch Capabilities Index126.56#153 of 213, top 72%Epoch AI2024-07-18
LiveBench41.3%#30 of 39, top 77%Epoch AI
PIQA88.7%Best of 27Epoch AI

Math

GPT-4o mini Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)0.7%#79 of 81, top 98%Epoch AI2026-08-27
OTIS Mock AIME 2024-20256.9%#141 of 173, top 82%Epoch AI2025-07-30
Omni-MATH28%#42 of 57, top 74%HELM Capabilities
LiveBench Math36.3%#30 of 39, top 77%Epoch AI
LMArena Math1267#205 of 285, top 72%LMArena2026-10-08
MATH Level 552.6%#44 of 79, top 56%Epoch AI2025-01-27
GSM8K91.3%#4 of 38, top 11%Epoch AI

Knowledge

GPT-4o mini Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond37.7%#154 of 186, top 83%Epoch AI2025-01-27
SimpleQA Verified8.3%#76 of 77, top 99%Epoch AI2026-08-31
MMLU-Pro60.3%#43 of 58, top 75%HELM Capabilities
Confabulations (lower is better)37.2%#50 of 51, top 99%Lech Mazur benchmarks
GPQA (HELM)36.8%#47 of 57, top 83%HELM Capabilities
LMArena Expert1235#206 of 273, top 76%LMArena2026-10-08
BoolQ88.7%#2 of 23, top 9%Epoch AI
MMLU81.8%#13 of 81, top 17%Epoch AI

Multimodal

GPT-4o mini Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1066#112 of 122, top 92%LMArena2026-10-09
Video-MME64.8%#9 of 15, top 60%Epoch AI
GeoBench64%#14 of 25, top 57%Epoch AI
VPCT34%#21 of 24, top 88%Epoch AI

Multilingual

GPT-4o mini Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1266#199 of 297, top 68%LMArena2026-10-08
LMArena Chinese1265#201 of 285, top 71%LMArena2026-10-08
LMArena French1297#163 of 223, top 74%LMArena2026-10-08
LMArena German1272#164 of 231, top 71%LMArena2026-10-08
LMArena Japanese1216#150 of 211, top 72%LMArena2026-10-08
LMArena Korean1195#163 of 213, top 77%LMArena2026-10-08
LMArena Russian1275#193 of 283, top 69%LMArena2026-10-08
LMArena Spanish1276#171 of 226, top 76%LMArena2026-10-08

Instruction Following

GPT-4o mini Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following56.8%#32 of 39, top 83%Epoch AI
IFEval78.2%#46 of 57, top 81%HELM Capabilities
LMArena Instruction Following1258#204 of 298, top 69%LMArena2026-10-08

Long Context

GPT-4o mini Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1289#197 of 291, top 68%LMArena2026-10-08

Writing & Preference

GPT-4o mini Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1286#203 of 297, top 69%LMArena2026-10-08
LMArena Creative Writing1268#193 of 295, top 66%LMArena2026-10-08
Short-Story Creative Writing67.2%#33 of 39, top 85%Epoch AI
EQ-Bench Creative Writing873#100 of 115, top 87%EQ-Bench
WildBench79.1%#35 of 57, top 62%HELM Capabilities
LMArena Multi-Turn1285#197 of 295, top 67%LMArena2026-10-08
LiveBench Language28.6%#28 of 39, top 72%Epoch AI

API pricing by provider

GPT-4o mini API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$0.15$0.60$0.0752026-10-10
openai$0.15$0.60$0.0752026-10-10
openrouter$0.15$0.60$0.0752026-10-10

Compare GPT-4o mini

Other OpenAI models

Frequently asked questions

How good is GPT-4o mini?

GPT-4o mini by OpenAI ranks 343rd of 354 ranked models on the Noometry Index as of October 2026, with a score of 25.5. Its strongest category is agentic & tool use, where it ranks 101st. API pricing starts at $0.15 per million input tokens and $0.60 per million output tokens, with a 128K-token context window.

How much does GPT-4o mini cost?

GPT-4o mini costs $0.15 per million input tokens and $0.60 per million output tokens on OpenAI's own API, with cached input at $0.075.

What is GPT-4o mini's context window?

GPT-4o mini accepts up to 128K tokens of input and can write up to 16K tokens in one response.

Is GPT-4o mini open source?

No. GPT-4o mini is proprietary and available only through OpenAI's API and partner platforms.

How fast is GPT-4o mini?

GPT-4o mini generated about 120 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are GPT-4o mini's strengths and weaknesses?

Relative to other ranked models, GPT-4o mini places best in long context, agentic & tool use, multilingual and lowest in reasoning, coding, math.

What is GPT-4o mini best at?

Its best category is agentic & tool use, where it ranks 101st on Noometry.