Model comparison

GPT-5 Nano vs Mercury

Mercury is the stronger model overall, scoring 37.6 to 33.5 on the Noometry Index.

Last verified . 9 shared benchmarks.

GPT-5 Nano OpenAI

33.5

Rank #241 Confirmed

Mercury Inception

37.6

Rank #199 Confirmed

Summary

  • They share 9 benchmarks with published results for both. GPT-5 Nano scores higher in 2 categories and Mercury in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where GPT-5 Nano leads 75.0 to 65.2.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 62.2% for GPT-5 Nano and 21.6% for Mercury.

Side by side

GPT-5 Nano and Mercury specifications
GPT-5 NanoMercury
ProviderOpenAIInception
Noometry Index33.537.6
Released2025-08-07—
WeightsProprietaryProprietary
Context window400K—
Max output128K—
Input $ / M tokens$0.05—
Output $ / M tokens$0.40—
Results tracked499

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury leads

GPT-5 Nano: 33.6 (#254), Mercury: 38.7 (#170)

Coding benchmarks
BenchmarkGPT-5 NanoMercury
LMArena Coding13511322
SWE-bench Verified (bash only)34.8%—
WeirdML38.1%—
ALE-Bench718.67—

Agentic & Tool Use Not comparable

GPT-5 Nano: 25.8 (#106), Mercury: —

Agentic & Tool Use benchmarks
BenchmarkGPT-5 NanoMercury
Terminal-Bench21.8%—
Berkeley Function Calling Leaderboard51.5%—

Reasoning Mercury leads

GPT-5 Nano: 16.3 (#306), Mercury: 17.5 (#293)

Reasoning benchmarks
BenchmarkGPT-5 NanoMercury
Kagi LLM Benchmark62.2%21.6%
LMArena Hard Prompts13281285
ARC-AGI-22.6%—
ARC-AGI-120.7%—
Chess Puzzles27%—
Mystery Game Puzzles9%—
DTBench62.7%—
LMCA7.9%—
Epoch Capabilities Index139.38—
ForecastBench59.1—

Math Not comparable

GPT-5 Nano: 29.4 (#241), Mercury: —

Knowledge Not comparable

GPT-5 Nano: 35.9 (#178), Mercury: —

Knowledge benchmarks
BenchmarkGPT-5 NanoMercury
GPQA Diamond69.4%—
SimpleQA Verified11.7%—
MMLU-Pro77.8%—
Vectara Hallucination Rate10.5%—
GPQA (HELM)67.9%—
LMArena Expert1321—

Multimodal Not comparable

GPT-5 Nano: 31.3 (#108), Mercury: —

Multimodal benchmarks
BenchmarkGPT-5 NanoMercury
LMArena Vision1159—
VPCT37.2%—

Multilingual GPT-5 Nano leads

GPT-5 Nano: 45.3 (#172), Mercury: 41.6 (#206)

Multilingual benchmarks
BenchmarkGPT-5 NanoMercury
LMArena Non-English13131260
LMArena Chinese1356—
LMArena German1327—
LMArena Japanese1226—
LMArena Korean1269—
LMArena Russian1296—
LMArena Spanish1360—

Instruction Following GPT-5 Nano leads

GPT-5 Nano: 75.0 (#79), Mercury: 65.2 (#224)

Instruction Following benchmarks
BenchmarkGPT-5 NanoMercury
LMArena Instruction Following13061239
IFEval93.2%—

Long Context Mercury leads

GPT-5 Nano: 31.3 (#281), Mercury: 38.4 (#198)

Long Context benchmarks
BenchmarkGPT-5 NanoMercury
LMArena Longer Query13121266
Fiction.LiveBench44.4%—

Writing & Preference Mercury leads

GPT-5 Nano: 39.1 (#249), Mercury: 46.2 (#221)

Writing & Preference benchmarks
BenchmarkGPT-5 NanoMercury
LMArena Text13201282
LMArena Creative Writing12491191
LMArena Multi-Turn13111282
EQ-Bench Creative Writing705—
WildBench80.6%—

Frequently asked questions

Is GPT-5 Nano better than Mercury?

Mercury is the stronger model overall, scoring 37.6 to 33.5 on the Noometry Index.

Is GPT-5 Nano or Mercury better for coding?

Mercury scores higher on coding benchmarks: 38.7 versus 33.6 in the Noometry coding category.

How many benchmarks do GPT-5 Nano and Mercury share?

9 benchmarks have published results for both models. GPT-5 Nano has 49 scored results on Noometry and Mercury has 9.

Related comparisons

Go deeper