Model comparison

GPT-4 Turbo vs Mercury 2

Mercury 2 is the stronger model overall, scoring 39.1 to 30.5 on the Noometry Index.

Last verified . 12 shared benchmarks.

GPT-4 Turbo OpenAI

30.5

Rank #292 Confirmed

Mercury 2 Inception

39.1

Rank #175 Confirmed

Summary

  • They share 12 benchmarks with published results for both. GPT-4 Turbo scores higher in 1 category and Mercury 2 in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Mercury 2 leads 36.2 to 24.3.
  • The biggest single-benchmark swing is WeirdML: 18% for GPT-4 Turbo and 43.2% for Mercury 2.
  • Mercury 2 is cheaper at $0.25 / $0.75 per million input/output tokens, against $10 / $30 for GPT-4 Turbo.

Side by side

GPT-4 Turbo and Mercury 2 specifications
GPT-4 TurboMercury 2
ProviderOpenAIInception
Noometry Index30.539.1
Released2023-11-062026-02-20
WeightsProprietaryProprietary
Context window128K128K
Max output4K50K
Input $ / M tokens$10$0.25
Output $ / M tokens$30$0.75
Results tracked3617

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GPT-4 Turbo: 33.8 (#249), Mercury 2: 33.5 (#255)

Coding benchmarks
BenchmarkGPT-4 TurboMercury 2
WeirdML18%43.2%
LMArena Coding12681391
LMArena WebDev—1171
SciCode—38.7%
BigCodeBench Instruct48.2%—
BigCodeBench Complete58.2%—
ALE-Bench—785.58
HumanEval+86.6%—
MBPP+73.3%—

Agentic & Tool Use Not comparable

GPT-4 Turbo: —, Mercury 2: —

Agentic & Tool Use benchmarks
BenchmarkGPT-4 TurboMercury 2
METR Time Horizons36.7%—

Reasoning Mercury 2 leads

GPT-4 Turbo: 15.3 (#317), Mercury 2: 23.8 (#170)

Reasoning benchmarks
BenchmarkGPT-4 TurboMercury 2
LMArena Hard Prompts12511362
SimpleBench25.1%—
CritPt—0.8%
Chess Puzzles6%—
DTBench61.6%—
LMCA9.8%—
Epoch Capabilities Index127.25—
ForecastBench59.4—

Math Not comparable

GPT-4 Turbo: 9.0 (#322), Mercury 2: —

Math benchmarks
BenchmarkGPT-4 TurboMercury 2
FrontierMath (Tiers 1-3)0.7%—
OTIS Mock AIME 2024-20256.7%—
LMArena Math1272—
MATH Level 546.7%—

Knowledge Mercury 2 leads

GPT-4 Turbo: 24.3 (#268), Mercury 2: 36.2 (#172)

Knowledge benchmarks
BenchmarkGPT-4 TurboMercury 2
LMArena Expert12231358
GPQA Diamond46.6%—
Confabulations28.4%—
Vectara Hallucination Rate—12.3%
MMLU81.3%—

Multimodal Not comparable

GPT-4 Turbo: 30.6 (#110), Mercury 2: —

Multimodal benchmarks
BenchmarkGPT-4 TurboMercury 2
LMArena Vision1090—

Multilingual Mercury 2 leads

GPT-4 Turbo: 40.5 (#216), Mercury 2: 46.6 (#157)

Multilingual benchmarks
BenchmarkGPT-4 TurboMercury 2
LMArena Non-English12451331
LMArena Chinese12421417
LMArena Russian12591304
LMArena French1276—
LMArena German1259—
LMArena Japanese1194—
LMArena Korean1187—
LMArena Spanish1260—

Instruction Following Mercury 2 leads

GPT-4 Turbo: 65.8 (#216), Mercury 2: 70.2 (#165)

Instruction Following benchmarks
BenchmarkGPT-4 TurboMercury 2
LMArena Instruction Following12491329

Long Context Mercury 2 leads

GPT-4 Turbo: 38.0 (#206), Mercury 2: 40.5 (#154)

Long Context benchmarks
BenchmarkGPT-4 TurboMercury 2
LMArena Longer Query12541330

Writing & Preference Mercury 2 leads

GPT-4 Turbo: 47.7 (#206), Mercury 2: 53.8 (#155)

Writing & Preference benchmarks
BenchmarkGPT-4 TurboMercury 2
LMArena Text12721355
LMArena Creative Writing12691289
LMArena Multi-Turn12671358

Frequently asked questions

Is GPT-4 Turbo better than Mercury 2?

Mercury 2 is the stronger model overall, scoring 39.1 to 30.5 on the Noometry Index.

Which is cheaper, GPT-4 Turbo or Mercury 2?

Mercury 2 is cheaper. It lists at $0.25 per million input tokens and $0.75 per million output tokens; GPT-4 Turbo lists at $10 and $30.

Is GPT-4 Turbo or Mercury 2 better for coding?

They score almost the same on coding (33.8 vs 33.5); test both on your own repository before choosing.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do GPT-4 Turbo and Mercury 2 share?

12 benchmarks have published results for both models. GPT-4 Turbo has 36 scored results on Noometry and Mercury 2 has 17.

Related comparisons

Go deeper