Model comparison

Grok 4.1 Fast vs Mercury

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 37.6 on the Noometry Index.

Last verified . 8 shared benchmarks.

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Mercury Inception

37.6

Rank #199 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Grok 4.1 Fast scores higher in 5 categories and Mercury in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.1 Fast leads 43.4 to 17.5.

Side by side

Grok 4.1 Fast and Mercury specifications
Grok 4.1 FastMercury
ProviderxAIInception
Noometry Index41.437.6
Released2025-06-27—
WeightsProprietaryProprietary
Context window128K—
Max output30K—
Input $ / M tokens$0.20—
Output $ / M tokens$0.50—
Results tracked329

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury leads

Grok 4.1 Fast: 34.1 (#245), Mercury: 38.7 (#170)

Coding benchmarks
BenchmarkGrok 4.1 FastMercury
LMArena Coding14111322
LMArena WebDev1242—
ALE-Bench394.93—

Agentic & Tool Use Not comparable

Grok 4.1 Fast: 36.3 (#39), Mercury: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1 FastMercury
Berkeley Function Calling Leaderboard69.6%—
τ²-bench Banking13.1%—
LMArena Search1171—
Vending-Bench 21,107—

Reasoning Grok 4.1 Fast leads

Grok 4.1 Fast: 43.4 (#49), Mercury: 17.5 (#293)

Reasoning benchmarks
BenchmarkGrok 4.1 FastMercury
LMArena Hard Prompts14071285
SimpleBench56%—
Kagi LLM Benchmark—21.6%
NYT Connections (extended)87.4%—
DTBench87.7%—
ForecastBench61—

Math Not comparable

Grok 4.1 Fast: 31.9 (#221), Mercury: —

Math benchmarks
BenchmarkGrok 4.1 FastMercury
MathArena Final-Answer Competitions60.9%—
ProofBench4%—
LMArena Math1408—

Knowledge Not comparable

Grok 4.1 Fast: 33.1 (#207), Mercury: —

Knowledge benchmarks
BenchmarkGrok 4.1 FastMercury
Vectara Hallucination Rate17.8%—
LMArena Expert1399—

Multimodal Not comparable

Grok 4.1 Fast: 37.0 (#76), Mercury: —

Multimodal benchmarks
BenchmarkGrok 4.1 FastMercury
LMArena Vision1201—

Multilingual Grok 4.1 Fast leads

Grok 4.1 Fast: 51.0 (#114), Mercury: 41.6 (#206)

Multilingual benchmarks
BenchmarkGrok 4.1 FastMercury
LMArena Non-English13911260
LMArena Chinese1441—
LMArena French1415—
LMArena German1404—
LMArena Japanese1349—
LMArena Korean1361—
LMArena Russian1387—
LMArena Spanish1413—

Instruction Following Grok 4.1 Fast leads

Grok 4.1 Fast: 72.7 (#133), Mercury: 65.2 (#224)

Instruction Following benchmarks
BenchmarkGrok 4.1 FastMercury
LMArena Instruction Following13761239

Long Context Grok 4.1 Fast leads

Grok 4.1 Fast: 42.4 (#126), Mercury: 38.4 (#198)

Long Context benchmarks
BenchmarkGrok 4.1 FastMercury
LMArena Longer Query13901266

Writing & Preference Grok 4.1 Fast leads

Grok 4.1 Fast: 57.2 (#131), Mercury: 46.2 (#221)

Writing & Preference benchmarks
BenchmarkGrok 4.1 FastMercury
LMArena Text14081282
LMArena Creative Writing13941191
LMArena Multi-Turn13891282
EQ-Bench Creative Writing1327—

Frequently asked questions

Is Grok 4.1 Fast better than Mercury?

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 37.6 on the Noometry Index.

Is Grok 4.1 Fast or Mercury better for coding?

Mercury scores higher on coding benchmarks: 38.7 versus 34.1 in the Noometry coding category.

How many benchmarks do Grok 4.1 Fast and Mercury share?

8 benchmarks have published results for both models. Grok 4.1 Fast has 32 scored results on Noometry and Mercury has 9.

Related comparisons

Go deeper