Model comparison

GPT-6.1 Sol vs Grok 4.7

GPT-6.1 Sol is the stronger model overall, scoring 65.6 to 53.1 on the Noometry Index.

Last verified . 30 shared benchmarks.

GPT-6.1 Sol OpenAI

65.6

Rank #6 Confirmed

Grok 4.7 xAI

53.1

Rank #37 Confirmed

Summary

  • They share 30 benchmarks with published results for both. GPT-6.1 Sol scores higher in 9 categories and Grok 4.7 in 1 category; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-6.1 Sol leads 93.7 to 57.8.
  • The biggest single-benchmark swing is FrontierMath Tier 4: 100% for GPT-6.1 Sol and 17.1% for Grok 4.7.
  • Grok 4.7 is cheaper at $2 / $6 per million input/output tokens, against $2 / $10 for GPT-6.1 Sol.
  • GPT-6.1 Sol accepts more context: 1.05M tokens versus 500K.

Side by side

GPT-6.1 Sol and Grok 4.7 specifications
GPT-6.1 SolGrok 4.7
ProviderOpenAIxAI
Noometry Index65.653.1
Released2026-09-292026-09-21
WeightsProprietaryProprietary
Context window1.05M500K
Max output128K500K
Input $ / M tokens$2$2
Output $ / M tokens$10$6
Results tracked3439

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-6.1 Sol leads

GPT-6.1 Sol: 63.2 (#8), Grok 4.7: 58.0 (#18)

Coding benchmarks
BenchmarkGPT-6.1 SolGrok 4.7
FrontierCode50.2%47.6%
LMArena WebDev17551639
SciCode55.8%57.8%
LMArena Coding14871427
DeepSWE75.2%—
CursorBench—46.3%
FrontierSWE—29.5%

Agentic & Tool Use GPT-6.1 Sol leads

GPT-6.1 Sol: 39.6 (#26), Grok 4.7: 36.7 (#37)

Agentic & Tool Use benchmarks
BenchmarkGPT-6.1 SolGrok 4.7
APEX-Agents60%54.6%
GDP.pdf32%22.8%
Vending-Bench 2—10,537

Reasoning GPT-6.1 Sol leads

GPT-6.1 Sol: 81.9 (#2), Grok 4.7: 49.1 (#40)

Reasoning benchmarks
BenchmarkGPT-6.1 SolGrok 4.7
NYT Connections (extended)95.5%76.8%
CritPt31.7%18%
Chess Puzzles61%38%
LMArena Hard Prompts14661413
Mystery Game Puzzles80%29%
Epoch Capabilities Index166.09153.53
ARC-AGI-294.2%—
ARC-AGI-198.5%—
EBR-Bench54.3%—
DTBench—96%
LMCA—49.4%

Math GPT-6.1 Sol leads

GPT-6.1 Sol: 93.7 (#1), Grok 4.7: 57.8 (#39)

Math benchmarks
BenchmarkGPT-6.1 SolGrok 4.7
FrontierMath (Tiers 1-3)93.7%53%
FrontierMath Tier 4100%17.1%
OTIS Mock AIME 2024-2025100%98.1%
ProofBench99%34%
LMArena Math14641407

Knowledge GPT-6.1 Sol leads

GPT-6.1 Sol: 71.8 (#4), Grok 4.7: 62.8 (#22)

Knowledge benchmarks
BenchmarkGPT-6.1 SolGrok 4.7
GPQA Diamond95.4%92.7%
SimpleQA Verified73.9%56%
LMArena Expert15021422

Multimodal GPT-6.1 Sol leads

GPT-6.1 Sol: 52.7 (#5), Grok 4.7: 35.5 (#87)

Multimodal benchmarks
BenchmarkGPT-6.1 SolGrok 4.7
LMArena Vision12881228
Furniture Assembly80%20.8%
Blueprint-Bench 2—32.5%

Multilingual GPT-6.1 Sol leads

GPT-6.1 Sol: 54.3 (#46), Grok 4.7: 50.8 (#116)

Multilingual benchmarks
BenchmarkGPT-6.1 SolGrok 4.7
LMArena Non-English14381389
LMArena Chinese14771455
LMArena Russian14551397
LMArena French—1455
LMArena Spanish—1400

Instruction Following GPT-6.1 Sol leads

GPT-6.1 Sol: 77.0 (#29), Grok 4.7: 74.1 (#105)

Instruction Following benchmarks
BenchmarkGPT-6.1 SolGrok 4.7
LMArena Instruction Following14681404

Long Context GPT-6.1 Sol leads

GPT-6.1 Sol: 44.9 (#54), Grok 4.7: 43.1 (#104)

Long Context benchmarks
BenchmarkGPT-6.1 SolGrok 4.7
LMArena Longer Query14651413

Writing & Preference Grok 4.7 leads

GPT-6.1 Sol: 63.6 (#63), Grok 4.7: 70.0 (#24)

Writing & Preference benchmarks
BenchmarkGPT-6.1 SolGrok 4.7
LMArena Text14471399
LMArena Creative Writing14321391
LMArena Multi-Turn14491393
EQ-Bench Creative Writing—2007

Frequently asked questions

Is GPT-6.1 Sol better than Grok 4.7?

GPT-6.1 Sol is the stronger model overall, scoring 65.6 to 53.1 on the Noometry Index.

Which is cheaper, GPT-6.1 Sol or Grok 4.7?

Grok 4.7 is cheaper. It lists at $2 per million input tokens and $6 per million output tokens; GPT-6.1 Sol lists at $2 and $10.

Is GPT-6.1 Sol or Grok 4.7 better for coding?

GPT-6.1 Sol scores higher on coding benchmarks: 63.2 versus 58.0 in the Noometry coding category.

Which has the bigger context window?

GPT-6.1 Sol does, with 1.05M tokens against 500K.

How many benchmarks do GPT-6.1 Sol and Grok 4.7 share?

30 benchmarks have published results for both models. GPT-6.1 Sol has 34 scored results on Noometry and Grok 4.7 has 39.

Related comparisons

Go deeper