Model comparison

Gemini 3 Pro vs Grok 4.6

Grok 4.6 is the stronger model overall, scoring 56.9 to 54.8 on the Noometry Index.

Last verified . 33 shared benchmarks.

Gemini 3 Pro Google

54.8

Rank #28 Confirmed

Grok 4.6 xAI

56.9

Rank #21 Confirmed

Summary

  • They share 33 benchmarks with published results for both. Gemini 3 Pro scores higher in 6 categories and Grok 4.6 in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok 4.6 leads 67.0 to 49.9.
  • The biggest single-benchmark swing is ARC-AGI-2: 31.1% for Gemini 3 Pro and 67.1% for Grok 4.6.

Side by side

Gemini 3 Pro and Grok 4.6 specifications
Gemini 3 ProGrok 4.6
ProviderGooglexAI
Noometry Index54.856.9
Released2025-11-182026-08-12
WeightsProprietaryProprietary
Context window—500K
Max output—500K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked6749

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.6 leads

Gemini 3 Pro: 51.6 (#39), Grok 4.6: 58.5 (#16)

Coding benchmarks
BenchmarkGemini 3 ProGrok 4.6
LMArena WebDev14401617
WeirdML69.9%67.3%
LMArena Coding14811465
ALE-Bench1,1771,508
SWE-bench Verified72.9%—
DeepSWE—67.5%
FrontierCode—48%
SWE-bench Verified (bash only)74.2%—
CursorBench—41.4%
SWE-bench Multilingual68.7%—
FrontierSWE—25.3%
SciCode—56.5%
GSO18.6%—
AlgoTune1.83—

Agentic & Tool Use Gemini 3 Pro leads

Gemini 3 Pro: 40.6 (#23), Grok 4.6: 39.4 (#27)

Agentic & Tool Use benchmarks
BenchmarkGemini 3 ProGrok 4.6
Vending-Bench 25,4789,047
Terminal-Bench69.4%—
APEX-Agents—65.3%
Berkeley Function Calling Leaderboard72.5%—
GDPval40.3%—
Remote Labor Index1.3%—
τ²-bench Airline80.5%—
τ²-bench Banking18%—
τ²-bench Retail75.9%—
τ²-bench Telecom91%—
DeepResearch Bench46.3%—
BALROG58.1%—
GDP.pdf—17.2%
LMArena Search1207—
METR Time Horizons71%—

Reasoning Grok 4.6 leads

Gemini 3 Pro: 52.5 (#31), Grok 4.6: 61.4 (#20)

Reasoning benchmarks
BenchmarkGemini 3 ProGrok 4.6
ARC-AGI-231.1%67.1%
SimpleBench76.4%75.9%
NYT Connections (extended)94.4%80%
ARC-AGI-175%87.5%
CritPt6.9%19.7%
Chess Puzzles31%40%
LMArena Hard Prompts14801447
Epoch Capabilities Index152.92156.44
Kagi LLM Benchmark80.1%—
EnigmaEval18.2%—
EBR-Bench—30.5%
Mystery Game Puzzles—34%
DTBench—97.3%
LMCA—48.5%
ForecastBench61.2—

Math Grok 4.6 leads

Gemini 3 Pro: 49.9 (#59), Grok 4.6: 67.0 (#24)

Knowledge Gemini 3 Pro leads

Gemini 3 Pro: 64.4 (#16), Grok 4.6: 63.3 (#20)

Knowledge benchmarks
BenchmarkGemini 3 ProGrok 4.6
GPQA Diamond92.6%94%
LMArena Expert14751467
Humanity's Last Exam37.5%—
SimpleQA Verified—49.3%
MMLU-Pro90.3%—
Vectara Hallucination Rate13.6%—
GPQA (HELM)80.3%—

Multimodal Gemini 3 Pro leads

Gemini 3 Pro: 57.6 (#2), Grok 4.6: 43.6 (#23)

Multimodal benchmarks
BenchmarkGemini 3 ProGrok 4.6
LMArena Vision13051263
LMArena Document14341452
GeoBench84%—
VPCT91%—
Blueprint-Bench 2—33.2%
Furniture Assembly—40%

Multilingual Gemini 3 Pro leads

Gemini 3 Pro: 56.9 (#16), Grok 4.6: 53.0 (#74)

Multilingual benchmarks
BenchmarkGemini 3 ProGrok 4.6
LMArena Non-English14741420
LMArena Chinese15231480
LMArena French14921461
LMArena German15151431
LMArena Japanese15101376
LMArena Korean14481397
LMArena Russian14931422
LMArena Spanish14701404

Instruction Following Too close to call

Gemini 3 Pro: 76.3 (#45), Grok 4.6: 75.4 (#63)

Instruction Following benchmarks
BenchmarkGemini 3 ProGrok 4.6
LMArena Instruction Following14581431
IFEval87.7%—

Long Context Too close to call

Gemini 3 Pro: 44.0 (#79), Grok 4.6: 44.5 (#66)

Long Context benchmarks
BenchmarkGemini 3 ProGrok 4.6
LMArena Longer Query14711454
CL-bench15.8%—

Writing & Preference Gemini 3 Pro leads

Gemini 3 Pro: 66.4 (#35), Grok 4.6: 62.3 (#80)

Writing & Preference benchmarks
BenchmarkGemini 3 ProGrok 4.6
LMArena Text14791428
LMArena Creative Writing14821428
LMArena Multi-Turn14841425
EQ-Bench Creative Writing1525—
WildBench85.9%—

Frequently asked questions

Is Gemini 3 Pro better than Grok 4.6?

Grok 4.6 is the stronger model overall, scoring 56.9 to 54.8 on the Noometry Index.

Is Gemini 3 Pro or Grok 4.6 better for coding?

Grok 4.6 scores higher on coding benchmarks: 58.5 versus 51.6 in the Noometry coding category.

How many benchmarks do Gemini 3 Pro and Grok 4.6 share?

33 benchmarks have published results for both models. Gemini 3 Pro has 67 scored results on Noometry and Grok 4.6 has 49.

Related comparisons

Go deeper