Model comparison

Grok 4.3 vs o3-pro

Grok 4.3 and o3-pro score almost the same on the Noometry Index (43.8 vs 42.9), so choose on price, context window or the category you care about most.

Last verified . 4 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

o3-pro OpenAI

42.9

Rank #105 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Grok 4.3 scores higher in 3 categories and o3-pro in 2 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in long context, where o3-pro leads 72.2 to 42.5.
  • The biggest single-benchmark swing is WeirdML: 49.9% for Grok 4.3 and 58.2% for o3-pro.
  • Grok 4.3 is cheaper at $1.25 / $2.50 per million input/output tokens, against $20 / $80 for o3-pro.
  • Grok 4.3 accepts more context: 1M tokens versus 200K.

Side by side

Grok 4.3 and o3-pro specifications
Grok 4.3o3-pro
ProviderxAIOpenAI
Noometry Index43.842.9
Released2026-04-172025-06-10
WeightsProprietaryProprietary
Context window1M200K
Max output30K100K
Input $ / M tokens$1.25$20
Output $ / M tokens$2.50$80
Results tracked4012

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3-pro leads

Grok 4.3: 41.6 (#121), o3-pro: 55.5 (#24)

Coding benchmarks
BenchmarkGrok 4.3o3-pro
WeirdML49.9%58.2%
Aider Polyglot—84.9%
LMArena WebDev1357—
SciCode47.3%—
LMArena Coding1415—
ALE-Bench944.17—

Agentic & Tool Use Not comparable

Grok 4.3: 27.7 (#99), o3-pro: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.3o3-pro
GDP.pdf8%—
LMArena Search1165—
Vending-Bench 235.26—

Reasoning Grok 4.3 leads

Grok 4.3: 35.9 (#68), o3-pro: 23.8 (#171)

Reasoning benchmarks
BenchmarkGrok 4.3o3-pro
DTBench90.7%86.9%
LMCA38.3%38.5%
Epoch Capabilities Index149.16147.42
ARC-AGI-2—4.9%
Kagi LLM Benchmark—72.1%
NYT Connections (extended)55.2%—
ARC-AGI-1—59.3%
CritPt8%—
Chess Puzzles25%—
LMArena Hard Prompts1396—
ForecastBench60.3—

Math Not comparable

Grok 4.3: 46.0 (#74), o3-pro: —

Math benchmarks
BenchmarkGrok 4.3o3-pro
FrontierMath (Tiers 1-3)42.8%—
FrontierMath Tier 414.6%—
OTIS Mock AIME 2024-202593.3%—
ProofBench11%—
LMArena Math1388—

Knowledge Grok 4.3 leads

Grok 4.3: 52.5 (#62), o3-pro: 29.5 (#238)

Knowledge benchmarks
BenchmarkGrok 4.3o3-pro
GPQA Diamond88.8%—
SimpleQA Verified33.2%—
Confabulations—14.2%
Vectara Hallucination Rate—23.3%
LMArena Expert1385—

Multimodal Not comparable

Grok 4.3: 31.6 (#104), o3-pro: —

Multimodal benchmarks
BenchmarkGrok 4.3o3-pro
LMArena Vision1229—
Blueprint-Bench 20%—

Multilingual Not comparable

Grok 4.3: 50.5 (#120), o3-pro: —

Multilingual benchmarks
BenchmarkGrok 4.3o3-pro
LMArena Non-English1385—
LMArena Chinese1422—
LMArena French1412—
LMArena German1395—
LMArena Japanese1379—
LMArena Korean1356—
LMArena Russian1399—
LMArena Spanish1398—

Instruction Following Not comparable

Grok 4.3: 72.1 (#140), o3-pro: —

Instruction Following benchmarks
BenchmarkGrok 4.3o3-pro
LMArena Instruction Following1366—

Long Context o3-pro leads

Grok 4.3: 42.5 (#123), o3-pro: 72.2 (#1)

Long Context benchmarks
BenchmarkGrok 4.3o3-pro
Fiction.LiveBench—97.2%
LMArena Longer Query1393—

Writing & Preference Grok 4.3 leads

Grok 4.3: 58.5 (#118), o3-pro: 57.1 (#133)

Writing & Preference benchmarks
BenchmarkGrok 4.3o3-pro
LMArena Text1397—
LMArena Creative Writing1380—
Short-Story Creative Writing—84.4%
EQ-Bench 41075—
LMArena Multi-Turn1406—

Frequently asked questions

Is Grok 4.3 better than o3-pro?

Grok 4.3 and o3-pro score almost the same on the Noometry Index (43.8 vs 42.9), so choose on price, context window or the category you care about most.

Which is cheaper, Grok 4.3 or o3-pro?

Grok 4.3 is cheaper. It lists at $1.25 per million input tokens and $2.50 per million output tokens; o3-pro lists at $20 and $80.

Is Grok 4.3 or o3-pro better for coding?

o3-pro scores higher on coding benchmarks: 55.5 versus 41.6 in the Noometry coding category.

Which has the bigger context window?

Grok 4.3 does, with 1M tokens against 200K.

How many benchmarks do Grok 4.3 and o3-pro share?

4 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and o3-pro has 12.

Related comparisons

Go deeper