Model comparison

Grok 4.3 vs o3

o3 is the stronger model overall, scoring 47.5 to 43.8 on the Noometry Index. Grok 4.3 costs 2.2× less per token, which makes it the better buy when o3's lead doesn't matter for your workload.

Last verified . 31 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 31 benchmarks with published results for both. Grok 4.3 scores higher in 1 category and o3 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in long context, where o3 leads 53.3 to 42.5.
  • The biggest single-benchmark swing is SimpleQA Verified: 33.2% for Grok 4.3 and 49.4% for o3.
  • Grok 4.3 is cheaper at $1.25 / $2.50 per million input/output tokens, against $2 / $8 for o3.
  • Grok 4.3 accepts more context: 1M tokens versus 200K.

Side by side

Grok 4.3 and o3 specifications
Grok 4.3o3
ProviderxAIOpenAI
Noometry Index43.847.5
Released2026-04-172025-04-16
WeightsProprietaryProprietary
Context window1M200K
Max output30K100K
Input $ / M tokens$1.25$2
Output $ / M tokens$2.50$8
Results tracked4063

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Grok 4.3: 41.6 (#121), o3: 46.8 (#64)

Coding benchmarks
BenchmarkGrok 4.3o3
WeirdML49.9%52.4%
LMArena Coding14151408
ALE-Bench944.17933.55
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
LMArena WebDev1357—
SciCode47.3%—
GSO—8.8%
CadEval—74%

Agentic & Tool Use o3 leads

Grok 4.3: 27.7 (#99), o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.3o3
LMArena Search11651144
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
GDP.pdf8%—
METR Time Horizons—65.4%
Vending-Bench 235.26—

Reasoning Grok 4.3 leads

Grok 4.3: 35.9 (#68), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkGrok 4.3o3
CritPt8%1.4%
Chess Puzzles25%38%
LMArena Hard Prompts13961402
DTBench90.7%84.8%
LMCA38.3%39.7%
Epoch Capabilities Index149.16146.86
ForecastBench60.362.5
ARC-AGI-2—6.5%
SimpleBench—53.1%
Kagi LLM Benchmark—67.6%
NYT Connections (extended)55.2%—
ARC-AGI-1—60.8%
EnigmaEval—13.1%
Mystery Game Puzzles—29%

Math o3 leads

Grok 4.3: 46.0 (#74), o3: 50.2 (#58)

Knowledge o3 leads

Grok 4.3: 52.5 (#62), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkGrok 4.3o3
GPQA Diamond88.8%81.8%
SimpleQA Verified33.2%49.4%
LMArena Expert13851402
Humanity's Last Exam—20.3%
MMLU-Pro—85.9%
Confabulations—14.4%
GPQA (HELM)—75.3%

Multimodal o3 leads

Grok 4.3: 31.6 (#104), o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkGrok 4.3o3
LMArena Vision12291214
GeoBench—74%
VPCT—52%
Blueprint-Bench 20%—

Multilingual o3 leads

Grok 4.3: 50.5 (#120), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkGrok 4.3o3
LMArena Non-English13851401
LMArena Chinese14221437
LMArena French14121430
LMArena German13951420
LMArena Japanese13791403
LMArena Korean13561370
LMArena Russian13991406
LMArena Spanish13981395

Instruction Following Too close to call

Grok 4.3: 72.1 (#140), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkGrok 4.3o3
LMArena Instruction Following13661368
IFEval—86.9%

Long Context o3 leads

Grok 4.3: 42.5 (#123), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkGrok 4.3o3
LMArena Longer Query13931372
Fiction.LiveBench—88.9%
CL-bench—17.8%

Writing & Preference o3 leads

Grok 4.3: 58.5 (#118), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkGrok 4.3o3
LMArena Text13971410
LMArena Creative Writing13801359
LMArena Multi-Turn14061405
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
WildBench—86.1%
EQ-Bench 41075—

Frequently asked questions

Is Grok 4.3 better than o3?

o3 is the stronger model overall, scoring 47.5 to 43.8 on the Noometry Index. Grok 4.3 costs 2.2× less per token, which makes it the better buy when o3's lead doesn't matter for your workload.

Which is cheaper, Grok 4.3 or o3?

Grok 4.3 is cheaper. It lists at $1.25 per million input tokens and $2.50 per million output tokens; o3 lists at $2 and $8.

Is Grok 4.3 or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 41.6 in the Noometry coding category.

Which has the bigger context window?

Grok 4.3 does, with 1M tokens against 200K.

How many benchmarks do Grok 4.3 and o3 share?

31 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper