Model comparison

Grok 4.3 vs Step 1o Turbo 202506

Grok 4.3 is the stronger model overall, scoring 43.8 to 39.7 on the Noometry Index.

Last verified . 14 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Step 1o Turbo 202506 StepFun

39.7

Rank #160 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Grok 4.3 scores higher in 8 categories and Step 1o Turbo 202506 in 1 category; 9 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.3 leads 52.5 to 36.1.

Side by side

Grok 4.3 and Step 1o Turbo 202506 specifications
Grok 4.3Step 1o Turbo 202506
ProviderxAIStepFun
Noometry Index43.839.7
Released2026-04-17—
WeightsProprietaryProprietary
Context window1M—
Max output30K—
Input $ / M tokens$1.25—
Output $ / M tokens$2.50—
Results tracked4014

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.3 leads

Grok 4.3: 41.6 (#121), Step 1o Turbo 202506: 39.2 (#160)

Coding benchmarks
BenchmarkGrok 4.3Step 1o Turbo 202506
LMArena Coding14151339
LMArena WebDev1357—
SciCode47.3%—
WeirdML49.9%—
ALE-Bench944.17—

Agentic & Tool Use Not comparable

Grok 4.3: 27.7 (#99), Step 1o Turbo 202506: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.3Step 1o Turbo 202506
GDP.pdf8%—
LMArena Search1165—
Vending-Bench 235.26—

Reasoning Grok 4.3 leads

Grok 4.3: 35.9 (#68), Step 1o Turbo 202506: 26.8 (#129)

Reasoning benchmarks
BenchmarkGrok 4.3Step 1o Turbo 202506
LMArena Hard Prompts13961335
NYT Connections (extended)55.2%—
CritPt8%—
Chess Puzzles25%—
DTBench90.7%—
LMCA38.3%—
Epoch Capabilities Index149.16—
ForecastBench60.3—

Math Grok 4.3 leads

Grok 4.3: 46.0 (#74), Step 1o Turbo 202506: 36.6 (#164)

Math benchmarks
BenchmarkGrok 4.3Step 1o Turbo 202506
LMArena Math13881318
FrontierMath (Tiers 1-3)42.8%—
FrontierMath Tier 414.6%—
OTIS Mock AIME 2024-202593.3%—
ProofBench11%—

Knowledge Grok 4.3 leads

Grok 4.3: 52.5 (#62), Step 1o Turbo 202506: 36.1 (#176)

Knowledge benchmarks
BenchmarkGrok 4.3Step 1o Turbo 202506
LMArena Expert13851308
GPQA Diamond88.8%—
SimpleQA Verified33.2%—

Multimodal Step 1o Turbo 202506 leads

Grok 4.3: 31.6 (#104), Step 1o Turbo 202506: 36.1 (#80)

Multimodal benchmarks
BenchmarkGrok 4.3Step 1o Turbo 202506
LMArena Vision12291186
Blueprint-Bench 20%—

Multilingual Grok 4.3 leads

Grok 4.3: 50.5 (#120), Step 1o Turbo 202506: 45.3 (#173)

Multilingual benchmarks
BenchmarkGrok 4.3Step 1o Turbo 202506
LMArena Non-English13851313
LMArena Chinese14221380
LMArena German13951314
LMArena Russian13991327
LMArena French1412—
LMArena Japanese1379—
LMArena Korean1356—
LMArena Spanish1398—

Instruction Following Grok 4.3 leads

Grok 4.3: 72.1 (#140), Step 1o Turbo 202506: 69.2 (#176)

Instruction Following benchmarks
BenchmarkGrok 4.3Step 1o Turbo 202506
LMArena Instruction Following13661310

Long Context Grok 4.3 leads

Grok 4.3: 42.5 (#123), Step 1o Turbo 202506: 41.0 (#146)

Long Context benchmarks
BenchmarkGrok 4.3Step 1o Turbo 202506
LMArena Longer Query13931348

Writing & Preference Grok 4.3 leads

Grok 4.3: 58.5 (#118), Step 1o Turbo 202506: 53.1 (#160)

Writing & Preference benchmarks
BenchmarkGrok 4.3Step 1o Turbo 202506
LMArena Text13971336
LMArena Creative Writing13801307
LMArena Multi-Turn14061340
EQ-Bench 41075—

Frequently asked questions

Is Grok 4.3 better than Step 1o Turbo 202506?

Grok 4.3 is the stronger model overall, scoring 43.8 to 39.7 on the Noometry Index.

Is Grok 4.3 or Step 1o Turbo 202506 better for coding?

Grok 4.3 scores higher on coding benchmarks: 41.6 versus 39.2 in the Noometry coding category.

How many benchmarks do Grok 4.3 and Step 1o Turbo 202506 share?

14 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and Step 1o Turbo 202506 has 14.

Related comparisons

Go deeper