Model comparison

Amazon Nova Experimental Chat 10 20 vs Grok 4.3

Grok 4.3 is the stronger model overall, scoring 43.8 to 42.1 on the Noometry Index.

Last verified . 17 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Amazon Nova Experimental Chat 10 20 scores higher in 0 categories and Grok 4.3 in 8 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.3 leads 52.5 to 38.5.

Side by side

Amazon Nova Experimental Chat 10 20 and Grok 4.3 specifications
Amazon Nova Experimental Chat 10 20Grok 4.3
ProviderAmazonxAI
Noometry Index42.143.8
Released—2026-04-17
WeightsProprietaryProprietary
Context window—1M
Max output—30K
Input $ / M tokens—$1.25
Output $ / M tokens—$2.50
Results tracked1740

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Amazon Nova Experimental Chat 10 20: 41.6 (#123), Grok 4.3: 41.6 (#121)

Coding benchmarks
BenchmarkAmazon Nova Experimental Chat 10 20Grok 4.3
LMArena Coding14121415
LMArena WebDev—1357
SciCode—47.3%
WeirdML—49.9%
ALE-Bench—944.17

Agentic & Tool Use Not comparable

Amazon Nova Experimental Chat 10 20: —, Grok 4.3: 27.7 (#99)

Agentic & Tool Use benchmarks
BenchmarkAmazon Nova Experimental Chat 10 20Grok 4.3
GDP.pdf—8%
LMArena Search—1165
Vending-Bench 2—35.26

Reasoning Grok 4.3 leads

Amazon Nova Experimental Chat 10 20: 28.4 (#106), Grok 4.3: 35.9 (#68)

Reasoning benchmarks
BenchmarkAmazon Nova Experimental Chat 10 20Grok 4.3
LMArena Hard Prompts13961396
NYT Connections (extended)—55.2%
CritPt—8%
Chess Puzzles—25%
DTBench—90.7%
LMCA—38.3%
Epoch Capabilities Index—149.16
ForecastBench—60.3

Math Grok 4.3 leads

Amazon Nova Experimental Chat 10 20: 39.0 (#119), Grok 4.3: 46.0 (#74)

Math benchmarks
BenchmarkAmazon Nova Experimental Chat 10 20Grok 4.3
LMArena Math14251388
FrontierMath (Tiers 1-3)—42.8%
FrontierMath Tier 4—14.6%
OTIS Mock AIME 2024-2025—93.3%
ProofBench—11%

Knowledge Grok 4.3 leads

Amazon Nova Experimental Chat 10 20: 38.5 (#144), Grok 4.3: 52.5 (#62)

Knowledge benchmarks
BenchmarkAmazon Nova Experimental Chat 10 20Grok 4.3
LMArena Expert13861385
GPQA Diamond—88.8%
SimpleQA Verified—33.2%

Multimodal Not comparable

Amazon Nova Experimental Chat 10 20: —, Grok 4.3: 31.6 (#104)

Multimodal benchmarks
BenchmarkAmazon Nova Experimental Chat 10 20Grok 4.3
LMArena Vision—1229
Blueprint-Bench 2—0%

Multilingual Too close to call

Amazon Nova Experimental Chat 10 20: 49.5 (#131), Grok 4.3: 50.5 (#120)

Multilingual benchmarks
BenchmarkAmazon Nova Experimental Chat 10 20Grok 4.3
LMArena Non-English13721385
LMArena Chinese14041422
LMArena French14251412
LMArena German13911395
LMArena Japanese13601379
LMArena Korean13241356
LMArena Russian13711399
LMArena Spanish13821398

Instruction Following Too close to call

Amazon Nova Experimental Chat 10 20: 72.1 (#141), Grok 4.3: 72.1 (#140)

Instruction Following benchmarks
BenchmarkAmazon Nova Experimental Chat 10 20Grok 4.3
LMArena Instruction Following13651366

Long Context Too close to call

Amazon Nova Experimental Chat 10 20: 41.7 (#133), Grok 4.3: 42.5 (#123)

Long Context benchmarks
BenchmarkAmazon Nova Experimental Chat 10 20Grok 4.3
LMArena Longer Query13701393

Writing & Preference Grok 4.3 leads

Amazon Nova Experimental Chat 10 20: 56.6 (#136), Grok 4.3: 58.5 (#118)

Writing & Preference benchmarks
BenchmarkAmazon Nova Experimental Chat 10 20Grok 4.3
LMArena Text13941397
LMArena Creative Writing13181380
LMArena Multi-Turn13641406
EQ-Bench 4—1075

Frequently asked questions

Is Amazon Nova Experimental Chat 10 20 better than Grok 4.3?

Grok 4.3 is the stronger model overall, scoring 43.8 to 42.1 on the Noometry Index.

Is Amazon Nova Experimental Chat 10 20 or Grok 4.3 better for coding?

They score almost the same on coding (41.6 vs 41.6); test both on your own repository before choosing.

How many benchmarks do Amazon Nova Experimental Chat 10 20 and Grok 4.3 share?

17 benchmarks have published results for both models. Amazon Nova Experimental Chat 10 20 has 17 scored results on Noometry and Grok 4.3 has 40.

Related comparisons

Go deeper