Model comparison

C4ai Aya Expanse 32b vs Grok 4.3

Grok 4.3 is the stronger model overall, scoring 43.8 to 35.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

C4ai Aya Expanse 32b Cohere

35.9

Rank #221 Confirmed

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Summary

  • They share 17 benchmarks with published results for both. C4ai Aya Expanse 32b scores higher in 0 categories and Grok 4.3 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.3 leads 52.5 to 33.2.
  • Grok 4.3 accepts more context: 1M tokens versus 128K.
  • C4ai Aya Expanse 32b has downloadable open weights; the other is API-only.

Side by side

C4ai Aya Expanse 32b and Grok 4.3 specifications
C4ai Aya Expanse 32bGrok 4.3
ProviderCoherexAI
Noometry Index35.943.8
Released2024-10-242026-04-17
WeightsOpenProprietary
Context window128K1M
Max output4K30K
Input $ / M tokens—$1.25
Output $ / M tokens—$2.50
Results tracked1840

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.3 leads

C4ai Aya Expanse 32b: 34.8 (#231), Grok 4.3: 41.6 (#121)

Coding benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.3
LMArena Coding11971415
LMArena WebDev—1357
SciCode—47.3%
WeirdML—49.9%
ALE-Bench—944.17

Agentic & Tool Use Not comparable

C4ai Aya Expanse 32b: —, Grok 4.3: 27.7 (#99)

Agentic & Tool Use benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.3
GDP.pdf—8%
LMArena Search—1165
Vending-Bench 2—35.26

Reasoning Grok 4.3 leads

C4ai Aya Expanse 32b: 23.3 (#180), Grok 4.3: 35.9 (#68)

Reasoning benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.3
LMArena Hard Prompts11931396
NYT Connections (extended)—55.2%
CritPt—8%
Chess Puzzles—25%
DTBench—90.7%
LMCA—38.3%
Epoch Capabilities Index—149.16
ForecastBench—60.3

Math Grok 4.3 leads

C4ai Aya Expanse 32b: 34.0 (#197), Grok 4.3: 46.0 (#74)

Math benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.3
LMArena Math12001388
FrontierMath (Tiers 1-3)—42.8%
FrontierMath Tier 4—14.6%
OTIS Mock AIME 2024-2025—93.3%
ProofBench—11%

Knowledge Grok 4.3 leads

C4ai Aya Expanse 32b: 33.2 (#206), Grok 4.3: 52.5 (#62)

Knowledge benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.3
LMArena Expert11821385
GPQA Diamond—88.8%
SimpleQA Verified—33.2%
Vectara Hallucination Rate10.9%—

Multimodal Not comparable

C4ai Aya Expanse 32b: —, Grok 4.3: 31.6 (#104)

Multimodal benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.3
LMArena Vision—1229
Blueprint-Bench 2—0%

Multilingual Grok 4.3 leads

C4ai Aya Expanse 32b: 38.4 (#230), Grok 4.3: 50.5 (#120)

Multilingual benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.3
LMArena Non-English12131385
LMArena Chinese12111422
LMArena French12481412
LMArena German11991395
LMArena Japanese11631379
LMArena Korean11581356
LMArena Russian12271399
LMArena Spanish11931398

Instruction Following Grok 4.3 leads

C4ai Aya Expanse 32b: 62.6 (#237), Grok 4.3: 72.1 (#140)

Instruction Following benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.3
LMArena Instruction Following11961366

Long Context Grok 4.3 leads

C4ai Aya Expanse 32b: 37.2 (#220), Grok 4.3: 42.5 (#123)

Long Context benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.3
LMArena Longer Query12281393

Writing & Preference Grok 4.3 leads

C4ai Aya Expanse 32b: 42.2 (#235), Grok 4.3: 58.5 (#118)

Writing & Preference benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.3
LMArena Text12241397
LMArena Creative Writing12001380
LMArena Multi-Turn11901406
EQ-Bench 4—1075

Frequently asked questions

Is C4ai Aya Expanse 32b better than Grok 4.3?

Grok 4.3 is the stronger model overall, scoring 43.8 to 35.9 on the Noometry Index.

Is C4ai Aya Expanse 32b or Grok 4.3 better for coding?

Grok 4.3 scores higher on coding benchmarks: 41.6 versus 34.8 in the Noometry coding category.

Which has the bigger context window?

Grok 4.3 does, with 1M tokens against 128K.

How many benchmarks do C4ai Aya Expanse 32b and Grok 4.3 share?

17 benchmarks have published results for both models. C4ai Aya Expanse 32b has 18 scored results on Noometry and Grok 4.3 has 40.

Related comparisons

Go deeper