Model comparison

Grok 4.3 vs Qwen3-Next 80B-A3B Instruct

Grok 4.3 and Qwen3-Next 80B-A3B Instruct score almost the same on the Noometry Index (43.8 vs 43.0), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Grok 4.3 scores higher in 6 categories and Qwen3-Next 80B-A3B Instruct in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.3 leads 52.5 to 41.8.
  • Qwen3-Next 80B-A3B Instruct is cheaper at $0.50 / $2 per million input/output tokens, against $1.25 / $2.50 for Grok 4.3.
  • Grok 4.3 accepts more context: 1M tokens versus 131K.
  • Qwen3-Next 80B-A3B Instruct has downloadable open weights; the other is API-only.

Side by side

Grok 4.3 and Qwen3-Next 80B-A3B Instruct specifications
Grok 4.3Qwen3-Next 80B-A3B Instruct
ProviderxAIAlibaba (Qwen)
Noometry Index43.843.0
Released2026-04-172025-09
WeightsProprietaryOpen
Context window1M131K
Max output30K33K
Input $ / M tokens$1.25$0.50
Output $ / M tokens$2.50$2
Results tracked4025

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Grok 4.3: 41.6 (#121), Qwen3-Next 80B-A3B Instruct: 42.5 (#98)

Coding benchmarks
BenchmarkGrok 4.3Qwen3-Next 80B-A3B Instruct
LMArena Coding14151440
LMArena WebDev1357—
SciCode47.3%—
WeirdML49.9%—
ALE-Bench944.17—

Agentic & Tool Use Not comparable

Grok 4.3: 27.7 (#99), Qwen3-Next 80B-A3B Instruct: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.3Qwen3-Next 80B-A3B Instruct
GDP.pdf8%—
LMArena Search1165—
Vending-Bench 235.26—

Reasoning Grok 4.3 leads

Grok 4.3: 35.9 (#68), Qwen3-Next 80B-A3B Instruct: 31.1 (#81)

Reasoning benchmarks
BenchmarkGrok 4.3Qwen3-Next 80B-A3B Instruct
LMArena Hard Prompts13961428
Kagi LLM Benchmark—66.7%
NYT Connections (extended)55.2%—
CritPt8%—
Chess Puzzles25%—
DTBench90.7%—
LMCA38.3%—
Epoch Capabilities Index149.16—
ForecastBench60.3—

Math Grok 4.3 leads

Grok 4.3: 46.0 (#74), Qwen3-Next 80B-A3B Instruct: 38.8 (#126)

Math benchmarks
BenchmarkGrok 4.3Qwen3-Next 80B-A3B Instruct
LMArena Math13881440
FrontierMath (Tiers 1-3)42.8%—
FrontierMath Tier 414.6%—
OTIS Mock AIME 2024-202593.3%—
ProofBench11%—
Omni-MATH—46.7%

Knowledge Grok 4.3 leads

Grok 4.3: 52.5 (#62), Qwen3-Next 80B-A3B Instruct: 41.8 (#106)

Knowledge benchmarks
BenchmarkGrok 4.3Qwen3-Next 80B-A3B Instruct
LMArena Expert13851417
GPQA Diamond88.8%—
SimpleQA Verified33.2%—
MMLU-Pro—78.6%
Vectara Hallucination Rate—9.3%
GPQA (HELM)—63%

Multimodal Not comparable

Grok 4.3: 31.6 (#104), Qwen3-Next 80B-A3B Instruct: —

Multimodal benchmarks
BenchmarkGrok 4.3Qwen3-Next 80B-A3B Instruct
LMArena Vision1229—
Blueprint-Bench 20%—

Multilingual Qwen3-Next 80B-A3B Instruct leads

Grok 4.3: 50.5 (#120), Qwen3-Next 80B-A3B Instruct: 52.1 (#93)

Multilingual benchmarks
BenchmarkGrok 4.3Qwen3-Next 80B-A3B Instruct
LMArena Non-English13851407
LMArena Chinese14221460
LMArena French14121413
LMArena German13951417
LMArena Japanese13791395
LMArena Korean13561364
LMArena Russian13991404
LMArena Spanish13981435

Instruction Following Grok 4.3 leads

Grok 4.3: 72.1 (#140), Qwen3-Next 80B-A3B Instruct: 70.8 (#159)

Instruction Following benchmarks
BenchmarkGrok 4.3Qwen3-Next 80B-A3B Instruct
LMArena Instruction Following13661389
IFEval—81%

Long Context Grok 4.3 leads

Grok 4.3: 42.5 (#123), Qwen3-Next 80B-A3B Instruct: 37.0 (#223)

Long Context benchmarks
BenchmarkGrok 4.3Qwen3-Next 80B-A3B Instruct
LMArena Longer Query13931403
Fiction.LiveBench—55.6%

Writing & Preference Too close to call

Grok 4.3: 58.5 (#118), Qwen3-Next 80B-A3B Instruct: 58.0 (#121)

Writing & Preference benchmarks
BenchmarkGrok 4.3Qwen3-Next 80B-A3B Instruct
LMArena Text13971417
LMArena Creative Writing13801334
LMArena Multi-Turn14061416
WildBench—80.7%
EQ-Bench 41075—

Frequently asked questions

Is Grok 4.3 better than Qwen3-Next 80B-A3B Instruct?

Grok 4.3 and Qwen3-Next 80B-A3B Instruct score almost the same on the Noometry Index (43.8 vs 43.0), so choose on price, context window or the category you care about most.

Which is cheaper, Grok 4.3 or Qwen3-Next 80B-A3B Instruct?

Qwen3-Next 80B-A3B Instruct is cheaper. It lists at $0.50 per million input tokens and $2 per million output tokens; Grok 4.3 lists at $1.25 and $2.50.

Is Grok 4.3 or Qwen3-Next 80B-A3B Instruct better for coding?

They score almost the same on coding (41.6 vs 42.5); test both on your own repository before choosing.

Which has the bigger context window?

Grok 4.3 does, with 1M tokens against 131K.

How many benchmarks do Grok 4.3 and Qwen3-Next 80B-A3B Instruct share?

17 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and Qwen3-Next 80B-A3B Instruct has 25.

Related comparisons

Go deeper