Model comparison
Grok 4.3 vs Qwen3-Next 80B-A3B Instruct
Grok 4.3 and Qwen3-Next 80B-A3B Instruct score almost the same on the Noometry Index (43.8 vs 43.0), so choose on price, context window or the category you care about most.
Last verified . 17 shared benchmarks.
Summary
- They share 17 benchmarks with published results for both. Grok 4.3 scores higher in 6 categories and Qwen3-Next 80B-A3B Instruct in 2 categories; 6 gaps are clear of the uncertainty.
- The widest gap is in knowledge, where Grok 4.3 leads 52.5 to 41.8.
- Qwen3-Next 80B-A3B Instruct is cheaper at $0.50 / $2 per million input/output tokens, against $1.25 / $2.50 for Grok 4.3.
- Grok 4.3 accepts more context: 1M tokens versus 131K.
- Qwen3-Next 80B-A3B Instruct has downloadable open weights; the other is API-only.
Side by side
| Grok 4.3 | Qwen3-Next 80B-A3B Instruct | |
|---|---|---|
| Provider | xAI | Alibaba (Qwen) |
| Noometry Index | 43.8 | 43.0 |
| Released | 2026-04-17 | 2025-09 |
| Weights | Proprietary | Open |
| Context window | 1M | 131K |
| Max output | 30K | 33K |
| Input $ / M tokens | $1.25 | $0.50 |
| Output $ / M tokens | $2.50 | $2 |
| Results tracked | 40 | 25 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Too close to call
Grok 4.3: 41.6 (#121), Qwen3-Next 80B-A3B Instruct: 42.5 (#98)
| Benchmark | Grok 4.3 | Qwen3-Next 80B-A3B Instruct |
|---|---|---|
| LMArena Coding | 1415 | 1440 |
| LMArena WebDev | 1357 | — |
| SciCode | 47.3% | — |
| WeirdML | 49.9% | — |
| ALE-Bench | 944.17 | — |
Agentic & Tool Use Not comparable
Grok 4.3: 27.7 (#99), Qwen3-Next 80B-A3B Instruct: —
| Benchmark | Grok 4.3 | Qwen3-Next 80B-A3B Instruct |
|---|---|---|
| GDP.pdf | 8% | — |
| LMArena Search | 1165 | — |
| Vending-Bench 2 | 35.26 | — |
Reasoning Grok 4.3 leads
Grok 4.3: 35.9 (#68), Qwen3-Next 80B-A3B Instruct: 31.1 (#81)
| Benchmark | Grok 4.3 | Qwen3-Next 80B-A3B Instruct |
|---|---|---|
| LMArena Hard Prompts | 1396 | 1428 |
| Kagi LLM Benchmark | — | 66.7% |
| NYT Connections (extended) | 55.2% | — |
| CritPt | 8% | — |
| Chess Puzzles | 25% | — |
| DTBench | 90.7% | — |
| LMCA | 38.3% | — |
| Epoch Capabilities Index | 149.16 | — |
| ForecastBench | 60.3 | — |
Math Grok 4.3 leads
Grok 4.3: 46.0 (#74), Qwen3-Next 80B-A3B Instruct: 38.8 (#126)
| Benchmark | Grok 4.3 | Qwen3-Next 80B-A3B Instruct |
|---|---|---|
| LMArena Math | 1388 | 1440 |
| FrontierMath (Tiers 1-3) | 42.8% | — |
| FrontierMath Tier 4 | 14.6% | — |
| OTIS Mock AIME 2024-2025 | 93.3% | — |
| ProofBench | 11% | — |
| Omni-MATH | — | 46.7% |
Knowledge Grok 4.3 leads
Grok 4.3: 52.5 (#62), Qwen3-Next 80B-A3B Instruct: 41.8 (#106)
| Benchmark | Grok 4.3 | Qwen3-Next 80B-A3B Instruct |
|---|---|---|
| LMArena Expert | 1385 | 1417 |
| GPQA Diamond | 88.8% | — |
| SimpleQA Verified | 33.2% | — |
| MMLU-Pro | — | 78.6% |
| Vectara Hallucination Rate | — | 9.3% |
| GPQA (HELM) | — | 63% |
Multimodal Not comparable
Grok 4.3: 31.6 (#104), Qwen3-Next 80B-A3B Instruct: —
| Benchmark | Grok 4.3 | Qwen3-Next 80B-A3B Instruct |
|---|---|---|
| LMArena Vision | 1229 | — |
| Blueprint-Bench 2 | 0% | — |
Multilingual Qwen3-Next 80B-A3B Instruct leads
Grok 4.3: 50.5 (#120), Qwen3-Next 80B-A3B Instruct: 52.1 (#93)
| Benchmark | Grok 4.3 | Qwen3-Next 80B-A3B Instruct |
|---|---|---|
| LMArena Non-English | 1385 | 1407 |
| LMArena Chinese | 1422 | 1460 |
| LMArena French | 1412 | 1413 |
| LMArena German | 1395 | 1417 |
| LMArena Japanese | 1379 | 1395 |
| LMArena Korean | 1356 | 1364 |
| LMArena Russian | 1399 | 1404 |
| LMArena Spanish | 1398 | 1435 |
Instruction Following Grok 4.3 leads
Grok 4.3: 72.1 (#140), Qwen3-Next 80B-A3B Instruct: 70.8 (#159)
| Benchmark | Grok 4.3 | Qwen3-Next 80B-A3B Instruct |
|---|---|---|
| LMArena Instruction Following | 1366 | 1389 |
| IFEval | — | 81% |
Long Context Grok 4.3 leads
Grok 4.3: 42.5 (#123), Qwen3-Next 80B-A3B Instruct: 37.0 (#223)
| Benchmark | Grok 4.3 | Qwen3-Next 80B-A3B Instruct |
|---|---|---|
| LMArena Longer Query | 1393 | 1403 |
| Fiction.LiveBench | — | 55.6% |
Writing & Preference Too close to call
Grok 4.3: 58.5 (#118), Qwen3-Next 80B-A3B Instruct: 58.0 (#121)
| Benchmark | Grok 4.3 | Qwen3-Next 80B-A3B Instruct |
|---|---|---|
| LMArena Text | 1397 | 1417 |
| LMArena Creative Writing | 1380 | 1334 |
| LMArena Multi-Turn | 1406 | 1416 |
| WildBench | — | 80.7% |
| EQ-Bench 4 | 1075 | — |
Frequently asked questions
Is Grok 4.3 better than Qwen3-Next 80B-A3B Instruct?
Grok 4.3 and Qwen3-Next 80B-A3B Instruct score almost the same on the Noometry Index (43.8 vs 43.0), so choose on price, context window or the category you care about most.
Which is cheaper, Grok 4.3 or Qwen3-Next 80B-A3B Instruct?
Qwen3-Next 80B-A3B Instruct is cheaper. It lists at $0.50 per million input tokens and $2 per million output tokens; Grok 4.3 lists at $1.25 and $2.50.
Is Grok 4.3 or Qwen3-Next 80B-A3B Instruct better for coding?
They score almost the same on coding (41.6 vs 42.5); test both on your own repository before choosing.
Which has the bigger context window?
Grok 4.3 does, with 1M tokens against 131K.
How many benchmarks do Grok 4.3 and Qwen3-Next 80B-A3B Instruct share?
17 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and Qwen3-Next 80B-A3B Instruct has 25.