Model comparison
Grok 4.1 vs Nova 2.0 Pro Preview
Grok 4.1 is the stronger model overall, scoring 41.5 to 33.4 on the Noometry Index.
Last verified . 0 shared benchmarks.
Summary
- The widest gap is in agentic & tool use, where Grok 4.1 leads 34.1 to 21.9.
Side by side
| Grok 4.1 | Nova 2.0 Pro Preview | |
|---|---|---|
| Provider | xAI | Amazon |
| Noometry Index | 41.5 | 33.4 |
| Released | 2025-11-17 | 2025-12-02 |
| Weights | Proprietary | Proprietary |
| Context window | — | — |
| Max output | — | — |
| Input $ / M tokens | — | — |
| Output $ / M tokens | — | — |
| Results tracked | 19 | 3 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Nova 2.0 Pro Preview leads
Grok 4.1: 33.7 (#253), Nova 2.0 Pro Preview: 40.8 (#133)
| Benchmark | Grok 4.1 | Nova 2.0 Pro Preview |
|---|---|---|
| LMArena WebDev | 1214 | — |
| SciCode | — | 42.7% |
| LMArena Coding | 1445 | — |
Agentic & Tool Use Grok 4.1 leads
Grok 4.1: 34.1 (#49), Nova 2.0 Pro Preview: 21.9 (#136)
Reasoning Grok 4.1 leads
Grok 4.1: 29.5 (#91), Nova 2.0 Pro Preview: 22.4 (#194)
| Benchmark | Grok 4.1 | Nova 2.0 Pro Preview |
|---|---|---|
| CritPt | — | 0% |
| LMArena Hard Prompts | 1435 | — |
Math Not comparable
Grok 4.1: 38.9 (#120), Nova 2.0 Pro Preview: —
| Benchmark | Grok 4.1 | Nova 2.0 Pro Preview |
|---|---|---|
| LMArena Math | 1422 | — |
Knowledge Not comparable
Grok 4.1: 39.5 (#133), Nova 2.0 Pro Preview: —
| Benchmark | Grok 4.1 | Nova 2.0 Pro Preview |
|---|---|---|
| LMArena Expert | 1417 | — |
Multilingual Not comparable
Grok 4.1: 53.4 (#68), Nova 2.0 Pro Preview: —
| Benchmark | Grok 4.1 | Nova 2.0 Pro Preview |
|---|---|---|
| LMArena Non-English | 1425 | — |
| LMArena Chinese | 1465 | — |
| LMArena French | 1448 | — |
| LMArena German | 1446 | — |
| LMArena Japanese | 1397 | — |
| LMArena Korean | 1407 | — |
| LMArena Russian | 1434 | — |
| LMArena Spanish | 1438 | — |
Instruction Following Not comparable
Grok 4.1: 73.8 (#111), Nova 2.0 Pro Preview: —
| Benchmark | Grok 4.1 | Nova 2.0 Pro Preview |
|---|---|---|
| LMArena Instruction Following | 1400 | — |
Long Context Not comparable
Grok 4.1: 43.2 (#100), Nova 2.0 Pro Preview: —
| Benchmark | Grok 4.1 | Nova 2.0 Pro Preview |
|---|---|---|
| LMArena Longer Query | 1416 | — |
Writing & Preference Not comparable
Grok 4.1: 62.4 (#75), Nova 2.0 Pro Preview: —
| Benchmark | Grok 4.1 | Nova 2.0 Pro Preview |
|---|---|---|
| LMArena Text | 1437 | — |
| LMArena Creative Writing | 1411 | — |
| LMArena Multi-Turn | 1437 | — |
Frequently asked questions
Is Grok 4.1 better than Nova 2.0 Pro Preview?
Grok 4.1 is the stronger model overall, scoring 41.5 to 33.4 on the Noometry Index.
Is Grok 4.1 or Nova 2.0 Pro Preview better for coding?
Nova 2.0 Pro Preview scores higher on coding benchmarks: 40.8 versus 33.7 in the Noometry coding category.
How many benchmarks do Grok 4.1 and Nova 2.0 Pro Preview share?
0 benchmarks have published results for both models. Grok 4.1 has 19 scored results on Noometry and Nova 2.0 Pro Preview has 3.