Model comparison
Grok-2 (Dec 2024) vs Nova 2.0 Pro Preview
Grok-2 (Dec 2024) and Nova 2.0 Pro Preview score almost the same on the Noometry Index (33.7 vs 33.4), so choose on price, context window or the category you care about most.
Last verified . 0 shared benchmarks.
Summary
- The widest gap is in coding, where Nova 2.0 Pro Preview leads 40.8 to 33.3.
Side by side
| Grok-2 (Dec 2024) | Nova 2.0 Pro Preview | |
|---|---|---|
| Provider | xAI | Amazon |
| Noometry Index | 33.7 | 33.4 |
| Released | 2024-08-13 | 2025-12-02 |
| Weights | Proprietary | Proprietary |
| Context window | — | — |
| Max output | — | — |
| Input $ / M tokens | — | — |
| Output $ / M tokens | — | — |
| Results tracked | 34 | 3 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Nova 2.0 Pro Preview leads
Grok-2 (Dec 2024): 33.3 (#258), Nova 2.0 Pro Preview: 40.8 (#133)
| Benchmark | Grok-2 (Dec 2024) | Nova 2.0 Pro Preview |
|---|---|---|
| SciCode | — | 42.7% |
| WeirdML | 22.2% | — |
| LiveBench Coding | 46.4% | — |
| LMArena Coding | 1287 | — |
Agentic & Tool Use Not comparable
Grok-2 (Dec 2024): —, Nova 2.0 Pro Preview: 21.9 (#136)
| Benchmark | Grok-2 (Dec 2024) | Nova 2.0 Pro Preview |
|---|---|---|
| GDP.pdf | — | 2% |
Reasoning Nova 2.0 Pro Preview leads
Grok-2 (Dec 2024): 16.9 (#299), Nova 2.0 Pro Preview: 22.4 (#194)
| Benchmark | Grok-2 (Dec 2024) | Nova 2.0 Pro Preview |
|---|---|---|
| SimpleBench | 22.7% | — |
| CritPt | — | 0% |
| LiveBench Reasoning | 54.8% | — |
| LMArena Hard Prompts | 1272 | — |
| DTBench | 65.2% | — |
| LiveBench Data Analysis | 54.5% | — |
| Epoch Capabilities Index | 130.48 | — |
| LiveBench | 54.3% | — |
Math Not comparable
Grok-2 (Dec 2024): 20.8 (#284), Nova 2.0 Pro Preview: —
| Benchmark | Grok-2 (Dec 2024) | Nova 2.0 Pro Preview |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 11.5% | — |
| LiveBench Math | 54.9% | — |
| LMArena Math | 1283 | — |
| MATH Level 5 | 63.5% | — |
| FrontierMath (Feb 2025 set) | 0.7% | — |
Knowledge Not comparable
Grok-2 (Dec 2024): 29.8 (#233), Nova 2.0 Pro Preview: —
| Benchmark | Grok-2 (Dec 2024) | Nova 2.0 Pro Preview |
|---|---|---|
| GPQA Diamond | 53.8% | — |
| Confabulations | 20.1% | — |
| LMArena Expert | 1254 | — |
Multilingual Not comparable
Grok-2 (Dec 2024): 43.1 (#188), Nova 2.0 Pro Preview: —
| Benchmark | Grok-2 (Dec 2024) | Nova 2.0 Pro Preview |
|---|---|---|
| LMArena Non-English | 1282 | — |
| LMArena Chinese | 1289 | — |
| LMArena French | 1318 | — |
| LMArena German | 1287 | — |
| LMArena Japanese | 1244 | — |
| LMArena Korean | 1237 | — |
| LMArena Russian | 1286 | — |
| LMArena Spanish | 1281 | — |
Instruction Following Not comparable
Grok-2 (Dec 2024): 66.9 (#202), Nova 2.0 Pro Preview: —
| Benchmark | Grok-2 (Dec 2024) | Nova 2.0 Pro Preview |
|---|---|---|
| LiveBench Instruction Following | 69.6% | — |
| LMArena Instruction Following | 1270 | — |
Long Context Not comparable
Grok-2 (Dec 2024): 38.8 (#190), Nova 2.0 Pro Preview: —
| Benchmark | Grok-2 (Dec 2024) | Nova 2.0 Pro Preview |
|---|---|---|
| LMArena Longer Query | 1276 | — |
Writing & Preference Not comparable
Grok-2 (Dec 2024): 48.6 (#198), Nova 2.0 Pro Preview: —
| Benchmark | Grok-2 (Dec 2024) | Nova 2.0 Pro Preview |
|---|---|---|
| LMArena Text | 1305 | — |
| LMArena Creative Writing | 1284 | — |
| Short-Story Creative Writing | 63.6% | — |
| LMArena Multi-Turn | 1290 | — |
| LiveBench Language | 45.6% | — |
Frequently asked questions
Is Grok-2 (Dec 2024) better than Nova 2.0 Pro Preview?
Grok-2 (Dec 2024) and Nova 2.0 Pro Preview score almost the same on the Noometry Index (33.7 vs 33.4), so choose on price, context window or the category you care about most.
Is Grok-2 (Dec 2024) or Nova 2.0 Pro Preview better for coding?
Nova 2.0 Pro Preview scores higher on coding benchmarks: 40.8 versus 33.3 in the Noometry coding category.
How many benchmarks do Grok-2 (Dec 2024) and Nova 2.0 Pro Preview share?
0 benchmarks have published results for both models. Grok-2 (Dec 2024) has 34 scored results on Noometry and Nova 2.0 Pro Preview has 3.