xAI, proprietary
Grok-2 (Dec 2024)
Grok-2 (Dec 2024) by xAI ranks 239th of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.7. Its strongest category is multilingual, where it ranks 188th.
Last verified
Specifications
- Noometry rank
- #239 of 354
- Index score
- 33.7
- Evidence
- Confirmed 34 results
- Provider
- xAI
- Released
- August 13, 2024
- Weights
- Proprietary
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- Not measured
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Coding 33.3
- Reasoning 16.9
- Math 20.8
- Knowledge 29.8
- Multilingual 43.1
- Instruction Following 66.9
- Long Context 38.8
- Writing & Preference 48.6
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 33.3 | #258 | 3 |
| Reasoning | 16.9 | #299 | 5 |
| Math | 20.8 | #284 | 4 |
| Knowledge | 29.8 | #233 | 3 |
| Multilingual | 43.1 | #188 | 1 |
| Instruction Following | 66.9 | #202 | 2 |
| Long Context | 38.8 | #190 | 1 |
| Writing & Preference | 48.6 | #198 | 5 |
Strengths and weaknesses
Categories where Grok-2 (Dec 2024) places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multilingual | 43.1 | −4.3 | #188 of 297, top 64% |
| Writing & Preference | 48.6 | −5.1 | #198 of 312, top 64% |
| Long Context | 38.8 | −2.2 | #190 of 296, top 65% |
Closest competitors
The models ranked just above and below Grok-2 (Dec 2024). When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| o1-mini | #235 | 34.0 | — | — | Compare |
| Qwen3.5-9B | #236 | 33.8 | $0.11 | — | Compare |
| Codellama 70b Instruct | #237 | 33.7 | — | — | Compare |
| Qwen3 8B | #238 | 33.7 | $0.31 | — | Compare |
| GPT-4.1 mini | #240 | 33.6 | $0.70 | 86 | Compare |
| GPT-5 Nano | #241 | 33.5 | $0.14 | 4 | Compare |
| Mercury 2.5 | #242 | 33.5 | $0.0675 | — | Compare |
| Mistral Small | #243 | 33.4 | $0.26 | 120 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| WeirdML | 22.2% | #102 of 119, top 86% | Epoch AI | ||
| LiveBench Coding | 46.4% | #21 of 39, top 54% | Epoch AI | ||
| LMArena Coding | 1287 | #207 of 294, top 71% | LMArena | 2026-10-08 |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SimpleBench | 22.7% | #70 of 77, top 91% | Epoch AI | ||
| LiveBench Reasoning | 54.8% | #16 of 39, top 42% | Epoch AI | ||
| LMArena Hard Prompts | 1272 | #204 of 297, top 69% | LMArena | 2026-10-08 | |
| DTBench | 65.2% | #100 of 151, top 67% | Epoch AI | ||
| LiveBench Data Analysis | 54.5% | #19 of 39, top 49% | Epoch AI | ||
| Epoch Capabilities Index | 130.48 | #136 of 213, top 64% | Epoch AI | 2024-12-12 | |
| LiveBench | 54.3% | #17 of 39, top 44% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 11.5% | #132 of 173, top 77% | Epoch AI | 2025-02-25 | |
| LiveBench Math | 54.9% | #18 of 39, top 47% | Epoch AI | ||
| LMArena Math | 1283 | #190 of 285, top 67% | LMArena | 2026-10-08 | |
| MATH Level 5 | 63.5% | #36 of 79, top 46% | Epoch AI | 2025-02-17 | |
| FrontierMath (Feb 2025 set) | 0.7% | #61 of 68, top 90% | Epoch AI | 2025-03-06 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 53.8% | #123 of 186, top 67% | Epoch AI | 2025-01-27 | |
| Confabulations (lower is better) | 20.1% | #33 of 51, top 65% | Lech Mazur benchmarks | ||
| LMArena Expert | 1254 | #193 of 273, top 71% | LMArena | 2026-10-08 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1282 | #188 of 297, top 64% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1289 | #191 of 285, top 68% | LMArena | 2026-10-08 | |
| LMArena French | 1318 | #156 of 223, top 70% | LMArena | 2026-10-08 | |
| LMArena German | 1287 | #155 of 231, top 68% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1244 | #143 of 211, top 68% | LMArena | 2026-10-08 | |
| LMArena Korean | 1237 | #148 of 213, top 70% | LMArena | 2026-10-08 | |
| LMArena Russian | 1286 | #185 of 283, top 66% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1281 | #168 of 226, top 75% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LiveBench Instruction Following | 69.6% | #18 of 39, top 47% | Epoch AI | ||
| LMArena Instruction Following | 1270 | #196 of 298, top 66% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1276 | #205 of 291, top 71% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1305 | #188 of 297, top 64% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1284 | #184 of 295, top 63% | LMArena | 2026-10-08 | |
| Short-Story Creative Writing | 63.6% | #35 of 39, top 90% | Epoch AI | ||
| LMArena Multi-Turn | 1290 | #194 of 295, top 66% | LMArena | 2026-10-08 | |
| LiveBench Language | 45.6% | #15 of 39, top 39% | Epoch AI |
Compare Grok-2 (Dec 2024)
- Grok-2 (Dec 2024) vs Qwen3 8B
- Grok-2 (Dec 2024) vs GPT-4.1 mini
- Grok-2 (Dec 2024) vs Codellama 70b Instruct
- Grok-2 (Dec 2024) vs GPT-5 Nano
- Grok-2 (Dec 2024) vs Qwen3.5-9B
- Grok-2 (Dec 2024) vs Mercury 2.5
- Grok-2 (Dec 2024) vs GPT-6 Astra
- Grok-2 (Dec 2024) vs Claude Fable 5.1
- Grok-2 (Dec 2024) vs Gemini 3.8 Flash
- Grok-2 (Dec 2024) vs Kimi K3
- Grok-2 (Dec 2024) vs Qwen3.8 Max
- Grok-2 (Dec 2024) vs GLM-5.3
- Grok-2 (Dec 2024) vs Muse Spark 1.3
- Grok-2 (Dec 2024) vs DeepSeek V4 Pro
Other xAI models
- Grok 4.656.9
- Grok 4.555.0
- Grok 4.753.1
- Grok 4.20 (Non-Reasoning)48.6
- Grok 448.1
- Grok 4.20 Multi-Agent46.2
- Grok 4.343.8
- Grok 4.141.5
Frequently asked questions
How good is Grok-2 (Dec 2024)?
Grok-2 (Dec 2024) by xAI ranks 239th of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.7. Its strongest category is multilingual, where it ranks 188th.
Is Grok-2 (Dec 2024) open source?
No. Grok-2 (Dec 2024) is proprietary and available only through xAI's API and partner platforms.
What are Grok-2 (Dec 2024)'s strengths and weaknesses?
Relative to other ranked models, Grok-2 (Dec 2024) places best in multilingual, writing & preference, long context and lowest in math, reasoning, coding.
What is Grok-2 (Dec 2024) best at?
Its best category is multilingual, where it ranks 188th on Noometry.