xAI, proprietary

Grok-2 (Dec 2024)

Grok-2 (Dec 2024) by xAI ranks 239th of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.7. Its strongest category is multilingual, where it ranks 188th.

Last verified

Specifications

Noometry rank
#239 of 354
Index score
33.7
Evidence
Confirmed 34 results
Provider
xAI
Released
August 13, 2024
Weights
Proprietary
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
Not measured
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Grok-2 (Dec 2024) category scores
  1. Coding 33.3
  2. Reasoning 16.9
  3. Math 20.8
  4. Knowledge 29.8
  5. Multilingual 43.1
  6. Instruction Following 66.9
  7. Long Context 38.8
  8. Writing & Preference 48.6
Grok-2 (Dec 2024) category ranks
CategoryScoreRankResults
Coding33.3#2583
Reasoning16.9#2995
Math20.8#2844
Knowledge29.8#2333
Multilingual43.1#1881
Instruction Following66.9#2022
Long Context38.8#1901
Writing & Preference48.6#1985

Strengths and weaknesses

Categories where Grok-2 (Dec 2024) places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Grok-2 (Dec 2024): strongest categories
CategoryScorevs medianRank
Multilingual43.1−4.3#188 of 297, top 64%
Writing & Preference48.6−5.1#198 of 312, top 64%
Long Context38.8−2.2#190 of 296, top 65%

Weakest categories

Grok-2 (Dec 2024): weakest categories
CategoryScorevs medianRank
Math20.8−15.8#284 of 327, top 87%
Reasoning16.9−6.7#299 of 350, top 86%
Coding33.3−5.4#258 of 340, top 76%

Closest competitors

The models ranked just above and below Grok-2 (Dec 2024). When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Grok-2 (Dec 2024)
ModelRankScoreBlended $/MSpeed
o1-mini#23534.0——Compare
Qwen3.5-9B#23633.8$0.11—Compare
Codellama 70b Instruct#23733.7——Compare
Qwen3 8B#23833.7$0.31—Compare
GPT-4.1 mini#24033.6$0.7086Compare
GPT-5 Nano#24133.5$0.144Compare
Mercury 2.5#24233.5$0.0675—Compare
Mistral Small#24333.4$0.26120Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Grok-2 (Dec 2024) Coding benchmark results
BenchmarkScorePositionSettingSourceDate
WeirdML22.2%#102 of 119, top 86%Epoch AI
LiveBench Coding46.4%#21 of 39, top 54%Epoch AI
LMArena Coding1287#207 of 294, top 71%LMArena2026-10-08

Reasoning

Grok-2 (Dec 2024) Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
SimpleBench22.7%#70 of 77, top 91%Epoch AI
LiveBench Reasoning54.8%#16 of 39, top 42%Epoch AI
LMArena Hard Prompts1272#204 of 297, top 69%LMArena2026-10-08
DTBench65.2%#100 of 151, top 67%Epoch AI
LiveBench Data Analysis54.5%#19 of 39, top 49%Epoch AI
Epoch Capabilities Index130.48#136 of 213, top 64%Epoch AI2024-12-12
LiveBench54.3%#17 of 39, top 44%Epoch AI

Math

Grok-2 (Dec 2024) Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-202511.5%#132 of 173, top 77%Epoch AI2025-02-25
LiveBench Math54.9%#18 of 39, top 47%Epoch AI
LMArena Math1283#190 of 285, top 67%LMArena2026-10-08
MATH Level 563.5%#36 of 79, top 46%Epoch AI2025-02-17
FrontierMath (Feb 2025 set)0.7%#61 of 68, top 90%Epoch AI2025-03-06

Knowledge

Grok-2 (Dec 2024) Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond53.8%#123 of 186, top 67%Epoch AI2025-01-27
Confabulations (lower is better)20.1%#33 of 51, top 65%Lech Mazur benchmarks
LMArena Expert1254#193 of 273, top 71%LMArena2026-10-08

Multilingual

Grok-2 (Dec 2024) Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1282#188 of 297, top 64%LMArena2026-10-08
LMArena Chinese1289#191 of 285, top 68%LMArena2026-10-08
LMArena French1318#156 of 223, top 70%LMArena2026-10-08
LMArena German1287#155 of 231, top 68%LMArena2026-10-08
LMArena Japanese1244#143 of 211, top 68%LMArena2026-10-08
LMArena Korean1237#148 of 213, top 70%LMArena2026-10-08
LMArena Russian1286#185 of 283, top 66%LMArena2026-10-08
LMArena Spanish1281#168 of 226, top 75%LMArena2026-10-08

Instruction Following

Grok-2 (Dec 2024) Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following69.6%#18 of 39, top 47%Epoch AI
LMArena Instruction Following1270#196 of 298, top 66%LMArena2026-10-08

Long Context

Grok-2 (Dec 2024) Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1276#205 of 291, top 71%LMArena2026-10-08

Writing & Preference

Grok-2 (Dec 2024) Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1305#188 of 297, top 64%LMArena2026-10-08
LMArena Creative Writing1284#184 of 295, top 63%LMArena2026-10-08
Short-Story Creative Writing63.6%#35 of 39, top 90%Epoch AI
LMArena Multi-Turn1290#194 of 295, top 66%LMArena2026-10-08
LiveBench Language45.6%#15 of 39, top 39%Epoch AI

Compare Grok-2 (Dec 2024)

Other xAI models

Frequently asked questions

How good is Grok-2 (Dec 2024)?

Grok-2 (Dec 2024) by xAI ranks 239th of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.7. Its strongest category is multilingual, where it ranks 188th.

Is Grok-2 (Dec 2024) open source?

No. Grok-2 (Dec 2024) is proprietary and available only through xAI's API and partner platforms.

What are Grok-2 (Dec 2024)'s strengths and weaknesses?

Relative to other ranked models, Grok-2 (Dec 2024) places best in multilingual, writing & preference, long context and lowest in math, reasoning, coding.

What is Grok-2 (Dec 2024) best at?

Its best category is multilingual, where it ranks 188th on Noometry.