Model comparison

Claude 2 vs Grok 4.1

Grok 4.1 is the stronger model overall, scoring 41.5 to 25.0 on the Noometry Index.

Last verified . 0 shared benchmarks.

Claude 2 Anthropic

25.0

Rank #346 Reported

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Summary

  • The widest gap is in math, where Grok 4.1 leads 38.9 to 9.3.

Side by side

Claude 2 and Grok 4.1 specifications
Claude 2Grok 4.1
ProviderAnthropicxAI
Noometry Index25.041.5
Released2023-07-112025-11-17
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked819

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Claude 2: —, Grok 4.1: 33.7 (#253)

Coding benchmarks
BenchmarkClaude 2Grok 4.1
LMArena WebDev—1214
LMArena Coding—1445
HumanEval+61.6%—

Agentic & Tool Use Not comparable

Claude 2: —, Grok 4.1: 34.1 (#49)

Agentic & Tool Use benchmarks
BenchmarkClaude 2Grok 4.1
Cybench—39%

Reasoning Grok 4.1 leads

Claude 2: 21.7 (#216), Grok 4.1: 29.5 (#91)

Reasoning benchmarks
BenchmarkClaude 2Grok 4.1
LMArena Hard Prompts—1435
DTBench51.9%—
Epoch Capabilities Index120.13—

Math Grok 4.1 leads

Claude 2: 9.3 (#320), Grok 4.1: 38.9 (#120)

Math benchmarks
BenchmarkClaude 2Grok 4.1
OTIS Mock AIME 2024-20252.5%—
LMArena Math—1422
MATH Level 511.7%—

Knowledge Grok 4.1 leads

Claude 2: 16.9 (#287), Grok 4.1: 39.5 (#133)

Knowledge benchmarks
BenchmarkClaude 2Grok 4.1
GPQA Diamond34.7%—
LMArena Expert—1417
MMLU78.5%—
TriviaQA87.5%—

Multilingual Not comparable

Claude 2: —, Grok 4.1: 53.4 (#68)

Multilingual benchmarks
BenchmarkClaude 2Grok 4.1
LMArena Non-English—1425
LMArena Chinese—1465
LMArena French—1448
LMArena German—1446
LMArena Japanese—1397
LMArena Korean—1407
LMArena Russian—1434
LMArena Spanish—1438

Instruction Following Not comparable

Claude 2: —, Grok 4.1: 73.8 (#111)

Instruction Following benchmarks
BenchmarkClaude 2Grok 4.1
LMArena Instruction Following—1400

Long Context Not comparable

Claude 2: —, Grok 4.1: 43.2 (#100)

Long Context benchmarks
BenchmarkClaude 2Grok 4.1
LMArena Longer Query—1416

Writing & Preference Not comparable

Claude 2: —, Grok 4.1: 62.4 (#75)

Writing & Preference benchmarks
BenchmarkClaude 2Grok 4.1
LMArena Text—1437
LMArena Creative Writing—1411
LMArena Multi-Turn—1437

Frequently asked questions

Is Claude 2 better than Grok 4.1?

Grok 4.1 is the stronger model overall, scoring 41.5 to 25.0 on the Noometry Index.

How many benchmarks do Claude 2 and Grok 4.1 share?

0 benchmarks have published results for both models. Claude 2 has 8 scored results on Noometry and Grok 4.1 has 19.

Related comparisons

Go deeper