Model comparison

Claude 2.1 vs Grok 4.1 Fast

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 25.2 on the Noometry Index.

Last verified . 2 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Claude 2.1 scores higher in 0 categories and Grok 4.1 Fast in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.1 Fast leads 43.4 to 21.4.
  • The biggest single-benchmark swing is DTBench: 51% for Claude 2.1 and 87.7% for Grok 4.1 Fast.

Side by side

Claude 2.1 and Grok 4.1 Fast specifications
Claude 2.1Grok 4.1 Fast
ProviderAnthropicxAI
Noometry Index25.241.4
Released2023-11-212025-06-27
WeightsProprietaryProprietary
Context window—128K
Max output—30K
Input $ / M tokens—$0.20
Output $ / M tokens—$0.50
Results tracked732

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.1 Fast leads

Claude 2.1: 26.2 (#327), Grok 4.1 Fast: 34.1 (#245)

Coding benchmarks
BenchmarkClaude 2.1Grok 4.1 Fast
LMArena WebDev—1242
WeirdML7.1%—
LMArena Coding—1411
ALE-Bench—394.93

Agentic & Tool Use Not comparable

Claude 2.1: —, Grok 4.1 Fast: 36.3 (#39)

Agentic & Tool Use benchmarks
BenchmarkClaude 2.1Grok 4.1 Fast
Berkeley Function Calling Leaderboard—69.6%
τ²-bench Banking—13.1%
LMArena Search—1171
Vending-Bench 2—1,107

Reasoning Grok 4.1 Fast leads

Claude 2.1: 21.4 (#221), Grok 4.1 Fast: 43.4 (#49)

Reasoning benchmarks
BenchmarkClaude 2.1Grok 4.1 Fast
DTBench51%87.7%
ForecastBench54.261
SimpleBench—56%
NYT Connections (extended)—87.4%
LMArena Hard Prompts—1407
Epoch Capabilities Index119.27—

Math Grok 4.1 Fast leads

Claude 2.1: 10.2 (#315), Grok 4.1 Fast: 31.9 (#221)

Math benchmarks
BenchmarkClaude 2.1Grok 4.1 Fast
MathArena Final-Answer Competitions—60.9%
OTIS Mock AIME 2024-20251.9%—
ProofBench—4%
LMArena Math—1408

Knowledge Grok 4.1 Fast leads

Claude 2.1: 15.4 (#292), Grok 4.1 Fast: 33.1 (#207)

Knowledge benchmarks
BenchmarkClaude 2.1Grok 4.1 Fast
GPQA Diamond33%—
Vectara Hallucination Rate—17.8%
LMArena Expert—1399
MMLU73.5%—

Multimodal Not comparable

Claude 2.1: —, Grok 4.1 Fast: 37.0 (#76)

Multimodal benchmarks
BenchmarkClaude 2.1Grok 4.1 Fast
LMArena Vision—1201

Multilingual Not comparable

Claude 2.1: —, Grok 4.1 Fast: 51.0 (#114)

Multilingual benchmarks
BenchmarkClaude 2.1Grok 4.1 Fast
LMArena Non-English—1391
LMArena Chinese—1441
LMArena French—1415
LMArena German—1404
LMArena Japanese—1349
LMArena Korean—1361
LMArena Russian—1387
LMArena Spanish—1413

Instruction Following Not comparable

Claude 2.1: —, Grok 4.1 Fast: 72.7 (#133)

Instruction Following benchmarks
BenchmarkClaude 2.1Grok 4.1 Fast
LMArena Instruction Following—1376

Long Context Not comparable

Claude 2.1: —, Grok 4.1 Fast: 42.4 (#126)

Long Context benchmarks
BenchmarkClaude 2.1Grok 4.1 Fast
LMArena Longer Query—1390

Writing & Preference Not comparable

Claude 2.1: —, Grok 4.1 Fast: 57.2 (#131)

Writing & Preference benchmarks
BenchmarkClaude 2.1Grok 4.1 Fast
LMArena Text—1408
LMArena Creative Writing—1394
EQ-Bench Creative Writing—1327
LMArena Multi-Turn—1389

Frequently asked questions

Is Claude 2.1 better than Grok 4.1 Fast?

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 25.2 on the Noometry Index.

Is Claude 2.1 or Grok 4.1 Fast better for coding?

Grok 4.1 Fast scores higher on coding benchmarks: 34.1 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and Grok 4.1 Fast share?

2 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Grok 4.1 Fast has 32.

Related comparisons

Go deeper