Model comparison

Grok Build 0.1 vs Llama 3.1 Nemotron Ultra 253b v1

Grok Build 0.1 and Llama 3.1 Nemotron Ultra 253b v1 score almost the same on the Noometry Index (36.4 vs 36.7), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Summary

  • The widest gap is in agentic & tool use, where Grok Build 0.1 leads 22.7 to 15.7.
  • Llama 3.1 Nemotron Ultra 253b v1 has downloadable open weights; the other is API-only.

Side by side

Grok Build 0.1 and Llama 3.1 Nemotron Ultra 253b v1 specifications
Grok Build 0.1Llama 3.1 Nemotron Ultra 253b v1
ProviderxAINVIDIA
Noometry Index36.436.7
Released2026-04-16—
WeightsProprietaryOpen
Context window256K—
Max output256K—
Input $ / M tokens$1—
Output $ / M tokens$2—
Results tracked311

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

Grok Build 0.1: 43.1 (#91), Llama 3.1 Nemotron Ultra 253b v1: 38.4 (#177)

Coding benchmarks
BenchmarkGrok Build 0.1Llama 3.1 Nemotron Ultra 253b v1
SciCode50.2%—
LMArena Coding—1312

Agentic & Tool Use Grok Build 0.1 leads

Grok Build 0.1: 22.7 (#129), Llama 3.1 Nemotron Ultra 253b v1: 15.7 (#149)

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1Llama 3.1 Nemotron Ultra 253b v1
Berkeley Function Calling Leaderboard—10%
GBAEval2.4%—

Reasoning Grok Build 0.1 leads

Grok Build 0.1: 32.2 (#77), Llama 3.1 Nemotron Ultra 253b v1: 26.3 (#134)

Reasoning benchmarks
BenchmarkGrok Build 0.1Llama 3.1 Nemotron Ultra 253b v1
CritPt9.1%—
LMArena Hard Prompts—1316

Math Not comparable

Grok Build 0.1: —, Llama 3.1 Nemotron Ultra 253b v1: 37.5 (#152)

Math benchmarks
BenchmarkGrok Build 0.1Llama 3.1 Nemotron Ultra 253b v1
LMArena Math—1360

Multilingual Not comparable

Grok Build 0.1: —, Llama 3.1 Nemotron Ultra 253b v1: 43.1 (#187)

Multilingual benchmarks
BenchmarkGrok Build 0.1Llama 3.1 Nemotron Ultra 253b v1
LMArena Non-English—1282
LMArena Russian—1284

Instruction Following Not comparable

Grok Build 0.1: —, Llama 3.1 Nemotron Ultra 253b v1: 69.0 (#178)

Instruction Following benchmarks
BenchmarkGrok Build 0.1Llama 3.1 Nemotron Ultra 253b v1
LMArena Instruction Following—1308

Long Context Not comparable

Grok Build 0.1: —, Llama 3.1 Nemotron Ultra 253b v1: 39.5 (#177)

Long Context benchmarks
BenchmarkGrok Build 0.1Llama 3.1 Nemotron Ultra 253b v1
LMArena Longer Query—1299

Writing & Preference Not comparable

Grok Build 0.1: —, Llama 3.1 Nemotron Ultra 253b v1: 52.2 (#175)

Writing & Preference benchmarks
BenchmarkGrok Build 0.1Llama 3.1 Nemotron Ultra 253b v1
LMArena Text—1320
LMArena Creative Writing—1314
LMArena Multi-Turn—1317

Frequently asked questions

Is Grok Build 0.1 better than Llama 3.1 Nemotron Ultra 253b v1?

Grok Build 0.1 and Llama 3.1 Nemotron Ultra 253b v1 score almost the same on the Noometry Index (36.4 vs 36.7), so choose on price, context window or the category you care about most.

Is Grok Build 0.1 or Llama 3.1 Nemotron Ultra 253b v1 better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 38.4 in the Noometry coding category.

How many benchmarks do Grok Build 0.1 and Llama 3.1 Nemotron Ultra 253b v1 share?

0 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and Llama 3.1 Nemotron Ultra 253b v1 has 11.

Related comparisons

Go deeper