Model comparison

Grok Build 0.1 vs Llama 3.1-70B

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 29.6 on the Noometry Index. Llama 3.1-70B costs 3.1× less per token, which makes it the better buy when Grok Build 0.1's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Llama 3.1-70B Meta

29.6

Rank #308 Confirmed

Summary

  • The widest gap is in coding, where Grok Build 0.1 leads 43.1 to 30.3.
  • Llama 3.1-70B is cheaper at $0.40 / $0.40 per million input/output tokens, against $1 / $2 for Grok Build 0.1.
  • Grok Build 0.1 accepts more context: 256K tokens versus 128K.
  • Llama 3.1-70B has downloadable open weights; the other is API-only.

Side by side

Grok Build 0.1 and Llama 3.1-70B specifications
Grok Build 0.1Llama 3.1-70B
ProviderxAIMeta
Noometry Index36.429.6
Released2026-04-162024-07-23
WeightsProprietaryOpen
Context window256K128K
Max output256K4K
Input $ / M tokens$1$0.40
Output $ / M tokens$2$0.40
Results tracked335

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

Grok Build 0.1: 43.1 (#91), Llama 3.1-70B: 30.3 (#296)

Coding benchmarks
BenchmarkGrok Build 0.1Llama 3.1-70B
SciCode50.2%—
WeirdML—9%
BigCodeBench Instruct—46.1%
LMArena Coding—1260
BigCodeBench Complete—54.8%

Agentic & Tool Use Llama 3.1-70B leads

Grok Build 0.1: 22.7 (#129), Llama 3.1-70B: 25.1 (#112)

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1Llama 3.1-70B
TheAgentCompany—6.9%
BALROG—27.9%
GBAEval2.4%—

Reasoning Grok Build 0.1 leads

Grok Build 0.1: 32.2 (#77), Llama 3.1-70B: 21.6 (#220)

Reasoning benchmarks
BenchmarkGrok Build 0.1Llama 3.1-70B
CritPt9.1%—
LMArena Hard Prompts—1241
DTBench—60%
LMCA—14.8%
Epoch Capabilities Index—125.92

Math Not comparable

Grok Build 0.1: —, Llama 3.1-70B: 13.5 (#304)

Math benchmarks
BenchmarkGrok Build 0.1Llama 3.1-70B
OTIS Mock AIME 2024-2025—3.6%
Omni-MATH—21%
LMArena Math—1252
MATH Level 5—36.7%

Knowledge Not comparable

Grok Build 0.1: —, Llama 3.1-70B: 24.2 (#269)

Knowledge benchmarks
BenchmarkGrok Build 0.1Llama 3.1-70B
GPQA Diamond—44.2%
MMLU-Pro—65.3%
GPQA (HELM)—42.6%
LMArena Expert—1209
MMLU—80.1%

Multilingual Not comparable

Grok Build 0.1: —, Llama 3.1-70B: 38.8 (#225)

Multilingual benchmarks
BenchmarkGrok Build 0.1Llama 3.1-70B
LMArena Non-English—1219
LMArena Chinese—1215
LMArena French—1261
LMArena German—1222
LMArena Japanese—1132
LMArena Korean—1140
LMArena Russian—1234
LMArena Spanish—1253

Instruction Following Not comparable

Grok Build 0.1: —, Llama 3.1-70B: 65.3 (#223)

Instruction Following benchmarks
BenchmarkGrok Build 0.1Llama 3.1-70B
IFEval—82.1%
LMArena Instruction Following—1231

Long Context Not comparable

Grok Build 0.1: —, Llama 3.1-70B: 37.6 (#214)

Long Context benchmarks
BenchmarkGrok Build 0.1Llama 3.1-70B
LMArena Longer Query—1241

Writing & Preference Not comparable

Grok Build 0.1: —, Llama 3.1-70B: 35.4 (#267)

Writing & Preference benchmarks
BenchmarkGrok Build 0.1Llama 3.1-70B
LMArena Text—1261
LMArena Creative Writing—1232
EQ-Bench Creative Writing—784
WildBench—75.8%
LMArena Multi-Turn—1256

Frequently asked questions

Is Grok Build 0.1 better than Llama 3.1-70B?

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 29.6 on the Noometry Index. Llama 3.1-70B costs 3.1× less per token, which makes it the better buy when Grok Build 0.1's lead doesn't matter for your workload.

Which is cheaper, Grok Build 0.1 or Llama 3.1-70B?

Llama 3.1-70B is cheaper. It lists at $0.40 per million input tokens and $0.40 per million output tokens; Grok Build 0.1 lists at $1 and $2.

Is Grok Build 0.1 or Llama 3.1-70B better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 30.3 in the Noometry coding category.

Which has the bigger context window?

Grok Build 0.1 does, with 256K tokens against 128K.

How many benchmarks do Grok Build 0.1 and Llama 3.1-70B share?

0 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and Llama 3.1-70B has 35.

Related comparisons

Go deeper