Model comparison

Codellama 34b Instruct vs Grok Build 0.1

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 30.8 on the Noometry Index.

Last verified . 0 shared benchmarks.

Codellama 34b Instruct Meta

30.8

Rank #287 Confirmed

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Summary

  • The widest gap is in coding, where Grok Build 0.1 leads 43.1 to 28.5.
  • Codellama 34b Instruct has downloadable open weights; the other is API-only.

Side by side

Codellama 34b Instruct and Grok Build 0.1 specifications
Codellama 34b InstructGrok Build 0.1
ProviderMetaxAI
Noometry Index30.836.4
Released—2026-04-16
WeightsOpenProprietary
Context window—256K
Max output—256K
Input $ / M tokens—$1
Output $ / M tokens—$2
Results tracked143

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

Codellama 34b Instruct: 28.5 (#314), Grok Build 0.1: 43.1 (#91)

Coding benchmarks
BenchmarkCodellama 34b InstructGrok Build 0.1
SciCode—50.2%
BigCodeBench Instruct29%—
LMArena Coding1046—
BigCodeBench Complete37.1%—
HumanEval+43.9%—
MBPP+56.3%—

Agentic & Tool Use Not comparable

Codellama 34b Instruct: —, Grok Build 0.1: 22.7 (#129)

Agentic & Tool Use benchmarks
BenchmarkCodellama 34b InstructGrok Build 0.1
GBAEval—2.4%

Reasoning Grok Build 0.1 leads

Codellama 34b Instruct: 19.6 (#255), Grok Build 0.1: 32.2 (#77)

Reasoning benchmarks
BenchmarkCodellama 34b InstructGrok Build 0.1
CritPt—9.1%
LMArena Hard Prompts1032—

Math Not comparable

Codellama 34b Instruct: 31.0 (#230), Grok Build 0.1: —

Math benchmarks
BenchmarkCodellama 34b InstructGrok Build 0.1
LMArena Math1056—

Multilingual Not comparable

Codellama 34b Instruct: 25.8 (#284), Grok Build 0.1: —

Multilingual benchmarks
BenchmarkCodellama 34b InstructGrok Build 0.1
LMArena Non-English1011—
LMArena Chinese976—

Instruction Following Not comparable

Codellama 34b Instruct: 52.2 (#291), Grok Build 0.1: —

Instruction Following benchmarks
BenchmarkCodellama 34b InstructGrok Build 0.1
LMArena Instruction Following1028—

Long Context Not comparable

Codellama 34b Instruct: 30.9 (#284), Grok Build 0.1: —

Long Context benchmarks
BenchmarkCodellama 34b InstructGrok Build 0.1
LMArena Longer Query1013—

Writing & Preference Not comparable

Codellama 34b Instruct: 28.2 (#297), Grok Build 0.1: —

Writing & Preference benchmarks
BenchmarkCodellama 34b InstructGrok Build 0.1
LMArena Text1066—
LMArena Creative Writing1032—
LMArena Multi-Turn1015—

Frequently asked questions

Is Codellama 34b Instruct better than Grok Build 0.1?

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 30.8 on the Noometry Index.

Is Codellama 34b Instruct or Grok Build 0.1 better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 28.5 in the Noometry coding category.

How many benchmarks do Codellama 34b Instruct and Grok Build 0.1 share?

0 benchmarks have published results for both models. Codellama 34b Instruct has 14 scored results on Noometry and Grok Build 0.1 has 3.

Related comparisons

Go deeper