Model comparison

Codellama 34b Instruct vs Qwen2.5 Plus 1127

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 30.8 on the Noometry Index.

Last verified . 10 shared benchmarks.

Codellama 34b Instruct Meta

30.8

Rank #287 Confirmed

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Codellama 34b Instruct scores higher in 0 categories and Qwen2.5 Plus 1127 in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen2.5 Plus 1127 leads 49.4 to 28.2.
  • Codellama 34b Instruct has downloadable open weights; the other is API-only.

Side by side

Codellama 34b Instruct and Qwen2.5 Plus 1127 specifications
Codellama 34b InstructQwen2.5 Plus 1127
ProviderMetaAlibaba (Qwen)
Noometry Index30.838.8
Released——
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1414

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 Plus 1127 leads

Codellama 34b Instruct: 28.5 (#314), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkCodellama 34b InstructQwen2.5 Plus 1127
LMArena Coding10461314
BigCodeBench Instruct29%—
BigCodeBench Complete37.1%—
HumanEval+43.9%—
MBPP+56.3%—

Reasoning Qwen2.5 Plus 1127 leads

Codellama 34b Instruct: 19.6 (#255), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkCodellama 34b InstructQwen2.5 Plus 1127
LMArena Hard Prompts10321299

Math Qwen2.5 Plus 1127 leads

Codellama 34b Instruct: 31.0 (#230), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkCodellama 34b InstructQwen2.5 Plus 1127
LMArena Math10561298

Knowledge Not comparable

Codellama 34b Instruct: —, Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkCodellama 34b InstructQwen2.5 Plus 1127
LMArena Expert—1289

Multilingual Qwen2.5 Plus 1127 leads

Codellama 34b Instruct: 25.8 (#284), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkCodellama 34b InstructQwen2.5 Plus 1127
LMArena Non-English10111265
LMArena Chinese9761314
LMArena German—1231
LMArena Japanese—1207
LMArena Russian—1271

Instruction Following Qwen2.5 Plus 1127 leads

Codellama 34b Instruct: 52.2 (#291), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkCodellama 34b InstructQwen2.5 Plus 1127
LMArena Instruction Following10281275

Long Context Qwen2.5 Plus 1127 leads

Codellama 34b Instruct: 30.9 (#284), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkCodellama 34b InstructQwen2.5 Plus 1127
LMArena Longer Query10131292

Writing & Preference Qwen2.5 Plus 1127 leads

Codellama 34b Instruct: 28.2 (#297), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkCodellama 34b InstructQwen2.5 Plus 1127
LMArena Text10661299
LMArena Creative Writing10321262
LMArena Multi-Turn10151299

Frequently asked questions

Is Codellama 34b Instruct better than Qwen2.5 Plus 1127?

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 30.8 on the Noometry Index.

Is Codellama 34b Instruct or Qwen2.5 Plus 1127 better for coding?

Qwen2.5 Plus 1127 scores higher on coding benchmarks: 38.5 versus 28.5 in the Noometry coding category.

How many benchmarks do Codellama 34b Instruct and Qwen2.5 Plus 1127 share?

10 benchmarks have published results for both models. Codellama 34b Instruct has 14 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper