Model comparison

Codellama 34b Instruct vs Qwen1.5 4b Chat

Codellama 34b Instruct is the stronger model overall, scoring 30.8 to 28.8 on the Noometry Index.

Last verified . 10 shared benchmarks.

Codellama 34b Instruct Meta

30.8

Rank #287 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Codellama 34b Instruct scores higher in 6 categories and Qwen1.5 4b Chat in 1 category; 4 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Codellama 34b Instruct leads 28.2 to 23.8.

Side by side

Codellama 34b Instruct and Qwen1.5 4b Chat specifications
Codellama 34b InstructQwen1.5 4b Chat
ProviderMetaAlibaba (Qwen)
Noometry Index30.828.8
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1413

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Codellama 34b Instruct: 28.5 (#314), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkCodellama 34b InstructQwen1.5 4b Chat
LMArena Coding1046999
BigCodeBench Instruct29%—
BigCodeBench Complete37.1%—
HumanEval+43.9%—
MBPP+56.3%—

Reasoning Codellama 34b Instruct leads

Codellama 34b Instruct: 19.6 (#255), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkCodellama 34b InstructQwen1.5 4b Chat
LMArena Hard Prompts1032976

Math Too close to call

Codellama 34b Instruct: 31.0 (#230), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkCodellama 34b InstructQwen1.5 4b Chat
LMArena Math10561026

Knowledge Not comparable

Codellama 34b Instruct: —, Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkCodellama 34b InstructQwen1.5 4b Chat
LMArena Expert—980

Multilingual Codellama 34b Instruct leads

Codellama 34b Instruct: 25.8 (#284), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkCodellama 34b InstructQwen1.5 4b Chat
LMArena Non-English1011979
LMArena Chinese9761024
LMArena German—902
LMArena Russian—952

Instruction Following Codellama 34b Instruct leads

Codellama 34b Instruct: 52.2 (#291), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkCodellama 34b InstructQwen1.5 4b Chat
LMArena Instruction Following1028978

Long Context Too close to call

Codellama 34b Instruct: 30.9 (#284), Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkCodellama 34b InstructQwen1.5 4b Chat
LMArena Longer Query1013988

Writing & Preference Codellama 34b Instruct leads

Codellama 34b Instruct: 28.2 (#297), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkCodellama 34b InstructQwen1.5 4b Chat
LMArena Text1066997
LMArena Creative Writing1032969
LMArena Multi-Turn1015977

Frequently asked questions

Is Codellama 34b Instruct better than Qwen1.5 4b Chat?

Codellama 34b Instruct is the stronger model overall, scoring 30.8 to 28.8 on the Noometry Index.

Is Codellama 34b Instruct or Qwen1.5 4b Chat better for coding?

They score almost the same on coding (28.5 vs 29.1); test both on your own repository before choosing.

How many benchmarks do Codellama 34b Instruct and Qwen1.5 4b Chat share?

10 benchmarks have published results for both models. Codellama 34b Instruct has 14 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper