Model comparison

Codellama 70b Instruct vs Qwen1.5 4b Chat

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 28.8 on the Noometry Index.

Last verified . 4 shared benchmarks.

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Codellama 70b Instruct scores higher in 5 categories and Qwen1.5 4b Chat in 0 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Codellama 70b Instruct leads 33.4 to 23.8.

Side by side

Codellama 70b Instruct and Qwen1.5 4b Chat specifications
Codellama 70b InstructQwen1.5 4b Chat
ProviderMetaAlibaba (Qwen)
Noometry Index33.728.8
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked713

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Codellama 70b Instruct leads

Codellama 70b Instruct: 37.6 (#193), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkCodellama 70b InstructQwen1.5 4b Chat
BigCodeBench Instruct40.7%—
LMArena Coding—999
BigCodeBench Complete49.6%—
HumanEval+65.9%—

Reasoning Codellama 70b Instruct leads

Codellama 70b Instruct: 20.1 (#242), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkCodellama 70b InstructQwen1.5 4b Chat
LMArena Hard Prompts1052976

Math Not comparable

Codellama 70b Instruct: —, Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkCodellama 70b InstructQwen1.5 4b Chat
LMArena Math—1026

Knowledge Not comparable

Codellama 70b Instruct: —, Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkCodellama 70b InstructQwen1.5 4b Chat
LMArena Expert—980

Multilingual Too close to call

Codellama 70b Instruct: 24.8 (#288), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkCodellama 70b InstructQwen1.5 4b Chat
LMArena Non-English992979
LMArena Chinese—1024
LMArena German—902
LMArena Russian—952

Instruction Following Codellama 70b Instruct leads

Codellama 70b Instruct: 51.9 (#293), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkCodellama 70b InstructQwen1.5 4b Chat
LMArena Instruction Following1024978

Long Context Not comparable

Codellama 70b Instruct: —, Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkCodellama 70b InstructQwen1.5 4b Chat
LMArena Longer Query—988

Writing & Preference Codellama 70b Instruct leads

Codellama 70b Instruct: 33.4 (#277), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkCodellama 70b InstructQwen1.5 4b Chat
LMArena Text1057997
LMArena Creative Writing—969
LMArena Multi-Turn—977

Frequently asked questions

Is Codellama 70b Instruct better than Qwen1.5 4b Chat?

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 28.8 on the Noometry Index.

Is Codellama 70b Instruct or Qwen1.5 4b Chat better for coding?

Codellama 70b Instruct scores higher on coding benchmarks: 37.6 versus 29.1 in the Noometry coding category.

How many benchmarks do Codellama 70b Instruct and Qwen1.5 4b Chat share?

4 benchmarks have published results for both models. Codellama 70b Instruct has 7 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper