Model comparison

Gemma 2 27B vs Llama 3.2 90B

Gemma 2 27B is the stronger model overall, scoring 29.4 to 27.5 on the Noometry Index.

Last verified . 5 shared benchmarks.

Gemma 2 27B Google

29.4

Rank #312 Confirmed

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Gemma 2 27B scores higher in 0 categories and Llama 3.2 90B in 3 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Llama 3.2 90B leads 21.7 to 15.3.
  • The biggest single-benchmark swing is MATH Level 5: 27.9% for Gemma 2 27B and 39.4% for Llama 3.2 90B.

Side by side

Gemma 2 27B and Llama 3.2 90B specifications
Gemma 2 27BLlama 3.2 90B
ProviderGoogleMeta
Noometry Index29.427.5
Released2024-06-242024-09-24
WeightsOpenOpen
Context window8K—
Max output2K—
Input $ / M tokens$0.65—
Output $ / M tokens$0.65—
Results tracked349

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Gemma 2 27B: 34.1 (#246), Llama 3.2 90B: —

Coding benchmarks
BenchmarkGemma 2 27BLlama 3.2 90B
BigCodeBench Instruct42.8%—
LiveBench Coding36%—
LMArena Coding1211—
BigCodeBench Complete52.5%—

Agentic & Tool Use Not comparable

Gemma 2 27B: —, Llama 3.2 90B: 30.0 (#80)

Agentic & Tool Use benchmarks
BenchmarkGemma 2 27BLlama 3.2 90B
BALROG—27.3%

Reasoning Llama 3.2 90B leads

Gemma 2 27B: 15.3 (#315), Llama 3.2 90B: 21.7 (#217)

Reasoning benchmarks
BenchmarkGemma 2 27BLlama 3.2 90B
Epoch Capabilities Index122.08125.5
EnigmaEval—0.4%
LiveBench Reasoning28.1%—
LMArena Hard Prompts1198—
DTBench48%—
LiveBench Data Analysis47.9%—
LMCA7.1%—
LiveBench38.2%—

Math Too close to call

Gemma 2 27B: 10.7 (#311), Llama 3.2 90B: 11.1 (#308)

Math benchmarks
BenchmarkGemma 2 27BLlama 3.2 90B
OTIS Mock AIME 2024-20251.4%2.6%
MATH Level 527.9%39.4%
LiveBench Math26.5%—
LMArena Math1212—

Knowledge Llama 3.2 90B leads

Gemma 2 27B: 19.0 (#280), Llama 3.2 90B: 21.7 (#274)

Knowledge benchmarks
BenchmarkGemma 2 27BLlama 3.2 90B
GPQA Diamond36.5%41%
MMLU75.7%80.3%
Confabulations27.1%—
LMArena Expert1172—

Multimodal Not comparable

Gemma 2 27B: —, Llama 3.2 90B: 25.4 (#124)

Multimodal benchmarks
BenchmarkGemma 2 27BLlama 3.2 90B
LMArena Vision—1000
GeoBench—52%

Multilingual Not comparable

Gemma 2 27B: 38.6 (#226), Llama 3.2 90B: —

Multilingual benchmarks
BenchmarkGemma 2 27BLlama 3.2 90B
LMArena Non-English1217—
LMArena Chinese1221—
LMArena French1247—
LMArena German1209—
LMArena Japanese1175—
LMArena Korean1174—
LMArena Russian1234—
LMArena Spanish1228—

Instruction Following Not comparable

Gemma 2 27B: 60.5 (#249), Llama 3.2 90B: —

Instruction Following benchmarks
BenchmarkGemma 2 27BLlama 3.2 90B
LiveBench Instruction Following58.1%—
LMArena Instruction Following1206—

Long Context Not comparable

Gemma 2 27B: 37.3 (#218), Llama 3.2 90B: —

Long Context benchmarks
BenchmarkGemma 2 27BLlama 3.2 90B
LMArena Longer Query1231—

Writing & Preference Not comparable

Gemma 2 27B: 44.2 (#225), Llama 3.2 90B: —

Writing & Preference benchmarks
BenchmarkGemma 2 27BLlama 3.2 90B
LMArena Text1231—
LMArena Creative Writing1241—
LMArena Multi-Turn1224—
LiveBench Language32.6%—

Frequently asked questions

Is Gemma 2 27B better than Llama 3.2 90B?

Gemma 2 27B is the stronger model overall, scoring 29.4 to 27.5 on the Noometry Index.

How many benchmarks do Gemma 2 27B and Llama 3.2 90B share?

5 benchmarks have published results for both models. Gemma 2 27B has 34 scored results on Noometry and Llama 3.2 90B has 9.

Related comparisons

Go deeper