Model comparison

Gemma 2 27B vs Llama-3.3-70B-Instruct

Llama-3.3-70B-Instruct is the stronger model overall, scoring 30.6 to 29.4 on the Noometry Index.

Last verified . 34 shared benchmarks.

Gemma 2 27B Google

29.4

Rank #312 Confirmed

Llama-3.3-70B-Instruct Meta

30.6

Rank #291 Confirmed

Summary

  • They share 34 benchmarks with published results for both. Gemma 2 27B scores higher in 3 categories and Llama-3.3-70B-Instruct in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Llama-3.3-70B-Instruct leads 30.6 to 19.0.
  • The biggest single-benchmark swing is LiveBench Instruction Following: 58.1% for Gemma 2 27B and 82.7% for Llama-3.3-70B-Instruct.
  • Llama-3.3-70B-Instruct is cheaper at $0.10 / $0.32 per million input/output tokens, against $0.65 / $0.65 for Gemma 2 27B.
  • Llama-3.3-70B-Instruct accepts more context: 128K tokens versus 8K.

Side by side

Gemma 2 27B and Llama-3.3-70B-Instruct specifications
Gemma 2 27BLlama-3.3-70B-Instruct
ProviderGoogleMeta
Noometry Index29.430.6
Released2024-06-242024-12-06
WeightsOpenOpen
Context window8K128K
Max output2K4K
Input $ / M tokens$0.65$0.10
Output $ / M tokens$0.65$0.32
Results tracked3443

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 2 27B leads

Gemma 2 27B: 34.1 (#246), Llama-3.3-70B-Instruct: 31.0 (#290)

Coding benchmarks
BenchmarkGemma 2 27BLlama-3.3-70B-Instruct
BigCodeBench Instruct42.8%46.9%
LiveBench Coding36%36.6%
LMArena Coding12111268
BigCodeBench Complete52.5%57.5%
SciCode—26%
WeirdML—14.4%

Agentic & Tool Use Not comparable

Gemma 2 27B: —, Llama-3.3-70B-Instruct: 25.8 (#105)

Agentic & Tool Use benchmarks
BenchmarkGemma 2 27BLlama-3.3-70B-Instruct
Berkeley Function Calling Leaderboard—31.9%
BALROG—23%

Reasoning Gemma 2 27B leads

Gemma 2 27B: 15.3 (#315), Llama-3.3-70B-Instruct: 14.1 (#327)

Reasoning benchmarks
BenchmarkGemma 2 27BLlama-3.3-70B-Instruct
LiveBench Reasoning28.1%50.8%
LMArena Hard Prompts11981257
DTBench48%59.5%
LiveBench Data Analysis47.9%49.5%
LMCA7.1%17.5%
Epoch Capabilities Index122.08127.33
LiveBench38.2%50.2%
SimpleBench—19.9%
CritPt—0%
ForecastBench—58.6

Math Llama-3.3-70B-Instruct leads

Gemma 2 27B: 10.7 (#311), Llama-3.3-70B-Instruct: 15.3 (#298)

Math benchmarks
BenchmarkGemma 2 27BLlama-3.3-70B-Instruct
OTIS Mock AIME 2024-20251.4%5.1%
LiveBench Math26.5%42.2%
LMArena Math12121267
MATH Level 527.9%41.6%

Knowledge Llama-3.3-70B-Instruct leads

Gemma 2 27B: 19.0 (#280), Llama-3.3-70B-Instruct: 30.6 (#226)

Knowledge benchmarks
BenchmarkGemma 2 27BLlama-3.3-70B-Instruct
GPQA Diamond36.5%47.4%
Confabulations27.1%22.8%
LMArena Expert11721225
MMLU75.7%86.3%
Vectara Hallucination Rate—4.1%

Multilingual Llama-3.3-70B-Instruct leads

Gemma 2 27B: 38.6 (#226), Llama-3.3-70B-Instruct: 39.9 (#220)

Multilingual benchmarks
BenchmarkGemma 2 27BLlama-3.3-70B-Instruct
LMArena Non-English12171236
LMArena Chinese12211217
LMArena French12471281
LMArena German12091251
LMArena Japanese11751150
LMArena Korean11741143
LMArena Russian12341252
LMArena Spanish12281270

Instruction Following Llama-3.3-70B-Instruct leads

Gemma 2 27B: 60.5 (#249), Llama-3.3-70B-Instruct: 71.1 (#157)

Instruction Following benchmarks
BenchmarkGemma 2 27BLlama-3.3-70B-Instruct
LiveBench Instruction Following58.1%82.7%
LMArena Instruction Following12061242

Long Context Gemma 2 27B leads

Gemma 2 27B: 37.3 (#218), Llama-3.3-70B-Instruct: 26.4 (#295)

Long Context benchmarks
BenchmarkGemma 2 27BLlama-3.3-70B-Instruct
LMArena Longer Query12311256
Fiction.LiveBench—33.3%

Writing & Preference Llama-3.3-70B-Instruct leads

Gemma 2 27B: 44.2 (#225), Llama-3.3-70B-Instruct: 47.6 (#207)

Writing & Preference benchmarks
BenchmarkGemma 2 27BLlama-3.3-70B-Instruct
LMArena Text12311274
LMArena Creative Writing12411250
LMArena Multi-Turn12241280
LiveBench Language32.6%39.2%

Frequently asked questions

Is Gemma 2 27B better than Llama-3.3-70B-Instruct?

Llama-3.3-70B-Instruct is the stronger model overall, scoring 30.6 to 29.4 on the Noometry Index.

Which is cheaper, Gemma 2 27B or Llama-3.3-70B-Instruct?

Llama-3.3-70B-Instruct is cheaper. It lists at $0.10 per million input tokens and $0.32 per million output tokens; Gemma 2 27B lists at $0.65 and $0.65.

Is Gemma 2 27B or Llama-3.3-70B-Instruct better for coding?

Gemma 2 27B scores higher on coding benchmarks: 34.1 versus 31.0 in the Noometry coding category.

Which has the bigger context window?

Llama-3.3-70B-Instruct does, with 128K tokens against 8K.

How many benchmarks do Gemma 2 27B and Llama-3.3-70B-Instruct share?

34 benchmarks have published results for both models. Gemma 2 27B has 34 scored results on Noometry and Llama-3.3-70B-Instruct has 43.

Related comparisons

Go deeper