Model comparison

Claude 3.7 Sonnet vs Gemma 2B

Claude 3.7 Sonnet is the stronger model overall, scoring 39.5 to 29.6 on the Noometry Index.

Last verified . 12 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

Gemma 2B Google

29.6

Rank #307 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 6 categories and Gemma 2B in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Claude 3.7 Sonnet leads 54.4 to 24.0.
  • Gemma 2B has downloadable open weights; the other is API-only.

Side by side

Claude 3.7 Sonnet and Gemma 2B specifications
Claude 3.7 SonnetGemma 2B
ProviderAnthropicGoogle
Noometry Index39.529.6
Released2025-02-242024-02-21
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked5823

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 40.6 (#136), Gemma 2B: 29.4 (#305)

Coding benchmarks
BenchmarkClaude 3.7 SonnetGemma 2B
LMArena Coding13611010
SWE-bench Verified61%—
SWE-bench Verified (bash only)52.8%—
Aider Polyglot64.9%—
GSO3.8%—
LiveBench Coding74.5%—
CadEval54%—
HumanEval+—20.7%
MBPP+—34.1%

Agentic & Tool Use Not comparable

Claude 3.7 Sonnet: 34.1 (#50), Gemma 2B: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 SonnetGemma 2B
TheAgentCompany30.9%—
Cybench20%—
DeepResearch Bench43.6%—
OSWorld35.8%—
METR Time Horizons60%—

Reasoning Too close to call

Claude 3.7 Sonnet: 18.6 (#277), Gemma 2B: 18.8 (#275)

Reasoning benchmarks
BenchmarkClaude 3.7 SonnetGemma 2B
LMArena Hard Prompts1333989
Epoch Capabilities Index141.1694.2
ARC-AGI-20.9%—
SimpleBench46.4%—
ARC-AGI-128.6%—
EnigmaEval4.2%—
LiveBench Reasoning87.8%—
LiveBench Data Analysis74%—
BIG-Bench Hard—35.2%
ForecastBench61.8—
HellaSwag—71.4%
LiveBench76.1%—
PIQA—77.3%
WinoGrande—65.4%

Math Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 37.5 (#153), Gemma 2B: 30.0 (#239)

Math benchmarks
BenchmarkClaude 3.7 SonnetGemma 2B
LMArena Math13371009
OTIS Mock AIME 2024-202557.8%—
Omni-MATH33%—
LiveBench Math79%—
MATH Level 591.2%—
FrontierMath (Feb 2025 set)4.1%—
GSM8K—17.7%

Knowledge Not comparable

Claude 3.7 Sonnet: 39.8 (#130), Gemma 2B: —

Knowledge benchmarks
BenchmarkClaude 3.7 SonnetGemma 2B
GPQA Diamond79.7%—
Humanity's Last Exam8%—
MMLU-Pro78.4%—
Confabulations14.7%—
GPQA (HELM)60.8%—
LMArena Expert1321—
ARC (AI2) Challenge—42.1%
BoolQ—69.4%
MMLU—42.3%
TriviaQA—53.2%

Multimodal Not comparable

Claude 3.7 Sonnet: 33.7 (#95), Gemma 2B: —

Multimodal benchmarks
BenchmarkClaude 3.7 SonnetGemma 2B
LMArena Vision1169—
GeoBench68%—
VPCT39%—
SpatialViz-Bench33.9%—

Multilingual Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 44.1 (#179), Gemma 2B: 23.0 (#294)

Multilingual benchmarks
BenchmarkClaude 3.7 SonnetGemma 2B
LMArena Non-English1296958
LMArena Chinese1299986
LMArena Russian1311937
LMArena French1303—
LMArena German1301—
LMArena Japanese1267—
LMArena Korean1249—
LMArena Spanish1298—

Instruction Following Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 72.9 (#125), Gemma 2B: 48.5 (#302)

Instruction Following benchmarks
BenchmarkClaude 3.7 SonnetGemma 2B
LMArena Instruction Following1352970
LiveBench Instruction Following81.3%—
IFEval83.4%—

Long Context Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 50.3 (#10), Gemma 2B: 29.9 (#291)

Long Context benchmarks
BenchmarkClaude 3.7 SonnetGemma 2B
LMArena Longer Query1373981
Fiction.LiveBench83.3%—

Writing & Preference Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 54.4 (#150), Gemma 2B: 24.0 (#308)

Writing & Preference benchmarks
BenchmarkClaude 3.7 SonnetGemma 2B
LMArena Text13141002
LMArena Creative Writing1332987
LMArena Multi-Turn1339945
Short-Story Creative Writing81.1%—
EQ-Bench Creative Writing1412—
WildBench81.4%—
LiveBench Language59.9%—

Frequently asked questions

Is Claude 3.7 Sonnet better than Gemma 2B?

Claude 3.7 Sonnet is the stronger model overall, scoring 39.5 to 29.6 on the Noometry Index.

Is Claude 3.7 Sonnet or Gemma 2B better for coding?

Claude 3.7 Sonnet scores higher on coding benchmarks: 40.6 versus 29.4 in the Noometry coding category.

How many benchmarks do Claude 3.7 Sonnet and Gemma 2B share?

12 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and Gemma 2B has 23.

Related comparisons

Go deeper