Model comparison

Gemma 2B vs o3

o3 is the stronger model overall, scoring 47.5 to 29.6 on the Noometry Index.

Last verified . 12 shared benchmarks.

Gemma 2B Google

29.6

Rank #307 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Gemma 2B scores higher in 0 categories and o3 in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where o3 leads 63.5 to 24.0.
  • Gemma 2B has downloadable open weights; the other is API-only.

Side by side

Gemma 2B and o3 specifications
Gemma 2Bo3
ProviderGoogleOpenAI
Noometry Index29.647.5
Released2024-02-212025-04-16
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked2363

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Gemma 2B: 29.4 (#305), o3: 46.8 (#64)

Coding benchmarks
BenchmarkGemma 2Bo3
LMArena Coding10101408
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
GSO—8.8%
WeirdML—52.4%
CadEval—74%
ALE-Bench—933.55
HumanEval+20.7%—
MBPP+34.1%—

Agentic & Tool Use Not comparable

Gemma 2B: —, o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkGemma 2Bo3
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

Gemma 2B: 18.8 (#275), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkGemma 2Bo3
LMArena Hard Prompts9891402
Epoch Capabilities Index94.2146.86
ARC-AGI-2—6.5%
SimpleBench—53.1%
Kagi LLM Benchmark—67.6%
ARC-AGI-1—60.8%
CritPt—1.4%
Chess Puzzles—38%
EnigmaEval—13.1%
Mystery Game Puzzles—29%
DTBench—84.8%
LMCA—39.7%
BIG-Bench Hard35.2%—
ForecastBench—62.5
HellaSwag71.4%—
PIQA77.3%—
WinoGrande65.4%—

Math o3 leads

Gemma 2B: 30.0 (#239), o3: 50.2 (#58)

Math benchmarks
BenchmarkGemma 2Bo3
LMArena Math10091426
FrontierMath (Tiers 1-3)—33.3%
OTIS Mock AIME 2024-2025—84.4%
Omni-MATH—71.4%
MATH Level 5—97.8%
FrontierMath (Feb 2025 set)—18.7%
FrontierMath Tier 4 (v1)—2.1%
GSM8K17.7%—

Knowledge Not comparable

Gemma 2B: —, o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkGemma 2Bo3
GPQA Diamond—81.8%
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
Confabulations—14.4%
GPQA (HELM)—75.3%
LMArena Expert—1402
ARC (AI2) Challenge42.1%—
BoolQ69.4%—
MMLU42.3%—
TriviaQA53.2%—

Multimodal Not comparable

Gemma 2B: —, o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkGemma 2Bo3
LMArena Vision—1214
GeoBench—74%
VPCT—52%

Multilingual o3 leads

Gemma 2B: 23.0 (#294), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkGemma 2Bo3
LMArena Non-English9581401
LMArena Chinese9861437
LMArena Russian9371406
LMArena French—1430
LMArena German—1420
LMArena Japanese—1403
LMArena Korean—1370
LMArena Spanish—1395

Instruction Following o3 leads

Gemma 2B: 48.5 (#302), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkGemma 2Bo3
LMArena Instruction Following9701368
IFEval—86.9%

Long Context o3 leads

Gemma 2B: 29.9 (#291), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkGemma 2Bo3
LMArena Longer Query9811372
Fiction.LiveBench—88.9%
CL-bench—17.8%

Writing & Preference o3 leads

Gemma 2B: 24.0 (#308), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkGemma 2Bo3
LMArena Text10021410
LMArena Creative Writing9871359
LMArena Multi-Turn9451405
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
WildBench—86.1%

Frequently asked questions

Is Gemma 2B better than o3?

o3 is the stronger model overall, scoring 47.5 to 29.6 on the Noometry Index.

Is Gemma 2B or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 29.4 in the Noometry coding category.

How many benchmarks do Gemma 2B and o3 share?

12 benchmarks have published results for both models. Gemma 2B has 23 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper