Model comparison

Gemma 3n E4b IT vs o3

o3 is the stronger model overall, scoring 47.5 to 37.3 on the Noometry Index.

Last verified . 18 shared benchmarks.

Gemma 3n E4b IT Google

37.3

Rank #206 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Gemma 3n E4b IT scores higher in 0 categories and o3 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o3 leads 54.6 to 34.2.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 31.5% for Gemma 3n E4b IT and 67.6% for o3.
  • Gemma 3n E4b IT has downloadable open weights; the other is API-only.

Side by side

Gemma 3n E4b IT and o3 specifications
Gemma 3n E4b ITo3
ProviderGoogleOpenAI
Noometry Index37.347.5
Released—2025-04-16
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked1863

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Gemma 3n E4b IT: 37.0 (#198), o3: 46.8 (#64)

Coding benchmarks
BenchmarkGemma 3n E4b ITo3
LMArena Coding12681408
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
GSO—8.8%
WeirdML—52.4%
CadEval—74%
ALE-Bench—933.55

Agentic & Tool Use Not comparable

Gemma 3n E4b IT: —, o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkGemma 3n E4b ITo3
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

Gemma 3n E4b IT: 19.9 (#247), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkGemma 3n E4b ITo3
Kagi LLM Benchmark31.5%67.6%
LMArena Hard Prompts12841402
ARC-AGI-2—6.5%
SimpleBench—53.1%
ARC-AGI-1—60.8%
CritPt—1.4%
Chess Puzzles—38%
EnigmaEval—13.1%
Mystery Game Puzzles—29%
DTBench—84.8%
LMCA—39.7%
Epoch Capabilities Index—146.86
ForecastBench—62.5

Math o3 leads

Gemma 3n E4b IT: 35.1 (#188), o3: 50.2 (#58)

Math benchmarks
BenchmarkGemma 3n E4b ITo3
LMArena Math12511426
FrontierMath (Tiers 1-3)—33.3%
OTIS Mock AIME 2024-2025—84.4%
Omni-MATH—71.4%
MATH Level 5—97.8%
FrontierMath (Feb 2025 set)—18.7%
FrontierMath Tier 4 (v1)—2.1%

Knowledge o3 leads

Gemma 3n E4b IT: 34.2 (#198), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkGemma 3n E4b ITo3
LMArena Expert12461402
GPQA Diamond—81.8%
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
Confabulations—14.4%
GPQA (HELM)—75.3%

Multimodal Not comparable

Gemma 3n E4b IT: —, o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkGemma 3n E4b ITo3
LMArena Vision—1214
GeoBench—74%
VPCT—52%

Multilingual o3 leads

Gemma 3n E4b IT: 43.4 (#183), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkGemma 3n E4b ITo3
LMArena Non-English12851401
LMArena Chinese13091437
LMArena French13301430
LMArena German13111420
LMArena Japanese12721403
LMArena Korean12591370
LMArena Russian12881406
LMArena Spanish13051395

Instruction Following o3 leads

Gemma 3n E4b IT: 66.1 (#210), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkGemma 3n E4b ITo3
LMArena Instruction Following12551368
IFEval—86.9%

Long Context o3 leads

Gemma 3n E4b IT: 38.7 (#191), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkGemma 3n E4b ITo3
LMArena Longer Query12761372
Fiction.LiveBench—88.9%
CL-bench—17.8%

Writing & Preference o3 leads

Gemma 3n E4b IT: 50.1 (#186), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkGemma 3n E4b ITo3
LMArena Text13061410
LMArena Creative Writing12871359
LMArena Multi-Turn12761405
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
WildBench—86.1%

Frequently asked questions

Is Gemma 3n E4b IT better than o3?

o3 is the stronger model overall, scoring 47.5 to 37.3 on the Noometry Index.

Is Gemma 3n E4b IT or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 37.0 in the Noometry coding category.

How many benchmarks do Gemma 3n E4b IT and o3 share?

18 benchmarks have published results for both models. Gemma 3n E4b IT has 18 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper