Model comparison

o3 vs Olmo 7b Instruct

o3 is the stronger model overall, scoring 47.5 to 30.3 on the Noometry Index.

Last verified . 10 shared benchmarks.

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 10 benchmarks with published results for both. o3 scores higher in 6 categories and Olmo 7b Instruct in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where o3 leads 63.5 to 25.8.
  • Olmo 7b Instruct has downloadable open weights; the other is API-only.

Side by side

o3 and Olmo 7b Instruct specifications
o3Olmo 7b Instruct
ProviderOpenAIAllen Institute for AI (Ai2)
Noometry Index47.530.3
Released2025-04-16—
WeightsProprietaryOpen
Context window200K—
Max output100K—
Input $ / M tokens$2—
Output $ / M tokens$8—
Results tracked6310

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

o3: 46.8 (#64), Olmo 7b Instruct: 29.6 (#303)

Coding benchmarks
Benchmarko3Olmo 7b Instruct
LMArena Coding14081016
SWE-bench Verified62.3%—
SWE-bench Verified (bash only)58.4%—
Aider Polyglot81.3%—
GSO8.8%—
WeirdML52.4%—
CadEval74%—
ALE-Bench933.55—

Agentic & Tool Use Not comparable

o3: 34.5 (#44), Olmo 7b Instruct: —

Agentic & Tool Use benchmarks
Benchmarko3Olmo 7b Instruct
Berkeley Function Calling Leaderboard63%—
GDPval30.8%—
DeepResearch Bench45.2%—
OSWorld23%—
LMArena Search1144—
METR Time Horizons65.4%—

Reasoning o3 leads

o3: 32.0 (#78), Olmo 7b Instruct: 18.8 (#274)

Reasoning benchmarks
Benchmarko3Olmo 7b Instruct
LMArena Hard Prompts1402993
ARC-AGI-26.5%—
SimpleBench53.1%—
Kagi LLM Benchmark67.6%—
ARC-AGI-160.8%—
CritPt1.4%—
Chess Puzzles38%—
EnigmaEval13.1%—
Mystery Game Puzzles29%—
DTBench84.8%—
LMCA39.7%—
Epoch Capabilities Index146.86—
ForecastBench62.5—

Math o3 leads

o3: 50.2 (#58), Olmo 7b Instruct: 30.2 (#237)

Math benchmarks
Benchmarko3Olmo 7b Instruct
LMArena Math14261018
FrontierMath (Tiers 1-3)33.3%—
OTIS Mock AIME 2024-202584.4%—
Omni-MATH71.4%—
MATH Level 597.8%—
FrontierMath (Feb 2025 set)18.7%—
FrontierMath Tier 4 (v1)2.1%—

Knowledge Not comparable

o3: 54.6 (#52), Olmo 7b Instruct: —

Knowledge benchmarks
Benchmarko3Olmo 7b Instruct
GPQA Diamond81.8%—
Humanity's Last Exam20.3%—
SimpleQA Verified49.4%—
MMLU-Pro85.9%—
Confabulations14.4%—
GPQA (HELM)75.3%—
LMArena Expert1402—

Multimodal Not comparable

o3: 41.4 (#36), Olmo 7b Instruct: —

Multimodal benchmarks
Benchmarko3Olmo 7b Instruct
LMArena Vision1214—
GeoBench74%—
VPCT52%—

Multilingual o3 leads

o3: 51.7 (#105), Olmo 7b Instruct: 24.0 (#291)

Multilingual benchmarks
Benchmarko3Olmo 7b Instruct
LMArena Non-English1401977
LMArena Chinese14371014
LMArena Russian1406947
LMArena French1430—
LMArena German1420—
LMArena Japanese1403—
LMArena Korean1370—
LMArena Spanish1395—

Instruction Following o3 leads

o3: 72.8 (#127), Olmo 7b Instruct: 49.0 (#301)

Instruction Following benchmarks
Benchmarko3Olmo 7b Instruct
LMArena Instruction Following1368978
IFEval86.9%—

Long Context Not comparable

o3: 53.3 (#6), Olmo 7b Instruct: —

Long Context benchmarks
Benchmarko3Olmo 7b Instruct
Fiction.LiveBench88.9%—
CL-bench17.8%—
LMArena Longer Query1372—

Writing & Preference o3 leads

o3: 63.5 (#64), Olmo 7b Instruct: 25.8 (#303)

Writing & Preference benchmarks
Benchmarko3Olmo 7b Instruct
LMArena Text14101032
LMArena Creative Writing1359990
LMArena Multi-Turn14051007
Short-Story Creative Writing83.9%—
EQ-Bench Creative Writing1676—
WildBench86.1%—

Frequently asked questions

Is o3 better than Olmo 7b Instruct?

o3 is the stronger model overall, scoring 47.5 to 30.3 on the Noometry Index.

Is o3 or Olmo 7b Instruct better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 29.6 in the Noometry coding category.

How many benchmarks do o3 and Olmo 7b Instruct share?

10 benchmarks have published results for both models. o3 has 63 scored results on Noometry and Olmo 7b Instruct has 10.

Related comparisons

Go deeper