Model comparison

Llama 4 Maverick vs Olmo 7b Instruct

Llama 4 Maverick and Olmo 7b Instruct score almost the same on the Noometry Index (30.9 vs 30.3), so choose on price, context window or the category you care about most.

Last verified . 10 shared benchmarks.

Llama 4 Maverick Meta

30.9

Rank #282 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Llama 4 Maverick scores higher in 3 categories and Olmo 7b Instruct in 3 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Llama 4 Maverick leads 71.7 to 49.0.

Side by side

Llama 4 Maverick and Olmo 7b Instruct specifications
Llama 4 MaverickOlmo 7b Instruct
ProviderMetaAllen Institute for AI (Ai2)
Noometry Index30.930.3
Released2025-04-05—
WeightsOpenOpen
Context window128K—
Max output4K—
Input $ / M tokens$0.19—
Output $ / M tokens$0.65—
Results tracked5410

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 7b Instruct leads

Llama 4 Maverick: 26.6 (#324), Olmo 7b Instruct: 29.6 (#303)

Coding benchmarks
BenchmarkLlama 4 MaverickOlmo 7b Instruct
LMArena Coding13021016
SWE-bench Verified (bash only)21%—
Aider Polyglot15.6%—
SciCode33.1%—
WeirdML24.5%—
BigCodeBench Instruct49.7%—
BigCodeBench Complete61.4%—
ALE-Bench172.97—

Agentic & Tool Use Not comparable

Llama 4 Maverick: 28.2 (#91), Olmo 7b Instruct: —

Agentic & Tool Use benchmarks
BenchmarkLlama 4 MaverickOlmo 7b Instruct
Berkeley Function Calling Leaderboard37.3%—

Reasoning Olmo 7b Instruct leads

Llama 4 Maverick: 10.1 (#342), Olmo 7b Instruct: 18.8 (#274)

Reasoning benchmarks
BenchmarkLlama 4 MaverickOlmo 7b Instruct
LMArena Hard Prompts1281993
ARC-AGI-20%—
SimpleBench27.7%—
Kagi LLM Benchmark55.9%—
NYT Connections (extended)8%—
ARC-AGI-14.4%—
CritPt0%—
EnigmaEval0.6%—
DTBench61.9%—
LMCA15.9%—
Epoch Capabilities Index132.2—
ForecastBench57.5—

Math Olmo 7b Instruct leads

Llama 4 Maverick: 26.0 (#262), Olmo 7b Instruct: 30.2 (#237)

Math benchmarks
BenchmarkLlama 4 MaverickOlmo 7b Instruct
LMArena Math12991018
OTIS Mock AIME 2024-202520.6%—
Omni-MATH42.2%—
MATH Level 573%—
FrontierMath (Feb 2025 set)0.7%—

Knowledge Not comparable

Llama 4 Maverick: 33.4 (#204), Olmo 7b Instruct: —

Knowledge benchmarks
BenchmarkLlama 4 MaverickOlmo 7b Instruct
GPQA Diamond67%—
Humanity's Last Exam5.7%—
MMLU-Pro81%—
Confabulations22.6%—
Vectara Hallucination Rate8.2%—
GPQA (HELM)65%—
LMArena Expert1259—

Multimodal Not comparable

Llama 4 Maverick: 31.6 (#105), Olmo 7b Instruct: —

Multimodal benchmarks
BenchmarkLlama 4 MaverickOlmo 7b Instruct
LMArena Vision1142—
GeoBench52%—
SpatialViz-Bench31.8%—

Multilingual Llama 4 Maverick leads

Llama 4 Maverick: 42.2 (#195), Olmo 7b Instruct: 24.0 (#291)

Multilingual benchmarks
BenchmarkLlama 4 MaverickOlmo 7b Instruct
LMArena Non-English1269977
LMArena Chinese12771014
LMArena Russian1286947
LMArena French1259—
LMArena German1291—
LMArena Japanese1207—
LMArena Korean1203—
LMArena Spanish1293—

Instruction Following Llama 4 Maverick leads

Llama 4 Maverick: 71.7 (#146), Olmo 7b Instruct: 49.0 (#301)

Instruction Following benchmarks
BenchmarkLlama 4 MaverickOlmo 7b Instruct
LMArena Instruction Following1267978
IFEval90.8%—

Long Context Not comparable

Llama 4 Maverick: 31.4 (#279), Olmo 7b Instruct: —

Long Context benchmarks
BenchmarkLlama 4 MaverickOlmo 7b Instruct
Fiction.LiveBench46.2%—
LMArena Longer Query1280—

Writing & Preference Llama 4 Maverick leads

Llama 4 Maverick: 38.8 (#252), Olmo 7b Instruct: 25.8 (#303)

Writing & Preference benchmarks
BenchmarkLlama 4 MaverickOlmo 7b Instruct
LMArena Text12871032
LMArena Creative Writing1267990
LMArena Multi-Turn12891007
Short-Story Creative Writing62%—
EQ-Bench Creative Writing860—
WildBench80%—

Frequently asked questions

Is Llama 4 Maverick better than Olmo 7b Instruct?

Llama 4 Maverick and Olmo 7b Instruct score almost the same on the Noometry Index (30.9 vs 30.3), so choose on price, context window or the category you care about most.

Is Llama 4 Maverick or Olmo 7b Instruct better for coding?

Olmo 7b Instruct scores higher on coding benchmarks: 29.6 versus 26.6 in the Noometry coding category.

How many benchmarks do Llama 4 Maverick and Olmo 7b Instruct share?

10 benchmarks have published results for both models. Llama 4 Maverick has 54 scored results on Noometry and Olmo 7b Instruct has 10.

Related comparisons

Go deeper