Model comparison

Mistral Large vs Olmo 7b Instruct

Mistral Large is the stronger model overall, scoring 31.9 to 30.3 on the Noometry Index.

Last verified . 10 shared benchmarks.

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Mistral Large scores higher in 4 categories and Olmo 7b Instruct in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Mistral Large leads 67.9 to 49.0.

Side by side

Mistral Large and Olmo 7b Instruct specifications
Mistral LargeOlmo 7b Instruct
ProviderMistral AIAllen Institute for AI (Ai2)
Noometry Index31.930.3
Released2024-02-26—
WeightsOpenOpen
Context window131K—
Max output16K—
Input $ / M tokens$2—
Output $ / M tokens$6—
Results tracked5110

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large leads

Mistral Large: 34.3 (#240), Olmo 7b Instruct: 29.6 (#303)

Coding benchmarks
BenchmarkMistral LargeOlmo 7b Instruct
LMArena Coding12771016
SciCode36.2%—
BigCodeBench Instruct30%—
LiveBench Coding47.1%—
BigCodeBench Complete38.3%—
ALE-Bench264.7—
HumanEval+62.2%—
MBPP+59.5%—

Agentic & Tool Use Not comparable

Mistral Large: 28.6 (#89), Olmo 7b Instruct: —

Agentic & Tool Use benchmarks
BenchmarkMistral LargeOlmo 7b Instruct
Berkeley Function Calling Leaderboard38.4%—

Reasoning Olmo 7b Instruct leads

Mistral Large: 15.8 (#310), Olmo 7b Instruct: 18.8 (#274)

Reasoning benchmarks
BenchmarkMistral LargeOlmo 7b Instruct
LMArena Hard Prompts1257993
SimpleBench22.5%—
CritPt0%—
LiveBench Reasoning43.5%—
DTBench65.1%—
LiveBench Data Analysis50.1%—
LMCA16.7%—
Epoch Capabilities Index128.52—
ForecastBench57.1—
LiveBench48.4%—

Math Olmo 7b Instruct leads

Mistral Large: 18.2 (#291), Olmo 7b Instruct: 30.2 (#237)

Math benchmarks
BenchmarkMistral LargeOlmo 7b Instruct
LMArena Math12621018
OTIS Mock AIME 2024-20258.5%—
Omni-MATH28.1%—
LiveBench Math42.5%—
MATH Level 550.3%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Not comparable

Mistral Large: 30.1 (#230), Olmo 7b Instruct: —

Knowledge benchmarks
BenchmarkMistral LargeOlmo 7b Instruct
GPQA Diamond51.3%—
MMLU-Pro59.9%—
Confabulations21.4%—
Vectara Hallucination Rate4.5%—
GPQA (HELM)43.5%—
LMArena Expert1232—
MMLU80%—

Multilingual Mistral Large leads

Mistral Large: 40.0 (#219), Olmo 7b Instruct: 24.0 (#291)

Multilingual benchmarks
BenchmarkMistral LargeOlmo 7b Instruct
LMArena Non-English1237977
LMArena Chinese12401014
LMArena Russian1257947
LMArena French1325—
LMArena German1254—
LMArena Japanese1188—
LMArena Korean1202—
LMArena Spanish1268—

Instruction Following Mistral Large leads

Mistral Large: 67.9 (#191), Olmo 7b Instruct: 49.0 (#301)

Instruction Following benchmarks
BenchmarkMistral LargeOlmo 7b Instruct
LMArena Instruction Following1249978
LiveBench Instruction Following67.9%—
IFEval87.7%—

Long Context Not comparable

Mistral Large: 38.3 (#199), Olmo 7b Instruct: —

Long Context benchmarks
BenchmarkMistral LargeOlmo 7b Instruct
LMArena Longer Query1261—

Writing & Preference Mistral Large leads

Mistral Large: 40.7 (#242), Olmo 7b Instruct: 25.8 (#303)

Writing & Preference benchmarks
BenchmarkMistral LargeOlmo 7b Instruct
LMArena Text12661032
LMArena Creative Writing1243990
LMArena Multi-Turn12601007
Short-Story Creative Writing69%—
EQ-Bench Creative Writing985—
WildBench80.1%—
LiveBench Language39.4%—

Frequently asked questions

Is Mistral Large better than Olmo 7b Instruct?

Mistral Large is the stronger model overall, scoring 31.9 to 30.3 on the Noometry Index.

Is Mistral Large or Olmo 7b Instruct better for coding?

Mistral Large scores higher on coding benchmarks: 34.3 versus 29.6 in the Noometry coding category.

How many benchmarks do Mistral Large and Olmo 7b Instruct share?

10 benchmarks have published results for both models. Mistral Large has 51 scored results on Noometry and Olmo 7b Instruct has 10.

Related comparisons

Go deeper