Model comparison

Mistral 7B vs Olmo 3.1 32b Think

Olmo 3.1 32b Think is the stronger model overall, scoring 37.9 to 23.0 on the Noometry Index.

Last verified . 15 shared benchmarks.

Mistral 7B Mistral AI

23.0

Rank #351 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Mistral 7B scores higher in 0 categories and Olmo 3.1 32b Think in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Olmo 3.1 32b Think leads 35.7 to 7.4.

Side by side

Mistral 7B and Olmo 3.1 32b Think specifications
Mistral 7BOlmo 3.1 32b Think
ProviderMistral AIAllen Institute for AI (Ai2)
Noometry Index23.037.9
Released2023-09-27—
WeightsOpenOpen
Context window8K—
Max output8K—
Input $ / M tokens$0.25—
Output $ / M tokens$0.25—
Results tracked3715

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Think leads

Mistral 7B: 26.4 (#326), Olmo 3.1 32b Think: 37.7 (#189)

Coding benchmarks
BenchmarkMistral 7BOlmo 3.1 32b Think
LMArena Coding10821291
BigCodeBench Instruct19.5%—
BigCodeBench Complete27.3%—
HumanEval+36%—
MBPP+42.1%—

Reasoning Olmo 3.1 32b Think leads

Mistral 7B: 13.1 (#336), Olmo 3.1 32b Think: 25.2 (#150)

Reasoning benchmarks
BenchmarkMistral 7BOlmo 3.1 32b Think
LMArena Hard Prompts10671272
Chess Puzzles0%—
DTBench42.5%—
Adversarial NLI47.1%—
BIG-Bench Hard56.1%—
Epoch Capabilities Index112.21—
HellaSwag81%—
PIQA83%—
WinoGrande75.3%—

Math Olmo 3.1 32b Think leads

Mistral 7B: 8.1 (#325), Olmo 3.1 32b Think: 36.3 (#168)

Math benchmarks
BenchmarkMistral 7BOlmo 3.1 32b Think
LMArena Math10851305
OTIS Mock AIME 2024-20250.3%—
MATH Level 53.7%—
GSM8K54.4%—

Knowledge Olmo 3.1 32b Think leads

Mistral 7B: 7.4 (#311), Olmo 3.1 32b Think: 35.7 (#181)

Knowledge benchmarks
BenchmarkMistral 7BOlmo 3.1 32b Think
LMArena Expert10361295
GPQA Diamond15.2%—
ARC (AI2) Challenge78.6%—
BoolQ87.4%—
MMLU62.5%—
OpenBookQA79.8%—
TriviaQA75.2%—

Multilingual Olmo 3.1 32b Think leads

Mistral 7B: 25.8 (#283), Olmo 3.1 32b Think: 38.1 (#231)

Multilingual benchmarks
BenchmarkMistral 7BOlmo 3.1 32b Think
LMArena Non-English10121209
LMArena Chinese10091242
LMArena French10371260
LMArena German9871262
LMArena Russian10181193
LMArena Spanish10261289
LMArena Japanese878—

Instruction Following Olmo 3.1 32b Think leads

Mistral 7B: 54.2 (#280), Olmo 3.1 32b Think: 65.6 (#218)

Instruction Following benchmarks
BenchmarkMistral 7BOlmo 3.1 32b Think
LMArena Instruction Following10601247

Long Context Olmo 3.1 32b Think leads

Mistral 7B: 32.2 (#271), Olmo 3.1 32b Think: 38.6 (#195)

Long Context benchmarks
BenchmarkMistral 7BOlmo 3.1 32b Think
LMArena Longer Query10601272

Writing & Preference Olmo 3.1 32b Think leads

Mistral 7B: 30.7 (#286), Olmo 3.1 32b Think: 46.2 (#220)

Writing & Preference benchmarks
BenchmarkMistral 7BOlmo 3.1 32b Think
LMArena Text10901272
LMArena Creative Writing10681226
LMArena Multi-Turn10621252

Frequently asked questions

Is Mistral 7B better than Olmo 3.1 32b Think?

Olmo 3.1 32b Think is the stronger model overall, scoring 37.9 to 23.0 on the Noometry Index.

Is Mistral 7B or Olmo 3.1 32b Think better for coding?

Olmo 3.1 32b Think scores higher on coding benchmarks: 37.7 versus 26.4 in the Noometry coding category.

How many benchmarks do Mistral 7B and Olmo 3.1 32b Think share?

15 benchmarks have published results for both models. Mistral 7B has 37 scored results on Noometry and Olmo 3.1 32b Think has 15.

Related comparisons

Go deeper