Model comparison

Mixtral 8x7B vs Olmo 3.1 32b Think

Olmo 3.1 32b Think is the stronger model overall, scoring 37.9 to 27.1 on the Noometry Index.

Last verified . 15 shared benchmarks.

Mixtral 8x7B Mistral AI

27.1

Rank #334 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Mixtral 8x7B scores higher in 0 categories and Olmo 3.1 32b Think in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Olmo 3.1 32b Think leads 35.7 to 11.0.

Side by side

Mixtral 8x7B and Olmo 3.1 32b Think specifications
Mixtral 8x7BOlmo 3.1 32b Think
ProviderMistral AIAllen Institute for AI (Ai2)
Noometry Index27.137.9
Released2023-12-11—
WeightsOpenOpen
Context window32K—
Max output32K—
Input $ / M tokens$0.70—
Output $ / M tokens$0.70—
Results tracked3815

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Think leads

Mixtral 8x7B: 32.8 (#269), Olmo 3.1 32b Think: 37.7 (#189)

Coding benchmarks
BenchmarkMixtral 8x7BOlmo 3.1 32b Think
LMArena Coding11261291
HumanEval+39.6%—
MBPP+49.7%—

Reasoning Olmo 3.1 32b Think leads

Mixtral 8x7B: 18.2 (#285), Olmo 3.1 32b Think: 25.2 (#150)

Reasoning benchmarks
BenchmarkMixtral 8x7BOlmo 3.1 32b Think
LMArena Hard Prompts11151272
DTBench49.6%—
Adversarial NLI55.2%—
Epoch Capabilities Index118.47—
ForecastBench56.3—
HellaSwag86.7%—
PIQA83.6%—
WinoGrande77.2%—

Math Olmo 3.1 32b Think leads

Mixtral 8x7B: 18.8 (#289), Olmo 3.1 32b Think: 36.3 (#168)

Math benchmarks
BenchmarkMixtral 8x7BOlmo 3.1 32b Think
LMArena Math11471305
Omni-MATH10.5%—
MATH Level 510%—
GSM8K74.4%—

Knowledge Olmo 3.1 32b Think leads

Mixtral 8x7B: 11.0 (#301), Olmo 3.1 32b Think: 35.7 (#181)

Knowledge benchmarks
BenchmarkMixtral 8x7BOlmo 3.1 32b Think
LMArena Expert10881295
GPQA Diamond30.6%—
MMLU-Pro33.5%—
GPQA (HELM)29.6%—
ARC (AI2) Challenge87.3%—
MMLU70.6%—
OpenBookQA85.8%—
TriviaQA82.2%—

Multilingual Olmo 3.1 32b Think leads

Mixtral 8x7B: 29.6 (#266), Olmo 3.1 32b Think: 38.1 (#231)

Multilingual benchmarks
BenchmarkMixtral 8x7BOlmo 3.1 32b Think
LMArena Non-English10771209
LMArena Chinese10551242
LMArena French11661260
LMArena German11141262
LMArena Russian10901193
LMArena Spanish11111289
LMArena Japanese931—
LMArena Korean968—

Instruction Following Olmo 3.1 32b Think leads

Mixtral 8x7B: 51.0 (#297), Olmo 3.1 32b Think: 65.6 (#218)

Instruction Following benchmarks
BenchmarkMixtral 8x7BOlmo 3.1 32b Think
LMArena Instruction Following11091247
IFEval57.5%—

Long Context Olmo 3.1 32b Think leads

Mixtral 8x7B: 33.4 (#260), Olmo 3.1 32b Think: 38.6 (#195)

Long Context benchmarks
BenchmarkMixtral 8x7BOlmo 3.1 32b Think
LMArena Longer Query11031272

Writing & Preference Olmo 3.1 32b Think leads

Mixtral 8x7B: 34.2 (#270), Olmo 3.1 32b Think: 46.2 (#220)

Writing & Preference benchmarks
BenchmarkMixtral 8x7BOlmo 3.1 32b Think
LMArena Text11321272
LMArena Creative Writing11091226
LMArena Multi-Turn11151252
WildBench67.3%—

Frequently asked questions

Is Mixtral 8x7B better than Olmo 3.1 32b Think?

Olmo 3.1 32b Think is the stronger model overall, scoring 37.9 to 27.1 on the Noometry Index.

Is Mixtral 8x7B or Olmo 3.1 32b Think better for coding?

Olmo 3.1 32b Think scores higher on coding benchmarks: 37.7 versus 32.8 in the Noometry coding category.

How many benchmarks do Mixtral 8x7B and Olmo 3.1 32b Think share?

15 benchmarks have published results for both models. Mixtral 8x7B has 38 scored results on Noometry and Olmo 3.1 32b Think has 15.

Related comparisons

Go deeper