Model comparison
Mistral Nemo vs Olmo 2 0325 32b Instruct
Olmo 2 0325 32b Instruct is the stronger model overall, scoring 32.7 to 26.4 on the Noometry Index.
Last verified . 0 shared benchmarks.
Summary
- The widest gap is in writing & preference, where Olmo 2 0325 32b Instruct leads 42.1 to 28.5.
Side by side
| Mistral Nemo | Olmo 2 0325 32b Instruct | |
|---|---|---|
| Provider | Mistral AI | Allen Institute for AI (Ai2) |
| Noometry Index | 26.4 | 32.7 |
| Released | 2024-07-01 | — |
| Weights | Open | Open |
| Context window | 128K | — |
| Max output | 128K | — |
| Input $ / M tokens | $0.15 | — |
| Output $ / M tokens | $0.15 | — |
| Results tracked | 10 | 16 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Not comparable
Mistral Nemo: —, Olmo 2 0325 32b Instruct: 35.2 (#227)
| Benchmark | Mistral Nemo | Olmo 2 0325 32b Instruct |
|---|---|---|
| LMArena Coding | — | 1210 |
Agentic & Tool Use Not comparable
Mistral Nemo: 23.5 (#125), Olmo 2 0325 32b Instruct: —
| Benchmark | Mistral Nemo | Olmo 2 0325 32b Instruct |
|---|---|---|
| Berkeley Function Calling Leaderboard | 27.6% | — |
| BALROG | 17.6% | — |
Reasoning Olmo 2 0325 32b Instruct leads
Mistral Nemo: 20.7 (#232), Olmo 2 0325 32b Instruct: 23.6 (#175)
| Benchmark | Mistral Nemo | Olmo 2 0325 32b Instruct |
|---|---|---|
| LMArena Hard Prompts | — | 1208 |
| DTBench | 48.6% | — |
| Epoch Capabilities Index | 118.68 | — |
| PIQA | 83.5% | — |
Math Olmo 2 0325 32b Instruct leads
Mistral Nemo: 25.5 (#268), Olmo 2 0325 32b Instruct: 26.8 (#255)
| Benchmark | Mistral Nemo | Olmo 2 0325 32b Instruct |
|---|---|---|
| Omni-MATH | — | 16.1% |
| LMArena Math | — | 1208 |
| MATH Level 5 | 10.8% | — |
| GSM8K | 84.2% | — |
Knowledge Olmo 2 0325 32b Instruct leads
Mistral Nemo: 12.3 (#298), Olmo 2 0325 32b Instruct: 19.5 (#279)
| Benchmark | Mistral Nemo | Olmo 2 0325 32b Instruct |
|---|---|---|
| GPQA Diamond | 29.9% | — |
| MMLU-Pro | — | 41.4% |
| GPQA (HELM) | — | 28.7% |
| BoolQ | 82.5% | — |
Multilingual Not comparable
Mistral Nemo: —, Olmo 2 0325 32b Instruct: 34.8 (#248)
| Benchmark | Mistral Nemo | Olmo 2 0325 32b Instruct |
|---|---|---|
| LMArena Non-English | — | 1160 |
| LMArena Chinese | — | 1192 |
| LMArena Russian | — | 1187 |
Instruction Following Not comparable
Mistral Nemo: —, Olmo 2 0325 32b Instruct: 61.5 (#244)
| Benchmark | Mistral Nemo | Olmo 2 0325 32b Instruct |
|---|---|---|
| IFEval | — | 78% |
| LMArena Instruction Following | — | 1186 |
Long Context Not comparable
Mistral Nemo: —, Olmo 2 0325 32b Instruct: 36.2 (#234)
| Benchmark | Mistral Nemo | Olmo 2 0325 32b Instruct |
|---|---|---|
| LMArena Longer Query | — | 1194 |
Writing & Preference Olmo 2 0325 32b Instruct leads
Mistral Nemo: 28.5 (#296), Olmo 2 0325 32b Instruct: 42.1 (#236)
| Benchmark | Mistral Nemo | Olmo 2 0325 32b Instruct |
|---|---|---|
| LMArena Text | — | 1218 |
| LMArena Creative Writing | — | 1199 |
| EQ-Bench Creative Writing | 881 | — |
| WildBench | — | 73.4% |
| LMArena Multi-Turn | — | 1221 |
Frequently asked questions
Is Mistral Nemo better than Olmo 2 0325 32b Instruct?
Olmo 2 0325 32b Instruct is the stronger model overall, scoring 32.7 to 26.4 on the Noometry Index.
How many benchmarks do Mistral Nemo and Olmo 2 0325 32b Instruct share?
0 benchmarks have published results for both models. Mistral Nemo has 10 scored results on Noometry and Olmo 2 0325 32b Instruct has 16.