Model comparison

Mistral Nemo vs Qwen2.5 7B Instruct

Qwen2.5 7B Instruct is the stronger model overall, scoring 29.0 to 26.4 on the Noometry Index. Mistral Nemo costs 2.0× less per token, which makes it the better buy when Qwen2.5 7B Instruct's lead doesn't matter for your workload.

Last verified . 4 shared benchmarks.

Mistral Nemo Mistral AI

26.4

Rank #337 Confirmed

Qwen2.5 7B Instruct Alibaba (Qwen)

29.0

Rank #320 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Mistral Nemo scores higher in 2 categories and Qwen2.5 7B Instruct in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen2.5 7B Instruct leads 48.8 to 28.5.
  • The biggest single-benchmark swing is BALROG: 17.6% for Mistral Nemo and 7.8% for Qwen2.5 7B Instruct.
  • Mistral Nemo is cheaper at $0.15 / $0.15 per million input/output tokens, against $0.17 / $0.70 for Qwen2.5 7B Instruct.
  • Qwen2.5 7B Instruct accepts more context: 131K tokens versus 128K.

Side by side

Mistral Nemo and Qwen2.5 7B Instruct specifications
Mistral NemoQwen2.5 7B Instruct
ProviderMistral AIAlibaba (Qwen)
Noometry Index26.429.0
Released2024-07-012024-09
WeightsOpenOpen
Context window128K131K
Max output128K8K
Input $ / M tokens$0.15$0.17
Output $ / M tokens$0.15$0.70
Results tracked1015

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral Nemo: —, Qwen2.5 7B Instruct: 36.5 (#208)

Coding benchmarks
BenchmarkMistral NemoQwen2.5 7B Instruct
BigCodeBench Instruct—37.6%
BigCodeBench Complete—46.1%

Agentic & Tool Use Too close to call

Mistral Nemo: 23.5 (#125), Qwen2.5 7B Instruct: 23.8 (#124)

Agentic & Tool Use benchmarks
BenchmarkMistral NemoQwen2.5 7B Instruct
BALROG17.6%7.8%
Berkeley Function Calling Leaderboard27.6%—

Reasoning Mistral Nemo leads

Mistral Nemo: 20.7 (#232), Qwen2.5 7B Instruct: 14.8 (#322)

Reasoning benchmarks
BenchmarkMistral NemoQwen2.5 7B Instruct
DTBench48.6%47.7%
Epoch Capabilities Index118.68118.51
Chess Puzzles—0%
LMCA—6.4%
PIQA83.5%—

Math Mistral Nemo leads

Mistral Nemo: 25.5 (#268), Qwen2.5 7B Instruct: 12.6 (#306)

Math benchmarks
BenchmarkMistral NemoQwen2.5 7B Instruct
OTIS Mock AIME 2024-2025—2.5%
Omni-MATH—29.4%
MATH Level 510.8%—
GSM8K84.2%—

Knowledge Qwen2.5 7B Instruct leads

Mistral Nemo: 12.3 (#298), Qwen2.5 7B Instruct: 17.0 (#286)

Knowledge benchmarks
BenchmarkMistral NemoQwen2.5 7B Instruct
GPQA Diamond29.9%35.5%
MMLU-Pro—53.9%
GPQA (HELM)—34.1%
BoolQ82.5%—
MMLU—72.9%

Instruction Following Not comparable

Mistral Nemo: —, Qwen2.5 7B Instruct: 63.2 (#231)

Instruction Following benchmarks
BenchmarkMistral NemoQwen2.5 7B Instruct
IFEval—74.1%

Writing & Preference Qwen2.5 7B Instruct leads

Mistral Nemo: 28.5 (#296), Qwen2.5 7B Instruct: 48.8 (#195)

Writing & Preference benchmarks
BenchmarkMistral NemoQwen2.5 7B Instruct
EQ-Bench Creative Writing881—
WildBench—73.1%

Frequently asked questions

Is Mistral Nemo better than Qwen2.5 7B Instruct?

Qwen2.5 7B Instruct is the stronger model overall, scoring 29.0 to 26.4 on the Noometry Index. Mistral Nemo costs 2.0× less per token, which makes it the better buy when Qwen2.5 7B Instruct's lead doesn't matter for your workload.

Which is cheaper, Mistral Nemo or Qwen2.5 7B Instruct?

Mistral Nemo is cheaper. It lists at $0.15 per million input tokens and $0.15 per million output tokens; Qwen2.5 7B Instruct lists at $0.17 and $0.70.

Which has the bigger context window?

Qwen2.5 7B Instruct does, with 131K tokens against 128K.

How many benchmarks do Mistral Nemo and Qwen2.5 7B Instruct share?

4 benchmarks have published results for both models. Mistral Nemo has 10 scored results on Noometry and Qwen2.5 7B Instruct has 15.

Related comparisons

Go deeper