Model comparison

Mistral 7B vs Qwen Turbo

Qwen Turbo is the stronger model overall, scoring 27.1 to 23.0 on the Noometry Index.

Last verified . 3 shared benchmarks.

Mistral 7B Mistral AI

23.0

Rank #351 Confirmed

Qwen Turbo Alibaba (Qwen)

27.1

Rank #335 Reported

Summary

  • They share 3 benchmarks with published results for both. Mistral 7B scores higher in 0 categories and Qwen Turbo in 2 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen Turbo leads 22.2 to 7.4.
  • The biggest single-benchmark swing is MATH Level 5: 3.7% for Mistral 7B and 56.2% for Qwen Turbo.
  • Qwen Turbo is cheaper at $0.05 / $0.20 per million input/output tokens, against $0.25 / $0.25 for Mistral 7B.
  • Qwen Turbo accepts more context: 1M tokens versus 8K.
  • Mistral 7B has downloadable open weights; the other is API-only.

Side by side

Mistral 7B and Qwen Turbo specifications
Mistral 7BQwen Turbo
ProviderMistral AIAlibaba (Qwen)
Noometry Index23.027.1
Released2023-09-272024-11-01
WeightsOpenProprietary
Context window8K1M
Max output8K16K
Input $ / M tokens$0.25$0.05
Output $ / M tokens$0.25$0.20
Results tracked373

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral 7B: 26.4 (#326), Qwen Turbo: —

Coding benchmarks
BenchmarkMistral 7BQwen Turbo
BigCodeBench Instruct19.5%—
LMArena Coding1082—
BigCodeBench Complete27.3%—
HumanEval+36%—
MBPP+42.1%—

Reasoning Not comparable

Mistral 7B: 13.1 (#336), Qwen Turbo: —

Reasoning benchmarks
BenchmarkMistral 7BQwen Turbo
Chess Puzzles0%—
LMArena Hard Prompts1067—
DTBench42.5%—
Adversarial NLI47.1%—
BIG-Bench Hard56.1%—
Epoch Capabilities Index112.21—
HellaSwag81%—
PIQA83%—
WinoGrande75.3%—

Math Qwen Turbo leads

Mistral 7B: 8.1 (#325), Qwen Turbo: 15.3 (#297)

Math benchmarks
BenchmarkMistral 7BQwen Turbo
OTIS Mock AIME 2024-20250.3%6.1%
MATH Level 53.7%56.2%
LMArena Math1085—
GSM8K54.4%—

Knowledge Qwen Turbo leads

Mistral 7B: 7.4 (#311), Qwen Turbo: 22.2 (#272)

Knowledge benchmarks
BenchmarkMistral 7BQwen Turbo
GPQA Diamond15.2%41.8%
LMArena Expert1036—
ARC (AI2) Challenge78.6%—
BoolQ87.4%—
MMLU62.5%—
OpenBookQA79.8%—
TriviaQA75.2%—

Multilingual Not comparable

Mistral 7B: 25.8 (#283), Qwen Turbo: —

Multilingual benchmarks
BenchmarkMistral 7BQwen Turbo
LMArena Non-English1012—
LMArena Chinese1009—
LMArena French1037—
LMArena German987—
LMArena Japanese878—
LMArena Russian1018—
LMArena Spanish1026—

Instruction Following Not comparable

Mistral 7B: 54.2 (#280), Qwen Turbo: —

Instruction Following benchmarks
BenchmarkMistral 7BQwen Turbo
LMArena Instruction Following1060—

Long Context Not comparable

Mistral 7B: 32.2 (#271), Qwen Turbo: —

Long Context benchmarks
BenchmarkMistral 7BQwen Turbo
LMArena Longer Query1060—

Writing & Preference Not comparable

Mistral 7B: 30.7 (#286), Qwen Turbo: —

Writing & Preference benchmarks
BenchmarkMistral 7BQwen Turbo
LMArena Text1090—
LMArena Creative Writing1068—
LMArena Multi-Turn1062—

Frequently asked questions

Is Mistral 7B better than Qwen Turbo?

Qwen Turbo is the stronger model overall, scoring 27.1 to 23.0 on the Noometry Index.

Which is cheaper, Mistral 7B or Qwen Turbo?

Qwen Turbo is cheaper. It lists at $0.05 per million input tokens and $0.20 per million output tokens; Mistral 7B lists at $0.25 and $0.25.

Which has the bigger context window?

Qwen Turbo does, with 1M tokens against 8K.

How many benchmarks do Mistral 7B and Qwen Turbo share?

3 benchmarks have published results for both models. Mistral 7B has 37 scored results on Noometry and Qwen Turbo has 3.

Related comparisons

Go deeper