Model comparison

Mistral vs Qwen3.7 Flash

Qwen3.7 Flash is the stronger model overall, scoring 39.9 to 29.9 on the Noometry Index.

Last verified . 0 shared benchmarks.

Mistral Mistral AI

29.9

Rank #303 Confirmed

Qwen3.7 Flash Alibaba (Qwen)

39.9

Rank #156 Confirmed

Summary

  • The widest gap is in knowledge, where Qwen3.7 Flash leads 48.9 to 16.6.

Side by side

Mistral and Qwen3.7 Flash specifications
MistralQwen3.7 Flash
ProviderMistral AIAlibaba (Qwen)
Noometry Index29.939.9
Released—2026-07-15
WeightsProprietaryProprietary
Context window—1M
Max output—131K
Input $ / M tokens—$0.03
Output $ / M tokens—$0.13
Results tracked227

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral: 33.8 (#250), Qwen3.7 Flash: —

Coding benchmarks
BenchmarkMistralQwen3.7 Flash
LMArena Coding1162—

Reasoning Qwen3.7 Flash leads

Mistral: 22.2 (#200), Qwen3.7 Flash: 28.2 (#108)

Reasoning benchmarks
BenchmarkMistralQwen3.7 Flash
NYT Connections (extended)—43.8%
Chess Puzzles—23%
LMArena Hard Prompts1149—
Mystery Game Puzzles—15%
Epoch Capabilities Index—144.64

Math Qwen3.7 Flash leads

Mistral: 22.3 (#278), Qwen3.7 Flash: 38.3 (#140)

Math benchmarks
BenchmarkMistralQwen3.7 Flash
FrontierMath (Tiers 1-3)—19.3%
OTIS Mock AIME 2024-2025—86.7%
Omni-MATH7.2%—
LMArena Math1180—

Knowledge Qwen3.7 Flash leads

Mistral: 16.6 (#288), Qwen3.7 Flash: 48.9 (#75)

Knowledge benchmarks
BenchmarkMistralQwen3.7 Flash
GPQA Diamond—82.3%
MMLU-Pro27.7%—
GPQA (HELM)30.3%—
LMArena Expert1125—

Multilingual Not comparable

Mistral: 32.8 (#254), Qwen3.7 Flash: —

Multilingual benchmarks
BenchmarkMistralQwen3.7 Flash
LMArena Non-English1129—
LMArena Chinese1109—
LMArena French1180—
LMArena German1155—
LMArena Japanese1013—
LMArena Korean1032—
LMArena Russian1168—
LMArena Spanish1143—

Instruction Following Not comparable

Mistral: 52.6 (#288), Qwen3.7 Flash: —

Instruction Following benchmarks
BenchmarkMistralQwen3.7 Flash
IFEval56.8%—
LMArena Instruction Following1152—

Long Context Not comparable

Mistral: 35.0 (#245), Qwen3.7 Flash: —

Long Context benchmarks
BenchmarkMistralQwen3.7 Flash
LMArena Longer Query1153—

Writing & Preference Not comparable

Mistral: 37.0 (#260), Qwen3.7 Flash: —

Writing & Preference benchmarks
BenchmarkMistralQwen3.7 Flash
LMArena Text1165—
LMArena Creative Writing1158—
WildBench66%—
LMArena Multi-Turn1147—

Frequently asked questions

Is Mistral better than Qwen3.7 Flash?

Qwen3.7 Flash is the stronger model overall, scoring 39.9 to 29.9 on the Noometry Index.

How many benchmarks do Mistral and Qwen3.7 Flash share?

0 benchmarks have published results for both models. Mistral has 22 scored results on Noometry and Qwen3.7 Flash has 7.

Related comparisons

Go deeper