Model comparison

Hy4 preview vs Mistral Small 3.1

Hy4 preview is the stronger model overall, scoring 45.3 to 31.7 on the Noometry Index. Mistral Small 3.1 costs 3.1× less per token, which makes it the better buy when Hy4 preview's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Hy4 preview Tencent

45.3

Rank #73 Reported

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • The widest gap is in math, where Hy4 preview leads 55.7 to 14.7.
  • Mistral Small 3.1 is cheaper at $0.35 / $0.56 per million input/output tokens, against $0.83 / $2.50 for Hy4 preview.
  • Hy4 preview accepts more context: 1.05M tokens versus 128K.

Side by side

Hy4 preview and Mistral Small 3.1 specifications
Hy4 previewMistral Small 3.1
ProviderTencentMistral AI
Noometry Index45.331.7
Released2026-08-282025-03-17
WeightsOpenOpen
Context window1.05M128K
Max output64K102K
Input $ / M tokens$0.83$0.35
Output $ / M tokens$2.50$0.56
Results tracked328

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy4 preview leads

Hy4 preview: 51.6 (#38), Mistral Small 3.1: 38.3 (#179)

Coding benchmarks
BenchmarkHy4 previewMistral Small 3.1
LMArena WebDev1632—
LMArena Coding—1309

Reasoning Hy4 preview leads

Hy4 preview: 31.9 (#79), Mistral Small 3.1: 19.7 (#254)

Reasoning benchmarks
BenchmarkHy4 previewMistral Small 3.1
NYT Connections (extended)68.2%—
Chess Puzzles—1%
LMArena Hard Prompts—1278
Epoch Capabilities Index—127.48

Math Hy4 preview leads

Hy4 preview: 55.7 (#42), Mistral Small 3.1: 14.7 (#301)

Math benchmarks
BenchmarkHy4 previewMistral Small 3.1
OTIS Mock AIME 2024-2025—3.9%
ProofBench75%—
Omni-MATH—24.8%
LMArena Math—1262

Knowledge Not comparable

Hy4 preview: —, Mistral Small 3.1: 22.6 (#271)

Knowledge benchmarks
BenchmarkHy4 previewMistral Small 3.1
GPQA Diamond—41.9%
MMLU-Pro—61%
GPQA (HELM)—39.2%
LMArena Expert—1257

Multimodal Not comparable

Hy4 preview: —, Mistral Small 3.1: 33.2 (#99)

Multimodal benchmarks
BenchmarkHy4 previewMistral Small 3.1
LMArena Vision—1136

Multilingual Not comparable

Hy4 preview: —, Mistral Small 3.1: 41.2 (#209)

Multilingual benchmarks
BenchmarkHy4 previewMistral Small 3.1
LMArena Non-English—1255
LMArena Chinese—1253
LMArena French—1273
LMArena German—1266
LMArena Japanese—1208
LMArena Korean—1206
LMArena Russian—1263
LMArena Spanish—1283

Instruction Following Not comparable

Hy4 preview: —, Mistral Small 3.1: 63.6 (#230)

Instruction Following benchmarks
BenchmarkHy4 previewMistral Small 3.1
IFEval—75%
LMArena Instruction Following—1264

Long Context Not comparable

Hy4 preview: —, Mistral Small 3.1: 39.5 (#178)

Long Context benchmarks
BenchmarkHy4 previewMistral Small 3.1
LMArena Longer Query—1299

Writing & Preference Not comparable

Hy4 preview: —, Mistral Small 3.1: 37.0 (#259)

Writing & Preference benchmarks
BenchmarkHy4 previewMistral Small 3.1
LMArena Text—1277
LMArena Creative Writing—1253
EQ-Bench Creative Writing—761
WildBench—78.8%
LMArena Multi-Turn—1270

Frequently asked questions

Is Hy4 preview better than Mistral Small 3.1?

Hy4 preview is the stronger model overall, scoring 45.3 to 31.7 on the Noometry Index. Mistral Small 3.1 costs 3.1× less per token, which makes it the better buy when Hy4 preview's lead doesn't matter for your workload.

Which is cheaper, Hy4 preview or Mistral Small 3.1?

Mistral Small 3.1 is cheaper. It lists at $0.35 per million input tokens and $0.56 per million output tokens; Hy4 preview lists at $0.83 and $2.50.

Is Hy4 preview or Mistral Small 3.1 better for coding?

Hy4 preview scores higher on coding benchmarks: 51.6 versus 38.3 in the Noometry coding category.

Which has the bigger context window?

Hy4 preview does, with 1.05M tokens against 128K.

How many benchmarks do Hy4 preview and Mistral Small 3.1 share?

0 benchmarks have published results for both models. Hy4 preview has 3 scored results on Noometry and Mistral Small 3.1 has 28.

Related comparisons

Go deeper