Model comparison

Hy3 vs Mistral Small 3.1

Hy3 is the stronger model overall, scoring 44.2 to 31.7 on the Noometry Index.

Last verified . 17 shared benchmarks.

Hy3 Tencent

44.2

Rank #79 Confirmed

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Hy3 scores higher in 8 categories and Mistral Small 3.1 in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Hy3 leads 40.1 to 14.7.
  • Hy3 is cheaper at $0.0825 / $0.33 per million input/output tokens, against $0.35 / $0.56 for Mistral Small 3.1.
  • Hy3 accepts more context: 262K tokens versus 128K.

Side by side

Hy3 and Mistral Small 3.1 specifications
Hy3Mistral Small 3.1
ProviderTencentMistral AI
Noometry Index44.231.7
Released2026-07-062025-03-17
WeightsOpenOpen
Context window262K128K
Max output128K102K
Input $ / M tokens$0.0825$0.35
Output $ / M tokens$0.33$0.56
Results tracked1928

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy3 leads

Hy3: 46.8 (#63), Mistral Small 3.1: 38.3 (#179)

Coding benchmarks
BenchmarkHy3Mistral Small 3.1
LMArena Coding14641309
LMArena WebDev1508—

Reasoning Hy3 leads

Hy3: 26.1 (#136), Mistral Small 3.1: 19.7 (#254)

Reasoning benchmarks
BenchmarkHy3Mistral Small 3.1
LMArena Hard Prompts14471278
NYT Connections (extended)41.2%—
Chess Puzzles—1%
Epoch Capabilities Index—127.48

Math Hy3 leads

Hy3: 40.1 (#93), Mistral Small 3.1: 14.7 (#301)

Math benchmarks
BenchmarkHy3Mistral Small 3.1
LMArena Math14751262
OTIS Mock AIME 2024-2025—3.9%
Omni-MATH—24.8%

Knowledge Hy3 leads

Hy3: 40.8 (#114), Mistral Small 3.1: 22.6 (#271)

Knowledge benchmarks
BenchmarkHy3Mistral Small 3.1
LMArena Expert14601257
GPQA Diamond—41.9%
MMLU-Pro—61%
GPQA (HELM)—39.2%

Multimodal Not comparable

Hy3: —, Mistral Small 3.1: 33.2 (#99)

Multimodal benchmarks
BenchmarkHy3Mistral Small 3.1
LMArena Vision—1136

Multilingual Hy3 leads

Hy3: 53.5 (#65), Mistral Small 3.1: 41.2 (#209)

Multilingual benchmarks
BenchmarkHy3Mistral Small 3.1
LMArena Non-English14261255
LMArena Chinese14931253
LMArena French14611273
LMArena German14391266
LMArena Japanese13921208
LMArena Korean13951206
LMArena Russian14321263
LMArena Spanish14561283

Instruction Following Hy3 leads

Hy3: 75.1 (#70), Mistral Small 3.1: 63.6 (#230)

Instruction Following benchmarks
BenchmarkHy3Mistral Small 3.1
LMArena Instruction Following14261264
IFEval—75%

Long Context Hy3 leads

Hy3: 44.1 (#75), Mistral Small 3.1: 39.5 (#178)

Long Context benchmarks
BenchmarkHy3Mistral Small 3.1
LMArena Longer Query14421299

Writing & Preference Hy3 leads

Hy3: 62.2 (#81), Mistral Small 3.1: 37.0 (#259)

Writing & Preference benchmarks
BenchmarkHy3Mistral Small 3.1
LMArena Text14391277
LMArena Creative Writing14021253
LMArena Multi-Turn14361270
EQ-Bench Creative Writing—761
WildBench—78.8%

Frequently asked questions

Is Hy3 better than Mistral Small 3.1?

Hy3 is the stronger model overall, scoring 44.2 to 31.7 on the Noometry Index.

Which is cheaper, Hy3 or Mistral Small 3.1?

Hy3 is cheaper. It lists at $0.0825 per million input tokens and $0.33 per million output tokens; Mistral Small 3.1 lists at $0.35 and $0.56.

Is Hy3 or Mistral Small 3.1 better for coding?

Hy3 scores higher on coding benchmarks: 46.8 versus 38.3 in the Noometry coding category.

Which has the bigger context window?

Hy3 does, with 262K tokens against 128K.

How many benchmarks do Hy3 and Mistral Small 3.1 share?

17 benchmarks have published results for both models. Hy3 has 19 scored results on Noometry and Mistral Small 3.1 has 28.

Related comparisons

Go deeper