Model comparison

Mistral Medium 3.5 vs Tulu 3 (Tülu 3) 70B

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 33.0 on the Noometry Index.

Last verified . 11 shared benchmarks.

Summary

  • They share 11 benchmarks with published results for both. Mistral Medium 3.5 scores higher in 7 categories and Tulu 3 (Tülu 3) 70B in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mistral Medium 3.5 leads 39.1 to 14.2.

Side by side

Mistral Medium 3.5 and Tulu 3 (Tülu 3) 70B specifications
Mistral Medium 3.5Tulu 3 (Tülu 3) 70B
ProviderMistral AIAllen Institute for AI (Ai2)
Noometry Index40.233.0
Released—2024-11-21
WeightsOpenOpen
Context window262K—
Max output210K—
Input $ / M tokens$1.50—
Output $ / M tokens$7.50—
Results tracked2214

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mistral Medium 3.5: 36.0 (#213), Tulu 3 (Tülu 3) 70B: 36.0 (#214)

Coding benchmarks
BenchmarkMistral Medium 3.5Tulu 3 (Tülu 3) 70B
LMArena Coding14611235
LMArena WebDev1264—

Reasoning Tulu 3 (Tülu 3) 70B leads

Mistral Medium 3.5: 17.3 (#295), Tulu 3 (Tülu 3) 70B: 23.9 (#169)

Reasoning benchmarks
BenchmarkMistral Medium 3.5Tulu 3 (Tülu 3) 70B
LMArena Hard Prompts14361220
Kagi LLM Benchmark41.4%—
NYT Connections (extended)12.9%—
Epoch Capabilities Index141.35—

Math Mistral Medium 3.5 leads

Mistral Medium 3.5: 39.1 (#113), Tulu 3 (Tülu 3) 70B: 14.2 (#303)

Math benchmarks
BenchmarkMistral Medium 3.5Tulu 3 (Tülu 3) 70B
LMArena Math14311242
OTIS Mock AIME 2024-2025—4.4%
MATH Level 5—42.7%

Knowledge Mistral Medium 3.5 leads

Mistral Medium 3.5: 40.0 (#126), Tulu 3 (Tülu 3) 70B: 25.0 (#264)

Knowledge benchmarks
BenchmarkMistral Medium 3.5Tulu 3 (Tülu 3) 70B
GPQA Diamond—46.3%
LMArena Expert1432—

Multimodal Not comparable

Mistral Medium 3.5: 38.3 (#65), Tulu 3 (Tülu 3) 70B: —

Multimodal benchmarks
BenchmarkMistral Medium 3.5Tulu 3 (Tülu 3) 70B
LMArena Vision1223—

Multilingual Mistral Medium 3.5 leads

Mistral Medium 3.5: 51.9 (#100), Tulu 3 (Tülu 3) 70B: 39.9 (#222)

Multilingual benchmarks
BenchmarkMistral Medium 3.5Tulu 3 (Tülu 3) 70B
LMArena Non-English14041236
LMArena Chinese14421249
LMArena Russian13951246
LMArena French1448—
LMArena German1451—
LMArena Korean1385—
LMArena Spanish1409—

Instruction Following Mistral Medium 3.5 leads

Mistral Medium 3.5: 74.6 (#90), Tulu 3 (Tülu 3) 70B: 64.8 (#227)

Instruction Following benchmarks
BenchmarkMistral Medium 3.5Tulu 3 (Tülu 3) 70B
LMArena Instruction Following14151233

Long Context Mistral Medium 3.5 leads

Mistral Medium 3.5: 43.2 (#103), Tulu 3 (Tülu 3) 70B: 37.1 (#222)

Long Context benchmarks
BenchmarkMistral Medium 3.5Tulu 3 (Tülu 3) 70B
LMArena Longer Query14151224

Writing & Preference Mistral Medium 3.5 leads

Mistral Medium 3.5: 58.5 (#117), Tulu 3 (Tülu 3) 70B: 45.6 (#223)

Writing & Preference benchmarks
BenchmarkMistral Medium 3.5Tulu 3 (Tülu 3) 70B
LMArena Text14211256
LMArena Creative Writing13741231
LMArena Multi-Turn14231252
EQ-Bench 4993—

Frequently asked questions

Is Mistral Medium 3.5 better than Tulu 3 (Tülu 3) 70B?

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 33.0 on the Noometry Index.

Is Mistral Medium 3.5 or Tulu 3 (Tülu 3) 70B better for coding?

They score almost the same on coding (36.0 vs 36.0); test both on your own repository before choosing.

How many benchmarks do Mistral Medium 3.5 and Tulu 3 (Tülu 3) 70B share?

11 benchmarks have published results for both models. Mistral Medium 3.5 has 22 scored results on Noometry and Tulu 3 (Tülu 3) 70B has 14.

Related comparisons

Go deeper