Model comparison

Claude 3.5 Sonnet vs Mistral Medium 3.5

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 34.6 on the Noometry Index.

Last verified . 18 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Mistral Medium 3.5 Mistral AI

40.2

Rank #152 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 2 categories and Mistral Medium 3.5 in 7 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mistral Medium 3.5 leads 39.1 to 19.2.
  • Mistral Medium 3.5 has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Sonnet and Mistral Medium 3.5 specifications
Claude 3.5 SonnetMistral Medium 3.5
ProviderAnthropicMistral AI
Noometry Index34.640.2
Released2024-06-20—
WeightsProprietaryOpen
Context window—262K
Max output—210K
Input $ / M tokens—$1.50
Output $ / M tokens—$7.50
Results tracked6022

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 39.0 (#165), Mistral Medium 3.5: 36.0 (#213)

Coding benchmarks
BenchmarkClaude 3.5 SonnetMistral Medium 3.5
LMArena Coding13421461
Aider Polyglot51.6%—
LMArena WebDev—1264
GSO4.6%—
WeirdML40%—
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
CadEval48%—
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Not comparable

Claude 3.5 Sonnet: 32.3 (#67), Mistral Medium 3.5: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetMistral Medium 3.5
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 23.1 (#183), Mistral Medium 3.5: 17.3 (#295)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetMistral Medium 3.5
LMArena Hard Prompts13051436
Epoch Capabilities Index133.55141.35
SimpleBench41.4%—
Kagi LLM Benchmark—41.4%
NYT Connections (extended)—12.9%
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
DTBench67.8%—
LiveBench Data Analysis55%—
ForecastBench60.7—
LiveBench59%—

Math Mistral Medium 3.5 leads

Claude 3.5 Sonnet: 19.2 (#288), Mistral Medium 3.5: 39.1 (#113)

Math benchmarks
BenchmarkClaude 3.5 SonnetMistral Medium 3.5
LMArena Math13071431
OTIS Mock AIME 2024-20258.5%—
Omni-MATH27.6%—
LiveBench Math52.3%—
MATH Level 556.9%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge Mistral Medium 3.5 leads

Claude 3.5 Sonnet: 28.6 (#245), Mistral Medium 3.5: 40.0 (#126)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetMistral Medium 3.5
LMArena Expert12651432
GPQA Diamond55.3%—
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
Confabulations19.9%—
GPQA (HELM)56.5%—
MMLU87.3%—

Multimodal Mistral Medium 3.5 leads

Claude 3.5 Sonnet: 26.5 (#120), Mistral Medium 3.5: 38.3 (#65)

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetMistral Medium 3.5
LMArena Vision11251223
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual Mistral Medium 3.5 leads

Claude 3.5 Sonnet: 43.2 (#185), Mistral Medium 3.5: 51.9 (#100)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetMistral Medium 3.5
LMArena Non-English12831404
LMArena Chinese12721442
LMArena French13051448
LMArena German12971451
LMArena Korean12001385
LMArena Russian13061395
LMArena Spanish12901409
LMArena Japanese1234—

Instruction Following Mistral Medium 3.5 leads

Claude 3.5 Sonnet: 68.8 (#182), Mistral Medium 3.5: 74.6 (#90)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetMistral Medium 3.5
LMArena Instruction Following12971415
LiveBench Instruction Following69.3%—
IFEval85.5%—

Long Context Mistral Medium 3.5 leads

Claude 3.5 Sonnet: 39.9 (#167), Mistral Medium 3.5: 43.2 (#103)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetMistral Medium 3.5
LMArena Longer Query13111415

Writing & Preference Mistral Medium 3.5 leads

Claude 3.5 Sonnet: 52.9 (#164), Mistral Medium 3.5: 58.5 (#117)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetMistral Medium 3.5
LMArena Text12981421
LMArena Creative Writing12921374
LMArena Multi-Turn13261423
Short-Story Creative Writing80.3%—
EQ-Bench Creative Writing1451—
WildBench79.2%—
EQ-Bench 4—993
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Mistral Medium 3.5?

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 34.6 on the Noometry Index.

Is Claude 3.5 Sonnet or Mistral Medium 3.5 better for coding?

Claude 3.5 Sonnet scores higher on coding benchmarks: 39.0 versus 36.0 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and Mistral Medium 3.5 share?

18 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Mistral Medium 3.5 has 22.

Related comparisons

Go deeper