Model comparison

Claude 3.7 Sonnet vs Mistral Large 3

Claude 3.7 Sonnet and Mistral Large 3 score almost the same on the Noometry Index (39.5 vs 39.1), so choose on price, context window or the category you care about most.

Last verified . 19 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

Mistral Large 3 Mistral AI

39.1

Rank #176 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 4 categories and Mistral Large 3 in 5 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Mistral Large 3 leads 52.5 to 44.1.
  • Mistral Large 3 has downloadable open weights; the other is API-only.

Side by side

Claude 3.7 Sonnet and Mistral Large 3 specifications
Claude 3.7 SonnetMistral Large 3
ProviderAnthropicMistral AI
Noometry Index39.539.1
Released2025-02-242025-12-02
WeightsProprietaryOpen
Context window—262K
Max output—8K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.75
Results tracked5824

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 40.6 (#136), Mistral Large 3: 34.4 (#237)

Coding benchmarks
BenchmarkClaude 3.7 SonnetMistral Large 3
LMArena Coding13611448
SWE-bench Verified61%—
SWE-bench Verified (bash only)52.8%—
Aider Polyglot64.9%—
LMArena WebDev—1230
GSO3.8%—
LiveBench Coding74.5%—
CadEval54%—

Agentic & Tool Use Not comparable

Claude 3.7 Sonnet: 34.1 (#50), Mistral Large 3: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 SonnetMistral Large 3
TheAgentCompany30.9%—
Cybench20%—
DeepResearch Bench43.6%—
OSWorld35.8%—
METR Time Horizons60%—

Reasoning Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 18.6 (#277), Mistral Large 3: 15.2 (#319)

Reasoning benchmarks
BenchmarkClaude 3.7 SonnetMistral Large 3
LMArena Hard Prompts13331429
ARC-AGI-20.9%—
SimpleBench46.4%—
Kagi LLM Benchmark—50.9%
NYT Connections (extended)—7.5%
ARC-AGI-128.6%—
EnigmaEval4.2%—
Thematic Generalization—23%
LiveBench Reasoning87.8%—
LiveBench Data Analysis74%—
Epoch Capabilities Index141.16—
ForecastBench61.8—
LiveBench76.1%—

Math Mistral Large 3 leads

Claude 3.7 Sonnet: 37.5 (#153), Mistral Large 3: 38.7 (#129)

Math benchmarks
BenchmarkClaude 3.7 SonnetMistral Large 3
LMArena Math13371414
OTIS Mock AIME 2024-202557.8%—
Omni-MATH33%—
LiveBench Math79%—
MATH Level 591.2%—
FrontierMath (Feb 2025 set)4.1%—

Knowledge Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 39.8 (#130), Mistral Large 3: 36.0 (#177)

Knowledge benchmarks
BenchmarkClaude 3.7 SonnetMistral Large 3
LMArena Expert13211421
GPQA Diamond79.7%—
Humanity's Last Exam8%—
MMLU-Pro78.4%—
Confabulations14.7%—
Vectara Hallucination Rate—14.5%
GPQA (HELM)60.8%—

Multimodal Mistral Large 3 leads

Claude 3.7 Sonnet: 33.7 (#95), Mistral Large 3: 38.2 (#66)

Multimodal benchmarks
BenchmarkClaude 3.7 SonnetMistral Large 3
LMArena Vision11691221
GeoBench68%—
VPCT39%—
SpatialViz-Bench33.9%—

Multilingual Mistral Large 3 leads

Claude 3.7 Sonnet: 44.1 (#179), Mistral Large 3: 52.5 (#84)

Multilingual benchmarks
BenchmarkClaude 3.7 SonnetMistral Large 3
LMArena Non-English12961413
LMArena Chinese12991447
LMArena French13031455
LMArena German13011437
LMArena Japanese12671394
LMArena Korean12491384
LMArena Russian13111411
LMArena Spanish12981440

Instruction Following Mistral Large 3 leads

Claude 3.7 Sonnet: 72.9 (#125), Mistral Large 3: 74.0 (#108)

Instruction Following benchmarks
BenchmarkClaude 3.7 SonnetMistral Large 3
LMArena Instruction Following13521403
LiveBench Instruction Following81.3%—
IFEval83.4%—

Long Context Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 50.3 (#10), Mistral Large 3: 43.1 (#105)

Long Context benchmarks
BenchmarkClaude 3.7 SonnetMistral Large 3
LMArena Longer Query13731413
Fiction.LiveBench83.3%—

Writing & Preference Mistral Large 3 leads

Claude 3.7 Sonnet: 54.4 (#150), Mistral Large 3: 60.0 (#101)

Writing & Preference benchmarks
BenchmarkClaude 3.7 SonnetMistral Large 3
LMArena Text13141428
LMArena Creative Writing13321386
EQ-Bench Creative Writing14121412
LMArena Multi-Turn13391429
Short-Story Creative Writing81.1%—
WildBench81.4%—
LiveBench Language59.9%—

Frequently asked questions

Is Claude 3.7 Sonnet better than Mistral Large 3?

Claude 3.7 Sonnet and Mistral Large 3 score almost the same on the Noometry Index (39.5 vs 39.1), so choose on price, context window or the category you care about most.

Is Claude 3.7 Sonnet or Mistral Large 3 better for coding?

Claude 3.7 Sonnet scores higher on coding benchmarks: 40.6 versus 34.4 in the Noometry coding category.

How many benchmarks do Claude 3.7 Sonnet and Mistral Large 3 share?

19 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and Mistral Large 3 has 24.

Related comparisons

Go deeper