Model comparison

Claude 3.7 Sonnet vs Magistral Small

Claude 3.7 Sonnet is the stronger model overall, scoring 39.5 to 30.2 on the Noometry Index.

Last verified . 5 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 4 categories and Magistral Small in 0 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Claude 3.7 Sonnet leads 18.6 to 6.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 57.8% for Claude 3.7 Sonnet and 30% for Magistral Small.
  • Magistral Small has downloadable open weights; the other is API-only.

Side by side

Claude 3.7 Sonnet and Magistral Small specifications
Claude 3.7 SonnetMagistral Small
ProviderAnthropicMistral AI
Noometry Index39.530.2
Released2025-02-242025-06-10
WeightsProprietaryOpen
Context window—128K
Max output—40K
Input $ / M tokens—$0.50
Output $ / M tokens—$1.50
Results tracked5810

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 40.6 (#136), Magistral Small: 38.4 (#176)

Coding benchmarks
BenchmarkClaude 3.7 SonnetMagistral Small
SWE-bench Verified61%—
SWE-bench Verified (bash only)52.8%—
Aider Polyglot64.9%—
SciCode—35.2%
GSO3.8%—
LiveBench Coding74.5%—
LMArena Coding1361—
CadEval54%—

Agentic & Tool Use Not comparable

Claude 3.7 Sonnet: 34.1 (#50), Magistral Small: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 SonnetMagistral Small
TheAgentCompany30.9%—
Cybench20%—
DeepResearch Bench43.6%—
OSWorld35.8%—
METR Time Horizons60%—

Reasoning Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 18.6 (#277), Magistral Small: 6.8 (#350)

Reasoning benchmarks
BenchmarkClaude 3.7 SonnetMagistral Small
ARC-AGI-20.9%0%
ARC-AGI-128.6%5%
Epoch Capabilities Index141.16133.19
SimpleBench46.4%—
Kagi LLM Benchmark—6.3%
CritPt—0.3%
Chess Puzzles—3%
EnigmaEval4.2%—
LiveBench Reasoning87.8%—
LMArena Hard Prompts1333—
DTBench—61.3%
LiveBench Data Analysis74%—
ForecastBench61.8—
LiveBench76.1%—

Math Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 37.5 (#153), Magistral Small: 26.2 (#261)

Math benchmarks
BenchmarkClaude 3.7 SonnetMagistral Small
OTIS Mock AIME 2024-202557.8%30%
Omni-MATH33%—
LiveBench Math79%—
LMArena Math1337—
MATH Level 591.2%—
FrontierMath (Feb 2025 set)4.1%—

Knowledge Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 39.8 (#130), Magistral Small: 30.9 (#223)

Knowledge benchmarks
BenchmarkClaude 3.7 SonnetMagistral Small
GPQA Diamond79.7%56.1%
Humanity's Last Exam8%—
MMLU-Pro78.4%—
Confabulations14.7%—
GPQA (HELM)60.8%—
LMArena Expert1321—

Multimodal Not comparable

Claude 3.7 Sonnet: 33.7 (#95), Magistral Small: —

Multimodal benchmarks
BenchmarkClaude 3.7 SonnetMagistral Small
LMArena Vision1169—
GeoBench68%—
VPCT39%—
SpatialViz-Bench33.9%—

Multilingual Not comparable

Claude 3.7 Sonnet: 44.1 (#179), Magistral Small: —

Multilingual benchmarks
BenchmarkClaude 3.7 SonnetMagistral Small
LMArena Non-English1296—
LMArena Chinese1299—
LMArena French1303—
LMArena German1301—
LMArena Japanese1267—
LMArena Korean1249—
LMArena Russian1311—
LMArena Spanish1298—

Instruction Following Not comparable

Claude 3.7 Sonnet: 72.9 (#125), Magistral Small: —

Instruction Following benchmarks
BenchmarkClaude 3.7 SonnetMagistral Small
LiveBench Instruction Following81.3%—
IFEval83.4%—
LMArena Instruction Following1352—

Long Context Not comparable

Claude 3.7 Sonnet: 50.3 (#10), Magistral Small: —

Long Context benchmarks
BenchmarkClaude 3.7 SonnetMagistral Small
Fiction.LiveBench83.3%—
LMArena Longer Query1373—

Writing & Preference Not comparable

Claude 3.7 Sonnet: 54.4 (#150), Magistral Small: —

Writing & Preference benchmarks
BenchmarkClaude 3.7 SonnetMagistral Small
LMArena Text1314—
LMArena Creative Writing1332—
Short-Story Creative Writing81.1%—
EQ-Bench Creative Writing1412—
WildBench81.4%—
LMArena Multi-Turn1339—
LiveBench Language59.9%—

Frequently asked questions

Is Claude 3.7 Sonnet better than Magistral Small?

Claude 3.7 Sonnet is the stronger model overall, scoring 39.5 to 30.2 on the Noometry Index.

Is Claude 3.7 Sonnet or Magistral Small better for coding?

Claude 3.7 Sonnet scores higher on coding benchmarks: 40.6 versus 38.4 in the Noometry coding category.

How many benchmarks do Claude 3.7 Sonnet and Magistral Small share?

5 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and Magistral Small has 10.

Related comparisons

Go deeper