Model comparison

Claude 2.1 vs Mistral Small

Mistral Small is the stronger model overall, scoring 33.4 to 25.2 on the Noometry Index.

Last verified . 4 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Claude 2.1 scores higher in 1 category and Mistral Small in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Mistral Small leads 31.0 to 15.4.
  • The biggest single-benchmark swing is DTBench: 51% for Claude 2.1 and 70.9% for Mistral Small.
  • Mistral Small has downloadable open weights; the other is API-only.

Side by side

Claude 2.1 and Mistral Small specifications
Claude 2.1Mistral Small
ProviderAnthropicMistral AI
Noometry Index25.233.4
Released2023-11-212024-02-26
WeightsProprietaryOpen
Context window—262K
Max output—256K
Input $ / M tokens—$0.15
Output $ / M tokens—$0.60
Results tracked739

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small leads

Claude 2.1: 26.2 (#327), Mistral Small: 34.0 (#247)

Coding benchmarks
BenchmarkClaude 2.1Mistral Small
SciCode—26.5%
WeirdML7.1%—
BigCodeBench Instruct—36.1%
LiveBench Coding—36.2%
LMArena Coding—1362
BigCodeBench Complete—46.6%
ALE-Bench—497.62

Agentic & Tool Use Not comparable

Claude 2.1: —, Mistral Small: 28.1 (#93)

Agentic & Tool Use benchmarks
BenchmarkClaude 2.1Mistral Small
Berkeley Function Calling Leaderboard—37.1%

Reasoning Claude 2.1 leads

Claude 2.1: 21.4 (#221), Mistral Small: 19.8 (#250)

Reasoning benchmarks
BenchmarkClaude 2.1Mistral Small
DTBench51%70.9%
Kagi LLM Benchmark—37.8%
CritPt—0%
LiveBench Reasoning—44.8%
LMArena Hard Prompts—1335
LiveBench Data Analysis—53.7%
LMCA—20.6%
Epoch Capabilities Index119.27—
ForecastBench54.2—
LiveBench—44%

Math Mistral Small leads

Claude 2.1: 10.2 (#315), Mistral Small: 16.4 (#293)

Math benchmarks
BenchmarkClaude 2.1Mistral Small
OTIS Mock AIME 2024-20251.9%5.8%
LiveBench Math—39.9%
LMArena Math—1341
MATH Level 5—46.8%

Knowledge Mistral Small leads

Claude 2.1: 15.4 (#292), Mistral Small: 31.0 (#222)

Knowledge benchmarks
BenchmarkClaude 2.1Mistral Small
GPQA Diamond33%47.5%
MMLU73.5%68.7%
Vectara Hallucination Rate—5.1%
LMArena Expert—1291

Multimodal Not comparable

Claude 2.1: —, Mistral Small: 33.5 (#96)

Multimodal benchmarks
BenchmarkClaude 2.1Mistral Small
LMArena Vision—1142

Multilingual Not comparable

Claude 2.1: —, Mistral Small: 45.5 (#169)

Multilingual benchmarks
BenchmarkClaude 2.1Mistral Small
LMArena Non-English—1315
LMArena Chinese—1340
LMArena French—1337
LMArena German—1340
LMArena Japanese—1275
LMArena Korean—1259
LMArena Russian—1324
LMArena Spanish—1346

Instruction Following Not comparable

Claude 2.1: —, Mistral Small: 66.4 (#209)

Instruction Following benchmarks
BenchmarkClaude 2.1Mistral Small
LiveBench Instruction Following—63.7%
LMArena Instruction Following—1310

Long Context Not comparable

Claude 2.1: —, Mistral Small: 40.4 (#156)

Long Context benchmarks
BenchmarkClaude 2.1Mistral Small
LMArena Longer Query—1327

Writing & Preference Not comparable

Claude 2.1: —, Mistral Small: 52.5 (#171)

Writing & Preference benchmarks
BenchmarkClaude 2.1Mistral Small
LMArena Text—1338
LMArena Creative Writing—1305
LMArena Multi-Turn—1344
LiveBench Language—30.5%

Frequently asked questions

Is Claude 2.1 better than Mistral Small?

Mistral Small is the stronger model overall, scoring 33.4 to 25.2 on the Noometry Index.

Is Claude 2.1 or Mistral Small better for coding?

Mistral Small scores higher on coding benchmarks: 34.0 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and Mistral Small share?

4 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Mistral Small has 39.

Related comparisons

Go deeper