Model comparison

Claude 2.1 vs Pixtral Large

Pixtral Large is the stronger model overall, scoring 32.2 to 25.2 on the Noometry Index.

Last verified . 0 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Pixtral Large Mistral AI

32.2

Rank #259 Reported

Summary

  • Pixtral Large has downloadable open weights; the other is API-only.

Side by side

Claude 2.1 and Pixtral Large specifications
Claude 2.1Pixtral Large
ProviderAnthropicMistral AI
Noometry Index25.232.2
Released2023-11-212024-11-01
WeightsProprietaryOpen
Context window—128K
Max output—128K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked73

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Claude 2.1: 26.2 (#327), Pixtral Large: —

Coding benchmarks
BenchmarkClaude 2.1Pixtral Large
WeirdML7.1%—

Reasoning Too close to call

Claude 2.1: 21.4 (#221), Pixtral Large: 21.7 (#218)

Reasoning benchmarks
BenchmarkClaude 2.1Pixtral Large
EnigmaEval—0.8%
DTBench51%—
Epoch Capabilities Index119.27—
ForecastBench54.2—

Math Not comparable

Claude 2.1: 10.2 (#315), Pixtral Large: —

Math benchmarks
BenchmarkClaude 2.1Pixtral Large
OTIS Mock AIME 2024-20251.9%—

Knowledge Not comparable

Claude 2.1: 15.4 (#292), Pixtral Large: —

Knowledge benchmarks
BenchmarkClaude 2.1Pixtral Large
GPQA Diamond33%—
MMLU73.5%—

Multimodal Not comparable

Claude 2.1: —, Pixtral Large: 30.6 (#111)

Multimodal benchmarks
BenchmarkClaude 2.1Pixtral Large
LMArena Vision—1089

Writing & Preference Not comparable

Claude 2.1: —, Pixtral Large: 32.9 (#278)

Writing & Preference benchmarks
BenchmarkClaude 2.1Pixtral Large
EQ-Bench Creative Writing—988

Frequently asked questions

Is Claude 2.1 better than Pixtral Large?

Pixtral Large is the stronger model overall, scoring 32.2 to 25.2 on the Noometry Index.

How many benchmarks do Claude 2.1 and Pixtral Large share?

0 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Pixtral Large has 3.

Related comparisons

Go deeper