Model comparison

Pixtral Large vs Qwen1.5 4b Chat

Pixtral Large is the stronger model overall, scoring 32.2 to 28.8 on the Noometry Index.

Last verified . 0 shared benchmarks.

Pixtral Large Mistral AI

32.2

Rank #259 Reported

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • The widest gap is in writing & preference, where Pixtral Large leads 32.9 to 23.8.

Side by side

Pixtral Large and Qwen1.5 4b Chat specifications
Pixtral LargeQwen1.5 4b Chat
ProviderMistral AIAlibaba (Qwen)
Noometry Index32.228.8
Released2024-11-01—
WeightsOpenOpen
Context window128K—
Max output128K—
Input $ / M tokens$2—
Output $ / M tokens$6—
Results tracked313

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Pixtral Large: —, Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkPixtral LargeQwen1.5 4b Chat
LMArena Coding—999

Reasoning Pixtral Large leads

Pixtral Large: 21.7 (#218), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkPixtral LargeQwen1.5 4b Chat
EnigmaEval0.8%—
LMArena Hard Prompts—976

Math Not comparable

Pixtral Large: —, Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkPixtral LargeQwen1.5 4b Chat
LMArena Math—1026

Knowledge Not comparable

Pixtral Large: —, Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkPixtral LargeQwen1.5 4b Chat
LMArena Expert—980

Multimodal Not comparable

Pixtral Large: 30.6 (#111), Qwen1.5 4b Chat: —

Multimodal benchmarks
BenchmarkPixtral LargeQwen1.5 4b Chat
LMArena Vision1089—

Multilingual Not comparable

Pixtral Large: —, Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkPixtral LargeQwen1.5 4b Chat
LMArena Non-English—979
LMArena Chinese—1024
LMArena German—902
LMArena Russian—952

Instruction Following Not comparable

Pixtral Large: —, Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkPixtral LargeQwen1.5 4b Chat
LMArena Instruction Following—978

Long Context Not comparable

Pixtral Large: —, Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkPixtral LargeQwen1.5 4b Chat
LMArena Longer Query—988

Writing & Preference Pixtral Large leads

Pixtral Large: 32.9 (#278), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkPixtral LargeQwen1.5 4b Chat
LMArena Text—997
LMArena Creative Writing—969
EQ-Bench Creative Writing988—
LMArena Multi-Turn—977

Frequently asked questions

Is Pixtral Large better than Qwen1.5 4b Chat?

Pixtral Large is the stronger model overall, scoring 32.2 to 28.8 on the Noometry Index.

How many benchmarks do Pixtral Large and Qwen1.5 4b Chat share?

0 benchmarks have published results for both models. Pixtral Large has 3 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper