Model comparison

Pixtral Large vs Qwen3 14B

Qwen3 14B is the stronger model overall, scoring 35.5 to 32.2 on the Noometry Index.

Last verified . 0 shared benchmarks.

Pixtral Large Mistral AI

32.2

Rank #259 Reported

Qwen3 14B Alibaba (Qwen)

35.5

Rank #225 Confirmed

Summary

  • The widest gap is in reasoning, where Pixtral Large leads 21.7 to 18.5.
  • Qwen3 14B is cheaper at $0.35 / $1.40 per million input/output tokens, against $2 / $6 for Pixtral Large.
  • Qwen3 14B accepts more context: 131K tokens versus 128K.

Side by side

Pixtral Large and Qwen3 14B specifications
Pixtral LargeQwen3 14B
ProviderMistral AIAlibaba (Qwen)
Noometry Index32.235.5
Released2024-11-012025-04
WeightsOpenOpen
Context window128K131K
Max output128K8K
Input $ / M tokens$2$0.35
Output $ / M tokens$6$1.40
Results tracked312

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Pixtral Large: —, Qwen3 14B: 37.3 (#195)

Coding benchmarks
BenchmarkPixtral LargeQwen3 14B
SciCode—31.6%

Agentic & Tool Use Not comparable

Pixtral Large: —, Qwen3 14B: 29.6 (#83)

Agentic & Tool Use benchmarks
BenchmarkPixtral LargeQwen3 14B
Berkeley Function Calling Leaderboard—41%

Reasoning Pixtral Large leads

Pixtral Large: 21.7 (#218), Qwen3 14B: 18.5 (#280)

Reasoning benchmarks
BenchmarkPixtral LargeQwen3 14B
Kagi LLM Benchmark—49.1%
CritPt—0%
Chess Puzzles—4%
EnigmaEval0.8%—
DTBench—64%
LMCA—18.2%
Epoch Capabilities Index—138.23

Math Not comparable

Pixtral Large: —, Qwen3 14B: 38.6 (#133)

Math benchmarks
BenchmarkPixtral LargeQwen3 14B
OTIS Mock AIME 2024-2025—66.4%

Knowledge Not comparable

Pixtral Large: —, Qwen3 14B: 39.3 (#134)

Knowledge benchmarks
BenchmarkPixtral LargeQwen3 14B
GPQA Diamond—63.8%
Vectara Hallucination Rate—5.4%

Multimodal Not comparable

Pixtral Large: 30.6 (#111), Qwen3 14B: —

Multimodal benchmarks
BenchmarkPixtral LargeQwen3 14B
LMArena Vision1089—

Long Context Not comparable

Pixtral Large: —, Qwen3 14B: 38.1 (#204)

Long Context benchmarks
BenchmarkPixtral LargeQwen3 14B
Fiction.LiveBench—62.5%

Writing & Preference Not comparable

Pixtral Large: 32.9 (#278), Qwen3 14B: —

Writing & Preference benchmarks
BenchmarkPixtral LargeQwen3 14B
EQ-Bench Creative Writing988—

Frequently asked questions

Is Pixtral Large better than Qwen3 14B?

Qwen3 14B is the stronger model overall, scoring 35.5 to 32.2 on the Noometry Index.

Which is cheaper, Pixtral Large or Qwen3 14B?

Qwen3 14B is cheaper. It lists at $0.35 per million input tokens and $1.40 per million output tokens; Pixtral Large lists at $2 and $6.

Which has the bigger context window?

Qwen3 14B does, with 131K tokens against 128K.

How many benchmarks do Pixtral Large and Qwen3 14B share?

0 benchmarks have published results for both models. Pixtral Large has 3 scored results on Noometry and Qwen3 14B has 12.

Related comparisons

Go deeper