Model comparison
Pixtral Large vs Qwen3 14B
Qwen3 14B is the stronger model overall, scoring 35.5 to 32.2 on the Noometry Index.
Last verified . 0 shared benchmarks.
Summary
- The widest gap is in reasoning, where Pixtral Large leads 21.7 to 18.5.
- Qwen3 14B is cheaper at $0.35 / $1.40 per million input/output tokens, against $2 / $6 for Pixtral Large.
- Qwen3 14B accepts more context: 131K tokens versus 128K.
Side by side
| Pixtral Large | Qwen3 14B | |
|---|---|---|
| Provider | Mistral AI | Alibaba (Qwen) |
| Noometry Index | 32.2 | 35.5 |
| Released | 2024-11-01 | 2025-04 |
| Weights | Open | Open |
| Context window | 128K | 131K |
| Max output | 128K | 8K |
| Input $ / M tokens | $2 | $0.35 |
| Output $ / M tokens | $6 | $1.40 |
| Results tracked | 3 | 12 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Not comparable
Pixtral Large: —, Qwen3 14B: 37.3 (#195)
| Benchmark | Pixtral Large | Qwen3 14B |
|---|---|---|
| SciCode | — | 31.6% |
Agentic & Tool Use Not comparable
Pixtral Large: —, Qwen3 14B: 29.6 (#83)
| Benchmark | Pixtral Large | Qwen3 14B |
|---|---|---|
| Berkeley Function Calling Leaderboard | — | 41% |
Reasoning Pixtral Large leads
Pixtral Large: 21.7 (#218), Qwen3 14B: 18.5 (#280)
| Benchmark | Pixtral Large | Qwen3 14B |
|---|---|---|
| Kagi LLM Benchmark | — | 49.1% |
| CritPt | — | 0% |
| Chess Puzzles | — | 4% |
| EnigmaEval | 0.8% | — |
| DTBench | — | 64% |
| LMCA | — | 18.2% |
| Epoch Capabilities Index | — | 138.23 |
Math Not comparable
Pixtral Large: —, Qwen3 14B: 38.6 (#133)
| Benchmark | Pixtral Large | Qwen3 14B |
|---|---|---|
| OTIS Mock AIME 2024-2025 | — | 66.4% |
Knowledge Not comparable
Pixtral Large: —, Qwen3 14B: 39.3 (#134)
| Benchmark | Pixtral Large | Qwen3 14B |
|---|---|---|
| GPQA Diamond | — | 63.8% |
| Vectara Hallucination Rate | — | 5.4% |
Multimodal Not comparable
Pixtral Large: 30.6 (#111), Qwen3 14B: —
| Benchmark | Pixtral Large | Qwen3 14B |
|---|---|---|
| LMArena Vision | 1089 | — |
Long Context Not comparable
Pixtral Large: —, Qwen3 14B: 38.1 (#204)
| Benchmark | Pixtral Large | Qwen3 14B |
|---|---|---|
| Fiction.LiveBench | — | 62.5% |
Writing & Preference Not comparable
Pixtral Large: 32.9 (#278), Qwen3 14B: —
| Benchmark | Pixtral Large | Qwen3 14B |
|---|---|---|
| EQ-Bench Creative Writing | 988 | — |
Frequently asked questions
Is Pixtral Large better than Qwen3 14B?
Qwen3 14B is the stronger model overall, scoring 35.5 to 32.2 on the Noometry Index.
Which is cheaper, Pixtral Large or Qwen3 14B?
Qwen3 14B is cheaper. It lists at $0.35 per million input tokens and $1.40 per million output tokens; Pixtral Large lists at $2 and $6.
Which has the bigger context window?
Qwen3 14B does, with 131K tokens against 128K.
How many benchmarks do Pixtral Large and Qwen3 14B share?
0 benchmarks have published results for both models. Pixtral Large has 3 scored results on Noometry and Qwen3 14B has 12.