Model comparison

Mistral Large 4 vs Qwen3.5 Plus

Mistral Large 4 and Qwen3.5 Plus score almost the same on the Noometry Index (43.1 vs 42.9), so choose on price, context window or the category you care about most.

Last verified . 1 shared benchmarks.

Mistral Large 4 Mistral AI

43.1

Rank #99 Confirmed

Qwen3.5 Plus Alibaba (Qwen)

42.9

Rank #106 Confirmed

Summary

  • They share 1 benchmark with published results for both. Mistral Large 4 scores higher in 1 category and Qwen3.5 Plus in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.5 Plus leads 32.8 to 22.5.
  • The biggest single-benchmark swing is SimpleQA Verified: 20% for Mistral Large 4 and 25.4% for Qwen3.5 Plus.
  • Qwen3.5 Plus is cheaper at $0.40 / $2.40 per million input/output tokens, against $0.68 / $2.09 for Mistral Large 4.
  • Mistral Large 4 accepts more context: 1.05M tokens versus 1M.

Side by side

Mistral Large 4 and Qwen3.5 Plus specifications
Mistral Large 4Qwen3.5 Plus
ProviderMistral AIAlibaba (Qwen)
Noometry Index43.142.9
Released2026-10-062026-02-16
WeightsProprietaryProprietary
Context window1.05M1M
Max output262K66K
Input $ / M tokens$0.68$0.40
Output $ / M tokens$2.09$2.40
Results tracked1515

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral Large 4: 48.6 (#57), Qwen3.5 Plus: —

Coding benchmarks
BenchmarkMistral Large 4Qwen3.5 Plus
LMArena WebDev1541—
LMArena Coding1475—
ALE-Bench—621.92

Agentic & Tool Use Not comparable

Mistral Large 4: —, Qwen3.5 Plus: —

Agentic & Tool Use benchmarks
BenchmarkMistral Large 4Qwen3.5 Plus
Vending-Bench 2—0.54

Reasoning Qwen3.5 Plus leads

Mistral Large 4: 22.5 (#192), Qwen3.5 Plus: 32.8 (#74)

Reasoning benchmarks
BenchmarkMistral Large 4Qwen3.5 Plus
NYT Connections (extended)27.4%—
Chess Puzzles—22%
LMArena Hard Prompts1444—
Mystery Game Puzzles—17%
DTBench—80.5%
LMCA—36.4%
Epoch Capabilities Index—146.78

Math Qwen3.5 Plus leads

Mistral Large 4: 40.4 (#91), Qwen3.5 Plus: 49.6 (#61)

Math benchmarks
BenchmarkMistral Large 4Qwen3.5 Plus
OTIS Mock AIME 2024-2025—86.7%
LMArena Math1488—
FrontierMath (Feb 2025 set)—21%
FrontierMath Tier 4 (v1)—2.1%

Knowledge Qwen3.5 Plus leads

Mistral Large 4: 36.6 (#166), Qwen3.5 Plus: 46.0 (#83)

Knowledge benchmarks
BenchmarkMistral Large 4Qwen3.5 Plus
SimpleQA Verified20%25.4%
GPQA Diamond—84.8%
Vectara Hallucination Rate—10.7%
LMArena Expert1447—

Multilingual Not comparable

Mistral Large 4: 52.6 (#82), Qwen3.5 Plus: —

Multilingual benchmarks
BenchmarkMistral Large 4Qwen3.5 Plus
LMArena Non-English1415—
LMArena Chinese1491—
LMArena Russian1414—

Instruction Following Not comparable

Mistral Large 4: 75.0 (#76), Qwen3.5 Plus: —

Instruction Following benchmarks
BenchmarkMistral Large 4Qwen3.5 Plus
LMArena Instruction Following1424—

Long Context Too close to call

Mistral Large 4: 43.6 (#89), Qwen3.5 Plus: 43.0 (#113)

Long Context benchmarks
BenchmarkMistral Large 4Qwen3.5 Plus
CL-bench—19.8%
CL-bench Life—12.4%
LMArena Longer Query1429—

Writing & Preference Not comparable

Mistral Large 4: 60.4 (#97), Qwen3.5 Plus: —

Writing & Preference benchmarks
BenchmarkMistral Large 4Qwen3.5 Plus
LMArena Text1427—
LMArena Creative Writing1361—
LMArena Multi-Turn1424—

Frequently asked questions

Is Mistral Large 4 better than Qwen3.5 Plus?

Mistral Large 4 and Qwen3.5 Plus score almost the same on the Noometry Index (43.1 vs 42.9), so choose on price, context window or the category you care about most.

Which is cheaper, Mistral Large 4 or Qwen3.5 Plus?

Qwen3.5 Plus is cheaper. It lists at $0.40 per million input tokens and $2.40 per million output tokens; Mistral Large 4 lists at $0.68 and $2.09.

Which has the bigger context window?

Mistral Large 4 does, with 1.05M tokens against 1M.

How many benchmarks do Mistral Large 4 and Qwen3.5 Plus share?

1 benchmark has published results for both models. Mistral Large 4 has 15 scored results on Noometry and Qwen3.5 Plus has 15.

Related comparisons

Go deeper