Model comparison

Mixtral 8x7B vs Qwen3-VL 235B-A22B

Qwen3-VL 235B-A22B is the stronger model overall, scoring 43.2 to 27.1 on the Noometry Index. Mixtral 8x7B costs 1.8× less per token, which makes it the better buy when Qwen3-VL 235B-A22B's lead doesn't matter for your workload.

Last verified . 17 shared benchmarks.

Mixtral 8x7B Mistral AI

27.1

Rank #334 Confirmed

Qwen3-VL 235B-A22B Alibaba (Qwen)

43.2

Rank #95 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Mixtral 8x7B scores higher in 0 categories and Qwen3-VL 235B-A22B in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3-VL 235B-A22B leads 40.3 to 11.0.
  • Mixtral 8x7B is cheaper at $0.70 / $0.70 per million input/output tokens, against $0.70 / $2.80 for Qwen3-VL 235B-A22B.
  • Qwen3-VL 235B-A22B accepts more context: 131K tokens versus 32K.

Side by side

Mixtral 8x7B and Qwen3-VL 235B-A22B specifications
Mixtral 8x7BQwen3-VL 235B-A22B
ProviderMistral AIAlibaba (Qwen)
Noometry Index27.143.2
Released2023-12-112025-04
WeightsOpenOpen
Context window32K131K
Max output32K33K
Input $ / M tokens$0.70$0.70
Output $ / M tokens$0.70$2.80
Results tracked3818

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3-VL 235B-A22B leads

Mixtral 8x7B: 32.8 (#269), Qwen3-VL 235B-A22B: 42.4 (#100)

Coding benchmarks
BenchmarkMixtral 8x7BQwen3-VL 235B-A22B
LMArena Coding11261439
HumanEval+39.6%—
MBPP+49.7%—

Reasoning Qwen3-VL 235B-A22B leads

Mixtral 8x7B: 18.2 (#285), Qwen3-VL 235B-A22B: 29.3 (#92)

Reasoning benchmarks
BenchmarkMixtral 8x7BQwen3-VL 235B-A22B
LMArena Hard Prompts11151428
DTBench49.6%—
Adversarial NLI55.2%—
Epoch Capabilities Index118.47—
ForecastBench56.3—
HellaSwag86.7%—
PIQA83.6%—
WinoGrande77.2%—

Math Qwen3-VL 235B-A22B leads

Mixtral 8x7B: 18.8 (#289), Qwen3-VL 235B-A22B: 39.0 (#118)

Math benchmarks
BenchmarkMixtral 8x7BQwen3-VL 235B-A22B
LMArena Math11471426
Omni-MATH10.5%—
MATH Level 510%—
GSM8K74.4%—

Knowledge Qwen3-VL 235B-A22B leads

Mixtral 8x7B: 11.0 (#301), Qwen3-VL 235B-A22B: 40.3 (#121)

Knowledge benchmarks
BenchmarkMixtral 8x7BQwen3-VL 235B-A22B
LMArena Expert10881442
GPQA Diamond30.6%—
MMLU-Pro33.5%—
GPQA (HELM)29.6%—
ARC (AI2) Challenge87.3%—
MMLU70.6%—
OpenBookQA85.8%—
TriviaQA82.2%—

Multimodal Not comparable

Mixtral 8x7B: —, Qwen3-VL 235B-A22B: 39.8 (#55)

Multimodal benchmarks
BenchmarkMixtral 8x7BQwen3-VL 235B-A22B
LMArena Vision—1247

Multilingual Qwen3-VL 235B-A22B leads

Mixtral 8x7B: 29.6 (#266), Qwen3-VL 235B-A22B: 51.9 (#97)

Multilingual benchmarks
BenchmarkMixtral 8x7BQwen3-VL 235B-A22B
LMArena Non-English10771405
LMArena Chinese10551463
LMArena French11661452
LMArena German11141424
LMArena Japanese9311385
LMArena Korean9681394
LMArena Russian10901408
LMArena Spanish11111428

Instruction Following Qwen3-VL 235B-A22B leads

Mixtral 8x7B: 51.0 (#297), Qwen3-VL 235B-A22B: 74.2 (#101)

Instruction Following benchmarks
BenchmarkMixtral 8x7BQwen3-VL 235B-A22B
LMArena Instruction Following11091406
IFEval57.5%—

Long Context Qwen3-VL 235B-A22B leads

Mixtral 8x7B: 33.4 (#260), Qwen3-VL 235B-A22B: 43.4 (#98)

Long Context benchmarks
BenchmarkMixtral 8x7BQwen3-VL 235B-A22B
LMArena Longer Query11031420

Writing & Preference Qwen3-VL 235B-A22B leads

Mixtral 8x7B: 34.2 (#270), Qwen3-VL 235B-A22B: 60.2 (#99)

Writing & Preference benchmarks
BenchmarkMixtral 8x7BQwen3-VL 235B-A22B
LMArena Text11321420
LMArena Creative Writing11091366
LMArena Multi-Turn11151428
WildBench67.3%—

Frequently asked questions

Is Mixtral 8x7B better than Qwen3-VL 235B-A22B?

Qwen3-VL 235B-A22B is the stronger model overall, scoring 43.2 to 27.1 on the Noometry Index. Mixtral 8x7B costs 1.8× less per token, which makes it the better buy when Qwen3-VL 235B-A22B's lead doesn't matter for your workload.

Which is cheaper, Mixtral 8x7B or Qwen3-VL 235B-A22B?

Mixtral 8x7B is cheaper. It lists at $0.70 per million input tokens and $0.70 per million output tokens; Qwen3-VL 235B-A22B lists at $0.70 and $2.80.

Is Mixtral 8x7B or Qwen3-VL 235B-A22B better for coding?

Qwen3-VL 235B-A22B scores higher on coding benchmarks: 42.4 versus 32.8 in the Noometry coding category.

Which has the bigger context window?

Qwen3-VL 235B-A22B does, with 131K tokens against 32K.

How many benchmarks do Mixtral 8x7B and Qwen3-VL 235B-A22B share?

17 benchmarks have published results for both models. Mixtral 8x7B has 38 scored results on Noometry and Qwen3-VL 235B-A22B has 18.

Related comparisons

Go deeper