Model comparison

Mixtral 8x22B vs Qwen3 235B-A22B

Qwen3 235B-A22B is the stronger model overall, scoring 43.5 to 27.1 on the Noometry Index.

Last verified . 28 shared benchmarks.

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Qwen3 235B-A22B Alibaba (Qwen)

43.5

Rank #91 Confirmed

Summary

  • They share 28 benchmarks with published results for both. Mixtral 8x22B scores higher in 1 category and Qwen3 235B-A22B in 8 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3 235B-A22B leads 49.6 to 15.1.
  • The biggest single-benchmark swing is Omni-MATH: 16.3% for Mixtral 8x22B and 71.8% for Qwen3 235B-A22B.
  • Qwen3 235B-A22B is cheaper at $0.70 / $2.80 per million input/output tokens, against $2 / $6 for Mixtral 8x22B.
  • Qwen3 235B-A22B accepts more context: 131K tokens versus 64K.

Side by side

Mixtral 8x22B and Qwen3 235B-A22B specifications
Mixtral 8x22BQwen3 235B-A22B
ProviderMistral AIAlibaba (Qwen)
Noometry Index27.143.5
Released2024-04-172025-04
WeightsOpenOpen
Context window64K131K
Max output64K16K
Input $ / M tokens$2$0.70
Output $ / M tokens$6$2.80
Results tracked3449

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 235B-A22B leads

Mixtral 8x22B: 24.2 (#329), Qwen3 235B-A22B: 44.3 (#75)

Coding benchmarks
BenchmarkMixtral 8x22BQwen3 235B-A22B
WeirdML3.2%41%
LMArena Coding11661445
Aider Polyglot—59.6%
SciCode—42.4%
BigCodeBench Instruct40.6%—
BigCodeBench Complete50.2%—
HumanEval+72%—
MBPP+64.3%—

Agentic & Tool Use Qwen3 235B-A22B leads

Mixtral 8x22B: 23.1 (#127), Qwen3 235B-A22B: 33.9 (#51)

Agentic & Tool Use benchmarks
BenchmarkMixtral 8x22BQwen3 235B-A22B
Berkeley Function Calling Leaderboard—52.1%
Cybench7.5%—
Vending-Bench 2—-11.34

Reasoning Mixtral 8x22B leads

Mixtral 8x22B: 19.9 (#248), Qwen3 235B-A22B: 15.7 (#311)

Reasoning benchmarks
BenchmarkMixtral 8x22BQwen3 235B-A22B
LMArena Hard Prompts11501433
DTBench55.1%80.3%
Epoch Capabilities Index122.03143.85
ForecastBench56.359.7
ARC-AGI-2—1.3%
SimpleBench—31%
Kagi LLM Benchmark—69.4%
ARC-AGI-1—11%
CritPt—0%
Chess Puzzles—12%
Mystery Game Puzzles—9%
LMCA—29.3%

Math Qwen3 235B-A22B leads

Mixtral 8x22B: 22.9 (#275), Qwen3 235B-A22B: 50.4 (#57)

Math benchmarks
BenchmarkMixtral 8x22BQwen3 235B-A22B
Omni-MATH16.3%71.8%
LMArena Math11841432
MATH Level 524.2%68.9%
OTIS Mock AIME 2024-2025—86.7%
FrontierMath (Feb 2025 set)—8.5%
FrontierMath Tier 4 (v1)—0%

Knowledge Qwen3 235B-A22B leads

Mixtral 8x22B: 15.1 (#293), Qwen3 235B-A22B: 49.6 (#73)

Knowledge benchmarks
BenchmarkMixtral 8x22BQwen3 235B-A22B
GPQA Diamond34.1%80.1%
MMLU-Pro46%84.4%
GPQA (HELM)33.4%72.7%
LMArena Expert11131463
SimpleQA Verified—40.4%
Confabulations—15.6%
Vectara Hallucination Rate—9.3%
MMLU77.8%—

Multilingual Qwen3 235B-A22B leads

Mixtral 8x22B: 32.8 (#255), Qwen3 235B-A22B: 52.3 (#89)

Multilingual benchmarks
BenchmarkMixtral 8x22BQwen3 235B-A22B
LMArena Non-English11281409
LMArena Chinese11161481
LMArena French11661445
LMArena German11411433
LMArena Japanese10371399
LMArena Korean10571391
LMArena Russian11581411
LMArena Spanish11511430

Instruction Following Qwen3 235B-A22B leads

Mixtral 8x22B: 57.7 (#266), Qwen3 235B-A22B: 72.6 (#136)

Instruction Following benchmarks
BenchmarkMixtral 8x22BQwen3 235B-A22B
IFEval72.4%83.5%
LMArena Instruction Following11471408

Long Context Qwen3 235B-A22B leads

Mixtral 8x22B: 34.7 (#247), Qwen3 235B-A22B: 46.1 (#26)

Long Context benchmarks
BenchmarkMixtral 8x22BQwen3 235B-A22B
LMArena Longer Query11441426
Fiction.LiveBench—75%

Writing & Preference Qwen3 235B-A22B leads

Mixtral 8x22B: 36.9 (#262), Qwen3 235B-A22B: 59.6 (#108)

Writing & Preference benchmarks
BenchmarkMixtral 8x22BQwen3 235B-A22B
LMArena Text11621419
LMArena Creative Writing11411384
WildBench71.1%86.6%
LMArena Multi-Turn11301432
Short-Story Creative Writing—83%
EQ-Bench Creative Writing—1366

Frequently asked questions

Is Mixtral 8x22B better than Qwen3 235B-A22B?

Qwen3 235B-A22B is the stronger model overall, scoring 43.5 to 27.1 on the Noometry Index.

Which is cheaper, Mixtral 8x22B or Qwen3 235B-A22B?

Qwen3 235B-A22B is cheaper. It lists at $0.70 per million input tokens and $2.80 per million output tokens; Mixtral 8x22B lists at $2 and $6.

Is Mixtral 8x22B or Qwen3 235B-A22B better for coding?

Qwen3 235B-A22B scores higher on coding benchmarks: 44.3 versus 24.2 in the Noometry coding category.

Which has the bigger context window?

Qwen3 235B-A22B does, with 131K tokens against 64K.

How many benchmarks do Mixtral 8x22B and Qwen3 235B-A22B share?

28 benchmarks have published results for both models. Mixtral 8x22B has 34 scored results on Noometry and Qwen3 235B-A22B has 49.

Related comparisons

Go deeper