Model comparison

MiniMax-M3 vs Qwen3.8 27B

Qwen3.8 27B is the stronger model overall, scoring 46.0 to 43.8 on the Noometry Index. MiniMax-M3 costs 2.1× less per token, which makes it the better buy when Qwen3.8 27B's lead doesn't matter for your workload.

Last verified . 28 shared benchmarks.

MiniMax-M3 MiniMax

43.8

Rank #85 Confirmed

Qwen3.8 27B Alibaba (Qwen)

46.0

Rank #68 Confirmed

Summary

  • They share 28 benchmarks with published results for both. MiniMax-M3 scores higher in 2 categories and Qwen3.8 27B in 8 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where MiniMax-M3 leads 58.4 to 41.6.
  • The biggest single-benchmark swing is NYT Connections (extended): 65.1% for MiniMax-M3 and 54.5% for Qwen3.8 27B.
  • MiniMax-M3 is cheaper at $0.30 / $1.20 per million input/output tokens, against $0.99 / $1.49 for Qwen3.8 27B.
  • MiniMax-M3 accepts more context: 1M tokens versus 262K.

Side by side

MiniMax-M3 and Qwen3.8 27B specifications
MiniMax-M3Qwen3.8 27B
ProviderMiniMaxAlibaba (Qwen)
Noometry Index43.846.0
Released2026-06-012026-08-14
WeightsOpenOpen
Context window1M262K
Max output512K33K
Input $ / M tokens$0.30$0.99
Output $ / M tokens$1.20$1.49
Results tracked4131

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.8 27B leads

MiniMax-M3: 41.8 (#118), Qwen3.8 27B: 50.5 (#44)

Coding benchmarks
BenchmarkMiniMax-M3Qwen3.8 27B
LMArena WebDev14821593
SciCode47.1%46.6%
LMArena Coding14691482
FrontierCode14.7%—
ALE-Bench640.02—

Agentic & Tool Use Qwen3.8 27B leads

MiniMax-M3: 22.6 (#130), Qwen3.8 27B: 32.9 (#57)

Agentic & Tool Use benchmarks
BenchmarkMiniMax-M3Qwen3.8 27B
APEX-Agents37.7%47.5%
OSWorld 2.04.6%—
GBAEval0.9%—
Vending-Bench 22,158—

Reasoning Qwen3.8 27B leads

MiniMax-M3: 30.1 (#87), Qwen3.8 27B: 41.0 (#54)

Reasoning benchmarks
BenchmarkMiniMax-M3Qwen3.8 27B
NYT Connections (extended)65.1%54.5%
CritPt3.7%5.4%
LMArena Hard Prompts14471460
DTBench78.9%88%
LMCA33.7%41.4%
Surface Evolver Bench55%45%
Epoch Capabilities Index146.95149.38
ARC-AGI-2—42.4%
SimpleBench45.8%—
ARC-AGI-1—87.5%
Chess Puzzles14%—
Mystery Game Puzzles8%—
ForecastBench61.4—

Math MiniMax-M3 leads

MiniMax-M3: 40.0 (#95), Qwen3.8 27B: 37.1 (#161)

Math benchmarks
BenchmarkMiniMax-M3Qwen3.8 27B
ProofBench18%16%
LMArena Math14291456
OTIS Mock AIME 2024-202571.1%—

Knowledge MiniMax-M3 leads

MiniMax-M3: 58.4 (#35), Qwen3.8 27B: 41.6 (#109)

Knowledge benchmarks
BenchmarkMiniMax-M3Qwen3.8 27B
LMArena Expert14611482
GPQA Diamond90.9%—

Multimodal Qwen3.8 27B leads

MiniMax-M3: 40.2 (#51), Qwen3.8 27B: 41.3 (#37)

Multimodal benchmarks
BenchmarkMiniMax-M3Qwen3.8 27B
LMArena Vision12531271
LMArena Document1435—

Multilingual Too close to call

MiniMax-M3: 53.0 (#75), Qwen3.8 27B: 53.7 (#60)

Multilingual benchmarks
BenchmarkMiniMax-M3Qwen3.8 27B
LMArena Non-English14201430
LMArena Chinese14631504
LMArena French14471465
LMArena German14261438
LMArena Japanese13811384
LMArena Korean13721393
LMArena Russian14281415
LMArena Spanish14321448

Instruction Following Too close to call

MiniMax-M3: 75.5 (#62), Qwen3.8 27B: 75.8 (#53)

Instruction Following benchmarks
BenchmarkMiniMax-M3Qwen3.8 27B
LMArena Instruction Following14331439

Long Context Too close to call

MiniMax-M3: 44.2 (#72), Qwen3.8 27B: 44.3 (#70)

Long Context benchmarks
BenchmarkMiniMax-M3Qwen3.8 27B
LMArena Longer Query14451450

Writing & Preference Qwen3.8 27B leads

MiniMax-M3: 62.1 (#83), Qwen3.8 27B: 65.8 (#43)

Writing & Preference benchmarks
BenchmarkMiniMax-M3Qwen3.8 27B
LMArena Text14331441
LMArena Creative Writing14041384
LMArena Multi-Turn14421441
EQ-Bench Creative Writing—1671
EQ-Bench 41150—

Frequently asked questions

Is MiniMax-M3 better than Qwen3.8 27B?

Qwen3.8 27B is the stronger model overall, scoring 46.0 to 43.8 on the Noometry Index. MiniMax-M3 costs 2.1× less per token, which makes it the better buy when Qwen3.8 27B's lead doesn't matter for your workload.

Which is cheaper, MiniMax-M3 or Qwen3.8 27B?

MiniMax-M3 is cheaper. It lists at $0.30 per million input tokens and $1.20 per million output tokens; Qwen3.8 27B lists at $0.99 and $1.49.

Is MiniMax-M3 or Qwen3.8 27B better for coding?

Qwen3.8 27B scores higher on coding benchmarks: 50.5 versus 41.8 in the Noometry coding category.

Which has the bigger context window?

MiniMax-M3 does, with 1M tokens against 262K.

How many benchmarks do MiniMax-M3 and Qwen3.8 27B share?

28 benchmarks have published results for both models. MiniMax-M3 has 41 scored results on Noometry and Qwen3.8 27B has 31.

Related comparisons

Go deeper