Model comparison

Qwen3.8 27B vs Qwen3 Max

Qwen3.8 27B is the stronger model overall, scoring 46.0 to 43.7 on the Noometry Index.

Last verified . 21 shared benchmarks.

Qwen3.8 27B Alibaba (Qwen)

46.0

Rank #68 Confirmed

Qwen3 Max Alibaba (Qwen)

43.7

Rank #87 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Qwen3.8 27B scores higher in 6 categories and Qwen3 Max in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.8 27B leads 41.0 to 22.6.
  • The biggest single-benchmark swing is NYT Connections (extended): 54.5% for Qwen3.8 27B and 30.1% for Qwen3 Max.
  • Qwen3.8 27B is cheaper at $0.99 / $1.49 per million input/output tokens, against $1.20 / $6 for Qwen3 Max.
  • Qwen3.8 27B has downloadable open weights; the other is API-only.

Side by side

Qwen3.8 27B and Qwen3 Max specifications
Qwen3.8 27BQwen3 Max
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index46.043.7
Released2026-08-142025-09-23
WeightsOpenProprietary
Context window262K262K
Max output33K66K
Input $ / M tokens$0.99$1.20
Output $ / M tokens$1.49$6
Results tracked3133

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.8 27B leads

Qwen3.8 27B: 50.5 (#44), Qwen3 Max: 43.0 (#93)

Coding benchmarks
BenchmarkQwen3.8 27BQwen3 Max
LMArena Coding14821456
LMArena WebDev1593—
SciCode46.6%—
ALE-Bench—370.45

Agentic & Tool Use Not comparable

Qwen3.8 27B: 32.9 (#57), Qwen3 Max: —

Agentic & Tool Use benchmarks
BenchmarkQwen3.8 27BQwen3 Max
APEX-Agents47.5%—
Vending-Bench 2—71.56

Reasoning Qwen3.8 27B leads

Qwen3.8 27B: 41.0 (#54), Qwen3 Max: 22.6 (#190)

Reasoning benchmarks
BenchmarkQwen3.8 27BQwen3 Max
NYT Connections (extended)54.5%30.1%
LMArena Hard Prompts14601448
DTBench88%82.1%
LMCA41.4%28.3%
Epoch Capabilities Index149.38142.38
ARC-AGI-242.4%—
Kagi LLM Benchmark—72.5%
ARC-AGI-187.5%—
CritPt5.4%—
Chess Puzzles—4%
Mystery Game Puzzles—5%
Surface Evolver Bench45%—

Math Qwen3 Max leads

Qwen3.8 27B: 37.1 (#161), Qwen3 Max: 38.7 (#131)

Math benchmarks
BenchmarkQwen3.8 27BQwen3 Max
LMArena Math14561446
FrontierMath (Tiers 1-3)—18.9%
OTIS Mock AIME 2024-2025—73.3%
ProofBench16%—
MATH Level 5—97.1%

Knowledge Qwen3 Max leads

Qwen3.8 27B: 41.6 (#109), Qwen3 Max: 48.1 (#78)

Knowledge benchmarks
BenchmarkQwen3.8 27BQwen3 Max
LMArena Expert14821455
GPQA Diamond—72.6%
SimpleQA Verified—48.7%

Multimodal Not comparable

Qwen3.8 27B: 41.3 (#37), Qwen3 Max: —

Multimodal benchmarks
BenchmarkQwen3.8 27BQwen3 Max
LMArena Vision1271—

Multilingual Too close to call

Qwen3.8 27B: 53.7 (#60), Qwen3 Max: 53.7 (#62)

Multilingual benchmarks
BenchmarkQwen3.8 27BQwen3 Max
LMArena Non-English14301429
LMArena Chinese15041478
LMArena French14651449
LMArena German14381463
LMArena Japanese13841397
LMArena Korean13931399
LMArena Russian14151428
LMArena Spanish14481462

Instruction Following Too close to call

Qwen3.8 27B: 75.8 (#53), Qwen3 Max: 74.8 (#87)

Instruction Following benchmarks
BenchmarkQwen3.8 27BQwen3 Max
LMArena Instruction Following14391419

Long Context Qwen3.8 27B leads

Qwen3.8 27B: 44.3 (#70), Qwen3 Max: 41.6 (#134)

Long Context benchmarks
BenchmarkQwen3.8 27BQwen3 Max
LMArena Longer Query14501438
Fiction.LiveBench—66.7%
CL-bench—14.5%

Writing & Preference Qwen3.8 27B leads

Qwen3.8 27B: 65.8 (#43), Qwen3 Max: 62.4 (#76)

Writing & Preference benchmarks
BenchmarkQwen3.8 27BQwen3 Max
LMArena Text14411439
LMArena Creative Writing13841402
LMArena Multi-Turn14411446
EQ-Bench Creative Writing1671—

Frequently asked questions

Is Qwen3.8 27B better than Qwen3 Max?

Qwen3.8 27B is the stronger model overall, scoring 46.0 to 43.7 on the Noometry Index.

Which is cheaper, Qwen3.8 27B or Qwen3 Max?

Qwen3.8 27B is cheaper. It lists at $0.99 per million input tokens and $1.49 per million output tokens; Qwen3 Max lists at $1.20 and $6.

Is Qwen3.8 27B or Qwen3 Max better for coding?

Qwen3.8 27B scores higher on coding benchmarks: 50.5 versus 43.0 in the Noometry coding category.

Which has the bigger context window?

Both accept 262K tokens.

How many benchmarks do Qwen3.8 27B and Qwen3 Max share?

21 benchmarks have published results for both models. Qwen3.8 27B has 31 scored results on Noometry and Qwen3 Max has 33.

Related comparisons

Go deeper