Alibaba (Qwen), proprietary

Qwen3.8 Max

Qwen3.8 Max by Alibaba (Qwen) ranks 22nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 56.8. Its strongest category is agentic & tool use, where it ranks 14th. API pricing starts at $2 per million input tokens and $6 per million output tokens, with a 1M-token context window.

Last verified

Specifications

Noometry rank
#22 of 354
Index score
56.8
Evidence
Confirmed 39 results
Released
August 2, 2026
Weights
Proprietary
Reasoning
Yes
Context window
1M
Max output
131K
Input price
$2 / M
Output price
$6 / M
Blended price
$3 / M
Output speed
Not measured
Value
#154 of 219
Knowledge cutoff
Unknown
Input
text, image, video, pdf

Category scores

Each category score combines every public result we have in that category.

Qwen3.8 Max category scores
  1. Coding 53.5
  2. Agentic & Tool Use 45.4
  3. Reasoning 54.4
  4. Math 73.2
  5. Knowledge 61.7
  6. Multimodal 37.2
  7. Multilingual 56.7
  8. Instruction Following 77.6
  9. Long Context 45.6
  10. Writing & Preference 67.1
Qwen3.8 Max category ranks
CategoryScoreRankResults
Coding53.5#295
Agentic & Tool Use45.4#143
Reasoning54.4#267
Math73.2#205
Knowledge61.7#273
Multimodal37.2#752
Multilingual56.7#181
Instruction Following77.6#171
Long Context45.6#311
Writing & Preference67.1#303

Strengths and weaknesses

Categories where Qwen3.8 Max places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Qwen3.8 Max: strongest categories
CategoryScorevs medianRank
Instruction Following77.6+6.3#17 of 305, top 6%
Multilingual56.7+9.3#18 of 297, top 7%
Math73.2+36.6#20 of 327, top 7%

Weakest categories

Qwen3.8 Max: weakest categories
CategoryScorevs medianRank
Multimodal37.2−1.4#75 of 128, top 59%
Long Context45.6+4.7#31 of 296, top 11%
Writing & Preference67.1+13.3#30 of 312, top 10%

Closest competitors

The models ranked just above and below Qwen3.8 Max. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Qwen3.8 Max
ModelRankScoreBlended $/MSpeed
GPT-5.4 Pro#1858.9$67.50—Compare
Claude Opus 4.7#1958.3$1033Compare
Claude Opus 4.6#2058.2$1019Compare
Grok 4.6#2156.9$3—Compare
Gemini 3.1 Pro Preview#2356.7$4.50—Compare
Gemini 4 Argon#2456.5——Compare
Grok 4.5#2555.0$34Compare
GLM-5.3#2654.8$2.15—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Qwen3.8 Max Coding benchmark results
BenchmarkScorePositionSettingSourceDate
DeepSWE57.5%#18 of 29, top 63%xhighEpoch AI
LMArena WebDev1674#9 of 113, top 8%LMArena2026-10-08
LMArena WebDev1672LMArena2026-10-08
FrontierSWE17.8%#16 of 18, top 89%xhighEpoch AI
FrontierSWE15.8%xhighEpoch AI
SciCode52.1%Epoch AI
SciCode53.2%#33 of 121, top 28%Epoch AI
LMArena Coding1502#19 of 294, top 7%LMArena2026-10-08

Agentic & Tool Use

Qwen3.8 Max Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
APEX-Agents63.3%#11 of 49, top 23%Epoch AI
τ²-bench Banking55.1%Best of 26xhighτ²-bench2026-08-04
GDP.pdf23.2%#15 of 36, top 42%xhighEpoch AI

Reasoning

Qwen3.8 Max Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
NYT Connections (extended)88.3%#24 of 91, top 27%Lech Mazur benchmarks
CritPt17.7%Epoch AI
CritPt20%#24 of 134, top 18%Epoch AI
Chess Puzzles29%xhighEpoch AI2026-08-04
Chess Puzzles40%#22 of 129, top 18%xhighEpoch AI2026-09-02
LMArena Hard Prompts1496#11 of 297, top 4%LMArena2026-10-08
Mystery Game Puzzles38%#14 of 74, top 19%xhighEpoch AI2026-08-05
DTBench92%#29 of 151, top 20%xhighEpoch AI
LMCA46.2%#30 of 125, top 24%xhighEpoch AI
Epoch Capabilities Index156.41#22 of 213, top 11%Epoch AI2026-08-02
Epoch Capabilities Index155.05Epoch AI2026-09-01

Math

Qwen3.8 Max Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)65.6%xhighEpoch AI2026-09-02
FrontierMath (Tiers 1-3)74.7%#19 of 81, top 24%xhighEpoch AI2026-08-04
FrontierMath Tier 434.1%xhighEpoch AI2026-09-02
FrontierMath Tier 446.3%#21 of 63, top 34%xhighEpoch AI2026-08-04
OTIS Mock AIME 2024-202599.4%xhighEpoch AI2026-08-04
OTIS Mock AIME 2024-2025100%#11 of 173, top 7%xhighEpoch AI2026-09-02
ProofBench58%#22 of 77, top 29%Epoch AI
LMArena Math1499#13 of 285, top 5%LMArena2026-10-08

Knowledge

Qwen3.8 Max Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond92.3%xhighEpoch AI2026-09-02
GPQA Diamond92.7%#21 of 186, top 12%xhighEpoch AI2026-08-04
SimpleQA Verified45.8%xhighEpoch AI2026-08-27
SimpleQA Verified47.3%#31 of 77, top 41%xhighEpoch AI2026-09-02
LMArena Expert1507#18 of 273, top 7%LMArena2026-10-08

Multimodal

Qwen3.8 Max Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1314#9 of 122, top 8%LMArena2026-10-09
Furniture Assembly20%#31 of 31, top 100%xhighEpoch AI2026-09-10

Multilingual

Qwen3.8 Max Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1472#18 of 297, top 7%LMArena2026-10-08
LMArena Chinese1538#9 of 285, top 4%LMArena2026-10-08
LMArena French1503#11 of 223, top 5%LMArena2026-10-08
LMArena German1483#19 of 231, top 9%LMArena2026-10-08
LMArena Japanese1467#19 of 211, top 10%LMArena2026-10-08
LMArena Korean1461#9 of 213, top 5%LMArena2026-10-08
LMArena Russian1481#19 of 283, top 7%LMArena2026-10-08
LMArena Spanish1492#10 of 226, top 5%LMArena2026-10-08

Instruction Following

Qwen3.8 Max Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1479#14 of 298, top 5%LMArena2026-10-08

Long Context

Qwen3.8 Max Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1489#13 of 291, top 5%LMArena2026-10-08

Writing & Preference

Qwen3.8 Max Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1483#12 of 297, top 5%LMArena2026-10-08
LMArena Creative Writing1479#12 of 295, top 5%LMArena2026-10-08
LMArena Multi-Turn1489#11 of 295, top 4%LMArena2026-10-08

API pricing by provider

Qwen3.8 Max API prices
RouteInput $/MOutput $/MCached input $/MChecked
alibaba$2$6$0.252026-10-10
deepinfra$1.65$4.95$0.212026-10-10
fireworks$2$6$0.252026-10-10
openrouter$2$6$0.252026-10-10

Compare Qwen3.8 Max

Other Alibaba (Qwen) models

Frequently asked questions

How good is Qwen3.8 Max?

Qwen3.8 Max by Alibaba (Qwen) ranks 22nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 56.8. Its strongest category is agentic & tool use, where it ranks 14th. API pricing starts at $2 per million input tokens and $6 per million output tokens, with a 1M-token context window.

How much does Qwen3.8 Max cost?

Qwen3.8 Max costs $2 per million input tokens and $6 per million output tokens on Alibaba (Qwen)'s own API, with cached input at $0.25.

What is Qwen3.8 Max's context window?

Qwen3.8 Max accepts up to 1M tokens of input and can write up to 131K tokens in one response.

Is Qwen3.8 Max open source?

No. Qwen3.8 Max is proprietary and available only through Alibaba (Qwen)'s API and partner platforms.

What are Qwen3.8 Max's strengths and weaknesses?

Relative to other ranked models, Qwen3.8 Max places best in instruction following, multilingual, math and lowest in multimodal, long context, writing & preference.

What is Qwen3.8 Max best at?

Its best category is agentic & tool use, where it ranks 14th on Noometry.