Alibaba (Qwen), proprietary
Qwen3.8 Max
Qwen3.8 Max by Alibaba (Qwen) ranks 22nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 56.8. Its strongest category is agentic & tool use, where it ranks 14th. API pricing starts at $2 per million input tokens and $6 per million output tokens, with a 1M-token context window.
Last verified
Specifications
- Noometry rank
- #22 of 354
- Index score
- 56.8
- Evidence
- Confirmed 39 results
- Provider
Alibaba (Qwen)
- Released
- August 2, 2026
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 1M
- Max output
- 131K
- Input price
- $2 / M
- Output price
- $6 / M
- Blended price
- $3 / M
- Output speed
- Not measured
- Value
- #154 of 219
- Knowledge cutoff
- Unknown
- Input
- text, image, video, pdf
Category scores
Each category score combines every public result we have in that category.
- Coding 53.5
- Agentic & Tool Use 45.4
- Reasoning 54.4
- Math 73.2
- Knowledge 61.7
- Multimodal 37.2
- Multilingual 56.7
- Instruction Following 77.6
- Long Context 45.6
- Writing & Preference 67.1
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 53.5 | #29 | 5 |
| Agentic & Tool Use | 45.4 | #14 | 3 |
| Reasoning | 54.4 | #26 | 7 |
| Math | 73.2 | #20 | 5 |
| Knowledge | 61.7 | #27 | 3 |
| Multimodal | 37.2 | #75 | 2 |
| Multilingual | 56.7 | #18 | 1 |
| Instruction Following | 77.6 | #17 | 1 |
| Long Context | 45.6 | #31 | 1 |
| Writing & Preference | 67.1 | #30 | 3 |
Strengths and weaknesses
Categories where Qwen3.8 Max places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Instruction Following | 77.6 | +6.3 | #17 of 305, top 6% |
| Multilingual | 56.7 | +9.3 | #18 of 297, top 7% |
| Math | 73.2 | +36.6 | #20 of 327, top 7% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multimodal | 37.2 | −1.4 | #75 of 128, top 59% |
| Long Context | 45.6 | +4.7 | #31 of 296, top 11% |
| Writing & Preference | 67.1 | +13.3 | #30 of 312, top 10% |
Closest competitors
The models ranked just above and below Qwen3.8 Max. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| GPT-5.4 Pro | #18 | 58.9 | $67.50 | — | Compare |
| Claude Opus 4.7 | #19 | 58.3 | $10 | 33 | Compare |
| Claude Opus 4.6 | #20 | 58.2 | $10 | 19 | Compare |
| Grok 4.6 | #21 | 56.9 | $3 | — | Compare |
| Gemini 3.1 Pro Preview | #23 | 56.7 | $4.50 | — | Compare |
| Gemini 4 Argon | #24 | 56.5 | — | — | Compare |
| Grok 4.5 | #25 | 55.0 | $3 | 4 | Compare |
| GLM-5.3 | #26 | 54.8 | $2.15 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| DeepSWE | 57.5% | #18 of 29, top 63% | xhigh | Epoch AI | |
| LMArena WebDev | 1674 | #9 of 113, top 8% | LMArena | 2026-10-08 | |
| LMArena WebDev | 1672 | LMArena | 2026-10-08 | ||
| FrontierSWE | 17.8% | #16 of 18, top 89% | xhigh | Epoch AI | |
| FrontierSWE | 15.8% | xhigh | Epoch AI | ||
| SciCode | 52.1% | Epoch AI | |||
| SciCode | 53.2% | #33 of 121, top 28% | Epoch AI | ||
| LMArena Coding | 1502 | #19 of 294, top 7% | LMArena | 2026-10-08 |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| APEX-Agents | 63.3% | #11 of 49, top 23% | Epoch AI | ||
| τ²-bench Banking | 55.1% | Best of 26 | xhigh | τ²-bench | 2026-08-04 |
| GDP.pdf | 23.2% | #15 of 36, top 42% | xhigh | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| NYT Connections (extended) | 88.3% | #24 of 91, top 27% | Lech Mazur benchmarks | ||
| CritPt | 17.7% | Epoch AI | |||
| CritPt | 20% | #24 of 134, top 18% | Epoch AI | ||
| Chess Puzzles | 29% | xhigh | Epoch AI | 2026-08-04 | |
| Chess Puzzles | 40% | #22 of 129, top 18% | xhigh | Epoch AI | 2026-09-02 |
| LMArena Hard Prompts | 1496 | #11 of 297, top 4% | LMArena | 2026-10-08 | |
| Mystery Game Puzzles | 38% | #14 of 74, top 19% | xhigh | Epoch AI | 2026-08-05 |
| DTBench | 92% | #29 of 151, top 20% | xhigh | Epoch AI | |
| LMCA | 46.2% | #30 of 125, top 24% | xhigh | Epoch AI | |
| Epoch Capabilities Index | 156.41 | #22 of 213, top 11% | Epoch AI | 2026-08-02 | |
| Epoch Capabilities Index | 155.05 | Epoch AI | 2026-09-01 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 65.6% | xhigh | Epoch AI | 2026-09-02 | |
| FrontierMath (Tiers 1-3) | 74.7% | #19 of 81, top 24% | xhigh | Epoch AI | 2026-08-04 |
| FrontierMath Tier 4 | 34.1% | xhigh | Epoch AI | 2026-09-02 | |
| FrontierMath Tier 4 | 46.3% | #21 of 63, top 34% | xhigh | Epoch AI | 2026-08-04 |
| OTIS Mock AIME 2024-2025 | 99.4% | xhigh | Epoch AI | 2026-08-04 | |
| OTIS Mock AIME 2024-2025 | 100% | #11 of 173, top 7% | xhigh | Epoch AI | 2026-09-02 |
| ProofBench | 58% | #22 of 77, top 29% | Epoch AI | ||
| LMArena Math | 1499 | #13 of 285, top 5% | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 92.3% | xhigh | Epoch AI | 2026-09-02 | |
| GPQA Diamond | 92.7% | #21 of 186, top 12% | xhigh | Epoch AI | 2026-08-04 |
| SimpleQA Verified | 45.8% | xhigh | Epoch AI | 2026-08-27 | |
| SimpleQA Verified | 47.3% | #31 of 77, top 41% | xhigh | Epoch AI | 2026-09-02 |
| LMArena Expert | 1507 | #18 of 273, top 7% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1314 | #9 of 122, top 8% | LMArena | 2026-10-09 | |
| Furniture Assembly | 20% | #31 of 31, top 100% | xhigh | Epoch AI | 2026-09-10 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1472 | #18 of 297, top 7% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1538 | #9 of 285, top 4% | LMArena | 2026-10-08 | |
| LMArena French | 1503 | #11 of 223, top 5% | LMArena | 2026-10-08 | |
| LMArena German | 1483 | #19 of 231, top 9% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1467 | #19 of 211, top 10% | LMArena | 2026-10-08 | |
| LMArena Korean | 1461 | #9 of 213, top 5% | LMArena | 2026-10-08 | |
| LMArena Russian | 1481 | #19 of 283, top 7% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1492 | #10 of 226, top 5% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1479 | #14 of 298, top 5% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1489 | #13 of 291, top 5% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1483 | #12 of 297, top 5% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1479 | #12 of 295, top 5% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1489 | #11 of 295, top 4% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| alibaba | $2 | $6 | $0.25 | 2026-10-10 |
| deepinfra | $1.65 | $4.95 | $0.21 | 2026-10-10 |
| fireworks | $2 | $6 | $0.25 | 2026-10-10 |
| openrouter | $2 | $6 | $0.25 | 2026-10-10 |
Compare Qwen3.8 Max
- Qwen3.8 Max vs Qwen3.7 Max
- Qwen3.8 Max vs Grok 4.6
- Qwen3.8 Max vs Gemini 3.1 Pro Preview
- Qwen3.8 Max vs Claude Opus 4.6
- Qwen3.8 Max vs Gemini 4 Argon
- Qwen3.8 Max vs Claude Opus 4.7
- Qwen3.8 Max vs Grok 4.5
- Qwen3.8 Max vs GPT-6 Astra
- Qwen3.8 Max vs Claude Fable 5.1
- Qwen3.8 Max vs Gemini 3.8 Flash
- Qwen3.8 Max vs Kimi K3
- Qwen3.8 Max vs GLM-5.3
- Qwen3.8 Max vs Muse Spark 1.3
- Qwen3.8 Max vs DeepSeek V4 Pro
Other Alibaba (Qwen) models
- Qwen3.7 Max51.5
- Qwen3.6 Max Preview51.5
- Qwen3.6 Plus47.5
- Qwen3.5 397B-A17B46.0
- Qwen3.8 27B46.0
- Qwen3.5 Max Preview45.3
- Qwen3.7 Plus45.3
- Qwen3 Max43.7
Frequently asked questions
How good is Qwen3.8 Max?
Qwen3.8 Max by Alibaba (Qwen) ranks 22nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 56.8. Its strongest category is agentic & tool use, where it ranks 14th. API pricing starts at $2 per million input tokens and $6 per million output tokens, with a 1M-token context window.
How much does Qwen3.8 Max cost?
Qwen3.8 Max costs $2 per million input tokens and $6 per million output tokens on Alibaba (Qwen)'s own API, with cached input at $0.25.
What is Qwen3.8 Max's context window?
Qwen3.8 Max accepts up to 1M tokens of input and can write up to 131K tokens in one response.
Is Qwen3.8 Max open source?
No. Qwen3.8 Max is proprietary and available only through Alibaba (Qwen)'s API and partner platforms.
What are Qwen3.8 Max's strengths and weaknesses?
Relative to other ranked models, Qwen3.8 Max places best in instruction following, multilingual, math and lowest in multimodal, long context, writing & preference.
What is Qwen3.8 Max best at?
Its best category is agentic & tool use, where it ranks 14th on Noometry.