Model comparison

Gemini 3.1 Pro Preview vs Qwen3.6 Max Preview

Gemini 3.1 Pro Preview is the stronger model overall, scoring 56.7 to 51.5 on the Noometry Index. Qwen3.6 Max Preview costs 1.5× less per token, which makes it the better buy when Gemini 3.1 Pro Preview's lead doesn't matter for your workload.

Last verified . 29 shared benchmarks.

Gemini 3.1 Pro Preview Google

56.7

Rank #23 Confirmed

Qwen3.6 Max Preview Alibaba (Qwen)

51.5

Rank #43 Confirmed

Summary

  • They share 29 benchmarks with published results for both. Gemini 3.1 Pro Preview scores higher in 7 categories and Qwen3.6 Max Preview in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 3.1 Pro Preview leads 71.7 to 41.7.
  • The biggest single-benchmark swing is Chess Puzzles: 55% for Gemini 3.1 Pro Preview and 20% for Qwen3.6 Max Preview.
  • Qwen3.6 Max Preview is cheaper at $1.30 / $7.80 per million input/output tokens, against $2 / $12 for Gemini 3.1 Pro Preview.
  • Gemini 3.1 Pro Preview accepts more context: 1.05M tokens versus 262K.

Side by side

Gemini 3.1 Pro Preview and Qwen3.6 Max Preview specifications
Gemini 3.1 Pro PreviewQwen3.6 Max Preview
ProviderGoogleAlibaba (Qwen)
Noometry Index56.751.5
Released2026-02-192026-04-20
WeightsProprietaryProprietary
Context window1.05M262K
Max output66K66K
Input $ / M tokens$2$1.30
Output $ / M tokens$12$7.80
Results tracked7129

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.6 Max Preview leads

Gemini 3.1 Pro Preview: 42.5 (#99), Qwen3.6 Max Preview: 48.7 (#54)

Coding benchmarks
BenchmarkGemini 3.1 Pro PreviewQwen3.6 Max Preview
SWE-bench Verified75.6%76.7%
LMArena WebDev14471482
LMArena Coding14841471
DeepSWE11.7%—
SciCode58.9%—
GSO22.6%—
WeirdML72.1%—
MirrorCode8.9%—
ALE-Bench1,161—
AlgoTune2.02—

Agentic & Tool Use Not comparable

Gemini 3.1 Pro Preview: 37.7 (#34), Qwen3.6 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkGemini 3.1 Pro PreviewQwen3.6 Max Preview
Vending-Bench 23,7744,254
Terminal-Bench80.2%—
APEX-Agents35.3%—
τ²-bench Banking26%—
DeepResearch Bench47.8%—
PostTrainBench22%—
BALROG57%—
ExploitBench26.1%—
GBAEval0.8%—
GDP.pdf17%—
LMArena Search1211—
METR Time Horizons77%—

Reasoning Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 71.7 (#12), Qwen3.6 Max Preview: 41.7 (#53)

Reasoning benchmarks
BenchmarkGemini 3.1 Pro PreviewQwen3.6 Max Preview
SimpleBench79.6%63%
NYT Connections (extended)97.4%74.1%
Chess Puzzles55%20%
LMArena Hard Prompts14851457
Mystery Game Puzzles34%19%
DTBench97.1%87.2%
LMCA53.8%42.5%
Epoch Capabilities Index154.77149.24
ARC-AGI-277.1%—
ARC-AGI-198%—
CritPt17.7%—
EnigmaEval36.8%—
Thematic Generalization79.4%—
EBR-Bench14.3%—
ForecastBench59—

Math Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 62.1 (#34), Qwen3.6 Max Preview: 54.1 (#46)

Math benchmarks
BenchmarkGemini 3.1 Pro PreviewQwen3.6 Max Preview
OTIS Mock AIME 2024-202595.6%91.1%
LMArena Math14851465
FrontierMath (Feb 2025 set)36.9%23.1%
FrontierMath Tier 4 (v1)16.7%4.2%
FrontierMath (Tiers 1-3)59.6%—
FrontierMath Tier 426.8%—
MathArena Final-Answer Competitions86.5%—
ProofBench26%—

Knowledge Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 71.8 (#3), Qwen3.6 Max Preview: 57.6 (#39)

Knowledge benchmarks
BenchmarkGemini 3.1 Pro PreviewQwen3.6 Max Preview
GPQA Diamond94.4%87.4%
SimpleQA Verified73.5%52%
LMArena Expert14851478
Humanity's Last Exam46.4%—
Vectara Hallucination Rate10.4%—

Multimodal Not comparable

Gemini 3.1 Pro Preview: 37.9 (#69), Qwen3.6 Max Preview: —

Multimodal benchmarks
BenchmarkGemini 3.1 Pro PreviewQwen3.6 Max Preview
LMArena Vision1296—
Blueprint-Bench 226.5%—
Furniture Assembly26.7%—
LMArena Document1444—

Multilingual Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 57.0 (#12), Qwen3.6 Max Preview: 54.2 (#48)

Multilingual benchmarks
BenchmarkGemini 3.1 Pro PreviewQwen3.6 Max Preview
LMArena Non-English14771437
LMArena Chinese15291487
LMArena French14871449
LMArena Russian14981445
LMArena Spanish14791454
LMArena German1491—
LMArena Japanese1493—
LMArena Korean1455—

Instruction Following Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 77.0 (#32), Qwen3.6 Max Preview: 75.7 (#55)

Instruction Following benchmarks
BenchmarkGemini 3.1 Pro PreviewQwen3.6 Max Preview
LMArena Instruction Following14661438

Long Context Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 47.4 (#18), Qwen3.6 Max Preview: 44.6 (#61)

Long Context benchmarks
BenchmarkGemini 3.1 Pro PreviewQwen3.6 Max Preview
LMArena Longer Query14831457
CL-bench20.8%—
CL-bench Life16.9%—

Writing & Preference Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 66.1 (#37), Qwen3.6 Max Preview: 63.8 (#60)

Writing & Preference benchmarks
BenchmarkGemini 3.1 Pro PreviewQwen3.6 Max Preview
LMArena Text14811447
LMArena Creative Writing14821435
LMArena Multi-Turn14881456
EQ-Bench Creative Writing1491—
EQ-Bench 41142—

Frequently asked questions

Is Gemini 3.1 Pro Preview better than Qwen3.6 Max Preview?

Gemini 3.1 Pro Preview is the stronger model overall, scoring 56.7 to 51.5 on the Noometry Index. Qwen3.6 Max Preview costs 1.5× less per token, which makes it the better buy when Gemini 3.1 Pro Preview's lead doesn't matter for your workload.

Which is cheaper, Gemini 3.1 Pro Preview or Qwen3.6 Max Preview?

Qwen3.6 Max Preview is cheaper. It lists at $1.30 per million input tokens and $7.80 per million output tokens; Gemini 3.1 Pro Preview lists at $2 and $12.

Is Gemini 3.1 Pro Preview or Qwen3.6 Max Preview better for coding?

Qwen3.6 Max Preview scores higher on coding benchmarks: 48.7 versus 42.5 in the Noometry coding category.

Which has the bigger context window?

Gemini 3.1 Pro Preview does, with 1.05M tokens against 262K.

How many benchmarks do Gemini 3.1 Pro Preview and Qwen3.6 Max Preview share?

29 benchmarks have published results for both models. Gemini 3.1 Pro Preview has 71 scored results on Noometry and Qwen3.6 Max Preview has 29.

Related comparisons

Go deeper