Model comparison

Llama 4 Maverick vs Qwen Max

Qwen Max is the stronger model overall, scoring 34.7 to 30.9 on the Noometry Index. Llama 4 Maverick costs 9.2× less per token, which makes it the better buy when Qwen Max's lead doesn't matter for your workload.

Last verified . 23 shared benchmarks.

Llama 4 Maverick Meta

30.9

Rank #282 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 23 benchmarks with published results for both. Llama 4 Maverick scores higher in 4 categories and Qwen Max in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen Max leads 25.1 to 10.1.
  • The biggest single-benchmark swing is Fiction.LiveBench: 46.2% for Llama 4 Maverick and 66.7% for Qwen Max.
  • Llama 4 Maverick is cheaper at $0.19 / $0.65 per million input/output tokens, against $1.60 / $6.40 for Qwen Max.
  • Llama 4 Maverick accepts more context: 128K tokens versus 33K.
  • Llama 4 Maverick has downloadable open weights; the other is API-only.

Side by side

Llama 4 Maverick and Qwen Max specifications
Llama 4 MaverickQwen Max
ProviderMetaAlibaba (Qwen)
Noometry Index30.934.7
Released2025-04-052024-04-03
WeightsOpenProprietary
Context window128K33K
Max output4K8K
Input $ / M tokens$0.19$1.60
Output $ / M tokens$0.65$6.40
Results tracked5423

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen Max leads

Llama 4 Maverick: 26.6 (#324), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkLlama 4 MaverickQwen Max
Aider Polyglot15.6%21.8%
LMArena Coding13021288
SWE-bench Verified (bash only)21%—
SciCode33.1%—
WeirdML24.5%—
BigCodeBench Instruct49.7%—
BigCodeBench Complete61.4%—
ALE-Bench172.97—

Agentic & Tool Use Not comparable

Llama 4 Maverick: 28.2 (#91), Qwen Max: —

Agentic & Tool Use benchmarks
BenchmarkLlama 4 MaverickQwen Max
Berkeley Function Calling Leaderboard37.3%—

Reasoning Qwen Max leads

Llama 4 Maverick: 10.1 (#342), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkLlama 4 MaverickQwen Max
LMArena Hard Prompts12811269
ARC-AGI-20%—
SimpleBench27.7%—
Kagi LLM Benchmark55.9%—
NYT Connections (extended)8%—
ARC-AGI-14.4%—
CritPt0%—
EnigmaEval0.6%—
DTBench61.9%—
LMCA15.9%—
Epoch Capabilities Index132.2—
ForecastBench57.5—

Math Llama 4 Maverick leads

Llama 4 Maverick: 26.0 (#262), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkLlama 4 MaverickQwen Max
OTIS Mock AIME 2024-202520.6%16.1%
LMArena Math12991275
MATH Level 573%67.2%
FrontierMath (Feb 2025 set)0.7%1%
Omni-MATH42.2%—

Knowledge Llama 4 Maverick leads

Llama 4 Maverick: 33.4 (#204), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkLlama 4 MaverickQwen Max
GPQA Diamond67%56.1%
LMArena Expert12591248
Humanity's Last Exam5.7%—
MMLU-Pro81%—
Confabulations22.6%—
Vectara Hallucination Rate8.2%—
GPQA (HELM)65%—

Multimodal Not comparable

Llama 4 Maverick: 31.6 (#105), Qwen Max: —

Multimodal benchmarks
BenchmarkLlama 4 MaverickQwen Max
LMArena Vision1142—
GeoBench52%—
SpatialViz-Bench31.8%—

Multilingual Too close to call

Llama 4 Maverick: 42.2 (#195), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkLlama 4 MaverickQwen Max
LMArena Non-English12691263
LMArena Chinese12771254
LMArena French12591330
LMArena German12911254
LMArena Japanese12071205
LMArena Korean12031142
LMArena Russian12861274
LMArena Spanish12931290

Instruction Following Llama 4 Maverick leads

Llama 4 Maverick: 71.7 (#146), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkLlama 4 MaverickQwen Max
LMArena Instruction Following12671262
IFEval90.8%—

Long Context Qwen Max leads

Llama 4 Maverick: 31.4 (#279), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkLlama 4 MaverickQwen Max
Fiction.LiveBench46.2%66.7%
LMArena Longer Query12801288

Writing & Preference Qwen Max leads

Llama 4 Maverick: 38.8 (#252), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkLlama 4 MaverickQwen Max
LMArena Text12871282
LMArena Creative Writing12671248
LMArena Multi-Turn12891277
Short-Story Creative Writing62%—
EQ-Bench Creative Writing860—
WildBench80%—

Frequently asked questions

Is Llama 4 Maverick better than Qwen Max?

Qwen Max is the stronger model overall, scoring 34.7 to 30.9 on the Noometry Index. Llama 4 Maverick costs 9.2× less per token, which makes it the better buy when Qwen Max's lead doesn't matter for your workload.

Which is cheaper, Llama 4 Maverick or Qwen Max?

Llama 4 Maverick is cheaper. It lists at $0.19 per million input tokens and $0.65 per million output tokens; Qwen Max lists at $1.60 and $6.40.

Is Llama 4 Maverick or Qwen Max better for coding?

Qwen Max scores higher on coding benchmarks: 30.7 versus 26.6 in the Noometry coding category.

Which has the bigger context window?

Llama 4 Maverick does, with 128K tokens against 33K.

How many benchmarks do Llama 4 Maverick and Qwen Max share?

23 benchmarks have published results for both models. Llama 4 Maverick has 54 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper