Model comparison

Llama 4 Scout vs Qwen2.5-Max

Qwen2.5-Max is the stronger model overall, scoring 40.7 to 27.7 on the Noometry Index.

Last verified . 18 shared benchmarks.

Llama 4 Scout Meta

27.7

Rank #330 Confirmed

Qwen2.5-Max Alibaba (Qwen)

40.7

Rank #146 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Llama 4 Scout scores higher in 0 categories and Qwen2.5-Max in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Qwen2.5-Max leads 41.8 to 20.2.
  • Llama 4 Scout has downloadable open weights; the other is API-only.

Side by side

Llama 4 Scout and Qwen2.5-Max specifications
Llama 4 ScoutQwen2.5-Max
ProviderMetaAlibaba (Qwen)
Noometry Index27.740.7
Released2025-04-052025-01-25
WeightsOpenProprietary
Context window128K—
Max output4K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.30—
Results tracked4327

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5-Max leads

Llama 4 Scout: 20.2 (#339), Qwen2.5-Max: 41.8 (#117)

Coding benchmarks
BenchmarkLlama 4 ScoutQwen2.5-Max
LMArena Coding12861359
SWE-bench Verified (bash only)9.1%—
SciCode17%—
LiveBench Coding—64.4%
BigCodeBench Complete43.1%—

Agentic & Tool Use Not comparable

Llama 4 Scout: 24.6 (#119), Qwen2.5-Max: —

Agentic & Tool Use benchmarks
BenchmarkLlama 4 ScoutQwen2.5-Max
Berkeley Function Calling Leaderboard28.1%—

Reasoning Qwen2.5-Max leads

Llama 4 Scout: 9.1 (#345), Qwen2.5-Max: 25.6 (#147)

Reasoning benchmarks
BenchmarkLlama 4 ScoutQwen2.5-Max
LMArena Hard Prompts12661360
Epoch Capabilities Index129.64132.53
ARC-AGI-20%—
Kagi LLM Benchmark36.9%—
ARC-AGI-10.5%—
CritPt0%—
LiveBench Reasoning—51.4%
DTBench57.9%—
LiveBench Data Analysis—67.9%
LMCA12%—
ForecastBench57.5—
LiveBench—62.3%

Math Qwen2.5-Max leads

Llama 4 Scout: 19.6 (#286), Qwen2.5-Max: 36.9 (#162)

Math benchmarks
BenchmarkLlama 4 ScoutQwen2.5-Max
LMArena Math12871369
OTIS Mock AIME 2024-20257.8%—
Omni-MATH37.3%—
LiveBench Math—58.4%
MATH Level 562.3%—
FrontierMath (Feb 2025 set)0%—

Knowledge Qwen2.5-Max leads

Llama 4 Scout: 31.9 (#217), Qwen2.5-Max: 35.3 (#186)

Knowledge benchmarks
BenchmarkLlama 4 ScoutQwen2.5-Max
LMArena Expert12351337
GPQA Diamond51.8%—
MMLU-Pro74.2%—
Confabulations—21.8%
Vectara Hallucination Rate7.7%—
GPQA (HELM)50.7%—

Multimodal Not comparable

Llama 4 Scout: 32.2 (#102), Qwen2.5-Max: —

Multimodal benchmarks
BenchmarkLlama 4 ScoutQwen2.5-Max
LMArena Vision1118—
SpatialViz-Bench34.2%—

Multilingual Qwen2.5-Max leads

Llama 4 Scout: 41.0 (#212), Qwen2.5-Max: 48.1 (#146)

Multilingual benchmarks
BenchmarkLlama 4 ScoutQwen2.5-Max
LMArena Non-English12521352
LMArena Chinese12551382
LMArena French12821396
LMArena German12721350
LMArena Japanese12061300
LMArena Korean12071304
LMArena Russian12631353
LMArena Spanish12781377

Instruction Following Qwen2.5-Max leads

Llama 4 Scout: 65.8 (#217), Qwen2.5-Max: 71.3 (#152)

Instruction Following benchmarks
BenchmarkLlama 4 ScoutQwen2.5-Max
LMArena Instruction Following12481335
LiveBench Instruction Following—75.3%
IFEval81.8%—

Long Context Qwen2.5-Max leads

Llama 4 Scout: 27.5 (#294), Qwen2.5-Max: 41.4 (#142)

Long Context benchmarks
BenchmarkLlama 4 ScoutQwen2.5-Max
LMArena Longer Query12651358
Fiction.LiveBench36%—

Writing & Preference Qwen2.5-Max leads

Llama 4 Scout: 37.0 (#261), Qwen2.5-Max: 55.4 (#146)

Writing & Preference benchmarks
BenchmarkLlama 4 ScoutQwen2.5-Max
LMArena Text12791367
LMArena Creative Writing12491339
LMArena Multi-Turn12801364
Short-Story Creative Writing—72.9%
EQ-Bench Creative Writing783—
WildBench78%—
LiveBench Language—56.3%

Frequently asked questions

Is Llama 4 Scout better than Qwen2.5-Max?

Qwen2.5-Max is the stronger model overall, scoring 40.7 to 27.7 on the Noometry Index.

Is Llama 4 Scout or Qwen2.5-Max better for coding?

Qwen2.5-Max scores higher on coding benchmarks: 41.8 versus 20.2 in the Noometry coding category.

How many benchmarks do Llama 4 Scout and Qwen2.5-Max share?

18 benchmarks have published results for both models. Llama 4 Scout has 43 scored results on Noometry and Qwen2.5-Max has 27.

Related comparisons

Go deeper