Model comparison

Llama 3.2 1B vs Qwen3.5 Max Preview

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 20.1 on the Noometry Index.

Last verified . 13 shared benchmarks.

Llama 3.2 1B Meta

20.1

Rank #354 Confirmed

Qwen3.5 Max Preview Alibaba (Qwen)

45.3

Rank #71 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Llama 3.2 1B scores higher in 0 categories and Qwen3.5 Max Preview in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen3.5 Max Preview leads 66.0 to 21.3.
  • Llama 3.2 1B has downloadable open weights; the other is API-only.

Side by side

Llama 3.2 1B and Qwen3.5 Max Preview specifications
Llama 3.2 1BQwen3.5 Max Preview
ProviderMetaAlibaba (Qwen)
Noometry Index20.145.3
Released2024-09-24—
WeightsOpenProprietary
Context window60K—
Max output54K—
Input $ / M tokens$0.027—
Output $ / M tokens$0.20—
Results tracked2217

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5 Max Preview leads

Llama 3.2 1B: 21.1 (#338), Qwen3.5 Max Preview: 44.0 (#77)

Coding benchmarks
BenchmarkLlama 3.2 1BQwen3.5 Max Preview
LMArena Coding10701487
BigCodeBench Instruct8.2%—
BigCodeBench Complete11.3%—

Agentic & Tool Use Not comparable

Llama 3.2 1B: 14.6 (#150), Qwen3.5 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 1BQwen3.5 Max Preview
Berkeley Function Calling Leaderboard10.8%—
BALROG6.6%—

Reasoning Qwen3.5 Max Preview leads

Llama 3.2 1B: 16.2 (#308), Qwen3.5 Max Preview: 30.8 (#84)

Reasoning benchmarks
BenchmarkLlama 3.2 1BQwen3.5 Max Preview
LMArena Hard Prompts10441483
Chess Puzzles0%—
Epoch Capabilities Index101.99—

Math Qwen3.5 Max Preview leads

Llama 3.2 1B: 10.4 (#313), Qwen3.5 Max Preview: 40.1 (#94)

Math benchmarks
BenchmarkLlama 3.2 1BQwen3.5 Max Preview
LMArena Math10861474
OTIS Mock AIME 2024-20250.6%—

Knowledge Qwen3.5 Max Preview leads

Llama 3.2 1B: 7.2 (#312), Qwen3.5 Max Preview: 41.8 (#107)

Knowledge benchmarks
BenchmarkLlama 3.2 1BQwen3.5 Max Preview
LMArena Expert10071489
GPQA Diamond23.9%—

Multilingual Qwen3.5 Max Preview leads

Llama 3.2 1B: 23.8 (#292), Qwen3.5 Max Preview: 56.2 (#22)

Multilingual benchmarks
BenchmarkLlama 3.2 1BQwen3.5 Max Preview
LMArena Non-English9731465
LMArena Chinese9591534
LMArena German10141487
LMArena Russian9411471
LMArena French—1484
LMArena Japanese—1495
LMArena Korean—1438
LMArena Spanish—1470

Instruction Following Qwen3.5 Max Preview leads

Llama 3.2 1B: 52.4 (#290), Qwen3.5 Max Preview: 77.0 (#31)

Instruction Following benchmarks
BenchmarkLlama 3.2 1BQwen3.5 Max Preview
LMArena Instruction Following10311467

Long Context Qwen3.5 Max Preview leads

Llama 3.2 1B: 31.9 (#274), Qwen3.5 Max Preview: 45.2 (#45)

Long Context benchmarks
BenchmarkLlama 3.2 1BQwen3.5 Max Preview
LMArena Longer Query10501476

Writing & Preference Qwen3.5 Max Preview leads

Llama 3.2 1B: 21.3 (#310), Qwen3.5 Max Preview: 66.0 (#41)

Writing & Preference benchmarks
BenchmarkLlama 3.2 1BQwen3.5 Max Preview
LMArena Text10551470
LMArena Creative Writing10331464
LMArena Multi-Turn10301478
EQ-Bench Creative Writing200—

Frequently asked questions

Is Llama 3.2 1B better than Qwen3.5 Max Preview?

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 20.1 on the Noometry Index.

Is Llama 3.2 1B or Qwen3.5 Max Preview better for coding?

Qwen3.5 Max Preview scores higher on coding benchmarks: 44.0 versus 21.1 in the Noometry coding category.

How many benchmarks do Llama 3.2 1B and Qwen3.5 Max Preview share?

13 benchmarks have published results for both models. Llama 3.2 1B has 22 scored results on Noometry and Qwen3.5 Max Preview has 17.

Related comparisons

Go deeper