Model comparison

Qwen2.5 72B Instruct vs Qwen3-30B-A3B

Qwen3-30B-A3B is the stronger model overall, scoring 38.9 to 31.9 on the Noometry Index.

Last verified . 24 shared benchmarks.

Qwen2.5 72B Instruct Alibaba (Qwen)

31.9

Rank #267 Confirmed

Qwen3-30B-A3B Alibaba (Qwen)

38.9

Rank #179 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Qwen2.5 72B Instruct scores higher in 2 categories and Qwen3-30B-A3B in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3-30B-A3B leads 37.4 to 19.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 8.1% for Qwen2.5 72B Instruct and 70.3% for Qwen3-30B-A3B.
  • Qwen3-30B-A3B is cheaper at $0.12 / $0.50 per million input/output tokens, against $1.40 / $5.60 for Qwen2.5 72B Instruct.
  • Qwen2.5 72B Instruct accepts more context: 131K tokens versus 41K.

Side by side

Qwen2.5 72B Instruct and Qwen3-30B-A3B specifications
Qwen2.5 72B InstructQwen3-30B-A3B
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index31.938.9
Released2024-092025-04-28
WeightsOpenOpen
Context window131K41K
Max output8K16K
Input $ / M tokens$1.40$0.12
Output $ / M tokens$5.60$0.50
Results tracked4332

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3-30B-A3B leads

Qwen2.5 72B Instruct: 33.2 (#260), Qwen3-30B-A3B: 37.5 (#194)

Coding benchmarks
BenchmarkQwen2.5 72B InstructQwen3-30B-A3B
WeirdML16%29.8%
LMArena Coding12921416
SciCode—33.3%
BigCodeBench Instruct45.8%—
BigCodeBench Complete55.9%—

Agentic & Tool Use Qwen3-30B-A3B leads

Qwen2.5 72B Instruct: 22.1 (#133), Qwen3-30B-A3B: 29.8 (#82)

Agentic & Tool Use benchmarks
BenchmarkQwen2.5 72B InstructQwen3-30B-A3B
Berkeley Function Calling Leaderboard—41.4%
TheAgentCompany5.7%—
BALROG16.2%—
METR Time Horizons35.8%—

Reasoning Too close to call

Qwen2.5 72B Instruct: 22.3 (#199), Qwen3-30B-A3B: 22.2 (#204)

Reasoning benchmarks
BenchmarkQwen2.5 72B InstructQwen3-30B-A3B
LMArena Hard Prompts12711398
DTBench62.9%69.3%
LMCA13.4%22.4%
Epoch Capabilities Index129139.63
Kagi LLM Benchmark—54.9%
CritPt—0.3%
Chess Puzzles—8%
BIG-Bench Hard79.8%—
ForecastBench57.5—
HellaSwag84.8%—
PIQA82.6%—
WinoGrande82.3%—

Math Qwen3-30B-A3B leads

Qwen2.5 72B Instruct: 19.3 (#287), Qwen3-30B-A3B: 37.4 (#157)

Math benchmarks
BenchmarkQwen2.5 72B InstructQwen3-30B-A3B
OTIS Mock AIME 2024-20258.1%70.3%
LMArena Math12831394
MathArena Final-Answer Competitions—47.8%
Omni-MATH33%—
MATH Level 563.2%—

Knowledge Qwen3-30B-A3B leads

Qwen2.5 72B Instruct: 27.0 (#253), Qwen3-30B-A3B: 41.8 (#105)

Knowledge benchmarks
BenchmarkQwen2.5 72B InstructQwen3-30B-A3B
GPQA Diamond49.1%70.1%
Confabulations19.1%12.3%
LMArena Expert12451396
MMLU-Pro63.1%—
GPQA (HELM)42.6%—
ARC (AI2) Challenge94.5%—
MMLU85.3%—
TriviaQA71.9%—

Multilingual Qwen3-30B-A3B leads

Qwen2.5 72B Instruct: 41.0 (#213), Qwen3-30B-A3B: 49.5 (#132)

Multilingual benchmarks
BenchmarkQwen2.5 72B InstructQwen3-30B-A3B
LMArena Non-English12521372
LMArena Chinese12721433
LMArena French12801418
LMArena German12341380
LMArena Japanese11801337
LMArena Korean11881331
LMArena Russian12641370
LMArena Spanish12561404

Instruction Following Qwen3-30B-A3B leads

Qwen2.5 72B Instruct: 65.5 (#221), Qwen3-30B-A3B: 72.0 (#142)

Instruction Following benchmarks
BenchmarkQwen2.5 72B InstructQwen3-30B-A3B
LMArena Instruction Following12541363
IFEval80.6%—

Long Context Qwen2.5 72B Instruct leads

Qwen2.5 72B Instruct: 38.9 (#188), Qwen3-30B-A3B: 31.0 (#283)

Long Context benchmarks
BenchmarkQwen2.5 72B InstructQwen3-30B-A3B
LMArena Longer Query12821379
Fiction.LiveBench—40.6%

Writing & Preference Qwen3-30B-A3B leads

Qwen2.5 72B Instruct: 46.7 (#215), Qwen3-30B-A3B: 55.6 (#143)

Writing & Preference benchmarks
BenchmarkQwen2.5 72B InstructQwen3-30B-A3B
LMArena Text12691384
LMArena Creative Writing12211317
LMArena Multi-Turn12721378
Short-Story Creative Writing—75.3%
WildBench80.2%—

Frequently asked questions

Is Qwen2.5 72B Instruct better than Qwen3-30B-A3B?

Qwen3-30B-A3B is the stronger model overall, scoring 38.9 to 31.9 on the Noometry Index.

Which is cheaper, Qwen2.5 72B Instruct or Qwen3-30B-A3B?

Qwen3-30B-A3B is cheaper. It lists at $0.12 per million input tokens and $0.50 per million output tokens; Qwen2.5 72B Instruct lists at $1.40 and $5.60.

Is Qwen2.5 72B Instruct or Qwen3-30B-A3B better for coding?

Qwen3-30B-A3B scores higher on coding benchmarks: 37.5 versus 33.2 in the Noometry coding category.

Which has the bigger context window?

Qwen2.5 72B Instruct does, with 131K tokens against 41K.

How many benchmarks do Qwen2.5 72B Instruct and Qwen3-30B-A3B share?

24 benchmarks have published results for both models. Qwen2.5 72B Instruct has 43 scored results on Noometry and Qwen3-30B-A3B has 32.

Related comparisons

Go deeper