Model comparison

Claude 2.1 vs Qwen3.8 27B

Qwen3.8 27B is the stronger model overall, scoring 46.0 to 25.2 on the Noometry Index.

Last verified . 2 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Qwen3.8 27B Alibaba (Qwen)

46.0

Rank #68 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Claude 2.1 scores higher in 0 categories and Qwen3.8 27B in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3.8 27B leads 37.1 to 10.2.
  • The biggest single-benchmark swing is DTBench: 51% for Claude 2.1 and 88% for Qwen3.8 27B.
  • Qwen3.8 27B has downloadable open weights; the other is API-only.

Side by side

Claude 2.1 and Qwen3.8 27B specifications
Claude 2.1Qwen3.8 27B
ProviderAnthropicAlibaba (Qwen)
Noometry Index25.246.0
Released2023-11-212026-08-14
WeightsProprietaryOpen
Context window—262K
Max output—33K
Input $ / M tokens—$0.99
Output $ / M tokens—$1.49
Results tracked731

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.8 27B leads

Claude 2.1: 26.2 (#327), Qwen3.8 27B: 50.5 (#44)

Coding benchmarks
BenchmarkClaude 2.1Qwen3.8 27B
LMArena WebDev—1593
SciCode—46.6%
WeirdML7.1%—
LMArena Coding—1482

Agentic & Tool Use Not comparable

Claude 2.1: —, Qwen3.8 27B: 32.9 (#57)

Agentic & Tool Use benchmarks
BenchmarkClaude 2.1Qwen3.8 27B
APEX-Agents—47.5%

Reasoning Qwen3.8 27B leads

Claude 2.1: 21.4 (#221), Qwen3.8 27B: 41.0 (#54)

Reasoning benchmarks
BenchmarkClaude 2.1Qwen3.8 27B
DTBench51%88%
Epoch Capabilities Index119.27149.38
ARC-AGI-2—42.4%
NYT Connections (extended)—54.5%
ARC-AGI-1—87.5%
CritPt—5.4%
LMArena Hard Prompts—1460
LMCA—41.4%
Surface Evolver Bench—45%
ForecastBench54.2—

Math Qwen3.8 27B leads

Claude 2.1: 10.2 (#315), Qwen3.8 27B: 37.1 (#161)

Math benchmarks
BenchmarkClaude 2.1Qwen3.8 27B
OTIS Mock AIME 2024-20251.9%—
ProofBench—16%
LMArena Math—1456

Knowledge Qwen3.8 27B leads

Claude 2.1: 15.4 (#292), Qwen3.8 27B: 41.6 (#109)

Knowledge benchmarks
BenchmarkClaude 2.1Qwen3.8 27B
GPQA Diamond33%—
LMArena Expert—1482
MMLU73.5%—

Multimodal Not comparable

Claude 2.1: —, Qwen3.8 27B: 41.3 (#37)

Multimodal benchmarks
BenchmarkClaude 2.1Qwen3.8 27B
LMArena Vision—1271

Multilingual Not comparable

Claude 2.1: —, Qwen3.8 27B: 53.7 (#60)

Multilingual benchmarks
BenchmarkClaude 2.1Qwen3.8 27B
LMArena Non-English—1430
LMArena Chinese—1504
LMArena French—1465
LMArena German—1438
LMArena Japanese—1384
LMArena Korean—1393
LMArena Russian—1415
LMArena Spanish—1448

Instruction Following Not comparable

Claude 2.1: —, Qwen3.8 27B: 75.8 (#53)

Instruction Following benchmarks
BenchmarkClaude 2.1Qwen3.8 27B
LMArena Instruction Following—1439

Long Context Not comparable

Claude 2.1: —, Qwen3.8 27B: 44.3 (#70)

Long Context benchmarks
BenchmarkClaude 2.1Qwen3.8 27B
LMArena Longer Query—1450

Writing & Preference Not comparable

Claude 2.1: —, Qwen3.8 27B: 65.8 (#43)

Writing & Preference benchmarks
BenchmarkClaude 2.1Qwen3.8 27B
LMArena Text—1441
LMArena Creative Writing—1384
EQ-Bench Creative Writing—1671
LMArena Multi-Turn—1441

Frequently asked questions

Is Claude 2.1 better than Qwen3.8 27B?

Qwen3.8 27B is the stronger model overall, scoring 46.0 to 25.2 on the Noometry Index.

Is Claude 2.1 or Qwen3.8 27B better for coding?

Qwen3.8 27B scores higher on coding benchmarks: 50.5 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and Qwen3.8 27B share?

2 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Qwen3.8 27B has 31.

Related comparisons

Go deeper