Model comparison

Claude 2.1 vs Qwen Turbo

Qwen Turbo is the stronger model overall, scoring 27.1 to 25.2 on the Noometry Index.

Last verified . 2 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Qwen Turbo Alibaba (Qwen)

27.1

Rank #335 Reported

Summary

  • They share 2 benchmarks with published results for both. Claude 2.1 scores higher in 0 categories and Qwen Turbo in 2 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen Turbo leads 22.2 to 15.4.
  • The biggest single-benchmark swing is GPQA Diamond: 33% for Claude 2.1 and 41.8% for Qwen Turbo.

Side by side

Claude 2.1 and Qwen Turbo specifications
Claude 2.1Qwen Turbo
ProviderAnthropicAlibaba (Qwen)
Noometry Index25.227.1
Released2023-11-212024-11-01
WeightsProprietaryProprietary
Context window—1M
Max output—16K
Input $ / M tokens—$0.05
Output $ / M tokens—$0.20
Results tracked73

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Claude 2.1: 26.2 (#327), Qwen Turbo: —

Coding benchmarks
BenchmarkClaude 2.1Qwen Turbo
WeirdML7.1%—

Reasoning Not comparable

Claude 2.1: 21.4 (#221), Qwen Turbo: —

Reasoning benchmarks
BenchmarkClaude 2.1Qwen Turbo
DTBench51%—
Epoch Capabilities Index119.27—
ForecastBench54.2—

Math Qwen Turbo leads

Claude 2.1: 10.2 (#315), Qwen Turbo: 15.3 (#297)

Math benchmarks
BenchmarkClaude 2.1Qwen Turbo
OTIS Mock AIME 2024-20251.9%6.1%
MATH Level 5—56.2%

Knowledge Qwen Turbo leads

Claude 2.1: 15.4 (#292), Qwen Turbo: 22.2 (#272)

Knowledge benchmarks
BenchmarkClaude 2.1Qwen Turbo
GPQA Diamond33%41.8%
MMLU73.5%—

Frequently asked questions

Is Claude 2.1 better than Qwen Turbo?

Qwen Turbo is the stronger model overall, scoring 27.1 to 25.2 on the Noometry Index.

How many benchmarks do Claude 2.1 and Qwen Turbo share?

2 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Qwen Turbo has 3.

Related comparisons

Go deeper