Model comparison
Claude 2.1 vs Qwen Turbo
Qwen Turbo is the stronger model overall, scoring 27.1 to 25.2 on the Noometry Index.
Last verified . 2 shared benchmarks.
Summary
- They share 2 benchmarks with published results for both. Claude 2.1 scores higher in 0 categories and Qwen Turbo in 2 categories; 2 gaps are clear of the uncertainty.
- The widest gap is in knowledge, where Qwen Turbo leads 22.2 to 15.4.
- The biggest single-benchmark swing is GPQA Diamond: 33% for Claude 2.1 and 41.8% for Qwen Turbo.
Side by side
| Claude 2.1 | Qwen Turbo | |
|---|---|---|
| Provider | Anthropic | Alibaba (Qwen) |
| Noometry Index | 25.2 | 27.1 |
| Released | 2023-11-21 | 2024-11-01 |
| Weights | Proprietary | Proprietary |
| Context window | — | 1M |
| Max output | — | 16K |
| Input $ / M tokens | — | $0.05 |
| Output $ / M tokens | — | $0.20 |
| Results tracked | 7 | 3 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Not comparable
Claude 2.1: 26.2 (#327), Qwen Turbo: —
| Benchmark | Claude 2.1 | Qwen Turbo |
|---|---|---|
| WeirdML | 7.1% | — |
Reasoning Not comparable
Claude 2.1: 21.4 (#221), Qwen Turbo: —
| Benchmark | Claude 2.1 | Qwen Turbo |
|---|---|---|
| DTBench | 51% | — |
| Epoch Capabilities Index | 119.27 | — |
| ForecastBench | 54.2 | — |
Math Qwen Turbo leads
Claude 2.1: 10.2 (#315), Qwen Turbo: 15.3 (#297)
| Benchmark | Claude 2.1 | Qwen Turbo |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 1.9% | 6.1% |
| MATH Level 5 | — | 56.2% |
Knowledge Qwen Turbo leads
Claude 2.1: 15.4 (#292), Qwen Turbo: 22.2 (#272)
| Benchmark | Claude 2.1 | Qwen Turbo |
|---|---|---|
| GPQA Diamond | 33% | 41.8% |
| MMLU | 73.5% | — |
Frequently asked questions
Is Claude 2.1 better than Qwen Turbo?
Qwen Turbo is the stronger model overall, scoring 27.1 to 25.2 on the Noometry Index.
How many benchmarks do Claude 2.1 and Qwen Turbo share?
2 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Qwen Turbo has 3.