Model comparison

DeepSeek-R1-Distill-Qwen-1.5B vs GPT-5.3 Chat

GPT-5.3 Chat is the stronger model overall, scoring 42.8 to 26.1 on the Noometry Index.

Last verified . 0 shared benchmarks.

GPT-5.3 Chat OpenAI

42.8

Rank #109 Confirmed

Summary

  • The widest gap is in knowledge, where GPT-5.3 Chat leads 38.8 to 16.0.
  • DeepSeek-R1-Distill-Qwen-1.5B has downloadable open weights; the other is API-only.

Side by side

DeepSeek-R1-Distill-Qwen-1.5B and GPT-5.3 Chat specifications
DeepSeek-R1-Distill-Qwen-1.5BGPT-5.3 Chat
ProviderDeepSeekOpenAI
Noometry Index26.142.8
Released2025-01-202026-03-03
WeightsOpenProprietary
Context window—128K
Max output—16K
Input $ / M tokens—$1.75
Output $ / M tokens—$14
Results tracked518

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.3 Chat leads

DeepSeek-R1-Distill-Qwen-1.5B: 21.8 (#336), GPT-5.3 Chat: 41.4 (#124)

Coding benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BGPT-5.3 Chat
BigCodeBench Instruct7%—
LMArena Coding—1408
BigCodeBench Complete7.9%—

Reasoning GPT-5.3 Chat leads

DeepSeek-R1-Distill-Qwen-1.5B: 19.2 (#262), GPT-5.3 Chat: 28.5 (#102)

Reasoning benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BGPT-5.3 Chat
Chess Puzzles0%—
LMArena Hard Prompts—1399

Math GPT-5.3 Chat leads

DeepSeek-R1-Distill-Qwen-1.5B: 23.0 (#274), GPT-5.3 Chat: 38.2 (#142)

Math benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BGPT-5.3 Chat
OTIS Mock AIME 2024-202521.4%—
LMArena Math—1389

Knowledge GPT-5.3 Chat leads

DeepSeek-R1-Distill-Qwen-1.5B: 16.0 (#290), GPT-5.3 Chat: 38.8 (#140)

Knowledge benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BGPT-5.3 Chat
GPQA Diamond33.6%—
LMArena Expert—1397

Multilingual Not comparable

DeepSeek-R1-Distill-Qwen-1.5B: —, GPT-5.3 Chat: 50.3 (#124)

Multilingual benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BGPT-5.3 Chat
LMArena Non-English—1382
LMArena Chinese—1432
LMArena French—1397
LMArena German—1384
LMArena Japanese—1352
LMArena Korean—1346
LMArena Russian—1400
LMArena Spanish—1371

Instruction Following Not comparable

DeepSeek-R1-Distill-Qwen-1.5B: —, GPT-5.3 Chat: 72.8 (#129)

Instruction Following benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BGPT-5.3 Chat
LMArena Instruction Following—1378

Long Context Not comparable

DeepSeek-R1-Distill-Qwen-1.5B: —, GPT-5.3 Chat: 42.6 (#120)

Long Context benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BGPT-5.3 Chat
LMArena Longer Query—1396

Writing & Preference Not comparable

DeepSeek-R1-Distill-Qwen-1.5B: —, GPT-5.3 Chat: 63.1 (#68)

Writing & Preference benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BGPT-5.3 Chat
LMArena Text—1389
LMArena Creative Writing—1355
EQ-Bench Creative Writing—1690
LMArena Multi-Turn—1412

Frequently asked questions

Is DeepSeek-R1-Distill-Qwen-1.5B better than GPT-5.3 Chat?

GPT-5.3 Chat is the stronger model overall, scoring 42.8 to 26.1 on the Noometry Index.

Is DeepSeek-R1-Distill-Qwen-1.5B or GPT-5.3 Chat better for coding?

GPT-5.3 Chat scores higher on coding benchmarks: 41.4 versus 21.8 in the Noometry coding category.

How many benchmarks do DeepSeek-R1-Distill-Qwen-1.5B and GPT-5.3 Chat share?

0 benchmarks have published results for both models. DeepSeek-R1-Distill-Qwen-1.5B has 5 scored results on Noometry and GPT-5.3 Chat has 18.

Related comparisons

Go deeper