Model comparison

DeepSeek-V3.1 vs Sonar

DeepSeek-V3.1 is the stronger model overall, scoring 42.8 to 38.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

DeepSeek-V3.1 DeepSeek

42.8

Rank #108 Confirmed

Sonar Perplexity

38.5

Rank #187 Confirmed

Summary

  • The widest gap is in writing & preference, where DeepSeek-V3.1 leads 60.3 to 52.6.
  • DeepSeek-V3.1 is cheaper at $0.25 / $0.95 per million input/output tokens, against $1 / $1 for Sonar.
  • DeepSeek-V3.1 accepts more context: 164K tokens versus 128K.
  • DeepSeek-V3.1 has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3.1 and Sonar specifications
DeepSeek-V3.1Sonar
ProviderDeepSeekPerplexity
Noometry Index42.838.5
Released2025-08-212024-01-01
WeightsOpenProprietary
Context window164K128K
Max output8K4K
Input $ / M tokens$0.25$1
Output $ / M tokens$0.95$1
Results tracked277

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.1 leads

DeepSeek-V3.1: 40.3 (#144), Sonar: 35.7 (#221)

Coding benchmarks
BenchmarkDeepSeek-V3.1Sonar
WeirdML38.4%—
LiveBench Coding—35.1%
LMArena Coding1417—

Reasoning DeepSeek-V3.1 leads

DeepSeek-V3.1: 27.9 (#110), Sonar: 21.1 (#227)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1Sonar
SimpleBench40%—
Kagi LLM Benchmark53.2%—
LiveBench Reasoning—46.3%
LMArena Hard Prompts1417—
DTBench82.7%—
LiveBench Data Analysis—37.9%
LMCA24.3%—
Epoch Capabilities Index139.92—
ForecastBench58—
LiveBench—46.9%

Math DeepSeek-V3.1 leads

DeepSeek-V3.1: 38.9 (#122), Sonar: 33.7 (#200)

Math benchmarks
BenchmarkDeepSeek-V3.1Sonar
LiveBench Math—41.6%
LMArena Math1420—

Knowledge Not comparable

DeepSeek-V3.1: 43.7 (#90), Sonar: —

Knowledge benchmarks
BenchmarkDeepSeek-V3.1Sonar
Vectara Hallucination Rate5.5%—
LMArena Expert1405—

Multilingual Not comparable

DeepSeek-V3.1: 51.6 (#106), Sonar: —

Multilingual benchmarks
BenchmarkDeepSeek-V3.1Sonar
LMArena Non-English1400—
LMArena Chinese1469—
LMArena French1447—
LMArena German1411—
LMArena Japanese1378—
LMArena Korean1337—
LMArena Russian1405—
LMArena Spanish1431—

Instruction Following DeepSeek-V3.1 leads

DeepSeek-V3.1: 73.9 (#110), Sonar: 71.4 (#150)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1Sonar
LiveBench Instruction Following—76.2%
LMArena Instruction Following1400—

Long Context Not comparable

DeepSeek-V3.1: 36.3 (#232), Sonar: —

Long Context benchmarks
BenchmarkDeepSeek-V3.1Sonar
Fiction.LiveBench52.8%—
LMArena Longer Query1422—

Writing & Preference DeepSeek-V3.1 leads

DeepSeek-V3.1: 60.3 (#98), Sonar: 52.6 (#167)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1Sonar
LMArena Text1420—
LMArena Creative Writing1401—
EQ-Bench Creative Writing1436—
LMArena Multi-Turn1408—
LiveBench Language—44.1%

Frequently asked questions

Is DeepSeek-V3.1 better than Sonar?

DeepSeek-V3.1 is the stronger model overall, scoring 42.8 to 38.5 on the Noometry Index.

Which is cheaper, DeepSeek-V3.1 or Sonar?

DeepSeek-V3.1 is cheaper. It lists at $0.25 per million input tokens and $0.95 per million output tokens; Sonar lists at $1 and $1.

Is DeepSeek-V3.1 or Sonar better for coding?

DeepSeek-V3.1 scores higher on coding benchmarks: 40.3 versus 35.7 in the Noometry coding category.

Which has the bigger context window?

DeepSeek-V3.1 does, with 164K tokens against 128K.

How many benchmarks do DeepSeek-V3.1 and Sonar share?

0 benchmarks have published results for both models. DeepSeek-V3.1 has 27 scored results on Noometry and Sonar has 7.

Related comparisons

Go deeper