Model comparison

Qwen3.8 27B vs Sonar

Qwen3.8 27B is the stronger model overall, scoring 46.0 to 38.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

Qwen3.8 27B Alibaba (Qwen)

46.0

Rank #68 Confirmed

Sonar Perplexity

38.5

Rank #187 Confirmed

Summary

  • The widest gap is in reasoning, where Qwen3.8 27B leads 41.0 to 21.1.
  • Qwen3.8 27B is cheaper at $0.04 / $2.30 per million input/output tokens, against $1 / $1 for Sonar.
  • Qwen3.8 27B accepts more context: 262K tokens versus 128K.
  • Qwen3.8 27B has downloadable open weights; the other is API-only.

Side by side

Qwen3.8 27B and Sonar specifications
Qwen3.8 27BSonar
ProviderAlibaba (Qwen)Perplexity
Noometry Index46.038.5
Released2026-08-142024-01-01
WeightsOpenProprietary
Context window262K128K
Max output33K4K
Input $ / M tokens$0.04$1
Output $ / M tokens$2.30$1
Results tracked317

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.8 27B leads

Qwen3.8 27B: 50.5 (#44), Sonar: 35.7 (#221)

Coding benchmarks
BenchmarkQwen3.8 27BSonar
LMArena WebDev1593—
SciCode46.6%—
LiveBench Coding—35.1%
LMArena Coding1482—

Agentic & Tool Use Not comparable

Qwen3.8 27B: 32.9 (#57), Sonar: —

Agentic & Tool Use benchmarks
BenchmarkQwen3.8 27BSonar
APEX-Agents47.5%—

Reasoning Qwen3.8 27B leads

Qwen3.8 27B: 41.0 (#54), Sonar: 21.1 (#227)

Reasoning benchmarks
BenchmarkQwen3.8 27BSonar
ARC-AGI-242.4%—
NYT Connections (extended)54.5%—
ARC-AGI-187.5%—
CritPt5.4%—
LiveBench Reasoning—46.3%
LMArena Hard Prompts1460—
DTBench88%—
LiveBench Data Analysis—37.9%
LMCA41.4%—
Surface Evolver Bench45%—
Epoch Capabilities Index149.38—
LiveBench—46.9%

Math Qwen3.8 27B leads

Qwen3.8 27B: 37.1 (#161), Sonar: 33.7 (#200)

Math benchmarks
BenchmarkQwen3.8 27BSonar
ProofBench16%—
LiveBench Math—41.6%
LMArena Math1456—

Knowledge Not comparable

Qwen3.8 27B: 41.6 (#109), Sonar: —

Knowledge benchmarks
BenchmarkQwen3.8 27BSonar
LMArena Expert1482—

Multimodal Not comparable

Qwen3.8 27B: 41.3 (#37), Sonar: —

Multimodal benchmarks
BenchmarkQwen3.8 27BSonar
LMArena Vision1271—

Multilingual Not comparable

Qwen3.8 27B: 53.7 (#60), Sonar: —

Multilingual benchmarks
BenchmarkQwen3.8 27BSonar
LMArena Non-English1430—
LMArena Chinese1504—
LMArena French1465—
LMArena German1438—
LMArena Japanese1384—
LMArena Korean1393—
LMArena Russian1415—
LMArena Spanish1448—

Instruction Following Qwen3.8 27B leads

Qwen3.8 27B: 75.8 (#53), Sonar: 71.4 (#150)

Instruction Following benchmarks
BenchmarkQwen3.8 27BSonar
LiveBench Instruction Following—76.2%
LMArena Instruction Following1439—

Long Context Not comparable

Qwen3.8 27B: 44.3 (#70), Sonar: —

Long Context benchmarks
BenchmarkQwen3.8 27BSonar
LMArena Longer Query1450—

Writing & Preference Qwen3.8 27B leads

Qwen3.8 27B: 65.8 (#43), Sonar: 52.6 (#167)

Writing & Preference benchmarks
BenchmarkQwen3.8 27BSonar
LMArena Text1441—
LMArena Creative Writing1384—
EQ-Bench Creative Writing1671—
LMArena Multi-Turn1441—
LiveBench Language—44.1%

Frequently asked questions

Is Qwen3.8 27B better than Sonar?

Qwen3.8 27B is the stronger model overall, scoring 46.0 to 38.5 on the Noometry Index.

Which is cheaper, Qwen3.8 27B or Sonar?

Qwen3.8 27B is cheaper. It lists at $0.04 per million input tokens and $2.30 per million output tokens; Sonar lists at $1 and $1.

Is Qwen3.8 27B or Sonar better for coding?

Qwen3.8 27B scores higher on coding benchmarks: 50.5 versus 35.7 in the Noometry coding category.

Which has the bigger context window?

Qwen3.8 27B does, with 262K tokens against 128K.

How many benchmarks do Qwen3.8 27B and Sonar share?

0 benchmarks have published results for both models. Qwen3.8 27B has 31 scored results on Noometry and Sonar has 7.

Related comparisons

Go deeper