Model comparison

Falcon-180B vs Qwen3-VL 235B-A22B

Qwen3-VL 235B-A22B is the stronger model overall, scoring 43.2 to 32.2 on the Noometry Index.

Last verified . 6 shared benchmarks.

Summary

  • They share 6 benchmarks with published results for both. Falcon-180B scores higher in 0 categories and Qwen3-VL 235B-A22B in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen3-VL 235B-A22B leads 60.2 to 29.1.

Side by side

Falcon-180B and Qwen3-VL 235B-A22B specifications
Falcon-180BQwen3-VL 235B-A22B
ProviderTechnology Innovation InstituteAlibaba (Qwen)
Noometry Index32.243.2
Released2023-09-062025-04
WeightsOpenOpen
Context window—131K
Max output—33K
Input $ / M tokens—$0.70
Output $ / M tokens—$2.80
Results tracked1618

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Falcon-180B: —, Qwen3-VL 235B-A22B: 42.4 (#100)

Coding benchmarks
BenchmarkFalcon-180BQwen3-VL 235B-A22B
LMArena Coding—1439

Reasoning Qwen3-VL 235B-A22B leads

Falcon-180B: 19.1 (#269), Qwen3-VL 235B-A22B: 29.3 (#92)

Reasoning benchmarks
BenchmarkFalcon-180BQwen3-VL 235B-A22B
LMArena Hard Prompts10071428
Epoch Capabilities Index112.13—
HellaSwag89%—
LAMBADA79.8%—
PIQA84.9%—
WinoGrande87.1%—

Math Not comparable

Falcon-180B: —, Qwen3-VL 235B-A22B: 39.0 (#118)

Math benchmarks
BenchmarkFalcon-180BQwen3-VL 235B-A22B
LMArena Math—1426
GSM8K54.4%—

Knowledge Not comparable

Falcon-180B: —, Qwen3-VL 235B-A22B: 40.3 (#121)

Knowledge benchmarks
BenchmarkFalcon-180BQwen3-VL 235B-A22B
LMArena Expert—1442
ARC (AI2) Challenge67.8%—
BoolQ89%—
MMLU70.6%—
OpenBookQA64.2%—

Multimodal Not comparable

Falcon-180B: —, Qwen3-VL 235B-A22B: 39.8 (#55)

Multimodal benchmarks
BenchmarkFalcon-180BQwen3-VL 235B-A22B
LMArena Vision—1247

Multilingual Qwen3-VL 235B-A22B leads

Falcon-180B: 25.2 (#286), Qwen3-VL 235B-A22B: 51.9 (#97)

Multilingual benchmarks
BenchmarkFalcon-180BQwen3-VL 235B-A22B
LMArena Non-English10001405
LMArena Chinese—1463
LMArena French—1452
LMArena German—1424
LMArena Japanese—1385
LMArena Korean—1394
LMArena Russian—1408
LMArena Spanish—1428

Instruction Following Qwen3-VL 235B-A22B leads

Falcon-180B: 53.4 (#286), Qwen3-VL 235B-A22B: 74.2 (#101)

Instruction Following benchmarks
BenchmarkFalcon-180BQwen3-VL 235B-A22B
LMArena Instruction Following10471406

Long Context Not comparable

Falcon-180B: —, Qwen3-VL 235B-A22B: 43.4 (#98)

Long Context benchmarks
BenchmarkFalcon-180BQwen3-VL 235B-A22B
LMArena Longer Query—1420

Writing & Preference Qwen3-VL 235B-A22B leads

Falcon-180B: 29.1 (#295), Qwen3-VL 235B-A22B: 60.2 (#99)

Writing & Preference benchmarks
BenchmarkFalcon-180BQwen3-VL 235B-A22B
LMArena Text10541420
LMArena Creative Writing10891366
LMArena Multi-Turn10131428

Frequently asked questions

Is Falcon-180B better than Qwen3-VL 235B-A22B?

Qwen3-VL 235B-A22B is the stronger model overall, scoring 43.2 to 32.2 on the Noometry Index.

How many benchmarks do Falcon-180B and Qwen3-VL 235B-A22B share?

6 benchmarks have published results for both models. Falcon-180B has 16 scored results on Noometry and Qwen3-VL 235B-A22B has 18.

Related comparisons

Go deeper