Model comparison

Gemma 4 26B A4B IT vs Qwen3.5 27B

Gemma 4 26B A4B IT is the stronger model overall, scoring 43.5 to 41.9 on the Noometry Index.

Last verified . 21 shared benchmarks.

Gemma 4 26B A4B IT Google

43.5

Rank #92 Confirmed

Qwen3.5 27B Alibaba (Qwen)

41.9

Rank #127 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Gemma 4 26B A4B IT scores higher in 7 categories and Qwen3.5 27B in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemma 4 26B A4B IT leads 47.6 to 38.8.
  • The biggest single-benchmark swing is DTBench: 74.9% for Gemma 4 26B A4B IT and 82.4% for Qwen3.5 27B.
  • Gemma 4 26B A4B IT is cheaper at $0.0675 / $0.23 per million input/output tokens, against $0.30 / $2.40 for Qwen3.5 27B.

Side by side

Gemma 4 26B A4B IT and Qwen3.5 27B specifications
Gemma 4 26B A4B ITQwen3.5 27B
ProviderGoogleAlibaba (Qwen)
Noometry Index43.541.9
Released2026-04-022026-02-23
WeightsOpenOpen
Context window262K262K
Max output33K66K
Input $ / M tokens$0.0675$0.30
Output $ / M tokens$0.23$2.40
Results tracked2828

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemma 4 26B A4B IT: 39.0 (#164), Qwen3.5 27B: 38.9 (#168)

Coding benchmarks
BenchmarkGemma 4 26B A4B ITQwen3.5 27B
LMArena WebDev13591358
WeirdML35.2%39.5%
LMArena Coding14471427
ALE-Bench927.17349.45
SciCode40%—

Agentic & Tool Use Not comparable

Gemma 4 26B A4B IT: —, Qwen3.5 27B: —

Agentic & Tool Use benchmarks
BenchmarkGemma 4 26B A4B ITQwen3.5 27B
Vending-Bench 2—201.98

Reasoning Qwen3.5 27B leads

Gemma 4 26B A4B IT: 21.8 (#213), Qwen3.5 27B: 27.5 (#117)

Reasoning benchmarks
BenchmarkGemma 4 26B A4B ITQwen3.5 27B
LMArena Hard Prompts14391414
DTBench74.9%82.4%
LMCA29.7%34%
NYT Connections (extended)—47.9%
CritPt0%—
Chess Puzzles6%—
Thematic Generalization—45.5%
Epoch Capabilities Index141.85—

Math Gemma 4 26B A4B IT leads

Gemma 4 26B A4B IT: 47.6 (#67), Qwen3.5 27B: 38.8 (#127)

Math benchmarks
BenchmarkGemma 4 26B A4B ITQwen3.5 27B
LMArena Math14701429
MathArena Final-Answer Competitions—56.7%
OTIS Mock AIME 2024-202582.2%—

Knowledge Gemma 4 26B A4B IT leads

Gemma 4 26B A4B IT: 45.7 (#85), Qwen3.5 27B: 38.0 (#150)

Knowledge benchmarks
BenchmarkGemma 4 26B A4B ITQwen3.5 27B
Vectara Hallucination Rate5.2%12.1%
LMArena Expert14471428
GPQA Diamond73.2%—

Multimodal Gemma 4 26B A4B IT leads

Gemma 4 26B A4B IT: 40.6 (#46), Qwen3.5 27B: 39.4 (#59)

Multimodal benchmarks
BenchmarkGemma 4 26B A4B ITQwen3.5 27B
LMArena Vision12601241

Multilingual Gemma 4 26B A4B IT leads

Gemma 4 26B A4B IT: 53.1 (#71), Qwen3.5 27B: 50.8 (#115)

Multilingual benchmarks
BenchmarkGemma 4 26B A4B ITQwen3.5 27B
LMArena Non-English14211390
LMArena Chinese14951478
LMArena French14601410
LMArena Russian14341390
LMArena Spanish14171407
LMArena German—1393
LMArena Japanese—1345
LMArena Korean—1358

Instruction Following Gemma 4 26B A4B IT leads

Gemma 4 26B A4B IT: 74.8 (#82), Qwen3.5 27B: 73.5 (#119)

Instruction Following benchmarks
BenchmarkGemma 4 26B A4B ITQwen3.5 27B
LMArena Instruction Following14201393

Long Context Too close to call

Gemma 4 26B A4B IT: 43.6 (#91), Qwen3.5 27B: 43.1 (#106)

Long Context benchmarks
BenchmarkGemma 4 26B A4B ITQwen3.5 27B
LMArena Longer Query14281413

Writing & Preference Too close to call

Gemma 4 26B A4B IT: 58.6 (#115), Qwen3.5 27B: 59.3 (#111)

Writing & Preference benchmarks
BenchmarkGemma 4 26B A4B ITQwen3.5 27B
LMArena Text14341409
LMArena Creative Writing14021362
LMArena Multi-Turn14411410
EQ-Bench Creative Writing1305—

Frequently asked questions

Is Gemma 4 26B A4B IT better than Qwen3.5 27B?

Gemma 4 26B A4B IT is the stronger model overall, scoring 43.5 to 41.9 on the Noometry Index.

Which is cheaper, Gemma 4 26B A4B IT or Qwen3.5 27B?

Gemma 4 26B A4B IT is cheaper. It lists at $0.0675 per million input tokens and $0.23 per million output tokens; Qwen3.5 27B lists at $0.30 and $2.40.

Is Gemma 4 26B A4B IT or Qwen3.5 27B better for coding?

They score almost the same on coding (39.0 vs 38.9); test both on your own repository before choosing.

Which has the bigger context window?

Both accept 262K tokens.

How many benchmarks do Gemma 4 26B A4B IT and Qwen3.5 27B share?

21 benchmarks have published results for both models. Gemma 4 26B A4B IT has 28 scored results on Noometry and Qwen3.5 27B has 28.

Related comparisons

Go deeper