Model comparison

GPT-4.1 mini vs Qwen3 Coder Next

GPT-4.1 mini and Qwen3 Coder Next score almost the same on the Noometry Index (33.6 vs 34.3), so choose on price, context window or the category you care about most.

Last verified . 3 shared benchmarks.

GPT-4.1 mini OpenAI

33.6

Rank #240 Confirmed

Qwen3 Coder Next Alibaba (Qwen)

34.3

Rank #232 Reported

Summary

  • They share 3 benchmarks with published results for both. GPT-4.1 mini scores higher in 0 categories and Qwen3 Coder Next in 2 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3 Coder Next leads 22.4 to 10.8.
  • The biggest single-benchmark swing is SciCode: 40.4% for GPT-4.1 mini and 32.3% for Qwen3 Coder Next.
  • Qwen3 Coder Next is cheaper at $0.12 / $0.80 per million input/output tokens, against $0.40 / $1.60 for GPT-4.1 mini.
  • GPT-4.1 mini accepts more context: 1.05M tokens versus 262K.
  • Qwen3 Coder Next has downloadable open weights; the other is API-only.

Side by side

GPT-4.1 mini and Qwen3 Coder Next specifications
GPT-4.1 miniQwen3 Coder Next
ProviderOpenAIAlibaba (Qwen)
Noometry Index33.634.3
Released2025-04-142026-02-02
WeightsProprietaryOpen
Context window1.05M262K
Max output33K66K
Input $ / M tokens$0.40$0.12
Output $ / M tokens$1.60$0.80
Results tracked473

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 Coder Next leads

GPT-4.1 mini: 30.6 (#293), Qwen3 Coder Next: 36.3 (#210)

Coding benchmarks
BenchmarkGPT-4.1 miniQwen3 Coder Next
SciCode40.4%32.3%
WeirdML37.6%34.4%
SWE-bench Verified (bash only)23.9%—
Aider Polyglot32.4%—
BigCodeBench Instruct48.9%—
LMArena Coding1367—
CadEval16%—

Agentic & Tool Use Not comparable

GPT-4.1 mini: 33.3 (#55), Qwen3 Coder Next: —

Agentic & Tool Use benchmarks
BenchmarkGPT-4.1 miniQwen3 Coder Next
Berkeley Function Calling Leaderboard50.5%—

Reasoning Qwen3 Coder Next leads

GPT-4.1 mini: 10.8 (#340), Qwen3 Coder Next: 22.4 (#196)

Reasoning benchmarks
BenchmarkGPT-4.1 miniQwen3 Coder Next
CritPt0%0%
ARC-AGI-20%—
Kagi LLM Benchmark48.6%—
ARC-AGI-13.5%—
Chess Puzzles7%—
LMArena Hard Prompts1349—
Mystery Game Puzzles7%—
DTBench68.8%—
LMCA21.1%—
Epoch Capabilities Index135.01—

Math Not comparable

GPT-4.1 mini: 24.1 (#270), Qwen3 Coder Next: —

Math benchmarks
BenchmarkGPT-4.1 miniQwen3 Coder Next
FrontierMath (Tiers 1-3)6.7%—
OTIS Mock AIME 2024-202544.7%—
Omni-MATH49.1%—
LMArena Math1343—
MATH Level 587.3%—
FrontierMath (Feb 2025 set)4.5%—

Knowledge Not comparable

GPT-4.1 mini: 34.7 (#194), Qwen3 Coder Next: —

Knowledge benchmarks
BenchmarkGPT-4.1 miniQwen3 Coder Next
GPQA Diamond65.8%—
SimpleQA Verified12.7%—
MMLU-Pro78.3%—
GPQA (HELM)61.4%—
LMArena Expert1338—

Multimodal Not comparable

GPT-4.1 mini: 35.8 (#82), Qwen3 Coder Next: —

Multimodal benchmarks
BenchmarkGPT-4.1 miniQwen3 Coder Next
LMArena Vision1181—

Multilingual Not comparable

GPT-4.1 mini: 45.7 (#166), Qwen3 Coder Next: —

Multilingual benchmarks
BenchmarkGPT-4.1 miniQwen3 Coder Next
LMArena Non-English1318—
LMArena Chinese1329—
LMArena French1358—
LMArena German1351—
LMArena Japanese1290—
LMArena Korean1298—
LMArena Russian1324—
LMArena Spanish1319—

Instruction Following Not comparable

GPT-4.1 mini: 73.7 (#118), Qwen3 Coder Next: —

Instruction Following benchmarks
BenchmarkGPT-4.1 miniQwen3 Coder Next
IFEval90.4%—
LMArena Instruction Following1333—

Long Context Not comparable

GPT-4.1 mini: 31.8 (#275), Qwen3 Coder Next: —

Long Context benchmarks
BenchmarkGPT-4.1 miniQwen3 Coder Next
Fiction.LiveBench44.4%—
LMArena Longer Query1344—

Writing & Preference Not comparable

GPT-4.1 mini: 48.6 (#199), Qwen3 Coder Next: —

Writing & Preference benchmarks
BenchmarkGPT-4.1 miniQwen3 Coder Next
LMArena Text1340—
LMArena Creative Writing1300—
EQ-Bench Creative Writing1147—
WildBench83.8%—
LMArena Multi-Turn1354—

Frequently asked questions

Is GPT-4.1 mini better than Qwen3 Coder Next?

GPT-4.1 mini and Qwen3 Coder Next score almost the same on the Noometry Index (33.6 vs 34.3), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-4.1 mini or Qwen3 Coder Next?

Qwen3 Coder Next is cheaper. It lists at $0.12 per million input tokens and $0.80 per million output tokens; GPT-4.1 mini lists at $0.40 and $1.60.

Is GPT-4.1 mini or Qwen3 Coder Next better for coding?

Qwen3 Coder Next scores higher on coding benchmarks: 36.3 versus 30.6 in the Noometry coding category.

Which has the bigger context window?

GPT-4.1 mini does, with 1.05M tokens against 262K.

How many benchmarks do GPT-4.1 mini and Qwen3 Coder Next share?

3 benchmarks have published results for both models. GPT-4.1 mini has 47 scored results on Noometry and Qwen3 Coder Next has 3.

Related comparisons

Go deeper