Model comparison

Claude Opus 5.5 vs GPT-4.1 mini

Claude Opus 5.5 is the stronger model overall, scoring 68.6 to 33.6 on the Noometry Index. GPT-4.1 mini costs 11× less per token, which makes it the better buy when Claude Opus 5.5's lead doesn't matter for your workload.

Last verified . 28 shared benchmarks.

Claude Opus 5.5 Anthropic

68.6

Rank #3 Confirmed

GPT-4.1 mini OpenAI

33.6

Rank #240 Confirmed

Summary

  • They share 28 benchmarks with published results for both. Claude Opus 5.5 scores higher in 10 categories and GPT-4.1 mini in 0 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Claude Opus 5.5 leads 80.2 to 10.8.
  • The biggest single-benchmark swing is ARC-AGI-1: 98.5% for Claude Opus 5.5 and 3.5% for GPT-4.1 mini.
  • GPT-4.1 mini is cheaper at $0.40 / $1.60 per million input/output tokens, against $4 / $20 for Claude Opus 5.5.
  • GPT-4.1 mini accepts more context: 1.05M tokens versus 1M.

Side by side

Claude Opus 5.5 and GPT-4.1 mini specifications
Claude Opus 5.5GPT-4.1 mini
ProviderAnthropicOpenAI
Noometry Index68.633.6
Released2026-09-222025-04-14
WeightsProprietaryProprietary
Context window1M1.05M
Max output128K33K
Input $ / M tokens$4$0.40
Output $ / M tokens$20$1.60
Results tracked4447

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 5.5 leads

Claude Opus 5.5: 71.9 (#3), GPT-4.1 mini: 30.6 (#293)

Coding benchmarks
BenchmarkClaude Opus 5.5GPT-4.1 mini
SciCode66.9%40.4%
LMArena Coding15471367
FrontierCode54.6%—
SWE-bench Verified (bash only)—23.9%
Aider Polyglot—32.4%
CursorBench57.8%—
LMArena WebDev1813—
FrontierSWE62.3%—
WeirdML—37.6%
BigCodeBench Instruct—48.9%
MirrorCode77.4%—
CadEval—16%
ALE-Bench2,147—

Agentic & Tool Use Claude Opus 5.5 leads

Claude Opus 5.5: 45.3 (#15), GPT-4.1 mini: 33.3 (#55)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 5.5GPT-4.1 mini
APEX-Agents73.5%—
Berkeley Function Calling Leaderboard—50.5%
GDP.pdf30.6%—
Vending-Bench 29,235—

Reasoning Claude Opus 5.5 leads

Claude Opus 5.5: 80.2 (#3), GPT-4.1 mini: 10.8 (#340)

Reasoning benchmarks
BenchmarkClaude Opus 5.5GPT-4.1 mini
ARC-AGI-293.3%0%
ARC-AGI-198.5%3.5%
CritPt31.7%0%
LMArena Hard Prompts15351349
Mystery Game Puzzles71%7%
DTBench98.9%68.8%
LMCA68.2%21.1%
Epoch Capabilities Index167.33135.01
Kagi LLM Benchmark—48.6%
NYT Connections (extended)88.5%—
Chess Puzzles—7%
EBR-Bench71.4%—

Math Claude Opus 5.5 leads

Claude Opus 5.5: 91.8 (#3), GPT-4.1 mini: 24.1 (#270)

Math benchmarks
BenchmarkClaude Opus 5.5GPT-4.1 mini
FrontierMath (Tiers 1-3)91.2%6.7%
OTIS Mock AIME 2024-2025100%44.7%
LMArena Math15061343
FrontierMath Tier 495%—
ProofBench100%—
Omni-MATH—49.1%
MATH Level 5—87.3%
FrontierMath (Feb 2025 set)—4.5%
FrontierMath Erdős2.9%—

Knowledge Claude Opus 5.5 leads

Claude Opus 5.5: 66.4 (#10), GPT-4.1 mini: 34.7 (#194)

Knowledge benchmarks
BenchmarkClaude Opus 5.5GPT-4.1 mini
GPQA Diamond90.6%65.8%
SimpleQA Verified72.2%12.7%
LMArena Expert15471338
MMLU-Pro—78.3%
GPQA (HELM)—61.4%

Multimodal Claude Opus 5.5 leads

Claude Opus 5.5: 57.8 (#1), GPT-4.1 mini: 35.8 (#82)

Multimodal benchmarks
BenchmarkClaude Opus 5.5GPT-4.1 mini
LMArena Vision13211181
Blueprint-Bench 251.2%—
Furniture Assembly83.3%—

Multilingual Claude Opus 5.5 leads

Claude Opus 5.5: 59.1 (#2), GPT-4.1 mini: 45.7 (#166)

Multilingual benchmarks
BenchmarkClaude Opus 5.5GPT-4.1 mini
LMArena Non-English15071318
LMArena Chinese15881329
LMArena French15141358
LMArena Russian15201324
LMArena Spanish15071319
LMArena German—1351
LMArena Japanese—1290
LMArena Korean—1298

Instruction Following Claude Opus 5.5 leads

Claude Opus 5.5: 80.0 (#3), GPT-4.1 mini: 73.7 (#118)

Instruction Following benchmarks
BenchmarkClaude Opus 5.5GPT-4.1 mini
LMArena Instruction Following15371333
IFEval—90.4%

Long Context Claude Opus 5.5 leads

Claude Opus 5.5: 47.1 (#19), GPT-4.1 mini: 31.8 (#275)

Long Context benchmarks
BenchmarkClaude Opus 5.5GPT-4.1 mini
LMArena Longer Query15321344
Fiction.LiveBench—44.4%

Writing & Preference Claude Opus 5.5 leads

Claude Opus 5.5: 78.2 (#3), GPT-4.1 mini: 48.6 (#199)

Writing & Preference benchmarks
BenchmarkClaude Opus 5.5GPT-4.1 mini
LMArena Text15151340
LMArena Creative Writing15331300
EQ-Bench Creative Writing20501147
LMArena Multi-Turn14991354
WildBench—83.8%

Frequently asked questions

Is Claude Opus 5.5 better than GPT-4.1 mini?

Claude Opus 5.5 is the stronger model overall, scoring 68.6 to 33.6 on the Noometry Index. GPT-4.1 mini costs 11× less per token, which makes it the better buy when Claude Opus 5.5's lead doesn't matter for your workload.

Which is cheaper, Claude Opus 5.5 or GPT-4.1 mini?

GPT-4.1 mini is cheaper. It lists at $0.40 per million input tokens and $1.60 per million output tokens; Claude Opus 5.5 lists at $4 and $20.

Is Claude Opus 5.5 or GPT-4.1 mini better for coding?

Claude Opus 5.5 scores higher on coding benchmarks: 71.9 versus 30.6 in the Noometry coding category.

Which has the bigger context window?

GPT-4.1 mini does, with 1.05M tokens against 1M.

How many benchmarks do Claude Opus 5.5 and GPT-4.1 mini share?

28 benchmarks have published results for both models. Claude Opus 5.5 has 44 scored results on Noometry and GPT-4.1 mini has 47.

Related comparisons

Go deeper