Model comparison

Gemini 3.1 Flash Lite vs GPT-5.5 Instant

GPT-5.5 Instant is the stronger model overall, scoring 42.7 to 40.8 on the Noometry Index.

Last verified . 25 shared benchmarks.

Gemini 3.1 Flash Lite Google

40.8

Rank #144 Confirmed

GPT-5.5 Instant OpenAI

42.7

Rank #110 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Gemini 3.1 Flash Lite scores higher in 1 category and GPT-5.5 Instant in 8 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemini 3.1 Flash Lite leads 40.7 to 26.5.
  • The biggest single-benchmark swing is Chess Puzzles: 25% for Gemini 3.1 Flash Lite and 12% for GPT-5.5 Instant.

Side by side

Gemini 3.1 Flash Lite and GPT-5.5 Instant specifications
Gemini 3.1 Flash LiteGPT-5.5 Instant
ProviderGoogleOpenAI
Noometry Index40.842.7
Released2026-03-032026-05-05
WeightsProprietaryProprietary
Context window1.05M—
Max output66K—
Input $ / M tokens$0.25—
Output $ / M tokens$1.50—
Results tracked3827

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.5 Instant leads

Gemini 3.1 Flash Lite: 37.8 (#188), GPT-5.5 Instant: 44.3 (#74)

Coding benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.5 Instant
SciCode41.9%48.6%
LMArena Coding14001433
LMArena WebDev1256—
WeirdML52.2%—
ALE-Bench797.73—

Agentic & Tool Use Not comparable

Gemini 3.1 Flash Lite: 30.2 (#79), GPT-5.5 Instant: —

Agentic & Tool Use benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.5 Instant
DeepResearch Bench37.3%—

Reasoning GPT-5.5 Instant leads

Gemini 3.1 Flash Lite: 22.9 (#186), GPT-5.5 Instant: 24.9 (#155)

Reasoning benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.5 Instant
CritPt1.1%0%
Chess Puzzles25%12%
LMArena Hard Prompts14071426
Epoch Capabilities Index144.47142.52
Kagi LLM Benchmark67.2%—
NYT Connections (extended)8.2%—
EnigmaEval3%—
Thematic Generalization63.3%—
DTBench76.8%—
LMCA35%—
ForecastBench54.4—

Math Gemini 3.1 Flash Lite leads

Gemini 3.1 Flash Lite: 40.7 (#90), GPT-5.5 Instant: 26.5 (#259)

Math benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.5 Instant
FrontierMath (Tiers 1-3)27.7%26.3%
OTIS Mock AIME 2024-202580%68.1%
LMArena Math14281420
FrontierMath Tier 4—2.4%

Knowledge GPT-5.5 Instant leads

Gemini 3.1 Flash Lite: 41.9 (#104), GPT-5.5 Instant: 48.9 (#74)

Knowledge benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.5 Instant
GPQA Diamond81.8%82.5%
LMArena Expert13981409
Humanity's Last Exam8.6%—
Vectara Hallucination Rate8.2%—

Multimodal Too close to call

Gemini 3.1 Flash Lite: 39.4 (#60), GPT-5.5 Instant: 40.0 (#52)

Multimodal benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.5 Instant
LMArena Vision12401250
LMArena Document—1403

Multilingual Too close to call

Gemini 3.1 Flash Lite: 52.3 (#86), GPT-5.5 Instant: 52.8 (#80)

Multilingual benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.5 Instant
LMArena Non-English14111417
LMArena Chinese14611456
LMArena French14241428
LMArena German14291411
LMArena Japanese14131408
LMArena Korean13921392
LMArena Russian14201431
LMArena Spanish14211429

Instruction Following GPT-5.5 Instant leads

Gemini 3.1 Flash Lite: 72.7 (#131), GPT-5.5 Instant: 74.2 (#100)

Instruction Following benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.5 Instant
LMArena Instruction Following13771406

Long Context Too close to call

Gemini 3.1 Flash Lite: 42.5 (#122), GPT-5.5 Instant: 43.4 (#96)

Long Context benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.5 Instant
LMArena Longer Query13941422

Writing & Preference Too close to call

Gemini 3.1 Flash Lite: 60.9 (#94), GPT-5.5 Instant: 61.8 (#85)

Writing & Preference benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.5 Instant
LMArena Text14161419
LMArena Creative Writing14011419
LMArena Multi-Turn14171433

Frequently asked questions

Is Gemini 3.1 Flash Lite better than GPT-5.5 Instant?

GPT-5.5 Instant is the stronger model overall, scoring 42.7 to 40.8 on the Noometry Index.

Is Gemini 3.1 Flash Lite or GPT-5.5 Instant better for coding?

GPT-5.5 Instant scores higher on coding benchmarks: 44.3 versus 37.8 in the Noometry coding category.

How many benchmarks do Gemini 3.1 Flash Lite and GPT-5.5 Instant share?

25 benchmarks have published results for both models. Gemini 3.1 Flash Lite has 38 scored results on Noometry and GPT-5.5 Instant has 27.

Related comparisons

Go deeper