Model comparison

Gemini 3.5 Flash vs GPT-5.6 Sol

GPT-5.6 Sol is the stronger model overall, scoring 65.0 to 54.2 on the Noometry Index. Gemini 3.5 Flash costs 2.4× less per token, which makes it the better buy when GPT-5.6 Sol's lead doesn't matter for your workload.

Last verified . 49 shared benchmarks.

Gemini 3.5 Flash Google

54.2

Rank #32 Confirmed

GPT-5.6 Sol OpenAI

65.0

Rank #7 Confirmed

Summary

  • They share 49 benchmarks with published results for both. Gemini 3.5 Flash scores higher in 3 categories and GPT-5.6 Sol in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where GPT-5.6 Sol leads 50.3 to 24.7.
  • The biggest single-benchmark swing is FrontierMath Tier 4: 26.8% for Gemini 3.5 Flash and 82.9% for GPT-5.6 Sol.
  • Gemini 3.5 Flash is cheaper at $1.50 / $9 per million input/output tokens, against $4 / $20 for GPT-5.6 Sol.
  • GPT-5.6 Sol accepts more context: 1.05M tokens versus 1.05M.

Side by side

Gemini 3.5 Flash and GPT-5.6 Sol specifications
Gemini 3.5 FlashGPT-5.6 Sol
ProviderGoogleOpenAI
Noometry Index54.265.0
Released2026-05-192026-07-09
WeightsProprietaryProprietary
Context window1.05M1.05M
Max output66K128K
Input $ / M tokens$1.50$4
Output $ / M tokens$9$20
Results tracked5465

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.6 Sol leads

Gemini 3.5 Flash: 49.4 (#49), GPT-5.6 Sol: 65.1 (#7)

Coding benchmarks
BenchmarkGemini 3.5 FlashGPT-5.6 Sol
DeepSWE37.4%72.7%
LMArena WebDev14991618
SciCode53.1%57.1%
WeirdML62.6%89.4%
LMArena Coding14921498
ALE-Bench911.022,177
SWE-bench Verified79.3%—
FrontierCode—47.5%
CursorBench—41.7%
FrontierSWE—32.2%
GSO—76.5%
MirrorCode—20%

Agentic & Tool Use GPT-5.6 Sol leads

Gemini 3.5 Flash: 24.7 (#114), GPT-5.6 Sol: 50.3 (#7)

Agentic & Tool Use benchmarks
BenchmarkGemini 3.5 FlashGPT-5.6 Sol
APEX-Agents27.5%51.4%
GBAEval6.7%52.6%
GDP.pdf14%30.7%
Vending-Bench 25,3969,619
OSWorld 2.0—27.3%
τ²-bench Banking—46.9%
PostTrainBench—36.2%
BALROG—60%
LMArena Search—1257

Reasoning GPT-5.6 Sol leads

Gemini 3.5 Flash: 62.8 (#18), GPT-5.6 Sol: 74.8 (#8)

Reasoning benchmarks
BenchmarkGemini 3.5 FlashGPT-5.6 Sol
ARC-AGI-272.1%92.5%
SimpleBench76.7%71.7%
NYT Connections (extended)92.6%93.8%
ARC-AGI-192.5%97.5%
CritPt13.1%32.3%
Chess Puzzles50%64%
EnigmaEval25.4%37.1%
EBR-Bench4.8%44.8%
LMArena Hard Prompts14881484
Mystery Game Puzzles32%58%
DTBench94.7%96%
LMCA47.1%59.2%
Surface Evolver Bench58.1%93.1%
Epoch Capabilities Index154.46161.66
Kagi LLM Benchmark—67%
Bench to the Future 3—0.14
ForecastBench59—

Math GPT-5.6 Sol leads

Gemini 3.5 Flash: 60.7 (#36), GPT-5.6 Sol: 85.6 (#9)

Knowledge Gemini 3.5 Flash leads

Gemini 3.5 Flash: 66.3 (#11), GPT-5.6 Sol: 64.3 (#18)

Knowledge benchmarks
BenchmarkGemini 3.5 FlashGPT-5.6 Sol
GPQA Diamond92.8%93.5%
SimpleQA Verified66.2%69.7%
LMArena Expert14951516
Vectara Hallucination Rate—12.4%

Multimodal GPT-5.6 Sol leads

Gemini 3.5 Flash: 45.7 (#15), GPT-5.6 Sol: 48.6 (#9)

Multimodal benchmarks
BenchmarkGemini 3.5 FlashGPT-5.6 Sol
LMArena Vision13101281
Blueprint-Bench 233.6%33.6%
LMArena Document14631483
Furniture Assembly—56.7%

Multilingual Gemini 3.5 Flash leads

Gemini 3.5 Flash: 57.0 (#13), GPT-5.6 Sol: 55.3 (#32)

Multilingual benchmarks
BenchmarkGemini 3.5 FlashGPT-5.6 Sol
LMArena Non-English14761452
LMArena Chinese15261527
LMArena French14901477
LMArena German14921476
LMArena Japanese14861471
LMArena Korean14511442
LMArena Russian14931468
LMArena Spanish14801441

Instruction Following Too close to call

Gemini 3.5 Flash: 77.0 (#30), GPT-5.6 Sol: 77.7 (#16)

Instruction Following benchmarks
BenchmarkGemini 3.5 FlashGPT-5.6 Sol
LMArena Instruction Following14671482

Long Context Too close to call

Gemini 3.5 Flash: 45.4 (#38), GPT-5.6 Sol: 45.4 (#42)

Long Context benchmarks
BenchmarkGemini 3.5 FlashGPT-5.6 Sol
LMArena Longer Query14821480

Writing & Preference GPT-5.6 Sol leads

Gemini 3.5 Flash: 65.5 (#47), GPT-5.6 Sol: 73.3 (#12)

Writing & Preference benchmarks
BenchmarkGemini 3.5 FlashGPT-5.6 Sol
LMArena Text14821457
LMArena Creative Writing14701448
EQ-Bench 410871250
LMArena Multi-Turn14811460
EQ-Bench Creative Writing—1972

Frequently asked questions

Is Gemini 3.5 Flash better than GPT-5.6 Sol?

GPT-5.6 Sol is the stronger model overall, scoring 65.0 to 54.2 on the Noometry Index. Gemini 3.5 Flash costs 2.4× less per token, which makes it the better buy when GPT-5.6 Sol's lead doesn't matter for your workload.

Which is cheaper, Gemini 3.5 Flash or GPT-5.6 Sol?

Gemini 3.5 Flash is cheaper. It lists at $1.50 per million input tokens and $9 per million output tokens; GPT-5.6 Sol lists at $4 and $20.

Is Gemini 3.5 Flash or GPT-5.6 Sol better for coding?

GPT-5.6 Sol scores higher on coding benchmarks: 65.1 versus 49.4 in the Noometry coding category.

Which has the bigger context window?

GPT-5.6 Sol does, with 1.05M tokens against 1.05M.

How many benchmarks do Gemini 3.5 Flash and GPT-5.6 Sol share?

49 benchmarks have published results for both models. Gemini 3.5 Flash has 54 scored results on Noometry and GPT-5.6 Sol has 65.

Related comparisons

Go deeper