Model comparison

Gemini 2.0 Flash (Feb 2025) vs Step 1o Turbo 202506

Step 1o Turbo 202506 is the stronger model overall, scoring 39.7 to 35.1 on the Noometry Index.

Last verified . 14 shared benchmarks.

Gemini 2.0 Flash (Feb 2025) Google

35.1

Rank #228 Confirmed

Step 1o Turbo 202506 StepFun

39.7

Rank #160 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Gemini 2.0 Flash (Feb 2025) scores higher in 4 categories and Step 1o Turbo 202506 in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Step 1o Turbo 202506 leads 26.8 to 15.2.

Side by side

Gemini 2.0 Flash (Feb 2025) and Step 1o Turbo 202506 specifications
Gemini 2.0 Flash (Feb 2025)Step 1o Turbo 202506
ProviderGoogleStepFun
Noometry Index35.139.7
Released2024-12-06—
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked5414

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 1o Turbo 202506 leads

Gemini 2.0 Flash (Feb 2025): 28.4 (#315), Step 1o Turbo 202506: 39.2 (#160)

Coding benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Step 1o Turbo 202506
LMArena Coding13501339
SWE-bench Verified (bash only)13.5%—
Aider Polyglot38.2%—
WeirdML25.8%—
BigCodeBench Instruct45.9%—
LiveBench Coding63.4%—
BigCodeBench Complete59.9%—
CadEval30%—

Agentic & Tool Use Not comparable

Gemini 2.0 Flash (Feb 2025): 28.1 (#92), Step 1o Turbo 202506: —

Agentic & Tool Use benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Step 1o Turbo 202506
TheAgentCompany11.4%—

Reasoning Step 1o Turbo 202506 leads

Gemini 2.0 Flash (Feb 2025): 15.2 (#318), Step 1o Turbo 202506: 26.8 (#129)

Reasoning benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Step 1o Turbo 202506
LMArena Hard Prompts13461335
ARC-AGI-21.3%—
SimpleBench31.1%—
Kagi LLM Benchmark37.8%—
EnigmaEval1.1%—
LiveBench Reasoning78.2%—
DTBench63.2%—
LiveBench Data Analysis69.4%—
Epoch Capabilities Index135.36—
LiveBench66.9%—

Math Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 37.9 (#146), Step 1o Turbo 202506: 36.6 (#164)

Math benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Step 1o Turbo 202506
LMArena Math13521318
OTIS Mock AIME 2024-202557.8%—
Omni-MATH45.9%—
LiveBench Math75.8%—
MATH Level 582.2%—
FrontierMath (Feb 2025 set)1.7%—

Knowledge Step 1o Turbo 202506 leads

Gemini 2.0 Flash (Feb 2025): 32.0 (#213), Step 1o Turbo 202506: 36.1 (#176)

Knowledge benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Step 1o Turbo 202506
LMArena Expert13391308
GPQA Diamond64.1%—
Humanity's Last Exam6.6%—
MMLU-Pro73.7%—
Confabulations12.4%—
GPQA (HELM)55.6%—
MMLU79.7%—

Multimodal Too close to call

Gemini 2.0 Flash (Feb 2025): 36.5 (#79), Step 1o Turbo 202506: 36.1 (#80)

Multimodal benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Step 1o Turbo 202506
LMArena Vision11581186
GeoBench77%—

Multilingual Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 47.4 (#149), Step 1o Turbo 202506: 45.3 (#173)

Multilingual benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Step 1o Turbo 202506
LMArena Non-English13421313
LMArena Chinese13731380
LMArena German13531314
LMArena Russian13511327
LMArena French1391—
LMArena Japanese1294—
LMArena Korean1313—
LMArena Spanish1363—

Instruction Following Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 74.4 (#97), Step 1o Turbo 202506: 69.2 (#176)

Instruction Following benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Step 1o Turbo 202506
LMArena Instruction Following13361310
LiveBench Instruction Following85.8%—
IFEval84.1%—

Long Context Step 1o Turbo 202506 leads

Gemini 2.0 Flash (Feb 2025): 38.1 (#203), Step 1o Turbo 202506: 41.0 (#146)

Long Context benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Step 1o Turbo 202506
LMArena Longer Query13441348
Fiction.LiveBench61.1%—

Writing & Preference Step 1o Turbo 202506 leads

Gemini 2.0 Flash (Feb 2025): 49.5 (#190), Step 1o Turbo 202506: 53.1 (#160)

Writing & Preference benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Step 1o Turbo 202506
LMArena Text13541336
LMArena Creative Writing13401307
LMArena Multi-Turn13501340
Short-Story Creative Writing73.8%—
EQ-Bench Creative Writing1128—
WildBench80%—
LiveBench Language51.3%—

Frequently asked questions

Is Gemini 2.0 Flash (Feb 2025) better than Step 1o Turbo 202506?

Step 1o Turbo 202506 is the stronger model overall, scoring 39.7 to 35.1 on the Noometry Index.

Is Gemini 2.0 Flash (Feb 2025) or Step 1o Turbo 202506 better for coding?

Step 1o Turbo 202506 scores higher on coding benchmarks: 39.2 versus 28.4 in the Noometry coding category.

How many benchmarks do Gemini 2.0 Flash (Feb 2025) and Step 1o Turbo 202506 share?

14 benchmarks have published results for both models. Gemini 2.0 Flash (Feb 2025) has 54 scored results on Noometry and Step 1o Turbo 202506 has 14.

Related comparisons

Go deeper