Model comparison

Gemini 2.5 Flash-Lite vs Llama 3.1-405B

Gemini 2.5 Flash-Lite is the stronger model overall, scoring 37.0 to 30.7 on the Noometry Index.

Last verified . 26 shared benchmarks.

Gemini 2.5 Flash-Lite Google

37.0

Rank #211 Confirmed

Llama 3.1-405B Meta

30.7

Rank #288 Confirmed

Summary

  • They share 26 benchmarks with published results for both. Gemini 2.5 Flash-Lite scores higher in 8 categories and Llama 3.1-405B in 1 category; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemini 2.5 Flash-Lite leads 38.0 to 18.4.
  • The biggest single-benchmark swing is Omni-MATH: 48% for Gemini 2.5 Flash-Lite and 24.9% for Llama 3.1-405B.
  • Llama 3.1-405B has downloadable open weights; the other is API-only.

Side by side

Gemini 2.5 Flash-Lite and Llama 3.1-405B specifications
Gemini 2.5 Flash-LiteLlama 3.1-405B
ProviderGoogleMeta
Noometry Index37.030.7
Released2025-06-172024-07-23
WeightsProprietaryOpen
Context window1.05M—
Max output66K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.40—
Results tracked3342

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 38.5 (#173), Llama 3.1-405B: 33.1 (#262)

Coding benchmarks
BenchmarkGemini 2.5 Flash-LiteLlama 3.1-405B
WeirdML35.2%21.4%
LMArena Coding13731291
ALE-Bench325.9—

Agentic & Tool Use Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 28.0 (#96), Llama 3.1-405B: 21.0 (#140)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.5 Flash-LiteLlama 3.1-405B
Berkeley Function Calling Leaderboard36.9%—
TheAgentCompany—7.4%
Cybench—7.5%

Reasoning Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 22.2 (#205), Llama 3.1-405B: 16.8 (#300)

Reasoning benchmarks
BenchmarkGemini 2.5 Flash-LiteLlama 3.1-405B
Kagi LLM Benchmark40.5%45%
LMArena Hard Prompts13771269
DTBench62.8%61.4%
Epoch Capabilities Index133.94128.75
SimpleBench—23%
LMCA18.1%—
BIG-Bench Hard—82.9%
ForecastBench—59.9
HellaSwag—89.2%
PIQA—85.9%
WinoGrande—89.2%

Math Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 38.0 (#144), Llama 3.1-405B: 18.4 (#290)

Math benchmarks
BenchmarkGemini 2.5 Flash-LiteLlama 3.1-405B
Omni-MATH48%24.9%
LMArena Math13731281
OTIS Mock AIME 2024-2025—9.7%
MATH Level 5—49.8%

Knowledge Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 32.5 (#210), Llama 3.1-405B: 30.4 (#227)

Knowledge benchmarks
BenchmarkGemini 2.5 Flash-LiteLlama 3.1-405B
MMLU-Pro53.7%72.3%
GPQA (HELM)30.9%52.2%
LMArena Expert13731243
GPQA Diamond—50.9%
Confabulations—17.6%
Vectara Hallucination Rate3.3%—
ARC (AI2) Challenge—95.3%
MMLU—84.5%
TriviaQA—82.7%

Multimodal Not comparable

Gemini 2.5 Flash-Lite: 29.1 (#114), Llama 3.1-405B: —

Multimodal benchmarks
BenchmarkGemini 2.5 Flash-LiteLlama 3.1-405B
LMArena Vision1198—
VPCT30%—

Multilingual Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 49.3 (#134), Llama 3.1-405B: 40.7 (#214)

Multilingual benchmarks
BenchmarkGemini 2.5 Flash-LiteLlama 3.1-405B
LMArena Non-English13691248
LMArena Chinese14041242
LMArena French13881279
LMArena German13891252
LMArena Japanese13591208
LMArena Korean13601184
LMArena Russian13731265
LMArena Spanish13961260

Instruction Following Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 70.0 (#168), Llama 3.1-405B: 65.9 (#214)

Instruction Following benchmarks
BenchmarkGemini 2.5 Flash-LiteLlama 3.1-405B
IFEval81%81.1%
LMArena Instruction Following13671259

Long Context Llama 3.1-405B leads

Gemini 2.5 Flash-Lite: 33.3 (#262), Llama 3.1-405B: 38.4 (#197)

Long Context benchmarks
BenchmarkGemini 2.5 Flash-LiteLlama 3.1-405B
LMArena Longer Query13731266
Fiction.LiveBench47.2%—

Writing & Preference Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 56.8 (#135), Llama 3.1-405B: 38.9 (#251)

Writing & Preference benchmarks
BenchmarkGemini 2.5 Flash-LiteLlama 3.1-405B
LMArena Text13791284
LMArena Creative Writing13671262
WildBench81.8%78.3%
LMArena Multi-Turn13661297
EQ-Bench Creative Writing—870

Frequently asked questions

Is Gemini 2.5 Flash-Lite better than Llama 3.1-405B?

Gemini 2.5 Flash-Lite is the stronger model overall, scoring 37.0 to 30.7 on the Noometry Index.

Is Gemini 2.5 Flash-Lite or Llama 3.1-405B better for coding?

Gemini 2.5 Flash-Lite scores higher on coding benchmarks: 38.5 versus 33.1 in the Noometry coding category.

How many benchmarks do Gemini 2.5 Flash-Lite and Llama 3.1-405B share?

26 benchmarks have published results for both models. Gemini 2.5 Flash-Lite has 33 scored results on Noometry and Llama 3.1-405B has 42.

Related comparisons

Go deeper