Model comparison

Claude Fable 5 vs Gemini 3.1 Pro Preview

Claude Fable 5 is the stronger model overall, scoring 66.8 to 56.7 on the Noometry Index. Gemini 3.1 Pro Preview costs 4.4× less per token, which makes it the better buy when Claude Fable 5's lead doesn't matter for your workload.

Last verified . 55 shared benchmarks.

Claude Fable 5 Anthropic

66.8

Rank #5 Confirmed

Gemini 3.1 Pro Preview Google

56.7

Rank #23 Confirmed

Summary

  • They share 55 benchmarks with published results for both. Claude Fable 5 scores higher in 8 categories and Gemini 3.1 Pro Preview in 2 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Claude Fable 5 leads 70.6 to 42.5.
  • The biggest single-benchmark swing is GBAEval: 74.5% for Claude Fable 5 and 0.8% for Gemini 3.1 Pro Preview.
  • Gemini 3.1 Pro Preview is cheaper at $2 / $12 per million input/output tokens, against $10 / $50 for Claude Fable 5.
  • Gemini 3.1 Pro Preview accepts more context: 1.05M tokens versus 1M.

Side by side

Claude Fable 5 and Gemini 3.1 Pro Preview specifications
Claude Fable 5Gemini 3.1 Pro Preview
ProviderAnthropicGoogle
Noometry Index66.856.7
Released2026-06-072026-02-19
WeightsProprietaryProprietary
Context window1M1.05M
Max output128K66K
Input $ / M tokens$10$2
Output $ / M tokens$50$12
Results tracked6271

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Fable 5 leads

Claude Fable 5: 70.6 (#4), Gemini 3.1 Pro Preview: 42.5 (#99)

Coding benchmarks
BenchmarkClaude Fable 5Gemini 3.1 Pro Preview
DeepSWE69.9%11.7%
LMArena WebDev16251447
SciCode61%58.9%
GSO78.4%22.6%
WeirdML91.9%72.1%
LMArena Coding15191484
MirrorCode63.9%8.9%
ALE-Bench2,0411,161
SWE-bench Verified—75.6%
FrontierCode53.5%—
FrontierSWE47%—
AlgoTune—2.02

Agentic & Tool Use Claude Fable 5 leads

Claude Fable 5: 54.0 (#2), Gemini 3.1 Pro Preview: 37.7 (#34)

Agentic & Tool Use benchmarks
BenchmarkClaude Fable 5Gemini 3.1 Pro Preview
APEX-Agents63.6%35.3%
τ²-bench Banking39.7%26%
PostTrainBench41.8%22%
GBAEval74.5%0.8%
GDP.pdf30%17%
LMArena Search12301211
Vending-Bench 25,6803,774
Terminal-Bench—80.2%
Remote Labor Index16.1%—
DeepResearch Bench—47.8%
BALROG—57%
ExploitBench—26.1%
METR Time Horizons—77%

Reasoning Claude Fable 5 leads

Claude Fable 5: 76.8 (#6), Gemini 3.1 Pro Preview: 71.7 (#12)

Reasoning benchmarks
BenchmarkClaude Fable 5Gemini 3.1 Pro Preview
ARC-AGI-289.2%77.1%
SimpleBench81.9%79.6%
NYT Connections (extended)92.7%97.4%
ARC-AGI-198.5%98%
CritPt28.6%17.7%
Chess Puzzles41%55%
EnigmaEval39.3%36.8%
EBR-Bench39.5%14.3%
LMArena Hard Prompts15081485
Mystery Game Puzzles52%34%
DTBench98.4%97.1%
LMCA61.1%53.8%
Epoch Capabilities Index162.06154.77
Kagi LLM Benchmark91.4%—
Thematic Generalization—79.4%
Surface Evolver Bench95%—
Bench to the Future 30.13—
ForecastBench—59

Math Claude Fable 5 leads

Claude Fable 5: 88.5 (#5), Gemini 3.1 Pro Preview: 62.1 (#34)

Math benchmarks
BenchmarkClaude Fable 5Gemini 3.1 Pro Preview
FrontierMath (Tiers 1-3)87%59.6%
FrontierMath Tier 490.2%26.8%
OTIS Mock AIME 2024-2025100%95.6%
ProofBench95%26%
LMArena Math15191485
MathArena Final-Answer Competitions—86.5%
FrontierMath (Feb 2025 set)—36.9%
FrontierMath Erdős0%—
FrontierMath Tier 4 (v1)—16.7%

Knowledge Gemini 3.1 Pro Preview leads

Claude Fable 5: 62.2 (#25), Gemini 3.1 Pro Preview: 71.8 (#3)

Knowledge benchmarks
BenchmarkClaude Fable 5Gemini 3.1 Pro Preview
GPQA Diamond85.9%94.4%
SimpleQA Verified70.7%73.5%
LMArena Expert15341485
Humanity's Last Exam—46.4%
Vectara Hallucination Rate—10.4%

Multimodal Claude Fable 5 leads

Claude Fable 5: 45.3 (#17), Gemini 3.1 Pro Preview: 37.9 (#69)

Multimodal benchmarks
BenchmarkClaude Fable 5Gemini 3.1 Pro Preview
LMArena Vision13241296
Blueprint-Bench 238.6%26.5%
Furniture Assembly35.8%26.7%
LMArena Document14961444

Multilingual Too close to call

Claude Fable 5: 57.3 (#9), Gemini 3.1 Pro Preview: 57.0 (#12)

Multilingual benchmarks
BenchmarkClaude Fable 5Gemini 3.1 Pro Preview
LMArena Non-English14811477
LMArena Chinese15431529
LMArena French15051487
LMArena German14861491
LMArena Japanese15061493
LMArena Korean14881455
LMArena Russian15041498
LMArena Spanish14981479

Instruction Following Claude Fable 5 leads

Claude Fable 5: 78.6 (#8), Gemini 3.1 Pro Preview: 77.0 (#32)

Instruction Following benchmarks
BenchmarkClaude Fable 5Gemini 3.1 Pro Preview
LMArena Instruction Following15021466

Long Context Gemini 3.1 Pro Preview leads

Claude Fable 5: 46.3 (#23), Gemini 3.1 Pro Preview: 47.4 (#18)

Long Context benchmarks
BenchmarkClaude Fable 5Gemini 3.1 Pro Preview
LMArena Longer Query15091483
CL-bench—20.8%
CL-bench Life—16.9%

Writing & Preference Claude Fable 5 leads

Claude Fable 5: 75.9 (#5), Gemini 3.1 Pro Preview: 66.1 (#37)

Writing & Preference benchmarks
BenchmarkClaude Fable 5Gemini 3.1 Pro Preview
LMArena Text14911481
LMArena Creative Writing14941482
EQ-Bench Creative Writing19431491
EQ-Bench 413401142
LMArena Multi-Turn15041488

Frequently asked questions

Is Claude Fable 5 better than Gemini 3.1 Pro Preview?

Claude Fable 5 is the stronger model overall, scoring 66.8 to 56.7 on the Noometry Index. Gemini 3.1 Pro Preview costs 4.4× less per token, which makes it the better buy when Claude Fable 5's lead doesn't matter for your workload.

Which is cheaper, Claude Fable 5 or Gemini 3.1 Pro Preview?

Gemini 3.1 Pro Preview is cheaper. It lists at $2 per million input tokens and $12 per million output tokens; Claude Fable 5 lists at $10 and $50.

Is Claude Fable 5 or Gemini 3.1 Pro Preview better for coding?

Claude Fable 5 scores higher on coding benchmarks: 70.6 versus 42.5 in the Noometry coding category.

Which has the bigger context window?

Gemini 3.1 Pro Preview does, with 1.05M tokens against 1M.

How many benchmarks do Claude Fable 5 and Gemini 3.1 Pro Preview share?

55 benchmarks have published results for both models. Claude Fable 5 has 62 scored results on Noometry and Gemini 3.1 Pro Preview has 71.

Related comparisons

Go deeper