Model comparison

GPT-6 Sol vs Llama 4 Maverick

GPT-6 Sol is the stronger model overall, scoring 61.8 to 30.9 on the Noometry Index. Llama 4 Maverick costs 13× less per token, which makes it the better buy when GPT-6 Sol's lead doesn't matter for your workload.

Last verified . 31 shared benchmarks.

GPT-6 Sol OpenAI

61.8

Rank #12 Confirmed

Llama 4 Maverick Meta

30.9

Rank #282 Confirmed

Summary

  • They share 31 benchmarks with published results for both. GPT-6 Sol scores higher in 10 categories and Llama 4 Maverick in 0 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GPT-6 Sol leads 74.0 to 10.1.
  • The biggest single-benchmark swing is ARC-AGI-1: 95.5% for GPT-6 Sol and 4.4% for Llama 4 Maverick.
  • Llama 4 Maverick is cheaper at $0.19 / $0.65 per million input/output tokens, against $2 / $10 for GPT-6 Sol.
  • GPT-6 Sol accepts more context: 1.05M tokens versus 128K.
  • Llama 4 Maverick has downloadable open weights; the other is API-only.

Side by side

GPT-6 Sol and Llama 4 Maverick specifications
GPT-6 SolLlama 4 Maverick
ProviderOpenAIMeta
Noometry Index61.830.9
Released2026-09-222025-04-05
WeightsProprietaryOpen
Context window1.05M128K
Max output128K4K
Input $ / M tokens$2$0.19
Output $ / M tokens$10$0.65
Results tracked4554

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-6 Sol leads

GPT-6 Sol: 60.1 (#11), Llama 4 Maverick: 26.6 (#324)

Coding benchmarks
BenchmarkGPT-6 SolLlama 4 Maverick
SciCode57.6%33.1%
LMArena Coding14471302
ALE-Bench2,462172.97
DeepSWE68.8%—
FrontierCode49.3%—
SWE-bench Verified (bash only)—21%
Aider Polyglot—15.6%
LMArena WebDev1688—
WeirdML—24.5%
BigCodeBench Instruct—49.7%
BigCodeBench Complete—61.4%

Agentic & Tool Use GPT-6 Sol leads

GPT-6 Sol: 37.2 (#36), Llama 4 Maverick: 28.2 (#91)

Agentic & Tool Use benchmarks
BenchmarkGPT-6 SolLlama 4 Maverick
APEX-Agents54.3%—
Berkeley Function Calling Leaderboard—37.3%
GDP.pdf26.4%—
Vending-Bench 214,428—

Reasoning GPT-6 Sol leads

GPT-6 Sol: 74.0 (#9), Llama 4 Maverick: 10.1 (#342)

Reasoning benchmarks
BenchmarkGPT-6 SolLlama 4 Maverick
ARC-AGI-289.6%0%
NYT Connections (extended)90.1%8%
ARC-AGI-195.5%4.4%
CritPt30.9%0%
LMArena Hard Prompts14181281
DTBench97.3%61.9%
LMCA59.1%15.9%
Epoch Capabilities Index162.72132.2
SimpleBench—27.7%
Kagi LLM Benchmark—55.9%
EnigmaEval—0.6%
EBR-Bench53.3%—
Mystery Game Puzzles56%—
ForecastBench—57.5

Math GPT-6 Sol leads

GPT-6 Sol: 87.2 (#7), Llama 4 Maverick: 26.0 (#262)

Math benchmarks
BenchmarkGPT-6 SolLlama 4 Maverick
OTIS Mock AIME 2024-2025100%20.6%
LMArena Math14021299
FrontierMath (Tiers 1-3)89.8%—
FrontierMath Tier 490%—
ProofBench83%—
Omni-MATH—42.2%
MATH Level 5—73%
FrontierMath (Feb 2025 set)—0.7%

Knowledge GPT-6 Sol leads

GPT-6 Sol: 64.8 (#15), Llama 4 Maverick: 33.4 (#204)

Knowledge benchmarks
BenchmarkGPT-6 SolLlama 4 Maverick
GPQA Diamond94.3%67%
Vectara Hallucination Rate6.5%8.2%
LMArena Expert14391259
Humanity's Last Exam—5.7%
SimpleQA Verified60.7%—
MMLU-Pro—81%
Confabulations—22.6%
GPQA (HELM)—65%

Multimodal GPT-6 Sol leads

GPT-6 Sol: 47.6 (#10), Llama 4 Maverick: 31.6 (#105)

Multimodal benchmarks
BenchmarkGPT-6 SolLlama 4 Maverick
LMArena Vision12451142
GeoBench—52%
Blueprint-Bench 236.9%—
Furniture Assembly58.3%—
SpatialViz-Bench—31.8%

Multilingual GPT-6 Sol leads

GPT-6 Sol: 50.5 (#118), Llama 4 Maverick: 42.2 (#195)

Multilingual benchmarks
BenchmarkGPT-6 SolLlama 4 Maverick
LMArena Non-English13851269
LMArena Chinese14051277
LMArena French14101259
LMArena German13901291
LMArena Japanese13851207
LMArena Korean13411203
LMArena Russian14011286
LMArena Spanish13841293

Instruction Following GPT-6 Sol leads

GPT-6 Sol: 74.5 (#94), Llama 4 Maverick: 71.7 (#146)

Instruction Following benchmarks
BenchmarkGPT-6 SolLlama 4 Maverick
LMArena Instruction Following14121267
IFEval—90.8%

Long Context GPT-6 Sol leads

GPT-6 Sol: 43.1 (#108), Llama 4 Maverick: 31.4 (#279)

Long Context benchmarks
BenchmarkGPT-6 SolLlama 4 Maverick
LMArena Longer Query14111280
Fiction.LiveBench—46.2%

Writing & Preference GPT-6 Sol leads

GPT-6 Sol: 71.9 (#18), Llama 4 Maverick: 38.8 (#252)

Writing & Preference benchmarks
BenchmarkGPT-6 SolLlama 4 Maverick
LMArena Text13951287
LMArena Creative Writing13781267
EQ-Bench Creative Writing2125860
LMArena Multi-Turn14121289
Short-Story Creative Writing—62%
WildBench—80%

Frequently asked questions

Is GPT-6 Sol better than Llama 4 Maverick?

GPT-6 Sol is the stronger model overall, scoring 61.8 to 30.9 on the Noometry Index. Llama 4 Maverick costs 13× less per token, which makes it the better buy when GPT-6 Sol's lead doesn't matter for your workload.

Which is cheaper, GPT-6 Sol or Llama 4 Maverick?

Llama 4 Maverick is cheaper. It lists at $0.19 per million input tokens and $0.65 per million output tokens; GPT-6 Sol lists at $2 and $10.

Is GPT-6 Sol or Llama 4 Maverick better for coding?

GPT-6 Sol scores higher on coding benchmarks: 60.1 versus 26.6 in the Noometry coding category.

Which has the bigger context window?

GPT-6 Sol does, with 1.05M tokens against 128K.

How many benchmarks do GPT-6 Sol and Llama 4 Maverick share?

31 benchmarks have published results for both models. GPT-6 Sol has 45 scored results on Noometry and Llama 4 Maverick has 54.

Related comparisons

Go deeper