Model comparison

INTELLECT-1 vs Llama 3.1-405B

Llama 3.1-405B has enough public results to be ranked (#288); INTELLECT-1 does not yet, so treat this comparison as directional.

Last verified . 6 shared benchmarks.

Llama 3.1-405B Meta

30.7

Rank #288 Confirmed

Summary

  • They share 6 benchmarks with published results for both.
  • Llama 3.1-405B has downloadable open weights; the other is API-only.

Side by side

INTELLECT-1 and Llama 3.1-405B specifications
INTELLECT-1Llama 3.1-405B
ProviderHugging FaceMeta
Noometry Index—30.7
Released2024-11-292024-07-23
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked742

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

INTELLECT-1: —, Llama 3.1-405B: 33.1 (#262)

Coding benchmarks
BenchmarkINTELLECT-1Llama 3.1-405B
WeirdML—21.4%
LMArena Coding—1291

Agentic & Tool Use Not comparable

INTELLECT-1: —, Llama 3.1-405B: 21.0 (#140)

Agentic & Tool Use benchmarks
BenchmarkINTELLECT-1Llama 3.1-405B
TheAgentCompany—7.4%
Cybench—7.5%

Reasoning Not comparable

INTELLECT-1: —, Llama 3.1-405B: 16.8 (#300)

Reasoning benchmarks
BenchmarkINTELLECT-1Llama 3.1-405B
BIG-Bench Hard34.9%82.9%
Epoch Capabilities Index100.86128.75
HellaSwag71.4%89.2%
WinoGrande65.8%89.2%
SimpleBench—23%
Kagi LLM Benchmark—45%
LMArena Hard Prompts—1269
DTBench—61.4%
ForecastBench—59.9
PIQA—85.9%

Math Not comparable

INTELLECT-1: —, Llama 3.1-405B: 18.4 (#290)

Math benchmarks
BenchmarkINTELLECT-1Llama 3.1-405B
OTIS Mock AIME 2024-2025—9.7%
Omni-MATH—24.9%
LMArena Math—1281
MATH Level 5—49.8%
GSM8K38.6%—

Knowledge Not comparable

INTELLECT-1: —, Llama 3.1-405B: 30.4 (#227)

Knowledge benchmarks
BenchmarkINTELLECT-1Llama 3.1-405B
ARC (AI2) Challenge54.5%95.3%
MMLU49.9%84.5%
GPQA Diamond—50.9%
MMLU-Pro—72.3%
Confabulations—17.6%
GPQA (HELM)—52.2%
LMArena Expert—1243
TriviaQA—82.7%

Multilingual Not comparable

INTELLECT-1: —, Llama 3.1-405B: 40.7 (#214)

Multilingual benchmarks
BenchmarkINTELLECT-1Llama 3.1-405B
LMArena Non-English—1248
LMArena Chinese—1242
LMArena French—1279
LMArena German—1252
LMArena Japanese—1208
LMArena Korean—1184
LMArena Russian—1265
LMArena Spanish—1260

Instruction Following Not comparable

INTELLECT-1: —, Llama 3.1-405B: 65.9 (#214)

Instruction Following benchmarks
BenchmarkINTELLECT-1Llama 3.1-405B
IFEval—81.1%
LMArena Instruction Following—1259

Long Context Not comparable

INTELLECT-1: —, Llama 3.1-405B: 38.4 (#197)

Long Context benchmarks
BenchmarkINTELLECT-1Llama 3.1-405B
LMArena Longer Query—1266

Writing & Preference Not comparable

INTELLECT-1: —, Llama 3.1-405B: 38.9 (#251)

Writing & Preference benchmarks
BenchmarkINTELLECT-1Llama 3.1-405B
LMArena Text—1284
LMArena Creative Writing—1262
EQ-Bench Creative Writing—870
WildBench—78.3%
LMArena Multi-Turn—1297

Frequently asked questions

Is INTELLECT-1 better than Llama 3.1-405B?

Llama 3.1-405B has enough public results to be ranked (#288); INTELLECT-1 does not yet, so treat this comparison as directional.

How many benchmarks do INTELLECT-1 and Llama 3.1-405B share?

6 benchmarks have published results for both models. INTELLECT-1 has 7 scored results on Noometry and Llama 3.1-405B has 42.

Related comparisons

Go deeper