Meta, open weights

Llama 3-70B

Llama 3-70B by Meta ranks 323rd of 354 ranked models on the Noometry Index as of October 2026, with a score of 28.8. Its strongest category is agentic & tool use, where it ranks 139th.

Last verified

Specifications

Noometry rank
#323 of 354
Index score
28.8
Evidence
Confirmed 31 results
Provider
Meta
Released
April 18, 2024
Weights
Open weights
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
104 tokens/s Kagi
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Llama 3-70B category scores
  1. Coding 35.8
  2. Agentic & Tool Use 21.1
  3. Reasoning 18.0
  4. Math 12.8
  5. Knowledge 20.8
  6. Multilingual 33.6
  7. Instruction Following 62.5
  8. Long Context 35.6
  9. Writing & Preference 42.8
Llama 3-70B category ranks
CategoryScoreRankResults
Coding35.8#2183
Agentic & Tool Use21.1#1391
Reasoning18.0#2883
Math12.8#3053
Knowledge20.8#2772
Multilingual33.6#2511
Instruction Following62.5#2381
Long Context35.6#2401
Writing & Preference42.8#2313

Strengths and weaknesses

Categories where Llama 3-70B places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Llama 3-70B: strongest categories
CategoryScorevs medianRank
Coding35.8−2.9#218 of 340, top 65%
Writing & Preference42.8−10.9#231 of 312, top 75%
Instruction Following62.5−8.7#238 of 305, top 79%

Weakest categories

Llama 3-70B: weakest categories
CategoryScorevs medianRank
Math12.8−23.7#305 of 327, top 94%
Agentic & Tool Use21.1−9.2#139 of 154, top 91%
Knowledge20.8−16.5#277 of 314, top 89%

Closest competitors

The models ranked just above and below Llama 3-70B. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Llama 3-70B
ModelRankScoreBlended $/MSpeed
Claude 3 Sonnet#31929.0——Compare
Qwen2.5 7B Instruct#32029.0$0.31—Compare
Llama 3.2 3B#32128.9$0.12—Compare
Qwen1.5 4b Chat#32228.8——Compare
GPT-4o#32428.6$4.38—Compare
Ministral 8B#32528.2$0.15—Compare
Gemma 3 4B#32628.1$0.0572Compare
GPT-4.1 nano#32727.9$0.18135Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Llama 3-70B Coding benchmark results
BenchmarkScorePositionSettingSourceDate
BigCodeBench Instruct43.6%#25 of 64, top 40%BigCodeBench2024-04-18
LMArena Coding1206#239 of 294, top 82%LMArena2026-10-08
BigCodeBench Complete43.3%BigCodeBench2024-04-18
BigCodeBench Complete54.5%#21 of 66, top 32%BigCodeBench2024-04-18
HumanEval+72%#17 of 45, top 38%EvalPlus
MBPP+69%#16 of 38, top 43%EvalPlus

Agentic & Tool Use

Llama 3-70B Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Cybench5%#21 of 21, top 100%Epoch AI

Reasoning

Llama 3-70B Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
Kagi LLM Benchmark35.1%#88 of 99, top 89%Kagi LLM Benchmark
LMArena Hard Prompts1195#239 of 297, top 81%LMArena2026-10-08
DTBench54.2%#125 of 151, top 83%Epoch AI
Epoch Capabilities Index122.93#161 of 213, top 76%Epoch AI2024-04-18
ForecastBench57.1#61 of 72, top 85%Epoch AI
WinoGrande83.5%#9 of 43, top 21%Epoch AI

Math

Llama 3-70B Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-20254.3%#152 of 173, top 88%Epoch AI2025-02-25
LMArena Math1218#225 of 285, top 79%LMArena2026-10-08
MATH Level 522.6%#62 of 79, top 79%Epoch AI2025-01-27

Knowledge

Llama 3-70B Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond40.6%#150 of 186, top 81%Epoch AI2025-01-27
LMArena Expert1149#235 of 273, top 87%LMArena2026-10-08
MMLU79.3%#20 of 81, top 25%Epoch AI

Multilingual

Llama 3-70B Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1142#251 of 297, top 85%LMArena2026-10-08
LMArena Chinese1114#255 of 285, top 90%LMArena2026-10-08
LMArena French1232#188 of 223, top 85%LMArena2026-10-08
LMArena German1169#199 of 231, top 87%LMArena2026-10-08
LMArena Japanese1017#198 of 211, top 94%LMArena2026-10-08
LMArena Korean1017#198 of 213, top 93%LMArena2026-10-08
LMArena Russian1159#247 of 283, top 88%LMArena2026-10-08
LMArena Spanish1241#187 of 226, top 83%LMArena2026-10-08

Instruction Following

Llama 3-70B Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1194#240 of 298, top 81%LMArena2026-10-08

Long Context

Llama 3-70B Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1174#249 of 291, top 86%LMArena2026-10-08

Writing & Preference

Llama 3-70B Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1221#237 of 297, top 80%LMArena2026-10-08
LMArena Creative Writing1210#228 of 295, top 78%LMArena2026-10-08
LMArena Multi-Turn1223#230 of 295, top 78%LMArena2026-10-08

Compare Llama 3-70B

Other Meta models

Frequently asked questions

How good is Llama 3-70B?

Llama 3-70B by Meta ranks 323rd of 354 ranked models on the Noometry Index as of October 2026, with a score of 28.8. Its strongest category is agentic & tool use, where it ranks 139th.

Is Llama 3-70B open source?

Yes. Llama 3-70B's weights are downloadable; check the license for commercial terms.

How fast is Llama 3-70B?

Llama 3-70B generated about 104 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Llama 3-70B's strengths and weaknesses?

Relative to other ranked models, Llama 3-70B places best in coding, writing & preference, instruction following and lowest in math, agentic & tool use, knowledge.

What is Llama 3-70B best at?

Its best category is agentic & tool use, where it ranks 139th on Noometry.