Meta, open weights

Llama 3.1-8B

Llama 3.1-8B by Meta ranks 352nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 23.0. Its strongest category is agentic & tool use, where it ranks 131st. API pricing starts at $0.05 per million input tokens and $0.08 per million output tokens, with a 128K-token context window.

Last verified

Specifications

Noometry rank
#352 of 354
Index score
23.0
Evidence
Confirmed 43 results
Provider
Meta
Released
July 23, 2024
Weights
Open weights
Reasoning
No
Context window
128K
Max output
4K
Input price
$0.05 / M
Output price
$0.08 / M
Blended price
$0.0575 / M
Output speed
Not measured
Value
#13 of 219
Knowledge cutoff
December 2023
Input
text

Category scores

Each category score combines every public result we have in that category.

Llama 3.1-8B category scores
  1. Coding 20.2
  2. Agentic & Tool Use 22.5
  3. Reasoning 14.9
  4. Math 10.2
  5. Knowledge 8.0
  6. Multilingual 34.0
  7. Instruction Following 58.9
  8. Long Context 35.8
  9. Writing & Preference 29.7
Llama 3.1-8B category ranks
CategoryScoreRankResults
Coding20.2#3405
Agentic & Tool Use22.5#1312
Reasoning14.9#3215
Math10.2#3174
Knowledge8.0#3074
Multilingual34.0#2491
Instruction Following58.9#2582
Long Context35.8#2381
Writing & Preference29.7#2905

Strengths and weaknesses

Categories where Llama 3.1-8B places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Llama 3.1-8B: strongest categories
CategoryScorevs medianRank
Long Context35.8−5.1#238 of 296, top 81%
Multilingual34.0−13.4#249 of 297, top 84%
Instruction Following58.9−12.4#258 of 305, top 85%

Weakest categories

Llama 3.1-8B: weakest categories
CategoryScorevs medianRank
Coding20.2−18.5#340 of 340, top 100%
Knowledge8.0−29.3#307 of 314, top 98%
Math10.2−26.4#317 of 327, top 97%

Closest competitors

The models ranked just above and below Llama 3.1-8B. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Llama 3.1-8B
ModelRankScoreBlended $/MSpeed
Claude 2#34625.0——Compare
DeepSeek LLM 67B#34724.9——Compare
Llama 13b#34824.4——Compare
Llama 2-70B#34924.4——Compare
GPT-3.5-turbo#35023.2$0.75—Compare
Mistral 7B#35123.0$0.25—Compare
Gemma 3 1B#35321.1——Compare
Llama 3.2 1B#35420.1$0.0705—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Llama 3.1-8B Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SciCode13.2%#120 of 121, top 100%Epoch AI
WeirdML1.7%#119 of 119, top 100%Epoch AI
BigCodeBench Instruct32.8%#51 of 64, top 80%BigCodeBench2024-07-23
LMArena Coding1195#244 of 294, top 83%LMArena2026-10-08
BigCodeBench Complete40.5%#52 of 66, top 79%BigCodeBench2024-07-23
HumanEval+62.8%#25 of 45, top 56%EvalPlus
MBPP+55.6%#28 of 38, top 74%EvalPlus

Agentic & Tool Use

Llama 3.1-8B Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Berkeley Function Calling Leaderboard25.8%#41 of 49, top 84%promptBerkeley Function Calling Leaderboard
BALROG15.1%#30 of 35, top 86%Epoch AI

Reasoning

Llama 3.1-8B Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
CritPt0%#116 of 134, top 87%Epoch AI
Chess Puzzles0%#120 of 129, top 94%Epoch AI2026-08-27
LMArena Hard Prompts1175#245 of 297, top 83%LMArena2026-10-08
DTBench50.9%#135 of 151, top 90%Epoch AI
LMCA5.4%#122 of 125, top 98%Epoch AI
Epoch Capabilities Index116.57#182 of 213, top 86%Epoch AI2024-07-23
PIQA81.2%#18 of 27, top 67%Epoch AI

Math

Llama 3.1-8B Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-20251.7%#164 of 173, top 95%Epoch AI2026-08-27
Omni-MATH13.7%#55 of 57, top 97%HELM Capabilities
LMArena Math1179#242 of 285, top 85%LMArena2026-10-08
MATH Level 522.9%#61 of 79, top 78%Epoch AI2025-01-27
GSM8K82.4%#12 of 38, top 32%Epoch AI

Knowledge

Llama 3.1-8B Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond27%#177 of 186, top 96%Epoch AI2026-08-27
MMLU-Pro40.6%#54 of 58, top 94%HELM Capabilities
GPQA (HELM)24.7%#57 of 57, top 100%HELM Capabilities
LMArena Expert1144#239 of 273, top 88%LMArena2026-10-08
BoolQ82.8%#13 of 23, top 57%Epoch AI
MMLU56.1%#67 of 81, top 83%Epoch AI

Multilingual

Llama 3.1-8B Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1148#249 of 297, top 84%LMArena2026-10-08
LMArena Chinese1151#246 of 285, top 87%LMArena2026-10-08
LMArena French1177#198 of 223, top 89%LMArena2026-10-08
LMArena German1144#203 of 231, top 88%LMArena2026-10-08
LMArena Japanese1061#190 of 211, top 91%LMArena2026-10-08
LMArena Korean1053#193 of 213, top 91%LMArena2026-10-08
LMArena Russian1158#248 of 283, top 88%LMArena2026-10-08
LMArena Spanish1169#200 of 226, top 89%LMArena2026-10-08

Instruction Following

Llama 3.1-8B Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval74.3%#51 of 57, top 90%HELM Capabilities
LMArena Instruction Following1159#250 of 298, top 84%LMArena2026-10-08

Long Context

Llama 3.1-8B Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1182#247 of 291, top 85%LMArena2026-10-08

Writing & Preference

Llama 3.1-8B Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1187#249 of 297, top 84%LMArena2026-10-08
LMArena Creative Writing1154#249 of 295, top 85%LMArena2026-10-08
EQ-Bench Creative Writing713#109 of 115, top 95%EQ-Bench
WildBench68.7%#53 of 57, top 93%HELM Capabilities
LMArena Multi-Turn1172#245 of 295, top 84%LMArena2026-10-08

API pricing by provider

Llama 3.1-8B API prices
RouteInput $/MOutput $/MCached input $/MChecked
bedrock$0.22$0.22—2026-10-10
groq$0.05$0.08—2026-10-10
openrouter$0.05$0.08$0.0252026-10-10

Compare Llama 3.1-8B

Other Meta models

Frequently asked questions

How good is Llama 3.1-8B?

Llama 3.1-8B by Meta ranks 352nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 23.0. Its strongest category is agentic & tool use, where it ranks 131st. API pricing starts at $0.05 per million input tokens and $0.08 per million output tokens, with a 128K-token context window.

How much does Llama 3.1-8B cost?

Llama 3.1-8B costs $0.05 per million input tokens and $0.08 per million output tokens on groq.

What is Llama 3.1-8B's context window?

Llama 3.1-8B accepts up to 128K tokens of input and can write up to 4K tokens in one response.

Is Llama 3.1-8B open source?

Yes. Llama 3.1-8B's weights are downloadable from Hugging Face (meta-llama/Meta-Llama-3.1-8B-Instruct); check the license for commercial terms.

What are Llama 3.1-8B's strengths and weaknesses?

Relative to other ranked models, Llama 3.1-8B places best in long context, multilingual, instruction following and lowest in coding, knowledge, math.

What is Llama 3.1-8B best at?

Its best category is agentic & tool use, where it ranks 131st on Noometry.