Meta, open weights

Llama-3.3-70B-Instruct

Llama-3.3-70B-Instruct by Meta ranks 291st of 354 ranked models on the Noometry Index as of October 2026, with a score of 30.6. Its strongest category is agentic & tool use, where it ranks 105th. API pricing starts at $0.10 per million input tokens and $0.32 per million output tokens, with a 128K-token context window.

Last verified

Specifications

Noometry rank
#291 of 354
Index score
30.6
Evidence
Confirmed 43 results
Provider
Meta
Released
December 6, 2024
Weights
Open weights
Reasoning
No
Context window
128K
Max output
4K
Input price
$0.10 / M
Output price
$0.32 / M
Blended price
$0.16 / M
Output speed
Not measured
Value
#43 of 219
Knowledge cutoff
December 2023
Input
text

Category scores

Each category score combines every public result we have in that category.

Llama-3.3-70B-Instruct category scores
  1. Coding 31.0
  2. Agentic & Tool Use 25.8
  3. Reasoning 14.1
  4. Math 15.3
  5. Knowledge 30.6
  6. Multilingual 39.9
  7. Instruction Following 71.1
  8. Long Context 26.4
  9. Writing & Preference 47.6
Llama-3.3-70B-Instruct category ranks
CategoryScoreRankResults
Coding31.0#2906
Agentic & Tool Use25.8#1052
Reasoning14.1#3277
Math15.3#2984
Knowledge30.6#2264
Multilingual39.9#2201
Instruction Following71.1#1572
Long Context26.4#2952
Writing & Preference47.6#2074

Strengths and weaknesses

Categories where Llama-3.3-70B-Instruct places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Llama-3.3-70B-Instruct: strongest categories
CategoryScorevs medianRank
Instruction Following71.1−0.2#157 of 305, top 52%
Writing & Preference47.6−6.2#207 of 312, top 67%
Agentic & Tool Use25.8−4.5#105 of 154, top 69%

Weakest categories

Llama-3.3-70B-Instruct: weakest categories
CategoryScorevs medianRank
Long Context26.4−14.5#295 of 296, top 100%
Reasoning14.1−9.5#327 of 350, top 94%
Math15.3−21.3#298 of 327, top 92%

Closest competitors

The models ranked just above and below Llama-3.3-70B-Instruct. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Llama-3.3-70B-Instruct
ModelRankScoreBlended $/MSpeed
Codellama 34b Instruct#28730.8——Compare
Llama 3.1-405B#28830.7—78Compare
Yi-1.5-34B#28930.6——Compare
Codestral#29030.6$0.45271Compare
GPT-4 Turbo#29230.5$15—Compare
Qwen1.5-32B#29330.5——Compare
Amazon Nova Micro#29430.4$0.0613—Compare
Olmo 7b Instruct#29530.3——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Llama-3.3-70B-Instruct Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SciCode26%#111 of 121, top 92%Epoch AI
WeirdML14.4%#109 of 119, top 92%Epoch AI
BigCodeBench Instruct46.9%#10 of 64, top 16%BigCodeBench2024-12-19
LiveBench Coding36.6%#26 of 39, top 67%Epoch AI
LMArena Coding1268#219 of 294, top 75%LMArena2026-10-08
BigCodeBench Complete57.5%#12 of 66, top 19%BigCodeBench2024-12-19

Agentic & Tool Use

Llama-3.3-70B-Instruct Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Berkeley Function Calling Leaderboard31.9%#32 of 49, top 66%fcBerkeley Function Calling Leaderboard
BALROG23%#22 of 35, top 63%Epoch AI

Reasoning

Llama-3.3-70B-Instruct Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
SimpleBench19.9%#73 of 77, top 95%Epoch AI
CritPt0%#117 of 134, top 88%Epoch AI
LiveBench Reasoning50.8%#19 of 39, top 49%Epoch AI
LMArena Hard Prompts1257#214 of 297, top 73%LMArena2026-10-08
DTBench59.5%#118 of 151, top 79%Epoch AI
LiveBench Data Analysis49.5%#25 of 39, top 65%Epoch AI
LMCA17.5%#97 of 125, top 78%Epoch AI
Epoch Capabilities Index127.33#148 of 213, top 70%Epoch AI2024-12-06
ForecastBench58.6#45 of 72, top 63%Epoch AI
LiveBench50.2%#19 of 39, top 49%Epoch AI

Math

Llama-3.3-70B-Instruct Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-20255.1%#147 of 173, top 85%Epoch AI2025-02-25
LiveBench Math42.2%#24 of 39, top 62%Epoch AI
LMArena Math1267#206 of 285, top 73%LMArena2026-10-08
MATH Level 541.6%#51 of 79, top 65%Epoch AI2025-01-27

Knowledge

Llama-3.3-70B-Instruct Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond47.4%#135 of 186, top 73%Epoch AI2025-01-27
Confabulations (lower is better)22.8%#40 of 51, top 79%Lech Mazur benchmarks
Vectara Hallucination Rate (lower is better)4.1%#4 of 96, top 5%Vectara Hallucination Leaderboard
LMArena Expert1225#209 of 273, top 77%LMArena2026-10-08
MMLU86.3%#6 of 81, top 8%Epoch AI

Multilingual

Llama-3.3-70B-Instruct Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1236#220 of 297, top 75%LMArena2026-10-08
LMArena Chinese1217#224 of 285, top 79%LMArena2026-10-08
LMArena French1281#170 of 223, top 77%LMArena2026-10-08
LMArena German1251#177 of 231, top 77%LMArena2026-10-08
LMArena Japanese1150#175 of 211, top 83%LMArena2026-10-08
LMArena Korean1143#180 of 213, top 85%LMArena2026-10-08
LMArena Russian1252#212 of 283, top 75%LMArena2026-10-08
LMArena Spanish1270#173 of 226, top 77%LMArena2026-10-08

Instruction Following

Llama-3.3-70B-Instruct Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following82.7%#5 of 39, top 13%Epoch AI
LMArena Instruction Following1242#219 of 298, top 74%LMArena2026-10-08

Long Context

Llama-3.3-70B-Instruct Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench33.3%#46 of 47, top 98%Epoch AI
LMArena Longer Query1256#217 of 291, top 75%LMArena2026-10-08

Writing & Preference

Llama-3.3-70B-Instruct Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1274#213 of 297, top 72%LMArena2026-10-08
LMArena Creative Writing1250#204 of 295, top 70%LMArena2026-10-08
LMArena Multi-Turn1280#201 of 295, top 69%LMArena2026-10-08
LiveBench Language39.2%#20 of 39, top 52%Epoch AI

API pricing by provider

Llama-3.3-70B-Instruct API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$0.71$0.71—2026-10-10
bedrock$0.72$0.72—2026-10-10
deepinfra$0.10$0.32—2026-10-10
groq$0.59$0.79—2026-10-10
openrouter$0.22$0.50$0.112026-10-10
together$1.04$1.04—2026-10-10
vertex$0.72$0.72—2026-10-10

Compare Llama-3.3-70B-Instruct

Other Meta models

Frequently asked questions

How good is Llama-3.3-70B-Instruct?

Llama-3.3-70B-Instruct by Meta ranks 291st of 354 ranked models on the Noometry Index as of October 2026, with a score of 30.6. Its strongest category is agentic & tool use, where it ranks 105th. API pricing starts at $0.10 per million input tokens and $0.32 per million output tokens, with a 128K-token context window.

How much does Llama-3.3-70B-Instruct cost?

Llama-3.3-70B-Instruct costs $0.10 per million input tokens and $0.32 per million output tokens on deepinfra.

What is Llama-3.3-70B-Instruct's context window?

Llama-3.3-70B-Instruct accepts up to 128K tokens of input and can write up to 4K tokens in one response.

Is Llama-3.3-70B-Instruct open source?

Yes. Llama-3.3-70B-Instruct's weights are downloadable from Hugging Face (meta-llama/Llama-3.3-70B-Instruct); check the license for commercial terms.

What are Llama-3.3-70B-Instruct's strengths and weaknesses?

Relative to other ranked models, Llama-3.3-70B-Instruct places best in instruction following, writing & preference, agentic & tool use and lowest in long context, reasoning, math.

What is Llama-3.3-70B-Instruct best at?

Its best category is agentic & tool use, where it ranks 105th on Noometry.