Allen Institute for AI (Ai2), open weights
Llama 3.1 Tulu 3 8b
Llama 3.1 Tulu 3 8b by Allen Institute for AI (Ai2) ranks 224th of 354 ranked models on the Noometry Index as of October 2026, with a score of 35.7. Its strongest category is reasoning, where it ranks 188th.
Last verified
Specifications
- Noometry rank
- #224 of 354
- Index score
- 35.7
- Evidence
- Confirmed 11 results
- Provider
Allen Institute for AI (Ai2)
- Released
- Unknown
- Weights
- Open weights
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- Not measured
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Coding 34.4
- Reasoning 22.8
- Math 33.9
- Multilingual 35.4
- Instruction Following 61.3
- Long Context 35.8
- Writing & Preference 39.7
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 34.4 | #235 | 1 |
| Reasoning | 22.8 | #188 | 1 |
| Math | 33.9 | #198 | 1 |
| Multilingual | 35.4 | #246 | 1 |
| Instruction Following | 61.3 | #246 | 1 |
| Long Context | 35.8 | #239 | 1 |
| Writing & Preference | 39.7 | #245 | 3 |
Strengths and weaknesses
Categories where Llama 3.1 Tulu 3 8b places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multilingual | 35.4 | −12.0 | #246 of 297, top 83% |
| Long Context | 35.8 | −5.2 | #239 of 296, top 81% |
| Instruction Following | 61.3 | −9.9 | #246 of 305, top 81% |
Closest competitors
The models ranked just above and below Llama 3.1 Tulu 3 8b. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Deepseek Coder v2 | #220 | 35.9 | — | — | Compare |
| C4ai Aya Expanse 32b | #221 | 35.9 | — | — | Compare |
| Llama 3.1 Nemotron 51b Instruct | #222 | 35.9 | — | — | Compare |
| Nemotron 4 340b Instruct | #223 | 35.9 | — | — | Compare |
| Qwen3 14B | #225 | 35.5 | $0.61 | 79 | Compare |
| DeepSeek-R1-Distill-Qwen-32B | #226 | 35.5 | — | — | Compare |
| Magistral Medium | #227 | 35.2 | $2.75 | 0 | Compare |
| Gemini 2.0 Flash (Feb 2025) | #228 | 35.1 | — | 92 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Coding | 1183 | #248 of 294, top 85% | LMArena | 2026-10-08 |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Hard Prompts | 1174 | #246 of 297, top 83% | LMArena | 2026-10-08 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Math | 1195 | #233 of 285, top 82% | LMArena | 2026-10-08 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1169 | #246 of 297, top 83% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1176 | #243 of 285, top 86% | LMArena | 2026-10-08 | |
| LMArena Russian | 1193 | #238 of 283, top 85% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1174 | #245 of 298, top 83% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1181 | #248 of 291, top 86% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1193 | #245 of 297, top 83% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1182 | #240 of 295, top 82% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1154 | #251 of 295, top 86% | LMArena | 2026-10-08 |
Compare Llama 3.1 Tulu 3 8b
- Llama 3.1 Tulu 3 8b vs Nemotron 4 340b Instruct
- Llama 3.1 Tulu 3 8b vs Qwen3 14B
- Llama 3.1 Tulu 3 8b vs Llama 3.1 Nemotron 51b Instruct
- Llama 3.1 Tulu 3 8b vs DeepSeek-R1-Distill-Qwen-32B
- Llama 3.1 Tulu 3 8b vs C4ai Aya Expanse 32b
- Llama 3.1 Tulu 3 8b vs Magistral Medium
- Llama 3.1 Tulu 3 8b vs GPT-6 Astra
- Llama 3.1 Tulu 3 8b vs Claude Fable 5.1
- Llama 3.1 Tulu 3 8b vs Gemini 3.8 Flash
- Llama 3.1 Tulu 3 8b vs Kimi K3
- Llama 3.1 Tulu 3 8b vs Grok 4.6
- Llama 3.1 Tulu 3 8b vs Qwen3.8 Max
- Llama 3.1 Tulu 3 8b vs GLM-5.3
- Llama 3.1 Tulu 3 8b vs Muse Spark 1.3
Other Allen Institute for AI (Ai2) models
Frequently asked questions
How good is Llama 3.1 Tulu 3 8b?
Llama 3.1 Tulu 3 8b by Allen Institute for AI (Ai2) ranks 224th of 354 ranked models on the Noometry Index as of October 2026, with a score of 35.7. Its strongest category is reasoning, where it ranks 188th.
Is Llama 3.1 Tulu 3 8b open source?
Yes. Llama 3.1 Tulu 3 8b's weights are downloadable; check the license for commercial terms.
What are Llama 3.1 Tulu 3 8b's strengths and weaknesses?
Relative to other ranked models, Llama 3.1 Tulu 3 8b places best in reasoning, math, coding and lowest in multilingual, long context, instruction following.
What is Llama 3.1 Tulu 3 8b best at?
Its best category is reasoning, where it ranks 188th on Noometry.