Allen Institute for AI (Ai2), open weights
Tulu 3 (Tülu 3) 70B
Tulu 3 (Tülu 3) 70B by Allen Institute for AI (Ai2) ranks 251st of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.0. Its strongest category is reasoning, where it ranks 169th.
Last verified
Specifications
- Noometry rank
- #251 of 354
- Index score
- 33.0
- Evidence
- Confirmed 14 results
- Provider
Allen Institute for AI (Ai2)
- Released
- November 21, 2024
- Weights
- Open weights
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- Not measured
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Coding 36.0
- Reasoning 23.9
- Math 14.2
- Knowledge 25.0
- Multilingual 39.9
- Instruction Following 64.8
- Long Context 37.1
- Writing & Preference 45.6
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 36.0 | #214 | 1 |
| Reasoning | 23.9 | #169 | 1 |
| Math | 14.2 | #303 | 3 |
| Knowledge | 25.0 | #264 | 1 |
| Multilingual | 39.9 | #222 | 1 |
| Instruction Following | 64.8 | #227 | 1 |
| Long Context | 37.1 | #222 | 1 |
| Writing & Preference | 45.6 | #223 | 3 |
Strengths and weaknesses
Categories where Tulu 3 (Tülu 3) 70B places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Reasoning | 23.9 | +0.3 | #169 of 350, top 49% |
| Coding | 36.0 | −2.7 | #214 of 340, top 63% |
| Writing & Preference | 45.6 | −8.2 | #223 of 312, top 72% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Math | 14.2 | −22.4 | #303 of 327, top 93% |
| Knowledge | 25.0 | −12.3 | #264 of 314, top 85% |
| Long Context | 37.1 | −3.8 | #222 of 296, top 75% |
Closest competitors
The models ranked just above and below Tulu 3 (Tülu 3) 70B. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Granite 3.1 2b Instruct | #247 | 33.2 | — | — | Compare |
| Gemma 2 2b IT | #248 | 33.1 | — | — | Compare |
| Wizardlm 70b | #249 | 33.0 | — | — | Compare |
| Phi 3 Medium 4k Instruct | #250 | 33.0 | — | — | Compare |
| DeepSeek-R1-Distill-Qwen-14B | #252 | 32.7 | — | — | Compare |
| Qwen1.5-14B | #253 | 32.7 | — | — | Compare |
| Olmo 2 0325 32b Instruct | #254 | 32.7 | — | — | Compare |
| gpt-oss-20b | #255 | 32.5 | $0.036 | 96 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Coding | 1235 | #229 of 294, top 78% | LMArena | 2026-10-08 |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Hard Prompts | 1220 | #229 of 297, top 78% | LMArena | 2026-10-08 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 4.4% | #150 of 173, top 87% | Epoch AI | 2025-03-07 | |
| LMArena Math | 1242 | #218 of 285, top 77% | LMArena | 2026-10-08 | |
| MATH Level 5 | 42.7% | #50 of 79, top 64% | Epoch AI | 2025-01-27 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 46.3% | #140 of 186, top 76% | Epoch AI | 2025-01-27 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1236 | #222 of 297, top 75% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1249 | #208 of 285, top 73% | LMArena | 2026-10-08 | |
| LMArena Russian | 1246 | #215 of 283, top 76% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1233 | #226 of 298, top 76% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1224 | #233 of 291, top 81% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1256 | #224 of 297, top 76% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1231 | #221 of 295, top 75% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1252 | #224 of 295, top 76% | LMArena | 2026-10-08 |
Compare Tulu 3 (Tülu 3) 70B
- Tulu 3 (Tülu 3) 70B vs Phi 3 Medium 4k Instruct
- Tulu 3 (Tülu 3) 70B vs DeepSeek-R1-Distill-Qwen-14B
- Tulu 3 (Tülu 3) 70B vs Wizardlm 70b
- Tulu 3 (Tülu 3) 70B vs Qwen1.5-14B
- Tulu 3 (Tülu 3) 70B vs Gemma 2 2b IT
- Tulu 3 (Tülu 3) 70B vs Olmo 2 0325 32b Instruct
- Tulu 3 (Tülu 3) 70B vs GPT-6 Astra
- Tulu 3 (Tülu 3) 70B vs Claude Fable 5.1
- Tulu 3 (Tülu 3) 70B vs Gemini 3.8 Flash
- Tulu 3 (Tülu 3) 70B vs Kimi K3
- Tulu 3 (Tülu 3) 70B vs Grok 4.6
- Tulu 3 (Tülu 3) 70B vs Qwen3.8 Max
- Tulu 3 (Tülu 3) 70B vs GLM-5.3
- Tulu 3 (Tülu 3) 70B vs Muse Spark 1.3
Other Allen Institute for AI (Ai2) models
Frequently asked questions
How good is Tulu 3 (Tülu 3) 70B?
Tulu 3 (Tülu 3) 70B by Allen Institute for AI (Ai2) ranks 251st of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.0. Its strongest category is reasoning, where it ranks 169th.
Is Tulu 3 (Tülu 3) 70B open source?
Yes. Tulu 3 (Tülu 3) 70B's weights are downloadable; check the license for commercial terms.
What are Tulu 3 (Tülu 3) 70B's strengths and weaknesses?
Relative to other ranked models, Tulu 3 (Tülu 3) 70B places best in reasoning, coding, writing & preference and lowest in math, knowledge, long context.
What is Tulu 3 (Tülu 3) 70B best at?
Its best category is reasoning, where it ranks 169th on Noometry.