OpenAI, proprietary
GPT-3.5-turbo
GPT-3.5-turbo by OpenAI ranks 350th of 354 ranked models on the Noometry Index as of October 2026, with a score of 23.2. Its strongest category is long context, where it ranks 254th. API pricing starts at $0.50 per million input tokens and $1.50 per million output tokens, with a 16K-token context window.
Last verified
Specifications
- Noometry rank
- #350 of 354
- Index score
- 23.2
- Evidence
- Confirmed 44 results
- Provider
- OpenAI
- Released
- March 1, 2023
- Weights
- Proprietary
- Reasoning
- No
- Context window
- 16K
- Max output
- 4K
- Input price
- $0.50 / M
- Output price
- $1.50 / M
- Blended price
- $0.75 / M
- Output speed
- Not measured
- Value
- #133 of 219
- Knowledge cutoff
- September 2021
- Input
- text
Category scores
Each category score combines every public result we have in that category.
- Coding 23.9
- Reasoning 13.8
- Math 6.3
- Knowledge 10.0
- Multilingual 31.5
- Instruction Following 57.9
- Long Context 34.0
- Writing & Preference 25.3
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 23.9 | #331 | 4 |
| Reasoning | 13.8 | #332 | 5 |
| Math | 6.3 | #327 | 4 |
| Knowledge | 10.0 | #303 | 2 |
| Multilingual | 31.5 | #258 | 1 |
| Instruction Following | 57.9 | #262 | 1 |
| Long Context | 34.0 | #254 | 1 |
| Writing & Preference | 25.3 | #305 | 4 |
Strengths and weaknesses
Categories where GPT-3.5-turbo places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 34.0 | −7.0 | #254 of 296, top 86% |
| Instruction Following | 57.9 | −13.4 | #262 of 305, top 86% |
| Multilingual | 31.5 | −15.9 | #258 of 297, top 87% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Math | 6.3 | −30.3 | #327 of 327, top 100% |
| Writing & Preference | 25.3 | −28.5 | #305 of 312, top 98% |
| Coding | 23.9 | −14.8 | #331 of 340, top 98% |
Closest competitors
The models ranked just above and below GPT-3.5-turbo. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Claude 2 | #346 | 25.0 | — | — | Compare |
| DeepSeek LLM 67B | #347 | 24.9 | — | — | Compare |
| Llama 13b | #348 | 24.4 | — | — | Compare |
| Llama 2-70B | #349 | 24.4 | — | — | Compare |
| Mistral 7B | #351 | 23.0 | $0.25 | — | Compare |
| Llama 3.1-8B | #352 | 23.0 | $0.0575 | — | Compare |
| Gemma 3 1B | #353 | 21.1 | — | — | Compare |
| Llama 3.2 1B | #354 | 20.1 | $0.0705 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| WeirdML | 3.5% | #117 of 119, top 99% | Epoch AI | ||
| BigCodeBench Instruct | 39.1% | #35 of 64, top 55% | BigCodeBench | 2024-01-25 | |
| LMArena Coding | 1136 | #260 of 294, top 89% | LMArena | 2026-10-08 | |
| LMArena Coding | 1116 | LMArena | 2026-10-08 | ||
| BigCodeBench Complete | 50.6% | #31 of 66, top 47% | BigCodeBench | 2024-01-25 | |
| HumanEval+ | 66.5% | may 2023 | EvalPlus | ||
| HumanEval+ | 70.7% | #20 of 45, top 45% | nov 2023 | EvalPlus | |
| MBPP+ | 69.7% | #14 of 38, top 37% | nov 2023 | EvalPlus |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| METR Time Horizons | 21.5% | #32 of 32, top 100% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Chess Puzzles | 0% | #115 of 129, top 90% | Epoch AI | 2026-08-07 | |
| LMArena Hard Prompts | 1099 | LMArena | 2026-10-08 | ||
| LMArena Hard Prompts | 1108 | #266 of 297, top 90% | LMArena | 2026-10-08 | |
| Mystery Game Puzzles | 3% | #73 of 74, top 99% | Epoch AI | 2026-08-27 | |
| DTBench | 48.5% | #141 of 151, top 94% | Epoch AI | ||
| LMCA | 9.7% | #113 of 125, top 91% | Epoch AI | ||
| Adversarial NLI | 58.1% | Best of 9 | Epoch AI | ||
| BIG-Bench Hard | 61.6% | #12 of 27, top 45% | Epoch AI | ||
| CommonsenseQA 2.0 | 57% | Best of 2 | Epoch AI | ||
| Epoch Capabilities Index | 115.7 | Epoch AI | 2024-01-25 | ||
| Epoch Capabilities Index | 113.36 | Epoch AI | 2023-06-13 | ||
| Epoch Capabilities Index | 118.55 | #174 of 213, top 82% | Epoch AI | 2023-11-06 | |
| ForecastBench | 50.4 | #72 of 72, top 100% | Epoch AI | ||
| WinoGrande | 81.6% | #11 of 43, top 26% | Epoch AI | ||
| WinoGrande | 68.8% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 0% | #81 of 81, top 100% | Epoch AI | 2026-08-27 | |
| OTIS Mock AIME 2024-2025 | 2.2% | #160 of 173, top 93% | Epoch AI | 2026-08-07 | |
| LMArena Math | 1141 | LMArena | 2026-10-08 | ||
| LMArena Math | 1142 | #256 of 285, top 90% | LMArena | 2026-10-08 | |
| MATH Level 5 | 11.6% | Epoch AI | 2025-01-27 | ||
| MATH Level 5 | 15.9% | #66 of 79, top 84% | Epoch AI | 2025-01-27 | |
| GSM8K | 57.8% | #17 of 38, top 45% | Epoch AI |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 28% | #173 of 186, top 94% | Epoch AI | 2025-01-27 | |
| GPQA Diamond | 27.2% | Epoch AI | 2025-01-27 | ||
| LMArena Expert | 1070 | #256 of 273, top 94% | LMArena | 2026-10-08 | |
| LMArena Expert | 1066 | LMArena | 2026-10-08 | ||
| ARC (AI2) Challenge | 87.4% | #7 of 39, top 18% | Epoch AI | ||
| BoolQ | 87% | #5 of 23, top 22% | Epoch AI | ||
| MMLU | 71.4% | #41 of 81, top 51% | Epoch AI | ||
| MMLU | 68.9% | Epoch AI | |||
| MMLU | 67.3% | Epoch AI | |||
| OpenBookQA | 86% | #4 of 19, top 22% | Epoch AI | ||
| TriviaQA | 85.8% | #4 of 25, top 16% | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1074 | LMArena | 2026-10-08 | ||
| LMArena Non-English | 1108 | #258 of 297, top 87% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1075 | #260 of 285, top 92% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1012 | LMArena | 2026-10-08 | ||
| LMArena French | 1066 | LMArena | 2026-10-08 | ||
| LMArena French | 1118 | #210 of 223, top 95% | LMArena | 2026-10-08 | |
| LMArena German | 1090 | #212 of 231, top 92% | LMArena | 2026-10-08 | |
| LMArena German | 1056 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1043 | #191 of 211, top 91% | LMArena | 2026-10-08 | |
| LMArena Korean | 1019 | #197 of 213, top 93% | LMArena | 2026-10-08 | |
| LMArena Russian | 1077 | LMArena | 2026-10-08 | ||
| LMArena Russian | 1123 | #252 of 283, top 90% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1093 | LMArena | 2026-10-08 | ||
| LMArena Spanish | 1121 | #211 of 226, top 94% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1119 | #259 of 298, top 87% | LMArena | 2026-10-08 | |
| LMArena Instruction Following | 1093 | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1070 | LMArena | 2026-10-08 | ||
| LMArena Longer Query | 1121 | #262 of 291, top 91% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1125 | #266 of 297, top 90% | LMArena | 2026-10-08 | |
| LMArena Text | 1094 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1092 | #266 of 295, top 91% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1035 | LMArena | 2026-10-08 | ||
| EQ-Bench Creative Writing | 451 | #114 of 115, top 100% | EQ-Bench | ||
| LMArena Multi-Turn | 1117 | #259 of 295, top 88% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1078 | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $0.50 | $1.50 | — | 2026-10-10 |
| openai | $0.50 | $1.50 | Free | 2026-10-10 |
| openrouter | $0.50 | $1.50 | — | 2026-10-10 |
Compare GPT-3.5-turbo
- GPT-3.5-turbo vs Llama 2-70B
- GPT-3.5-turbo vs Mistral 7B
- GPT-3.5-turbo vs Llama 13b
- GPT-3.5-turbo vs Llama 3.1-8B
- GPT-3.5-turbo vs DeepSeek LLM 67B
- GPT-3.5-turbo vs Gemma 3 1B
- GPT-3.5-turbo vs Claude Fable 5.1
- GPT-3.5-turbo vs Gemini 3.8 Flash
- GPT-3.5-turbo vs Kimi K3
- GPT-3.5-turbo vs Grok 4.6
- GPT-3.5-turbo vs Qwen3.8 Max
- GPT-3.5-turbo vs GLM-5.3
- GPT-3.5-turbo vs Muse Spark 1.3
- GPT-3.5-turbo vs DeepSeek V4 Pro
Other OpenAI models
- GPT-6 Astra70.8
- GPT-6.1 Sol65.6
- GPT-5.6 Sol65.0
- GPT-5.5 Pro64.3
- GPT-5.563.4
- GPT-6 Sol61.8
- GPT-5.459.4
- GPT-5.6 Terra59.2
Frequently asked questions
How good is GPT-3.5-turbo?
GPT-3.5-turbo by OpenAI ranks 350th of 354 ranked models on the Noometry Index as of October 2026, with a score of 23.2. Its strongest category is long context, where it ranks 254th. API pricing starts at $0.50 per million input tokens and $1.50 per million output tokens, with a 16K-token context window.
How much does GPT-3.5-turbo cost?
GPT-3.5-turbo costs $0.50 per million input tokens and $1.50 per million output tokens on OpenAI's own API.
What is GPT-3.5-turbo's context window?
GPT-3.5-turbo accepts up to 16K tokens of input and can write up to 4K tokens in one response.
Is GPT-3.5-turbo open source?
No. GPT-3.5-turbo is proprietary and available only through OpenAI's API and partner platforms.
What are GPT-3.5-turbo's strengths and weaknesses?
Relative to other ranked models, GPT-3.5-turbo places best in long context, instruction following, multilingual and lowest in math, writing & preference, coding.
What is GPT-3.5-turbo best at?
Its best category is long context, where it ranks 254th on Noometry.