Reasoning benchmark
Adversarial NLI leaderboard
As of October 2026, GPT-3.5-turbo has the highest published Adversarial NLI score on Noometry at 58.1%, out of 9 models with results.
Last verified
About Adversarial NLI
Natural-language inference examples collected adversarially against earlier models.
- Category
- Reasoning
- Introduced
- 2019
- Format
- Classification
- Unit
- Percent (random guessing ≈ 33.3%)
- Official site
- github.com
Top 9 models
- GPT-3.5-turbo 58.1%
- Phi 3 Small 8k Instruct 58.1%
- Llama 3-8B 57.3%
- phi-3-medium 14B 55.8%
- Mixtral 8x7B 55.2%
- Phi 3 Mini 4k Instruct 52.8%
- Gemma 7B 48.7%
- Mistral 7B 47.1%
- Phi-2 42.5%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
| # | Model | Provider | Score | Setting | Source | Date |
|---|---|---|---|---|---|---|
| 1 | GPT-3.5-turbo | OpenAI | 58.1% | Epoch AI | ||
| 2 | Phi 3 Small 8k Instruct | 58.1% | Epoch AI | |||
| 3 | Llama 3-8B | 57.3% | Epoch AI | |||
| 4 | phi-3-medium 14B | 55.8% | Epoch AI | |||
| 5 | Mixtral 8x7B | 55.2% | Epoch AI | |||
| 6 | Phi 3 Mini 4k Instruct | 52.8% | Epoch AI | |||
| 7 | Gemma 7B | 48.7% | Epoch AI | |||
| 8 | Mistral 7B | 47.1% | Epoch AI | |||
| 9 | Phi-2 | 42.5% | Epoch AI |
Compare the leaders
Frequently asked questions
What does Adversarial NLI measure?
Natural-language inference examples collected adversarially against earlier models.
Which model has the highest Adversarial NLI score?
As of October 2026, GPT-3.5-turbo has the highest published Adversarial NLI score on Noometry at 58.1%, out of 9 models with results.
What is the best open-weight model on Adversarial NLI?
Phi 3 Small 8k Instruct has the highest Adversarial NLI accuracy among open-weight models at 58.1%, ranking 2 of 9 overall.