Reasoning benchmark

Adversarial NLI leaderboard

As of October 2026, GPT-3.5-turbo has the highest published Adversarial NLI score on Noometry at 58.1%, out of 9 models with results.

Last verified

About Adversarial NLI

Natural-language inference examples collected adversarially against earlier models.

Category
Reasoning
Introduced
2019
Format
Classification
Unit
Percent (random guessing ≈ 33.3%)
Official site
github.com

Top 9 models

Top models on Adversarial NLI
  1. GPT-3.5-turbo 58.1%
  2. Phi 3 Small 8k Instruct 58.1%
  3. Llama 3-8B 57.3%
  4. phi-3-medium 14B 55.8%
  5. Mixtral 8x7B 55.2%
  6. Phi 3 Mini 4k Instruct 52.8%
  7. Gemma 7B 48.7%
  8. Mistral 7B 47.1%
  9. Phi-2 42.5%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

Compare the leaders

Other reasoning benchmarks

Frequently asked questions

What does Adversarial NLI measure?

Natural-language inference examples collected adversarially against earlier models.

Which model has the highest Adversarial NLI score?

As of October 2026, GPT-3.5-turbo has the highest published Adversarial NLI score on Noometry at 58.1%, out of 9 models with results.

What is the best open-weight model on Adversarial NLI?

Phi 3 Small 8k Instruct has the highest Adversarial NLI accuracy among open-weight models at 58.1%, ranking 2 of 9 overall.