Best model for your job
Best AI model for RAG
For RAG, GPT-5 scores highest on the current data (66.9), weighting long context 40%, knowledge 30%, instruction following 30%. The cheapest model in the top 10 is Gemini 3.8 Flash at $0.75 / $3.75 per million tokens.
Last verified
Weights: Long Context 40%, Knowledge 30%, Instruction Following 30%. Retrieval-augmented generation depends on reading long retrieved context faithfully and citing it.
| # | Model | Provider | Fit | Long Context | Knowledge | Instruction Following | Input $/M | Output $/M |
|---|---|---|---|---|---|---|---|---|
| 1 | GPT-5 | OpenAI | 66.9 | 69.5 | 56.6 | 73.8 | $1.25 | $10 |
| 2 | Grok 4 | xAI | 65.1 | 63.1 | 53.8 | 79.2 | — | — |
| 3 | Gemini 3.8 Flash | 64.4 | 46.3 | 74.8 | 78.0 | $0.75 | $3.75 | |
| 4 | Gemini 3.1 Pro Preview | 63.6 | 47.4 | 71.8 | 77.0 | $2 | $12 | |
| 5 | Claude Fable 5.1 | Anthropic | 63.3 | 46.7 | 69.6 | 79.2 | $10 | $50 |
| 6 | GPT-6 Astra | OpenAI | 63.3 | 44.5 | 75.3 | 76.3 | $10 | $50 |
| 7 | Gemini 2.5 Pro | 63.2 | 59.8 | 56.0 | 75.0 | $1.25 | $10 | |
| 8 | GPT-5.4 | OpenAI | 62.9 | 50.3 | 65.3 | 77.1 | $2.50 | $15 |
| 9 | Claude Opus 5.5 | Anthropic | 62.8 | 47.1 | 66.4 | 80.0 | $4 | $20 |
| 10 | GPT-6.1 Sol | OpenAI | 62.6 | 44.9 | 71.8 | 77.0 | $2 | $10 |
| 11 | Gemini 3.7 Flash | 62.5 | 45.7 | 69.7 | 77.7 | $0.75 | $3.75 | |
| 12 | Claude Opus 5 | Anthropic | 62.4 | 46.5 | 66.8 | 79.2 | $5 | $25 |
| 13 | GPT-5.5 | OpenAI | 61.9 | 48.3 | 64.4 | 77.5 | $5 | $30 |
| 14 | Claude Opus 4.6 | Anthropic | 61.6 | 48.1 | 61.9 | 79.5 | $5 | $25 |
| 15 | Claude Sonnet 5.5 | Anthropic | 61.6 | 45.9 | 66.0 | 78.3 | $2 | $10 |
| 16 | Gemini 3.6 Flash | 61.5 | 45.1 | 67.8 | 77.0 | $0.75 | $3.75 | |
| 17 | Gemini 3.5 Flash | 61.2 | 45.4 | 66.3 | 77.0 | $1.50 | $9 | |
| 18 | Claude Opus 4.7 | Anthropic | 60.8 | 46.2 | 62.6 | 78.4 | $5 | $25 |
| 19 | Claude Fable 5 | Anthropic | 60.8 | 46.3 | 62.2 | 78.6 | $10 | $50 |
| 20 | GPT-5.6 Sol | OpenAI | 60.7 | 45.4 | 64.3 | 77.7 | $4 | $20 |
| 21 | Kimi K3 (open weights) | Moonshot AI | 60.6 | 45.8 | 63.2 | 77.7 | $3 | $15 |
| 22 | Muse Spark | Meta | 60.2 | 44.4 | 65.7 | 75.9 | — | — |
| 23 | Qwen3.8 Max | Alibaba (Qwen) | 60.0 | 45.6 | 61.7 | 77.6 | $2 | $6 |
| 24 | Gemini 3 Pro | 59.8 | 44.0 | 64.4 | 76.3 | — | — | |
| 25 | Claude Opus 4.8 | Anthropic | 59.8 | 45.4 | 61.3 | 77.4 | $5 | $25 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Top two head to head: GPT-5 vs Grok 4
Frequently asked questions
What is the best ai model for rag?
For RAG, GPT-5 scores highest on the current data (66.9), weighting long context 40%, knowledge 30%, instruction following 30%. The cheapest model in the top 10 is Gemini 3.8 Flash at $0.75 / $3.75 per million tokens.
How is this shortlist built?
Retrieval-augmented generation depends on reading long retrieved context faithfully and citing it. Each model's category scores are blended with those weights; only ranked models with results in every needed category are listed.
Other jobs
- Best AI model for coding
- Best AI model for debugging
- Best AI model for building websites
- Best local LLM for coding
- Best AI model for agents
- Best AI model for research
- Best AI model for studying
- Best AI model for math
- Best AI model for hard reasoning
- Best AI model for data analysis
- Best AI model for Excel
- Best AI model for data extraction
- Best AI model for writing
- Best AI model for emails
- Best AI model for copywriting
- Best AI model for creative writing
- Best AI model for translation
- Best AI model for language learning
- Best AI model for customer service
- Best AI model for presentations
- Best AI model for PDFs
- Best AI assistant for everyday use