Best model for your job

Best AI model for RAG

For RAG, GPT-5 scores highest on the current data (66.9), weighting long context 40%, knowledge 30%, instruction following 30%. The cheapest model in the top 10 is Gemini 3.8 Flash at $0.75 / $3.75 per million tokens.

Last verified

Weights: Long Context 40%, Knowledge 30%, Instruction Following 30%. Retrieval-augmented generation depends on reading long retrieved context faithfully and citing it.

Best AI model for RAG
#ModelProviderFitLong ContextKnowledgeInstruction FollowingInput $/MOutput $/M
1GPT-5OpenAI66.969.556.673.8$1.25$10
2Grok 4xAI65.163.153.879.2——
3Gemini 3.8 FlashGoogle64.446.374.878.0$0.75$3.75
4Gemini 3.1 Pro PreviewGoogle63.647.471.877.0$2$12
5Claude Fable 5.1Anthropic63.346.769.679.2$10$50
6GPT-6 AstraOpenAI63.344.575.376.3$10$50
7Gemini 2.5 ProGoogle63.259.856.075.0$1.25$10
8GPT-5.4OpenAI62.950.365.377.1$2.50$15
9Claude Opus 5.5Anthropic62.847.166.480.0$4$20
10GPT-6.1 SolOpenAI62.644.971.877.0$2$10
11Gemini 3.7 FlashGoogle62.545.769.777.7$0.75$3.75
12Claude Opus 5Anthropic62.446.566.879.2$5$25
13GPT-5.5OpenAI61.948.364.477.5$5$30
14Claude Opus 4.6Anthropic61.648.161.979.5$5$25
15Claude Sonnet 5.5Anthropic61.645.966.078.3$2$10
16Gemini 3.6 FlashGoogle61.545.167.877.0$0.75$3.75
17Gemini 3.5 FlashGoogle61.245.466.377.0$1.50$9
18Claude Opus 4.7Anthropic60.846.262.678.4$5$25
19Claude Fable 5Anthropic60.846.362.278.6$10$50
20GPT-5.6 SolOpenAI60.745.464.377.7$4$20
21Kimi K3 (open weights)Moonshot AI60.645.863.277.7$3$15
22Muse SparkMeta60.244.465.775.9——
23Qwen3.8 MaxAlibaba (Qwen)60.045.661.777.6$2$6
24Gemini 3 ProGoogle59.844.064.476.3——
25Claude Opus 4.8Anthropic59.845.461.377.4$5$25

Sponsored placements are available on pages like this one. Advertise on Noometry

Top two head to head: GPT-5 vs Grok 4

Frequently asked questions

What is the best ai model for rag?

For RAG, GPT-5 scores highest on the current data (66.9), weighting long context 40%, knowledge 30%, instruction following 30%. The cheapest model in the top 10 is Gemini 3.8 Flash at $0.75 / $3.75 per million tokens.

How is this shortlist built?

Retrieval-augmented generation depends on reading long retrieved context faithfully and citing it. Each model's category scores are blended with those weights; only ranked models with results in every needed category are listed.

Other jobs