Agentic & Tool Use benchmark
Vending-Bench 2 leaderboard
As of October 2026, GPT-6 Astra has the highest published Vending-Bench 2 score on Noometry at 15,515, out of 60 models with results.
Last verified
About Vending-Bench 2
Running a simulated vending-machine business over a long horizon. The score is the final bank balance.
- Introduced
- 2025
- Format
- Long-horizon simulation
- Unit
- Raw score
Top 15 models
Top models on Vending-Bench 2 - GPT-6 Astra 15,515
- GPT-6 Sol 14,428
- Gemini 4 Argon 13,718
- Claude Opus 5 11,182
- Claude Opus 4.7 10,937
- Grok 4.7 10,537
- GPT-5.6 Sol 9,619
- Claude Opus 5.5 9,235
- Grok 4.6 9,047
- GLM-5.2 8,314
- GLM-5.3 8,164
- Claude Opus 4.6 8,018
- GPT-5.5 7,524
- GPT-5.6 Terra 7,343
- Claude Sonnet 4.6 7,204
- 05000100001500020000
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Vending-Bench 2 results by model| # | Model | Provider | Rating | Setting | Source | Date |
|---|
| 1 | GPT-6 Astra | OpenAI | 15,515 | | Epoch AI | |
| 2 | GPT-6 Sol | OpenAI | 14,428 | | Epoch AI | |
| 3 | Gemini 4 Argon | Google | 13,718 | | Epoch AI | |
| 4 | Claude Opus 5 | Anthropic | 11,182 | | Epoch AI | |
| 5 | Claude Opus 4.7 | Anthropic | 10,937 | | Epoch AI | |
| 6 | Grok 4.7 | xAI | 10,537 | | Epoch AI | |
| 7 | GPT-5.6 Sol | OpenAI | 9,619 | | Epoch AI | |
| 8 | Claude Opus 5.5 | Anthropic | 9,235 | | Epoch AI | |
| 9 | Grok 4.6 | xAI | 9,047 | | Epoch AI | |
| 10 | GLM-5.2 | Z.ai (Zhipu) | 8,314 | | Epoch AI | |
| 11 | GLM-5.3 | Z.ai (Zhipu) | 8,164 | | Epoch AI | |
| 12 | Claude Opus 4.6 | Anthropic | 8,018 | | Epoch AI | |
| 13 | GPT-5.5 | OpenAI | 7,524 | | Epoch AI | |
| 14 | GPT-5.6 Terra | OpenAI | 7,343 | | Epoch AI | |
| 15 | Claude Sonnet 4.6 | Anthropic | 7,204 | | Epoch AI | |
| 16 | Muse Spark 1.1 | Meta | 6,520 | | Epoch AI | |
| 17 | Claude Sonnet 5 | Anthropic | 6,378 | | Epoch AI | |
| 18 | Kimi K2.6 | Moonshot AI | 6,205 | | Epoch AI | |
| 19 | GPT-5.4 | OpenAI | 6,144 | | Epoch AI | |
| 20 | GPT-5.3 Codex | OpenAI | 5,940 | | Epoch AI | |
| 21 | Claude Opus 4.8 | Anthropic | 5,787 | | Epoch AI | |
| 22 | Claude Fable 5 | Anthropic | 5,680 | high | Epoch AI | |
| 23 | GLM-5.1 | Z.ai (Zhipu) | 5,634 | | Epoch AI | |
| 24 | Gemini 3 Pro | Google | 5,478 | | Epoch AI | |
| 25 | Claude Fable 5.1 | Anthropic | 5,422 | | Epoch AI | |
| 26 | Gemini 3.5 Flash | Google | 5,396 | | Epoch AI | |
| 27 | Kimi K3 | Moonshot AI | 5,165 | | Epoch AI | |
| 28 | Qwen3.6 Plus | Alibaba (Qwen) | 5,115 | | Epoch AI | |
| 29 | Gemini 3.8 Flash | Google | 5,094 | | Epoch AI | |
| 30 | Kimi K2.7 Code | Moonshot AI | 5,083 | | Epoch AI | |
| 31 | Claude Opus 4.5 | Anthropic | 4,967 | | Epoch AI | |
| 32 | Grok 4.20 (Non-Reasoning) | xAI | 4,663 | | Epoch AI | |
| 33 | GLM-5 | Z.ai (Zhipu) | 4,432 | | Epoch AI | |
| 34 | Qwen3.6 Max Preview | Alibaba (Qwen) | 4,254 | | Epoch AI | |
| 35 | GPT-5.6 Luna | OpenAI | 4,095 | | Epoch AI | |
| 36 | Grok 4.5 | xAI | 3,887 | | Epoch AI | |
| 37 | Claude Sonnet 4.5 | Anthropic | 3,839 | | Epoch AI | |
| 38 | Gemini 3.1 Pro Preview | Google | 3,774 | | Epoch AI | |
| 39 | Gemini 3 Flash Preview | Google | 3,635 | | Epoch AI | |
| 40 | GPT-5.2 | OpenAI | 3,591 | | Epoch AI | |
| 41 | DeepSeek V4 Pro | DeepSeek | 3,285 | | Epoch AI | |
| 42 | GLM-4.7 | Z.ai (Zhipu) | 2,377 | | Epoch AI | |
| 43 | MiniMax-M3 | MiniMax | 2,158 | | Epoch AI | |
| 44 | GPT-5.1 | OpenAI | 1,473 | | Epoch AI | |
| 45 | Kimi K2.5 | Moonshot AI | 1,198 | | Epoch AI | |
| 46 | Grok 4.1 Fast | xAI | 1,107 | | Epoch AI | |
| 47 | DeepSeek-V3.2-Exp | DeepSeek | 1,034 | | Epoch AI | |
| 48 | Gemini 2.5 Pro | Google | 573.64 | | Epoch AI | |
| 49 | Gemini 2.5 Flash | Google | 548.84 | | Epoch AI | |
| 50 | Qwen3.5-Flash | Alibaba (Qwen) | 462.69 | | Epoch AI | |
| 51 | Claude Haiku 4.5 | Anthropic | 458.89 | | Epoch AI | |
| 52 | Qwen3.5 27B | Alibaba (Qwen) | 201.98 | | Epoch AI | |
| 53 | MiniMax-M2 | MiniMax | 160.6 | | Epoch AI | |
| 54 | Qwen3 Max | Alibaba (Qwen) | 71.56 | | Epoch AI | |
| 55 | Grok 4.3 | xAI | 35.26 | | Epoch AI | |
| 56 | Qwen3.5 Plus | Alibaba (Qwen) | 0.54 | | Epoch AI | |
| 57 | Qwen3 235B-A22B | Alibaba (Qwen) | -11.34 | | Epoch AI | |
| 58 | gpt-oss-120b | OpenAI | -21.53 | | Epoch AI | |
| 59 | MiniMax-M2.5 | MiniMax | -23.16 | | Epoch AI | |
| 60 | GPT-5 Mini | OpenAI | -31.18 | | Epoch AI | |
Other agentic & tool use benchmarks
Frequently asked questions
What does Vending-Bench 2 measure?
Running a simulated vending-machine business over a long horizon. The score is the final bank balance.
Which model has the highest Vending-Bench 2 score?
As of October 2026, GPT-6 Astra has the highest published Vending-Bench 2 score on Noometry at 15,515, out of 60 models with results.
What is the best open-weight model on Vending-Bench 2?
GLM-5.2 has the highest Vending-Bench 2 score among open-weight models at 8,314, ranking 10 of 60 overall.