xAI, proprietary

# Grok 4

> Grok 4 by xAI, released July 2025. Ranked #56 of 354 with a Noometry Index of 48.1. Scores, sources and comparisons.
- Canonical page: https://noometry.com/models/grok-4
- Last updated: 2026-10-10
- Title: Grok 4 Benchmarks, Price & Rank (October 2026) | Noometry

Grok 4 by xAI ranks 56th of 354 ranked models on the Noometry Index as of October 2026, with a score of 48.1. Its strongest category is long context, where it ranks 4th.

Last verified October 10, 2026

## Specifications

- **Noometry rank:** #56 of 354
- **Index score:** 48.1
- **Evidence:** Confirmed 48 results
- **Provider:** [xAI](https://noometry.com/providers/xai)
- **Released:** July 9, 2025
- **Weights:** Proprietary
- **Reasoning:** Unknown
- **Context window:** —
- **Max output:** —
- **Input price:** Not listed
- **Output price:** Not listed
- **Blended price:** Not listed
- **Output speed:** 1 tokens/s [Kagi](https://help.kagi.com/kagi/ai/llm-benchmark.html)
- **Value:** Not ranked
- **Knowledge cutoff:** Unknown

## Category scores

Each category score combines every public result we have in that category.

Grok 4 category scores

1.  Coding 50.3
2.  Agentic & Tool Use 32.3
3.  Reasoning 36.7
4.  Math 48.4
5.  Knowledge 53.8
6.  Multimodal 33.7
7.  Multilingual 51.8
8.  Instruction Following 79.2
9.  Long Context 63.1
10.  Writing & Preference 58.5
11.  020406080

Grok 4 category ranks
| Category | Score | Rank | Results |
| --- | --- | --- | --- |
| [Coding](https://noometry.com/best/coding) | 50.3 | #46 | 3 |
| [Agentic & Tool Use](https://noometry.com/best/agentic) | 32.3 | #68 | 6 |
| [Reasoning](https://noometry.com/best/reasoning) | 36.7 | #65 | 6 |
| [Math](https://noometry.com/best/math) | 48.4 | #64 | 3 |
| [Knowledge](https://noometry.com/best/knowledge) | 53.8 | #55 | 5 |
| [Multimodal](https://noometry.com/best/multimodal) | 33.7 | #94 | 2 |
| [Multilingual](https://noometry.com/best/multilingual) | 51.8 | #103 | 1 |
| [Instruction Following](https://noometry.com/best/instruction-following) | 79.2 | #5 | 2 |
| [Long Context](https://noometry.com/best/long-context) | 63.1 | #4 | 2 |
| [Writing & Preference](https://noometry.com/best/writing) | 58.5 | #116 | 5 |

## Strengths and weaknesses

Categories where Grok 4 places highest and lowest among the models ranked in each, with its score against that category's median.

### Strongest categories

Grok 4: strongest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Long Context](https://noometry.com/best/long-context) | 63.1 | +22.2 | #4 of 296, top 2% |
| [Instruction Following](https://noometry.com/best/instruction-following) | 79.2 | +8.0 | #5 of 305, top 2% |
| [Coding](https://noometry.com/best/coding) | 50.3 | +11.6 | #46 of 340, top 14% |

### Weakest categories

Grok 4: weakest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Multimodal](https://noometry.com/best/multimodal) | 33.7 | −4.8 | #94 of 128, top 74% |
| [Agentic & Tool Use](https://noometry.com/best/agentic) | 32.3 | +1.9 | #68 of 154, top 45% |
| [Writing & Preference](https://noometry.com/best/writing) | 58.5 | +4.7 | #116 of 312, top 38% |

## Closest competitors

The models ranked just above and below Grok 4. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Grok 4
| Model | Rank | Score | Blended $/M | Speed |  |
| --- | --- | --- | --- | --- | --- |
| [Claude Haiku 5.5](https://noometry.com/models/claude-haiku-5-5) | #52 | 49.5 | $0.20 | — | [Compare](https://noometry.com/compare/claude-haiku-5-5-vs-grok-4) |
| [GPT-5.1](https://noometry.com/models/gpt-5-1) | #53 | 49.0 | $3.44 | — | [Compare](https://noometry.com/compare/gpt-5-1-vs-grok-4) |
| [Grok 4.20 (Non-Reasoning)](https://noometry.com/models/grok-4-20) | #54 | 48.6 | $1.56 | 61 | [Compare](https://noometry.com/compare/grok-4-vs-grok-4-20) |
| [MiMo-V2.6-Flash](https://noometry.com/models/mimo-v2-6-flash) | #55 | 48.5 | $0.18 | — | [Compare](https://noometry.com/compare/grok-4-vs-mimo-v2-6-flash) |
| [Kimi K2.5](https://noometry.com/models/kimi-k2-5) | #57 | 48.1 | $0.90 | 66 | [Compare](https://noometry.com/compare/grok-4-vs-kimi-k2-5) |
| [Step 5 Preview](https://noometry.com/models/step-5-preview) | #58 | 47.9 | $1.43 | — | [Compare](https://noometry.com/compare/grok-4-vs-step-5-preview) |
| [GLM-5.1](https://noometry.com/models/glm-5-1) | #59 | 47.8 | $2.15 | — | [Compare](https://noometry.com/compare/glm-5-1-vs-grok-4) |
| [Kimi K2.6](https://noometry.com/models/kimi-k2-6) | #60 | 47.7 | $1.71 | — | [Compare](https://noometry.com/compare/grok-4-vs-kimi-k2-6) |

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

### Coding

Grok 4 Coding benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [Aider Polyglot](https://noometry.com/benchmarks/aider-polyglot) | 79.6% | #5 of 44, top 12% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Aider Polyglot](https://noometry.com/benchmarks/aider-polyglot) | 79.6% |  | high | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [WeirdML](https://noometry.com/benchmarks/weirdml) | 45.7% | #59 of 119, top 50% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [LMArena Coding](https://noometry.com/benchmarks/arena-coding) | 1408 | #130 of 294, top 45% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Agentic & Tool Use

Grok 4 Agentic & Tool Use benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [Terminal-Bench](https://noometry.com/benchmarks/terminal-bench) | 27.2% | #33 of 41, top 81% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Berkeley Function Calling Leaderboard](https://noometry.com/benchmarks/bfcl) | 63% | #8 of 49, top 17% | prompt | [Berkeley Function Calling Leaderboard](https://gorilla.cs.berkeley.edu/leaderboard.html) |  |
| [GDPval](https://noometry.com/benchmarks/gdpval) | 21.1% | #10 of 11, top 91% | high | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Cybench](https://noometry.com/benchmarks/cybench) | 43% | #4 of 21, top 20% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [DeepResearch Bench](https://noometry.com/benchmarks/deepresearch-bench) | 47.3% | #11 of 24, top 46% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [BALROG](https://noometry.com/benchmarks/balrog) | 43.6% | #9 of 35, top 26% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [LMArena Search](https://noometry.com/benchmarks/arena-search) | 1142 | #27 of 32, top 85% |  | [LMArena](https://lmarena.ai/leaderboard/search) | 2026-08-24 |
| [METR Time Horizons](https://noometry.com/benchmarks/metr-time-horizons) | 66.6% | #13 of 32, top 41% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Reasoning

Grok 4 Reasoning benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [ARC-AGI-2](https://noometry.com/benchmarks/arc-agi-2) | 16% | #45 of 83, top 55% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [SimpleBench](https://noometry.com/benchmarks/simplebench) | 60.5% | #25 of 77, top 33% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Kagi LLM Benchmark](https://noometry.com/benchmarks/kagi-reasoning) | 73.6% | #16 of 99, top 17% |  | [Kagi LLM Benchmark](https://help.kagi.com/kagi/ai/llm-benchmark.html) |  |
| [ARC-AGI-1](https://noometry.com/benchmarks/arc-agi-1) | 66.7% | #44 of 83, top 54% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Chess Puzzles](https://noometry.com/benchmarks/chess-puzzles) | 28% | #38 of 129, top 30% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-01-30 |
| [LMArena Hard Prompts](https://noometry.com/benchmarks/arena-hard-prompts) | 1409 | #119 of 297, top 41% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [Epoch Capabilities Index](https://noometry.com/benchmarks/epoch-capabilities-index) | 146.44 | #73 of 213, top 35% |  | [Epoch AI](https://epoch.ai/eci) | 2025-07-09 |
| [ForecastBench](https://noometry.com/benchmarks/forecastbench) | 60.9 | #21 of 72, top 30% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Math

Grok 4 Math benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [OTIS Mock AIME 2024-2025](https://noometry.com/benchmarks/otis-mock-aime) | 84% | #74 of 173, top 43% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Omni-MATH](https://noometry.com/benchmarks/omni-math) | 60.3% | #9 of 57, top 16% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |
| [LMArena Math](https://noometry.com/benchmarks/arena-math) | 1422 | #98 of 285, top 35% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [FrontierMath (Feb 2025 set)](https://noometry.com/benchmarks/frontiermath-2025-02) | 19.7% | #31 of 68, top 46% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-11-13 |
| [FrontierMath Tier 4 (v1)](https://noometry.com/benchmarks/frontiermath-tier-4-v1) | 2.1% | #41 of 55, top 75% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-08-11 |

### Knowledge

Grok 4 Knowledge benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [GPQA Diamond](https://noometry.com/benchmarks/gpqa-diamond) | 87% | #56 of 186, top 31% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [MMLU-Pro](https://noometry.com/benchmarks/mmlu-pro) | 85.1% | #7 of 58, top 13% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |
| [Confabulations](https://noometry.com/benchmarks/confabulations) (lower is better) | 12.4% | #7 of 51, top 14% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/confabulations) |  |
| [GPQA (HELM)](https://noometry.com/benchmarks/helm-gpqa) | 72.7% | #7 of 57, top 13% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |
| [LMArena Expert](https://noometry.com/benchmarks/arena-expert) | 1415 | #111 of 273, top 41% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Multimodal

Grok 4 Multimodal benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Vision](https://noometry.com/benchmarks/arena-vision) | 1210 | #76 of 122, top 63% |  | [LMArena](https://lmarena.ai/leaderboard/vision) | 2026-10-09 |
| [GeoBench](https://noometry.com/benchmarks/geobench) | 45% | #22 of 25, top 88% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Multilingual

Grok 4 Multilingual benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Non-English](https://noometry.com/benchmarks/arena-non-english) | 1403 | #103 of 297, top 35% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Chinese](https://noometry.com/benchmarks/arena-chinese) | 1427 | #120 of 285, top 43% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena French](https://noometry.com/benchmarks/arena-french) | 1418 | #103 of 223, top 47% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena German](https://noometry.com/benchmarks/arena-german) | 1429 | #67 of 231, top 30% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Japanese](https://noometry.com/benchmarks/arena-japanese) | 1394 | #65 of 211, top 31% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Korean](https://noometry.com/benchmarks/arena-korean) | 1377 | #82 of 213, top 39% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Russian](https://noometry.com/benchmarks/arena-russian) | 1410 | #95 of 283, top 34% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Spanish](https://noometry.com/benchmarks/arena-spanish) | 1420 | #92 of 226, top 41% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Instruction Following

Grok 4 Instruction Following benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [IFEval](https://noometry.com/benchmarks/ifeval) | 94.9% | #2 of 57, top 4% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |
| [LMArena Instruction Following](https://noometry.com/benchmarks/arena-instruction-following) | 1387 | #116 of 298, top 39% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Long Context

Grok 4 Long Context benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [Fiction.LiveBench](https://noometry.com/benchmarks/fiction-livebench) | 94.4% | #3 of 47, top 7% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [LMArena Longer Query](https://noometry.com/benchmarks/arena-longer-query) | 1409 | #109 of 291, top 38% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Writing & Preference

Grok 4 Writing & Preference benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Text](https://noometry.com/benchmarks/arena-text) | 1411 | #108 of 297, top 37% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Creative Writing](https://noometry.com/benchmarks/arena-creative-writing) | 1397 | #85 of 295, top 29% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [Short-Story Creative Writing](https://noometry.com/benchmarks/lech-mazur-writing) | 76.9% | #20 of 39, top 52% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [WildBench](https://noometry.com/benchmarks/wildbench) | 79.7% | #32 of 57, top 57% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |
| [LMArena Multi-Turn](https://noometry.com/benchmarks/arena-multi-turn) | 1416 | #99 of 295, top 34% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

## Compare Grok 4

-   [Grok 4 vs Grok 3](https://noometry.com/compare/grok-3-vs-grok-4)
-   [Grok 4 vs MiMo-V2.6-Flash](https://noometry.com/compare/grok-4-vs-mimo-v2-6-flash)
-   [Grok 4 vs Kimi K2.5](https://noometry.com/compare/grok-4-vs-kimi-k2-5)
-   [Grok 4 vs Grok 4.20 (Non-Reasoning)](https://noometry.com/compare/grok-4-vs-grok-4-20)
-   [Grok 4 vs Step 5 Preview](https://noometry.com/compare/grok-4-vs-step-5-preview)
-   [Grok 4 vs GPT-5.1](https://noometry.com/compare/gpt-5-1-vs-grok-4)
-   [Grok 4 vs GLM-5.1](https://noometry.com/compare/glm-5-1-vs-grok-4)
-   [Grok 4 vs GPT-6 Astra](https://noometry.com/compare/gpt-6-astra-vs-grok-4)
-   [Grok 4 vs Claude Fable 5.1](https://noometry.com/compare/claude-fable-5-1-vs-grok-4)
-   [Grok 4 vs Gemini 3.8 Flash](https://noometry.com/compare/gemini-3-8-flash-vs-grok-4)
-   [Grok 4 vs Kimi K3](https://noometry.com/compare/grok-4-vs-kimi-k3)
-   [Grok 4 vs Qwen3.8 Max](https://noometry.com/compare/grok-4-vs-qwen3-8-max)
-   [Grok 4 vs GLM-5.3](https://noometry.com/compare/glm-5-3-vs-grok-4)
-   [Grok 4 vs Muse Spark 1.3](https://noometry.com/compare/grok-4-vs-muse-spark-1-3)

## Other xAI models

-   [Grok 4.6](https://noometry.com/models/grok-4-6)56.9
-   [Grok 4.5](https://noometry.com/models/grok-4-5)55.0
-   [Grok 4.7](https://noometry.com/models/grok-4-7)53.1
-   [Grok 4.20 (Non-Reasoning)](https://noometry.com/models/grok-4-20)48.6
-   [Grok 4.20 Multi-Agent](https://noometry.com/models/grok-4-20-multi-agent)46.2
-   [Grok 4.3](https://noometry.com/models/grok-4-3)43.8
-   [Grok 4.1](https://noometry.com/models/grok-4-1)41.5
-   [Grok 4.1 Fast](https://noometry.com/models/grok-4-1-fast)41.4

## Frequently asked questions

### How good is Grok 4?

Grok 4 by xAI ranks 56th of 354 ranked models on the Noometry Index as of October 2026, with a score of 48.1. Its strongest category is long context, where it ranks 4th.

### Is Grok 4 open source?

No. Grok 4 is proprietary and available only through xAI's API and partner platforms.

### How fast is Grok 4?

Grok 4 generated about 1 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

### What are Grok 4's strengths and weaknesses?

Relative to other ranked models, Grok 4 places best in long context, instruction following, coding and lowest in multimodal, agentic & tool use, writing & preference.

### What is Grok 4 best at?

Its best category is long context, where it ranks 4th on Noometry.

### Cite this page

Noometry. (2026). Grok 4 benchmarks and pricing. Retrieved October 10, 2026, from https://noometry.com/models/grok-4

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/models/grok-4.md).
