xAI, proprietary

# Grok 3

> Grok 3 by xAI, released April 2025. Ranked #157 of 354 with a Noometry Index of 39.9. Scores, sources and comparisons.
- Canonical page: https://noometry.com/models/grok-3
- Last updated: 2026-10-10
- Title: Grok 3 Benchmarks, Price & Rank (October 2026) | Noometry

Grok 3 by xAI ranks 157th of 354 ranked models on the Noometry Index as of October 2026, with a score of 39.9. Its strongest category is instruction following, where it ranks 73rd.

Last verified October 10, 2026

## Specifications

- **Noometry rank:** #157 of 354
- **Index score:** 39.9
- **Evidence:** Confirmed 40 results
- **Provider:** [xAI](https://noometry.com/providers/xai)
- **Released:** April 9, 2025
- **Weights:** Proprietary
- **Reasoning:** Unknown
- **Context window:** —
- **Max output:** —
- **Input price:** Not listed
- **Output price:** Not listed
- **Blended price:** Not listed
- **Output speed:** 42 tokens/s [Kagi](https://help.kagi.com/kagi/ai/llm-benchmark.html)
- **Value:** Not ranked
- **Knowledge cutoff:** Unknown

## Category scores

Each category score combines every public result we have in that category.

Grok 3 category scores

1.  Coding 41.9
2.  Agentic & Tool Use 30.5
3.  Reasoning 13.7
4.  Math 38.0
5.  Knowledge 46.2
6.  Multilingual 52.3
7.  Instruction Following 75.0
8.  Long Context 38.7
9.  Writing & Preference 55.8
10.  020406080

Grok 3 category ranks
| Category | Score | Rank | Results |
| --- | --- | --- | --- |
| [Coding](https://noometry.com/best/coding) | 41.9 | #115 | 3 |
| [Agentic & Tool Use](https://noometry.com/best/agentic) | 30.5 | #76 | 1 |
| [Reasoning](https://noometry.com/best/reasoning) | 13.7 | #333 | 5 |
| [Math](https://noometry.com/best/math) | 38.0 | #145 | 4 |
| [Knowledge](https://noometry.com/best/knowledge) | 46.2 | #82 | 6 |
| [Multilingual](https://noometry.com/best/multilingual) | 52.3 | #87 | 1 |
| [Instruction Following](https://noometry.com/best/instruction-following) | 75.0 | #73 | 2 |
| [Long Context](https://noometry.com/best/long-context) | 38.7 | #192 | 2 |
| [Writing & Preference](https://noometry.com/best/writing) | 55.8 | #141 | 6 |

## Strengths and weaknesses

Categories where Grok 3 places highest and lowest among the models ranked in each, with its score against that category's median.

### Strongest categories

Grok 3: strongest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Instruction Following](https://noometry.com/best/instruction-following) | 75.0 | +3.8 | #73 of 305, top 24% |
| [Knowledge](https://noometry.com/best/knowledge) | 46.2 | +8.8 | #82 of 314, top 27% |
| [Multilingual](https://noometry.com/best/multilingual) | 52.3 | +4.9 | #87 of 297, top 30% |

### Weakest categories

Grok 3: weakest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Reasoning](https://noometry.com/best/reasoning) | 13.7 | −10.0 | #333 of 350, top 96% |
| [Long Context](https://noometry.com/best/long-context) | 38.7 | −2.2 | #192 of 296, top 65% |
| [Agentic & Tool Use](https://noometry.com/best/agentic) | 30.5 | +0.2 | #76 of 154, top 50% |

## Closest competitors

The models ranked just above and below Grok 3. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Grok 3
| Model | Rank | Score | Blended $/M | Speed |  |
| --- | --- | --- | --- | --- | --- |
| [Nemotron 3 Super](https://noometry.com/models/nemotron-3-super) | #153 | 40.1 | $0.17 | — | [Compare](https://noometry.com/compare/grok-3-vs-nemotron-3-super) |
| [Llama 3.3 Nemotron 49b Super v1](https://noometry.com/models/llama-3-3-nemotron-49b-super-v1) | #154 | 40.1 | — | — | [Compare](https://noometry.com/compare/grok-3-vs-llama-3-3-nemotron-49b-super-v1) |
| [Nemotron 3.5 Lightning](https://noometry.com/models/nemotron-3-5-lightning) | #155 | 40.0 | $0.0875 | — | [Compare](https://noometry.com/compare/grok-3-vs-nemotron-3-5-lightning) |
| [Qwen3.7 Flash](https://noometry.com/models/qwen3-7-flash) | #156 | 39.9 | $0.055 | — | [Compare](https://noometry.com/compare/grok-3-vs-qwen3-7-flash) |
| [GLM-4.5V](https://noometry.com/models/glm-4-5v) | #158 | 39.8 | $0.90 | 34 | [Compare](https://noometry.com/compare/glm-4-5v-vs-grok-3) |
| [QwQ-32B](https://noometry.com/models/qwq-32b) | #159 | 39.8 | — | — | [Compare](https://noometry.com/compare/grok-3-vs-qwq-32b) |
| [Step 1o Turbo 202506](https://noometry.com/models/step-1o-turbo-202506) | #160 | 39.7 | — | — | [Compare](https://noometry.com/compare/grok-3-vs-step-1o-turbo-202506) |
| [Nova 2 Lite](https://noometry.com/models/nova-2-lite) | #161 | 39.7 | $0.85 | — | [Compare](https://noometry.com/compare/grok-3-vs-nova-2-lite) |

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

### Coding

Grok 3 Coding benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [Aider Polyglot](https://noometry.com/benchmarks/aider-polyglot) | 53.3% | #20 of 44, top 46% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [WeirdML](https://noometry.com/benchmarks/weirdml) | 37.2% | #87 of 119, top 74% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [LMArena Coding](https://noometry.com/benchmarks/arena-coding) | 1432 | #109 of 294, top 38% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Agentic & Tool Use

Grok 3 Agentic & Tool Use benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [BALROG](https://noometry.com/benchmarks/balrog) | 29.5% | #18 of 35, top 52% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Reasoning

Grok 3 Reasoning benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [ARC-AGI-2](https://noometry.com/benchmarks/arc-agi-2) | 0% | #79 of 83, top 96% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [SimpleBench](https://noometry.com/benchmarks/simplebench) | 36.1% | #56 of 77, top 73% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Kagi LLM Benchmark](https://noometry.com/benchmarks/kagi-reasoning) | 61.3% | #40 of 99, top 41% |  | [Kagi LLM Benchmark](https://help.kagi.com/kagi/ai/llm-benchmark.html) |  |
| [ARC-AGI-1](https://noometry.com/benchmarks/arc-agi-1) | 5.5% | #77 of 83, top 93% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [LMArena Hard Prompts](https://noometry.com/benchmarks/arena-hard-prompts) | 1434 | #87 of 297, top 30% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [Epoch Capabilities Index](https://noometry.com/benchmarks/epoch-capabilities-index) | 138.33 | #114 of 213, top 54% |  | [Epoch AI](https://epoch.ai/eci) | 2025-04-09 |

### Math

Grok 3 Math benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [OTIS Mock AIME 2024-2025](https://noometry.com/benchmarks/otis-mock-aime) | 55.6% | #110 of 173, top 64% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-04-10 |
| [Omni-MATH](https://noometry.com/benchmarks/omni-math) | 46.4% | #20 of 57, top 36% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |
| [LMArena Math](https://noometry.com/benchmarks/arena-math) | 1391 | #135 of 285, top 48% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [MATH Level 5](https://noometry.com/benchmarks/math-level-5) | 88.7% | #17 of 79, top 22% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-04-10 |
| [FrontierMath (Feb 2025 set)](https://noometry.com/benchmarks/frontiermath-2025-02) | 3.8% | #52 of 68, top 77% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-04-10 |
| [FrontierMath Tier 4 (v1)](https://noometry.com/benchmarks/frontiermath-tier-4-v1) | 0% | #51 of 55, top 93% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-07-01 |

### Knowledge

Grok 3 Knowledge benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [GPQA Diamond](https://noometry.com/benchmarks/gpqa-diamond) | 75.8% | #93 of 186, top 50% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-05-26 |
| [MMLU-Pro](https://noometry.com/benchmarks/mmlu-pro) | 78.8% | #19 of 58, top 33% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |
| [Confabulations](https://noometry.com/benchmarks/confabulations) (lower is better) | 14.2% | #14 of 51, top 28% | no reasoning | [Lech Mazur benchmarks](https://github.com/lechmazur/confabulations) |  |
| [Vectara Hallucination Rate](https://noometry.com/benchmarks/vectara-hallucination) (lower is better) | 5.8% | #20 of 96, top 21% |  | [Vectara Hallucination Leaderboard](https://github.com/vectara/hallucination-leaderboard) |  |
| [GPQA (HELM)](https://noometry.com/benchmarks/helm-gpqa) | 65% | #18 of 57, top 32% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |
| [LMArena Expert](https://noometry.com/benchmarks/arena-expert) | 1421 | #106 of 273, top 39% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Multilingual

Grok 3 Multilingual benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Non-English](https://noometry.com/benchmarks/arena-non-english) | 1410 | #87 of 297, top 30% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Chinese](https://noometry.com/benchmarks/arena-chinese) | 1448 | #103 of 285, top 37% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena French](https://noometry.com/benchmarks/arena-french) | 1460 | #49 of 223, top 22% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena German](https://noometry.com/benchmarks/arena-german) | 1431 | #65 of 231, top 29% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Japanese](https://noometry.com/benchmarks/arena-japanese) | 1387 | #71 of 211, top 34% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Korean](https://noometry.com/benchmarks/arena-korean) | 1373 | #83 of 213, top 39% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Russian](https://noometry.com/benchmarks/arena-russian) | 1416 | #86 of 283, top 31% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Spanish](https://noometry.com/benchmarks/arena-spanish) | 1417 | #94 of 226, top 42% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Instruction Following

Grok 3 Instruction Following benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [IFEval](https://noometry.com/benchmarks/ifeval) | 88.4% | #12 of 57, top 22% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |
| [LMArena Instruction Following](https://noometry.com/benchmarks/arena-instruction-following) | 1409 | #89 of 298, top 30% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Long Context

Grok 3 Long Context benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [Fiction.LiveBench](https://noometry.com/benchmarks/fiction-livebench) | 58.3% | #31 of 47, top 66% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [LMArena Longer Query](https://noometry.com/benchmarks/arena-longer-query) | 1439 | #66 of 291, top 23% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Writing & Preference

Grok 3 Writing & Preference benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Text](https://noometry.com/benchmarks/arena-text) | 1426 | #86 of 297, top 29% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Creative Writing](https://noometry.com/benchmarks/arena-creative-writing) | 1414 | #60 of 295, top 21% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [Short-Story Creative Writing](https://noometry.com/benchmarks/lech-mazur-writing) | 76.4% | #22 of 39, top 57% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [EQ-Bench Creative Writing](https://noometry.com/benchmarks/eqbench-creative-writing) | 1186 | #86 of 115, top 75% |  | [EQ-Bench](https://eqbench.com/creative_writing.html) |  |
| [WildBench](https://noometry.com/benchmarks/wildbench) | 84.9% | #13 of 57, top 23% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |
| [LMArena Multi-Turn](https://noometry.com/benchmarks/arena-multi-turn) | 1425 | #91 of 295, top 31% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

## Compare Grok 3

-   [Grok 3 vs Qwen3.7 Flash](https://noometry.com/compare/grok-3-vs-qwen3-7-flash)
-   [Grok 3 vs GLM-4.5V](https://noometry.com/compare/glm-4-5v-vs-grok-3)
-   [Grok 3 vs Nemotron 3.5 Lightning](https://noometry.com/compare/grok-3-vs-nemotron-3-5-lightning)
-   [Grok 3 vs QwQ-32B](https://noometry.com/compare/grok-3-vs-qwq-32b)
-   [Grok 3 vs Llama 3.3 Nemotron 49b Super v1](https://noometry.com/compare/grok-3-vs-llama-3-3-nemotron-49b-super-v1)
-   [Grok 3 vs Step 1o Turbo 202506](https://noometry.com/compare/grok-3-vs-step-1o-turbo-202506)
-   [Grok 3 vs GPT-6 Astra](https://noometry.com/compare/gpt-6-astra-vs-grok-3)
-   [Grok 3 vs Claude Fable 5.1](https://noometry.com/compare/claude-fable-5-1-vs-grok-3)
-   [Grok 3 vs Gemini 3.8 Flash](https://noometry.com/compare/gemini-3-8-flash-vs-grok-3)
-   [Grok 3 vs Kimi K3](https://noometry.com/compare/grok-3-vs-kimi-k3)
-   [Grok 3 vs Qwen3.8 Max](https://noometry.com/compare/grok-3-vs-qwen3-8-max)
-   [Grok 3 vs GLM-5.3](https://noometry.com/compare/glm-5-3-vs-grok-3)
-   [Grok 3 vs Muse Spark 1.3](https://noometry.com/compare/grok-3-vs-muse-spark-1-3)
-   [Grok 3 vs DeepSeek V4 Pro](https://noometry.com/compare/deepseek-v4-pro-vs-grok-3)

## Other xAI models

-   [Grok 4.6](https://noometry.com/models/grok-4-6)56.9
-   [Grok 4.5](https://noometry.com/models/grok-4-5)55.0
-   [Grok 4.7](https://noometry.com/models/grok-4-7)53.1
-   [Grok 4.20 (Non-Reasoning)](https://noometry.com/models/grok-4-20)48.6
-   [Grok 4](https://noometry.com/models/grok-4)48.1
-   [Grok 4.20 Multi-Agent](https://noometry.com/models/grok-4-20-multi-agent)46.2
-   [Grok 4.3](https://noometry.com/models/grok-4-3)43.8
-   [Grok 4.1](https://noometry.com/models/grok-4-1)41.5

## Frequently asked questions

### How good is Grok 3?

Grok 3 by xAI ranks 157th of 354 ranked models on the Noometry Index as of October 2026, with a score of 39.9. Its strongest category is instruction following, where it ranks 73rd.

### Is Grok 3 open source?

No. Grok 3 is proprietary and available only through xAI's API and partner platforms.

### How fast is Grok 3?

Grok 3 generated about 42 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

### What are Grok 3's strengths and weaknesses?

Relative to other ranked models, Grok 3 places best in instruction following, knowledge, multilingual and lowest in reasoning, long context, agentic & tool use.

### What is Grok 3 best at?

Its best category is instruction following, where it ranks 73rd on Noometry.

### Cite this page

Noometry. (2026). Grok 3 benchmarks and pricing. Retrieved October 10, 2026, from https://noometry.com/models/grok-3

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/models/grok-3.md).
