Model comparison
Codestral vs Grok 4.1 Fast
Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 30.6 on the Noometry Index.
Last verified . 1 shared benchmarks.
Summary
- They share 1 benchmark with published results for both. Codestral scores higher in 0 categories and Grok 4.1 Fast in 2 categories; 2 gaps are clear of the uncertainty.
- The widest gap is in reasoning, where Grok 4.1 Fast leads 43.4 to 19.8.
- Grok 4.1 Fast is cheaper at $0.20 / $0.50 per million input/output tokens, against $0.30 / $0.90 for Codestral.
- Codestral accepts more context: 256K tokens versus 128K.
Side by side
| Codestral | Grok 4.1 Fast | |
|---|---|---|
| Provider | Mistral AI | xAI |
| Noometry Index | 30.6 | 41.4 |
| Released | 2024-05-29 | 2025-06-27 |
| Weights | Proprietary | Proprietary |
| Context window | 256K | 128K |
| Max output | 8K | 30K |
| Input $ / M tokens | $0.30 | $0.20 |
| Output $ / M tokens | $0.90 | $0.50 |
| Results tracked | 7 | 32 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Grok 4.1 Fast leads
Codestral: 27.3 (#321), Grok 4.1 Fast: 34.1 (#245)
| Benchmark | Codestral | Grok 4.1 Fast |
|---|---|---|
| ALE-Bench | 137.78 | 394.93 |
| Aider Polyglot | 11.1% | — |
| LMArena WebDev | — | 1242 |
| BigCodeBench Instruct | 41.8% | — |
| LMArena Coding | — | 1411 |
| BigCodeBench Complete | 52.5% | — |
| HumanEval+ | 73.8% | — |
| MBPP+ | 61.9% | — |
Agentic & Tool Use Not comparable
Codestral: —, Grok 4.1 Fast: 36.3 (#39)
| Benchmark | Codestral | Grok 4.1 Fast |
|---|---|---|
| Berkeley Function Calling Leaderboard | — | 69.6% |
| τ²-bench Banking | — | 13.1% |
| LMArena Search | — | 1171 |
| Vending-Bench 2 | — | 1,107 |
Reasoning Grok 4.1 Fast leads
Codestral: 19.8 (#251), Grok 4.1 Fast: 43.4 (#49)
| Benchmark | Codestral | Grok 4.1 Fast |
|---|---|---|
| SimpleBench | — | 56% |
| Kagi LLM Benchmark | 32.5% | — |
| NYT Connections (extended) | — | 87.4% |
| LMArena Hard Prompts | — | 1407 |
| DTBench | — | 87.7% |
| ForecastBench | — | 61 |
Math Not comparable
Codestral: —, Grok 4.1 Fast: 31.9 (#221)
| Benchmark | Codestral | Grok 4.1 Fast |
|---|---|---|
| MathArena Final-Answer Competitions | — | 60.9% |
| ProofBench | — | 4% |
| LMArena Math | — | 1408 |
Knowledge Not comparable
Codestral: —, Grok 4.1 Fast: 33.1 (#207)
| Benchmark | Codestral | Grok 4.1 Fast |
|---|---|---|
| Vectara Hallucination Rate | — | 17.8% |
| LMArena Expert | — | 1399 |
Multimodal Not comparable
Codestral: —, Grok 4.1 Fast: 37.0 (#76)
| Benchmark | Codestral | Grok 4.1 Fast |
|---|---|---|
| LMArena Vision | — | 1201 |
Multilingual Not comparable
Codestral: —, Grok 4.1 Fast: 51.0 (#114)
| Benchmark | Codestral | Grok 4.1 Fast |
|---|---|---|
| LMArena Non-English | — | 1391 |
| LMArena Chinese | — | 1441 |
| LMArena French | — | 1415 |
| LMArena German | — | 1404 |
| LMArena Japanese | — | 1349 |
| LMArena Korean | — | 1361 |
| LMArena Russian | — | 1387 |
| LMArena Spanish | — | 1413 |
Instruction Following Not comparable
Codestral: —, Grok 4.1 Fast: 72.7 (#133)
| Benchmark | Codestral | Grok 4.1 Fast |
|---|---|---|
| LMArena Instruction Following | — | 1376 |
Long Context Not comparable
Codestral: —, Grok 4.1 Fast: 42.4 (#126)
| Benchmark | Codestral | Grok 4.1 Fast |
|---|---|---|
| LMArena Longer Query | — | 1390 |
Writing & Preference Not comparable
Codestral: —, Grok 4.1 Fast: 57.2 (#131)
| Benchmark | Codestral | Grok 4.1 Fast |
|---|---|---|
| LMArena Text | — | 1408 |
| LMArena Creative Writing | — | 1394 |
| EQ-Bench Creative Writing | — | 1327 |
| LMArena Multi-Turn | — | 1389 |
Frequently asked questions
Is Codestral better than Grok 4.1 Fast?
Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 30.6 on the Noometry Index.
Which is cheaper, Codestral or Grok 4.1 Fast?
Grok 4.1 Fast is cheaper. It lists at $0.20 per million input tokens and $0.50 per million output tokens; Codestral lists at $0.30 and $0.90.
Is Codestral or Grok 4.1 Fast better for coding?
Grok 4.1 Fast scores higher on coding benchmarks: 34.1 versus 27.3 in the Noometry coding category.
Which has the bigger context window?
Codestral does, with 256K tokens against 128K.
How many benchmarks do Codestral and Grok 4.1 Fast share?
1 benchmark has published results for both models. Codestral has 7 scored results on Noometry and Grok 4.1 Fast has 32.