# Grok 4.1 vs Magistral Medium

> Grok 4.1 is the stronger model overall, scoring 41.5 to 35.2 on the Noometry Index.

- Canonical page: https://noometry.com/compare/grok-4-1-vs-magistral-medium
- Last updated: 2026-10-10
- Shared benchmarks: 17

## Summary

- They share 17 benchmarks with published results for both. Grok 4.1 scores higher in 7 categories and Magistral Medium in 1 category; 8 gaps are clear of the uncertainty.
- The widest gap is in reasoning, where Grok 4.1 leads 29.5 to 8.6.
- Magistral Medium has downloadable open weights; the other is API-only.

## Snapshot

| | Grok 4.1 | Magistral Medium |
|---|---|---|
| Provider | xAI | Mistral AI |
| Noometry Index | 41.5 | 35.2 |
| Rank | 134 | 227 |
| Context | — | 262K |
| Input $/M | — | $2 |
| Output $/M | — | $5 |
| Weights | Proprietary | Open |

## Coding

- Grok 4.1: 33.7 (#253)
- Magistral Medium: 39.1 (#161)

| Benchmark | Grok 4.1 | Magistral Medium |
|---|---|---|
| LMArena Coding | 1445 | 1319 |
| LMArena WebDev | 1214 | — |
| SciCode | — | 39.2% |

## Agentic & Tool Use

- Grok 4.1: 34.1 (#49)
- Magistral Medium: —

| Benchmark | Grok 4.1 | Magistral Medium |
|---|---|---|
| Cybench | 39% | — |

## Reasoning

- Grok 4.1: 29.5 (#91)
- Magistral Medium: 8.6 (#348)

| Benchmark | Grok 4.1 | Magistral Medium |
|---|---|---|
| LMArena Hard Prompts | 1435 | 1267 |
| ARC-AGI-2 | — | 0% |
| Kagi LLM Benchmark | — | 16.2% |
| ARC-AGI-1 | — | 6.1% |
| CritPt | — | 0.3% |

## Math

- Grok 4.1: 38.9 (#120)
- Magistral Medium: 35.1 (#189)

| Benchmark | Grok 4.1 | Magistral Medium |
|---|---|---|
| LMArena Math | 1422 | 1250 |

## Knowledge

- Grok 4.1: 39.5 (#133)
- Magistral Medium: 33.5 (#202)

| Benchmark | Grok 4.1 | Magistral Medium |
|---|---|---|
| LMArena Expert | 1417 | 1223 |

## Multilingual

- Grok 4.1: 53.4 (#68)
- Magistral Medium: 39.6 (#224)

| Benchmark | Grok 4.1 | Magistral Medium |
|---|---|---|
| LMArena Non-English | 1425 | 1232 |
| LMArena Chinese | 1465 | 1227 |
| LMArena French | 1448 | 1267 |
| LMArena German | 1446 | 1248 |
| LMArena Japanese | 1397 | 1175 |
| LMArena Korean | 1407 | 1125 |
| LMArena Russian | 1434 | 1224 |
| LMArena Spanish | 1438 | 1271 |

## Instruction Following

- Grok 4.1: 73.8 (#111)
- Magistral Medium: 66.0 (#211)

| Benchmark | Grok 4.1 | Magistral Medium |
|---|---|---|
| LMArena Instruction Following | 1400 | 1254 |

## Long Context

- Grok 4.1: 43.2 (#100)
- Magistral Medium: 39.3 (#183)

| Benchmark | Grok 4.1 | Magistral Medium |
|---|---|---|
| LMArena Longer Query | 1416 | 1295 |

## Writing & Preference

- Grok 4.1: 62.4 (#75)
- Magistral Medium: 46.3 (#219)

| Benchmark | Grok 4.1 | Magistral Medium |
|---|---|---|
| LMArena Text | 1437 | 1255 |
| LMArena Creative Writing | 1411 | 1245 |
| LMArena Multi-Turn | 1437 | 1275 |

## FAQ

### Is Grok 4.1 better than Magistral Medium?

Grok 4.1 is the stronger model overall, scoring 41.5 to 35.2 on the Noometry Index.

### Is Grok 4.1 or Magistral Medium better for coding?

Magistral Medium scores higher on coding benchmarks: 39.1 versus 33.7 in the Noometry coding category.

### How many benchmarks do Grok 4.1 and Magistral Medium share?

17 benchmarks have published results for both models. Grok 4.1 has 19 scored results on Noometry and Magistral Medium has 22.
