Model comparison

GPT-4 Turbo vs Ministral 3B

GPT-4 Turbo is the stronger model overall, scoring 30.5 to 26.2 on the Noometry Index. Ministral 3B costs 150× less per token, which makes it the better buy when GPT-4 Turbo's lead doesn't matter for your workload.

Last verified . 5 shared benchmarks.

GPT-4 Turbo OpenAI

30.5

Rank #292 Confirmed

Ministral 3B Mistral AI

26.2

Rank #338 Confirmed

Summary

  • They share 5 benchmarks with published results for both. GPT-4 Turbo scores higher in 1 category and Ministral 3B in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Ministral 3B leads 26.6 to 9.0.
  • The biggest single-benchmark swing is MATH Level 5: 46.7% for GPT-4 Turbo and 14.4% for Ministral 3B.
  • Ministral 3B is cheaper at $0.10 / $0.10 per million input/output tokens, against $10 / $30 for GPT-4 Turbo.
  • Ministral 3B accepts more context: 131K tokens versus 128K.
  • Ministral 3B has downloadable open weights; the other is API-only.

Side by side

GPT-4 Turbo and Ministral 3B specifications
GPT-4 TurboMinistral 3B
ProviderOpenAIMistral AI
Noometry Index30.526.2
Released2023-11-062024-10-01
WeightsProprietaryOpen
Context window128K131K
Max output4K262K
Input $ / M tokens$10$0.10
Output $ / M tokens$30$0.10
Results tracked366

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

GPT-4 Turbo: 33.8 (#249), Ministral 3B: —

Coding benchmarks
BenchmarkGPT-4 TurboMinistral 3B
WeirdML18%—
BigCodeBench Instruct48.2%—
LMArena Coding1268—
BigCodeBench Complete58.2%—
HumanEval+86.6%—
MBPP+73.3%—

Agentic & Tool Use Not comparable

GPT-4 Turbo: —, Ministral 3B: —

Agentic & Tool Use benchmarks
BenchmarkGPT-4 TurboMinistral 3B
METR Time Horizons36.7%—

Reasoning Ministral 3B leads

GPT-4 Turbo: 15.3 (#317), Ministral 3B: 18.4 (#282)

Reasoning benchmarks
BenchmarkGPT-4 TurboMinistral 3B
DTBench61.6%51.7%
LMCA9.8%5.5%
Epoch Capabilities Index127.25118.1
SimpleBench25.1%—
Chess Puzzles6%—
LMArena Hard Prompts1251—
ForecastBench59.4—

Math Ministral 3B leads

GPT-4 Turbo: 9.0 (#322), Ministral 3B: 26.6 (#258)

Math benchmarks
BenchmarkGPT-4 TurboMinistral 3B
MATH Level 546.7%14.4%
FrontierMath (Tiers 1-3)0.7%—
OTIS Mock AIME 2024-20256.7%—
LMArena Math1272—

Knowledge GPT-4 Turbo leads

GPT-4 Turbo: 24.3 (#268), Ministral 3B: 10.4 (#302)

Knowledge benchmarks
BenchmarkGPT-4 TurboMinistral 3B
GPQA Diamond46.6%25.3%
Confabulations28.4%—
Vectara Hallucination Rate—7.3%
LMArena Expert1223—
MMLU81.3%—

Multimodal Not comparable

GPT-4 Turbo: 30.6 (#110), Ministral 3B: —

Multimodal benchmarks
BenchmarkGPT-4 TurboMinistral 3B
LMArena Vision1090—

Multilingual Not comparable

GPT-4 Turbo: 40.5 (#216), Ministral 3B: —

Multilingual benchmarks
BenchmarkGPT-4 TurboMinistral 3B
LMArena Non-English1245—
LMArena Chinese1242—
LMArena French1276—
LMArena German1259—
LMArena Japanese1194—
LMArena Korean1187—
LMArena Russian1259—
LMArena Spanish1260—

Instruction Following Not comparable

GPT-4 Turbo: 65.8 (#216), Ministral 3B: —

Instruction Following benchmarks
BenchmarkGPT-4 TurboMinistral 3B
LMArena Instruction Following1249—

Long Context Not comparable

GPT-4 Turbo: 38.0 (#206), Ministral 3B: —

Long Context benchmarks
BenchmarkGPT-4 TurboMinistral 3B
LMArena Longer Query1254—

Writing & Preference Not comparable

GPT-4 Turbo: 47.7 (#206), Ministral 3B: —

Writing & Preference benchmarks
BenchmarkGPT-4 TurboMinistral 3B
LMArena Text1272—
LMArena Creative Writing1269—
LMArena Multi-Turn1267—

Frequently asked questions

Is GPT-4 Turbo better than Ministral 3B?

GPT-4 Turbo is the stronger model overall, scoring 30.5 to 26.2 on the Noometry Index. Ministral 3B costs 150× less per token, which makes it the better buy when GPT-4 Turbo's lead doesn't matter for your workload.

Which is cheaper, GPT-4 Turbo or Ministral 3B?

Ministral 3B is cheaper. It lists at $0.10 per million input tokens and $0.10 per million output tokens; GPT-4 Turbo lists at $10 and $30.

Which has the bigger context window?

Ministral 3B does, with 131K tokens against 128K.

How many benchmarks do GPT-4 Turbo and Ministral 3B share?

5 benchmarks have published results for both models. GPT-4 Turbo has 36 scored results on Noometry and Ministral 3B has 6.

Related comparisons

Go deeper