Model comparison

Chatgpt 4o Latest 20250326 vs Olmo 3.1 32b Think

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 37.9 on the Noometry Index.

Last verified . 15 shared benchmarks.

Summary

  • They share 15 benchmarks with published results for both. Chatgpt 4o Latest 20250326 scores higher in 8 categories and Olmo 3.1 32b Think in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Chatgpt 4o Latest 20250326 leads 62.6 to 46.2.
  • Olmo 3.1 32b Think has downloadable open weights; the other is API-only.

Side by side

Chatgpt 4o Latest 20250326 and Olmo 3.1 32b Think specifications
Chatgpt 4o Latest 20250326Olmo 3.1 32b Think
ProviderOpenAIAllen Institute for AI (Ai2)
Noometry Index43.837.9
Released——
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2115

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 41.6 (#122), Olmo 3.1 32b Think: 37.7 (#189)

Coding benchmarks
BenchmarkChatgpt 4o Latest 20250326Olmo 3.1 32b Think
LMArena Coding14131291

Reasoning Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 33.8 (#71), Olmo 3.1 32b Think: 25.2 (#150)

Reasoning benchmarks
BenchmarkChatgpt 4o Latest 20250326Olmo 3.1 32b Think
LMArena Hard Prompts14241272
Kagi LLM Benchmark75%—

Math Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 38.6 (#134), Olmo 3.1 32b Think: 36.3 (#168)

Math benchmarks
BenchmarkChatgpt 4o Latest 20250326Olmo 3.1 32b Think
LMArena Math14071305

Knowledge Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 39.2 (#137), Olmo 3.1 32b Think: 35.7 (#181)

Knowledge benchmarks
BenchmarkChatgpt 4o Latest 20250326Olmo 3.1 32b Think
LMArena Expert14011295
Confabulations16.6%—

Multimodal Not comparable

Chatgpt 4o Latest 20250326: 39.6 (#58), Olmo 3.1 32b Think: —

Multimodal benchmarks
BenchmarkChatgpt 4o Latest 20250326Olmo 3.1 32b Think
LMArena Vision1243—

Multilingual Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 52.9 (#76), Olmo 3.1 32b Think: 38.1 (#231)

Multilingual benchmarks
BenchmarkChatgpt 4o Latest 20250326Olmo 3.1 32b Think
LMArena Non-English14191209
LMArena Chinese14571242
LMArena French14461260
LMArena German14231262
LMArena Russian14291193
LMArena Spanish14341289
LMArena Japanese1405—
LMArena Korean1396—

Instruction Following Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 74.0 (#107), Olmo 3.1 32b Think: 65.6 (#218)

Instruction Following benchmarks
BenchmarkChatgpt 4o Latest 20250326Olmo 3.1 32b Think
LMArena Instruction Following14031247

Long Context Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 43.1 (#107), Olmo 3.1 32b Think: 38.6 (#195)

Long Context benchmarks
BenchmarkChatgpt 4o Latest 20250326Olmo 3.1 32b Think
LMArena Longer Query14131272

Writing & Preference Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 62.6 (#72), Olmo 3.1 32b Think: 46.2 (#220)

Writing & Preference benchmarks
BenchmarkChatgpt 4o Latest 20250326Olmo 3.1 32b Think
LMArena Text14291272
LMArena Creative Writing14051226
LMArena Multi-Turn14541252
EQ-Bench Creative Writing1501—

Frequently asked questions

Is Chatgpt 4o Latest 20250326 better than Olmo 3.1 32b Think?

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 37.9 on the Noometry Index.

Is Chatgpt 4o Latest 20250326 or Olmo 3.1 32b Think better for coding?

Chatgpt 4o Latest 20250326 scores higher on coding benchmarks: 41.6 versus 37.7 in the Noometry coding category.

How many benchmarks do Chatgpt 4o Latest 20250326 and Olmo 3.1 32b Think share?

15 benchmarks have published results for both models. Chatgpt 4o Latest 20250326 has 21 scored results on Noometry and Olmo 3.1 32b Think has 15.

Related comparisons

Go deeper