Model comparison

Hunyuan Vision 1.5 vs o3

o3 is the stronger model overall, scoring 47.5 to 43.1 on the Noometry Index.

Last verified . 9 shared benchmarks.

Hunyuan Vision 1.5 Tencent

43.1

Rank #98 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Hunyuan Vision 1.5 scores higher in 1 category and o3 in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in long context, where o3 leads 53.3 to 42.9.

Side by side

Hunyuan Vision 1.5 and o3 specifications
Hunyuan Vision 1.5o3
ProviderTencentOpenAI
Noometry Index43.147.5
Released—2025-04-16
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked963

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Hunyuan Vision 1.5: 41.8 (#119), o3: 46.8 (#64)

Coding benchmarks
BenchmarkHunyuan Vision 1.5o3
LMArena Coding14181408
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
GSO—8.8%
WeirdML—52.4%
CadEval—74%
ALE-Bench—933.55

Agentic & Tool Use Not comparable

Hunyuan Vision 1.5: —, o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkHunyuan Vision 1.5o3
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

Hunyuan Vision 1.5: 28.9 (#97), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkHunyuan Vision 1.5o3
LMArena Hard Prompts14161402
ARC-AGI-2—6.5%
SimpleBench—53.1%
Kagi LLM Benchmark—67.6%
ARC-AGI-1—60.8%
CritPt—1.4%
Chess Puzzles—38%
EnigmaEval—13.1%
Mystery Game Puzzles—29%
DTBench—84.8%
LMCA—39.7%
Epoch Capabilities Index—146.86
ForecastBench—62.5

Math Not comparable

Hunyuan Vision 1.5: —, o3: 50.2 (#58)

Math benchmarks
BenchmarkHunyuan Vision 1.5o3
FrontierMath (Tiers 1-3)—33.3%
OTIS Mock AIME 2024-2025—84.4%
Omni-MATH—71.4%
LMArena Math—1426
MATH Level 5—97.8%
FrontierMath (Feb 2025 set)—18.7%
FrontierMath Tier 4 (v1)—2.1%

Knowledge Not comparable

Hunyuan Vision 1.5: —, o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkHunyuan Vision 1.5o3
GPQA Diamond—81.8%
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
Confabulations—14.4%
GPQA (HELM)—75.3%
LMArena Expert—1402

Multimodal o3 leads

Hunyuan Vision 1.5: 35.7 (#84), o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkHunyuan Vision 1.5o3
LMArena Vision11801214
GeoBench—74%
VPCT—52%

Multilingual o3 leads

Hunyuan Vision 1.5: 50.3 (#125), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkHunyuan Vision 1.5o3
LMArena Non-English13821401
LMArena Chinese—1437
LMArena French—1430
LMArena German—1420
LMArena Japanese—1403
LMArena Korean—1370
LMArena Russian—1406
LMArena Spanish—1395

Instruction Following Too close to call

Hunyuan Vision 1.5: 73.7 (#117), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkHunyuan Vision 1.5o3
LMArena Instruction Following13971368
IFEval—86.9%

Long Context o3 leads

Hunyuan Vision 1.5: 42.9 (#115), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkHunyuan Vision 1.5o3
LMArena Longer Query14041372
Fiction.LiveBench—88.9%
CL-bench—17.8%

Writing & Preference o3 leads

Hunyuan Vision 1.5: 60.1 (#100), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkHunyuan Vision 1.5o3
LMArena Text14051410
LMArena Creative Writing13881359
LMArena Multi-Turn14221405
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
WildBench—86.1%

Frequently asked questions

Is Hunyuan Vision 1.5 better than o3?

o3 is the stronger model overall, scoring 47.5 to 43.1 on the Noometry Index.

Is Hunyuan Vision 1.5 or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 41.8 in the Noometry coding category.

How many benchmarks do Hunyuan Vision 1.5 and o3 share?

9 benchmarks have published results for both models. Hunyuan Vision 1.5 has 9 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper