Model comparison

Hunyuan Large Vision vs o3

o3 is the stronger model overall, scoring 47.5 to 37.6 on the Noometry Index.

Last verified . 13 shared benchmarks.

Hunyuan Large Vision Tencent

37.6

Rank #202 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Hunyuan Large Vision scores higher in 0 categories and o3 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o3 leads 54.6 to 34.4.

Side by side

Hunyuan Large Vision and o3 specifications
Hunyuan Large Visiono3
ProviderTencentOpenAI
Noometry Index37.647.5
Released—2025-04-16
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked1363

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Hunyuan Large Vision: 38.3 (#180), o3: 46.8 (#64)

Coding benchmarks
BenchmarkHunyuan Large Visiono3
LMArena Coding13081408
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
GSO—8.8%
WeirdML—52.4%
CadEval—74%
ALE-Bench—933.55

Agentic & Tool Use Not comparable

Hunyuan Large Vision: —, o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkHunyuan Large Visiono3
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

Hunyuan Large Vision: 24.8 (#158), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkHunyuan Large Visiono3
LMArena Hard Prompts12571402
ARC-AGI-2—6.5%
SimpleBench—53.1%
Kagi LLM Benchmark—67.6%
ARC-AGI-1—60.8%
CritPt—1.4%
Chess Puzzles—38%
EnigmaEval—13.1%
Mystery Game Puzzles—29%
DTBench—84.8%
LMCA—39.7%
Epoch Capabilities Index—146.86
ForecastBench—62.5

Math o3 leads

Hunyuan Large Vision: 35.5 (#183), o3: 50.2 (#58)

Math benchmarks
BenchmarkHunyuan Large Visiono3
LMArena Math12681426
FrontierMath (Tiers 1-3)—33.3%
OTIS Mock AIME 2024-2025—84.4%
Omni-MATH—71.4%
MATH Level 5—97.8%
FrontierMath (Feb 2025 set)—18.7%
FrontierMath Tier 4 (v1)—2.1%

Knowledge o3 leads

Hunyuan Large Vision: 34.4 (#196), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkHunyuan Large Visiono3
LMArena Expert12521402
GPQA Diamond—81.8%
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
Confabulations—14.4%
GPQA (HELM)—75.3%

Multimodal o3 leads

Hunyuan Large Vision: 35.7 (#83), o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkHunyuan Large Visiono3
LMArena Vision11801214
GeoBench—74%
VPCT—52%

Multilingual o3 leads

Hunyuan Large Vision: 39.9 (#221), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkHunyuan Large Visiono3
LMArena Non-English12361401
LMArena Chinese12871437
LMArena Russian12431406
LMArena French—1430
LMArena German—1420
LMArena Japanese—1403
LMArena Korean—1370
LMArena Spanish—1395

Instruction Following o3 leads

Hunyuan Large Vision: 65.8 (#215), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkHunyuan Large Visiono3
LMArena Instruction Following12511368
IFEval—86.9%

Long Context o3 leads

Hunyuan Large Vision: 38.9 (#189), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkHunyuan Large Visiono3
LMArena Longer Query12801372
Fiction.LiveBench—88.9%
CL-bench—17.8%

Writing & Preference o3 leads

Hunyuan Large Vision: 46.5 (#218), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkHunyuan Large Visiono3
LMArena Text12631410
LMArena Creative Writing12481359
LMArena Multi-Turn12541405
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
WildBench—86.1%

Frequently asked questions

Is Hunyuan Large Vision better than o3?

o3 is the stronger model overall, scoring 47.5 to 37.6 on the Noometry Index.

Is Hunyuan Large Vision or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 38.3 in the Noometry coding category.

How many benchmarks do Hunyuan Large Vision and o3 share?

13 benchmarks have published results for both models. Hunyuan Large Vision has 13 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper