Model comparison

Grok 4.7 vs Hunyuan Turbos 20250226

Grok 4.7 is the stronger model overall, scoring 53.1 to 41.3 on the Noometry Index.

Last verified . 13 shared benchmarks.

Grok 4.7 xAI

53.1

Rank #37 Confirmed

Hunyuan Turbos 20250226 Tencent

41.3

Rank #139 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Grok 4.7 scores higher in 8 categories and Hunyuan Turbos 20250226 in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.7 leads 62.8 to 37.0.

Side by side

Grok 4.7 and Hunyuan Turbos 20250226 specifications
Grok 4.7Hunyuan Turbos 20250226
ProviderxAITencent
Noometry Index53.141.3
Released2026-09-21—
WeightsProprietaryProprietary
Context window500K—
Max output500K—
Input $ / M tokens$2—
Output $ / M tokens$6—
Results tracked3916

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.7 leads

Grok 4.7: 58.0 (#18), Hunyuan Turbos 20250226: 40.0 (#152)

Coding benchmarks
BenchmarkGrok 4.7Hunyuan Turbos 20250226
LMArena Coding14271361
FrontierCode47.6%—
CursorBench46.3%—
LMArena WebDev1639—
FrontierSWE29.5%—
SciCode57.8%—

Agentic & Tool Use Not comparable

Grok 4.7: 36.7 (#37), Hunyuan Turbos 20250226: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.7Hunyuan Turbos 20250226
APEX-Agents54.6%—
GDP.pdf22.8%—
Vending-Bench 210,537—

Reasoning Grok 4.7 leads

Grok 4.7: 49.1 (#40), Hunyuan Turbos 20250226: 27.8 (#113)

Reasoning benchmarks
BenchmarkGrok 4.7Hunyuan Turbos 20250226
LMArena Hard Prompts14131374
NYT Connections (extended)76.8%—
CritPt18%—
Chess Puzzles38%—
Mystery Game Puzzles29%—
DTBench96%—
LMCA49.4%—
Epoch Capabilities Index153.53—

Math Grok 4.7 leads

Grok 4.7: 57.8 (#39), Hunyuan Turbos 20250226: 37.5 (#154)

Math benchmarks
BenchmarkGrok 4.7Hunyuan Turbos 20250226
LMArena Math14071359
FrontierMath (Tiers 1-3)53%—
FrontierMath Tier 417.1%—
OTIS Mock AIME 2024-202598.1%—
ProofBench34%—

Knowledge Grok 4.7 leads

Grok 4.7: 62.8 (#22), Hunyuan Turbos 20250226: 37.0 (#161)

Knowledge benchmarks
BenchmarkGrok 4.7Hunyuan Turbos 20250226
LMArena Expert14221339
GPQA Diamond92.7%—
SimpleQA Verified56%—

Multimodal Not comparable

Grok 4.7: 35.5 (#87), Hunyuan Turbos 20250226: —

Multimodal benchmarks
BenchmarkGrok 4.7Hunyuan Turbos 20250226
LMArena Vision1228—
Blueprint-Bench 232.5%—
Furniture Assembly20.8%—

Multilingual Grok 4.7 leads

Grok 4.7: 50.8 (#116), Hunyuan Turbos 20250226: 48.9 (#136)

Multilingual benchmarks
BenchmarkGrok 4.7Hunyuan Turbos 20250226
LMArena Non-English13891363
LMArena Chinese14551417
LMArena French14551391
LMArena Russian13971368
LMArena German—1355
LMArena Japanese—1342
LMArena Korean—1351
LMArena Spanish1400—

Instruction Following Grok 4.7 leads

Grok 4.7: 74.1 (#105), Hunyuan Turbos 20250226: 71.0 (#158)

Instruction Following benchmarks
BenchmarkGrok 4.7Hunyuan Turbos 20250226
LMArena Instruction Following14041344

Long Context Grok 4.7 leads

Grok 4.7: 43.1 (#104), Hunyuan Turbos 20250226: 41.6 (#136)

Long Context benchmarks
BenchmarkGrok 4.7Hunyuan Turbos 20250226
LMArena Longer Query14131366

Writing & Preference Grok 4.7 leads

Grok 4.7: 70.0 (#24), Hunyuan Turbos 20250226: 57.4 (#128)

Writing & Preference benchmarks
BenchmarkGrok 4.7Hunyuan Turbos 20250226
LMArena Text13991377
LMArena Creative Writing13911359
LMArena Multi-Turn13931387
EQ-Bench Creative Writing2007—

Frequently asked questions

Is Grok 4.7 better than Hunyuan Turbos 20250226?

Grok 4.7 is the stronger model overall, scoring 53.1 to 41.3 on the Noometry Index.

Is Grok 4.7 or Hunyuan Turbos 20250226 better for coding?

Grok 4.7 scores higher on coding benchmarks: 58.0 versus 40.0 in the Noometry coding category.

How many benchmarks do Grok 4.7 and Hunyuan Turbos 20250226 share?

13 benchmarks have published results for both models. Grok 4.7 has 39 scored results on Noometry and Hunyuan Turbos 20250226 has 16.

Related comparisons

Go deeper