Model comparison

Llama 13b vs Nemotron 3.5 Lightning

Nemotron 3.5 Lightning is the stronger model overall, scoring 40.0 to 24.4 on the Noometry Index.

Last verified . 8 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Nemotron 3.5 Lightning NVIDIA

40.0

Rank #155 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Llama 13b scores higher in 0 categories and Nemotron 3.5 Lightning in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Nemotron 3.5 Lightning leads 48.5 to 13.8.

Side by side

Llama 13b and Nemotron 3.5 Lightning specifications
Llama 13bNemotron 3.5 Lightning
ProviderMetaNVIDIA
Noometry Index24.440.0
Released2023-02-242026-08-11
WeightsOpenOpen
Context window—262K
Max output—262K
Input $ / M tokens—$0.05
Output $ / M tokens—$0.20
Results tracked2118

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Nemotron 3.5 Lightning leads

Llama 13b: 21.4 (#337), Nemotron 3.5 Lightning: 40.4 (#141)

Coding benchmarks
BenchmarkLlama 13bNemotron 3.5 Lightning
LMArena Coding6831375

Reasoning Nemotron 3.5 Lightning leads

Llama 13b: 14.0 (#329), Nemotron 3.5 Lightning: 26.8 (#127)

Reasoning benchmarks
BenchmarkLlama 13bNemotron 3.5 Lightning
LMArena Hard Prompts7281337
BIG-Bench Hard37.9%—
Epoch Capabilities Index100.58—
HellaSwag79.2%—
LAMBADA75.2%—
PIQA80.1%—
WinoGrande73%—

Math Nemotron 3.5 Lightning leads

Llama 13b: 26.7 (#256), Nemotron 3.5 Lightning: 37.5 (#155)

Math benchmarks
BenchmarkLlama 13bNemotron 3.5 Lightning
LMArena Math8381359
GSM8K20.6%—

Knowledge Not comparable

Llama 13b: —, Nemotron 3.5 Lightning: 37.5 (#154)

Knowledge benchmarks
BenchmarkLlama 13bNemotron 3.5 Lightning
LMArena Expert—1356
ARC (AI2) Challenge52.7%—
BoolQ78.7%—
MMLU47.7%—
OpenBookQA56.4%—
TriviaQA77.9%—

Multimodal Not comparable

Llama 13b: —, Nemotron 3.5 Lightning: —

Multimodal benchmarks
BenchmarkLlama 13bNemotron 3.5 Lightning
ScienceQA43.3%—

Multilingual Nemotron 3.5 Lightning leads

Llama 13b: 16.6 (#297), Nemotron 3.5 Lightning: 44.0 (#180)

Multilingual benchmarks
BenchmarkLlama 13bNemotron 3.5 Lightning
LMArena Non-English8191295
LMArena Chinese—1359
LMArena French—1366
LMArena German—1282
LMArena Japanese—1206
LMArena Korean—1238
LMArena Russian—1253
LMArena Spanish—1345

Instruction Following Nemotron 3.5 Lightning leads

Llama 13b: 36.7 (#305), Nemotron 3.5 Lightning: 69.6 (#170)

Instruction Following benchmarks
BenchmarkLlama 13bNemotron 3.5 Lightning
LMArena Instruction Following7811318

Long Context Not comparable

Llama 13b: —, Nemotron 3.5 Lightning: 39.9 (#165)

Long Context benchmarks
BenchmarkLlama 13bNemotron 3.5 Lightning
LMArena Longer Query—1314

Writing & Preference Nemotron 3.5 Lightning leads

Llama 13b: 13.8 (#312), Nemotron 3.5 Lightning: 48.5 (#201)

Writing & Preference benchmarks
BenchmarkLlama 13bNemotron 3.5 Lightning
LMArena Text8341327
LMArena Creative Writing7941254
LMArena Multi-Turn7531328
EQ-Bench Creative Writing—1280

Frequently asked questions

Is Llama 13b better than Nemotron 3.5 Lightning?

Nemotron 3.5 Lightning is the stronger model overall, scoring 40.0 to 24.4 on the Noometry Index.

Is Llama 13b or Nemotron 3.5 Lightning better for coding?

Nemotron 3.5 Lightning scores higher on coding benchmarks: 40.4 versus 21.4 in the Noometry coding category.

How many benchmarks do Llama 13b and Nemotron 3.5 Lightning share?

8 benchmarks have published results for both models. Llama 13b has 21 scored results on Noometry and Nemotron 3.5 Lightning has 18.

Related comparisons

Go deeper