Model comparison
Gemma 3 1B vs Llama 3.2 90B
Llama 3.2 90B is the stronger model overall, scoring 27.5 to 21.1 on the Noometry Index.
Last verified . 2 shared benchmarks.
Summary
- They share 2 benchmarks with published results for both. Gemma 3 1B scores higher in 0 categories and Llama 3.2 90B in 4 categories; 3 gaps are clear of the uncertainty.
- The widest gap is in agentic & tool use, where Llama 3.2 90B leads 30.0 to 13.7.
- The biggest single-benchmark swing is GPQA Diamond: 19.9% for Gemma 3 1B and 41% for Llama 3.2 90B.
Side by side
| Gemma 3 1B | Llama 3.2 90B | |
|---|---|---|
| Provider | Meta | |
| Noometry Index | 21.1 | 27.5 |
| Released | 2025-03-12 | 2024-09-24 |
| Weights | Open | Open |
| Context window | — | — |
| Max output | — | — |
| Input $ / M tokens | — | — |
| Output $ / M tokens | — | — |
| Results tracked | 4 | 9 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Agentic & Tool Use Llama 3.2 90B leads
Gemma 3 1B: 13.7 (#152), Llama 3.2 90B: 30.0 (#80)
| Benchmark | Gemma 3 1B | Llama 3.2 90B |
|---|---|---|
| Berkeley Function Calling Leaderboard | 7.2% | — |
| BALROG | — | 27.3% |
Reasoning Llama 3.2 90B leads
Gemma 3 1B: 19.2 (#264), Llama 3.2 90B: 21.7 (#217)
| Benchmark | Gemma 3 1B | Llama 3.2 90B |
|---|---|---|
| Chess Puzzles | 0% | — |
| EnigmaEval | — | 0.4% |
| Epoch Capabilities Index | — | 125.5 |
Math Too close to call
Gemma 3 1B: 10.2 (#316), Llama 3.2 90B: 11.1 (#308)
| Benchmark | Gemma 3 1B | Llama 3.2 90B |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 1.1% | 2.6% |
| MATH Level 5 | — | 39.4% |
Knowledge Llama 3.2 90B leads
Gemma 3 1B: 7.0 (#314), Llama 3.2 90B: 21.7 (#274)
| Benchmark | Gemma 3 1B | Llama 3.2 90B |
|---|---|---|
| GPQA Diamond | 19.9% | 41% |
| MMLU | — | 80.3% |
Multimodal Not comparable
Gemma 3 1B: —, Llama 3.2 90B: 25.4 (#124)
| Benchmark | Gemma 3 1B | Llama 3.2 90B |
|---|---|---|
| LMArena Vision | — | 1000 |
| GeoBench | — | 52% |
Frequently asked questions
Is Gemma 3 1B better than Llama 3.2 90B?
Llama 3.2 90B is the stronger model overall, scoring 27.5 to 21.1 on the Noometry Index.
How many benchmarks do Gemma 3 1B and Llama 3.2 90B share?
2 benchmarks have published results for both models. Gemma 3 1B has 4 scored results on Noometry and Llama 3.2 90B has 9.