VRAM · Local model
Nemotron 3 Nano 30B-A3B VRAM requirements
Nemotron 3 Nano 30B-A3B has 31.6 billion parameters. With an 8K context it needs about 20 GB of memory at Q4_K_M, which fits a 24 GB GPU such as the GeForce RTX 3090, and 65 GB at full precision.
- Parameters
- 31.6B
- Layers
- 29
- Max context
- 262,144
- Experts
- 128 (6 active)
Attention: 6 full-attention layers, 23 Mamba or linear-attention layers.
Requirements
VRAM by quantisation and context length
| Quantisation | 4K context | 32K context | 128K context | 256K context |
|---|---|---|---|---|
| FP16 / BF16 | 65 GB | 65 GB | 66 GB | 66 GB |
| FP8 | 32 GB | 33 GB | 33 GB | 34 GB |
| Q8_0 (8-bit) | 34 GB | 35 GB | 35 GB | 36 GB |
| Q6_K | 27 GB | 27 GB | 27 GB | 28 GB |
| Q5_K_M | 23 GB | 23 GB | 24 GB | 25 GB |
| Q4_K_M | 20 GB | 20 GB | 21 GB | 21 GB |
| MXFP4 | 17 GB | 17 GB | 18 GB | 19 GB |
| Q3_K_M | 16 GB | 16 GB | 17 GB | 18 GB |
| Q2_K | 13 GB | 13 GB | 14 GB | 14 GB |
Weights + FP16 KV cache + overhead (10%, at least 1 GB), batch 1. See the formula and assumptions.
Hardware
What can run it at Q4_K_M
On one GPU: GeForce RTX 3090, GeForce RTX 4090, Radeon RX 7900 XTX, GeForce RTX 5090, L4, L40S, A100 80GB, H100 80GB, RTX PRO 6000 Blackwell, H200, Instinct MI300X. Split across consumer cards: 3 × GeForce RTX 4060, 2 × GeForce RTX 3060, 2 × GeForce RTX 4070, 2 × GeForce RTX 5070. On a Mac: 32 GB of unified memory or more.
It’s a mixture-of-experts model: all 128 experts must be in memory, but only 6 run for each token, so it generates faster than a dense model of the same size.
Calculator
Try other settings
31.6B parameters, 128 experts (6 used per token) · 262,144-token context · model card
GGUF; the most popular balance of size and quality.
Requests served at the same time.
- Weights
- 18 GB
- KV cache
- 0.0 GB
- Overhead
- 1.8 GB
- Mixture of experts: all 128 experts must be in memory, but only 6 run for each token, so it generates faster than a dense model of this size.
- 23 of its layers use Mamba or linear attention, which keep a small fixed state instead of a growing cache (counted in the overhead).
| GPU | Memory | Runs it? |
|---|---|---|
| GeForce RTX 4060 | 8 GB | With 3 GPUs |
| GeForce RTX 3060 | 12 GB | With 2 GPUs |
| GeForce RTX 4070 | 12 GB | With 2 GPUs |
| GeForce RTX 5070 | 12 GB | With 2 GPUs |
| GeForce RTX 4060 Ti 16GB | 16 GB | With 2 GPUs |
| GeForce RTX 4080 Super | 16 GB | With 2 GPUs |
| GeForce RTX 5060 Ti 16GB | 16 GB | With 2 GPUs |
| GeForce RTX 5070 Ti | 16 GB | With 2 GPUs |
| GeForce RTX 5080 | 16 GB | With 2 GPUs |
| GeForce RTX 3090 | 24 GB | Yes |
| GeForce RTX 4090 | 24 GB | Yes |
| Radeon RX 7900 XTX | 24 GB | Yes |
| GeForce RTX 5090 | 32 GB | Yes |
| L4 | 24 GB | Yes |
| L40S | 48 GB | Yes |
| A100 80GB | 80 GB | Yes |
| H100 80GB | 80 GB | Yes |
| RTX PRO 6000 Blackwell | 96 GB | Yes |
| H200 | 141 GB | Yes |
| Instinct MI300X | 192 GB | Yes |
Apple Silicon Macs (unified memory)
- 16 GB: doesn’t fit
- 24 GB: doesn’t fit
- 32 GB: fits
- 36 GB: fits
- 48 GB: fits
- 64 GB: fits
- 96 GB: fits
- 128 GB: fits
- 192 GB: fits
- 256 GB: fits
- 512 GB: fits
Filled: fits as is. Outlined: fits after raising the GPU memory limit (leaving 8 GB for macOS). macOS gives the GPU two-thirds of memory up to 32 GB and three-quarters above by default.
FAQ
Frequently asked questions
How much VRAM does Nemotron 3 Nano 30B-A3B need?
With an 8K context, about 20 GB at Q4_K_M, 34 GB at 8-bit (Q8_0) and 65 GB at full 16-bit precision. That covers the weights, the KV cache and runtime overhead.
Can Nemotron 3 Nano 30B-A3B run on a 24 GB GPU?
Yes, at Q5_K_M or smaller, with an 8K context (about 23 GB). Longer contexts need more memory for the KV cache.
Can I run Nemotron 3 Nano 30B-A3B on a Mac?
Yes: at Q4_K_M and an 8K context it fits a Mac with 32 GB of unified memory using macOS’s default GPU memory limit.
How much memory does Nemotron 3 Nano 30B-A3B’s context use?
Its KV cache grows by about 6 MB for every 1,000 tokens of context in FP16. At its full 262,144-token context the cache is about 1.5 GB. Quantising the cache to 8-bit roughly halves it.
Related