VRAM · Local model
Llama 4 Scout VRAM requirements
Llama 4 Scout has 108.6 billion parameters. With an 8K context it needs about 70 GB of memory at Q4_K_M, which fits an 80 GB GPU such as the A100 80GB, and 224 GB at full precision.
- Parameters
- 108.6B
- Layers
- 48
- Max context
- 10,485,760
- Experts
- 16 (1 active)
Attention: 12 full-attention layers, 36 chunked-attention layers (8,192-token chunks).
Requirements
VRAM by quantisation and context length
| Quantisation | 4K context | 32K context | 128K context |
|---|---|---|---|
| FP16 / BF16 | 223 GB | 225 GB | 230 GB |
| FP8 | 112 GB | 114 GB | 119 GB |
| Q8_0 (8-bit) | 119 GB | 121 GB | 126 GB |
| Q6_K | 92 GB | 94 GB | 99 GB |
| Q5_K_M | 80 GB | 82 GB | 87 GB |
| Q4_K_M | 69 GB | 71 GB | 76 GB |
| MXFP4 | 60 GB | 62 GB | 67 GB |
| Q3_K_M | 56 GB | 59 GB | 63 GB |
| Q2_K | 45 GB | 47 GB | 52 GB |
Weights + FP16 KV cache + overhead (10%, at least 1 GB), batch 1. See the formula and assumptions.
Hardware
What can run it at Q4_K_M
On one GPU: A100 80GB, H100 80GB, RTX PRO 6000 Blackwell, H200, Instinct MI300X. Split across consumer cards: 6 × GeForce RTX 3060, 6 × GeForce RTX 4070, 6 × GeForce RTX 5070, 5 × GeForce RTX 4060 Ti 16GB. On a Mac: 96 GB of unified memory or more.
It’s a mixture-of-experts model: all 16 experts must be in memory, but only 1 run for each token, so it generates faster than a dense model of the same size.
Calculator
Try other settings
108.6B parameters, 16 experts (1 used per token) · 10,485,760-token context · model card
GGUF; the most popular balance of size and quality.
Requests served at the same time.
- Weights
- 62 GB
- KV cache
- 1.5 GB
- Overhead
- 6.3 GB
- Mixture of experts: all 16 experts must be in memory, but only 1 run for each token, so it generates faster than a dense model of this size.
- Some layers only look at the last 8,192 tokens, so the cache grows more slowly with context.
| GPU | Memory | Runs it? |
|---|---|---|
| GeForce RTX 4060 | 8 GB | No |
| GeForce RTX 3060 | 12 GB | With 6 GPUs |
| GeForce RTX 4070 | 12 GB | With 6 GPUs |
| GeForce RTX 5070 | 12 GB | With 6 GPUs |
| GeForce RTX 4060 Ti 16GB | 16 GB | With 5 GPUs |
| GeForce RTX 4080 Super | 16 GB | With 5 GPUs |
| GeForce RTX 5060 Ti 16GB | 16 GB | With 5 GPUs |
| GeForce RTX 5070 Ti | 16 GB | With 5 GPUs |
| GeForce RTX 5080 | 16 GB | With 5 GPUs |
| GeForce RTX 3090 | 24 GB | With 3 GPUs |
| GeForce RTX 4090 | 24 GB | With 3 GPUs |
| Radeon RX 7900 XTX | 24 GB | With 3 GPUs |
| GeForce RTX 5090 | 32 GB | With 3 GPUs |
| L4 | 24 GB | With 3 GPUs |
| L40S | 48 GB | With 2 GPUs |
| A100 80GB | 80 GB | Yes |
| H100 80GB | 80 GB | Yes |
| RTX PRO 6000 Blackwell | 96 GB | Yes |
| H200 | 141 GB | Yes |
| Instinct MI300X | 192 GB | Yes |
Apple Silicon Macs (unified memory)
- 16 GB: doesn’t fit
- 24 GB: doesn’t fit
- 32 GB: doesn’t fit
- 36 GB: doesn’t fit
- 48 GB: doesn’t fit
- 64 GB: doesn’t fit
- 96 GB: fits
- 128 GB: fits
- 192 GB: fits
- 256 GB: fits
- 512 GB: fits
Filled: fits as is. Outlined: fits after raising the GPU memory limit (leaving 8 GB for macOS). macOS gives the GPU two-thirds of memory up to 32 GB and three-quarters above by default.
FAQ
Frequently asked questions
How much VRAM does Llama 4 Scout need?
With an 8K context, about 70 GB at Q4_K_M, 120 GB at 8-bit (Q8_0) and 224 GB at full 16-bit precision. That covers the weights, the KV cache and runtime overhead.
Can Llama 4 Scout run on a 24 GB GPU?
Not on a single 24 GB card: even at Q2_K it needs about 46 GB. At Q4_K_M you’d need 3 × 24 GB cards.
Can I run Llama 4 Scout on a Mac?
Yes: at Q4_K_M and an 8K context it fits a Mac with 96 GB of unified memory using macOS’s default GPU memory limit.
How much memory does Llama 4 Scout’s context use?
Its KV cache grows by about 47 MB for every 1,000 tokens of context in FP16 (its sliding-window or chunked layers stop growing once their window is full). At its full 10,485,760-token context the cache is about 481 GB. Quantising the cache to 8-bit roughly halves it.
Related
Similar-sized models
- gpt-oss 120B65 GB
- Qwen2.5 72B48 GB
- Llama 3.3 70B47 GB
- Qwen3 235B-A22B149 GB
- Qwen2.5 Coder 32B23 GB
- DeepSeek R1 Distill Qwen 32B23 GB
Prefer the API? See Llama 4 Scout pricing.