VRAM · Local model
DeepSeek R1 (671B) VRAM requirements
DeepSeek R1 (671B) has 671.0 billion parameters. With an 8K context it needs about 421 GB of memory at Q4_K_M, and 1375 GB at full precision.
- Parameters
- 671.0B
- Layers
- 61
- Max context
- 163,840
- Experts
- 256 (8 active)
Attention: 61 layers with multi-head latent attention (one compressed 576-value vector per token).
Requirements
VRAM by quantisation and context length
| Quantisation | 4K context | 32K context | 128K context | 160K context |
|---|---|---|---|---|
| FP16 / BF16 | 1375 GB | 1377 GB | 1384 GB | 1387 GB |
| FP8 | 688 GB | 690 GB | 697 GB | 699 GB |
| Q8_0 (8-bit) | 731 GB | 733 GB | 740 GB | 742 GB |
| Q6_K | 564 GB | 566 GB | 573 GB | 575 GB |
| Q5_K_M | 490 GB | 492 GB | 499 GB | 502 GB |
| Q4_K_M | 420 GB | 423 GB | 430 GB | 432 GB |
| MXFP4 | 365 GB | 368 GB | 375 GB | 377 GB |
| Q3_K_M | 344 GB | 346 GB | 353 GB | 355 GB |
| Q2_K | 272 GB | 274 GB | 281 GB | 283 GB |
Weights + FP16 KV cache + overhead (10%, at least 1 GB), batch 1. See the formula and assumptions.
Hardware
What can run it at Q4_K_M
No single GPU in our list has enough memory at Q4_K_M with an 8K context. No Mac can hold it at this setting with the default GPU memory limit.
It’s a mixture-of-experts model: all 256 experts must be in memory, but only 8 run for each token, so it generates faster than a dense model of the same size.
Calculator
Try other settings
671.0B parameters, 256 experts (8 used per token) · 163,840-token context · model card
GGUF; the most popular balance of size and quality.
Requests served at the same time.
- Weights
- 382 GB
- KV cache
- 0.5 GB
- Overhead
- 38 GB
- Mixture of experts: all 256 experts must be in memory, but only 8 run for each token, so it generates faster than a dense model of this size.
| GPU | Memory | Runs it? |
|---|---|---|
| GeForce RTX 4060 | 8 GB | No |
| GeForce RTX 3060 | 12 GB | No |
| GeForce RTX 4070 | 12 GB | No |
| GeForce RTX 5070 | 12 GB | No |
| GeForce RTX 4060 Ti 16GB | 16 GB | No |
| GeForce RTX 4080 Super | 16 GB | No |
| GeForce RTX 5060 Ti 16GB | 16 GB | No |
| GeForce RTX 5070 Ti | 16 GB | No |
| GeForce RTX 5080 | 16 GB | No |
| GeForce RTX 3090 | 24 GB | No |
| GeForce RTX 4090 | 24 GB | No |
| Radeon RX 7900 XTX | 24 GB | No |
| GeForce RTX 5090 | 32 GB | No |
| L4 | 24 GB | No |
| L40S | 48 GB | No |
| A100 80GB | 80 GB | With 6 GPUs |
| H100 80GB | 80 GB | With 6 GPUs |
| RTX PRO 6000 Blackwell | 96 GB | With 5 GPUs |
| H200 | 141 GB | With 3 GPUs |
| Instinct MI300X | 192 GB | With 3 GPUs |
Apple Silicon Macs (unified memory)
- 16 GB: doesn’t fit
- 24 GB: doesn’t fit
- 32 GB: doesn’t fit
- 36 GB: doesn’t fit
- 48 GB: doesn’t fit
- 64 GB: doesn’t fit
- 96 GB: doesn’t fit
- 128 GB: doesn’t fit
- 192 GB: doesn’t fit
- 256 GB: doesn’t fit
- 512 GB: fits after raising the GPU limit
Filled: fits as is. Outlined: fits after raising the GPU memory limit (leaving 8 GB for macOS). macOS gives the GPU two-thirds of memory up to 32 GB and three-quarters above by default.
FAQ
Frequently asked questions
How much VRAM does DeepSeek R1 (671B) need?
With an 8K context, about 421 GB at Q4_K_M, 731 GB at 8-bit (Q8_0) and 1375 GB at full 16-bit precision. That covers the weights, the KV cache and runtime overhead.
Can DeepSeek R1 (671B) run on a 24 GB GPU?
Not on a single 24 GB card: even at Q2_K it needs about 272 GB. At Q4_K_M you’d need 18 × 24 GB cards.
Can I run DeepSeek R1 (671B) on a Mac?
Not at Q4_K_M: it needs about 421 GB, more than the largest Mac can give the GPU by default.
How much memory does DeepSeek R1 (671B)’s context use?
Its KV cache grows by about 67 MB for every 1,000 tokens of context in FP16. At its full 163,840-token context the cache is about 11 GB. Quantising the cache to 8-bit roughly halves it.
Related