VRAM · Local model
DeepSeek R1 Distill Qwen 32B VRAM requirements
DeepSeek R1 Distill Qwen 32B has 32.8 billion parameters. With an 8K context it needs about 23 GB of memory at Q4_K_M, which fits a 24 GB GPU such as the GeForce RTX 3090, and 69 GB at full precision.
- Parameters
- 32.8B
- Layers
- 64
- Max context
- 131,072
- Experts
- Dense
Attention: 64 full-attention layers.
Requirements
VRAM by quantisation and context length
| Quantisation | 4K context | 32K context | 128K context |
|---|---|---|---|
| FP16 / BF16 | 68 GB | 76 GB | 102 GB |
| FP8 | 35 GB | 42 GB | 69 GB |
| Q8_0 (8-bit) | 37 GB | 44 GB | 71 GB |
| Q6_K | 29 GB | 36 GB | 63 GB |
| Q5_K_M | 25 GB | 33 GB | 59 GB |
| Q4_K_M | 22 GB | 29 GB | 56 GB |
| MXFP4 | 19 GB | 27 GB | 53 GB |
| Q3_K_M | 18 GB | 26 GB | 52 GB |
| Q2_K | 14 GB | 22 GB | 48 GB |
Weights + FP16 KV cache + overhead (10%, at least 1 GB), batch 1. See the formula and assumptions.
Hardware
What can run it at Q4_K_M
On one GPU: GeForce RTX 3090, GeForce RTX 4090, Radeon RX 7900 XTX, GeForce RTX 5090, L4, L40S, A100 80GB, H100 80GB, RTX PRO 6000 Blackwell, H200, Instinct MI300X. Split across consumer cards: 3 × GeForce RTX 4060, 2 × GeForce RTX 3060, 2 × GeForce RTX 4070, 2 × GeForce RTX 5070. On a Mac: 36 GB of unified memory or more.
Calculator
Try other settings
32.8B parameters · 131,072-token context · model card
GGUF; the most popular balance of size and quality.
Requests served at the same time.
- Weights
- 19 GB
- KV cache
- 2.0 GB
- Overhead
- 2.1 GB
| GPU | Memory | Runs it? |
|---|---|---|
| GeForce RTX 4060 | 8 GB | With 3 GPUs |
| GeForce RTX 3060 | 12 GB | With 2 GPUs |
| GeForce RTX 4070 | 12 GB | With 2 GPUs |
| GeForce RTX 5070 | 12 GB | With 2 GPUs |
| GeForce RTX 4060 Ti 16GB | 16 GB | With 2 GPUs |
| GeForce RTX 4080 Super | 16 GB | With 2 GPUs |
| GeForce RTX 5060 Ti 16GB | 16 GB | With 2 GPUs |
| GeForce RTX 5070 Ti | 16 GB | With 2 GPUs |
| GeForce RTX 5080 | 16 GB | With 2 GPUs |
| GeForce RTX 3090 | 24 GB | Yes |
| GeForce RTX 4090 | 24 GB | Yes |
| Radeon RX 7900 XTX | 24 GB | Yes |
| GeForce RTX 5090 | 32 GB | Yes |
| L4 | 24 GB | Yes |
| L40S | 48 GB | Yes |
| A100 80GB | 80 GB | Yes |
| H100 80GB | 80 GB | Yes |
| RTX PRO 6000 Blackwell | 96 GB | Yes |
| H200 | 141 GB | Yes |
| Instinct MI300X | 192 GB | Yes |
Apple Silicon Macs (unified memory)
- 16 GB: doesn’t fit
- 24 GB: doesn’t fit
- 32 GB: fits after raising the GPU limit
- 36 GB: fits
- 48 GB: fits
- 64 GB: fits
- 96 GB: fits
- 128 GB: fits
- 192 GB: fits
- 256 GB: fits
- 512 GB: fits
Filled: fits as is. Outlined: fits after raising the GPU memory limit (leaving 8 GB for macOS). macOS gives the GPU two-thirds of memory up to 32 GB and three-quarters above by default.
FAQ
Frequently asked questions
How much VRAM does DeepSeek R1 Distill Qwen 32B need?
With an 8K context, about 23 GB at Q4_K_M, 38 GB at 8-bit (Q8_0) and 69 GB at full 16-bit precision. That covers the weights, the KV cache and runtime overhead.
Can DeepSeek R1 Distill Qwen 32B run on a 24 GB GPU?
Yes, at Q4_K_M or smaller, with an 8K context (about 23 GB). Longer contexts need more memory for the KV cache.
Can I run DeepSeek R1 Distill Qwen 32B on a Mac?
Yes: at Q4_K_M and an 8K context it fits a Mac with 36 GB of unified memory using macOS’s default GPU memory limit.
How much memory does DeepSeek R1 Distill Qwen 32B’s context use?
Its KV cache grows by about 250 MB for every 1,000 tokens of context in FP16. At its full 131,072-token context the cache is about 32 GB. Quantising the cache to 8-bit roughly halves it.
Related