VRAM · Local model
Qwen3 30B-A3B VRAM requirements
Qwen3 30B-A3B has 30.5 billion parameters. With an 8K context it needs about 20 GB of memory at Q4_K_M, which fits a 24 GB GPU such as the GeForce RTX 3090, and 63 GB at full precision.
- Parameters
- 30.5B
- Layers
- 48
- Max context
- 40,960
- Experts
- 128 (8 active)
Attention: 48 full-attention layers.
Requirements
VRAM by quantisation and context length
| Quantisation | 4K context | 32K context | 40K context |
|---|---|---|---|
| FP16 / BF16 | 63 GB | 66 GB | 67 GB |
| FP8 | 32 GB | 35 GB | 35 GB |
| Q8_0 (8-bit) | 34 GB | 37 GB | 37 GB |
| Q6_K | 26 GB | 29 GB | 30 GB |
| Q5_K_M | 23 GB | 26 GB | 26 GB |
| Q4_K_M | 20 GB | 22 GB | 23 GB |
| MXFP4 | 17 GB | 20 GB | 21 GB |
| Q3_K_M | 16 GB | 19 GB | 20 GB |
| Q2_K | 13 GB | 16 GB | 16 GB |
Weights + FP16 KV cache + overhead (10%, at least 1 GB), batch 1. See the formula and assumptions.
Hardware
What can run it at Q4_K_M
On one GPU: GeForce RTX 3090, GeForce RTX 4090, Radeon RX 7900 XTX, GeForce RTX 5090, L4, L40S, A100 80GB, H100 80GB, RTX PRO 6000 Blackwell, H200, Instinct MI300X. Split across consumer cards: 3 × GeForce RTX 4060, 2 × GeForce RTX 3060, 2 × GeForce RTX 4070, 2 × GeForce RTX 5070. On a Mac: 32 GB of unified memory or more.
It’s a mixture-of-experts model: all 128 experts must be in memory, but only 8 run for each token, so it generates faster than a dense model of the same size.
Calculator
Try other settings
30.5B parameters, 128 experts (8 used per token) · 40,960-token context · model card
GGUF; the most popular balance of size and quality.
Requests served at the same time.
- Weights
- 17 GB
- KV cache
- 0.8 GB
- Overhead
- 1.8 GB
- Mixture of experts: all 128 experts must be in memory, but only 8 run for each token, so it generates faster than a dense model of this size.
| GPU | Memory | Runs it? |
|---|---|---|
| GeForce RTX 4060 | 8 GB | With 3 GPUs |
| GeForce RTX 3060 | 12 GB | With 2 GPUs |
| GeForce RTX 4070 | 12 GB | With 2 GPUs |
| GeForce RTX 5070 | 12 GB | With 2 GPUs |
| GeForce RTX 4060 Ti 16GB | 16 GB | With 2 GPUs |
| GeForce RTX 4080 Super | 16 GB | With 2 GPUs |
| GeForce RTX 5060 Ti 16GB | 16 GB | With 2 GPUs |
| GeForce RTX 5070 Ti | 16 GB | With 2 GPUs |
| GeForce RTX 5080 | 16 GB | With 2 GPUs |
| GeForce RTX 3090 | 24 GB | Yes |
| GeForce RTX 4090 | 24 GB | Yes |
| Radeon RX 7900 XTX | 24 GB | Yes |
| GeForce RTX 5090 | 32 GB | Yes |
| L4 | 24 GB | Yes |
| L40S | 48 GB | Yes |
| A100 80GB | 80 GB | Yes |
| H100 80GB | 80 GB | Yes |
| RTX PRO 6000 Blackwell | 96 GB | Yes |
| H200 | 141 GB | Yes |
| Instinct MI300X | 192 GB | Yes |
Apple Silicon Macs (unified memory)
- 16 GB: doesn’t fit
- 24 GB: doesn’t fit
- 32 GB: fits
- 36 GB: fits
- 48 GB: fits
- 64 GB: fits
- 96 GB: fits
- 128 GB: fits
- 192 GB: fits
- 256 GB: fits
- 512 GB: fits
Filled: fits as is. Outlined: fits after raising the GPU memory limit (leaving 8 GB for macOS). macOS gives the GPU two-thirds of memory up to 32 GB and three-quarters above by default.
FAQ
Frequently asked questions
How much VRAM does Qwen3 30B-A3B need?
With an 8K context, about 20 GB at Q4_K_M, 34 GB at 8-bit (Q8_0) and 63 GB at full 16-bit precision. That covers the weights, the KV cache and runtime overhead.
Can Qwen3 30B-A3B run on a 24 GB GPU?
Yes, at Q5_K_M or smaller, with an 8K context (about 23 GB). Longer contexts need more memory for the KV cache.
Can I run Qwen3 30B-A3B on a Mac?
Yes: at Q4_K_M and an 8K context it fits a Mac with 32 GB of unified memory using macOS’s default GPU memory limit.
How much memory does Qwen3 30B-A3B’s context use?
Its KV cache grows by about 94 MB for every 1,000 tokens of context in FP16. At its full 40,960-token context the cache is about 3.8 GB. Quantising the cache to 8-bit roughly halves it.
Related