VRAM · Local model
Qwen3 235B-A22B VRAM requirements
Qwen3 235B-A22B has 235.1 billion parameters. With an 8K context it needs about 149 GB of memory at Q4_K_M, which fits a 192 GB GPU such as the Instinct MI300X, and 483 GB at full precision.
- Parameters
- 235.1B
- Layers
- 94
- Max context
- 262,144
- Experts
- 128 (8 active)
Attention: 94 full-attention layers.
Requirements
VRAM by quantisation and context length
| Quantisation | 4K context | 32K context | 128K context | 256K context |
|---|---|---|---|---|
| FP16 / BF16 | 482 GB | 488 GB | 508 GB | 533 GB |
| FP8 | 242 GB | 247 GB | 267 GB | 293 GB |
| Q8_0 (8-bit) | 257 GB | 262 GB | 282 GB | 308 GB |
| Q6_K | 198 GB | 204 GB | 223 GB | 249 GB |
| Q5_K_M | 172 GB | 178 GB | 197 GB | 223 GB |
| Q4_K_M | 148 GB | 154 GB | 173 GB | 199 GB |
| MXFP4 | 129 GB | 134 GB | 154 GB | 180 GB |
| Q3_K_M | 121 GB | 127 GB | 146 GB | 172 GB |
| Q2_K | 96 GB | 102 GB | 121 GB | 147 GB |
Weights + FP16 KV cache + overhead (10%, at least 1 GB), batch 1. See the formula and assumptions.
Hardware
What can run it at Q4_K_M
On one GPU: Instinct MI300X. Split across consumer cards: 7 × GeForce RTX 3090, 7 × GeForce RTX 4090, 7 × Radeon RX 7900 XTX, 5 × GeForce RTX 5090. On a Mac: 256 GB of unified memory or more.
It’s a mixture-of-experts model: all 128 experts must be in memory, but only 8 run for each token, so it generates faster than a dense model of the same size.
Calculator
Try other settings
235.1B parameters, 128 experts (8 used per token) · 262,144-token context · model card
GGUF; the most popular balance of size and quality.
Requests served at the same time.
- Weights
- 134 GB
- KV cache
- 1.5 GB
- Overhead
- 14 GB
- Mixture of experts: all 128 experts must be in memory, but only 8 run for each token, so it generates faster than a dense model of this size.
| GPU | Memory | Runs it? |
|---|---|---|
| GeForce RTX 4060 | 8 GB | No |
| GeForce RTX 3060 | 12 GB | No |
| GeForce RTX 4070 | 12 GB | No |
| GeForce RTX 5070 | 12 GB | No |
| GeForce RTX 4060 Ti 16GB | 16 GB | No |
| GeForce RTX 4080 Super | 16 GB | No |
| GeForce RTX 5060 Ti 16GB | 16 GB | No |
| GeForce RTX 5070 Ti | 16 GB | No |
| GeForce RTX 5080 | 16 GB | No |
| GeForce RTX 3090 | 24 GB | With 7 GPUs |
| GeForce RTX 4090 | 24 GB | With 7 GPUs |
| Radeon RX 7900 XTX | 24 GB | With 7 GPUs |
| GeForce RTX 5090 | 32 GB | With 5 GPUs |
| L4 | 24 GB | With 7 GPUs |
| L40S | 48 GB | With 4 GPUs |
| A100 80GB | 80 GB | With 2 GPUs |
| H100 80GB | 80 GB | With 2 GPUs |
| RTX PRO 6000 Blackwell | 96 GB | With 2 GPUs |
| H200 | 141 GB | With 2 GPUs |
| Instinct MI300X | 192 GB | Yes |
Apple Silicon Macs (unified memory)
- 16 GB: doesn’t fit
- 24 GB: doesn’t fit
- 32 GB: doesn’t fit
- 36 GB: doesn’t fit
- 48 GB: doesn’t fit
- 64 GB: doesn’t fit
- 96 GB: doesn’t fit
- 128 GB: doesn’t fit
- 192 GB: fits after raising the GPU limit
- 256 GB: fits
- 512 GB: fits
Filled: fits as is. Outlined: fits after raising the GPU memory limit (leaving 8 GB for macOS). macOS gives the GPU two-thirds of memory up to 32 GB and three-quarters above by default.
FAQ
Frequently asked questions
How much VRAM does Qwen3 235B-A22B need?
With an 8K context, about 149 GB at Q4_K_M, 258 GB at 8-bit (Q8_0) and 483 GB at full 16-bit precision. That covers the weights, the KV cache and runtime overhead.
Can Qwen3 235B-A22B run on a 24 GB GPU?
Not on a single 24 GB card: even at Q2_K it needs about 97 GB. At Q4_K_M you’d need 7 × 24 GB cards.
Can I run Qwen3 235B-A22B on a Mac?
Yes: at Q4_K_M and an 8K context it fits a Mac with 256 GB of unified memory using macOS’s default GPU memory limit.
How much memory does Qwen3 235B-A22B’s context use?
Its KV cache grows by about 184 MB for every 1,000 tokens of context in FP16. At its full 262,144-token context the cache is about 47 GB. Quantising the cache to 8-bit roughly halves it.
Related