VRAM · Local model
gpt-oss 120B VRAM requirements
gpt-oss 120B has 116.8 billion parameters. With an 8K context it needs about 65 GB of memory with its official MXFP4 weights, which fits an 80 GB GPU such as the A100 80GB.
- Parameters
- 116.8B
- Layers
- 36
- Max context
- 131,072
- Experts
- 128 (4 active)
Attention: 18 full-attention layers, 18 sliding-window layers (last 128 tokens).
Requirements
VRAM by context length
| Quantisation | 4K context | 32K context | 128K context |
|---|---|---|---|
| MXFP4 (official file) | 65 GB | 66 GB | 70 GB |
gpt-oss 120B is released in MXFP4, and local builds of it stay about that size whatever their quantisation label, so these use the official file. Weights + FP16 KV cache + overhead (10%, at least 1 GB), batch 1. See the formula and assumptions.
Hardware
What can run it with its official MXFP4 weights
On one GPU: A100 80GB, H100 80GB, RTX PRO 6000 Blackwell, H200, Instinct MI300X. Split across consumer cards: 6 × GeForce RTX 3060, 6 × GeForce RTX 4070, 6 × GeForce RTX 5070, 5 × GeForce RTX 4060 Ti 16GB. On a Mac: 96 GB of unified memory or more.
It’s a mixture-of-experts model: all 128 experts must be in memory, but only 4 run for each token, so it generates faster than a dense model of the same size.
Calculator
Try other settings
116.8B parameters, 128 experts (4 used per token) · 131,072-token context · model card
gpt-oss 120B is released in MXFP4, and local builds of it stay about that size whatever their quantisation label, so this uses the official file (59.0 GB of weights).
Requests served at the same time.
- Weights
- 59 GB
- KV cache
- 0.3 GB
- Overhead
- 5.9 GB
- Mixture of experts: all 128 experts must be in memory, but only 4 run for each token, so it generates faster than a dense model of this size.
- Some layers only look at the last 128 tokens, so the cache grows more slowly with context.
| GPU | Memory | Runs it? |
|---|---|---|
| GeForce RTX 4060 | 8 GB | No |
| GeForce RTX 3060 | 12 GB | With 6 GPUs |
| GeForce RTX 4070 | 12 GB | With 6 GPUs |
| GeForce RTX 5070 | 12 GB | With 6 GPUs |
| GeForce RTX 4060 Ti 16GB | 16 GB | With 5 GPUs |
| GeForce RTX 4080 Super | 16 GB | With 5 GPUs |
| GeForce RTX 5060 Ti 16GB | 16 GB | With 5 GPUs |
| GeForce RTX 5070 Ti | 16 GB | With 5 GPUs |
| GeForce RTX 5080 | 16 GB | With 5 GPUs |
| GeForce RTX 3090 | 24 GB | With 3 GPUs |
| GeForce RTX 4090 | 24 GB | With 3 GPUs |
| Radeon RX 7900 XTX | 24 GB | With 3 GPUs |
| GeForce RTX 5090 | 32 GB | With 3 GPUs |
| L4 | 24 GB | With 3 GPUs |
| L40S | 48 GB | With 2 GPUs |
| A100 80GB | 80 GB | Yes |
| H100 80GB | 80 GB | Yes |
| RTX PRO 6000 Blackwell | 96 GB | Yes |
| H200 | 141 GB | Yes |
| Instinct MI300X | 192 GB | Yes |
Apple Silicon Macs (unified memory)
- 16 GB: doesn’t fit
- 24 GB: doesn’t fit
- 32 GB: doesn’t fit
- 36 GB: doesn’t fit
- 48 GB: doesn’t fit
- 64 GB: doesn’t fit
- 96 GB: fits
- 128 GB: fits
- 192 GB: fits
- 256 GB: fits
- 512 GB: fits
Filled: fits as is. Outlined: fits after raising the GPU memory limit (leaving 8 GB for macOS). macOS gives the GPU two-thirds of memory up to 32 GB and three-quarters above by default.
FAQ
Frequently asked questions
How much VRAM does gpt-oss 120B need?
With an 8K context, about 65 GB using the official MXFP4 file (59 GB of weights). Local builds of gpt-oss 120B stay about that size whatever their quantisation label. That covers the weights, the KV cache and runtime overhead.
Can gpt-oss 120B run on a 24 GB GPU?
Not on a single 24 GB card: with its official MXFP4 weights it needs about 65 GB even with an 8K context, so you’d need 3 × 24 GB cards.
Can I run gpt-oss 120B on a Mac?
Yes: with its official MXFP4 weights and an 8K context it fits a Mac with 96 GB of unified memory using macOS’s default GPU memory limit.
How much memory does gpt-oss 120B’s context use?
Its KV cache grows by about 35 MB for every 1,000 tokens of context in FP16 (its sliding-window or chunked layers stop growing once their window is full). At its full 131,072-token context the cache is about 4.5 GB. Quantising the cache to 8-bit roughly halves it.
Related
Similar-sized models
- Llama 4 Scout70 GB
- Qwen2.5 72B48 GB
- Llama 3.3 70B47 GB
- Qwen3 235B-A22B149 GB
- Qwen2.5 Coder 32B23 GB
- DeepSeek R1 Distill Qwen 32B23 GB
Prefer the API? See gpt-oss-120b pricing.