Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

VRAM · Local model

gpt-oss 20B VRAM requirements

gpt-oss 20B has 20.9 billion parameters. With an 8K context it needs about 13 GB of memory with its official MXFP4 weights, which fits a 16 GB GPU such as the GeForce RTX 4060 Ti 16GB.

Parameters
20.9B
Layers
24
Max context
131,072
Experts
32 (4 active)

Attention: 12 full-attention layers, 12 sliding-window layers (last 128 tokens).

Requirements

VRAM by context length

Quantisation4K context32K context128K context
MXFP4 (official file)13 GB13 GB16 GB

gpt-oss 20B is released in MXFP4, and local builds of it stay about that size whatever their quantisation label, so these use the official file. Weights + FP16 KV cache + overhead (10%, at least 1 GB), batch 1. See the formula and assumptions.

Hardware

What can run it with its official MXFP4 weights

On one GPU: GeForce RTX 4060 Ti 16GB, GeForce RTX 4080 Super, GeForce RTX 5060 Ti 16GB, GeForce RTX 5070 Ti, GeForce RTX 5080, GeForce RTX 3090, GeForce RTX 4090, Radeon RX 7900 XTX, GeForce RTX 5090, L4, L40S, A100 80GB, H100 80GB, RTX PRO 6000 Blackwell, H200, Instinct MI300X. Split across consumer cards: 2 × GeForce RTX 4060, 2 × GeForce RTX 3060, 2 × GeForce RTX 4070, 2 × GeForce RTX 5070. On a Mac: 24 GB of unified memory or more.

It’s a mixture-of-experts model: all 32 experts must be in memory, but only 4 run for each token, so it generates faster than a dense model of the same size.

Calculator

Try other settings

20.9B parameters, 32 experts (4 used per token) · 131,072-token context · model card

gpt-oss 20B is released in MXFP4, and local builds of it stay about that size whatever their quantisation label, so this uses the official file (11.3 GB of weights).

Requests served at the same time.

Estimated memory needed13 GB
Weights
11 GB
KV cache
0.2 GB
Overhead
1.1 GB
  • Mixture of experts: all 32 experts must be in memory, but only 4 run for each token, so it generates faster than a dense model of this size.
  • Some layers only look at the last 128 tokens, so the cache grows more slowly with context.
GPUMemoryRuns it?
GeForce RTX 40608 GBWith 2 GPUs
GeForce RTX 306012 GBWith 2 GPUs
GeForce RTX 407012 GBWith 2 GPUs
GeForce RTX 507012 GBWith 2 GPUs
GeForce RTX 4060 Ti 16GB16 GBYes
GeForce RTX 4080 Super16 GBYes
GeForce RTX 5060 Ti 16GB16 GBYes
GeForce RTX 5070 Ti16 GBYes
GeForce RTX 508016 GBYes
GeForce RTX 309024 GBYes
GeForce RTX 409024 GBYes
Radeon RX 7900 XTX24 GBYes
GeForce RTX 509032 GBYes
L424 GBYes
L40S48 GBYes
A100 80GB80 GBYes
H100 80GB80 GBYes
RTX PRO 6000 Blackwell96 GBYes
H200141 GBYes
Instinct MI300X192 GBYes

Apple Silicon Macs (unified memory)

  • 16 GB: doesn’t fit
  • 24 GB: fits
  • 32 GB: fits
  • 36 GB: fits
  • 48 GB: fits
  • 64 GB: fits
  • 96 GB: fits
  • 128 GB: fits
  • 192 GB: fits
  • 256 GB: fits
  • 512 GB: fits

Filled: fits as is. Outlined: fits after raising the GPU memory limit (leaving 8 GB for macOS). macOS gives the GPU two-thirds of memory up to 32 GB and three-quarters above by default.

Estimate · formula below

FAQ

Frequently asked questions

How much VRAM does gpt-oss 20B need?

With an 8K context, about 13 GB using the official MXFP4 file (11 GB of weights). Local builds of gpt-oss 20B stay about that size whatever their quantisation label. That covers the weights, the KV cache and runtime overhead.

Can gpt-oss 20B run on a 24 GB GPU?

Yes: with its official MXFP4 weights and an 8K context it needs about 13 GB. Longer contexts need more memory for the KV cache.

Can I run gpt-oss 20B on a Mac?

Yes: with its official MXFP4 weights and an 8K context it fits a Mac with 24 GB of unified memory using macOS’s default GPU memory limit.

How much memory does gpt-oss 20B’s context use?

Its KV cache grows by about 23 MB for every 1,000 tokens of context in FP16 (its sliding-window or chunked layers stop growing once their window is full). At its full 131,072-token context the cache is about 3.0 GB. Quantising the cache to 8-bit roughly halves it.

Related