Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

VRAM · Local model

Qwen3 235B-A22B VRAM requirements

Qwen3 235B-A22B has 235.1 billion parameters. With an 8K context it needs about 149 GB of memory at Q4_K_M, which fits a 192 GB GPU such as the Instinct MI300X, and 483 GB at full precision.

Parameters
235.1B
Layers
94
Max context
262,144
Experts
128 (8 active)

Attention: 94 full-attention layers.

Requirements

VRAM by quantisation and context length

Quantisation4K context32K context128K context256K context
FP16 / BF16482 GB488 GB508 GB533 GB
FP8242 GB247 GB267 GB293 GB
Q8_0 (8-bit)257 GB262 GB282 GB308 GB
Q6_K198 GB204 GB223 GB249 GB
Q5_K_M172 GB178 GB197 GB223 GB
Q4_K_M148 GB154 GB173 GB199 GB
MXFP4129 GB134 GB154 GB180 GB
Q3_K_M121 GB127 GB146 GB172 GB
Q2_K96 GB102 GB121 GB147 GB

Weights + FP16 KV cache + overhead (10%, at least 1 GB), batch 1. See the formula and assumptions.

Hardware

What can run it at Q4_K_M

On one GPU: Instinct MI300X. Split across consumer cards: 7 × GeForce RTX 3090, 7 × GeForce RTX 4090, 7 × Radeon RX 7900 XTX, 5 × GeForce RTX 5090. On a Mac: 256 GB of unified memory or more.

It’s a mixture-of-experts model: all 128 experts must be in memory, but only 8 run for each token, so it generates faster than a dense model of the same size.

Calculator

Try other settings

235.1B parameters, 128 experts (8 used per token) · 262,144-token context · model card

GGUF; the most popular balance of size and quality.

Requests served at the same time.

Estimated memory needed149 GB
Weights
134 GB
KV cache
1.5 GB
Overhead
14 GB
  • Mixture of experts: all 128 experts must be in memory, but only 8 run for each token, so it generates faster than a dense model of this size.
GPUMemoryRuns it?
GeForce RTX 40608 GBNo
GeForce RTX 306012 GBNo
GeForce RTX 407012 GBNo
GeForce RTX 507012 GBNo
GeForce RTX 4060 Ti 16GB16 GBNo
GeForce RTX 4080 Super16 GBNo
GeForce RTX 5060 Ti 16GB16 GBNo
GeForce RTX 5070 Ti16 GBNo
GeForce RTX 508016 GBNo
GeForce RTX 309024 GBWith 7 GPUs
GeForce RTX 409024 GBWith 7 GPUs
Radeon RX 7900 XTX24 GBWith 7 GPUs
GeForce RTX 509032 GBWith 5 GPUs
L424 GBWith 7 GPUs
L40S48 GBWith 4 GPUs
A100 80GB80 GBWith 2 GPUs
H100 80GB80 GBWith 2 GPUs
RTX PRO 6000 Blackwell96 GBWith 2 GPUs
H200141 GBWith 2 GPUs
Instinct MI300X192 GBYes

Apple Silicon Macs (unified memory)

  • 16 GB: doesn’t fit
  • 24 GB: doesn’t fit
  • 32 GB: doesn’t fit
  • 36 GB: doesn’t fit
  • 48 GB: doesn’t fit
  • 64 GB: doesn’t fit
  • 96 GB: doesn’t fit
  • 128 GB: doesn’t fit
  • 192 GB: fits after raising the GPU limit
  • 256 GB: fits
  • 512 GB: fits

Filled: fits as is. Outlined: fits after raising the GPU memory limit (leaving 8 GB for macOS). macOS gives the GPU two-thirds of memory up to 32 GB and three-quarters above by default.

Estimate · formula below

FAQ

Frequently asked questions

How much VRAM does Qwen3 235B-A22B need?

With an 8K context, about 149 GB at Q4_K_M, 258 GB at 8-bit (Q8_0) and 483 GB at full 16-bit precision. That covers the weights, the KV cache and runtime overhead.

Can Qwen3 235B-A22B run on a 24 GB GPU?

Not on a single 24 GB card: even at Q2_K it needs about 97 GB. At Q4_K_M you’d need 7 × 24 GB cards.

Can I run Qwen3 235B-A22B on a Mac?

Yes: at Q4_K_M and an 8K context it fits a Mac with 256 GB of unified memory using macOS’s default GPU memory limit.

How much memory does Qwen3 235B-A22B’s context use?

Its KV cache grows by about 184 MB for every 1,000 tokens of context in FP16. At its full 262,144-token context the cache is about 47 GB. Quantising the cache to 8-bit roughly halves it.

Related