Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Local AI

What is VRAM, and why does it limit local LLMs?

Also called: video RAM, GPU memory, graphics memory

Definition

VRAM is the memory on a graphics card, and for running AI models locally it is the main limit: a model runs at full GPU speed only when its weights, KV cache and working buffers all fit in it.

Explained

How it works

VRAM (video RAM) sits on the graphics card itself, separate from your computer’s system RAM. A model runs fastest when everything it needs is there. When it doesn’t fit, llama.cpp and Ollama can keep some layers in system RAM and run them on the CPU, which works but is much slower.

Three things fill it: the weights (parameters × bits per weight ÷ 8), the KV cache for your context, and working space for the runtime. Our calculator estimates that last part as 10% of the other two, at least 1 GB, which is an assumption rather than a measurement.

Apple Silicon Macs have no separate VRAM: the CPU and GPU share unified memory. By default macOS lets the GPU use only part of it: Apple’s figures for M1 Pro and M1 Max MacBook Pros are 21 GB of 32 GB and 48 GB of 64 GB, and llama.cpp users report about two-thirds on Macs with up to 32 GB and three-quarters above. You can raise the limit with sudo sysctl iogpu.wired_limit_mb=…; our calculator then keeps 8 GB free for macOS.

Example

Qwen3 32B at Q4_K_M on one GPU

Qwen3 32B at Q4_K_M has 19 GB of weights. With an 8K context and a batch of one, the cache and overhead bring it to 23 GB, which fits a 24 GB card such as the GeForce RTX 3090. At 32K the cache alone is 8.0 GB and the 29 GB total needs a 32 GB card. A 16K context with an 8-bit cache fits 24 GB again. On a Mac, the 8K setting needs 36 GB of unified memory with the default limit.

Qwen3 32B at Q4_K_M: what fills the memory
SettingWeightsKV cacheOverheadTotalSmallest single GPU
8K context, FP16 cache18.7 GB2.0 GB2.1 GB22.7 GB24 GB (RTX 3090)
32K context, FP16 cache18.7 GB8.0 GB2.7 GB29.3 GB32 GB (RTX 5090)
16K context, q8_0 cache18.7 GB2.1 GB2.1 GB22.9 GB24 GB (RTX 3090)

From our VRAM calculator, batch 1. GPUs from the list it checks; leave some headroom for your desktop and other apps. Model config fetched 2026-10-08. GB = 1,024³ bytes, the way GPU memory is quoted.

Cost and quality

Why it matters

Memory, more than raw GPU speed, decides which models you can run locally: too little and the model either won’t load or spills onto the much slower CPU. That is why quantisation and context length matter so much: both are ways of fitting the same model into less memory.

When buying hardware, the memory size sets the largest model you can run. Check a specific model and context in our VRAM calculator before you buy.

Don’t mix up

Common confusions

VRAM vs system RAM
System RAM is the memory on your motherboard. A PC with 64 GB of RAM and an 8 GB graphics card can hold at most 8 GB of weights, cache and buffers on the GPU; any layers that don’t fit run on the CPU, which slows the whole model down.
Unified memory vs VRAM
A Mac’s memory is shared, so a 32 GB Mac doesn’t give the GPU 32 GB. With the default limit it gets about 21 GB, and macOS and your apps need the rest.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary