Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

Tokens & Costs

LLM VRAM calculator: can your GPU run it?

Work out how much GPU memory a local model needs (weights, KV cache and overhead) and which graphics cards or Macs can run it. Free, and it runs in your browser.

8.0B parameters · 131,072-token context · model card

GGUF; the most popular balance of size and quality.

Requests served at the same time.

Estimated memory needed6.6 GB
Weights
4.6 GB
KV cache
1.0 GB
Overhead
1.0 GB
    GPUMemoryRuns it?
    GeForce RTX 40608 GBYes
    GeForce RTX 306012 GBYes
    GeForce RTX 407012 GBYes
    GeForce RTX 507012 GBYes
    GeForce RTX 4060 Ti 16GB16 GBYes
    GeForce RTX 4080 Super16 GBYes
    GeForce RTX 5060 Ti 16GB16 GBYes
    GeForce RTX 5070 Ti16 GBYes
    GeForce RTX 508016 GBYes
    GeForce RTX 309024 GBYes
    GeForce RTX 409024 GBYes
    Radeon RX 7900 XTX24 GBYes
    GeForce RTX 509032 GBYes
    L424 GBYes
    L40S48 GBYes
    A100 80GB80 GBYes
    H100 80GB80 GBYes
    RTX PRO 6000 Blackwell96 GBYes
    H200141 GBYes
    Instinct MI300X192 GBYes

    Apple Silicon Macs (unified memory)

    • 16 GB: fits
    • 24 GB: fits
    • 32 GB: fits
    • 36 GB: fits
    • 48 GB: fits
    • 64 GB: fits
    • 96 GB: fits
    • 128 GB: fits
    • 192 GB: fits
    • 256 GB: fits
    • 512 GB: fits

    Filled: fits as is. Outlined: fits after raising the GPU memory limit (leaving 8 GB for macOS). macOS gives the GPU two-thirds of memory up to 32 GB and three-quarters above by default.

    Estimate · formula below

    Steps

    How to use the LLM VRAM calculator

    1. Pick a model, or choose “Custom size” and enter its parameter count.
    2. Choose the quantisation you plan to download, such as Q4_K_M.
    3. Set the context length you need, and the batch size if you serve several requests at once.
    4. Read the total and its breakdown, then check which GPUs and Mac memory sizes can run it.

    Method

    How it works

    Running a model locally needs memory for three things: the model’s weights, the KV cache that holds the conversation, and working space for the runtime. The calculator adds them up the same way the runtimes allocate them.

    The formula

    total = weights + KV cache + overhead

    • Weights = parameters × bits per weight ÷ 8. The bits per weight are llama.cpp’s measured figures, which include the scales that quantised formats store alongside the weights: for example 4.89 for Q4_K_M and 8.5 for Q8_0. They were measured on Llama 3.1 8B; mixed formats vary slightly between models.
    • KV cache = for each attention layer, the values it stores per token × the tokens it keeps × the batch size × bytes per value. A standard layer stores keys and values for each of its KV heads (2 × KV heads × head size). DeepSeek-style MLA layers store one compressed vector instead. Sliding-window and chunked layers only keep their last few thousand tokens, and Mamba or linear-attention layers keep a small fixed state rather than a cache.
    • Overhead = 10% of weights and cache, at least 1 GB, for the CUDA or Metal context, compute buffers and fragmentation. This is an assumption, not a measurement: the real figure depends on the runtime and its settings.

    Where the numbers come from

    Layer counts, KV heads, head sizes, attention windows and expert counts are read from each model’s official config.json on Hugging Face, pinned to a specific version, and parameter counts come from the published weights. Architectures the calculator can’t read reliably, such as the newest sparse-attention designs, are left out rather than estimated. GB here means 1,024³ bytes, the way GPU memory is quoted.

    Reading the result

    A model “fits” a GPU when the estimate is no larger than its memory. In practice, leave a little room: your desktop and other apps also use GPU memory, and long prompts need larger compute buffers. If you’re close to the limit, use a smaller quantisation, quantise the KV cache to 8-bit, or shorten the context. For a cost comparison with hosted models, see the LLM API cost calculator.

    Examples

    Worked examples

    Memory needed at an 8K context, batch 1

    ModelFP16 / BF16Q8_0 (8-bit)Q4_K_M
    Llama 3.1 8B18 GB9.9 GB6.6 GB
    Qwen3 14B32 GB17 GB11 GB
    Gemma 3 27B57 GB31 GB18 GB
    Qwen3 32B69 GB38 GB23 GB
    gpt-oss 20B13 GB13 GB13 GB
    Qwen3 30B-A3B63 GB34 GB20 GB
    Llama 3.3 70B147 GB80 GB47 GB
    gpt-oss 120B65 GB65 GB65 GB
    Qwen3 235B-A22B483 GB258 GB149 GB
    DeepSeek R1 (671B)1375 GB731 GB421 GB

    Weights + FP16 KV cache + overhead, from the formula above. Model configs fetched 2026-10-08.

    FAQ

    Frequently asked questions

    How much VRAM do I need for a 70B model?

    For Llama 3.3 70B with an 8K context, about 47 GB at Q4_K_M, 80 GB at 8-bit and 147 GB at full 16-bit precision. At Q4 that needs 2 × 24 GB cards or one 48 GB card, or a Mac with 64 GB of memory. Longer contexts add more: at 128K the cache alone is 40 GB.

    Can I run an 8B model on an 8 GB GPU?

    Yes, with quantisation. Llama 3.1 8B at Q4_K_M with an 8K context needs about 6.6 GB, which fits an 8 GB card. At full precision it needs about 18 GB, so a 24 GB card.

    What is quantisation, and how much quality does it cost?

    Quantisation stores each weight in fewer bits, for example about 4.9 bits instead of 16, which cuts memory by roughly two-thirds. 8-bit is close to lossless, Q5 and Q6 lose very little, and Q4_K_M is the usual sweet spot. Below 4 bits quality drops more quickly, especially for smaller models.

    Why does a longer context need more memory?

    For every token in the context, each attention layer keeps its keys and values in the KV cache, so the cache grows in proportion to the context length and the batch size. Models that use grouped-query attention, sliding windows or compressed (MLA) attention keep much smaller caches, which this calculator reads from each model’s configuration. Quantising the cache to 8-bit roughly halves it.

    Can I split a model across two GPUs?

    Yes. llama.cpp, vLLM and similar runtimes can split the layers across several GPUs, so two 24 GB cards hold about 48 GB of model. Each GPU needs a little extra for its own buffers, and splitting can be slower than one big card, but it is the usual way to run 70B models at home.

    What if the model doesn’t fit in VRAM?

    Runtimes such as llama.cpp and Ollama can keep some layers in normal system RAM and run them on the CPU. It works, but every layer on the CPU slows generation down considerably. A smaller quantisation, a shorter context or a smaller model usually gives a better experience.

    How much of a Mac’s memory can the GPU use?

    By default macOS lets the GPU use about two-thirds of unified memory on Macs with up to 32 GB and three-quarters on larger ones. You can raise the limit with the iogpu.wired_limit_mb setting; the calculator assumes you keep 8 GB for macOS when you do.

    Do mixture-of-experts models need less memory?

    No: every expert has to be loaded, so a 30B mixture-of-experts model needs as much memory as a 30B dense model. What it saves is compute: only a few experts run for each token, so it generates text much faster than a dense model of the same size.