Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

VRAM · Local model

Llama 3.2 3B VRAM requirements

Llama 3.2 3B has 3.2 billion parameters. With an 8K context it needs about 3.7 GB of memory at Q4_K_M, which fits an 8 GB GPU such as the GeForce RTX 4060, and 7.9 GB at full precision.

Parameters
3.2B
Layers
28
Max context
131,072
Experts
Dense

Attention: 28 full-attention layers.

Requirements

VRAM by quantisation and context length

Quantisation4K context32K context128K context
FP16 / BF167.4 GB10 GB22 GB
FP84.4 GB7.5 GB19 GB
Q8_0 (8-bit)4.6 GB7.7 GB19 GB
Q6_K3.9 GB7.0 GB18 GB
Q5_K_M3.6 GB6.6 GB18 GB
Q4_K_M3.3 GB6.3 GB17 GB
MXFP43.0 GB6.1 GB17 GB
Q3_K_M2.9 GB6.0 GB17 GB
Q2_K2.6 GB5.7 GB17 GB

Weights + FP16 KV cache + overhead (10%, at least 1 GB), batch 1. See the formula and assumptions.

Hardware

What can run it at Q4_K_M

On one GPU: GeForce RTX 4060, GeForce RTX 3060, GeForce RTX 4070, GeForce RTX 5070, GeForce RTX 4060 Ti 16GB, GeForce RTX 4080 Super, GeForce RTX 5060 Ti 16GB, GeForce RTX 5070 Ti, GeForce RTX 5080, GeForce RTX 3090, GeForce RTX 4090, Radeon RX 7900 XTX, GeForce RTX 5090, L4, L40S, A100 80GB, H100 80GB, RTX PRO 6000 Blackwell, H200, Instinct MI300X. On a Mac: 16 GB of unified memory or more.

Calculator

Try other settings

3.2B parameters · 131,072-token context · model card

GGUF; the most popular balance of size and quality.

Requests served at the same time.

Estimated memory needed3.7 GB
Weights
1.8 GB
KV cache
0.9 GB
Overhead
1.0 GB
    GPUMemoryRuns it?
    GeForce RTX 40608 GBYes
    GeForce RTX 306012 GBYes
    GeForce RTX 407012 GBYes
    GeForce RTX 507012 GBYes
    GeForce RTX 4060 Ti 16GB16 GBYes
    GeForce RTX 4080 Super16 GBYes
    GeForce RTX 5060 Ti 16GB16 GBYes
    GeForce RTX 5070 Ti16 GBYes
    GeForce RTX 508016 GBYes
    GeForce RTX 309024 GBYes
    GeForce RTX 409024 GBYes
    Radeon RX 7900 XTX24 GBYes
    GeForce RTX 509032 GBYes
    L424 GBYes
    L40S48 GBYes
    A100 80GB80 GBYes
    H100 80GB80 GBYes
    RTX PRO 6000 Blackwell96 GBYes
    H200141 GBYes
    Instinct MI300X192 GBYes

    Apple Silicon Macs (unified memory)

    • 16 GB: fits
    • 24 GB: fits
    • 32 GB: fits
    • 36 GB: fits
    • 48 GB: fits
    • 64 GB: fits
    • 96 GB: fits
    • 128 GB: fits
    • 192 GB: fits
    • 256 GB: fits
    • 512 GB: fits

    Filled: fits as is. Outlined: fits after raising the GPU memory limit (leaving 8 GB for macOS). macOS gives the GPU two-thirds of memory up to 32 GB and three-quarters above by default.

    Estimate · formula below

    FAQ

    Frequently asked questions

    How much VRAM does Llama 3.2 3B need?

    With an 8K context, about 3.7 GB at Q4_K_M, 5.1 GB at 8-bit (Q8_0) and 7.9 GB at full 16-bit precision. That covers the weights, the KV cache and runtime overhead.

    Can Llama 3.2 3B run on a 24 GB GPU?

    Yes, at FP16 / BF16 or smaller, with an 8K context (about 7.9 GB). Longer contexts need more memory for the KV cache.

    Can I run Llama 3.2 3B on a Mac?

    Yes: at Q4_K_M and an 8K context it fits a Mac with 16 GB of unified memory using macOS’s default GPU memory limit.

    How much memory does Llama 3.2 3B’s context use?

    Its KV cache grows by about 109 MB for every 1,000 tokens of context in FP16. At its full 131,072-token context the cache is about 14 GB. Quantising the cache to 8-bit roughly halves it.

    Related

    Updated 2026-10-08

    Architecture from config.json (a public copy of meta-llama/Llama-3.2-3B-Instruct); parameter count from the published weights; bits per weight from llama.cpp. GPU memory sizes from NVIDIA and AMD product pages.