Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

VRAM · Local model

Llama 4 Scout VRAM requirements

Llama 4 Scout has 108.6 billion parameters. With an 8K context it needs about 70 GB of memory at Q4_K_M, which fits an 80 GB GPU such as the A100 80GB, and 224 GB at full precision.

Parameters
108.6B
Layers
48
Max context
10,485,760
Experts
16 (1 active)

Attention: 12 full-attention layers, 36 chunked-attention layers (8,192-token chunks).

Requirements

VRAM by quantisation and context length

Quantisation4K context32K context128K context
FP16 / BF16223 GB225 GB230 GB
FP8112 GB114 GB119 GB
Q8_0 (8-bit)119 GB121 GB126 GB
Q6_K92 GB94 GB99 GB
Q5_K_M80 GB82 GB87 GB
Q4_K_M69 GB71 GB76 GB
MXFP460 GB62 GB67 GB
Q3_K_M56 GB59 GB63 GB
Q2_K45 GB47 GB52 GB

Weights + FP16 KV cache + overhead (10%, at least 1 GB), batch 1. See the formula and assumptions.

Hardware

What can run it at Q4_K_M

On one GPU: A100 80GB, H100 80GB, RTX PRO 6000 Blackwell, H200, Instinct MI300X. Split across consumer cards: 6 × GeForce RTX 3060, 6 × GeForce RTX 4070, 6 × GeForce RTX 5070, 5 × GeForce RTX 4060 Ti 16GB. On a Mac: 96 GB of unified memory or more.

It’s a mixture-of-experts model: all 16 experts must be in memory, but only 1 run for each token, so it generates faster than a dense model of the same size.

Calculator

Try other settings

108.6B parameters, 16 experts (1 used per token) · 10,485,760-token context · model card

GGUF; the most popular balance of size and quality.

Requests served at the same time.

Estimated memory needed70 GB
Weights
62 GB
KV cache
1.5 GB
Overhead
6.3 GB
  • Mixture of experts: all 16 experts must be in memory, but only 1 run for each token, so it generates faster than a dense model of this size.
  • Some layers only look at the last 8,192 tokens, so the cache grows more slowly with context.
GPUMemoryRuns it?
GeForce RTX 40608 GBNo
GeForce RTX 306012 GBWith 6 GPUs
GeForce RTX 407012 GBWith 6 GPUs
GeForce RTX 507012 GBWith 6 GPUs
GeForce RTX 4060 Ti 16GB16 GBWith 5 GPUs
GeForce RTX 4080 Super16 GBWith 5 GPUs
GeForce RTX 5060 Ti 16GB16 GBWith 5 GPUs
GeForce RTX 5070 Ti16 GBWith 5 GPUs
GeForce RTX 508016 GBWith 5 GPUs
GeForce RTX 309024 GBWith 3 GPUs
GeForce RTX 409024 GBWith 3 GPUs
Radeon RX 7900 XTX24 GBWith 3 GPUs
GeForce RTX 509032 GBWith 3 GPUs
L424 GBWith 3 GPUs
L40S48 GBWith 2 GPUs
A100 80GB80 GBYes
H100 80GB80 GBYes
RTX PRO 6000 Blackwell96 GBYes
H200141 GBYes
Instinct MI300X192 GBYes

Apple Silicon Macs (unified memory)

  • 16 GB: doesn’t fit
  • 24 GB: doesn’t fit
  • 32 GB: doesn’t fit
  • 36 GB: doesn’t fit
  • 48 GB: doesn’t fit
  • 64 GB: doesn’t fit
  • 96 GB: fits
  • 128 GB: fits
  • 192 GB: fits
  • 256 GB: fits
  • 512 GB: fits

Filled: fits as is. Outlined: fits after raising the GPU memory limit (leaving 8 GB for macOS). macOS gives the GPU two-thirds of memory up to 32 GB and three-quarters above by default.

Estimate · formula below

FAQ

Frequently asked questions

How much VRAM does Llama 4 Scout need?

With an 8K context, about 70 GB at Q4_K_M, 120 GB at 8-bit (Q8_0) and 224 GB at full 16-bit precision. That covers the weights, the KV cache and runtime overhead.

Can Llama 4 Scout run on a 24 GB GPU?

Not on a single 24 GB card: even at Q2_K it needs about 46 GB. At Q4_K_M you’d need 3 × 24 GB cards.

Can I run Llama 4 Scout on a Mac?

Yes: at Q4_K_M and an 8K context it fits a Mac with 96 GB of unified memory using macOS’s default GPU memory limit.

How much memory does Llama 4 Scout’s context use?

Its KV cache grows by about 47 MB for every 1,000 tokens of context in FP16 (its sliding-window or chunked layers stop growing once their window is full). At its full 10,485,760-token context the cache is about 481 GB. Quantising the cache to 8-bit roughly halves it.

Related

Prefer the API? See Llama 4 Scout pricing.

Updated 2026-10-08

Architecture from config.json (a public copy of meta-llama/Llama-4-Scout-17B-16E-Instruct); parameter count from the published weights; bits per weight from llama.cpp. GPU memory sizes from NVIDIA and AMD product pages.