Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

VRAM · Local model

DeepSeek R1 (671B) VRAM requirements

DeepSeek R1 (671B) has 671.0 billion parameters. With an 8K context it needs about 421 GB of memory at Q4_K_M, and 1375 GB at full precision.

Parameters
671.0B
Layers
61
Max context
163,840
Experts
256 (8 active)

Attention: 61 layers with multi-head latent attention (one compressed 576-value vector per token).

Requirements

VRAM by quantisation and context length

Quantisation4K context32K context128K context160K context
FP16 / BF161375 GB1377 GB1384 GB1387 GB
FP8688 GB690 GB697 GB699 GB
Q8_0 (8-bit)731 GB733 GB740 GB742 GB
Q6_K564 GB566 GB573 GB575 GB
Q5_K_M490 GB492 GB499 GB502 GB
Q4_K_M420 GB423 GB430 GB432 GB
MXFP4365 GB368 GB375 GB377 GB
Q3_K_M344 GB346 GB353 GB355 GB
Q2_K272 GB274 GB281 GB283 GB

Weights + FP16 KV cache + overhead (10%, at least 1 GB), batch 1. See the formula and assumptions.

Hardware

What can run it at Q4_K_M

No single GPU in our list has enough memory at Q4_K_M with an 8K context. No Mac can hold it at this setting with the default GPU memory limit.

It’s a mixture-of-experts model: all 256 experts must be in memory, but only 8 run for each token, so it generates faster than a dense model of the same size.

Calculator

Try other settings

671.0B parameters, 256 experts (8 used per token) · 163,840-token context · model card

GGUF; the most popular balance of size and quality.

Requests served at the same time.

Estimated memory needed421 GB
Weights
382 GB
KV cache
0.5 GB
Overhead
38 GB
  • Mixture of experts: all 256 experts must be in memory, but only 8 run for each token, so it generates faster than a dense model of this size.
GPUMemoryRuns it?
GeForce RTX 40608 GBNo
GeForce RTX 306012 GBNo
GeForce RTX 407012 GBNo
GeForce RTX 507012 GBNo
GeForce RTX 4060 Ti 16GB16 GBNo
GeForce RTX 4080 Super16 GBNo
GeForce RTX 5060 Ti 16GB16 GBNo
GeForce RTX 5070 Ti16 GBNo
GeForce RTX 508016 GBNo
GeForce RTX 309024 GBNo
GeForce RTX 409024 GBNo
Radeon RX 7900 XTX24 GBNo
GeForce RTX 509032 GBNo
L424 GBNo
L40S48 GBNo
A100 80GB80 GBWith 6 GPUs
H100 80GB80 GBWith 6 GPUs
RTX PRO 6000 Blackwell96 GBWith 5 GPUs
H200141 GBWith 3 GPUs
Instinct MI300X192 GBWith 3 GPUs

Apple Silicon Macs (unified memory)

  • 16 GB: doesn’t fit
  • 24 GB: doesn’t fit
  • 32 GB: doesn’t fit
  • 36 GB: doesn’t fit
  • 48 GB: doesn’t fit
  • 64 GB: doesn’t fit
  • 96 GB: doesn’t fit
  • 128 GB: doesn’t fit
  • 192 GB: doesn’t fit
  • 256 GB: doesn’t fit
  • 512 GB: fits after raising the GPU limit

Filled: fits as is. Outlined: fits after raising the GPU memory limit (leaving 8 GB for macOS). macOS gives the GPU two-thirds of memory up to 32 GB and three-quarters above by default.

Estimate · formula below

FAQ

Frequently asked questions

How much VRAM does DeepSeek R1 (671B) need?

With an 8K context, about 421 GB at Q4_K_M, 731 GB at 8-bit (Q8_0) and 1375 GB at full 16-bit precision. That covers the weights, the KV cache and runtime overhead.

Can DeepSeek R1 (671B) run on a 24 GB GPU?

Not on a single 24 GB card: even at Q2_K it needs about 272 GB. At Q4_K_M you’d need 18 × 24 GB cards.

Can I run DeepSeek R1 (671B) on a Mac?

Not at Q4_K_M: it needs about 421 GB, more than the largest Mac can give the GPU by default.

How much memory does DeepSeek R1 (671B)’s context use?

Its KV cache grows by about 67 MB for every 1,000 tokens of context in FP16. At its full 163,840-token context the cache is about 11 GB. Quantising the cache to 8-bit roughly halves it.

Related