Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

VRAM · Local model

Qwen 3.8 27B VRAM requirements

Qwen 3.8 27B has 27.8 billion parameters. With an 8K context it needs about 18 GB of memory at Q4_K_M, which fits a 24 GB GPU such as the GeForce RTX 3090, and 57 GB at full precision.

Parameters
27.8B
Layers
64
Max context
262,144
Experts
Dense

Attention: 16 full-attention layers, 48 Mamba or linear-attention layers.

Requirements

VRAM by quantisation and context length

Quantisation4K context32K context128K context256K context
FP16 / BF1657 GB59 GB66 GB75 GB
FP829 GB31 GB37 GB46 GB
Q8_0 (8-bit)31 GB32 GB39 GB48 GB
Q6_K24 GB26 GB32 GB41 GB
Q5_K_M21 GB22 GB29 GB38 GB
Q4_K_M18 GB20 GB26 GB35 GB
MXFP415 GB17 GB24 GB33 GB
Q3_K_M15 GB16 GB23 GB32 GB
Q2_K12 GB13 GB20 GB29 GB

Weights + FP16 KV cache + overhead (10%, at least 1 GB), batch 1. See the formula and assumptions.

Hardware

What can run it at Q4_K_M

On one GPU: GeForce RTX 3090, GeForce RTX 4090, Radeon RX 7900 XTX, GeForce RTX 5090, L4, L40S, A100 80GB, H100 80GB, RTX PRO 6000 Blackwell, H200, Instinct MI300X. Split across consumer cards: 3 × GeForce RTX 4060, 2 × GeForce RTX 3060, 2 × GeForce RTX 4070, 2 × GeForce RTX 5070. On a Mac: 32 GB of unified memory or more.

Calculator

Try other settings

27.8B parameters · 262,144-token context · model card

GGUF; the most popular balance of size and quality.

Requests served at the same time.

Estimated memory needed18 GB
Weights
16 GB
KV cache
0.5 GB
Overhead
1.6 GB
  • 48 of its layers use Mamba or linear attention, which keep a small fixed state instead of a growing cache (counted in the overhead).
GPUMemoryRuns it?
GeForce RTX 40608 GBWith 3 GPUs
GeForce RTX 306012 GBWith 2 GPUs
GeForce RTX 407012 GBWith 2 GPUs
GeForce RTX 507012 GBWith 2 GPUs
GeForce RTX 4060 Ti 16GB16 GBWith 2 GPUs
GeForce RTX 4080 Super16 GBWith 2 GPUs
GeForce RTX 5060 Ti 16GB16 GBWith 2 GPUs
GeForce RTX 5070 Ti16 GBWith 2 GPUs
GeForce RTX 508016 GBWith 2 GPUs
GeForce RTX 309024 GBYes
GeForce RTX 409024 GBYes
Radeon RX 7900 XTX24 GBYes
GeForce RTX 509032 GBYes
L424 GBYes
L40S48 GBYes
A100 80GB80 GBYes
H100 80GB80 GBYes
RTX PRO 6000 Blackwell96 GBYes
H200141 GBYes
Instinct MI300X192 GBYes

Apple Silicon Macs (unified memory)

  • 16 GB: doesn’t fit
  • 24 GB: doesn’t fit
  • 32 GB: fits
  • 36 GB: fits
  • 48 GB: fits
  • 64 GB: fits
  • 96 GB: fits
  • 128 GB: fits
  • 192 GB: fits
  • 256 GB: fits
  • 512 GB: fits

Filled: fits as is. Outlined: fits after raising the GPU memory limit (leaving 8 GB for macOS). macOS gives the GPU two-thirds of memory up to 32 GB and three-quarters above by default.

Estimate · formula below

FAQ

Frequently asked questions

How much VRAM does Qwen 3.8 27B need?

With an 8K context, about 18 GB at Q4_K_M, 31 GB at 8-bit (Q8_0) and 57 GB at full 16-bit precision. That covers the weights, the KV cache and runtime overhead.

Can Qwen 3.8 27B run on a 24 GB GPU?

Yes, at Q6_K or smaller, with an 8K context (about 24 GB). Longer contexts need more memory for the KV cache.

Can I run Qwen 3.8 27B on a Mac?

Yes: at Q4_K_M and an 8K context it fits a Mac with 32 GB of unified memory using macOS’s default GPU memory limit.

How much memory does Qwen 3.8 27B’s context use?

Its KV cache grows by about 63 MB for every 1,000 tokens of context in FP16. At its full 262,144-token context the cache is about 16 GB. Quantising the cache to 8-bit roughly halves it.

Related