VRAM · Local model
Gemma 4 26B-A4B VRAM requirements
Gemma 4 26B-A4B has 25.8 billion parameters. With an 8K context it needs about 17 GB of memory at Q4_K_M, which fits a 24 GB GPU such as the GeForce RTX 3090, and 53 GB at full precision.
- Parameters
- 25.8B
- Layers
- 30
- Max context
- 262,144
- Experts
- Dense
Attention: 5 full-attention layers, 25 sliding-window layers (last 1,024 tokens).
Requirements
VRAM by quantisation and context length
| Quantisation | 4K context | 32K context | 128K context | 256K context |
|---|---|---|---|---|
| FP16 / BF16 | 53 GB | 54 GB | 56 GB | 59 GB |
| FP8 | 27 GB | 27 GB | 29 GB | 32 GB |
| Q8_0 (8-bit) | 28 GB | 29 GB | 31 GB | 34 GB |
| Q6_K | 22 GB | 23 GB | 25 GB | 27 GB |
| Q5_K_M | 19 GB | 20 GB | 22 GB | 25 GB |
| Q4_K_M | 16 GB | 17 GB | 19 GB | 22 GB |
| MXFP4 | 14 GB | 15 GB | 17 GB | 20 GB |
| Q3_K_M | 14 GB | 14 GB | 16 GB | 19 GB |
| Q2_K | 11 GB | 11 GB | 13 GB | 16 GB |
Weights + FP16 KV cache + overhead (10%, at least 1 GB), batch 1. See the formula and assumptions.
Hardware
What can run it at Q4_K_M
On one GPU: GeForce RTX 3090, GeForce RTX 4090, Radeon RX 7900 XTX, GeForce RTX 5090, L4, L40S, A100 80GB, H100 80GB, RTX PRO 6000 Blackwell, H200, Instinct MI300X. Split across consumer cards: 3 × GeForce RTX 4060, 2 × GeForce RTX 3060, 2 × GeForce RTX 4070, 2 × GeForce RTX 5070. On a Mac: 32 GB of unified memory or more.
Calculator
Try other settings
25.8B parameters · 262,144-token context · model card
GGUF; the most popular balance of size and quality.
Requests served at the same time.
- Weights
- 15 GB
- KV cache
- 0.4 GB
- Overhead
- 1.5 GB
- Some layers only look at the last 1,024 tokens, so the cache grows more slowly with context.
| GPU | Memory | Runs it? |
|---|---|---|
| GeForce RTX 4060 | 8 GB | With 3 GPUs |
| GeForce RTX 3060 | 12 GB | With 2 GPUs |
| GeForce RTX 4070 | 12 GB | With 2 GPUs |
| GeForce RTX 5070 | 12 GB | With 2 GPUs |
| GeForce RTX 4060 Ti 16GB | 16 GB | With 2 GPUs |
| GeForce RTX 4080 Super | 16 GB | With 2 GPUs |
| GeForce RTX 5060 Ti 16GB | 16 GB | With 2 GPUs |
| GeForce RTX 5070 Ti | 16 GB | With 2 GPUs |
| GeForce RTX 5080 | 16 GB | With 2 GPUs |
| GeForce RTX 3090 | 24 GB | Yes |
| GeForce RTX 4090 | 24 GB | Yes |
| Radeon RX 7900 XTX | 24 GB | Yes |
| GeForce RTX 5090 | 32 GB | Yes |
| L4 | 24 GB | Yes |
| L40S | 48 GB | Yes |
| A100 80GB | 80 GB | Yes |
| H100 80GB | 80 GB | Yes |
| RTX PRO 6000 Blackwell | 96 GB | Yes |
| H200 | 141 GB | Yes |
| Instinct MI300X | 192 GB | Yes |
Apple Silicon Macs (unified memory)
- 16 GB: doesn’t fit
- 24 GB: doesn’t fit
- 32 GB: fits
- 36 GB: fits
- 48 GB: fits
- 64 GB: fits
- 96 GB: fits
- 128 GB: fits
- 192 GB: fits
- 256 GB: fits
- 512 GB: fits
Filled: fits as is. Outlined: fits after raising the GPU memory limit (leaving 8 GB for macOS). macOS gives the GPU two-thirds of memory up to 32 GB and three-quarters above by default.
FAQ
Frequently asked questions
How much VRAM does Gemma 4 26B-A4B need?
With an 8K context, about 17 GB at Q4_K_M, 28 GB at 8-bit (Q8_0) and 53 GB at full 16-bit precision. That covers the weights, the KV cache and runtime overhead.
Can Gemma 4 26B-A4B run on a 24 GB GPU?
Yes, at Q6_K or smaller, with an 8K context (about 22 GB). Longer contexts need more memory for the KV cache.
Can I run Gemma 4 26B-A4B on a Mac?
Yes: at Q4_K_M and an 8K context it fits a Mac with 32 GB of unified memory using macOS’s default GPU memory limit.
How much memory does Gemma 4 26B-A4B’s context use?
Its KV cache grows by about 20 MB for every 1,000 tokens of context in FP16 (its sliding-window or chunked layers stop growing once their window is full). At its full 262,144-token context the cache is about 5.2 GB. Quantising the cache to 8-bit roughly halves it.
Related