AI glossary · Local AI
What is a mixture-of-experts (MoE) model?
Also called: MoE, sparse mixture of experts, MoE model
Definition
A mixture-of-experts (MoE) model splits parts of each layer into many parallel sub-networks called experts and uses a small router to send each token through only a few of them, so it computes with a fraction of its parameters.
Explained
How it works
Mixture of experts dates back to 1991; today’s LLMs build on a 2017 paper, “Outrageously Large Neural Networks” (arXiv 1701.06538), which added a layer of many feed-forward sub-networks plus a trainable gating network that picks a sparse combination of them for each input. In today’s LLMs the experts take the place of the feed-forward block in most or all transformer layers, and a router chooses a few per token in each of those layers.
That gives two parameter counts. Total parameters include every expert; active parameters are the ones one token actually uses. Mixtral 8x7B has 8 experts per layer, routes each token to 2, and uses 13B of its 47B parameters per token. Qwen3 30B-A3B puts both in its name: 30.5B total, 3.3B activated, 8 of 128 experts per token.
The practical rule: memory follows total parameters, compute follows active ones. Any token can be routed to any expert, so all of them must be loaded; the Mixtral paper says its serving memory is proportional to the 47B. But each token only does the work of the active parameters.
Example
Qwen3 30B-A3B next to the dense Qwen3 32B
At Q4_K_M with an 8K context, Qwen3 30B-A3B needs 20 GB and the dense Qwen3 32B needs 23 GB: almost the same, because both load all their weights. But each token passes through 3.3B parameters in the MoE model and all 32.8B in the dense one, about 10 times fewer, so the MoE model generates text faster on the same hardware.
| Model | Total parameters | Active per token | Experts (used per token) | Memory at 8K context |
|---|---|---|---|---|
| Qwen3 30B-A3B | 30.5B | 3.3B | 128 (8) | 20 GB |
| gpt-oss 20B | 20.9B | 3.6B | 32 (4) | 13 GB |
| Qwen3 32B | 32.8B | 32.8B | Dense | 23 GB |
Active parameters from the model cards; experts from each config.json, fetched 2026-10-08. Memory from our VRAM calculator: Q4_K_M weights (gpt-oss 20B uses its official MXFP4 file), FP16 KV cache, batch 1, plus overhead. GB = 1,024³ bytes.
Cost and quality
Why it matters
MoE is how many recent open models get large capacity at a lower cost per token: 8 of the 28 models in our VRAM calculator use it. For local use, size your VRAM by total parameters and expect faster generation than a dense model of the same size.
llama.cpp can also keep the expert weights in system RAM (--cpu-moe, or --n-cpu-moe N for the first N layers) while the rest stays on the GPU, which lets an MoE model larger than your card still load.
Don’t mix up
Common confusions
- Active parameters vs memory
- A model called “A3B” doesn’t run in 3B worth of memory. Active parameters set the compute per token; every expert still has to be loaded.
- Experts vs topic specialists
- The experts aren’t subject specialists. In the Mixtral paper the authors saw no obvious pattern of experts by topic; routing followed syntax more than domain, and consecutive tokens often went to the same experts.
Go deeper
Try it and read more
- Free toolGPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.
- Free toolQwen3 30B-A3B VRAM requirementsQwen3 30B-A3B needs about 20 GB of VRAM at Q4_K_M and 34 GB at 8-bit with an 8K context.
- Guide · 12 min readHow much VRAM do you need to run an LLM locally?VRAM needed for 8B to 70B models at Q4, Q8 and FP16, what fits on 8 to 80 GB GPUs, and how context length, KV cache and CPU offload change it.
Related
Related terms
- VRAMVRAM is the memory on a graphics card, and for running AI models locally it is the main limit: a model runs at full GPU speed only when its weights, KV cache and working buffers all fit in it.
- QuantisationQuantisation is storing a model’s weights in fewer bits than the 16 per weight most models are released with, so the model needs less memory and runs on smaller hardware, at a small cost in accuracy.
- Open weightsOpen weights means a model’s trained parameters are published for anyone to download and run on their own hardware, under a licence that sets what they may do with them.
- KV cacheThe KV cache is the memory where a language model keeps the attention keys and values of every token it has already processed, so each new token is computed without reprocessing the whole sequence.