Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Local AI

What is a mixture-of-experts (MoE) model?

Also called: MoE, sparse mixture of experts, MoE model

Definition

A mixture-of-experts (MoE) model splits parts of each layer into many parallel sub-networks called experts and uses a small router to send each token through only a few of them, so it computes with a fraction of its parameters.

Explained

How it works

Mixture of experts dates back to 1991; today’s LLMs build on a 2017 paper, “Outrageously Large Neural Networks” (arXiv 1701.06538), which added a layer of many feed-forward sub-networks plus a trainable gating network that picks a sparse combination of them for each input. In today’s LLMs the experts take the place of the feed-forward block in most or all transformer layers, and a router chooses a few per token in each of those layers.

That gives two parameter counts. Total parameters include every expert; active parameters are the ones one token actually uses. Mixtral 8x7B has 8 experts per layer, routes each token to 2, and uses 13B of its 47B parameters per token. Qwen3 30B-A3B puts both in its name: 30.5B total, 3.3B activated, 8 of 128 experts per token.

The practical rule: memory follows total parameters, compute follows active ones. Any token can be routed to any expert, so all of them must be loaded; the Mixtral paper says its serving memory is proportional to the 47B. But each token only does the work of the active parameters.

Example

Qwen3 30B-A3B next to the dense Qwen3 32B

At Q4_K_M with an 8K context, Qwen3 30B-A3B needs 20 GB and the dense Qwen3 32B needs 23 GB: almost the same, because both load all their weights. But each token passes through 3.3B parameters in the MoE model and all 32.8B in the dense one, about 10 times fewer, so the MoE model generates text faster on the same hardware.

Total vs active parameters, and the memory each model needs
ModelTotal parametersActive per tokenExperts (used per token)Memory at 8K context
Qwen3 30B-A3B30.5B3.3B128 (8)20 GB
gpt-oss 20B20.9B3.6B32 (4)13 GB
Qwen3 32B32.8B32.8BDense23 GB

Active parameters from the model cards; experts from each config.json, fetched 2026-10-08. Memory from our VRAM calculator: Q4_K_M weights (gpt-oss 20B uses its official MXFP4 file), FP16 KV cache, batch 1, plus overhead. GB = 1,024³ bytes.

Cost and quality

Why it matters

MoE is how many recent open models get large capacity at a lower cost per token: 8 of the 28 models in our VRAM calculator use it. For local use, size your VRAM by total parameters and expect faster generation than a dense model of the same size.

llama.cpp can also keep the expert weights in system RAM (--cpu-moe, or --n-cpu-moe N for the first N layers) while the rest stays on the GPU, which lets an MoE model larger than your card still load.

Don’t mix up

Common confusions

Active parameters vs memory
A model called “A3B” doesn’t run in 3B worth of memory. Active parameters set the compute per token; every expert still has to be loaded.
Experts vs topic specialists
The experts aren’t subject specialists. In the Mixtral paper the authors saw no obvious pattern of experts by topic; routing followed syntax more than domain, and consecutive tokens often went to the same experts.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary