Lesson 6.2 · 27 min
Mixture of Experts: Routing Tokens to Specialists
How can a model have 671 billion parameters but only use 37 billion of them for each word it writes?
In short: A Mixture of Experts (MoE) layer replaces one big feed-forward network with many smaller "experts" and a small router that picks a few experts for each token. Total parameters (knowledge capacity) grow while compute per token stays close to that of a much smaller dense model. The price is memory for all experts, load-balancing work during training, and more complex serving.
Why Mixture of Experts was needed
In a dense Transformer every parameter takes part in processing every token. If we double the parameters to store more knowledge, we also double the arithmetic for each token, for training and for every answer we ever serve. Scaling laws told us bigger models are better, but the bill grows in lock-step.
The observation behind MoE: not every token needs all of the model's knowledge. A token in a Python snippet and a token in a French poem probably benefit from different stored patterns. So why pay to run all of them? Mixture of Experts splits part of the network into many sub-networks and runs only a few per token. This is called conditional computation: the amount of work depends on the input.
Think of it like a hospital A hospital employs dozens of specialists, but each patient sees only one or two. A receptionist (the router) looks at the patient and sends them to the right doctors. The hospital's total expertise is huge, but each visit only costs two consultations. MoE does this for every token, in every MoE layer.
Running example: our support chatbot gets the message "My invoice shows the wrong VAT rate". Each token of this sentence will be routed separately to a couple of experts in each MoE layer, and different tokens may visit different experts.
What an "expert" really means
An expert is simply a feed-forward network (FFN), the same kind of two-layer block that sits after attention in every Transformer layer. An MoE layer holds, say, 8, 64 or 256 of them side by side, each with its own weights and usually smaller than the single FFN it replaces.
Experts are not human-style specialists Nobody assigns "the math expert" or "the French expert". Specialization is learned and is usually much less tidy than the name suggests. The Mixtral authors, for example, reported no obvious topic-level specialization; routing patterns looked more tied to syntax and token types (such as indentation in code) than to subjects. Some newer designs, such as DeepSeek's fine-grained experts, aim for sharper specialization, but it is still learned, not designed.
Some models also add shared experts that every token always uses, to hold common knowledge, plus many routed experts chosen by the router. DeepSeek-V3, for instance, uses 1 shared expert and picks 8 of 256 routed experts per token.
The router and how it picks experts
The router (also called the gate) is a tiny linear layer. For a token vector x it computes one score per expert, turns the scores into probabilities with softmax, keeps the top-k experts, and uses their (renormalized) probabilities as mixing weights.
Routing one token, step by step
- Score: The token "VAT" (as a vector) is multiplied by the router weights, giving 8 scores, e.g. [0.1, 2.3, −0.4, 1.9, 0.0, 0.2, −1.1, 0.5].
- Softmax: Scores become probabilities that sum to 1. Expert 1 and expert 3 have the largest.
- Pick top-2: Keep experts 1 and 3; the other six are skipped entirely and cost no compute for this token.
- Renormalize: If their probabilities were 0.44 and 0.30, the gates become 0.44/0.74 ≈ 0.59 and 0.30/0.74 ≈ 0.41.
- Mix: Output = 0.59 × E₁(x) + 0.41 × E₃(x). The result goes on to the residual connection and the next layer, where routing happens again.
Where MoE sits inside a Transformer
A Transformer layer has two sub-blocks: attention (tokens exchange information) and the FFN (each token is processed on its own). MoE replaces only the FFN. Attention is unchanged and shared by all tokens. Since the FFN holds most of a model's parameters (roughly two thirds in a typical dense layer), this is where sparsity pays off most.
Not every layer has to be MoE. Some models alternate dense and MoE layers, or keep the first few layers dense, because early layers handle generic features where routing helps less.
Sparse activation and why it saves compute
Sparse activation means only a small fraction of parameters is used for each token. We now need two numbers to describe a model: total parameters (everything stored) and active parameters (what one token actually touches). Compute per token tracks the active count; knowledge capacity tracks, roughly, the total.
Mixtral 8x7B is a good worked example. Its name suggests 56B, but attention and embeddings are shared, so the total is about 46.7B. Each token uses 2 of 8 experts, so roughly 12.9B parameters are active: it runs at about the speed of a 13B dense model while drawing on the capacity of a much larger one.
Pause and think: A model has 8 experts of 1B parameters each, top-2 routing, plus 2B of shared attention and embedding parameters. What are its total and active parameter counts?
Total = 8 × 1B + 2B = 10B. Active = 2 × 1B + 2B = 4B. Compute per token is like a 4B dense model, but all 10B must be held in memory.
Sparse compute, not sparse memory The skipped experts still have to be loaded in GPU memory, because the next token may need them. An MoE with 671B total parameters needs memory for 671B parameters even though it computes like a ~37B model. That is the most common misunderstanding about MoE.
Load balancing across experts
Left alone, routers tend to collapse: a few experts get slightly better early, so they get more tokens, so they learn faster, so they get even more tokens. The rest starve and their parameters are wasted. On real hardware, experts sit on different GPUs (expert parallelism), so an overloaded expert also becomes a traffic jam.
- Auxiliary balance loss. Add a small extra loss that is lowest when tokens are spread evenly. The Switch Transformer form is
N · ∑ᵢ fᵢ · Pᵢ, wherefᵢis the fraction of tokens sent to expert i andPᵢis its average router probability; it equals 1.0 when perfectly balanced. - Capacity factor. Each expert accepts at most a fixed number of tokens per batch (e.g. 1.25 × the fair share). Overflow tokens are dropped (they skip the expert and pass through on the residual path) or sent to another expert.
- Bias-based balancing. DeepSeek-V3 adds a per-expert bias to the routing scores that is nudged down for busy experts and up for idle ones, avoiding most of the auxiliary loss (it calls this auxiliary-loss-free balancing).
- Noise in routing. Adding small random noise to router scores during training encourages exploration of under-used experts.
Code: a tiny top-2 MoE layer in numpy
This script routes 12 tokens through 8 tiny experts with top-2 gating, counts how many tokens each expert receives, and computes the Switch-style balance loss.
tiny_moe.py
import numpy as np
np.random.seed(42)
d, n_experts, top_k, n_tokens = 8, 8, 2, 12
X = np.random.randn(n_tokens, d) # 12 token vectors
W_router = np.random.randn(d, n_experts) * 0.5 # router: one score per expert
experts = [np.random.randn(d, d) * 0.3 for _ in range(n_experts)] # tiny FFNs
logits = X @ W_router # (12, 8) scores
probs = np.exp(logits - logits.max(1, keepdims=True))
probs /= probs.sum(1, keepdims=True) # softmax over experts
out = np.zeros_like(X)
load = np.zeros(n_experts, dtype=int)
for i in range(n_tokens):
top = np.argsort(probs[i])[-top_k:][::-1] # indices of the 2 best experts
gate = probs[i, top] / probs[i, top].sum() # renormalise the 2 weights
for e, g in zip(top, gate):
out[i] += g * np.maximum(0, X[i] @ experts[e]) # weighted expert outputs
load[e] += 1
if i < 3:
print(f"token {i}: experts {top.tolist()} weights {np.round(gate, 2).tolist()}")
print("tokens per expert:", load.tolist())
# Switch-style auxiliary loss: N * sum(fraction_routed * mean_prob); 1.0 = perfectly balanced
frac = load / load.sum()
aux = n_experts * np.sum(frac * probs.mean(0))
print(f"balance loss: {aux:.3f} (1.000 would be perfectly balanced)")
per_expert = d * d
print(f"total expert params: {n_experts * per_expert}, active per token: {top_k * per_expert}")Output:
token 0: experts [5, 4] weights [0.53, 0.47] token 1: experts [3, 6] weights [0.86, 0.14] token 2: experts [3, 7] weights [0.63, 0.37] tokens per expert: [1, 1, 6, 3, 5, 2, 5, 1] balance loss: 1.143 (1.000 would be perfectly balanced) total expert params: 512, active per token: 128
Advantages and challenges of MoE
Another subtle cost: at small batch sizes (one user, one token at a time) each token pulls in different experts, so the GPU spends its time loading expert weights from memory. MoE's compute savings show up most when many tokens are batched so each expert processes a decent group at once.
Pause and think: Our chatbot runs on one laptop GPU with 24 GB of memory. Someone proposes swapping our 7B dense model for a 47B-total / 13B-active MoE "because it is just as cheap". What is the flaw?
Compute per token is similar to a 13B model, but the weights of all 47B parameters must still be stored. At 16-bit precision that is about 94 GB, far beyond 24 GB, so it will not even fit without heavy quantization or offloading.
Why MoE powers many modern LLMs
At frontier scale, the binding constraint is compute: for training and for serving millions of users. MoE gives more quality per FLOP, so labs can build models with hundreds of billions to over a trillion total parameters while keeping each token affordable. Open examples include Mixtral, DeepSeek-V3 and V4, Qwen3's MoE variants, Llama 4 and gpt-oss. Several closed frontier models are reported to be MoE, though their makers usually do not publish details.
Real-world pattern Providers serving a large MoE spread experts across many GPUs (expert parallelism) and batch thousands of requests together, so every expert stays busy. This is why huge MoE models can be cheap per token through an API even though you could never run them on a single machine.
- Use MoE when you serve at scale with plenty of GPU memory and want the best quality per unit of compute.
- Prefer dense when memory is tight (phones, laptops), batch sizes are tiny, or you need simple fine-tuning and deployment.
Common mistakes and how to spot them
MoE models fail in a few typical ways, and most of them show up in one simple measurement: how many tokens each expert receives. Let us read one batch by hand.
Take 400 tokens, 4 experts and top-1 routing. The fair share is 400 / 4 = 100 tokens per expert. Suppose the counts are [288, 38, 35, 39] (illustrative; our practice script below produces them). The busiest expert holds 2.88 times its fair share. With a capacity factor of 1.25, each expert accepts at most 125 tokens, so 288 − 125 = 163 tokens overflow. That is about 41% of the batch skipping its expert. In a healthy layer the busiest expert stays close to the fair share and the overflow is small.
| What we see | Likely cause | What to check |
|---|---|---|
| A few experts receive most tokens | Router collapse: early winners keep winning | Tokens per expert in each batch; busiest count ÷ fair share |
| Quality drops after adding a capacity limit | Many tokens overflow and skip their expert | Overflow tokens per batch as a share of all tokens |
| Out of memory although active parameters are small | Every expert must be loaded, used or not | Plan memory from total parameters, not active ones |
| Slow with one user, fast with many | Each expert gets only a handful of tokens per step | Tokens per expert per step; batch more requests together |
| Balance looks fine on average, bad on one kind of input | Counts were averaged over mixed data | Tokens per expert separately for each kind of input |
One number to log For every MoE layer, log busiest expert count ÷ fair share. A value near 1 means balanced. A value that climbs during training is an early warning of collapse, long before quality metrics move.
Practice: try it yourself
We will build a router that starts out unfair, then fix it with the bias trick from the load-balancing section: push the score of busy experts down and the score of idle experts up, a little after every batch. We also count how many tokens would overflow a capacity limit.
practice_moe_balance.py
import numpy as np
rng = np.random.default_rng(0)
n_tokens, n_experts = 400, 4
# Router scores (illustrative): expert 0 starts with an unfair head start
scores = rng.normal(0, 1, (n_tokens, n_experts))
scores[:, 0] += 1.5
def loads(bias):
# top-1 routing: each token goes to its best expert after the bias is added
choice = (scores + bias).argmax(axis=1)
return np.bincount(choice, minlength=n_experts)
fair = n_tokens / n_experts # 100 tokens per expert
cap = int(1.25 * fair) # capacity factor 1.25 -> 125 tokens
bias = np.zeros(n_experts)
for step in range(6):
load = loads(bias)
dropped = int(np.maximum(load - cap, 0).sum()) # tokens over the cap
print(f"step {step}: load={load.tolist()} over cap {cap}: {dropped} "
f"bias={np.round(bias, 2).tolist()}")
# busy experts get a lower bias, idle experts a higher one
bias -= 0.5 * (load - fair) / fairOutput:
step 0: load=[288, 38, 35, 39] over cap 125: 163 bias=[0.0, 0.0, 0.0, 0.0] step 1: load=[139, 85, 77, 99] over cap 125: 14 bias=[-0.94, 0.31, 0.32, 0.3] step 2: load=[107, 98, 93, 102] over cap 125: 0 bias=[-1.14, 0.38, 0.44, 0.31] step 3: load=[101, 102, 97, 100] over cap 125: 0 bias=[-1.17, 0.4, 0.48, 0.3] step 4: load=[100, 101, 98, 101] over cap 125: 0 bias=[-1.17, 0.38, 0.49, 0.3] step 5: load=[100, 99, 100, 101] over cap 125: 0 bias=[-1.17, 0.38, 0.5, 0.3]
Now change it:
- Raise the head start on line 7 from
1.5to3.0. Predict the first load of expert 0, and whether 6 rounds are still enough to balance. - Change the step size on the last line from
0.5to4.0. Predict: do the loads settle faster, or do they swing back and forth? - Set the capacity factor to
1.0so the cap equals the fair share. Predict whether the overflow count can ever stay at exactly 0.
Pause and think: After balancing, expert 0 has a bias of −1.17. Does that mean expert 0 became a worse expert?
No. The bias does not touch the expert's own weights. It only cancels most of the unfair head start in the router scores. Tokens that prefer expert 0 by a wide margin still go there; tokens that preferred it only slightly now go elsewhere. The bias changes who gets which tokens, not what each expert computes.
Pause and think: In round 0 the load is [288, 38, 35, 39] with a cap of 125 and overflow tokens are dropped. How many tokens are actually processed by an expert, and what happens to the rest?
125 + 38 + 35 + 39 = 237 tokens are processed. The other 163 skip the expert and pass through on the residual path, so they get no feed-forward update in this layer. That is why heavy overflow hurts quality even though nothing crashes.
Key takeaways
- An MoE layer = many FFN experts + a router that sends each token to the top-k of them.
- Total parameters set capacity; active parameters set compute per token.
- All experts must sit in memory: MoE saves compute, not memory.
- Load balancing (aux loss, capacity limits, bias tweaks) prevents a few experts from taking all tokens.
- MoE shines at large-scale, batched serving; dense models are simpler for small or memory-limited deployments.
Key terms
- Expert: One of several feed-forward sub-networks inside an MoE layer.
- Router (gate): A small layer that scores experts for each token and picks the top-k.
- Top-k routing: Keeping only the k highest-scoring experts for a token and mixing their outputs.
- Active parameters: The parameters actually used to process one token.
- Load balancing: Techniques that spread tokens evenly across experts so none is overloaded or idle.
- Capacity factor: A cap on how many tokens each expert accepts per batch; overflow tokens are dropped or rerouted.
- Shared expert: An expert that every token uses, holding common knowledge alongside routed experts.
← 6.1 A Timeline of LLM Architecture Improvements · 6.3 Grouped Query Attention: Fewer KV Heads, Same Quality →