TL;DR
Mixture of Experts (MoE) is a conditional-computation architecture. A router scores each token representation, selects one or more expert subnetworks, dispatches token states to them, and combines their outputs. This can increase total parameter capacity without running every expert for every token.
Sparse activation is not free compute. A production MoE still pays for shared layers, expert-weight storage, routing, token permutation, memory movement, collective communication, uneven expert loads, and many small matrix multiplications. Therefore:
total parameters != active parameters != FLOPs != latency != quality
The useful engineering question is not whether MoE is universally more efficient. It is whether a specific MoE release delivers better accepted quality, latency, throughput, and cost under a fixed workload and hardware contract.
Table of Contents
- What MoE Changes in a Transformer
- The Batched Routing Data Path
- Total Parameters, Active Parameters, and Real Cost
- Training: Balance Without Erasing Useful Routing
- Serving: Memory, Communication, and Kernels
- Prefill and Decode Stress Different Paths
- What Published Architectures Actually Establish
- A Reproducible MoE Evaluation Contract
- Failure Modes and Diagnostics
- FAQ
- Summary
What MoE Changes in a Transformer
The original 1991 adaptive mixtures of local experts used separate networks and a learned gating network to divide training cases into subtasks. Modern sparse language models apply the same conditional-computation idea at much larger scale.
In a common Transformer design, selected dense feed-forward network (FFN) blocks are replaced by MoE blocks:
token states
-> normalization and attention
-> residual path
-> MoE router
-> selected expert FFNs
-> weighted combination
-> residual path
Attention, embeddings, normalization, output heads, and some FFN layers may remain dense or shared. Architectures also differ in how frequently they insert MoE layers, whether they include always-on shared experts, how many routed experts exist, and how many are selected.
An expert is usually an FFN with its own weights. It is not automatically a human-readable specialist. A router can learn repeatable preferences without producing clean categories such as "Python," "biology," or "grammar." Semantic specialization must be demonstrated with routing statistics, controlled interventions, and quality measurements.
The Batched Routing Data Path
A one-token loop hides the hard part of MoE. Real training and serving route a batch of token states:
- Score: the router maps each token state to logits over experts.
- Select: a routing policy chooses top-k experts, expert groups, or another sparse assignment.
- Normalize: selected scores become combination weights according to the architecture.
- Account for capacity: the implementation determines whether assignments are padded, dropped, rerouted, or executed with variable sizes.
- Permute: token states are grouped by destination expert.
- Dispatch: when experts live on other devices, collectives move token states to their owners.
- Compute: grouped or batched matrix multiplications run local experts.
- Combine and restore: expert outputs are weighted, returned, and restored to original token order.
The conceptual equation for token state x_t is:
MoE(x_t) = sum over selected experts i of g_i(x_t) * E_i(x_t)
That equation omits capacity, device placement, communication, padding, precision, and kernel scheduling. It describes model semantics, not a production implementation.
The 2017 sparsely-gated MoE paper established that a trainable gate could activate a sparse combination of many FFN experts. Switch Transformer then showed one-expert routing as a specific simplification. Neither result makes top-1 or top-2 a universal default.
Total Parameters Active Parameters and Real Cost
Four quantities must be reported separately:
| Quantity | What it describes | What it does not prove |
|---|---|---|
| Total parameters | All model weights, including inactive experts | Per-token arithmetic, latency, or quality |
| Active parameters | Weights used along a token's selected path | Actual FLOPs, memory traffic, or dense-model equivalence |
| FLOPs | Arithmetic under a stated sequence and batch shape | Kernel utilization, communication, or wall-clock latency |
| Resident memory | Weights, KV cache, activations, workspaces, runtime state | Whether the workload meets latency or throughput objectives |
Active parameters can be a useful architecture descriptor, but it is not a performance model. Shared attention and dense layers still run. Expert weights must be resident, sharded, streamed, or offloaded. Routing creates data movement and synchronization. Small or uneven expert batches can underuse accelerators.
For one release and workload, reason about:
request_cost =
dense_shared_compute
+ routed_expert_compute
+ router_and_permutation
+ local_memory_traffic
+ expert_parallel_communication
+ synchronization_and_padding
+ runtime_overhead
This is why "13B active parameters" does not establish "13B dense-model speed," and total parameters do not establish dense-model-equivalent quality.
Training Balance Without Erasing Useful Routing
Routers and experts are trained jointly, but discrete selection and uneven demand create several coupled problems:
- a few experts may receive most tokens;
- underused experts may learn slowly;
- overloaded devices can determine step time;
- capacity limits can drop assignments or waste padding;
- a strong balancing objective can distort the language-model objective;
- low-precision router decisions can change assignments near score boundaries.
Auxiliary load-balancing loss is one design, not a law. Switch Transformer used balancing and capacity mechanisms, while DeepSeek-V3 describes a dynamic expert-bias strategy intended to balance load without a conventional auxiliary loss. Production frameworks expose several alternatives, including auxiliary losses, sequence- or global-level balance, Sinkhorn-style routing, dynamic bias, and no balancing.
Track at least:
tokens per expert and per device
router probability mass per expert
coefficient of variation and peak-to-mean load
overflow, drop, reroute, and padding rates
expert batch-size distribution
router entropy and assignment churn
language-model loss and downstream quality by slice
communication and expert-compute time
Uniform routing is not the final objective. The target is stable quality and efficient execution without dead experts or uncontrolled hotspots.
Serving Memory Communication and Kernels
Expert parallelism is not tensor parallelism
- Expert parallelism (EP) places different experts on different workers and routes token states to the workers that own selected experts.
- Tensor parallelism (TP) shards tensors within a layer and synchronizes partial results.
- Pipeline parallelism (PP) assigns layer ranges to stages.
- Data parallelism (DP) replicates model execution across data shards.
Large deployments combine these dimensions. The best mapping depends on expert count, hidden sizes, sequence lengths, topology, memory, and traffic. It cannot be selected from total parameter count alone.
Dispatch is a systems operation
With EP, a typical MoE layer performs an all-to-all-style dispatch, local expert computation, and a return/combine collective. PyTorch's large-scale MoE account and the versioned Megatron-Core 0.16 MoE guide both expose this data movement as a first-class concern.
The runtime may use dropless variable-size execution, capacity padding, expert replication, grouped GEMMs, fused permutation kernels, communication overlap, or topology-aware routing. These are implementation choices. The old claim that serving frameworks universally use token dropping to avoid crashes confuses a capacity-based training option with all inference systems.
Current vLLM expert-parallel deployment guidance also treats backend selection, Expert Parallel Load Balancing, memory overhead, network configuration, and benchmarking as separate deployment decisions. Feature availability and performance remain release-, model-, topology-, and hardware-specific.
Local execution needs a measured memory plan
Quantization reduces weight storage, but local feasibility also depends on:
- exact checkpoint and quantization format;
- runtime metadata and temporary buffers;
- KV cache for context length, batch, and concurrency;
- CPU, GPU, or unified-memory placement;
- memory bandwidth and offload traffic;
- prompt processing and decode targets.
No fixed RAM number proves that an MoE model will run comfortably. Publish the checkpoint digest, runtime revision, context, concurrency, token rates, latency percentiles, peak memory, and quality checks.
Prefill and Decode Stress Different Paths
Prefill processes many prompt tokens together. It may create larger expert batches and more efficient grouped GEMMs, but long prompts increase attention, activation, KV-cache, dispatch, and network volume.
Decode usually advances each active sequence by one token per step. Continuous batching can aggregate work, but expert batches may remain smaller and more uneven. Kernel launches, synchronization, memory bandwidth, and interconnect latency can dominate.
Report both paths:
prefill: input-token throughput, time to first token, expert batch sizes
decode: output-token throughput, time per output token, inter-token latency
both: p50/p95/p99 latency, load imbalance, communication share, peak memory
An optimization that improves prefill throughput can leave decode unchanged or worse.
What Published Architectures Actually Establish
Use disclosed architectures as bounded examples:
- Mixtral 8x7B reports eight FFN experts per layer, two selected per token, 47B total parameters, and 13B active parameters for that checkpoint. These figures do not transfer to every MoE.
- DeepSeekMoE studies finer-grained routed experts and always-on shared experts to reduce redundancy and improve specialization in its evaluated designs.
- DeepSeek-V3 reports 671B total and 37B activated parameters, auxiliary-loss-free balancing, node-limited routing, no token dropping in its stated training setup, and separate prefill and decode deployment strategies.
- Switch Transformer evaluates top-1 routing and capacity controls in its own training configuration.
Closed-model internals should not be inferred from leaks or repeated speculation. OpenAI's public GPT-4 materials do not disclose an expert count or the leaked parameter arithmetic previously shown on this page, so those claims are not evidence for an MoE architecture.
A Reproducible MoE Evaluation Contract
Before comparing dense and sparse candidates, freeze the tokenizer and precision, workload trace, hardware topology, gate revision, and metric boundaries. This dependency-free Go fixture rejects unpaired evidence, failed quality or safety gates, and invalid counters before comparing accepted-output goodput.
package main
import (
"fmt"
"math"
"strings"
)
type Trial struct {
Name string
ArtifactRevision string
RuntimeRevision string
TokenizerRevision string
Precision string
WorkloadRevision string
HardwareRevision string
GateRevision string
AcceptedOutputs int
CompletedRequests int
TotalRequests int
DurationSeconds float64
TTFTP95MS float64
TPOTP95MS float64
PeakMemoryGB float64
CommunicationShare float64
CriticalFailures int
}
func validateTrial(trial Trial, minAcceptance, maxTTFT, maxTPOT float64) error {
if minAcceptance < 0 || minAcceptance > 1 || maxTTFT <= 0 || maxTPOT <= 0 {
return fmt.Errorf("release gate is invalid")
}
identity := []string{
trial.Name, trial.ArtifactRevision, trial.RuntimeRevision,
trial.TokenizerRevision, trial.Precision, trial.WorkloadRevision,
trial.HardwareRevision, trial.GateRevision,
}
for _, value := range identity {
if strings.TrimSpace(value) == "" {
return fmt.Errorf("%s has incomplete identity", trial.Name)
}
}
if trial.TotalRequests <= 0 ||
trial.AcceptedOutputs < 0 ||
trial.AcceptedOutputs > trial.CompletedRequests ||
trial.CompletedRequests > trial.TotalRequests ||
trial.CriticalFailures < 0 ||
trial.DurationSeconds <= 0 {
return fmt.Errorf("%s has invalid counters", trial.Name)
}
metrics := []float64{
trial.DurationSeconds, trial.TTFTP95MS, trial.TPOTP95MS,
trial.PeakMemoryGB, trial.CommunicationShare,
}
for _, value := range metrics {
if math.IsNaN(value) || math.IsInf(value, 0) || value < 0 {
return fmt.Errorf("%s has an invalid metric", trial.Name)
}
}
if trial.CommunicationShare > 1 {
return fmt.Errorf("%s has an invalid communication share", trial.Name)
}
acceptance := float64(trial.AcceptedOutputs) / float64(trial.TotalRequests)
if trial.CriticalFailures > 0 || acceptance < minAcceptance ||
trial.TTFTP95MS > maxTTFT || trial.TPOTP95MS > maxTPOT {
return fmt.Errorf("%s failed a release gate", trial.Name)
}
return nil
}
func compare(dense, moe Trial) (string, float64, float64, error) {
paired := dense.TokenizerRevision == moe.TokenizerRevision &&
dense.Precision == moe.Precision &&
dense.WorkloadRevision == moe.WorkloadRevision &&
dense.HardwareRevision == moe.HardwareRevision &&
dense.GateRevision == moe.GateRevision
if !paired {
return "", 0, 0, fmt.Errorf("trials are not paired")
}
for _, trial := range []Trial{dense, moe} {
if err := validateTrial(trial, 0.85, 900, 45); err != nil {
return "", 0, 0, err
}
}
denseGoodput := float64(dense.AcceptedOutputs) / dense.DurationSeconds
moeGoodput := float64(moe.AcceptedOutputs) / moe.DurationSeconds
winner := dense.Name
if moeGoodput > denseGoodput {
winner = moe.Name
}
return winner, denseGoodput, moeGoodput, nil
}
func main() {
common := Trial{
TokenizerRevision: "tokenizer/12", Precision: "bf16",
WorkloadRevision: "traffic/44", HardwareRevision: "cluster/7",
GateRevision: "quality-policy/9", TotalRequests: 100,
CompletedRequests: 100, DurationSeconds: 8,
}
dense := common
dense.Name = "dense"
dense.ArtifactRevision = "dense/17"
dense.RuntimeRevision = "runtime/31"
dense.AcceptedOutputs = 88
dense.TTFTP95MS = 640
dense.TPOTP95MS = 31
dense.PeakMemoryGB = 62
moe := common
moe.Name = "moe"
moe.ArtifactRevision = "moe/22"
moe.RuntimeRevision = "runtime/31"
moe.AcceptedOutputs = 92
moe.TTFTP95MS = 710
moe.TPOTP95MS = 34
moe.PeakMemoryGB = 78
moe.CommunicationShare = 0.18
winner, denseGoodput, moeGoodput, err := compare(dense, moe)
if err != nil {
panic(err)
}
fmt.Printf(
"dense_goodput=%.2f moe_goodput=%.2f winner_in_fixture=%s\n",
denseGoodput,
moeGoodput,
winner,
)
}
Expected output: dense_goodput=11.00 moe_goodput=11.50 winner_in_fixture=moe.
The two artifacts and their execution paths can differ because that is what the production selection compares; shared fields prove that the workload and acceptance contract are paired. The winner is only the result of this fixture, not a causal claim that the MoE architecture is universally faster. A useful benchmark reports rejected outputs and failed requests, not only raw tokens per second.
Failure Modes and Diagnostics
| Symptom | Evidence to collect | Possible causes |
|---|---|---|
| A few experts dominate | token counts, router mass, entropy, per-slice routes | router collapse, genuine skew, weak or excessive balance control |
| Tail latency rises | per-expert queue, per-device load, all-to-all time | hot experts, topology crossing, small GEMMs, synchronization |
| Training quality regresses | task slices, routing churn, drop rate, loss terms | overflow, unstable routing, balance objective interference |
| More experts do not help | iso-compute quality curves, utilization, memory | diminishing capacity returns, insufficient data, communication overhead |
| Local runtime is slow | bandwidth, offload bytes, page faults, peak memory | weights do not remain resident, KV-cache pressure, poor kernels |
| Quantized output changes | routing agreement, per-slice quality, calibration | router score perturbation, expert-weight error, runtime mismatch |
Do not diagnose an MoE only from average utilization. Correlate model quality with routing, expert batches, collectives, memory, and request-level latency.
FAQ
Is MoE always faster than a dense model
No. Sparse expert arithmetic may be lower than a same-total-parameter dense design, but communication, memory traffic, small expert batches, imbalance, and shared layers can dominate. Only an end-to-end benchmark under the target workload answers the question.
Must every MoE route two experts per token
No. Published systems use different top-k values, shared experts, grouped routing, and balancing strategies. Top-k is part of a versioned architecture, not an industry constant.
Are experts separate models
Usually not in a Transformer MoE. They are typically FFN modules inside selected layers and share the rest of the model. Their outputs participate in one forward pass.
Does expert capacity mean inference must drop tokens
No. Capacity and token dropping are implementation choices. Dropless kernels and variable-size expert batches exist, and some architectures explicitly avoid dropping. Record the actual runtime behavior.
Can active parameter count compare model quality
No. Quality depends on training data, optimization, architecture, total capacity, active path, tokenizer, post-training, and evaluation protocol. Compare quality directly on representative slices.
Summary
MoE turns some dense computation into conditional computation:
route
-> group and dispatch
-> execute selected experts
-> combine and restore
Its value comes from increasing model capacity without activating every expert for every token. Its cost appears in weight residency, routing stability, expert load, collective communication, kernel efficiency, and operational complexity. Treat total parameters, active parameters, FLOPs, memory, latency, throughput, cost, and quality as separate measurements.