Theme
MoE and Ultra-Large-Scale Models
MoE (Mixture of Experts) is a sparse architecture that replaces the feedforward layers in a network with "multiple experts + a router," letting each token activate only a few experts. It breaks the tight coupling of "model capability ≈ parameters × compute": total parameters can be very large (lots of memory), but each inference step computes only a small fraction (cost is controllable). MoE is the common choice for virtually every ultra-large-scale model post-2024 (DeepSeek-V3, Qwen3, Llama 4, Grok-1, and even GPT-4 as rumored). This page covers its mechanisms, milestones, and deployment practices; for mechanism-level knowledge, see MoE Sparse Expert Models.
I. Why MoE Is Needed: A Cost Contradiction
In dense models, parameter count is proportional to compute: to have 1 trillion parameters, you must compute 1 trillion multiply-accumulate operations per token, making training and inference costs unbearable. But large models need "lots of knowledge" (corresponding to total parameters), while each individual token usually only needs "a small portion of capability" (translating a word doesn't require full mathematical ability).
MoE's idea is: prepare many "experts" (each a group of parameters), and let the router pick 2–8 most relevant experts per token to compute.
Dense model: token → [ 1 giant FFN ] → all 175B parameters participate in computation
MoE model: token → router → pick Top-2 experts → only ~20% parameters activated
↘ remaining experts are skipped (not computed)The gain is decoupling "parameters × compute": DeepSeek-V3 has 671B total parameters, but activates only 37B per token — training compute approaches that of a 37B dense model, yet capability nears 671B's "knowledge capacity."
Here's a common terminology issue easily misled by names: Mixtral 8×7B does not mean "8 × 7B = 56B total." 8×7B means each layer's FFN has 8 experts, each roughly 7B in FFN parameters; due to shared parts like attention, total parameters are 46.7B, not 56B. Similarly, DeepSeek-V3's "671B" is all parameters (including shared and routing), while "37B" is what each token actually computes. When looking at MoE models, always distinguish between "total parameters" and "activated parameters" — otherwise comparisons lead to erroneous conclusions.
II. Core Technology: Router, Top-k, and Load Balancing
1. Router (Gating)
The router is a small linear layer: it takes a token's representation and outputs "affinity scores for each expert," applies softmax to get a distribution, and selects the top-k experts:
h = router(token) # score vector, length = number of experts E
p = softmax(h) # normalize
top_k = argsort(p)[:k] # pick top k experts
output = Σ p[i] * Expert_i(token) # weighted sum2. Top-1 vs Top-2 vs Top-8
| Approach | Representative Model | Characteristics |
|---|---|---|
| Top-1 | Switch Transformer (1.6T) | Cheapest compute, but most severe load imbalance |
| Top-2 | Mixtral 8×7B, Grok-1, GLaM | Smoother, higher utilization, mainstream in community |
| Top-8 + shared expert | DeepSeek-V3 (256 routing experts + 1 shared expert) | Finer-grained experts, more flexible activation |
| Top-1 + 2 shared experts | Qwen3-235B-A22B | Shared experts guarantee general capability |
3. Load Balancing: The Life-or-Death Line of MoE Training
The router has a "collapse" risk: once a few experts are always selected, the rest starve, and MoE degrades into a dense model. Mainstream countermeasures:
- Auxiliary load balancing loss: penalizes "routing distribution deviating from uniform";
- Capacity factor: limits how many tokens each expert can process at once, excess tokens take a residual shortcut (Switch's approach);
- No-auxiliary-loss balancing (DeepSeek-V3's contribution): adds bias to each expert, dynamically adjusting selection probability, eliminating the performance loss from auxiliary loss.
There's also a subtle tradeoff in load balancing between "capability vs. efficiency": strong load-balancing constraints push routing toward "averaging," suppressing expert specialization (each expert learns similar capabilities); without constraints, a few experts may monopolize. The ideal state is "specialized but not imbalanced" — DeepSeek's no-auxiliary-loss balancing approximated this via dynamic bias. In practice, the load-balancing loss weight is a hyperparameter requiring careful tuning: too large hurts capability, too small breaks training. This also explains why MoE training demands more engineering experience.
MoE isn't "free parameters"
Large total parameters don't automatically mean stronger capability. If the router learns poorly or experts are redundant, MoE may only be slightly better than a dense model of the same compute. The efficiency of activated parameters, expert diversity, and load balance are the true determinants of MoE quality.
III. Milestone Timeline
| Time | Work | Key Point |
|---|---|---|
| 2017 | Shazeer et al. sparse-gated MoE | First MoE layer applied to LSTM language model (1024 experts), but hard to scale engineeringly |
| 2020.06 | GShard | Brought MoE to Transformer, Top-2 gating, conditional compute + automatic sharding |
| 2021.01 | Switch Transformer | Top-1 routing simplified, up to 1.6T total parameters, ~7× faster pretraining |
| 2021.12 | GLaM | 1.2T total / ~97B activated (64 experts), training cost ~1/3 of GPT-3, comparable to GPT-3 on 29 benchmarks |
| 2023.12 | Mixtral 8×7B | First "truly usable" open-source MoE: 46.7B total / 12.9B activated, Apache 2.0 |
| 2024.05 | DeepSeek-V2 | MLA + DeepSeekMoE, pushed MoE cost revolution to market |
| 2024.12 | DeepSeek-V3 | 671B total / 37B activated, 256 experts + 1 shared, FP8 training, ~2000 H800s trained, went viral globally |
| 2025.04 | Qwen3-235B-A22B | 235B total / 22B activated, mixed reasoning mode (thoughtful/non-thoughtful) |
| Unconfirmed | GPT-4 (rumored) | Widely believed to use MoE (rumored 8 experts, ~1.8T total), not confirmed by OpenAI |
Looking at the timeline together, MoE milestones evolve along two threads: one is the simplification and robustness of "routing mechanisms" (Top-2 → Top-1 → Top-8+shared → no-auxiliary-loss), and the other is the advancement of "engineering" (GShard's parallelism, vLLM's inference support, FP8 low-precision training). Mechanism innovation lowers the training barrier, engineering innovation lowers the deployment barrier — both together moved MoE from papers to production. This gives practitioners one lesson: when an architecture "looks better" but doesn't go mainstream, suspect the engineering support before the architecture itself — MoE spent years waiting for its vLLM.
Milestone One: Switch Transformer — "The 1.6 Trillion Parameter Power-Saving Solution"
A brief history of failure first: MoE hasn't been a smooth ride. Shazeer et al.'s 2017 sparse-gated MoE was hard to scale engineeringly and was once considered "flashy but impractical"; it wasn't until GShard and Switch Transformer solved routing load and parallelism that MoE made a comeback. This history reminds us: the adoption speed of architectural innovation depends on engineering support — load balancing, parallel communication, inference framework support — miss one link and a good architecture dies young. Today's MoE maturity is the result of years of engineering accumulation, not a single paper's credit.
Switch Transformer used Top-1 routing to scale MoE to 1.6T parameters, improving pretraining speed by ~7×. Its core contribution was proving MoE can significantly reduce cost per unit of capability without losing quality — "sparse activation" went from academic toy to a scalable engineering option.
Milestone Two: Mixtral 8×7B — The Open-Source MoE Ignition Point
In December 2023, Mistral released Mixtral 8×7B: FFN layers replaced with 8 experts, 2 activated per token, 46.7B total parameters activating only 12.9B, yet matching the performance of 70B-class dense models (like Llama 2 70B) at the time — and fully open under Apache 2.0. It made "70B-level capability at 1/5 the inference cost" a reality, and the open-source community turned completely to MoE.
Milestone Three: DeepSeek-V3 — The MoE Cost Revolution
DeepSeek-V3 pushed MoE to new heights in December 2024:
| Dimension | Value |
|---|---|
| Total parameters | 671B |
| Activated parameters | 37B (per token) |
| Expert config | 256 fine-grained routing experts + 1 shared expert, 8 activated per token |
| Attention | MLA (multi-head latent attention, significantly compressing KV cache) |
| Training precision | FP8 mixed precision |
| Training data | 14.8T tokens |
| Training cost | ~2.788M H800 GPU-hours (publicly disclosed) |
Its innovation was "fine-grained experts + shared experts": experts split finer, more activated (8), combined more flexibly; the shared expert specifically handles general knowledge (grammar, commonsense), while routing experts carry domain capability. Paired with MLA and FP8, DeepSeek-V3 reached flagship-level capability at far less compute than competitors, turning "MoE = big lab exclusive" into "innovators can play too."
Milestone Four: Qwen3-MoE and Llama 4
- Qwen3-235B-A22B (2025.4): 235B total / 22B activated, activates 2 shared + 1 routing expert, supports "thoughtful/non-thoughtful" mixed reasoning mode, stands out on instruction following and programming benchmarks;
- Llama 4 (2025.4): Meta's first shift to MoE — Scout (109B-A17B) and Maverick (400B-A17B), activating ~17B.
IV. Comparison Across Schemes
| Model | Total Params | Activated Params | Expert Config | Routing | Highlights |
|---|---|---|---|---|---|
| Switch Transformer | 1.6T | ~tens of B | Multiple experts per FFN | Top-1 | First ultra-large MoE validation |
| GLaM | 1.2T | ~97B | 64 experts | Top-2 | Training cost only 1/3 of GPT-3 |
| Mixtral 8×7B | 46.7B | 12.9B | 8 experts | Top-2 | Open-source MoE benchmark |
| Mixtral 8×22B | 141B | 39B | 8 experts | Top-2 | Larger open-source MoE |
| DeepSeek-V3 | 671B | 37B | 256+1 experts | Top-8 | Fine-grained experts + no-aux-loss balance |
| Qwen3-235B-A22B | 235B | 22B | Routing + 2 shared | Top-1+shared | Mixed reasoning, bilingual |
| Grok-1 | 314B | 86B | 8 experts | Top-2 | xAI's first open-source |
| GPT-4 (rumored) | ~1.8T | Not disclosed | 8 experts (rumored) | Not disclosed | Unconfirmed by official |
The parameters in the comparison table are "paper numbers" — the real difference is in the "capability/cost" curve: even with the same 671B total parameters, different implementations' routing efficiency, data quality, and training duration create huge gaps. The correct approach to evaluating MoE models: fix the "activated parameter budget" (e.g., 20B-class), compare different models horizontally on the same benchmarks and prompts, and measure actual throughput and latency. Parameter tables only say "how much material there is"; evaluation and benchmarking say "how much work gets done." Evaluation methods in Evaluation and Benchmarks.
V. MoE's Cost-Effectiveness: Activated Parameters Are the Bill
| Perspective | Total Parameters (memory capacity) | Activated Parameters (per-token compute) |
|---|---|---|
| Training cost | Affects VRAM, communication | Determines FLOPs (≈ activated params × token count) |
| Inference latency | Affects model loading, VRAM residency | Determines forward pass computation time |
| Capability ceiling | More = more knowledge | Higher = finer per-token processing |
Conclusion: training cost is mainly about "activated params × data volume," the capability ceiling is mainly about "total params + data quality," and inference latency is affected by both. MoE's value is exchanging "put more params in, compute fewer params" for training cost-effectiveness; on the inference side (below), a new problem emerges: "saved compute, spent VRAM."
MoE's cost-effectiveness logic differs between training and inference, and many people confuse the two stages. On the training side, MoE's gain comes from "each token only activates some experts" — under the same FLOPs budget, it can accommodate a larger knowledge capacity, so training cost is calculated by activated params, an order of magnitude lower than equivalent dense models. On the inference side, compute is also determined by activated params, but VRAM and bandwidth are determined by total params — if you only have one GPU, 671B of MoE simply won't fit, no matter how "compute-saving" it is. This explains why MoE is especially suited for the "trainer with limited budget + server with ample resources" pattern (saves training, spreads VRAM across a large cluster) — and it's why teams like DeepSeek, "limited compute but strong engineering," can build flagship MoE models. Understanding that you need to balance the books separately on both sides of training/inference is the first step in evaluating any MoE scheme.
One more emphasis: "large total params" is the indicator for knowledge capacity, "small activated params" is the indicator for cost advantage — when selecting, calculate both numbers along with your VRAM and concurrency, and the MoE equation becomes clear.
For "is MoE worth it," the practice community has a consensus answer: on the training side, MoE almost always wins — same compute yields larger knowledge capacity; on the inference side, the answer depends on the scenario — high concurrency, large model, long-running services: MoE's amortized cost is lower; low concurrency, single-GPU, latency-sensitive deployment: dense is simpler and more reliable. Another underappreciated point is ecosystem maturity: post-2024, native MoE support from frameworks like vLLM has significantly reduced deployment complexity — the old impression that "MoE is hard to deploy" is changing. See Framework and Tool Selection.
VI. MoE Deployment Practice and Challenges
MoE deployment isn't just about "saving money" — it comes with a set of unique engineering problems (see Deployment and Servicing for details):
| Challenge | Cause | Common Solutions |
|---|---|---|
| VRAM explosion | All expert weights must reside in VRAM (regardless of whether activated) | Multi-layer quantization, expert offloading to CPU/SSD |
| Communication overhead | Experts distributed across GPUs, tokens need cross-GPU forwarding (All-to-All) | Expert parallelism (EP), topology-aware scheduling |
| Load imbalance | Routing hotspots cause some GPUs to queue | Load-balanced training + dynamic batching |
| Batching utilization | Single request activates few experts, GPU compute wasted | Continuous batching to stack throughput |
| KV cache | Attention cache still stored per total layer count for long context | MLA (DeepSeek's solution) |
1. Expert Parallelism (EP)
Split experts across different GPUs, forward tokens to the corresponding card per routing result for computation, then aggregate — this is the standard parallel strategy for MoE inference. EP's throughput ceiling is constrained by the bottleneck of the "hottest expert's card."
2. Inference Framework Support
vLLM, SGLang, TensorRT-LLM, llama.cpp, and other mainstream engines now natively support MoE weight loading and EP scheduling; local 8×7B-class MoE can run on consumer GPUs (total weights of 46.7B need ~30 GB VRAM, reducible with quantization).
One more point to separately emphasize about MoE and quantization: MoE models are more sensitive to quantization, because the router layer's numerical stability matters — router scoring determines which expert a token goes to, and quantization noise can change routing decisions, leading to "misrouted experts." Therefore, MoE quantization typically requires: keeping the router layer at high precision, per-expert calibration, and post-quantization verification that the routing distribution hasn't drifted. These details make MoE quantization deployment more demanding than dense models. Relevant tools in Deployment and Servicing.
3. When to Choose MoE
| Scenario | Advice |
|---|---|
| Need huge knowledge capacity, limited budget | MoE (saves training compute) |
| Single-GPU low-VRAM deployment, latency-sensitive | Dense small model or quantized mid-range MoE |
| High-concurrency API service | MoE + EP + continuous batching (best cost) |
| Research/control priority | Dense models are simpler and easier to tune |
MoE services need more monitoring than dense models: in addition to standard latency, throughput, and token consumption, also monitor "routing distribution" — if some experts have long-term low utilization, routing may be degrading or load imbalance worsening; also watch VRAM occupancy (total params resident) and KV cache peaks. Integrating these metrics into alerts lets you spot problems before "quality silently degrades." General monitoring methods in Deployment and Servicing.
VII. MoE Training Details and Engineering Essentials
1. Three Engineering Points of Router Training
| Engineering Point | Problem | Common Practice |
|---|---|---|
| Router initialization | Initial uniform random → uneven expert capability | Small random + uniform initialization for router, avoid early collapse |
| Load balancing | Hotspot experts over-selected | Auxiliary load balance loss, capacity factor, dynamic bias |
| Expert division of labor | Expert redundancy, knowledge overlap | Fine-grained experts, shared experts handle general capability (DeepSeek approach) |
Load balancing mechanism details in MoE Sparse Expert Models.
2. Comparison with Dense Training
| Dimension | Dense Model | MoE |
|---|---|---|
| Per-token compute | All parameters | Only activated portion (e.g., 5%–20%) |
| VRAM requirement | Small weights but heavy compute | All weights resident in VRAM, VRAM actually larger |
| Communication | Standard tensor/data parallelism | Needs expert parallelism + All-to-All |
| Convergence | Stable | Needs load balance tuning, slightly unstable |
| Cost-effectiveness | Capability linear with cost | Higher capability per same cost (determined by activated params) |
3. Representative Model Training Config Comparison
| Model | Total/Activated Params | Training Data | Context | Activated Experts |
|---|---|---|---|---|
| Mixtral 8×7B | 46.7B / 12.9B | ~8T tokens | 32K | 2 / 8 |
| Mixtral 8×22B | 141B / 39B | Multilingual enhanced | 64K | 2 / 8 |
| DeepSeek-V3 | 671B / 37B | 14.8T tokens | 128K | 8 / 256 |
| Qwen3-235B-A22B | 235B / 22B | Large-scale multilingual | 32K–131K | Routing + 2 shared |
| Grok-1 | 314B / 86B | Not disclosed | 8K | 2 / 8 |
Numbers in the table are subject to each official release; "context" refers to the main version capability at release.
4. MoE Forward Pseudocode
# Simplified MoE layer forward pass (per token)
def moe_forward(x, router, experts, k=2):
logits = router(x) # [E]
p = softmax(logits) # normalize
top_k_idx = argsort(p)[-k:] # pick top k experts
out = zeros_like(x)
for i in top_k_idx:
out += p[i] * experts[i](x) # weighted sum
return outThree MoE training tips
① Secure load balance before discussing results; ② shared experts + fine-grained experts is the new mainstream; ③ evaluation must simultaneously look at "activated parameter cost-effectiveness" not just total parameters. Data and loss-related mechanisms in Pretraining: Data and Objectives and Scaling Laws.
VIII. MoE and the Future
- Test-time scaling: MoE can stack with test-time scaling — large params + sparse activation + long chain-of-thought is the combo punch of 2025 flagship models;
- Finer granularity: expert count continues rising (DeepSeek-V3 already has 256+1), routing precision and knowledge division getting finer;
- Combining with multimodal: vision experts, language experts set by domain, making MoE the universal base for multimodal LLMs (see Multimodal LLMs);
- MoE vs. Dense: short-term, MoE is mainstream for ultra-large scale; but dense models retain irreplaceable positions in low-VRAM, low-latency, and interpretability scenarios. Selection principles in Scaling Laws and Model Compendium.
One sentence to remember MoE
MoE = decoupling of total parameters (capacity) and activated parameters (cost). It turned "trillion parameters" from a gimmick into an engineering reality, and let open-source models compete head-to-head with closed-source flagships on cost for the first time.
IX. MoE FAQ and Selection Checklist
1. Common MoE Misconceptions
| Misconception | Truth |
|---|---|
| "Larger total params = smarter" | Capability depends on activated param efficiency, routing quality, and data |
| "MoE always saves VRAM" | Quite the opposite — all weights reside in VRAM |
| "MoE free capability upgrade" | Training needs load balance tuning, higher engineering complexity |
| "MoE inference necessarily faster" | Single request may be slower; throughput comes from batching/parallelism |
| "More experts is better" | Too many experts worsen load imbalance and communication |
2. When Not to Use MoE
| Scenario | Reason |
|---|---|
| Single-GPU / low-VRAM deployment | Total weights too large to fit |
| Ultra-low-latency single request | Sparse routing not as efficient as dense direct |
| Team without distributed experience | Expert parallelism debugging threshold is high |
| Need minimal interpretability | Dense models are easier to analyze |
| Prototype/validation stage | Get it running first, then optimize cost |
3. How to Evaluate an MoE Model
| Metric | Ask | Look At |
|---|---|---|
| Activated parameter cost-effectiveness | Is capability higher at same cost? | Activated params × token count vs. benchmark score |
| Routing quality | Do each expert specialize in their role? | Routing distribution stats, expert utilization |
| Batching throughput | How many token/s at high concurrency? | Benchmark reports, vLLM live tests |
| Long-context performance | Does it drop at 128K? | Needle-in-a-haystack, LongBench-type evals |
| Quantization robustness | How many points dropped at 4-bit? | Pre- and post-quantization comparison |
4. Migration Checklist: Dense to MoE
- [ ] Confirm the real bottleneck is training compute or inference cost
- [ ] Compare MoE vs. dense throughput at same VRAM
- [ ] Verify inference framework (vLLM/SGLang) support for this MoE
- [ ] Evaluate quantization scheme and precision loss
- [ ] Build cost model: include activated params, KV cache, communication overhead
- [ ] Is the long-term maintenance and upgrade path clear?
5. FAQ Quick Answers
| Question | Quick Answer |
|---|---|
| Is Mixtral 8×7B 56B? | No, total 46.7B, activated 12.9B |
| How many cards for DeepSeek-V3 training? | Officially disclosed ~2000 H800s (subject to official) |
| Can a single GPU run MoE? | Small/medium MoE quantized can, but throughput limited |
| What are shared experts? | All-token-passing general experts handling grammar/commonsense |
| What is load balancing loss? | Auxiliary loss penalizing routing bias, ensuring all experts are utilized |
| Is GPT-4 an MoE? | Widely speculated in the community, OpenAI didn't confirm |
One sentence to remember MoE selection
MoE buys "training cost-effectiveness" and sells "VRAM and engineering complexity." Calculate the "activated params × data volume" equation before jumping in.
One last note: the "cost-effectiveness" label for MoE holds only at sufficient scale — in small scenarios, dense models are often more trouble-free.
X. MoE Deep Dive: From Papers to Engineering
1. Key Papers at a Glance
| Paper / Work | Year | Core Contribution |
|---|---|---|
| Sparsely-Gated MoE Layer | 2017 | MoE layer founding, first applied to LM |
| GShard | 2020 | MoE + Transformer, Top-2 routing |
| Switch Transformer | 2021 | Top-1 routing, 1.6T params |
| GLaM | 2021 | 1.2T total / 97B activated, training cost ~1/3 |
| ST-MoE | 2022 | Systematic study of routing load balance and expert diversity |
| Mixtral of Experts | 2024 | Open-source MoE landing benchmark |
| DeepSeek-V2 / V3 | 2024 | MLA + fine-grained experts + no-aux-loss balance |
2. How Gradients Flow Through the Router
A core MoE training detail: "routing is discontinuous" — expert selection is discrete (Top-k), and gradients can't be directly computed. Mainstream handling:
- Router soft weights backpropagate directly: gradient updates on the selected experts' weights p[i];
- Expert parameters update independently: only selected experts receive gradients, others stay unchanged;
- Load balancing loss gradients: penalize routing distribution deviating from uniform, driving expert division of labor.
This creates one phenomenon: when expert utilization is low, gradients are sparse and training slows — so load balance is both a stability problem and a training efficiency problem.
A frequently asked question: why does "let each token pick several experts" learn different capabilities? Intuitively: the router and experts are jointly trained, and experts only receive gradients when selected, so they're only "responsible" for "the types of tokens that select them" — the router learns to route math-type tokens to the math expert, the math expert gets stronger in math accordingly, creating a positive feedback loop. It's a bit like social division of labor: the clearer the division, the more specialized each point, but the more dependent on the router's "dispatch." If the router fails (load imbalance), the division collapses — which is precisely why load balancing matters so much. Understanding this "division + dispatch" metaphor lets you predict MoE's strengths (large capacity with diverse domains) and weaknesses (routing errors and batch utilization).
3. Experimental Observations on Expert Diversity
Engineering and research communities have observed several common phenomena:
- Experts aren't fully specialized — there's significant "general expert" usage (almost all tokens use them) — which is the design motivation for shared experts;
- Higher layers → more expert specialization (lower layers handle lexical, higher layers handle semantic);
- Routing quality improves with training but may overfit to the training distribution; monitor routing distribution drift at inference time.
4. Key Empirical Comparison Points: MoE vs. Dense
| Comparison Experiment | Key Finding |
|---|---|
| Same FLOPs | MoE usually beats dense (sparse activation gain) |
| Same total params | Dense is more stable; MoE's cost-effectiveness is in "capability/cost" |
| Small batch | MoE advantage shrinks (low batching utilization) |
| Long context | MoE's KV cache matches dense; advantage still in activated params |
| Post-quantization | MoE more sensitive to quantization, needs targeted calibration |
5. Two Counterintuitive Facts
Two of the most counterintuitive points about MoE. First, MoE's "total params" barely affect single-token inference latency (unless VRAM bandwidth is constrained); what really determines latency is activated params and KV cache — so "671B total params" is primarily about knowledge capacity, not speed. Second, MoE's advantage is not obvious in "low-concurrency" scenarios because sparse routing means each single request only uses a few experts, and GPUs can't be fully fed — while high-concurrency batching can stagger expert demands across requests, maximizing utilization. Therefore, evaluating MoE always requires benchmarking under "your actual concurrency profile," not looking at single-request benchmarks. These two points directly determine deployment selection (see Deployment and Servicing).
6. Quick Memory Formula Judgment
Quickly estimate whether an MoE model can be deployed: VRAM ≈ total params × bytes per param + KV cache + activations. Because MoE's total params far exceed activated params, the VRAM bottleneck is almost always "total params." Example: Mixtral 8×7B total 46.7B, FP16 needs ~93 GB, quantized to 4-bit ~25 GB — a single 24 GB card can barely run it with quantization and offloading; DeepSeek-V3 total 671B, FP16 ~1.3 TB, only deployable on clusters, quantized still needs 300 GB+. Calculate the total params bill first, then look at activated params efficiency — calculating these two bills separately is MoE deployment's first lesson.
XI. MoE Resource Checklist
| Resource | Use |
|---|---|
| DeepSeek-V3 / R1 technical report | Complete engineering details of MoE + MLA + RL |
| Mixtral paper and Mistral blog | Open-source MoE implementation reference |
| vLLM / SGLang docs | MoE deployment and expert parallelism |
| HF Open LLM Leaderboard | Open-source MoE leaderboard comparison |
| nanoMoE and community reproduction projects | Hands-on MoE layer from scratch |
Final reminder: the resource checklist is just the starting point. To truly understand MoE, you must run it yourself — deploy a Mixtral with vLLM, observe routing distribution and throughput, then compare cost-effectiveness with same-VRAM dense models. The gap between reading about it and hands-on verification is especially large with MoE.
The order to read MoE papers
First read Switch Transformer to build intuition → then read Mixtral to see landing → finally read DeepSeek-V3 for engineering limits. After reading these three, MoE's mechanisms, engineering, and costs all click.
A supplement to selection advice: if your team has no distributed training experience, the first time touching MoE should start with "directly using open-source MoE models" (Mixtral, Qwen3-235B-type), rather than training from scratch — first experience its deployment and cost characteristics, then decide whether to invest on the training side. MoE optimization on the training side (routing, load balancing, FP8) is deep water — without a billion-token experiment budget, it's hard to experience all its complexity.
XII. Further Reading
- MoE Sparse Expert Models — detailed routing and load balancing mechanisms
- Scaling Laws — the "parameter-compute" power law MoE breaks and its boundaries
- Llama and the Open-Source Ecosystem — MoE landing in open-source models
- Deployment and Servicing — expert parallelism, quantization, and batching
- Model Compendium — MoE vs. dense model selection comparison
- Frontier Progress — MoE and test-time scaling frontier trends
References
- Shazeer et al. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer (2017) — MoE founding paper (arXiv)
- Lepikhin et al. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding (2020) — GShard paper (arXiv)
- Fedus et al. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity (2021) — Switch Transformer paper (arXiv)
- Du et al. GLaM: Efficient Scaling of Language Models with Mixture-of-Experts (2021) — GLaM paper (arXiv)
- Jiang et al. Mixtral of Experts (2024) — Mixtral 8×7B paper (arXiv)
- DeepSeek-AI. DeepSeek-V3 Technical Report (2024.12) — DeepSeek-V3 paper (arXiv)
- Qwen Team. Qwen3 Technical Report (2025) — Qwen3 technical report (arXiv)
- xAI. Grok-1 open-source repo — Grok-1 weights and documentation