Skip to content

MoE and Ultra-Large-Scale Models

At a glance Mixture of Experts (MoE) breaks the strong coupling between parameter count and compute via "sparse activation," making trillions of total parameters feasible. This article breaks down MoE routing/load-balancing mechanisms, traces milestones from GShard → Switch → GLaM → Mixtral → DeepSeek-V3, and compares activated parameters and deployment challenges.

This page contains time-sensitive content, current as of 2025-06; job descriptions, rankings, product features, and other information may have changed. Please verify with original sources before citing.

MoE and Ultra-Large-Scale Models ​

MoE (Mixture of Experts) is a sparse architecture that replaces the feedforward layers in a network with "multiple experts + a router," letting each token activate only a few experts. It breaks the tight coupling of "model capability ≈ parameters × compute": total parameters can be very large (lots of memory), but each inference step computes only a small fraction (cost is controllable). MoE is the common choice for virtually every ultra-large-scale model post-2024 (DeepSeek-V3, Qwen3, Llama 4, Grok-1, and even GPT-4 as rumored). This page covers its mechanisms, milestones, and deployment practices; for mechanism-level knowledge, see MoE Sparse Expert Models.

I. Why MoE Is Needed: A Cost Contradiction ​

In dense models, parameter count is proportional to compute: to have 1 trillion parameters, you must compute 1 trillion multiply-accumulate operations per token, making training and inference costs unbearable. But large models need "lots of knowledge" (corresponding to total parameters), while each individual token usually only needs "a small portion of capability" (translating a word doesn't require full mathematical ability).

MoE's idea is: prepare many "experts" (each a group of parameters), and let the router pick 2–8 most relevant experts per token to compute.

Dense model:  token → [ 1 giant FFN ] → all 175B parameters participate in computation
MoE model:    token → router → pick Top-2 experts → only ~20% parameters activated
                      ↘ remaining experts are skipped (not computed)

The gain is decoupling "parameters × compute": DeepSeek-V3 has 671B total parameters, but activates only 37B per token — training compute approaches that of a 37B dense model, yet capability nears 671B's "knowledge capacity."

Here's a common terminology issue easily misled by names: Mixtral 8×7B does not mean "8 × 7B = 56B total." 8×7B means each layer's FFN has 8 experts, each roughly 7B in FFN parameters; due to shared parts like attention, total parameters are 46.7B, not 56B. Similarly, DeepSeek-V3's "671B" is all parameters (including shared and routing), while "37B" is what each token actually computes. When looking at MoE models, always distinguish between "total parameters" and "activated parameters" — otherwise comparisons lead to erroneous conclusions.

II. Core Technology: Router, Top-k, and Load Balancing ​

1. Router (Gating) ​

The router is a small linear layer: it takes a token's representation and outputs "affinity scores for each expert," applies softmax to get a distribution, and selects the top-k experts:

h = router(token)                    # score vector, length = number of experts E
p = softmax(h)                       # normalize
top_k = argsort(p)[:k]               # pick top k experts
output = Σ p[i] * Expert_i(token)    # weighted sum

2. Top-1 vs Top-2 vs Top-8 ​

ApproachRepresentative ModelCharacteristics
Top-1Switch Transformer (1.6T)Cheapest compute, but most severe load imbalance
Top-2Mixtral 8×7B, Grok-1, GLaMSmoother, higher utilization, mainstream in community
Top-8 + shared expertDeepSeek-V3 (256 routing experts + 1 shared expert)Finer-grained experts, more flexible activation
Top-1 + 2 shared expertsQwen3-235B-A22BShared experts guarantee general capability

3. Load Balancing: The Life-or-Death Line of MoE Training ​

The router has a "collapse" risk: once a few experts are always selected, the rest starve, and MoE degrades into a dense model. Mainstream countermeasures:

  • Auxiliary load balancing loss: penalizes "routing distribution deviating from uniform";
  • Capacity factor: limits how many tokens each expert can process at once, excess tokens take a residual shortcut (Switch's approach);
  • No-auxiliary-loss balancing (DeepSeek-V3's contribution): adds bias to each expert, dynamically adjusting selection probability, eliminating the performance loss from auxiliary loss.

There's also a subtle tradeoff in load balancing between "capability vs. efficiency": strong load-balancing constraints push routing toward "averaging," suppressing expert specialization (each expert learns similar capabilities); without constraints, a few experts may monopolize. The ideal state is "specialized but not imbalanced" — DeepSeek's no-auxiliary-loss balancing approximated this via dynamic bias. In practice, the load-balancing loss weight is a hyperparameter requiring careful tuning: too large hurts capability, too small breaks training. This also explains why MoE training demands more engineering experience.

MoE isn't "free parameters"

Large total parameters don't automatically mean stronger capability. If the router learns poorly or experts are redundant, MoE may only be slightly better than a dense model of the same compute. The efficiency of activated parameters, expert diversity, and load balance are the true determinants of MoE quality.

III. Milestone Timeline ​

TimeWorkKey Point
2017Shazeer et al. sparse-gated MoEFirst MoE layer applied to LSTM language model (1024 experts), but hard to scale engineeringly
2020.06GShardBrought MoE to Transformer, Top-2 gating, conditional compute + automatic sharding
2021.01Switch TransformerTop-1 routing simplified, up to 1.6T total parameters, ~7× faster pretraining
2021.12GLaM1.2T total / ~97B activated (64 experts), training cost ~1/3 of GPT-3, comparable to GPT-3 on 29 benchmarks
2023.12Mixtral 8×7BFirst "truly usable" open-source MoE: 46.7B total / 12.9B activated, Apache 2.0
2024.05DeepSeek-V2MLA + DeepSeekMoE, pushed MoE cost revolution to market
2024.12DeepSeek-V3671B total / 37B activated, 256 experts + 1 shared, FP8 training, ~2000 H800s trained, went viral globally
2025.04Qwen3-235B-A22B235B total / 22B activated, mixed reasoning mode (thoughtful/non-thoughtful)
UnconfirmedGPT-4 (rumored)Widely believed to use MoE (rumored 8 experts, ~1.8T total), not confirmed by OpenAI

Looking at the timeline together, MoE milestones evolve along two threads: one is the simplification and robustness of "routing mechanisms" (Top-2 → Top-1 → Top-8+shared → no-auxiliary-loss), and the other is the advancement of "engineering" (GShard's parallelism, vLLM's inference support, FP8 low-precision training). Mechanism innovation lowers the training barrier, engineering innovation lowers the deployment barrier — both together moved MoE from papers to production. This gives practitioners one lesson: when an architecture "looks better" but doesn't go mainstream, suspect the engineering support before the architecture itself — MoE spent years waiting for its vLLM.

Milestone One: Switch Transformer — "The 1.6 Trillion Parameter Power-Saving Solution" ​

A brief history of failure first: MoE hasn't been a smooth ride. Shazeer et al.'s 2017 sparse-gated MoE was hard to scale engineeringly and was once considered "flashy but impractical"; it wasn't until GShard and Switch Transformer solved routing load and parallelism that MoE made a comeback. This history reminds us: the adoption speed of architectural innovation depends on engineering support — load balancing, parallel communication, inference framework support — miss one link and a good architecture dies young. Today's MoE maturity is the result of years of engineering accumulation, not a single paper's credit.

Switch Transformer used Top-1 routing to scale MoE to 1.6T parameters, improving pretraining speed by ~7×. Its core contribution was proving MoE can significantly reduce cost per unit of capability without losing quality — "sparse activation" went from academic toy to a scalable engineering option.

Milestone Two: Mixtral 8×7B — The Open-Source MoE Ignition Point ​

In December 2023, Mistral released Mixtral 8×7B: FFN layers replaced with 8 experts, 2 activated per token, 46.7B total parameters activating only 12.9B, yet matching the performance of 70B-class dense models (like Llama 2 70B) at the time — and fully open under Apache 2.0. It made "70B-level capability at 1/5 the inference cost" a reality, and the open-source community turned completely to MoE.

Milestone Three: DeepSeek-V3 — The MoE Cost Revolution ​

DeepSeek-V3 pushed MoE to new heights in December 2024:

DimensionValue
Total parameters671B
Activated parameters37B (per token)
Expert config256 fine-grained routing experts + 1 shared expert, 8 activated per token
AttentionMLA (multi-head latent attention, significantly compressing KV cache)
Training precisionFP8 mixed precision
Training data14.8T tokens
Training cost~2.788M H800 GPU-hours (publicly disclosed)

Its innovation was "fine-grained experts + shared experts": experts split finer, more activated (8), combined more flexibly; the shared expert specifically handles general knowledge (grammar, commonsense), while routing experts carry domain capability. Paired with MLA and FP8, DeepSeek-V3 reached flagship-level capability at far less compute than competitors, turning "MoE = big lab exclusive" into "innovators can play too."

Milestone Four: Qwen3-MoE and Llama 4 ​

  • Qwen3-235B-A22B (2025.4): 235B total / 22B activated, activates 2 shared + 1 routing expert, supports "thoughtful/non-thoughtful" mixed reasoning mode, stands out on instruction following and programming benchmarks;
  • Llama 4 (2025.4): Meta's first shift to MoE — Scout (109B-A17B) and Maverick (400B-A17B), activating ~17B.

IV. Comparison Across Schemes ​

ModelTotal ParamsActivated ParamsExpert ConfigRoutingHighlights
Switch Transformer1.6T~tens of BMultiple experts per FFNTop-1First ultra-large MoE validation
GLaM1.2T~97B64 expertsTop-2Training cost only 1/3 of GPT-3
Mixtral 8×7B46.7B12.9B8 expertsTop-2Open-source MoE benchmark
Mixtral 8×22B141B39B8 expertsTop-2Larger open-source MoE
DeepSeek-V3671B37B256+1 expertsTop-8Fine-grained experts + no-aux-loss balance
Qwen3-235B-A22B235B22BRouting + 2 sharedTop-1+sharedMixed reasoning, bilingual
Grok-1314B86B8 expertsTop-2xAI's first open-source
GPT-4 (rumored)~1.8TNot disclosed8 experts (rumored)Not disclosedUnconfirmed by official

The parameters in the comparison table are "paper numbers" — the real difference is in the "capability/cost" curve: even with the same 671B total parameters, different implementations' routing efficiency, data quality, and training duration create huge gaps. The correct approach to evaluating MoE models: fix the "activated parameter budget" (e.g., 20B-class), compare different models horizontally on the same benchmarks and prompts, and measure actual throughput and latency. Parameter tables only say "how much material there is"; evaluation and benchmarking say "how much work gets done." Evaluation methods in Evaluation and Benchmarks.

V. MoE's Cost-Effectiveness: Activated Parameters Are the Bill ​

PerspectiveTotal Parameters (memory capacity)Activated Parameters (per-token compute)
Training costAffects VRAM, communicationDetermines FLOPs (≈ activated params × token count)
Inference latencyAffects model loading, VRAM residencyDetermines forward pass computation time
Capability ceilingMore = more knowledgeHigher = finer per-token processing

Conclusion: training cost is mainly about "activated params × data volume," the capability ceiling is mainly about "total params + data quality," and inference latency is affected by both. MoE's value is exchanging "put more params in, compute fewer params" for training cost-effectiveness; on the inference side (below), a new problem emerges: "saved compute, spent VRAM."

MoE's cost-effectiveness logic differs between training and inference, and many people confuse the two stages. On the training side, MoE's gain comes from "each token only activates some experts" — under the same FLOPs budget, it can accommodate a larger knowledge capacity, so training cost is calculated by activated params, an order of magnitude lower than equivalent dense models. On the inference side, compute is also determined by activated params, but VRAM and bandwidth are determined by total params — if you only have one GPU, 671B of MoE simply won't fit, no matter how "compute-saving" it is. This explains why MoE is especially suited for the "trainer with limited budget + server with ample resources" pattern (saves training, spreads VRAM across a large cluster) — and it's why teams like DeepSeek, "limited compute but strong engineering," can build flagship MoE models. Understanding that you need to balance the books separately on both sides of training/inference is the first step in evaluating any MoE scheme.

One more emphasis: "large total params" is the indicator for knowledge capacity, "small activated params" is the indicator for cost advantage — when selecting, calculate both numbers along with your VRAM and concurrency, and the MoE equation becomes clear.

For "is MoE worth it," the practice community has a consensus answer: on the training side, MoE almost always wins — same compute yields larger knowledge capacity; on the inference side, the answer depends on the scenario — high concurrency, large model, long-running services: MoE's amortized cost is lower; low concurrency, single-GPU, latency-sensitive deployment: dense is simpler and more reliable. Another underappreciated point is ecosystem maturity: post-2024, native MoE support from frameworks like vLLM has significantly reduced deployment complexity — the old impression that "MoE is hard to deploy" is changing. See Framework and Tool Selection.

VI. MoE Deployment Practice and Challenges ​

MoE deployment isn't just about "saving money" — it comes with a set of unique engineering problems (see Deployment and Servicing for details):

ChallengeCauseCommon Solutions
VRAM explosionAll expert weights must reside in VRAM (regardless of whether activated)Multi-layer quantization, expert offloading to CPU/SSD
Communication overheadExperts distributed across GPUs, tokens need cross-GPU forwarding (All-to-All)Expert parallelism (EP), topology-aware scheduling
Load imbalanceRouting hotspots cause some GPUs to queueLoad-balanced training + dynamic batching
Batching utilizationSingle request activates few experts, GPU compute wastedContinuous batching to stack throughput
KV cacheAttention cache still stored per total layer count for long contextMLA (DeepSeek's solution)

1. Expert Parallelism (EP) ​

Split experts across different GPUs, forward tokens to the corresponding card per routing result for computation, then aggregate — this is the standard parallel strategy for MoE inference. EP's throughput ceiling is constrained by the bottleneck of the "hottest expert's card."

2. Inference Framework Support ​

vLLM, SGLang, TensorRT-LLM, llama.cpp, and other mainstream engines now natively support MoE weight loading and EP scheduling; local 8×7B-class MoE can run on consumer GPUs (total weights of 46.7B need ~30 GB VRAM, reducible with quantization).

One more point to separately emphasize about MoE and quantization: MoE models are more sensitive to quantization, because the router layer's numerical stability matters — router scoring determines which expert a token goes to, and quantization noise can change routing decisions, leading to "misrouted experts." Therefore, MoE quantization typically requires: keeping the router layer at high precision, per-expert calibration, and post-quantization verification that the routing distribution hasn't drifted. These details make MoE quantization deployment more demanding than dense models. Relevant tools in Deployment and Servicing.

3. When to Choose MoE ​

ScenarioAdvice
Need huge knowledge capacity, limited budgetMoE (saves training compute)
Single-GPU low-VRAM deployment, latency-sensitiveDense small model or quantized mid-range MoE
High-concurrency API serviceMoE + EP + continuous batching (best cost)
Research/control priorityDense models are simpler and easier to tune

MoE services need more monitoring than dense models: in addition to standard latency, throughput, and token consumption, also monitor "routing distribution" — if some experts have long-term low utilization, routing may be degrading or load imbalance worsening; also watch VRAM occupancy (total params resident) and KV cache peaks. Integrating these metrics into alerts lets you spot problems before "quality silently degrades." General monitoring methods in Deployment and Servicing.

VII. MoE Training Details and Engineering Essentials ​

1. Three Engineering Points of Router Training ​

Engineering PointProblemCommon Practice
Router initializationInitial uniform random → uneven expert capabilitySmall random + uniform initialization for router, avoid early collapse
Load balancingHotspot experts over-selectedAuxiliary load balance loss, capacity factor, dynamic bias
Expert division of laborExpert redundancy, knowledge overlapFine-grained experts, shared experts handle general capability (DeepSeek approach)

Load balancing mechanism details in MoE Sparse Expert Models.

2. Comparison with Dense Training ​

DimensionDense ModelMoE
Per-token computeAll parametersOnly activated portion (e.g., 5%–20%)
VRAM requirementSmall weights but heavy computeAll weights resident in VRAM, VRAM actually larger
CommunicationStandard tensor/data parallelismNeeds expert parallelism + All-to-All
ConvergenceStableNeeds load balance tuning, slightly unstable
Cost-effectivenessCapability linear with costHigher capability per same cost (determined by activated params)

3. Representative Model Training Config Comparison ​

ModelTotal/Activated ParamsTraining DataContextActivated Experts
Mixtral 8×7B46.7B / 12.9B~8T tokens32K2 / 8
Mixtral 8×22B141B / 39BMultilingual enhanced64K2 / 8
DeepSeek-V3671B / 37B14.8T tokens128K8 / 256
Qwen3-235B-A22B235B / 22BLarge-scale multilingual32K–131KRouting + 2 shared
Grok-1314B / 86BNot disclosed8K2 / 8

Numbers in the table are subject to each official release; "context" refers to the main version capability at release.

4. MoE Forward Pseudocode ​

# Simplified MoE layer forward pass (per token)
def moe_forward(x, router, experts, k=2):
    logits = router(x)                 # [E]
    p = softmax(logits)                # normalize
    top_k_idx = argsort(p)[-k:]        # pick top k experts
    out = zeros_like(x)
    for i in top_k_idx:
        out += p[i] * experts[i](x)    # weighted sum
    return out

Three MoE training tips

① Secure load balance before discussing results; ② shared experts + fine-grained experts is the new mainstream; ③ evaluation must simultaneously look at "activated parameter cost-effectiveness" not just total parameters. Data and loss-related mechanisms in Pretraining: Data and Objectives and Scaling Laws.

VIII. MoE and the Future ​

  • Test-time scaling: MoE can stack with test-time scaling — large params + sparse activation + long chain-of-thought is the combo punch of 2025 flagship models;
  • Finer granularity: expert count continues rising (DeepSeek-V3 already has 256+1), routing precision and knowledge division getting finer;
  • Combining with multimodal: vision experts, language experts set by domain, making MoE the universal base for multimodal LLMs (see Multimodal LLMs);
  • MoE vs. Dense: short-term, MoE is mainstream for ultra-large scale; but dense models retain irreplaceable positions in low-VRAM, low-latency, and interpretability scenarios. Selection principles in Scaling Laws and Model Compendium.

One sentence to remember MoE

MoE = decoupling of total parameters (capacity) and activated parameters (cost). It turned "trillion parameters" from a gimmick into an engineering reality, and let open-source models compete head-to-head with closed-source flagships on cost for the first time.

IX. MoE FAQ and Selection Checklist ​

1. Common MoE Misconceptions ​

MisconceptionTruth
"Larger total params = smarter"Capability depends on activated param efficiency, routing quality, and data
"MoE always saves VRAM"Quite the opposite — all weights reside in VRAM
"MoE free capability upgrade"Training needs load balance tuning, higher engineering complexity
"MoE inference necessarily faster"Single request may be slower; throughput comes from batching/parallelism
"More experts is better"Too many experts worsen load imbalance and communication

2. When Not to Use MoE ​

ScenarioReason
Single-GPU / low-VRAM deploymentTotal weights too large to fit
Ultra-low-latency single requestSparse routing not as efficient as dense direct
Team without distributed experienceExpert parallelism debugging threshold is high
Need minimal interpretabilityDense models are easier to analyze
Prototype/validation stageGet it running first, then optimize cost

3. How to Evaluate an MoE Model ​

MetricAskLook At
Activated parameter cost-effectivenessIs capability higher at same cost?Activated params × token count vs. benchmark score
Routing qualityDo each expert specialize in their role?Routing distribution stats, expert utilization
Batching throughputHow many token/s at high concurrency?Benchmark reports, vLLM live tests
Long-context performanceDoes it drop at 128K?Needle-in-a-haystack, LongBench-type evals
Quantization robustnessHow many points dropped at 4-bit?Pre- and post-quantization comparison

4. Migration Checklist: Dense to MoE ​

  • [ ] Confirm the real bottleneck is training compute or inference cost
  • [ ] Compare MoE vs. dense throughput at same VRAM
  • [ ] Verify inference framework (vLLM/SGLang) support for this MoE
  • [ ] Evaluate quantization scheme and precision loss
  • [ ] Build cost model: include activated params, KV cache, communication overhead
  • [ ] Is the long-term maintenance and upgrade path clear?

5. FAQ Quick Answers ​

QuestionQuick Answer
Is Mixtral 8×7B 56B?No, total 46.7B, activated 12.9B
How many cards for DeepSeek-V3 training?Officially disclosed ~2000 H800s (subject to official)
Can a single GPU run MoE?Small/medium MoE quantized can, but throughput limited
What are shared experts?All-token-passing general experts handling grammar/commonsense
What is load balancing loss?Auxiliary loss penalizing routing bias, ensuring all experts are utilized
Is GPT-4 an MoE?Widely speculated in the community, OpenAI didn't confirm

One sentence to remember MoE selection

MoE buys "training cost-effectiveness" and sells "VRAM and engineering complexity." Calculate the "activated params × data volume" equation before jumping in.

One last note: the "cost-effectiveness" label for MoE holds only at sufficient scale — in small scenarios, dense models are often more trouble-free.

X. MoE Deep Dive: From Papers to Engineering ​

1. Key Papers at a Glance ​

Paper / WorkYearCore Contribution
Sparsely-Gated MoE Layer2017MoE layer founding, first applied to LM
GShard2020MoE + Transformer, Top-2 routing
Switch Transformer2021Top-1 routing, 1.6T params
GLaM20211.2T total / 97B activated, training cost ~1/3
ST-MoE2022Systematic study of routing load balance and expert diversity
Mixtral of Experts2024Open-source MoE landing benchmark
DeepSeek-V2 / V32024MLA + fine-grained experts + no-aux-loss balance

2. How Gradients Flow Through the Router ​

A core MoE training detail: "routing is discontinuous" — expert selection is discrete (Top-k), and gradients can't be directly computed. Mainstream handling:

  • Router soft weights backpropagate directly: gradient updates on the selected experts' weights p[i];
  • Expert parameters update independently: only selected experts receive gradients, others stay unchanged;
  • Load balancing loss gradients: penalize routing distribution deviating from uniform, driving expert division of labor.

This creates one phenomenon: when expert utilization is low, gradients are sparse and training slows — so load balance is both a stability problem and a training efficiency problem.

A frequently asked question: why does "let each token pick several experts" learn different capabilities? Intuitively: the router and experts are jointly trained, and experts only receive gradients when selected, so they're only "responsible" for "the types of tokens that select them" — the router learns to route math-type tokens to the math expert, the math expert gets stronger in math accordingly, creating a positive feedback loop. It's a bit like social division of labor: the clearer the division, the more specialized each point, but the more dependent on the router's "dispatch." If the router fails (load imbalance), the division collapses — which is precisely why load balancing matters so much. Understanding this "division + dispatch" metaphor lets you predict MoE's strengths (large capacity with diverse domains) and weaknesses (routing errors and batch utilization).

3. Experimental Observations on Expert Diversity ​

Engineering and research communities have observed several common phenomena:

  1. Experts aren't fully specialized — there's significant "general expert" usage (almost all tokens use them) — which is the design motivation for shared experts;
  2. Higher layers → more expert specialization (lower layers handle lexical, higher layers handle semantic);
  3. Routing quality improves with training but may overfit to the training distribution; monitor routing distribution drift at inference time.

4. Key Empirical Comparison Points: MoE vs. Dense ​

Comparison ExperimentKey Finding
Same FLOPsMoE usually beats dense (sparse activation gain)
Same total paramsDense is more stable; MoE's cost-effectiveness is in "capability/cost"
Small batchMoE advantage shrinks (low batching utilization)
Long contextMoE's KV cache matches dense; advantage still in activated params
Post-quantizationMoE more sensitive to quantization, needs targeted calibration

5. Two Counterintuitive Facts ​

Two of the most counterintuitive points about MoE. First, MoE's "total params" barely affect single-token inference latency (unless VRAM bandwidth is constrained); what really determines latency is activated params and KV cache — so "671B total params" is primarily about knowledge capacity, not speed. Second, MoE's advantage is not obvious in "low-concurrency" scenarios because sparse routing means each single request only uses a few experts, and GPUs can't be fully fed — while high-concurrency batching can stagger expert demands across requests, maximizing utilization. Therefore, evaluating MoE always requires benchmarking under "your actual concurrency profile," not looking at single-request benchmarks. These two points directly determine deployment selection (see Deployment and Servicing).

6. Quick Memory Formula Judgment ​

Quickly estimate whether an MoE model can be deployed: VRAM ≈ total params × bytes per param + KV cache + activations. Because MoE's total params far exceed activated params, the VRAM bottleneck is almost always "total params." Example: Mixtral 8×7B total 46.7B, FP16 needs ~93 GB, quantized to 4-bit ~25 GB — a single 24 GB card can barely run it with quantization and offloading; DeepSeek-V3 total 671B, FP16 ~1.3 TB, only deployable on clusters, quantized still needs 300 GB+. Calculate the total params bill first, then look at activated params efficiency — calculating these two bills separately is MoE deployment's first lesson.

XI. MoE Resource Checklist ​

ResourceUse
DeepSeek-V3 / R1 technical reportComplete engineering details of MoE + MLA + RL
Mixtral paper and Mistral blogOpen-source MoE implementation reference
vLLM / SGLang docsMoE deployment and expert parallelism
HF Open LLM LeaderboardOpen-source MoE leaderboard comparison
nanoMoE and community reproduction projectsHands-on MoE layer from scratch

Final reminder: the resource checklist is just the starting point. To truly understand MoE, you must run it yourself — deploy a Mixtral with vLLM, observe routing distribution and throughput, then compare cost-effectiveness with same-VRAM dense models. The gap between reading about it and hands-on verification is especially large with MoE.

The order to read MoE papers

First read Switch Transformer to build intuition → then read Mixtral to see landing → finally read DeepSeek-V3 for engineering limits. After reading these three, MoE's mechanisms, engineering, and costs all click.

A supplement to selection advice: if your team has no distributed training experience, the first time touching MoE should start with "directly using open-source MoE models" (Mixtral, Qwen3-235B-type), rather than training from scratch — first experience its deployment and cost characteristics, then decide whether to invest on the training side. MoE optimization on the training side (routing, load balancing, FP8) is deep water — without a billion-token experiment budget, it's hard to experience all its complexity.

XII. Further Reading ​

References ​