Below the API

MoE Research

Expert Advanced 45 min Difficulty 4/5

Prerequisites XIII.02, XIII.03, IX.05


1. Where the field is#

SOLVED
  ✓ top-k routing works; MoE trains stably at scale
  ✓ grouped GEMM is the right kernel structure
  ✓ expert parallelism with All-to-All
  ✓ auxiliary losses (and now loss-free methods) for load balance

ACTIVE
  ~ fine-grained experts (many small vs few large)
  ~ shared experts alongside routed ones
  ~ inference-time load balancing and expert placement
  ~ expert offloading for memory-constrained serving
  ~ reducing All-to-All cost

OPEN
  ? optimal E and k for a given compute budget
  ? whether expert specialization is interpretable/exploitable
  ? MoE at small scale (does it help below ~10B?)

2. The architectural findings that matter for serving#

FINE-GRAINED EXPERTS (DeepSeek)
  256 small experts with k=8, rather than 8 large with k=2
  → same active parameters, more combinations, better specialization
  → SERVING IMPACT: more experts means more All-to-All destinations
    and finer load-balance granularity. Section XIII.02's
    "experts touched" formula changes shape.

SHARED EXPERT (DeepSeek)
  one expert always active, plus k routed
  → captures common computation; routed experts specialize
  → SERVING IMPACT: the shared expert's weights are read every token,
    which is a fixed cost. But it makes routing less critical, so
    imbalance hurts less.

AUXILIARY-LOSS-FREE BALANCING (DeepSeek-V3)
  per-expert bias terms adjusted during training to equalize load,
  instead of an auxiliary loss that fights the main objective
  → SERVING IMPACT: better-balanced routing at inference, which
    directly improves throughput (Section IX.05)

NODE-LIMITED ROUTING (DeepSeek)
  a token's experts are constrained to at most M nodes
  → SERVING IMPACT: bounds cross-node All-to-All by construction
  → this is a TRAINING-TIME solution to a SERVING problem, and it's
    the kind of co-design that matters

Node-limited routing is the most instructive item here. It’s an architecture constraint introduced specifically to make serving efficient — an example of inference requirements influencing model design (Section XIV.09’s theme).


3. The serving-side research#

EXPERT OFFLOADING  [EMERGING]
  keep cold experts in CPU memory; fetch on demand.
  → the same arithmetic as KV offload (Section XIII.07): is the
    fetch faster than... what? You can't recompute an expert.
  → so the question is whether the fetch fits in the step time.
  
  expert size: d × d_ff × 2 matrices × bytes
    DeepSeek-scale small expert: ~50 MB at FP8
    fetch over PCIe Gen5 (55 GB/s): 0.9 ms
    decode step: ~10 ms
  → fetching a few cold experts per step is feasible IF you can
    predict which ones, and prefetch.
  → prediction is the hard part; routing is data-dependent.

  STATUS: works for low-batch, latency-tolerant serving.
          Not viable at high batch (you touch all experts anyway).

EXPERT CACHING / PLACEMENT  [EMERGING]
  measure expert popularity; keep hot experts resident, replicate them
  → straightforward and effective (Section XIII.03)

ALL-TO-ALL OPTIMIZATION  [ACTIVE]
  overlapping with compute, better algorithms, hardware support
  → 35% of MoE layer time is the target (Section XIII.03)

BATCHED EXPERT EXECUTION  [SOLVED]
  grouped GEMM, sorting. Done.

4. The open question that matters most#

DOES MoE HELP AT INFERENCE, OR ONLY AT TRAINING?

THE CASE FOR
  ✓ quality per FLOP is much better
  ✓ at very high batch, the FLOP advantage is real throughput
  ✓ enables model sizes that dense training couldn't reach

THE CASE AGAINST
  ✗ memory capacity requirement is the full parameter count
  ✗ the memory-traffic advantage disappears by batch ~32
    (Section XIII.02)
  ✗ All-to-All overhead is 30-40% of MoE layer time
  ✗ load imbalance costs throughput
  ✗ operationally much more complex

THE HONEST ANSWER
  MoE is a TRAINING-EFFICIENCY architecture that is SERVEABLE,
  not a serving-efficiency architecture.
  
  It's the right choice when:
    - you want quality that dense training can't afford
    - you have the memory
    - you serve at high batch
  
  It's the wrong choice when:
    - you're memory-constrained
    - you serve at low batch with tight latency

Being clear about this matters because “37B active parameters” is often quoted as if it were a serving-cost statement, and it isn’t (Section XIII.02).


5. What to watch#

1. WHETHER FRONTIER MODELS STAY MoE
   → if yes, serving infrastructure must handle it well
   → if the field returns to dense at the top, MoE serving
     becomes a niche skill

2. HARDWARE SUPPORT FOR All-to-All
   → the dominant MoE serving cost
   → watch: interconnect and collective-offload capabilities

3. EXPERT PREDICTION FOR PREFETCHING
   → would make expert offloading viable
   → watch: whether routing is predictable enough

4. MoE + LONG CONTEXT INTERACTION
   → both are memory-hungry; do they compose or conflict?
   → underexplored

5. SMALLER-SCALE MoE
   → does MoE help at 3-10B? If so, it changes edge and
     cost-sensitive serving

6. What is unlikely to matter#

✗ New routing functions (top-k variants)
    The space is well-explored; the gains are small.

✗ MoE papers evaluated only on training efficiency
    Training efficiency was never in doubt.

✗ Expert offloading evaluated at batch 1 only
    Section XIII.02: at batch 32+ you touch all experts, so
    offloading has nothing to skip.

✗ Interpretability of expert specialization
    Interesting, but no established serving application.

7. Hands-on exercise#

A. Verify the experts-touched formula. For an MoE model you can run, instrument the router and measure experts touched at batch 1, 4, 16, 64, 256. Compare to E × (1 - (1-k/E)^B). Does it hold?

B. The offloading question. For an MoE model, compute: expert size, fetch time over your interconnect, and decode step time. At what batch size does offloading stop being able to skip anything?

C. Load imbalance over time. Measure the expert load distribution over a day of realistic traffic. Is it stable, or does it shift? Would static expert replication work, or do you need dynamic rebalancing?

D. All-to-All overlap. Profile an MoE model. What fraction of layer time is All-to-All? Is it overlapped with compute? What would perfect overlap save?

E. Compare to dense. For an MoE model and a dense model of similar quality, compute cost per million tokens at batch 8 and batch 256 (Section XI.03). Where does each win?


8. Interview questions#

  1. What’s the serving cost of an MoE model — active or total parameters?
  2. Why does MoE’s memory-traffic advantage disappear with batch size?
  3. What is node-limited routing and what problem does it solve?
  4. What is a shared expert and how does it help serving?
  5. When is expert offloading viable?
  6. Is MoE a serving-efficiency architecture? Justify your answer.
  7. What’s the dominant cost in MoE serving, and what would fix it?

9. Further reading#

  • [ESTABLISHED] Shazeer et al. (2017); Lepikhin et al., “GShard”; Fedus et al., “Switch”
  • [ESTABLISHED] Jiang et al., “Mixtral of Experts” (2024)
  • [ESTABLISHED] DeepSeek-V2 and V3 technical reports — the most detailed public account of MoE inference engineering, including node-limited routing and loss-free balancing
  • [ESTABLISHED] Gale et al., “MegaBlocks” (2023) — block-sparse MoE kernels
  • [EMERGING] Expert offloading work (Mixtral-offloading and similar)
  • Next: 07 — Long-context research

↑↓ navigate ↵ open