PidokuInfra

Decoding Research

Expert Advanced 1h Difficulty 4/5 Topic 05 of 11

Prerequisites V.12, V.13, XIII.09


1. The two problems#

PROBLEM A — SPEED
  decode is serial and memory-bound (Section V.03).
  → speculative decoding and its variants (Section XIII.09)
  → multi-token prediction
  → parallel decoding schemes

PROBLEM B — QUALITY
  which token to pick, and how sampling affects output quality.
  → sampling methods (Section V.13)
  → constrained/structured decoding
  → inference-time scaling (thinking longer)

These are usually studied separately but interact: inference-time scaling makes decode speed matter much more.


2. Speed: the state of speculative decoding#

Covered in Section XIII.09. The research summary:

SETTLED
  ✓ the algorithm and its correctness proof (Leviathan, Chen 2023)
  ✓ that it trades compute for bandwidth
  ✓ that it degrades with batch size
  ✓ n-gram/lookup as a free variant

ACTIVE
  ~ better draft models (distillation, architecture)
  ~ tree construction (EAGLE-2's dynamic trees)
  ~ feature-level vs token-level drafting
  ~ making it work at high batch

DIMINISHING
  - marginally better draft heads
  - variants requiring training, without addressing the batch problem

The batch problem is the open question. Every variant degrades toward 1.0x speedup as batch grows, because verification compute stops being free. A method that works at batch 128 would be genuinely new.

WHY THE BATCH PROBLEM IS HARD
  speculation trades COMPUTE for BANDWIDTH.
  at batch 1: compute is 99% idle → the trade is free
  at batch 128: compute is 40% used → the trade costs real time
  
  → the only escape is a draft that costs approximately nothing
    (n-gram) or a target verification that's cheaper than k separate
    forward passes by more than the draft costs

3. Multi-token prediction — the cleaner path#

INSTEAD OF: bolting speculation onto a model trained for one token
DO:         train the model to predict several tokens

DeepSeek-V3 includes MTP modules.
Gloeckle et al. (2024) show multi-token training also improves
  the model's quality, not just its speed.

WHY THIS IS BETTER
  ✓ no separate draft model or trained heads
  ✓ the predictions are aligned by construction
  ✓ the training signal improves the base model
  ✗ requires training with it; can't retrofit

STATUS: shipping in some models. Likely to become standard.

This is where the field should go, and the fact that multi-token training also improves quality (not just speed) makes it likely.


4. Quality: sampling#

SETTLED
  ✓ pure greedy is bad for open-ended generation (Holtzman 2019)
  ✓ nucleus (top-p) sampling as the practical default
  ✓ temperature as the diversity control

REFINEMENTS THAT MATTER
  min-p sampling: threshold relative to the max probability
    → adapts to the model's confidence, unlike fixed top-p
    → genuinely better across confidence levels (Section V.13)
  
  typical sampling, eta/epsilon sampling, and others
    → marginal; not widely adopted

ACTIVE
  ~ sampling for reasoning tasks (does temperature help or hurt
    chain-of-thought?)
  ~ adaptive sampling based on position or content

min-p is the one refinement worth knowing. It’s simple, better-motivated than top-p, and increasingly supported.


5. Constrained and structured decoding#

THE PROBLEM: force output to match a grammar (JSON schema, regex,
             a programming language).

THE NAIVE APPROACH: mask disallowed tokens at each step.
  → 0.5-5 ms per token to compute the mask. Can dominate ITL.

THE ADVANCES
  compressed FSM (SGLang)     precompute masks per automaton state
  jump-forward decoding       when only one continuation is valid,
                              emit it WITHOUT running the model
  XGrammar, llguidance        fast mask computation, often as bitmasks
  → 0.02 ms per token, and fewer model steps

WHY JUMP-FORWARD MATTERS
  in JSON, much of the output is literal structure:
    {"name": "..."}
     ^^^^^^^^^^     the model doesn't need to generate this
  → skipping it saves time AND removes the chance of getting it wrong
STATUS: ESTABLISHED. This is a solved engineering problem with
        good implementations available.

OPEN
  ~ constrained decoding that doesn't distort the distribution
    (masking changes the probabilities — is the result still the
     model's "intent"?)
  ~ efficient handling of very large grammars
  ~ constraints that depend on generated content (semantic, not syntactic)

The distribution-distortion question is real and under-discussed: forcing a token the model assigned low probability to is a form of intervention, and the downstream generation is conditioned on a token the model wouldn’t have chosen.


6. Inference-time scaling — the big shift#

THE IDEA: spend more compute at inference to get better answers.
  chain-of-thought, self-consistency (sample n, take the majority),
  tree search, verifier-guided search, extended "thinking" before
  answering

WHY IT MATTERS FOR INFERENCE ENGINEERING
  it changes the WORKLOAD SHAPE fundamentally:
  
    before: 500 input, 200 output
    with reasoning: 500 input, 5,000-50,000 output
    
  → decode becomes overwhelmingly dominant
  → per-token cost matters far more
  → speculative decoding becomes far more valuable
  → KV cache grows much larger per request
  → latency budgets change (users accept 30 s for a hard question)
CONSEQUENCES FOR THE STACK

  1. DECODE OPTIMIZATION becomes the dominant concern
     everything in Sections VII and XIII that speeds decode
     is worth more

  2. KV CACHE PRESSURE increases dramatically
     a 50,000-token reasoning trace has a 50,000-token KV cache
     → GQA/MLA, KV quantization, and long-context techniques
       become critical

  3. BATCHING CHANGES
     long generations mean sequences occupy slots for minutes
     → concurrency is limited by slot-time, not arrival rate
     → Little's Law with much larger W

  4. NEW OPTIMIZATION OPPORTUNITIES
     ~ can you cache/reuse reasoning traces?
     ~ can you speculate on reasoning steps?
     ~ can you prune search branches early?
     ~ parallel sampling shares the prompt's KV (Section V.10) —
       self-consistency is cheap in KV terms

This is the most consequential shift in the field for inference engineers. A workload that was 30% prefill and 70% decode becomes 3% prefill and 97% decode, and every optimization’s value changes accordingly.


7. What to watch, ranked#

1. INFERENCE-TIME SCALING ADOPTION
   → changes workload shape more than any technique changes efficiency
   → watch: product behavior, not papers

2. NATIVE MULTI-TOKEN PREDICTION IN MODELS
   → cleaner than bolted-on speculation
   → watch: model releases

3. SPECULATIVE DECODING AT HIGH BATCH
   → currently unsolved; would be significant
   → watch: methods that address the compute cost of verification

4. REASONING-TRACE REUSE / CACHING
   → if reasoning traces for similar problems can be reused,
     the cost of inference-time scaling drops substantially
   → watch: early work in this area

5. CONSTRAINED DECODING WITHOUT DISTRIBUTION DISTORTION
   → a correctness question that becomes important as structured
     output becomes universal

8. What is unlikely to matter#

✗ Another sampling method
    top-p and min-p cover the practical space.

✗ Speculative decoding variants that don't address the batch problem
    They're all 1.1-1.3x at production batch sizes.

✗ Beam search revivals for open-ended generation
    Holtzman's finding stands; and the KV cost is prohibitive.

✗ Constrained decoding methods slower than the existing fast ones

9. Hands-on exercise#

A. Measure the workload shift. For a reasoning model versus a standard model on the same questions, measure the output-length distribution. How does the prefill/decode time split change?

B. Recompute your optimization priorities. With the reasoning workload’s shape, recompute which optimizations matter (Section VII.01’s ranking). What moved?

C. KV pressure. For a 20,000-token reasoning trace, compute the KV cache size and the resulting concurrency. Compare to a 200-token response.

D. Self-consistency KV sharing. Generate n=8 samples from one prompt with and without KV sharing (Section V.10). Measure the memory difference.

E. Constrained decoding cost. Measure ITL with a naive masking implementation and with a fast one (XGrammar or SGLang). Quantify the difference.

F. Jump-forward. For a JSON-schema-constrained generation, count how many output tokens were literal structure that jump-forward decoding could skip.


10. Interview questions#

  1. Why does speculative decoding degrade with batch size, and what would fix it?
  2. What is multi-token prediction and why is it cleaner than bolted-on speculation?
  3. What is jump-forward decoding?
  4. How does inference-time scaling change the inference workload?
  5. Which optimizations become more valuable with reasoning models?
  6. What is min-p sampling and why is it better than top-p?
  7. What’s the correctness concern with constrained decoding?

11. Further reading#

  • [FUNDAMENTAL] Holtzman et al., “The Curious Case of Neural Text Degeneration” (2019)
  • [ESTABLISHED] Leviathan et al. (2022); Chen et al. (2023)
  • [EMERGING] Gloeckle et al., “Better & Faster Large Language Models via Multi-token Prediction” (2024)
  • [ESTABLISHED] Zheng et al., “SGLang” — compressed FSM and jump-forward decoding
  • [EMERGING] Work on inference-time compute scaling and its serving implications
  • Next: 06 — MoE research

↑↓ navigate↵ openesc close