1. Why this is the most productive research area#
The systems papers in this field have had the highest hit rate for
production adoption:
Orca (2022) → continuous batching. UNIVERSAL.
PagedAttention (2023) → paged KV. UNIVERSAL.
Sarathi (2023-24) → chunked prefill. DEFAULT in major engines.
SGLang (2024) → RadixAttention. ADOPTED.
DistServe/Splitwise → disaggregation. EMERGING.
Compare to attention approximation (20+ papers, ~0 adopted) or
KV eviction (many papers, none standard).
WHY: systems techniques are usually EXACT — they don't change the
model's output, only how the computation is scheduled and stored.
No quality risk means fast adoption.This is the pattern from Section XIV.01: exact techniques with no downside adopt fast. Systems work is mostly exact.
2. What’s settled#
CONTINUOUS BATCHING (Orca, 2022)
iteration-level scheduling. 3-10x over static batching.
→ universal
PAGED KV (vLLM, 2023)
fixed-size blocks, block tables, sharing. 2.5-4x memory efficiency.
→ universal
CHUNKED PREFILL (Sarathi, 2023-24)
interleave prefill chunks with decode. Removes head-of-line blocking.
→ default in major engines
PREFIX CACHING / RADIX TREES (SGLang, 2023-24)
reuse KV for shared prefixes; tree structure for branching.
→ adopted
FUSED KERNELS AND CUDA GRAPHS
engineering rather than research, but the impact is comparable
→ universalThese five together are most of the 10-30x between naive PyTorch and a production engine.
3. What’s emerging#
DISAGGREGATED PREFILL/DECODE (DistServe, Splitwise, Mooncake)
→ Section XIII.06. Real gains at scale; substantial complexity.
→ the open question: does the benefit justify it below very
large scale?
KV-CENTRIC ARCHITECTURES (Mooncake)
→ treat the KV cache as a distributed store, not per-replica state
→ unifies prefix caching, disaggregation, and tiering
→ the conceptual framing may outlast the specific system
SLO-AWARE SCHEDULING
→ schedulers that reason about deadlines and per-request SLOs
rather than FCFS or simple priority
→ Section XIII.08. Practically valuable; underexplored.
LENGTH PREDICTION FOR SCHEDULING
→ if you could predict output length, SJF becomes possible
→ approaches: a small classifier, the model's own signal, historical
per-tenant distributions
→ accuracy is the problem; a noisy predictor makes scheduling worse
MULTI-MODEL AND MULTI-LORA SERVING
→ Punica, S-LoRA established the kernels
→ placement and routing at scale is less studied4. What’s open and would matter#
1. SCHEDULING UNDER UNCERTAINTY
The core difficulty: you don't know how long a request will run.
Everything downstream (admission, preemption, placement, SLO
guarantees) is harder because of it.
→ a good length predictor would improve: admission decisions,
queue ordering, memory reservation, and SLO attainment
→ nobody has one that's accurate enough
2. GOODPUT-OPTIMAL SCHEDULING
Most schedulers optimize throughput or latency. Goodput (requests
meeting their SLO) is what matters (Section XI.01) and is a
different objective.
→ DistServe frames it this way; more work would be valuable
3. MULTI-TENANT FAIRNESS WITH HETEROGENEOUS REQUESTS
Weighted fair queueing assumes comparable work units. LLM requests
vary 1000x.
→ Section XII.05's approach (weighted tokens) is approximate
→ a principled treatment would help
4. CROSS-REQUEST OPTIMIZATION
Requests share prefixes, and sometimes share computation beyond
prefixes. Could a scheduler exploit more sharing?
→ SGLang's RadixAttention is the current answer; there may be more
5. SCHEDULING FOR REASONING WORKLOADS
Very long generations change the problem: sequences occupy slots
for minutes, and their length is even less predictable.
→ Section XIV.05's shift makes this urgent5. The evaluation problem in systems research#
SYSTEMS PAPERS ARE HARDER TO EVALUATE THAN ALGORITHM PAPERS
the baseline must be a TUNED production system (Section XIV.01 Q1)
the workload must be realistic (Section X.07)
the metric must be goodput, not throughput
the hardware and configuration must be stated
COMMON PROBLEMS
✗ comparing against an untuned vLLM
✗ synthetic workloads with fixed lengths
✗ closed-loop load generation (hides queueing)
✗ reporting throughput without the latency at which it was achieved
✗ single hardware, single model
WHAT GOOD LOOKS LIKE
✓ multiple tuned baselines
✓ real or realistically-modeled traces
✓ open-loop arrival with a rate sweep
✓ goodput at stated SLOs
✓ ablations showing which component contributes whatOrca and PagedAttention are both good examples — strong baselines, clear ablations, and honest about what the technique does and doesn’t do.
6. The engineering/research boundary#
MUCH OF WHAT MATTERS IS ENGINEERING, NOT RESEARCH
vLLM's V1 rearchitecture (scheduler efficiency, async handling)
→ not a paper, but a substantial performance improvement
TensorRT-LLM's kernel selection
→ not novel algorithmically; enormously effective
Custom all-reduce for small messages (Section IX.07)
→ known technique, careful implementation
CUDA graph integration with continuous batching
→ an engineering problem with real payoff
→ if you're looking for impact, the engineering side of this field
has more available than the research side.This is worth saying plainly: an engineer who profiles a production system and fixes what they find will typically deliver more value than one who implements the latest paper. The curriculum’s Sections VII.01 and X.01 are the framework for that.
7. What to watch#
1. SCHEDULING FOR REASONING WORKLOADS
→ the workload shift (Section XIV.05) is the biggest change
to the systems problem in years
2. DISAGGREGATION AT SMALLER SCALE
→ if the complexity can be reduced, the benefit is real
3. LENGTH PREDICTION
→ would unlock a class of scheduling improvements
→ watch for accuracy claims that hold on real traffic
4. STANDARDIZED KV INTERCHANGE
→ would enable cross-engine and cross-tier KV movement
→ an ecosystem question more than a research one
5. GOODPUT-OPTIMAL FORMULATIONS
→ schedulers that directly optimize what you care about8. What is unlikely to matter#
✗ Scheduling papers with weak baselines
A 10x speedup over HuggingFace generate() is a 1.2x over vLLM.
✗ Techniques requiring a custom engine
If it can't be integrated with continuous batching and paged
attention, it won't be adopted.
✗ Optimizations for regimes nobody serves
(batch 1 only, fixed lengths, single-tenant)
✗ Yet another priority scheduling variant
Priority + aging + fair queueing covers the practical space.9. Hands-on exercise#
A. Ablate a production engine. Disable, one at a time: continuous batching, paged KV, chunked prefill, prefix caching, CUDA graphs. Measure throughput and goodput after each. Which contributes most? Does the ranking match Section XIV.08 section 2?
B. Build a length predictor. From request logs, build a predictor for output length using prompt features and tenant history. Measure its accuracy. Is it good enough for SJF scheduling?
C. Goodput vs throughput. Instrument both for a system under increasing load. Plot them together. Where do they diverge, and what does that tell you?
D. Evaluate a systems paper. Take a recent scheduling paper. Apply Section XIV.01’s questions, focusing on Q1 (baseline) and Q2 (configuration). Does its claim survive?
E. Reproduce a baseline. Implement a paper’s baseline properly (tuned, with continuous batching). How much of the reported improvement remains?
F. Reasoning workload. Simulate a scheduler under a reasoning workload (very long, high-variance generations). What breaks? What would you change?
10. Interview questions#
- Why have systems papers had a higher production hit rate than algorithm papers in this field?
- What are the five settled systems techniques, and roughly what does each contribute?
- What is the core difficulty in LLM scheduling?
- Why is goodput a better objective than throughput for a scheduler?
- What’s commonly wrong with systems paper evaluations?
- How does the reasoning-workload shift change the scheduling problem?
- Where would you look for impact: research or engineering? Justify.
11. Further reading#
- [ESTABLISHED] Yu et al., “Orca” (OSDI 2022) — read it as a model of how to do this well
- [ESTABLISHED] Kwon et al., “PagedAttention” (SOSP 2023)
- [ESTABLISHED] Agrawal et al., “Sarathi-Serve” (OSDI 2024)
- [ESTABLISHED] Zheng et al., “SGLang” (2024)
- [EMERGING] Zhong et al., “DistServe” (2024); Qin et al., “Mooncake” (2024)
- [REFERENCE] vLLM and SGLang design documents and release notes
- Next: 09 — NVIDIA architecture trends