PidokuInfra

Scheduling and Systems Research

Expert Advanced 1h Difficulty 4/5 Topic 08 of 11

Prerequisites V.09, VIII.03, XIII.06, XIII.08


1. Why this is the most productive research area#

The systems papers in this field have had the highest hit rate for
production adoption:

  Orca (2022)            → continuous batching. UNIVERSAL.
  PagedAttention (2023)  → paged KV. UNIVERSAL.
  Sarathi (2023-24)      → chunked prefill. DEFAULT in major engines.
  SGLang (2024)          → RadixAttention. ADOPTED.
  DistServe/Splitwise    → disaggregation. EMERGING.

Compare to attention approximation (20+ papers, ~0 adopted) or
KV eviction (many papers, none standard).

WHY: systems techniques are usually EXACT — they don't change the
     model's output, only how the computation is scheduled and stored.
     No quality risk means fast adoption.

This is the pattern from Section XIV.01: exact techniques with no downside adopt fast. Systems work is mostly exact.


2. What’s settled#

CONTINUOUS BATCHING (Orca, 2022)
  iteration-level scheduling. 3-10x over static batching.
  → universal

PAGED KV (vLLM, 2023)
  fixed-size blocks, block tables, sharing. 2.5-4x memory efficiency.
  → universal

CHUNKED PREFILL (Sarathi, 2023-24)
  interleave prefill chunks with decode. Removes head-of-line blocking.
  → default in major engines

PREFIX CACHING / RADIX TREES (SGLang, 2023-24)
  reuse KV for shared prefixes; tree structure for branching.
  → adopted

FUSED KERNELS AND CUDA GRAPHS
  engineering rather than research, but the impact is comparable
  → universal

These five together are most of the 10-30x between naive PyTorch and a production engine.


3. What’s emerging#

DISAGGREGATED PREFILL/DECODE (DistServe, Splitwise, Mooncake)
  → Section XIII.06. Real gains at scale; substantial complexity.
  → the open question: does the benefit justify it below very
    large scale?

KV-CENTRIC ARCHITECTURES (Mooncake)
  → treat the KV cache as a distributed store, not per-replica state
  → unifies prefix caching, disaggregation, and tiering
  → the conceptual framing may outlast the specific system

SLO-AWARE SCHEDULING
  → schedulers that reason about deadlines and per-request SLOs
    rather than FCFS or simple priority
  → Section XIII.08. Practically valuable; underexplored.

LENGTH PREDICTION FOR SCHEDULING
  → if you could predict output length, SJF becomes possible
  → approaches: a small classifier, the model's own signal, historical
    per-tenant distributions
  → accuracy is the problem; a noisy predictor makes scheduling worse

MULTI-MODEL AND MULTI-LORA SERVING
  → Punica, S-LoRA established the kernels
  → placement and routing at scale is less studied

4. What’s open and would matter#

1. SCHEDULING UNDER UNCERTAINTY
   The core difficulty: you don't know how long a request will run.
   Everything downstream (admission, preemption, placement, SLO
   guarantees) is harder because of it.
   
   → a good length predictor would improve: admission decisions,
     queue ordering, memory reservation, and SLO attainment
   → nobody has one that's accurate enough

2. GOODPUT-OPTIMAL SCHEDULING
   Most schedulers optimize throughput or latency. Goodput (requests
   meeting their SLO) is what matters (Section XI.01) and is a
   different objective.
   → DistServe frames it this way; more work would be valuable

3. MULTI-TENANT FAIRNESS WITH HETEROGENEOUS REQUESTS
   Weighted fair queueing assumes comparable work units. LLM requests
   vary 1000x.
   → Section XII.05's approach (weighted tokens) is approximate
   → a principled treatment would help

4. CROSS-REQUEST OPTIMIZATION
   Requests share prefixes, and sometimes share computation beyond
   prefixes. Could a scheduler exploit more sharing?
   → SGLang's RadixAttention is the current answer; there may be more

5. SCHEDULING FOR REASONING WORKLOADS
   Very long generations change the problem: sequences occupy slots
   for minutes, and their length is even less predictable.
   → Section XIV.05's shift makes this urgent

5. The evaluation problem in systems research#

SYSTEMS PAPERS ARE HARDER TO EVALUATE THAN ALGORITHM PAPERS

  the baseline must be a TUNED production system (Section XIV.01 Q1)
  the workload must be realistic (Section X.07)
  the metric must be goodput, not throughput
  the hardware and configuration must be stated

COMMON PROBLEMS
  ✗ comparing against an untuned vLLM
  ✗ synthetic workloads with fixed lengths
  ✗ closed-loop load generation (hides queueing)
  ✗ reporting throughput without the latency at which it was achieved
  ✗ single hardware, single model

WHAT GOOD LOOKS LIKE
  ✓ multiple tuned baselines
  ✓ real or realistically-modeled traces
  ✓ open-loop arrival with a rate sweep
  ✓ goodput at stated SLOs
  ✓ ablations showing which component contributes what

Orca and PagedAttention are both good examples — strong baselines, clear ablations, and honest about what the technique does and doesn’t do.


6. The engineering/research boundary#

MUCH OF WHAT MATTERS IS ENGINEERING, NOT RESEARCH

  vLLM's V1 rearchitecture (scheduler efficiency, async handling)
  → not a paper, but a substantial performance improvement

  TensorRT-LLM's kernel selection
  → not novel algorithmically; enormously effective

  Custom all-reduce for small messages (Section IX.07)
  → known technique, careful implementation

  CUDA graph integration with continuous batching
  → an engineering problem with real payoff

→ if you're looking for impact, the engineering side of this field
  has more available than the research side.

This is worth saying plainly: an engineer who profiles a production system and fixes what they find will typically deliver more value than one who implements the latest paper. The curriculum’s Sections VII.01 and X.01 are the framework for that.


7. What to watch#

1. SCHEDULING FOR REASONING WORKLOADS
   → the workload shift (Section XIV.05) is the biggest change
     to the systems problem in years

2. DISAGGREGATION AT SMALLER SCALE
   → if the complexity can be reduced, the benefit is real

3. LENGTH PREDICTION
   → would unlock a class of scheduling improvements
   → watch for accuracy claims that hold on real traffic

4. STANDARDIZED KV INTERCHANGE
   → would enable cross-engine and cross-tier KV movement
   → an ecosystem question more than a research one

5. GOODPUT-OPTIMAL FORMULATIONS
   → schedulers that directly optimize what you care about

8. What is unlikely to matter#

✗ Scheduling papers with weak baselines
    A 10x speedup over HuggingFace generate() is a 1.2x over vLLM.

✗ Techniques requiring a custom engine
    If it can't be integrated with continuous batching and paged
    attention, it won't be adopted.

✗ Optimizations for regimes nobody serves
    (batch 1 only, fixed lengths, single-tenant)

✗ Yet another priority scheduling variant
    Priority + aging + fair queueing covers the practical space.

9. Hands-on exercise#

A. Ablate a production engine. Disable, one at a time: continuous batching, paged KV, chunked prefill, prefix caching, CUDA graphs. Measure throughput and goodput after each. Which contributes most? Does the ranking match Section XIV.08 section 2?

B. Build a length predictor. From request logs, build a predictor for output length using prompt features and tenant history. Measure its accuracy. Is it good enough for SJF scheduling?

C. Goodput vs throughput. Instrument both for a system under increasing load. Plot them together. Where do they diverge, and what does that tell you?

D. Evaluate a systems paper. Take a recent scheduling paper. Apply Section XIV.01’s questions, focusing on Q1 (baseline) and Q2 (configuration). Does its claim survive?

E. Reproduce a baseline. Implement a paper’s baseline properly (tuned, with continuous batching). How much of the reported improvement remains?

F. Reasoning workload. Simulate a scheduler under a reasoning workload (very long, high-variance generations). What breaks? What would you change?


10. Interview questions#

  1. Why have systems papers had a higher production hit rate than algorithm papers in this field?
  2. What are the five settled systems techniques, and roughly what does each contribute?
  3. What is the core difficulty in LLM scheduling?
  4. Why is goodput a better objective than throughput for a scheduler?
  5. What’s commonly wrong with systems paper evaluations?
  6. How does the reasoning-workload shift change the scheduling problem?
  7. Where would you look for impact: research or engineering? Justify.

11. Further reading#

  • [ESTABLISHED] Yu et al., “Orca” (OSDI 2022) — read it as a model of how to do this well
  • [ESTABLISHED] Kwon et al., “PagedAttention” (SOSP 2023)
  • [ESTABLISHED] Agrawal et al., “Sarathi-Serve” (OSDI 2024)
  • [ESTABLISHED] Zheng et al., “SGLang” (2024)
  • [EMERGING] Zhong et al., “DistServe” (2024); Qin et al., “Mooncake” (2024)
  • [REFERENCE] vLLM and SGLang design documents and release notes
  • Next: 09 — NVIDIA architecture trends

↑↓ navigate↵ openesc close