Curated, organized by section. Each entry is labeled:
- [FUNDAMENTAL] — timeless; the physics/math doesn’t change
- [ESTABLISHED] — production-proven engineering
- [EMERGING] — promising, not universally production-ready
- [REFERENCE] — documentation you will return to
Prefer primary sources. When a blog post and a paper disagree, read the paper; when the paper and your profiler disagree, believe the profiler.
Cross-cutting / start here#
- [REFERENCE] NVIDIA CUDA C++ Programming Guide — https://docs.nvidia.com/cuda/cuda-c-programming-guide/
- [REFERENCE] NVIDIA CUDA C++ Best Practices Guide — https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/
- [REFERENCE] PyTorch docs — https://pytorch.org/docs/stable/index.html
- [REFERENCE] vLLM docs — https://docs.vllm.ai/
- [REFERENCE] SGLang docs — https://docs.sglang.ai/
- [REFERENCE] Hugging Face Transformers docs — https://huggingface.co/docs/transformers/
- [FUNDAMENTAL] Hennessy & Patterson, Computer Architecture: A Quantitative Approach — the appendix on memory hierarchy alone repays the price.
I — Fundamentals#
- [FUNDAMENTAL] Williams, Waterman, Patterson, “Roofline: An Insightful Visual Performance Model for Multicore Architectures” (CACM 2009) — the origin of arithmetic intensity as a design tool.
- [ESTABLISHED] Kaplan et al., “Scaling Laws for Neural Language Models” (2020) — read for the FLOPs-per-token accounting, not the scaling conclusions.
- [ESTABLISHED] “LLM Inference Performance Engineering: Best Practices” — Databricks engineering blog. Good first pass on TTFT/ITL thinking.
- [FUNDAMENTAL] Latency Numbers Every Programmer Should Know (Jeff Dean / Peter Norvig tables) — internalize the orders of magnitude.
II — Computer Systems Foundations#
- [FUNDAMENTAL] Ulrich Drepper, “What Every Programmer Should Know About Memory” (2007) — still the best treatment of caches and NUMA. Long; read parts 2, 3, 5.
- [FUNDAMENTAL] Brendan Gregg, Systems Performance (2nd ed.) — the USE method,
perf, flame graphs. - [REFERENCE] Brendan Gregg’s site — https://www.brendangregg.com/ (flame graphs, eBPF)
- [REFERENCE]
man 7 numa,numactl(8),perf-stat(1) - [REFERENCE] PCI-SIG specifications overview; for practical numbers, NVIDIA’s
bandwidthTestsample in cuda-samples. - [FUNDAMENTAL] Bovet & Cesati, Understanding the Linux Kernel — for scheduling and VM.
III — Machine Learning Fundamentals#
- [FUNDAMENTAL] Vaswani et al., “Attention Is All You Need” (2017) — the transformer paper.
- [FUNDAMENTAL] Jay Alammar, “The Illustrated Transformer” — the best visual introduction.
- [FUNDAMENTAL] Andrej Karpathy, “Let’s build GPT: from scratch, in code, spelled out”
(video) and
nanoGPT— https://github.com/karpathy/nanoGPT - [FUNDAMENTAL] 3Blue1Brown, Essence of Linear Algebra (video series).
- [ESTABLISHED] Micikevicius et al., “Mixed Precision Training” (2018) — where FP16 loss scaling and BF16 reasoning come from.
- [ESTABLISHED] Micikevicius et al., “FP8 Formats for Deep Learning” (2022) — E4M3/E5M2.
- [REFERENCE] IEEE 754 and the bfloat16 numerics note from Google Brain.
IV — Neural Network Inference#
- [REFERENCE] NVIDIA cuBLAS docs; the “Matrix Multiplication Background User’s Guide” in NVIDIA Deep Learning Performance documentation — explains tiling and tile quantization.
- [ESTABLISHED] “Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking” (Jia et al.) — how people actually determine what hardware does.
- [REFERENCE] ONNX operator specification — https://onnx.ai/onnx/operators/
- [ESTABLISHED] Chen et al., “TVM: An Automated End-to-End Optimizing Compiler for Deep Learning” (OSDI 2018) — graph + operator level optimization.
- [REFERENCE] PyTorch
torch.compile/ TorchInductor documentation.
V — LLM Inference (the core section)#
- [ESTABLISHED] Kwon et al., “Efficient Memory Management for Large Language Model Serving with PagedAttention” (SOSP 2023) — the vLLM paper. Read this one twice.
- [ESTABLISHED] Yu et al., “Orca: A Distributed Serving System for Transformer-Based Generative Models” (OSDI 2022) — the origin of iteration-level (continuous) batching.
- [ESTABLISHED] Pope et al., “Efficiently Scaling Transformer Inference” (2022) — the canonical analytical treatment of prefill/decode, partitioning, and the arithmetic.
- [ESTABLISHED] Leviathan et al., “Fast Inference from Transformers via Speculative Decoding” (2022); Chen et al., “Accelerating Large Language Model Decoding with Speculative Sampling” (2023).
- [ESTABLISHED] Shazeer, “Fast Transformer Decoding: One Write-Head is All You Need” (2019) — MQA, and the clearest statement of why KV bandwidth dominates decode.
- [ESTABLISHED] Ainslie et al., “GQA: Training Generalized Multi-Query Transformer Models” (2023).
- [REFERENCE] vLLM design docs on the scheduler and block manager.
- [ESTABLISHED] Holtzman et al., “The Curious Case of Neural Text Degeneration” (2019) — where top-p (nucleus) sampling comes from.
VI — GPU Computing#
- [REFERENCE] CUDA C++ Programming Guide (again — chapters 2, 5, and the appendix on compute capabilities).
- [FUNDAMENTAL] NVIDIA whitepapers per architecture: Volta, Ampere (A100), Hopper (H100), Blackwell. Each explains the SM, memory system, and tensor cores of that generation.
- [FUNDAMENTAL] “CUDA Refresher” blog series on NVIDIA Developer Blog.
- [ESTABLISHED] Mark Harris, “How to Optimize Data Transfers in CUDA C/C++” and the “CUDA Pro Tip” series.
- [REFERENCE] Nsight Systems and Nsight Compute user guides.
- [FUNDAMENTAL] Programming Massively Parallel Processors (Kirk & Hwu), 4th ed. — the standard textbook.
- [REFERENCE] CUTLASS — https://github.com/NVIDIA/cutlass — how production GEMMs are built.
VII — Inference Optimization#
- [ESTABLISHED] Dao et al., “FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness” (2022).
- [ESTABLISHED] Dao, “FlashAttention-2” (2023).
- [EMERGING→ESTABLISHED] Shah et al., “FlashAttention-3” (2024) — Hopper-specific (async, FP8).
- [ESTABLISHED] Frantar et al., “GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers” (2022).
- [ESTABLISHED] Lin et al., “AWQ: Activation-aware Weight Quantization” (2023).
- [ESTABLISHED] Xiao et al., “SmoothQuant” (2022).
- [ESTABLISHED] Dettmers et al., “LLM.int8()” (2022) — outlier features, and why naive INT8 fails.
- [REFERENCE] TensorRT-LLM docs and its
docs/source/performancepages. - [REFERENCE] NVIDIA Transformer Engine (FP8) documentation.
- [EMERGING] “KIVI”, “KVQuant” and related KV-cache quantization papers.
VIII — Inference Serving Systems#
- [ESTABLISHED] vLLM source:
vllm/core/scheduler.py,vllm/core/block_manager*,vllm/attention/. Reading real schedulers beats reading about schedulers. - [ESTABLISHED] Zheng et al., “SGLang: Efficient Execution of Structured Language Model Programs” (2023) — RadixAttention / prefix caching.
- [REFERENCE] Hugging Face TGI architecture docs.
- [REFERENCE] NVIDIA Triton Inference Server docs — especially dynamic batching and the backend API.
- [REFERENCE] ONNX Runtime performance tuning docs.
- [FUNDAMENTAL] “The Tail at Scale” (Dean & Barroso, CACM 2013) — required reading for anyone owning a latency SLO.
- [FUNDAMENTAL] Little’s Law — any queueing theory primer.
L = λW.
IX — Distributed Inference#
- [ESTABLISHED] Shoeybi et al., “Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism” (2019) — where TP layouts come from.
- [ESTABLISHED] Pope et al., “Efficiently Scaling Transformer Inference” (2022) — again; it is the best partitioning analysis available.
- [REFERENCE] NCCL documentation and the NCCL tests repo (
nccl-tests) for measuring real collective bandwidth. - [REFERENCE] NVIDIA NVLink / NVSwitch technical briefs.
- [FUNDAMENTAL] Any treatment of the ring-allreduce algorithm (Baidu’s original writeup).
- [ESTABLISHED] Huang et al., “GPipe” (2018) — pipeline bubbles.
X — Memory & Performance Engineering#
- [FUNDAMENTAL] Roofline paper (see I).
- [REFERENCE] Nsight Compute metrics reference — learn
dram__throughput,sm__throughput,l2_tex__t_bytes. - [FUNDAMENTAL] Brendan Gregg, flame graphs;
py-spyfor Python-side stalls. - [REFERENCE] PyTorch Profiler and
torch.cuda.memory_summary()/ memory snapshot tooling. - [REFERENCE] DCGM exporter + Prometheus + Grafana for GPU fleet metrics.
- [ESTABLISHED] “USE Method” (Utilization, Saturation, Errors) — Gregg.
XI — Production Inference Engineering#
- [FUNDAMENTAL] Google SRE Book, chapters on SLOs, overload, and handling cascading failures — https://sre.google/books/
- [FUNDAMENTAL] “The Tail at Scale” (again).
- [ESTABLISHED] Public engineering writeups on LLM cost-per-token modeling (treat specific numbers as dated; take the method).
- [REFERENCE] OpenTelemetry semantic conventions for GenAI spans.
- [REFERENCE] OWASP Top 10 for LLM Applications — for the security/abuse section.
XII — Inference Platform Engineering#
- [REFERENCE] Kubernetes docs: device plugins, scheduling framework, topology manager,
ResourceQuota, HPA/KEDA. - [REFERENCE] NVIDIA GPU Operator and DCGM.
- [ESTABLISHED] KServe and Ray Serve architecture docs — two different answers to the same problem.
- [EMERGING] Kubernetes Gateway API Inference Extension — model-aware routing.
- [FUNDAMENTAL] “Control plane vs data plane” as articulated in networking literature; the distinction transfers exactly.
XIII — Advanced LLM Inference#
- [ESTABLISHED] Fedus et al., “Switch Transformers” (2021); Lepikhin et al., “GShard” (2020) — MoE routing.
- [ESTABLISHED] DeepSeek-V2 / V3 technical reports — MLA, MoE at scale, and unusually candid inference engineering detail.
- [ESTABLISHED] Agrawal et al., “Sarathi-Serve: Taming Throughput-Latency Tradeoff in LLM Inference” (2024) — chunked prefill.
- [EMERGING] Zhong et al., “DistServe: Disaggregating Prefill and Decoding” (2024); Splitwise (Microsoft, 2024).
- [EMERGING] Cai et al., “Medusa” (2024); Li et al., “EAGLE” / “EAGLE-2” (2024).
- [REFERENCE] OpenAI Triton language docs — https://triton-lang.org/
- [REFERENCE]
torch.compileinternals; TorchInductor design notes.
XIV — Research & Frontier#
- [REFERENCE] arXiv
cs.LG+cs.DC— filter for “inference”, “serving”, “KV cache”. - [REFERENCE] MLSys, OSDI, SOSP, ASPLOS, ISCA proceedings — the venues where inference systems work actually lands.
- [REFERENCE] MLPerf Inference results — https://mlcommons.org/benchmarks/inference/ — the closest thing to apples-to-apples hardware comparison.
- [REFERENCE] AWS Neuron SDK docs (Inferentia/Trainium); Google Cloud TPU docs; AMD ROCm / composable_kernel docs.
- [EMERGING] Processing-in-memory and near-memory literature (UPMEM, HBM-PIM papers).
2025-2026 additions#
Last reviewed: 2026-10-03. The sections above are the durable core. This block is the part that ages: newer papers, engine write-ups, and the feeds worth following. Everything here is [EMERGING] unless marked otherwise — apply the paper-reading rules at the bottom of this file, especially to vendor blogs, which report their best configuration.
The book-length treatment#
- [ESTABLISHED] Philip Kiely, Inference Engineering (Baseten Books, 2026) — the first full book on the discipline; spans models, hardware, runtimes, and production. https://www.baseten.co/inference-engineering/book/
- [REFERENCE] DeepLearning.AI × vLLM, “Fast & Efficient LLM Inference with vLLM” (short course) — https://vllm.ai/blog/2026-06-03-deeplearning-ai-vllm-course
V / VIII — Engine internals (read alongside the vLLM and SGLang sections)#
- [ESTABLISHED] “Inside vLLM: Anatomy of a High-Throughput LLM Inference System” — the best single walkthrough of a real engine. https://vllm.ai/blog/2025-09-05-anatomy-of-vllm
- [ESTABLISHED] “vLLM V1: A Major Upgrade to Core Architecture” — https://vllm.ai/blog/2025-01-27-v1-alpha-release
- “Model Runner V2: A Modular and Faster Core for vLLM” — https://vllm.ai/blog/2026-03-24-mrv2
- “Structured Decoding in vLLM: a gentle introduction” — https://vllm.ai/blog/2025-01-14-struct-decode-intro
- “Zero-Reload Model Switching with vLLM Sleep Mode” (cold starts, multi-model) — https://vllm.ai/blog/2025-10-26-sleep-mode
- “Efficiently serve dozens of fine-tuned models with vLLM” (multi-LoRA) — https://vllm.ai/blog/2026-02-26-multi-lora
- “Keeping vLLM Production Quality: CI, Benchmarking, and Release Process” — https://vllm.ai/blog/2026-07-16-keeping-vllm-production-quality
- SGLang development and P/D disaggregation roadmaps — the most direct view of where an engine is heading: https://github.com/sgl-project/sglang/issues/22949 , https://github.com/sgl-project/sglang/issues/21703
VI / VII — Kernels, compilers, quantization#
- Zadouri, Dao et al., “FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling” (2026) — Blackwell-era attention; the bottleneck moves to the softmax exponential. https://arxiv.org/abs/2603.05451 · https://tridao.me/blog/2026/flash4/
- “vLLM Triton Attention Backend Deep Dive” — https://vllm.ai/blog/2026-03-04-vllm-triton-backend-deep-dive
- “Introduction to torch.compile and How It Works with vLLM” — https://vllm.ai/blog/2025-08-20-torch-compile
- “CUDA Core Dump: Debugging Memory Access Issues” and “Tracing Hanging and Complicated GPU Kernels Down To The Source Code” — practical GPU debugging. https://vllm.ai/blog/2025-08-11-cuda-debugging · https://vllm.ai/blog/2025-12-03-improved-cuda-debugging
- “The State of FP8 KV-Cache and Attention Quantization in vLLM” — https://vllm.ai/blog/2026-04-22-fp8-kvcache
- Zandieh et al., “TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate” (ICLR 2026) — https://arxiv.org/abs/2504.19874 . Read it together with the independent evaluation, “A First Comprehensive Study of TurboQuant” (https://vllm.ai/blog/2026-05-11-turboquant), and the prior-work notes (https://arxiv.org/abs/2604.18555 , https://arxiv.org/abs/2604.19528). A good exercise in checking a headline compression claim against baselines.
- “Advancing Low-Bit Quantization for LLMs: AutoRound x LLM Compressor” — https://vllm.ai/blog/2025-12-09-intel-autoround-llmc
V.12 / VII.13 / XIII.09 — Speculative decoding#
- Li et al., “EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test” (2025) — https://arxiv.org/abs/2503.01840
- “P-EAGLE: Faster LLM inference with Parallel Speculative Decoding” — https://vllm.ai/blog/2026-03-13-p-eagle
- “EAGLE 3.1: Advancing Speculative Decoding Through Collaboration” — https://vllm.ai/blog/2026-05-26-eagle-3-1
- “Parallel All the Way Down: Beyond Single-Token Generation with Speculative Decoding” — https://vllm.ai/blog/2026-07-28-speculators-parallel-drafting
- “Correctness Forensics for Batch Speculative Decoding: Diagnosing the Ragged Tensor Problem” — why batched speculation is hard to get right. https://arxiv.org/abs/2510.22876
IX / XIII — Disaggregation, wide expert parallelism, KV as a managed resource#
The largest shift since this curriculum’s core papers: prefill/decode disaggregation and tiered KV memory moved from research into the default large-scale configuration.
- “Taking vLLM Apart: A Practical Guide to Disaggregated Serving” — https://vllm.ai/blog/2026-09-29-disaggregated-serving-guide
- “Next-Level Inference: Single-Node vLLM Setup Needs PD Disaggregation” — https://vllm.ai/blog/2026-04-07-moriio-kv-connector
- “vLLM Router: A High-Performance Prefill/Decode Aware Load Balancer” — https://vllm.ai/blog/2025-12-13-vllm-router-release
- “vLLM Large Scale Serving: DeepSeek @ 2.2k tok/s/H200 with Wide-EP” — https://vllm.ai/blog/2025-12-17-large-scale-serving ; follow-ups on Blackwell (https://vllm.ai/blog/2026-02-03-dsr1-gb200-part1) and “Elastic Expert Parallelism” (https://vllm.ai/blog/2026-05-14-elastic-expert-parallelism)
- “Efficient Decode Context Parallelism with vLLM for Long Context Workloads” — https://vllm.ai/blog/2026-08-07-decode-context-parallelism
- “Inside vLLM’s New KV Offloading Connector” and “Tiered KV Cache Offloading in vLLM” — https://vllm.ai/blog/2026-01-08-kv-offloading-connector · https://vllm.ai/blog/2026-09-10-tiered-kv-offloading
- Cheng et al., “LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference” — https://arxiv.org/abs/2510.09665
- “Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs” — measured break-evens for offloading. https://arxiv.org/abs/2609.11744
- [REFERENCE] NVIDIA Dynamo — KV Block Manager architecture docs. https://docs.nvidia.com/dynamo/latest/architecture/kvbm_intro.html
- [REFERENCE] Survey + paper list: “Towards Efficient LLM Serving: A Survey on System-Aware KV Cache Optimization” — https://github.com/jjiantong/Awesome-KV-Cache-Optimization
- “vLLM x AgentX: Optimizing for Real-World Agentic Serving” and “Serving Agentic Workloads at Scale with vLLM x Mooncake” — agent traffic (long, growing, cache-heavy contexts) as a distinct workload. https://vllm.ai/blog/2026-09-08-vllm-agentx · https://vllm.ai/blog/2026-05-06-mooncake-store
XI / XII — Platform, routing, Kubernetes#
- [REFERENCE] “Introducing Gateway API Inference Extension” (Kubernetes blog) —
https://kubernetes.io/blog/2025/06/05/introducing-gateway-api-inference-extension/ .
InferencePoolhas since graduated to a stable v1 API, and the endpoint picker moved to the llm-d project; the releases page has the current state: https://github.com/kubernetes-sigs/gateway-api-inference-extension/releases - [REFERENCE] “Welcome llm-d to the CNCF” (March 2026) — https://www.cncf.io/blog/2026/03/24/welcome-llm-d-to-the-cncf-evolving-kubernetes-into-sota-ai-infrastructure/
- [REFERENCE] “Kubernetes v1.37” release post (August 2026) — DRA device taints reach stable. https://kubernetes.io/blog/2026/08/26/kubernetes-v1-37-release/
- [REFERENCE] “How NVIDIA Dynamo 1.0 Powers Multi-Node Inference at Production Scale” — https://developer.nvidia.com/blog/nvidia-dynamo-1-production-ready/ (vendor source)
- [REFERENCE] “Understanding dynamic resource allocation in Kubernetes” (CNCF) — DRA is the successor to device plugins for GPU requests. https://www.cncf.io/blog/2026/07/01/understanding-dynamic-resource-allocation-in-kubernetes/
- llm-d — Kubernetes-native distributed inference; the blog is consistently concrete:
- “KV-Cache Wins You Can See: From Prefix Caching in vLLM to Distributed Scheduling with llm-d” — https://llm-d.ai/blog/kvcache-wins-you-can-see
- “Sticky Until Saturated: Token-Aware Routing in llm-d” — https://llm-d.ai/blog/sticky-until-saturated-token-aware-routing
- “Pull, Don’t Recompute: Peer-to-Peer KV Cache Sharing in llm-d” — https://llm-d.ai/blog/p2p-kv-cache-sharing-llm-d
- “Seeing Through the Stack: End-to-End and Fine-Grained Tracing in llm-d” — https://llm-d.ai/blog/end-to-end-and-fine-grained-tracing-in-llm-d
- “llm-d v0.9: Hardened for Scale” — https://llm-d.ai/blog/llm-d-v0.9-hardened-for-scale
- “Introducing AIBrix: A Scalable, Cost-Effective Control Plane for vLLM” — https://vllm.ai/blog/2025-02-21-aibrix-release
- “Streamlined multi-node serving with Ray symmetric-run” — https://vllm.ai/blog/2025-11-22-ray-symmetric-run
X / XIV — Benchmarks and hardware#
- [REFERENCE] InferenceX (formerly InferenceMAX, SemiAnalysis) — continuously re-run, open-source benchmark across engines and accelerators. The most useful public view of throughput-vs-latency curves per GPU. https://inferencex.semianalysis.com/ · https://github.com/SemiAnalysisAI/InferenceX
- [REFERENCE] MLPerf Inference v6.0 results (April 2026) — adds a serving-style load generator for LLMs. https://mlcommons.org/2026/04/mlperf-inference-v6-0-results/
- “Inside the NVIDIA Vera Rubin Platform: Six New Chips, One AI Supercomputer” — https://developer.nvidia.com/blog/inside-the-nvidia-rubin-platform-six-new-chips-one-ai-supercomputer/ . Rubin has been in production shipment since August 2026; at GTC 2026 a seventh chip, the Groq-derived LPU, joined the platform (see XIV.09). Vendor figures — check them against InferenceX as measurements appear.
- “vLLM TPU”, “Optimizing vLLM on Arm CPUs”, “Serving LLMs on Tenstorrent Hardware”, and “Beyond Porting: … AMD ROCm” — how one engine adapts to non-NVIDIA targets; useful companions to XIV.10. https://vllm.ai/blog/2025-10-16-vllm-tpu · https://vllm.ai/blog/2026-07-29-optimizing-vllm-on-arm-cpus · https://vllm.ai/blog/2026-09-07-vllm-tt-plugin · https://vllm.ai/blog/2026-02-27-rocm-attention-backend
Observability#
- The Observability Engineering path in this library is the long treatment; its VI.03 lists the metrics, project statuses and feeds to re-check.
- [REFERENCE] vLLM metrics reference — https://docs.vllm.ai/en/latest/usage/metrics.html
- [REFERENCE] OpenTelemetry GenAI semantic conventions (all still “Development”) — https://github.com/open-telemetry/semantic-conventions-genai
Status changes worth knowing#
- TGI is archived. Hugging Face moved Text Generation Inference to maintenance mode in December 2025 and archived the repository in March 2026; its docs recommend vLLM or SGLang.
- vLLM was at v0.30 (22 September 2026) at the time of review.
Feeds to follow#
Check these roughly weekly; they are where changes to this curriculum’s [EMERGING] tier show up first.
| Feed | Why | URL |
|---|---|---|
| vLLM blog | Engine internals, new techniques landing in production | https://vllm.ai/blog |
| SGLang blog + releases | The other reference engine; often first on cache and P/D work | https://www.sglang.io/blog · https://github.com/sgl-project/sglang/releases |
| llm-d blog | Kubernetes-side routing, scheduling, autoscaling | https://llm-d.ai/blog |
| NVIDIA Developer blog | Hardware, Dynamo, TensorRT-LLM, CUDA | https://developer.nvidia.com/blog |
| Tri Dao’s blog | Attention kernels from the source | https://tridao.me/blog/ |
| InferenceX | Whether last month’s claims held up on real hardware | https://inferencex.semianalysis.com/ |
| SemiAnalysis newsletter | Hardware economics and roadmaps | https://newsletter.semianalysis.com/ |
| Baseten blog | Practitioner write-ups from a serving provider | https://www.baseten.co/blog/ |
| Brendan Gregg | Performance methodology, now applied to AI systems | https://www.brendangregg.com/ |
arXiv cs.DC + cs.LG | Filter: “LLM serving”, “KV cache”, “speculative decoding” | https://arxiv.org/list/cs.DC/recent |
| MLSys / OSDI / SOSP / ASPLOS programs | Where serving-systems work is peer reviewed | — |
Community study repos#
elizabetht/100-days-of-inference— daily experiments structured around the Kiely book. https://github.com/elizabetht/100-days-of-inferenceamitshekhariitbhu/llm-inference-engineering— a step-by-step outline covering similar ground to Sections V and VIII. https://github.com/amitshekhariitbhu/llm-inference-engineering
How to read a paper in this field#
- Read the abstract, then jump to the evaluation setup (hardware, model, batch sizes, sequence lengths). If the setup does not resemble your workload, the speedup number does not transfer.
- Find the baseline. Half of reported speedups are against a weak baseline.
- Find the bottleneck they claim to remove and check it against the roofline. If a paper claims a 3x decode speedup without reducing bytes moved or improving arithmetic intensity, be suspicious.
- Ask what breaks at long context, large batch, or under multi-tenancy. Papers optimize for one regime.
See XIV-research-and-frontier/01-reading-papers.md for the long version.