Below the API

GLOSSARY

Blunt definitions. Where a term is commonly misunderstood, the misunderstanding is called out. Format: Term (section where it is taught) — definition.


A#

Activation (I, III) — the numeric output of a layer, computed at runtime from the input. Unlike weights, activations are not stored in the model file; they exist only during a forward pass. Memory for activations scales with batch size and sequence length; memory for weights does not.

Activation quantization (VII) — quantizing the runtime tensors, not just weights. Harder than weight quantization because activations have outliers that vary per input.

Admission control (VIII) — deciding to reject or defer a request before it enters the system, to protect the requests already inside. The opposite of “accept everything and hope.”

AllGather (IX) — a collective where every rank ends up with the concatenation of all ranks’ data. Used in tensor parallelism when each GPU computes a slice of an output that all GPUs then need.

AllReduce (IX) — a collective where every rank ends up with the elementwise sum (or other reduction) of all ranks’ data. The dominant communication primitive in tensor-parallel inference.

Arithmetic intensity (I, X) — FLOPs performed per byte moved from memory. The single most useful number in inference performance. Low intensity = memory bound. Written as FLOP/byte.

Attention (III, IV, V) — the transformer operation where each token computes a weighted average over other tokens’ values, with weights derived from query·key similarity.

Autoregressive (V) — generating output one token at a time, where each token depends on all previously generated tokens. This serial dependency is the reason LLM decode is hard to make fast.

AWQ (Activation-aware Weight Quantization) (VII) — a weight-only quantization method that protects the small fraction of weight channels that matter most, identified by activation magnitude.

B#

Backpressure (VIII) — signalling upstream that you are overloaded so it slows down, instead of silently queueing forever. Usually implemented as bounded queues + rejection.

Bandwidth (memory) (I, VI) — bytes per second movable between memory and compute. HBM3e on an H200 is roughly 4.8 TB/s; DDR5 on a server CPU is roughly 0.3-0.5 TB/s. The difference is most of why GPUs win at inference.

Batch (I, V) — a group of requests processed together in one pass so that weights are loaded from memory once and amortized across all of them.

Beam search (V) — decoding that keeps the k best partial sequences instead of one. Common in translation, rare in modern LLM chat serving because it multiplies KV cache cost.

BF16 (bfloat16) (III) — 16-bit float with FP32’s exponent range and fewer mantissa bits. Trades precision for range; the default for modern training and much inference.

Block (CUDA) (VI) — a group of threads that run on one SM and can share shared memory and synchronize with each other.

C#

Chunked prefill (XIII) — splitting a long prompt’s prefill into pieces so it can be interleaved with decode steps, preventing one long prompt from stalling everyone else.

Coalescing (VI) — when the 32 threads of a warp access contiguous memory so the hardware services them with the minimum number of memory transactions. Uncoalesced access can cost 32x.

Cold start (VIII, XI) — the time from “no replica exists” to “replica serves traffic.” For LLMs dominated by weight loading (tens of GB), not container start.

Collective (IX) — a communication operation involving all ranks in a group: AllReduce, AllGather, ReduceScatter, Broadcast, All-to-All.

Continuous batching (V, VIII) — a scheduler that adds and removes requests from the running batch at every decode step, instead of waiting for the whole batch to finish. Also called iteration-level scheduling or in-flight batching. The single biggest throughput win in modern LLM serving.

Context length (V) — the maximum number of tokens (prompt + generation) a model can attend over. Directly sets the worst-case KV cache size per sequence.

CUDA core (VI) — a scalar ALU lane in an NVIDIA GPU. Not analogous to a CPU core; it is much closer to a single SIMD lane.

CUDA graph (VI, VII) — a recorded DAG of GPU operations replayed with one launch, eliminating per-kernel launch overhead. Critical for small-batch decode where kernels are microseconds long.

D#

Decode (V) — the phase where the model generates tokens one at a time, each step processing exactly one new token per sequence. Memory-bandwidth bound at low batch size.

Disaggregated serving (XIII) — running prefill and decode on separate GPU pools, moving the KV cache between them, so each pool can be sized and scheduled for its own bottleneck.

DMA (Direct Memory Access) (II) — hardware moving data between memories without CPU involvement. How host↔device copies actually happen.

Dynamic batching (VIII) — waiting a short window to gather multiple requests into one batch. Predates continuous batching; still correct for non-autoregressive models.

E#

Expert parallelism (EP) (IX, XIII) — placing different MoE experts on different GPUs and routing tokens to them with All-to-All communication.

F#

FlashAttention (VII) — an attention algorithm that never materializes the full N×N attention matrix in HBM, instead tiling and using online softmax in SRAM. IO-aware, exact (not an approximation).

FLOP (I) — one floating point operation. FLOPs = a count; FLOP/s = a rate. Confusing the two is the most common unit error in the field.

FP8 (III, VII) — 8-bit float, in two common variants: E4M3 (more precision, less range, used for weights/activations) and E5M2 (more range, used for gradients). Natively supported from NVIDIA Hopper onward.

Fragmentation (V, X) — memory that is free but unusable because it is not contiguous in large enough pieces. The problem PagedAttention solves for KV cache.

G#

GEMM (IV) — GEneral Matrix Multiply, C = alpha*A*B + beta*C. The workhorse operation of prefill and of batched decode.

GEMV (IV) — matrix-vector multiply. What decode at batch size 1 degenerates into. Hopeless arithmetic intensity — roughly 2 FLOPs per byte of weight read.

GPTQ (VII) — a one-shot post-training weight quantization method using second-order (Hessian) information to choose rounding that minimizes layer output error.

GQA (Grouped-Query Attention) (XIII) — attention where several query heads share one key/value head. Shrinks KV cache by the grouping factor. Used by Llama 2 70B onward.

Grid (VI) — the full set of thread blocks launched by one kernel.

H#

HBM (High Bandwidth Memory) (VI) — the stacked DRAM on a GPU package. What people mean by “GPU memory” or “VRAM.” Fast (TB/s) but still ~100x slower than on-chip SRAM.

Head (III) — one independent attention computation. A model with 32 heads runs 32 attention computations in parallel over different learned projections.

I#

InfiniBand (IX) — a low-latency network fabric used between nodes in GPU clusters, typically with RDMA.

Inter-token latency (ITL) (I, V) — time between successive generated tokens. Also called TPOT (time per output token). Determines whether streaming output “feels” fast.

INT8 / INT4 (III, VII) — integer quantization formats. INT8 is production-standard; INT4 weight-only is production-common for weight-bound decode; INT4 activations are still hard.

K#

Kernel (IV, VI) — a function that runs on the GPU. “Kernel launch” = the CPU telling the GPU to run one.

Kernel fusion (IV, VII) — combining several operations into one kernel so intermediate results stay in registers/SRAM instead of round-tripping to HBM.

KV cache (V) — the stored key and value tensors for all previous tokens, kept so that each new token does not recompute attention over the whole history. Turns O(n²) regeneration into O(n) per step, at the cost of memory that grows linearly with tokens and batch size.

L#

Latency (I) — time for one request. Always specify which latency: TTFT, ITL, or end-to-end.

Load balancing (VIII, XII) — distributing requests across replicas. Naive round-robin is usually wrong for LLMs because request cost varies by 1000x and because prefix cache locality matters.

M#

MLA (Multi-head Latent Attention) (XIII) — DeepSeek’s attention variant that stores a compressed latent KV representation, dramatically shrinking the KV cache.

MoE (Mixture of Experts) (XIII) — a model where each token is routed to a small subset of “expert” FFN blocks. Large total parameters, small active parameters per token.

MQA (Multi-Query Attention) (XIII) — all query heads share a single key/value head. The extreme of GQA; smallest KV cache, some quality cost.

N#

NCCL (IX) — NVIDIA Collective Communications Library. The implementation of AllReduce and friends used by essentially every multi-GPU inference stack.

NUMA (II) — Non-Uniform Memory Access: on multi-socket servers, memory attached to another socket is slower. Matters for CPU-side tokenization, data loading, and pinned buffers.

NVLink (IX) — NVIDIA’s high-bandwidth GPU-to-GPU interconnect; hundreds of GB/s to over a TB/s per GPU, versus ~64 GB/s for PCIe Gen5 x16. Whether your GPUs have it changes tensor parallelism from “good idea” to “bad idea.”

O#

Occupancy (VI) — the ratio of active warps on an SM to the maximum possible. High occupancy helps hide memory latency; it is a means, not a goal.

Operator fusion — see Kernel fusion.

OOM (Out Of Memory) (X) — on GPUs, usually caused by KV cache growth or an unexpectedly long request, not by weights. Weights are predictable; KV cache is not.

P#

PagedAttention (V) — vLLM’s technique of storing KV cache in fixed-size non-contiguous blocks with a block table, borrowing virtual-memory paging ideas. Eliminates KV fragmentation and enables cheap sharing/copy-on-write between sequences.

Parameter (I, III) — a learned number in the model. “7B parameters” = 7 billion learned numbers. Memory = parameters × bytes per parameter.

PCIe (II, IX) — the bus connecting CPU and GPU (and sometimes GPU to GPU). Gen4 x16 ≈ 32 GB/s theoretical, ~25 GB/s real; Gen5 x16 ≈ 64 GB/s theoretical.

Pipeline parallelism (PP) (IX) — splitting a model by layers across GPUs. Low communication, but introduces bubbles unless carefully scheduled with microbatches.

Prefill (V) — the phase that processes the entire input prompt in one pass to produce the first output token and populate the KV cache. Compute bound; parallel over all prompt tokens.

Prefix caching (V, XII) — reusing the KV cache of a shared prompt prefix (system prompt, few-shot examples, document) across requests. Often the cheapest large win available.

Q#

Quantization (III, VII) — representing weights and/or activations in fewer bits. Reduces memory footprint and, crucially for decode, memory traffic.

QPS (I) — queries per second. A poor headline metric for LLMs, because a “query” can be 30 tokens or 30,000. Prefer tokens/sec plus a distribution of request shapes.

R#

RDMA (IX) — Remote Direct Memory Access: a NIC writing directly into remote memory without involving the remote CPU. The basis of fast multi-node inference.

ReduceScatter (IX) — reduce across ranks, then scatter the result so each rank holds one slice. AllReduce = ReduceScatter + AllGather.

Roofline model (X) — a plot of achievable FLOP/s vs arithmetic intensity, bounded by a sloped bandwidth line and a flat compute ceiling. The standard tool for answering “is this kernel memory or compute bound?”

S#

Sampling (V) — choosing the next token from the model’s probability distribution. Temperature, top-k, top-p, min-p are its knobs.

Shared memory (CUDA) (VI) — fast programmer-managed SRAM scoped to a thread block. FlashAttention’s home.

SLO / SLI (XI) — Service Level Objective (the target) and Indicator (the measurement). For LLMs, define both TTFT and ITL SLOs, not a single “latency” SLO.

SM (Streaming Multiprocessor) (VI) — the fundamental compute unit of an NVIDIA GPU. An H100 has 132 of them. Each runs many warps concurrently.

SmoothQuant (VII) — a technique that migrates activation outlier magnitude into the weights via per-channel scaling, making INT8 activation quantization feasible.

Speculative decoding (V, VII, XIII) — a small draft model (or a head, or n-grams) proposes several tokens; the big model verifies them in one forward pass. Trades spare compute for reduced memory traffic per accepted token.

Streaming (V, VIII) — sending tokens to the client as they are generated (SSE, chunked HTTP, gRPC stream) rather than waiting for the full response.

T#

Tensor (I, III) — an n-dimensional array. Scalar (0D), vector (1D), matrix (2D), and beyond.

Tensor core (VI) — a hardware unit that performs a small matrix multiply-accumulate per instruction. Delivers an order of magnitude more FLOP/s than CUDA cores for GEMM in supported precisions.

Tensor parallelism (TP) (IX) — splitting individual weight matrices across GPUs so every GPU does part of each layer. Requires an AllReduce per transformer block — hence needs NVLink to be efficient.

Throughput (I) — work per unit time. For LLMs: output tokens/sec, or total tokens/sec across the fleet.

TTFT (Time To First Token) (I, V) — from request arrival to first token emitted. Governed by queueing + prefill.

Token (V) — the model’s unit of text; typically a subword. ~4 characters of English on average, but wildly variable across languages and code.

U#

Utilization (GPU) (X) — nvidia-smi’s “GPU-Util” is the fraction of time at least one kernel was resident. It is not a measure of how much of the chip is doing useful work. A completely memory-stalled kernel shows 100%.

V#

vLLM (VIII) — an open-source LLM inference engine; introduced PagedAttention and popularized continuous batching.

Virtual memory (II) — the OS abstraction that gives each process its own address space. The conceptual ancestor of PagedAttention.

W#

Warp (VI) — 32 threads that execute in lockstep on an SM. The real unit of scheduling on an NVIDIA GPU.

Warp divergence (VI) — when threads in a warp take different branches, forcing the hardware to execute both paths serially with masking.

Weight (I, III) — a learned parameter used to transform inputs. Weights are what you load from disk; activations are what you compute.

Weight-only quantization (VII) — storing weights in low precision but computing in higher precision after dequantizing on the fly. The dominant approach for memory-bound decode.

↑↓ navigate ↵ open