Below the API

Memory Planning for LLMs

Advanced Intermediate 1h 10m Difficulty 4/5

Prerequisites III.03, II.03

The idea in one minute#

An LLM server’s GPU memory holds three things: the weights (fixed), the KV cache (grows with every token of every active conversation), and working space (activations and overhead). The weights are a one-time cost. The KV cache is what actually limits how many users fit — and it is routinely larger than the model.

You can compute all of it from the model’s configuration before renting a single GPU.

A picture#

pie showData
  title 80 GB GPU serving an 8B model, FP16
  "Weights" : 16
  "KV cache pool" : 56
  "Activations and overhead" : 4
  "Reserved headroom" : 4

How it really works#

Weights#

weight bytes = parameters × bytes per parameter
8B × 2 bytes (FP16) = 16 GB

The KV cache#

To generate each new token, a transformer attends to every previous token. Rather than recompute the per-token keys and values each step, it stores them — that store is the KV cache. It is what makes generation affordable, and it costs memory per token:

KV bytes per token = 2 × layers × kv_heads × head_dim × bytes per value

(2 for a key and a value. kv_heads is the number of key/value heads, often fewer than the attention heads — a design called grouped-query attention.)

For a typical 8B model (32 layers, 8 KV heads, head dimension 128, FP16):

2 × 32 × 8 × 128 × 2 bytes = 131,072 bytes ≈ 128 KiB per token

A 4,000-token conversation therefore holds about 0.5 GB of KV cache. A hundred of them hold 50 GB — three times the weights.

The capacity formula#

concurrent sequences = (GPU memory − weights − overhead) ÷ (KV per token × tokens per sequence)

Everything about LLM capacity planning hangs on this line. Its consequences:

  • Context length is paid for in users. Eight times the context means one eighth the concurrency.
  • Quantizing weights frees memory for KV, which can raise concurrency more than it raises speed.
  • Model architecture matters. Fewer KV heads means proportionally smaller KV per token.

How servers manage it#

They apply III.03’s block pool directly: reserve most of the GPU at start-up and split the KV region into fixed-size blocks (for example 16 tokens each). A sequence takes blocks as it grows and returns them when it ends. No fragmentation, no allocation in the hot path, and identical prefixes can share blocks. This technique is known as paged attention.

When the pool is full#

Running out of KV blocks is the LLM server’s version of out-of-memory, and it is normal operation, not a crash. The server must choose: queue new requests, reject them, or pause a running sequence and resume it later. Which one happens is a policy decision made above the GPU — the subject of the inference course.

Code#

A capacity planner you can point at any model and GPU.

// plan.go — will it fit, and how many users at once?
package main

import "fmt"

type Model struct {
	Name                     string
	Params                   float64
	Layers, KVHeads, HeadDim int
	WeightBytes, KVBytes     float64 // bytes per weight, bytes per KV value
}

func (m Model) WeightsGB() float64 { return m.Params * m.WeightBytes / 1e9 }
func (m Model) KVPerToken() float64 {
	return 2 * float64(m.Layers*m.KVHeads*m.HeadDim) * m.KVBytes
}

func plan(m Model, gpuGB float64, gpus int, tokensPerSeq float64) {
	const overhead = 0.10 // activations, runtime, headroom
	total := gpuGB * float64(gpus)
	free := total*(1-overhead) - m.WeightsGB()
	fmt.Printf("%-22s on %dx%.0fGB  weights %5.1f GB  KV/token %4.0f KiB  ",
		m.Name, gpus, gpuGB, m.WeightsGB(), m.KVPerToken()/1024)
	if free <= 0 {
		fmt.Println("DOES NOT FIT")
		return
	}
	seqs := free * 1e9 / (m.KVPerToken() * tokensPerSeq)
	fmt.Printf("-> %4.0f sequences of %.0f tokens\n", seqs, tokensPerSeq)
}

func main() {
	m8 := Model{"8B FP16", 8e9, 32, 8, 128, 2, 2}
	m8q := Model{"8B 4-bit weights", 8e9, 32, 8, 128, 0.5, 2}
	m70 := Model{"70B FP16", 70e9, 80, 8, 128, 2, 2}
	m70q := Model{"70B 8-bit, FP8 KV", 70e9, 80, 8, 128, 1, 1}

	plan(m8, 24, 1, 4096)
	plan(m8q, 24, 1, 4096)
	plan(m8, 80, 1, 4096)
	plan(m8, 80, 1, 32768)
	plan(m70, 80, 1, 4096)
	plan(m70, 80, 2, 4096)
	plan(m70q, 80, 1, 4096)
	plan(m70, 180, 1, 4096)
}

Compare lines 3 and 4 (context length) and lines 1 and 2 (weight quantization on a small GPU). Those two comparisons are most of practical capacity planning.

Remember this#

  • GPU memory = weights + KV cache + working space.
  • KV per token = 2 × layers × kv_heads × head_dim × bytes. Often ~100 KiB–1 MiB.
  • Concurrency = free memory ÷ (KV per token × tokens per sequence).
  • Servers pre-allocate the KV region as a pool of fixed-size blocks.

Try it#

  1. Run plan.go. Look up the real configuration of a model you use and add it.
  2. For the 8B model on 80 GB: at what context length does one sequence’s KV cache equal the weights?
  3. Your traffic averages 2,000 tokens per sequence but the server allows 32,000. Should you plan for the average, the maximum, or something else? What happens at the pool limit?

Check yourself#

  1. Write the KV-cache-per-token formula and say what each factor is.
  2. Why does doubling the context halve concurrency?
  3. Why can quantizing weights raise the number of users served even if speed is unchanged?

↑↓ navigate ↵ open