1. Problem → Why → Optimization#
PROBLEM At long context and moderate batch, the KV cache exceeds the model
weights in both memory and bandwidth.
WHY KV scales with batch × context; weights are fixed.
OPTIMIZE Store K and V in FP8 or INT4 instead of FP16.
TRADE-OFFS
✓ 2x (FP8) or 4x (INT4) more concurrency at the same memory
✓ 2x or 4x less KV bandwidth per decode step
✓ quality cost is small for FP8, modest for INT4
✗ requires kernel support (dequantize inside the attention kernel)
✗ per-token scale storage overhead
✗ K is more sensitive than V (asymmetric treatment needed for INT4)
WHEN TO USE Long context (>8k), high batch, or any KV-bound workload.
FP8 KV is nearly free — enable it by default on Hopper.
WHEN NOT TO Short context and low batch, where weights dominate anyway.2. Why it’s a separate decision from weight quantization#
They attack different terms and their benefits are situational in opposite ways:
Weight quantization KV quantization
Helps when weights dominate bytes KV dominates bytes
(short ctx, low batch) (long ctx, high batch)
Memory saved fixed scales with load
Quality risk all layers affected attention only
Kernel changes GEMM prologue attention kernelA workload at 32k context with batch 64 gets more from KV quantization than from weight quantization. Compute your crossover (Section V.06) before choosing.
3. Simple analogy#
Compressing an archive versus compressing the working files.
Weight quantization compresses the reference library — a fixed one-time saving.
KV quantization compresses the growing pile of working notes. The saving grows with how much work is in flight, which is exactly when you need it.
4. Tiny example — the arithmetic#
Llama-3-8B, one H100 (80 GB), 32k context:
FP16 KV: 131,072 B/token × 32,768 = 4.29 GB per sequence
budget 60 GB → 13 concurrent sequences
decode step KV read at batch 13: 55.8 GB
+ 16 GB weights = 71.8 GB → 21.4 ms/step
FP8 KV: 65,536 B/token × 32,768 = 2.15 GB per sequence
budget 60 GB → 27 concurrent sequences 2.1x
decode step KV read at batch 27: 58.0 GB
+ 16 GB weights = 74 GB → 22.1 ms/step
→ 27 sequences at ~the same step time = 2.05x THROUGHPUT
INT4 KV: 32,768 B/token × 32,768 = 1.07 GB per sequence
budget 60 GB → 55 concurrent sequences 4.2x
(in practice, limited by other factors first)2x throughput from a config flag, at long context. This is one of the highest return-on-effort optimizations available, and it’s frequently overlooked because people think of quantization as a weights thing.
5. Technical explanation#
Where the scales live#
Per-token, per-head scaling is standard:
k_cache: (num_blocks, block_size, n_kv_heads, head_dim) in FP8/INT8/INT4
k_scale: (num_blocks, block_size, n_kv_heads) in FP16
Overhead for INT8, head_dim=128:
128 bytes of data + 2 bytes of scale = 1.6% overhead. Negligible.
For INT4, head_dim=128, group=32:
64 bytes of data + 4 groups × 2 bytes = 72 bytes vs 256 FP16 → 3.5x, not 4x.Some implementations use a single per-tensor scale (cheapest, worst quality) or per-channel (along head_dim — better for K, see below).
K is more sensitive than V#
An empirical finding that shapes the design:
K participates in: softmax(q·kᵀ/√d) — errors are AMPLIFIED by the exponential
V participates in: Σ p_i v_i — errors are AVERAGED, and damped
Measured: at INT4, quantizing only K costs ~3x more quality than quantizing only V.Consequences in practice:
- Per-channel (along head_dim) quantization for K, per-token for V. K’s outliers are channel-consistent, like activations (file 04).
- Asymmetric bit allocation: K at 8 bits, V at 4 bits is a common and effective choice.
- FP8 for K and INT4 for V.
KIVI and KVQuant both formalize versions of this.
FP8 vs INT4 for KV#
FP8 (E4M3)
✓ wide dynamic range absorbs outliers (same argument as file 05)
✓ per-tensor or per-token scaling is sufficient
✓ quality loss typically < 0.5%
✓ Hopper hardware support; simple kernels
→ THE DEFAULT CHOICE on modern hardware
INT4
✓ 4x instead of 2x
✗ needs group-wise scaling and careful K/V asymmetry
✗ quality loss 1-3%, more on long-context retrieval tasks
✗ dequantization work inside the attention kernel
→ use when memory-desperate and validatedThe kernel side#
The attention kernel must dequantize while reading:
for each KV block:
load quantized K block (FP8/INT4) from HBM ← 2-4x less traffic
load scales
dequantize in registers to FP16/FP32
compute q · kᵀ
...The dequantization is nearly free (attention decode is memory-bound), so the bandwidth saving translates almost fully into speed.
Support varies by engine and kernel. vLLM supports FP8 KV (--kv-cache-dtype fp8);
TensorRT-LLM supports FP8 and INT8. INT4 KV is less widely supported.
Interaction with prefix caching#
Quantized KV blocks are still shareable — the block hash covers the token ids, not the stored values. But note: a prefix computed and cached in FP8 and one recomputed will differ slightly (Section IV.12), so cache hits change results marginally. Expected and harmless, but it’s another source of nondeterminism to be aware of.
6-9. Under the hood, performance, production, mistakes#
Under the hood — measuring the benefit:
Speedup ≈ (weights + KV_fp16) / (weights + KV_fp16/2)
Llama-3-8B, batch 32:
context 1k: (16 + 4.3)/(16 + 2.1) = 1.12x ← small
context 8k: (16 + 34)/(16 + 17) = 1.52x
context 32k: (16 + 137)/(16 + 69) = 1.80x ← largeThe benefit grows with context and batch — exactly the regime where you need it.
Production:
- Enable FP8 KV cache by default on Hopper+ for any workload with context > 4k.
--kv-cache-dtype fp8in vLLM. - Validate on long-context retrieval tasks specifically — needle-in-a-haystack style evaluations are the sensitive ones. Perplexity barely moves; retrieval accuracy can.
- Report the effective concurrency gain in your capacity model.
- Consider asymmetric K/V precision if your engine supports it.
- Note the interaction with prefix caching: both consume the same block pool.
Mistakes:
- Not enabling it. The most common mistake — it’s a flag.
- Validating with perplexity only. KV quantization’s cost shows in long-range retrieval, which perplexity doesn’t measure.
- Quantizing K as aggressively as V. K is more sensitive.
- Per-tensor scales at INT4. Use groups.
- Assuming the benefit at short context. At 1k context it’s ~1.1x.
10. Hands-on exercise#
A. Measure the gain curve. Run a model with FP16 and FP8 KV cache. Measure max concurrency and throughput at context 1k, 4k, 16k, 32k. Plot the speedup vs context. Confirm it grows.
B. K vs V sensitivity. Quantize only K to INT4, then only V, then both. Evaluate on a long-context retrieval task. Quantify the asymmetry.
C. Retrieval validation. Build a needle-in-a-haystack test: place a fact at various depths in a 16k context and ask about it. Evaluate FP16, FP8, and INT4 KV. Plot retrieval accuracy vs depth for each. This is the evaluation that matters and the one usually skipped.
D. Capacity impact. Recompute your Section V.15 capacity plan with FP8 KV. How many fewer nodes do you need?
E. Overhead accounting. Compute the exact bytes per token including scales for FP8 per-token, INT8 per-token, and INT4 group-32. Compare to the naive 2x/4x claims.
11. Interview questions#
- When does KV quantization matter more than weight quantization?
- Why is K more sensitive to quantization than V? Give the mechanism.
- What granularity would you use for K and for V, and why do they differ?
- Why does the benefit of KV quantization grow with context length?
- What evaluation would you run to validate FP8 KV cache?
- Compute the effective bytes per token for INT4 KV with group size 32.
- Why is FP8 usually the right choice over INT8 for KV on Hopper?
12. Further reading#
- [EMERGING] Liu et al., “KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache” (2024)
- [EMERGING] Hooper et al., “KVQuant” (2024)
- [REFERENCE] vLLM
--kv-cache-dtypedocumentation - [REFERENCE] TensorRT-LLM FP8 KV cache documentation
- Next: 13 — Speculative decoding in practice