1. What is it?#
The maximum number of tokens (prompt + generation) a model can attend over. Modern models advertise 128k, 200k, or 1M. Serving those numbers is a fundamentally different engineering problem from serving 4k.
2. Why does it exist as a topic?#
Because “supports 128k context” and “can serve 128k context economically” are entirely different claims, and the gap between them is where a lot of money is lost.
4k context: KV cache is a manageable fraction of memory
32k context: KV cache exceeds the model weights
128k context: ONE sequence needs more memory than the model
1M context: requires a different architecture, not just more memory3. Simple analogy#
A conversation where you must remember every word.
At 10 sentences, trivial. At 1,000, you need notes. At 100,000, you need a filing system, a retrieval strategy, and probably an assistant — and you’ve fundamentally changed how the conversation works.
Long context isn’t “more of the same.” Past a point, it demands different mechanisms.
4. Tiny example#
The scaling, laid out:
package main
import "fmt"
func analysis(layers, nKV, headDim int, params, bytesPer, gpuGB float64) {
kvTok := 2 * float64(layers*nKV*headDim) * bytesPer
w := params * bytesPer / 1e9
budget := gpuGB - w - 4
fmt.Printf("weights %.1f GB, KV/token %.0f KiB, budget %.1f GB\n\n", w, kvTok/1024, budget)
fmt.Printf("%8s %10s %9s %14s\n", "context", "KV/seq", "max conc", "KV vs weights")
for _, S := range []float64{1024, 4096, 8192, 32768, 131072, 1000000} {
kvSeq := kvTok * S / 1e9
conc := 0
if kvSeq < budget {
conc = int(budget / kvSeq)
}
fmt.Printf("%8.0f %9.2fG %9d %13.2fx\n", S, kvSeq, conc, kvSeq/w)
}
}
func main() {
analysis(32, 8, 128, 8.03e9, 2, 80)
}
Output:
weights 16.1 GB, KV/token 128 KiB, budget 59.9 GB
context KV/seq max conc KV vs weights
1024 0.13G 446 0.01x
4096 0.54G 111 0.03x
8192 1.07G 55 0.07x
32768 4.29G 13 0.27x
131072 17.18G 3 1.07x
1000000 131.07G 0 8.16xThree things to notice. Concurrency falls linearly. At 128k, one sequence’s cache exceeds the model. At 1M, a single sequence doesn’t fit on an H100 at all — you need multiple GPUs to serve one user.
5. Technical explanation#
The four costs of long context#
Cost 1 — KV memory (linear in S). Covered above. Reduces concurrency proportionally.
Cost 2 — KV bandwidth (linear in S). Every decode step reads the whole cache:
8B model, batch 8:
context 4k: KV read = 4.3 GB/step, weights 16 GB → weights dominate
context 128k: KV read = 137 GB/step, weights 16 GB → KV dominates 8.6xITL at 128k is ~9x worse than at 4k even at the same batch size.
Cost 3 — prefill compute (quadratic in S).
prefill FLOPs ≈ 2·P·S + 2·L·S²·d
8B model:
S=4k: 2×8e9×4096 + 2×32×4096²×4096 = 65.5 + 4.4 TFLOP = 70 TFLOP
S=128k: 2×8e9×131072 + 2×32×131072²×4096 = 2,097 + 4,504 TFLOP = 6,601 TFLOP
32x more tokens → 94x more computeOn an H100 at 600 TFLOP/s achieved: 0.12 s vs 11 s. TTFT for a 128k prompt is 11 seconds of pure compute, before queueing.
Cost 4 — head-of-line blocking. An 11-second prefill blocks every decoding sequence unless chunked.
Why cost scales superlinearly overall#
Effect Scaling
KV memory S¹
KV bandwidth per step S¹
Prefill compute S² (attention term)
Concurrency S⁻¹
→ cost per token ≈ S¹ to S²Empirically, 4k → 32k typically costs 4-8x per token, not 8x-linear. This is why long-context API pricing is disproportionate and why “context caching” products exist.
The mitigations#
ARCHITECTURAL (chosen at training time)
GQA / MQA 4-8x smaller KV [ESTABLISHED]
MLA (DeepSeek) ~10-20x smaller KV [ESTABLISHED]
Sliding window KV capped at W, not S [ESTABLISHED]
Interleaved layers some global, most local [ESTABLISHED]
Linear/SSM hybrids O(1) state instead of O(S) [EMERGING]
SYSTEMS (chosen at serving time)
KV quantization 2-4x smaller (FP8, INT4 KV) [ESTABLISHED]
Chunked prefill removes HOL blocking [ESTABLISHED]
Prefix caching skip re-prefilling shared context [ESTABLISHED]
KV offload to CPU more capacity, needs fast link [EMERGING]
KV compression/eviction (H2O, SnapKV, StreamingLLM) [EMERGING]
Disaggregation separate prefill/decode pools [EMERGING]Sliding window attention — the biggest architectural lever#
Full attention: KV grows to S
Sliding window W: KV capped at W
Mistral-7B with W=4096:
context 4k: KV = 0.5 GB (same as full)
context 128k: KV = 0.5 GB (vs 16 GB for full attention!) ← 32x savingThe cost: information beyond W tokens back is only accessible indirectly, propagated through layers. With L layers and window W, the theoretical receptive field is L×W, but the practical ability to use distant information degrades.
Modern designs interleave: e.g. 1 global-attention layer for every 5 sliding-window layers. This caps most of the KV cache while retaining true long-range access in a few layers.
Context extension methods#
Models are trained at some context length and extended afterward:
Position Interpolation (PI) compress positions into the trained range
NTK-aware scaling adjust rope_theta by a scaling factor
YaRN NTK plus attention-temperature correction
LongRoPE search for per-dimension rescaling factorsAll of these change rope_theta or the RoPE computation. If you serve a long-context
fine-tune with the base model’s rope_theta, you get degraded output at long positions with no
error. Always read the config.
The “effective context” problem#
A model advertising 128k may not use 128k well. The “needle in a haystack” test and its successors measure retrieval at various depths; many models show degradation in the middle of long contexts (“lost in the middle”).
Practical implication: before building a product on 128k context, measure whether the model actually uses it for your task. You may be paying 8x for capability you don’t get.
6. Under the hood#
Attention at long context is dominated by KV reads, and the access pattern matters enormously:
Decode attention at S=128k, batch 1, GQA-8:
KV bytes = 2 × 32 × 8 × 128 × 2 × 131072 = 17.2 GB
At 3.35 TB/s: 5.1 ms JUST for attention
Plus 16 GB of weights: 4.8 ms
→ ~10 ms/token, half of it attentionContrast with S=4k where attention is 0.16 ms of a 5 ms step (3%).
FlashDecoding’s split-K parallelization is essential here: with one query and 128k keys, you must parallelize over keys to fill the GPU.
7. Performance implications#
Measured, Llama-3-8B on H100, batch 8:
context TTFT ITL max batch cost/M tokens (relative)
4k 0.15 s 5.2 ms 111 1.0x
16k 0.9 s 6.8 ms 27 2.6x
32k 3.1 s 9.5 ms 13 5.1x
128k 11.4 s 24.0 ms 3 22.0x22x cost per token at 128k vs 4k. Not 32x (linear in context) because the model weights are a fixed cost amortized differently, but far worse than free.
8. Production implications#
- Tier your context limits. Free tier 8k, pro 32k, enterprise 128k — and price accordingly.
- Separate pools for long-context requests. Mixing a 128k request into a pool of 4k requests destroys the pool’s efficiency: it consumes 32 slots’ worth of memory and blocks on prefill. Route long requests to dedicated capacity.
- Chunked prefill is mandatory if you accept long prompts and have ITL SLOs.
- Prefix caching is enormously valuable for long-context workloads — a 100k-token document re-prefilled per question is pure waste.
- Measure effective context for your task before promising it.
- Consider RAG as an alternative. Retrieving 4k of relevant context is often better and ~20x cheaper than stuffing 128k. The engineering question “long context or retrieval?” is frequently answered by this cost ratio.
9. Common mistakes#
Assuming context length is free capacity. It’s the denominator of your concurrency.
Setting max_model_len to the model’s maximum “just in case.” vLLM allocates KV blocks based
on it; a huge value reduces the blocks available per sequence and can prevent the server from
starting.
Mixing long and short requests in one pool.
Ignoring rope_theta when serving a long-context fine-tune.
Assuming advertised context = usable context.
Not chunking prefill. One 128k prompt freezes every other user for 11 seconds.
10. Hands-on exercise#
A. Build the scaling table. Run the script in section 4 for three models including one with sliding-window attention. Compare.
B. Measure the cost curve. On a real server, measure TTFT and ITL at context 1k, 4k, 16k, 64k (as far as your hardware allows). Plot both vs context on log axes. Fit the exponents.
C. Prefill quadratic. Measure prefill time vs prompt length from 128 to 32,768 tokens. Fit
a·S + b·S². At what S does the quadratic term exceed the linear one? Compare to your
prediction from Section I.08.
D. Sliding window. Compare KV memory and max concurrency for a full-attention model and a sliding-window model of similar size at 32k context.
E. RAG vs long context. For a document-QA task, compare: (i) stuffing a 50k-token document, (ii) retrieving the 3 most relevant 1k chunks. Measure cost per query and answer quality. Which wins?
11. Interview questions#
- What are the four costs of long context? Which scales quadratically?
- Why is 128k context ~20x more expensive per token than 4k, rather than 32x?
- What is sliding-window attention and how much does it save?
- Why does serving a long-context fine-tune require checking
rope_theta? - How would you architect a service that must support both 4k and 128k requests?
- When is RAG a better answer than long context? Give the cost argument.
- What is “effective context” and how would you measure it?
12. Further reading#
- [ESTABLISHED] Jiang et al., “Mistral 7B” (2023) — sliding window attention
- [ESTABLISHED] Peng et al., “YaRN” (2023) — context extension
- [ESTABLISHED] Liu et al., “Lost in the Middle” (2023) — effective context
- [EMERGING] Xiao et al., “StreamingLLM” (2023); Zhang et al., “H2O” (2023) — KV eviction
- [ESTABLISHED] DeepSeek-V2 technical report — MLA
- Next: 08 — Batching: static and dynamic