1. The four costs, restated#
Section V.07 established them; this file is about the techniques that attack each.
COST SCALING ATTACKED BY
KV memory O(S) GQA/MLA, KV quant, sliding window,
eviction, offload
KV read bandwidth per step O(S) same, plus sparse attention
Prefill compute O(S²) chunked prefill (doesn't reduce, but
makes it schedulable), sparse attention,
prefix caching
Concurrency O(1/S) everything above2. The techniques, honestly labelled#
[ESTABLISHED] — deploy these
GQA / MQA 4-8x KV reduction, no quality cost
MLA 10-60x KV reduction (DeepSeek architecture)
FP8 KV cache 2x, ~free
Sliding window attention caps KV at W (architectural)
Chunked prefill makes long prefill schedulable
Prefix caching eliminates redundant prefill
FlashAttention / FlashDecoding required for feasibility at all
[EMERGING] — evaluate carefully
INT4 KV cache 4x, measurable quality cost
Attention sinks (StreamingLLU) bounded memory for infinite streams
KV offload to CPU/NVMe capacity, not bandwidth
Disaggregated prefill/decode (file 06)
Cross-layer KV sharing 2x, limited production deployment
[RESEARCH] — interesting, not production
KV eviction (H2O, SnapKV) 3-5x, SILENT failure mode
Learned KV compression
Sparse attention patterns (beyond sliding window)
Linear attention / SSM hybrids O(1) state, different quality profileThe eviction row is the one to be careful about. It appears in many papers with impressive benchmark results and has a failure mode — silently losing information the model needed — that your monitoring cannot detect.
3. Context extension: how models get long context#
Models are trained at some length and extended. The methods matter to you because they change
rope_theta, and using the wrong value silently degrades output at long positions.
POSITION INTERPOLATION (PI)
Compress positions into the trained range: pos' = pos × (L_train/L_target)
✓ simple
✗ degrades short-context performance
NTK-AWARE SCALING
Scale rope_theta rather than positions: theta' = theta × s^(d/(d-2))
✓ preserves short-context performance better
→ what most models use
YARN
NTK scaling + attention temperature correction + selective
interpolation by frequency band
✓ better than NTK alone
✗ more parameters to get right
LONGROPE
Search for per-dimension rescaling factors
✓ best reported results
✗ requires the search// In config.json — READ THIS before deploying
{
"rope_theta": 500000.0, // not the base model's 10000!
"rope_scaling": {
"type": "yarn",
"factor": 8.0,
"original_max_position_embeddings": 8192
},
"max_position_embeddings": 131072
}
Verify your engine reads and applies rope_scaling. If it applies plain RoPE with the
original theta, output at position 50,000 will be degraded with no error.
4. Effective context vs advertised context#
A model advertising 128k may not USE 128k well.
MEASUREMENT: needle-in-a-haystack and its successors
place a fact at depth d in a context of length L
ask a question requiring it
measure retrieval accuracy as a function of (L, d)
TYPICAL FINDINGS
✓ accurate at the beginning and end of the context
✗ degraded in the middle ("lost in the middle", Liu et al. 2023)
✗ accuracy falls with L, often well before the advertised limit
✗ multi-fact retrieval degrades much faster than single-fact
✗ reasoning over long context degrades faster than retrievalBefore building a product on 128k context, measure whether the model uses it for your task. The cost is 20x (Section V.07); paying it for capability you don’t get is expensive.
A PRACTICAL TEST SUITE
1. single needle, varying depth and length (easiest)
2. multiple needles requiring aggregation
3. a question requiring reasoning over two distant facts
4. your actual task at varying context lengths
→ plot accuracy vs context length. Where does YOUR task break?5. RAG vs long context — the cost argument#
TASK: answer questions about a 200-page document (~150k tokens).
OPTION A — STUFF THE CONTEXT
prefill 150k tokens per question
prefill FLOPs ≈ 2×P×S + 2×L×S²×d
for an 8B model: 2.4e15 + 5.9e15 = 8.3 PFLOP
at 600 TFLOP/s: 14 seconds of TTFT
KV: 150k × 128 KiB = 19.2 GB per sequence
→ ~3 concurrent users on an H100
WITH PREFIX CACHING (same document, many questions):
first question: 14 s
subsequent: ~0.1 s (the document's KV is cached)
→ transforms the economics IF questions share the document
OPTION B — RAG
embed and index the document once (offline)
per question: retrieve 5 chunks × 800 tokens = 4k tokens
prefill 4k tokens: 0.06 s TTFT
KV: 0.5 GB per sequence → ~120 concurrent users
→ 40x cheaper per question
QUALITY
RAG wins when the answer is in a retrievable chunk.
Long context wins when the answer requires synthesis across the
whole document, or when retrieval fails.The right answer is usually: RAG for retrieval-shaped questions, long context with prefix caching for synthesis-shaped ones, and measure which your task is.
Note that prefix caching changes the calculation dramatically when many questions share one document — which is the common case. Long context with a warm prefix cache is competitive with RAG.
6. Sparse and windowed attention#
SLIDING WINDOW (Mistral) [ESTABLISHED]
each token attends to the last W tokens
→ KV capped at W regardless of S
→ information beyond W propagates through layers (L×W theoretical reach)
INTERLEAVED (Gemma 2, others) [ESTABLISHED]
most layers sliding-window, a few full attention
→ caps most of the KV while retaining true long-range access
→ e.g. 5 local layers per 1 global layer → ~83% KV reduction
ATTENTION SINKS (StreamingLLM) [EMERGING]
keep the first few tokens (which act as attention "sinks") plus a
recent window
→ enables unbounded streaming with bounded memory
→ the model doesn't GAIN long-range ability; it just doesn't break
→ some models now train with explicit sink tokens
STRUCTURED SPARSITY [RESEARCH]
strided, dilated, or learned sparse patterns
→ promising in papers; limited production deployment
→ the kernel efficiency of irregular patterns is the obstacleInterleaved local/global attention is the design most likely to become standard, because it captures most of the memory saving without giving up long-range capability.
7. KV eviction — why to be careful#
THE IDEA: not all cached tokens matter. Keep the important ones.
H2O: keep tokens with high cumulative attention scores
SnapKV: at the end of prefill, use the last queries' attention to
select which prompt tokens to keep
PyramidKV: allocate more KV budget to lower layers
REPORTED: 3-5x KV reduction with "minimal" quality loss.
THE PROBLEM
Quality loss is TASK-DEPENDENT and the failure mode is SILENT.
If the model needed an evicted token, it doesn't error — it
CONFABULATES. Your monitoring sees a normal response.
Benchmarks that show minimal loss typically test tasks where the
relevant information is recent or highly attended. Tasks requiring
a specific fact from the middle of a long document fail badly.
IF YOU DEPLOY IT
□ evaluate on YOUR task, specifically on cases requiring information
from evicted positions
□ needle-in-a-haystack at varying depths, WITH eviction enabled
□ have a fallback: detect low confidence and re-run without eviction
□ don't use it for tasks where a wrong answer is costlyThis is the technique in this curriculum most often presented as ready when it isn’t. It works; the question is whether it works for your task, and you must measure that specifically.
8. Production implications#
- Cap context by tier. Long context is genuinely expensive; price and limit it.
- Segregate long-context requests into their own pool (Section VIII.06). One 128k request poisons a pool tuned for 4k.
- Chunked prefill is mandatory if you accept long prompts and have an ITL SLO.
- Prefix caching is the highest-value long-context optimization when documents recur.
- Verify
rope_scalingis applied. - Measure effective context for your task before promising it.
- Evaluate RAG as an alternative. Often 40x cheaper.
- FP8 KV cache first, before considering eviction.
- Be conservative with eviction. Silent failures.
9. Common mistakes#
Deploying a long-context fine-tune without checking rope_scaling.
Promising 128k without measuring effective context.
Mixing long and short requests in one pool.
Not chunking prefill. An 11-second prefill freezes everyone.
Deploying eviction based on paper benchmarks.
Not evaluating RAG as an alternative.
Assuming cost scales linearly with context. It’s superlinear (Section V.07).
10. Hands-on exercise#
A. Measure effective context. Build a needle-in-a-haystack test: place a fact at depths 0%, 25%, 50%, 75%, 100% in contexts of 4k, 16k, 64k, 128k. Plot accuracy. Where does your model actually break?
B. Multi-fact. Extend A to require two facts at different depths. How much faster does accuracy degrade?
C. RAG comparison. For a document-QA task, implement both: full-context and RAG with retrieval. Measure cost per question and answer quality. Where’s the crossover?
D. Prefix caching effect. Measure the cost of 20 questions about one 100k-token document, with and without prefix caching. Quantify the difference.
E. Eviction risk. Implement a simple eviction scheme (keep first 4 + last N). Run your needle test with it enabled. At what eviction ratio does retrieval fail? Is the failure silent?
F. rope_scaling. Take a long-context model and deliberately serve it with the base
rope_theta. Measure output quality at position 1k vs 50k. How bad is it, and would you have
noticed?
11. Interview questions#
- What are the four costs of long context, and how does each scale?
- What is
rope_scalingand what happens if you ignore it? - What is “effective context” and how would you measure it?
- When is RAG better than long context? Give the cost argument.
- What is interleaved local/global attention and what does it save?
- Why are KV eviction methods risky in production?
- What would you deploy today for a 128k-context service?
12. Further reading#
- [ESTABLISHED] Peng et al., “YaRN” (2023); Su et al., “RoFormer” (RoPE, 2021)
- [ESTABLISHED] Liu et al., “Lost in the Middle” (2023)
- [ESTABLISHED] Jiang et al., “Mistral 7B” (2023) — sliding window
- [EMERGING] Xiao et al., “StreamingLLM” (2023)
- [RESEARCH] Zhang et al., “H2O” (2023); Li et al., “SnapKV” (2024)
- Next: 05 — Chunked prefill