1. The problem being attacked#
Attention costs:
compute: O(S²) for prefill
memory: O(S) KV cache, read every decode step
Everything in this area attacks one of those two.THE SOLUTION SPACE
EXACT, IO-AWARE same math, fewer memory trips
→ FlashAttention family. SOLVED and adopted.
ARCHITECTURAL fewer KV heads, or a compressed KV
→ MQA, GQA, MLA. SOLVED and adopted.
SPARSE attend to fewer positions
→ sliding window (adopted), learned sparsity (not adopted)
LINEAR / RECURRENT replace O(S²) with O(S) and O(1) state
→ the most active research area; not yet standard
APPROXIMATE low-rank, kernel methods, hashing
→ largely abandoned for LLMs; quality cost too high2. The exact/IO-aware line — settled#
FlashAttention (2022) tiling + online softmax; O(S) memory
FlashAttention-2 (2023) sequence-dim parallelism; better work partitioning
FlashDecoding (2023) split-K over keys for decode
FlashAttention-3 (2024) Hopper: TMA, warp specialization, FP8
FlashInfer a library of attention kernel variants,
including paged, with a JIT for shape
specialization
STATUS: SOLVED. Use these. There is no reason to write attention yourself.
REMAINING WORK
~ Blackwell-specific optimizations
~ better support for unusual head dimensions and mask patterns
~ FP8/FP4 attention numericsSection VII.10 covers what you need. This line of work is done from a practitioner’s perspective; the remaining research is kernel engineering for new hardware.
3. The architectural line — mostly settled#
MQA (2019) → GQA (2023) → MLA (2024)
STATUS: GQA is universal. MLA is DeepSeek's and is being evaluated
elsewhere.
ACTIVE QUESTIONS
? cross-layer KV sharing (CLA, YOCO) — 2x on top of GQA, limited
production validation
? how far can KV be compressed before quality suffers?
? can KV compression be learned per-model rather than architectural?Cross-layer sharing is the one to watch here. If it holds up, it’s another 2x on the constraint that binds most deployments (Section V.06).
4. The sparse line — partially adopted#
ADOPTED
✓ sliding window (Mistral, and now common)
✓ interleaved local/global layers (Gemma 2, others)
✓ attention sinks (StreamingLLM's insight, now trained in)
NOT ADOPTED (despite many papers)
✗ strided / dilated patterns (Sparse Transformer, Longformer)
✗ learned / content-based sparsity (Reformer, Routing Transformer)
✗ block-sparse with learned block selection
WHY NOT
1. KERNEL EFFICIENCY. Irregular sparsity doesn't map to tensor cores.
A 4x FLOP reduction with a 4x less efficient kernel is a wash.
2. QUALITY. Content-based sparsity means the model can't see
something it needed, and the failure is silent.
3. THE ALTERNATIVE. GQA + FlashAttention + sliding window already
solved most of the practical problem.The lesson generalizes: a FLOP reduction that doesn’t map to efficient hardware execution is not a speedup (Section I.07, Section X.03). Many sparse attention papers report FLOP reductions and modest wall-clock improvements for exactly this reason.
What’s newly active: dynamic sparse attention at inference time — selecting which KV blocks to attend to per query, using cheap approximations. This is closer to KV eviction (Section XIII.04) and carries the same silent-failure risk.
5. The linear/recurrent line — the interesting one#
THE IDEA: replace attention's O(S²) and O(S) state with something
O(S) and O(1).
LINEAR ATTENTION
softmax(QK^T)V ≈ φ(Q)(φ(K)^T V) with a feature map φ
→ compute the (d × d) matrix φ(K)^T V once, reuse it
→ O(1) state per layer, O(S) total compute
✗ quality gap vs softmax attention has been persistent
STATE SPACE MODELS (Mamba, S4, S6)
a recurrent state updated per token, with selective gating
→ O(1) state, O(S) compute
✓ competitive quality at moderate scale
✓ constant memory regardless of context length
✗ weaker at exact retrieval from long context ("what was the
phone number on page 40?")
✗ hardware efficiency: the recurrence is inherently sequential
(though parallel scan algorithms exist)
HYBRIDS (Jamba, Zamba, Samba, and others)
mostly SSM layers + a few attention layers
✓ SSM handles the bulk; attention layers provide exact retrieval
✓ KV cache only for the attention layers → much smaller
→ THE MOST PROMISING DIRECTIONWHY THIS MATTERS FOR INFERENCE
A pure attention model at 1M context:
KV cache: 130+ GB per sequence (Section V.07)
An SSM at 1M context:
state: a few MB, CONSTANT
A hybrid with 1 attention layer per 8:
KV cache: 1/8 of pure attention, plus a constant state
→ this is a 10-100x change in the binding constraint.If hybrids reach parity with pure-attention models at scale, it changes long-context serving completely. That’s why this is the area to watch.
Honest status: hybrids are shipping in real models (Jamba, and several others) and perform well, but the largest and strongest models remain pure attention. The question is whether that’s a fundamental gap or a matter of investment.
6. What to watch, ranked#
1. HYBRID SSM/ATTENTION AT SCALE
→ if a frontier-quality model ships as a hybrid, long-context
serving economics change by an order of magnitude
→ watch: model releases, not papers
2. CROSS-LAYER KV SHARING
→ 2x on the binding constraint, architectural, low risk
→ watch: whether major model families adopt it
3. DYNAMIC SPARSE ATTENTION AT INFERENCE
→ potentially large, but shares KV eviction's silent-failure risk
→ watch: whether anyone demonstrates a bounded failure mode
4. FP4/FP8 ATTENTION NUMERICS
→ attention in low precision is harder than GEMM (the softmax)
→ watch: FlashAttention-3+ and Blackwell kernels
5. NATIVE LONG-CONTEXT ARCHITECTURES
→ models designed for 1M+ from the start, rather than extended
→ watch: whether effective context catches up to advertised7. What is unlikely to matter#
Stated plainly, because reading time is finite:
✗ New approximate attention mechanisms (low-rank, kernel, hashing)
20+ papers, none adopted. The quality cost is real and the
kernel efficiency is poor.
✗ FLOP-reduction papers without wall-clock measurements
See section 4.
✗ Attention variants requiring training from scratch, unless a
frontier lab adopts them
The cost of retraining a frontier model is prohibitive for
an uncertain gain.
✗ Papers evaluated only at small scale (< 1B parameters)
Attention behavior changes with scale; small-scale results
frequently don't transfer.8. Hands-on exercise#
A. Apply the seven questions. Take three recent attention papers. Apply Section XIV.01’s framework. Which survive?
B. The FLOP/wall-clock gap. For a sparse attention paper, compute the claimed FLOP reduction and find the reported wall-clock improvement. What’s the ratio? Why?
C. Measure the hybrid benefit. For a hybrid model (Jamba or similar) and a pure-attention model of similar size, compute KV bytes per token at 128k context. What’s the ratio?
D. Test effective retrieval. Run a needle-in-a-haystack test (Section XIII.04) on a hybrid model and a pure-attention model. Does the hybrid’s weaker exact-retrieval show up?
E. Read FlashAttention-3. Identify the three Hopper-specific mechanisms it uses (Section VI.03). Which would not exist on Ampere?
9. Interview questions#
- Which attention research lines are settled and which are active?
- Why haven’t sparse attention methods been adopted despite many papers?
- What are SSMs and why do they matter for inference?
- Why are hybrid SSM/attention models the most promising direction?
- What would change about long-context serving if hybrids reached parity?
- Why is attention harder to do in low precision than GEMM?
10. Further reading#
- [ESTABLISHED] Dao et al., FlashAttention 1/2/3
- [ESTABLISHED] Ainslie et al., GQA; DeepSeek-V2 (MLA)
- [EMERGING] Gu & Dao, “Mamba” (2023); Dao & Gu, “Transformers are SSMs” (2024)
- [EMERGING] Lieber et al., “Jamba” (2024) — a production hybrid
- [EMERGING] Brandon et al., “Cross-Layer Attention” (2024)
- Next: 03 — KV cache research