Below the API

Attention Research

Expert Advanced 1h Difficulty 4/5

Prerequisites III.08, VII.10, XIII.01


1. The problem being attacked#

Attention costs:
  compute:  O(S²) for prefill
  memory:   O(S) KV cache, read every decode step
  
Everything in this area attacks one of those two.
THE SOLUTION SPACE

EXACT, IO-AWARE          same math, fewer memory trips
  → FlashAttention family. SOLVED and adopted.

ARCHITECTURAL            fewer KV heads, or a compressed KV
  → MQA, GQA, MLA. SOLVED and adopted.

SPARSE                   attend to fewer positions
  → sliding window (adopted), learned sparsity (not adopted)

LINEAR / RECURRENT       replace O(S²) with O(S) and O(1) state
  → the most active research area; not yet standard

APPROXIMATE             low-rank, kernel methods, hashing
  → largely abandoned for LLMs; quality cost too high

2. The exact/IO-aware line — settled#

FlashAttention (2022)     tiling + online softmax; O(S) memory
FlashAttention-2 (2023)   sequence-dim parallelism; better work partitioning
FlashDecoding (2023)      split-K over keys for decode
FlashAttention-3 (2024)   Hopper: TMA, warp specialization, FP8
FlashInfer                a library of attention kernel variants,
                          including paged, with a JIT for shape
                          specialization

STATUS: SOLVED. Use these. There is no reason to write attention yourself.

REMAINING WORK
  ~ Blackwell-specific optimizations
  ~ better support for unusual head dimensions and mask patterns
  ~ FP8/FP4 attention numerics

Section VII.10 covers what you need. This line of work is done from a practitioner’s perspective; the remaining research is kernel engineering for new hardware.


3. The architectural line — mostly settled#

MQA (2019) → GQA (2023) → MLA (2024)

STATUS: GQA is universal. MLA is DeepSeek's and is being evaluated
        elsewhere.

ACTIVE QUESTIONS
  ? cross-layer KV sharing (CLA, YOCO) — 2x on top of GQA, limited
    production validation
  ? how far can KV be compressed before quality suffers?
  ? can KV compression be learned per-model rather than architectural?

Cross-layer sharing is the one to watch here. If it holds up, it’s another 2x on the constraint that binds most deployments (Section V.06).


4. The sparse line — partially adopted#

ADOPTED
  ✓ sliding window (Mistral, and now common)
  ✓ interleaved local/global layers (Gemma 2, others)
  ✓ attention sinks (StreamingLLM's insight, now trained in)

NOT ADOPTED (despite many papers)
  ✗ strided / dilated patterns (Sparse Transformer, Longformer)
  ✗ learned / content-based sparsity (Reformer, Routing Transformer)
  ✗ block-sparse with learned block selection

WHY NOT
  1. KERNEL EFFICIENCY. Irregular sparsity doesn't map to tensor cores.
     A 4x FLOP reduction with a 4x less efficient kernel is a wash.
  2. QUALITY. Content-based sparsity means the model can't see
     something it needed, and the failure is silent.
  3. THE ALTERNATIVE. GQA + FlashAttention + sliding window already
     solved most of the practical problem.

The lesson generalizes: a FLOP reduction that doesn’t map to efficient hardware execution is not a speedup (Section I.07, Section X.03). Many sparse attention papers report FLOP reductions and modest wall-clock improvements for exactly this reason.

What’s newly active: dynamic sparse attention at inference time — selecting which KV blocks to attend to per query, using cheap approximations. This is closer to KV eviction (Section XIII.04) and carries the same silent-failure risk.


5. The linear/recurrent line — the interesting one#

THE IDEA: replace attention's O(S²) and O(S) state with something
          O(S) and O(1).

LINEAR ATTENTION
  softmax(QK^T)V ≈ φ(Q)(φ(K)^T V)     with a feature map φ
  → compute the (d × d) matrix φ(K)^T V once, reuse it
  → O(1) state per layer, O(S) total compute
  ✗ quality gap vs softmax attention has been persistent

STATE SPACE MODELS (Mamba, S4, S6)
  a recurrent state updated per token, with selective gating
  → O(1) state, O(S) compute
  ✓ competitive quality at moderate scale
  ✓ constant memory regardless of context length
  ✗ weaker at exact retrieval from long context ("what was the
    phone number on page 40?")
  ✗ hardware efficiency: the recurrence is inherently sequential
    (though parallel scan algorithms exist)

HYBRIDS (Jamba, Zamba, Samba, and others)
  mostly SSM layers + a few attention layers
  ✓ SSM handles the bulk; attention layers provide exact retrieval
  ✓ KV cache only for the attention layers → much smaller
  → THE MOST PROMISING DIRECTION
WHY THIS MATTERS FOR INFERENCE

  A pure attention model at 1M context:
    KV cache: 130+ GB per sequence (Section V.07)
    
  An SSM at 1M context:
    state: a few MB, CONSTANT
    
  A hybrid with 1 attention layer per 8:
    KV cache: 1/8 of pure attention, plus a constant state
    
  → this is a 10-100x change in the binding constraint.

If hybrids reach parity with pure-attention models at scale, it changes long-context serving completely. That’s why this is the area to watch.

Honest status: hybrids are shipping in real models (Jamba, and several others) and perform well, but the largest and strongest models remain pure attention. The question is whether that’s a fundamental gap or a matter of investment.


6. What to watch, ranked#

1. HYBRID SSM/ATTENTION AT SCALE
   → if a frontier-quality model ships as a hybrid, long-context
     serving economics change by an order of magnitude
   → watch: model releases, not papers

2. CROSS-LAYER KV SHARING
   → 2x on the binding constraint, architectural, low risk
   → watch: whether major model families adopt it

3. DYNAMIC SPARSE ATTENTION AT INFERENCE
   → potentially large, but shares KV eviction's silent-failure risk
   → watch: whether anyone demonstrates a bounded failure mode

4. FP4/FP8 ATTENTION NUMERICS
   → attention in low precision is harder than GEMM (the softmax)
   → watch: FlashAttention-3+ and Blackwell kernels

5. NATIVE LONG-CONTEXT ARCHITECTURES
   → models designed for 1M+ from the start, rather than extended
   → watch: whether effective context catches up to advertised

7. What is unlikely to matter#

Stated plainly, because reading time is finite:

✗ New approximate attention mechanisms (low-rank, kernel, hashing)
    20+ papers, none adopted. The quality cost is real and the
    kernel efficiency is poor.

✗ FLOP-reduction papers without wall-clock measurements
    See section 4.

✗ Attention variants requiring training from scratch, unless a
    frontier lab adopts them
    The cost of retraining a frontier model is prohibitive for
    an uncertain gain.

✗ Papers evaluated only at small scale (< 1B parameters)
    Attention behavior changes with scale; small-scale results
    frequently don't transfer.

8. Hands-on exercise#

A. Apply the seven questions. Take three recent attention papers. Apply Section XIV.01’s framework. Which survive?

B. The FLOP/wall-clock gap. For a sparse attention paper, compute the claimed FLOP reduction and find the reported wall-clock improvement. What’s the ratio? Why?

C. Measure the hybrid benefit. For a hybrid model (Jamba or similar) and a pure-attention model of similar size, compute KV bytes per token at 128k context. What’s the ratio?

D. Test effective retrieval. Run a needle-in-a-haystack test (Section XIII.04) on a hybrid model and a pure-attention model. Does the hybrid’s weaker exact-retrieval show up?

E. Read FlashAttention-3. Identify the three Hopper-specific mechanisms it uses (Section VI.03). Which would not exist on Ampere?


9. Interview questions#

  1. Which attention research lines are settled and which are active?
  2. Why haven’t sparse attention methods been adopted despite many papers?
  3. What are SSMs and why do they matter for inference?
  4. Why are hybrid SSM/attention models the most promising direction?
  5. What would change about long-context serving if hybrids reached parity?
  6. Why is attention harder to do in low precision than GEMM?

10. Further reading#

  • [ESTABLISHED] Dao et al., FlashAttention 1/2/3
  • [ESTABLISHED] Ainslie et al., GQA; DeepSeek-V2 (MLA)
  • [EMERGING] Gu & Dao, “Mamba” (2023); Dao & Gu, “Transformers are SSMs” (2024)
  • [EMERGING] Lieber et al., “Jamba” (2024) — a production hybrid
  • [EMERGING] Brandon et al., “Cross-Layer Attention” (2024)
  • Next: 03 — KV cache research

↑↓ navigate ↵ open