Below the API

FLOPs, Bandwidth, and Arithmetic Intensity

Foundations Intermediate 1h 30m Difficulty 3/5

Prerequisites 06, 07

This file is the quantitative heart of Section I. Everything here is arithmetic you should be able to do on a whiteboard. Do not skim it.


1. What is it?#

Three numbers that let you predict inference performance before writing any code.

FLOPs      how much arithmetic a computation requires        [operations]
Bytes      how much data it must move                        [bytes]
Intensity  I = FLOPs / Bytes                                 [FLOP/byte]

With those three plus two hardware constants (peak FLOP/s, peak bytes/s), you can estimate the runtime of any kernel to within a factor of ~2. That estimate is often more useful than a measurement, because it tells you what is possible.

Diagram — Which wall will a kernel hit?#

flowchart LR
  K["Kernel"] --> AI["Arithmetic intensity<br/>= FLOPs / bytes moved"]
  AI --> C{"AI vs ridge point<br/>peak FLOPs / bandwidth"}
  C -->|"below"| M["MEMORY-BOUND<br/>speed = bandwidth x AI"]
  C -->|"above"| P["COMPUTE-BOUND<br/>speed = peak FLOPs"]
  M --> M2["Fix: move fewer bytes<br/>batch, quantize, fuse"]
  P --> P2["Fix: do less math<br/>lower precision, better kernel"]

  class M,M2 memory
  class P,P2 compute
  class K,AI neutral
  class C queue

2. Why does it exist?#

Because “let’s try it and see” does not scale to a design space with dozens of dimensions (model size, precision, batch size, context length, parallelism degree, hardware SKU). You need a model that lets you reject 90% of the options on paper.

This is what distinguishes an inference engineer from someone who tunes flags. The engineer can say, before provisioning anything: “70B at FP8 on 4×H100 with TP=4 will give you about 90 tokens/sec per user at batch 1, and about 2,500 tokens/sec aggregate at batch 64, and you will run out of KV memory at 190 concurrent users at 8k context.” And then be right within 30%.


3. Simple analogy#

Estimating a road trip.

You do not drive the route to find out how long it takes. You compute: 600 miles ÷ 60 mph = 10 hours, plus stops. You are within 20% instantly.

FLOPs is the distance. Peak FLOP/s is the speed limit. Bandwidth is the fuel-tank size that forces stops. Arithmetic intensity is your miles-per-gallon: high intensity means fewer stops.

And just as with driving, the binding constraint changes with conditions: on an empty motorway the speed limit binds; in city traffic the stops bind. Knowing which one binds tells you whether to buy a faster car or a bigger tank.


4. Tiny example#

Count FLOPs for a matrix multiply by hand.

C = A @ B      where A is (M, K) and B is (K, N)

Each element of C is a dot product of length K:
    K multiplies + (K-1) adds ≈ 2K FLOPs

C has M×N elements.

Total FLOPs = 2 · M · N · K

Memorize 2MNK. It is the single most-used formula in this field.

Check with a tiny case:

A = [1 2]     B = [5 6]
    [3 4]         [7 8]

M=N=K=2  →  2·2·2·2 = 16 FLOPs

C[0][0] = 1·5 + 2·7 = 19       (2 mults, 1 add = 3... )

The formula says 16, exact counting says 4 elements × 3 ops = 12. The 2MNK convention counts a multiply-add as 2 FLOPs and ignores the off-by-one, which is standard and correct at scale (K=4096 makes the difference 0.02%).

Now bytes, in FP16:

Read A:   2·M·K
Read B:   2·K·N
Write C:  2·M·N
Total:    2(MK + KN + MN)

Intensity:

I = 2MNK / [2(MK + KN + MN)] = MNK / (MK + KN + MN)

Square case M=N=K=n:   I = n³/(3n²) = n/3

So a 4096³ matmul has intensity ~1365 FLOP/byte — deeply compute-bound. A 1×4096×4096 GEMV (M=1) has:

I = 1·4096·4096 / (4096 + 4096·4096 + 4096) ≈ 1.0

Intensity 1.0. Memory-bound by a factor of ~300 on an H100. That single calculation is why batch-1 decode is slow, and you just derived it from first principles.


5. Technical explanation#

FLOPs for a transformer#

Per token, forward pass, dense decoder-only transformer:

Per layer:
  QKV projection:      2 · d · (d + 2·d_kv)      where d_kv = n_kv_heads·head_dim
  Attention scores:    2 · S · d_head · n_heads   = 2·S·d
  Attention @ V:       2 · S · d_head · n_heads   = 2·S·d
  Output projection:   2 · d · d
  FFN (SwiGLU, 3 mats):2 · 3 · d · d_ff

Plus the final LM head: 2 · d · V   (once, not per layer)

For MHA with d_ff = 4d this collapses to the famous approximation:

FLOPs per token ≈ 2 · P + 4 · L · S · d
                  └────┘   └──────────┘
                  weights   attention (grows with context!)

The first term is “2 × parameters” — every weight does one multiply and one add. The second term is the attention over the KV cache, which grows linearly with sequence position.

When does the attention term matter? Set them equal:

2P = 4·L·S·d   →   S = P / (2·L·d)

Llama-3 8B:  P=8e9, L=32, d=4096  →  S = 8e9/(2·32·4096) ≈ 30,500 tokens
Llama-3 70B: P=7e10, L=80, d=8192 →  S = 7e10/(2·80·8192) ≈ 53,400 tokens

So below ~30k context, attention FLOPs are a minority of the total for an 8B model. Above it, attention dominates. This is why long-context inference is a qualitatively different problem and gets its own treatment in Section XIII.

For prefill of a prompt of length S, multiply per-token cost by S, but the attention term becomes quadratic:

prefill FLOPs ≈ 2·P·S + 2·L·S²·d

(The factor differs slightly from the decode form because prefill computes a full S×S attention matrix, half of it masked away.)

Bytes for a transformer#

Decode step, batch B, context S:

Weights read:    P · bytes_per_weight        ← once per step, regardless of B
KV cache read:   2 · L · n_kv · head_dim · bytes · S · B
Activations:     small, O(B·d·L)

Note the crucial asymmetry:

  • Weight bytes are independent of B. (This is why batching works.)
  • KV bytes are proportional to B×S. (This is why long context at high batch eventually becomes the bottleneck instead.)

The crossover, for Llama-3-70B FP16 (KV 320 KB/token):

Weight bytes = 140 GB
KV bytes     = 320 KB × S × B

Equal when S·B = 140e9 / 320e3 = 437,500 token-slots

e.g. B=64, S=6,800   or   B=16, S=27,000

Below that, weights dominate and quantizing weights is your best lever. Above it, KV dominates and GQA/MLA/KV-quantization are your best levers. Knowing which side of that line you are on tells you which optimization to invest in. Compute this for your workload.

Putting it together: predicting decode speed#

T_step = max( FLOPs_step / peak_FLOPs ,  Bytes_step / peak_BW )

tokens_per_sec_total   = B / T_step
tokens_per_sec_per_user = 1 / T_step

Worked example: Llama-3-8B, FP16, one H100, batch 32, context 2048.

FLOPs  = B · (2P + 4·L·S·d)
       = 32 · (2·8e9 + 4·32·2048·4096)
       = 32 · (16e9 + 1.07e9) = 32 · 17.1e9 = 547 GFLOPs

Bytes  = P·2 + 2·L·n_kv·head_dim·2·S·B
       = 16e9 + 2·32·8·128·2·2048·32
       = 16e9 + 4.3e9 = 20.3 GB

I = 547e9 / 20.3e9 = 27 FLOP/byte     → well below H100's ridge (296): MEMORY BOUND

T_step = 20.3e9 / 3.35e12 = 6.06 ms

Throughput = 32 / 0.00606 = 5,280 tokens/sec
Per-user   = 1 / 0.00606  = 165 tokens/sec       (ITL = 6.06 ms)

Real systems achieve perhaps 60-75% of this because of kernel inefficiency, launch overhead, sampling, and scheduling. So predict ~3,500-4,000 tok/s. That is a genuinely useful estimate made in five lines of arithmetic. Go do it for your own model and hardware, then measure and compare. Learning how far off the estimate is for your stack is what turns the formula into engineering judgment.


6. Under the hood#

Why real kernels miss peak:

Loss sourceTypical cost
Tile quantization (M, N not multiples of tile size)5-30%
Wave quantization (grid not a multiple of SM count)5-20%
Kernel launch overhead (decode: many tiny kernels)10-40% at batch 1
Non-GEMM ops (norms, activations, sampling)10-25%
Memory access patterns (uncoalesced, strided KV)10-50%
Power/thermal throttling0-20%
Tensor core idle during epilogue/prologue5-15%

Multiply these and 50-70% of theoretical is a good result. If you are at 20%, something is structurally wrong (usually launch overhead or an unfused chain of elementwise ops).

Precision changes both sides of the equation:

FP16 → FP8:   bytes halve  AND  peak FLOP/s doubles (Hopper tensor cores)
              → up to 2x on memory-bound, up to 2x on compute-bound
INT4 weights: bytes quarter, but compute stays FP16 (dequant on the fly)
              → up to 4x on memory-bound, ~0 on compute-bound

This is exactly why INT4 weight-only quantization helps decode enormously and prefill barely at all. Decode is memory-bound (you shrink the binding constraint); prefill is compute-bound (you don’t). A question that catches people out in interviews, and a real deployment decision.


7. Performance implications#

Build the table for your own hardware and model — it becomes your design reference:

Llama-3-70B on 8×H100 (TP=8), FP16, context 4096:

Per GPU: weights 17.5 GB, bandwidth 3.35 TB/s

B=1:   bytes=17.5GB + tiny KV  → T=5.3ms  → 189 tok/s/user, 189 total
B=8:   bytes=17.5GB + 0.13GB   → T=5.3ms  → 189 tok/s/user, 1,510 total
B=32:  bytes=17.5GB + 0.52GB   → T=5.4ms  → 185 tok/s/user, 5,920 total
B=128: bytes=17.5GB + 2.1GB    → T=5.8ms  → 172 tok/s/user, 22,000 total
B=512: bytes=17.5GB + 8.4GB    → T=7.7ms  → 130 tok/s/user, 66,500 total

(Ignoring AllReduce cost, which adds ~0.5-1 ms per layer group — see Section IX.)

Look at what this says: going from batch 1 to batch 128 costs each user 9% of their token rate and gains the system 116x throughput. That is the most favorable engineering tradeoff you will ever be offered, and it is why continuous batching is non-negotiable in production.


8. Production implications#

  • Do the estimate before the procurement. “Will 4 A100s serve this?” is answerable on paper in ten minutes. Doing it after you have bought the hardware is expensive.
  • Publish your model’s FLOPs and bytes per token internally. It becomes the shared language for capacity conversations with non-specialists.
  • Track achieved-vs-theoretical as a health metric. If your system historically achieves 65% of the roofline prediction and drops to 40%, something regressed — a kernel fell back, a fusion broke, a shape changed.
  • Use the crossover formula to decide your next optimization. Weight-bytes dominant → quantize weights. KV-bytes dominant → GQA, KV quantization, shorter effective context, paged/offloaded KV.

9. Common mistakes#

Using peak FLOP/s numbers that include sparsity. NVIDIA quotes both dense and 2:4-sparse numbers; the sparse ones are double and rarely achievable. Always check which you cited.

Forgetting that FLOPs “per token” for prefill means per prompt token. A 2,000-token prompt costs 2,000× the per-token decode FLOPs (plus quadratic attention), all at once. That is why TTFT is dominated by prefill.

Counting only weight bytes. At long context and high batch, KV reads can be the majority.

Comparing FLOPs across precisions naively. “FP8 gives 2x FLOPs” is a hardware statement about tensor-core throughput; it does not mean your memory-bound kernel gets 2x.

Ignoring that AllReduce consumes bandwidth too. In TP, communication competes with weight reads for the same memory system on some paths, and adds latency on the interconnect.

Assuming intensity is a property of the algorithm alone. It depends on the implementation. Fusing three elementwise ops triples the intensity of that chain without changing the math.


10. Hands-on exercise#

A. Derive and verify 2MNK. Write a script that, for various M/N/K, computes theoretical FLOPs, measures the actual time, and reports achieved TFLOP/s. Sweep M from 1 to 4096 with N=K=4096. Plot achieved TFLOP/s vs M. Mark the point where you cross 50% of peak.

B. Full model estimate. Pick a model you can run. Compute by hand:

  1. FLOPs per decode token at context 1k and 32k.
  2. Bytes per decode step at batch 1, 32, 256 for both contexts.
  3. Predicted ITL and total throughput for each of the six combinations.
  4. Then measure all six. Build a table of predicted vs actual and a column for the ratio.
  5. Explain the biggest discrepancy.

C. Find your crossover. For your model, compute the (B, S) curve along which weight bytes equal KV bytes. Plot it. Mark where your production traffic sits.

D. Precision study. Repeat B for FP16, FP8, and INT4 weights (predicted only). Which phase benefits most from each? Does your prediction match the community’s reported speedups?


11. Interview questions#

  1. Write the FLOP count for an (M,K)@(K,N) matmul and derive its arithmetic intensity.
  2. Give the approximate FLOPs per token for a transformer and explain both terms.
  3. At what context length do attention FLOPs exceed weight FLOPs for a 70B model?
  4. At what (batch, context) do KV cache bytes exceed weight bytes for a 70B GQA model?
  5. Why does INT4 weight quantization speed up decode ~4x but prefill almost not at all?
  6. Estimate tokens/sec for a 13B FP16 model at batch 16 on an A100. State your assumptions.
  7. Your measured throughput is 35% of the roofline prediction. Name five plausible causes and how you would distinguish them.

12. Further reading#

  • [FUNDAMENTAL] Williams et al., “Roofline” (CACM 2009)
  • [ESTABLISHED] Pope et al., “Efficiently Scaling Transformer Inference” (2022) — the best worked analytical treatment in the literature
  • [ESTABLISHED] Kaplan et al. (2020), Appendix — transformer FLOP accounting
  • [REFERENCE] NVIDIA architecture whitepapers for peak numbers (use dense, not sparse)
  • Next: 09 — CPU vs GPU

↑↓ navigate ↵ open