PidokuInfra

Where This Is Going

Expert Advanced 35 min Difficulty 2/5 Topic 04 of 04

Prerequisites all previous modules

The idea in one minute#

Specific products will be obsolete in a year. The pressures that shape them will not. Six trends follow directly from physics and economics you have already learned: memory is the bottleneck, precision keeps falling, the system is the product, power is the limit, inference is specializing, and software sets the pace. If you understand why each one exists, new announcements become predictable.

A picture#

flowchart TB
  P1["Arithmetic grows faster<br/>than memory speed"] --> T1["Memory-centric design<br/>HBM4, stacking, in-rack memory"]
  P2["Networks tolerate<br/>low precision"] --> T2["Smaller formats<br/>FP8, FP4, mixed"]
  P3["Models outgrow<br/>one chip"] --> T3["Rack-scale systems<br/>NVLink domains, optics"]
  P4["Watts per chip<br/>keep rising"] --> T4["Liquid cooling,<br/>power as the constraint"]
  P5["Inference outgrows<br/>training in volume"] --> T5["Inference-specific<br/>hardware and tiers"]
  P6["Hardware is useless<br/>without kernels"] --> T6["Compilers and<br/>software ecosystems decide"]
  class P1,P2,P3,P4,P5,P6 neutral
  class T1 memory
  class T2,T6 compute
  class T3,T5 io
  class T4 warn

How it really works#

1. Memory is the bottleneck#

The roofline (IV.01) says LLM inference at small batch is memory-bound, and II.03 showed that bandwidth has historically grown more slowly than arithmetic. Expect continued large jumps in HBM bandwidth and capacity; tiered memory, where cold data such as inactive KV cache lives in cheaper, slower memory and is paged in; and research into computing closer to, or inside, the memory itself.

2. Precision keeps falling#

Each generation’s tensor cores support a smaller format (FP16 → FP8 → FP4). The limit is accuracy, not hardware, and progress comes from better quantization methods and from training models to tolerate low precision from the start. Expect per-layer and per-tensor mixed precision to become normal, chosen automatically.

3. The system is the product#

Communication between chips is the tax on every large model (VI.02). So the trend is toward bigger tightly-coupled units: racks that behave as one computer, co-packaged optics replacing copper for longer high-speed links, and CPU, GPU, network processor and switches designed together. Comparing chips in isolation will become less and less meaningful.

4. Power is the limit#

Performance per watt improves each generation; total watts rise anyway. Data-center location, grid access and cooling design are now strategic decisions. Efficiency — tokens per joule — is increasingly the metric that matters, and it rewards exactly the software techniques this course and its sister course teach.

5. Inference is specializing#

Training and inference want different things (V.04), and inference is now the larger market. Expect more hardware tuned only for serving, and disaggregation within serving itself — using different hardware for the compute-bound phase (reading the prompt) and the memory-bound phase (generating tokens).

6. Software sets the pace#

A chip is only as fast as its worst kernel. Compilers that generate good kernels automatically for new hardware (MLIR-based toolchains, Triton, TVM-style systems) are what make alternative accelerators viable, and the main reason an incumbent stays ahead. When evaluating anything new, weigh the software at least as heavily as the silicon.

TrendWhat happened
Memory is the bottleneckHBM4 shipped: Rubin moved from ~8 to ~22 TB/s per GPU. NVIDIA added a shared, network-attached tier for KV cache to its rack design
Precision keeps fallingFP4 is the headline format of every 2026 flagship
The system is the productRubin is sold only as a liquid-cooled rack with NVIDIA’s own CPUs, switches and network cards; AMD shipped its first rack-scale system
Power is the limitRacks went from ~120 kW to roughly 200 kW, with ~600 kW announced; vendors now quote tokens per megawatt
Inference is specializingNVIDIA licensed Groq’s inference technology and sells a non-GPU inference rack; cloud vendors are splitting training and inference chips
Software sets the paceEvery newcomer is judged first on which serving engines run on it

None of this was a surprise to anyone holding the six trends. That is the use of them.

To watch these systems in operation — power, throttling, tokens per joule — see Observability Engineering, module V.

What does not change#

  • A processor is limited by arithmetic or by memory traffic. Find out which.
  • Moving data costs more than computing on it. Keep data where the work is.
  • Uniform, independent work parallelizes; branchy, dependent work does not.
  • Fixed costs per operation are amortized by batching.
  • Measure before optimizing, and compare against a predicted ceiling.

Those five sentences were true of the first programmable GPU and will be true of whatever replaces the current generation.

Reading a new announcement#

A checklist to apply to any launch:

  1. Capacity — GB per device. Does it change what fits?
  2. Bandwidth — TB/s. This is your memory-bound speedup.
  3. Arithmetic — in which format? Can you use that format?
  4. Interconnect — how many devices form one fast domain?
  5. Power and cooling — can your facility host it?
  6. Software — what runs on day one?
  7. Price and availability — cost per unit of your work, and when you can actually get it.

Remember this#

  • Trends follow from constraints: memory, precision, communication, power, workload mix, software.
  • Evaluate new hardware with the same seven questions every time.
  • The principles — regime, locality, uniformity, amortization, measurement — outlast every product.

Try it#

  1. Take the most recent accelerator announcement you can find and fill in the seven-point checklist. What is missing from the announcement?
  2. Pick one trend above and argue the opposite case: what would have to be true for it to stop?
  3. Write, in five sentences of your own, what a GPU is and what limits it. Compare with what you would have written before module I.

Check yourself#

  1. Why does the industry keep investing in memory bandwidth?
  2. What limits how far number formats can shrink?
  3. Which of the five durable principles do you find yourself using most?

You have finished the course. The natural next step is Inference Engineering, which builds the serving systems that sit on top of everything you now understand; then Observability Engineering, which teaches you to watch them — or the projects, if you have not built them yet.

↑↓ navigate↵ openesc close