Below the API

FP8 and INT8 Inference

Intermediate Advanced 1h 15m Difficulty 4/5

Prerequisites III.11, 04


1. Problem → Why → Optimization#

PROBLEM   BF16 inference uses 2 bytes per value and runs tensor cores at
          half their potential rate on Hopper+.
WHY       The hardware supports 8-bit tensor core operations at 2x the
          throughput, and 8-bit values halve memory traffic.
OPTIMIZE  Serve in FP8 (E4M3) or INT8, quantizing both weights and activations.

This is the production default for modern inference on Hopper and Blackwell. It is [ESTABLISHED], not experimental.


2. FP8 vs INT8 — the comparison#

                    FP8 E4M3                INT8
Representation      1 sign, 4 exp, 3 mant   1 sign, 7 magnitude
Max value           448                     127
Dynamic range       ~2^-9 to 448 (~10⁵)     -127 to 127 (~10²)
Precision           ~6% relative, uniform   absolute, uniform
                    across magnitudes       (relative error bad for small values)
Zero point          not needed              sometimes needed (asymmetric)
Outlier tolerance   GOOD (exponent absorbs) POOR (needs SmoothQuant)
Hardware            Hopper+                 Turing+
Throughput          2x BF16                 2x BF16
Typical quality Δ   -0.1 to -0.5%           -0.2 to -1.0%
Calibration         simple (per-tensor amax) needs care (SmoothQuant)

The key difference: FP8 has an exponent. That means its relative precision is roughly constant across magnitudes, so a channel with values around 60 and a channel with values around 0.5 are both represented with ~6% relative error. INT8’s uniform absolute spacing means the 0.5-magnitude channel is destroyed if the tensor’s max is 60.

This is why FP8 usually doesn’t need SmoothQuant and INT8 does. It is the single most practical thing to know about the two formats.


3. Simple analogy#

Measuring with a ruler versus with scientific notation.

INT8 is a ruler with 254 evenly spaced marks. If your longest object is 60 metres, each mark is 24 cm — and you cannot measure a 5 cm object at all.

FP8 is scientific notation with 3 significant figures. It measures 60 metres as 60.0 m and 5 cm as 5.00 cm, both to 3 figures. The relative precision is the same at every scale.

For data with a wide dynamic range — which activations have — scientific notation wins.


4. Tiny example#

package main

import (
	"fmt"
	"math"
)

// int8PerTensor: one scale for the whole tensor, 255 evenly spaced levels.
func int8PerTensor(x []float64) []float64 {
	var absmax float64
	for _, v := range x {
		absmax = math.Max(absmax, math.Abs(v))
	}
	s := absmax / 127
	out := make([]float64, len(x))
	for i, v := range x {
		out[i] = math.Round(v/s) * s
	}
	return out
}

// fp8E4M3: 1 sign bit, 4 exponent bits, 3 mantissa bits. Every power of two gets
// 8 evenly spaced values, so the RELATIVE error is about the same at every magnitude.
func fp8E4M3(v float64) float64 {
	if v == 0 {
		return 0
	}
	e := math.Max(math.Floor(math.Log2(math.Abs(v))), -6) // smallest normal exponent
	step := math.Pow(2, e-3)                              // 3 mantissa bits
	q := math.Round(v/step) * step
	return math.Max(-448, math.Min(448, q)) // largest representable value
}

func main() {
	x := []float64{60.0, 0.5, 2.0, 0.05, 30.0}

	i8 := int8PerTensor(x)
	fmt.Printf("INT8   : %.4f\nrel err:", i8)
	for i := range x {
		fmt.Printf(" %.3f", math.Abs((i8[i]-x[i])/x[i]))
	}
	fmt.Print("\nFP8    : [")
	for _, v := range x {
		fmt.Printf("%.4f ", fp8E4M3(v))
	}
	fmt.Print("]\nrel err:")
	for _, v := range x {
		fmt.Printf(" %.3f", math.Abs((fp8E4M3(v)-v)/v))
	}
	fmt.Println()
}

Output:

INT8   : [60.0000 0.4724 1.8898 0.0000 30.2362]
rel err: 0.000 0.055 0.055 1.000 0.008          ← 0.05 became ZERO
FP8    : [60.0000 0.5000 2.0000 0.0508 30.0000 ]
rel err: 0.000 0.000 0.000 0.016 0.000          ← every value within a few percent

INT8 annihilated the 0.05 value. FP8 kept it to within 2%. With outliers present, that’s the difference between a working and a broken quantization.


5. Technical explanation#

FP8 in practice#

Format:      E4M3 for weights and forward activations
             E5M2 for gradients (training only)

Scaling:     even with FP8's range, per-tensor scaling helps:
               x_fp8 = clamp(x / s, -448, 448)
               s chosen so max|x|/s ≈ 448
             This is much simpler than INT8's requirements.

Granularity: per-tensor is usually sufficient (contrast with INT8)
             per-channel weights + per-token activations for extra quality
             (some implementations use per-block, e.g. DeepSeek's 128×128 blocks)

Accumulation: FP32, always

NVIDIA’s Transformer Engine handles scale management, including “delayed scaling” (using a history of recent amax values to set the scale, avoiding a synchronization).

INT8 in practice#

Requires the full apparatus from file 04:

Weights:      per-channel symmetric, static (from the checkpoint)
Activations:  per-token symmetric, dynamic (computed at runtime)
Plus:         SmoothQuant preprocessing to make it work at all
Accumulation: INT32
Dequant:      Y_fp16 = Y_int32 × s_x[token] × s_w[channel]   (fused in the epilogue)

Choosing between them#

Hopper/Blackwell available?
├─ YES → FP8. Simpler, better quality, same speed. Done.
└─ NO (Ampere, Turing, or non-NVIDIA)
    └─ INT8 with SmoothQuant. Validate carefully.

Special cases:
  - AMD MI300: FP8 supported. Use it.
  - Older hardware: INT8 only.
  - Extreme memory constraints: neither; use INT4 weight-only.

What DeepSeek did (worth knowing)#

DeepSeek-V3 trained and serves in FP8 with fine-grained (per-128×128-block for weights, per-128-element-group for activations) scaling and FP32 accumulation at intervals. This is the most aggressive production FP8 deployment publicly documented, and their technical report is worth reading for the numerical details.

Takeaway: FP8 is not just a serving optimization; it’s becoming the native precision.

The blockwise/fine-grained trend#

Per-tensor:     1 scale        simplest, most quality loss
Per-channel:    C scales       standard for weights
Per-token:      T scales       standard for activations
Per-block:      (T/128)×(C/128) scales   ← DeepSeek-style, best quality

Finer granularity costs a little memory and a little kernel complexity, and buys quality. The trend is toward finer.


6-9. Under the hood, performance, production, mistakes#

Under the hood — a FP8 GEMM epilogue:

tensor core: FP8 × FP8 → FP32 accumulator
epilogue (in registers):
    acc *= (s_a × s_b)              dequantize
    acc += bias
    acc = activation(acc)
    if next layer is FP8:
        acc = clamp(acc / s_out, -448, 448)   requantize
        store as FP8                           ← never touches HBM in FP16!
    else:
        store as BF16

That last branch matters: with FP8 throughout, activations move between layers as 1 byte, halving activation traffic too. Mixed FP8/BF16 pipelines lose that.

Performance (70B, H100 node, TP=8):

                Prefill tok/s   Decode tok/s (b=64)   Memory   Max batch
BF16              28,000            1,530             141 GB     ~400
FP8               50,000            2,900              71 GB     ~800
INT8 (SmoothQ)    46,000            2,750              71 GB     ~800

Both roughly 1.8-1.9x. FP8 slightly ahead and much simpler to produce.

Production:

  • Default to FP8 on Hopper+. It is the current production standard.
  • Use per-tensor or finer scaling; use delayed scaling to avoid syncs.
  • Validate at all four levels (Section IV.12). FP8’s quality loss is small but nonzero.
  • Watch for NaN/inf. FP8’s max is 448; an unscaled activation exceeding it becomes inf. Scale management is where FP8 bugs live.
  • Ship the quantized checkpoint with its scales.
  • Note that FP8 KV cache is a separate decision (file 12) and often an additional 1.3-1.8x.

Mistakes:

  • Using INT8 on Hopper when FP8 is available. More work, worse quality.
  • Forgetting scale management. FP8 overflow → inf → NaN.
  • Per-tensor INT8 activations without SmoothQuant. Quality collapse.
  • Assuming FP8 is lossless. It isn’t; validate.
  • Mixing FP8 and BF16 unnecessarily, losing the activation-traffic benefit.

10. Hands-on exercise#

A. Compare the formats. Run the section 4 example. Extend it to a real activation tensor from a model. Plot per-value relative error for INT8 and FP8. Where does each fail?

B. Quantize and measure. Take a model, produce FP8 and INT8 versions. Measure: prefill throughput, decode throughput, memory, perplexity, and one downstream benchmark. Build the comparison table.

C. Scale management. Deliberately set an FP8 scale too small and observe the overflow to inf. Then implement amax-based scaling and confirm it’s fixed.

D. Activation traffic. Compare a fully-FP8 pipeline to one that dequantizes to BF16 between layers. Measure the activation memory traffic difference in a profile.

E. Granularity. If your tooling supports it, compare per-tensor, per-channel, and per-block FP8 scaling on quality. Is the finer granularity worth it for your model?


11. Interview questions#

  1. Compare FP8 E4M3 and INT8. Which handles outliers better and why?
  2. Why does FP8 often not need SmoothQuant?
  3. What is delayed scaling and what problem does it solve?
  4. What is FP8’s maximum value and what happens when you exceed it?
  5. Why does a fully-FP8 pipeline beat a mixed FP8/BF16 one?
  6. When would you use INT8 instead of FP8?
  7. What is fine-grained (block) scaling and what does it buy?

12. Further reading#

  • [ESTABLISHED] Micikevicius et al., “FP8 Formats for Deep Learning” (2022)
  • [REFERENCE] NVIDIA Transformer Engine documentation
  • [ESTABLISHED] DeepSeek-V3 technical report — FP8 training and serving at scale
  • [ESTABLISHED] Xiao et al., “SmoothQuant” (for the INT8 path)
  • Next: 06 — INT4 and low-bit inference

↑↓ navigate ↵ open