PidokuInfra

Quantization Research

Expert Advanced 1h Difficulty 4/5 Topic 04 of 11

Prerequisites VII.02-06, XIII.10


1. Where the field is#

SOLVED
  ✓ FP8 (Hopper+): production standard, ~2x, minimal quality cost
  ✓ INT8 W8A8 with SmoothQuant: production standard elsewhere
  ✓ INT4 weight-only (GPTQ/AWQ): production-common for decode
  ✓ the mechanism: block/group scaling is what makes low bits work
  ✓ the outlier problem: understood and addressed

ACTIVE
  ~ FP4/FP6 with hardware support (Blackwell)
  ~ sub-4-bit via codebooks (quality good, speed poor)
  ~ quantization-aware training vs post-training at 4 bits and below
  ~ mixed-precision allocation (which layers get which bits)
  ~ KV cache quantization below 4 bits

DIMINISHING
  - yet another post-training method at 4 bits
    (GPTQ and AWQ are close to the achievable limit for PTQ)

2. The three eras#

Useful framing for reading the literature:

ERA 1 (2022-2023): DISCOVERING THE PROBLEM
  LLM.int8() found emergent outliers
  → the reason naive quantization fails at scale
  GPTQ, AWQ solved weight quantization
  SmoothQuant solved activation quantization
  → the fundamentals were established

ERA 2 (2023-2024): REFINEMENT AND FORMATS
  FP8 hardware support arrived; formats standardized (E4M3/E5M2)
  block/microscaling formalized (MX)
  KV quantization studied separately
  the K/V asymmetry discovered
  → the practical toolkit reached maturity

ERA 3 (2024- ): HARDWARE-DRIVEN, AND BELOW 4 BITS
  FP4/FP6 with native tensor core support
  the question shifts from "can we represent it" to
    "can we train for it"
  codebook methods for extreme compression
  → the frontier moves from post-training to training-time

Reading a paper without knowing which era it’s from wastes time. An Era 1 paper proposing a weight quantization method is historically interesting; an Era 3 paper about FP4 training is actionable.


3. The open questions#

Q1: IS POST-TRAINING QUANTIZATION SUFFICIENT AT 4 BITS AND BELOW?
  evidence for: AWQ/GPTQ at 4 bits are widely deployed with
    acceptable quality
  evidence against: quality loss at 4 bits is measurable on
    reasoning and long-form tasks (Section VII.06)
  → the answer is probably "sufficient for many tasks, not all,
    and you must measure yours"

Q2: WHERE IS THE FLOOR FOR PTQ?
  4 bits with good methods: workable
  3 bits: 0.5-1.5 perplexity increase, significant task degradation
  2 bits: requires codebooks, and then the kernels are slow
  → 4 bits looks like the practical floor for pretrained models

Q3: DOES TRAINING IN LOW PRECISION CHANGE THE FLOOR?
  DeepSeek trains in FP8 and serves in FP8: no quantization step at all
  BitNet trains ternary: no quantization error by construction
  → this is the most likely path below 4 bits

Q4: HOW SHOULD BITS BE ALLOCATED ACROSS LAYERS?
  we know: embeddings, LM head, norms, routers need more
  we don't know: a principled allocation across the rest
  → mixed-precision papers show gains; the search cost is high

Q5: KV CACHE BELOW 4 BITS?
  KIVI demonstrates 2-bit with asymmetric treatment
  → but long-context retrieval evaluation is thin

4. What’s genuinely new and worth attention#

1. NATIVE LOW-PRECISION TRAINING
   DeepSeek-V3's FP8 training at scale is the significant datapoint:
   a frontier-scale model trained AND served in FP8, with published
   detail about the numerical techniques (fine-grained scaling,
   periodic FP32 accumulation).
   → removes the quantization step entirely
   → the likely path to FP4-native models

2. MICROSCALING FORMATS (MX)
   An industry standard (OCP) for block-scaled low-precision formats.
   → hardware support across vendors
   → matters because it means FP4 won't be NVIDIA-specific

3. CODEBOOK / VECTOR QUANTIZATION
   AQLM, QuIP#: excellent quality at 2-3 bits.
   → but the kernels are slow (lookup isn't GEMM-friendly)
   → useful for FITTING a model, not for going fast
   → watch for hardware or kernel work that fixes the speed

4. QUANTIZATION-AWARE FINE-TUNING AT SCALE
   Brief fine-tuning after quantization recovers much of the loss.
   → cheaper than full QAT, better than pure PTQ
   → increasingly standard in production pipelines

Item 1 is the important one. If models are trained in the precision they’ll be served in, the entire post-training quantization literature becomes a legacy concern.


5. What’s saturated#

✗ Another post-training weight quantization method at 4 bits
    GPTQ and AWQ are near the PTQ limit. Marginal improvements
    don't justify integration cost.

✗ Outlier-handling variations
    The problem is understood; SmoothQuant and FP8's range
    address it.

✗ Papers reporting only perplexity
    Perplexity is insensitive to what quantization actually breaks
    (Section VII.06).

✗ Methods requiring kernels that don't exist
    A quality result without a fast kernel is not deployable.

6. How to evaluate a quantization paper#

Applying Section XIV.01’s framework specifically:

Q1 BASELINE
  ✓ compared against GPTQ AND AWQ at the same bits and group size?
  ✗ compared only against round-to-nearest? (a trivially weak baseline)

Q2 CONFIGURATION
  ✓ multiple model sizes (7B, 70B — behavior differs)
  ✓ the group size stated (a 4-bit result at group 32 vs group 128
    is a different claim)
  ✓ effective bits-per-weight including scale overhead

Q3 MECHANISM
  ✓ does it reduce bytes, or bytes AND compute? (weight-only vs W×A×)
  ✓ what does the kernel look like? Does one exist?

Q5 QUALITY  ← the critical one for this area
  ✓ perplexity (necessary)
  ✓ standard benchmarks
  ✓ LONG-FORM GENERATION          ← usually missing
  ✓ REASONING / MULTI-STEP TASKS  ← usually missing
  ✓ code generation correctness   ← usually missing
  ✗ perplexity only → insufficient evidence

Q6 HIDDEN COSTS
  ✓ calibration data requirement and sensitivity
  ✓ quantization time
  ✓ does it need a custom kernel?
  ✓ does it compose with paged attention, CUDA graphs, TP?

Q7 SIMPLER ALTERNATIVE
  → is FP8 available? It's 2x, nearly free, and well-validated.
  → does the paper's method beat FP8 + AWQ-INT4 on the FFN?

The “long-form generation” gap is systematic in this literature. Quantization error compounds across autoregressive steps (Section III.10), so a method validated on short-answer benchmarks has not been validated for the workload most people run.


7. What to watch, ranked#

1. FP4-NATIVE MODELS
   trained in FP4, served in FP4, no quantization step
   → would be a 4x memory and ~4x compute change with no quality
     question
   → watch: model technical reports, not quantization papers

2. FP4 POST-TRAINING QUALITY AT SCALE
   can a pretrained BF16 model be served at FP4 acceptably?
   → watch: independent evaluations on Blackwell hardware

3. FAST CODEBOOK KERNELS
   would make 2-3 bit deployment practical
   → watch: kernel work, not quantization methods

4. STANDARDIZED MIXED-PRECISION ALLOCATION
   a principled answer to "which layers get which bits"
   → watch: whether anyone finds a rule rather than a search

5. KV QUANTIZATION BELOW 4 BITS, PROPERLY EVALUATED
   → watch: long-context retrieval evaluations, specifically

8. Hands-on exercise#

A. Reproduce the era boundaries. Take one paper from each era in section 2. What problem was each solving? What had changed between them?

B. Evaluate the long-form gap. Take a recent 4-bit quantization paper. Does it evaluate long generations? If not, run that evaluation yourself on their released model and see whether the claim holds.

C. Effective bits. For three quantization methods, compute the true bits-per-weight including scale and zero-point overhead. Do the papers’ headline bit counts match?

D. Compare to the simple alternative. Take a sophisticated 4-bit method and compare it against FP8 (if you have Hopper) on quality and speed. Does the extra complexity pay?

E. Test the compounding. For an INT4 model, measure output divergence from FP16 at token positions 10, 100, 500, and 2000. Does error compound as predicted (Section III.10)?


9. Interview questions#

  1. What’s solved and what’s open in quantization?
  2. Why is 4 bits the practical floor for post-training quantization?
  3. What would change if models were trained in the precision they’re served in?
  4. What’s systematically missing from quantization paper evaluations?
  5. Why are codebook methods good at quality and bad at speed?
  6. How would you evaluate a new quantization method?
  7. What is microscaling and why does an industry standard matter?

10. Further reading#

  • [ESTABLISHED] Dettmers et al., “LLM.int8()”; Frantar et al., “GPTQ”; Lin et al., “AWQ”; Xiao et al., “SmoothQuant”
  • [ESTABLISHED] Micikevicius et al., “FP8 Formats” (2022)
  • [ESTABLISHED] DeepSeek-V3 technical report — FP8 training at scale
  • [EMERGING] OCP Microscaling Formats specification
  • [EMERGING] Egiazarian et al., “AQLM”; Tseng et al., “QuIP#”
  • [RESEARCH] Ma et al., “BitNet b1.58”
  • Next: 05 — Decoding research

↑↓ navigate↵ openesc close