1. Problem → Why → Optimization#
PROBLEM Naive round-to-nearest INT4 quantization loses too much quality.
WHY Rounding error is uniform in magnitude, but weights are NOT uniformly
important. Some weights matter far more than others for the layer's output.
OPTIMIZE Choose the rounding to minimize OUTPUT error rather than weight error.Two approaches, both post-training (no retraining required):
GPTQ uses second-order information (the Hessian of the layer's output error)
to decide rounding, compensating for each rounding decision in the
remaining weights.
AWQ identifies the ~1% of weight channels that see the largest activations
and protects them by rescaling before quantization.2. Why they exist#
The naive approach quantizes each weight independently:
w_hat = round(w / s) * s
error is uniform, up to s/2 per weightBut the layer’s output error is X · (W - Ŵ). A weight column that multiplies large
activations contributes proportionally more error. Two insights follow:
- The error should be measured on the output, not the weights. (GPTQ)
- Weight importance depends on the activation magnitude it sees. (AWQ)
Both were discovered around the same time and remain the two dominant approaches.
3. Simple analogy#
Rounding a budget.
Naive: round every line item to the nearest £1,000. Total error accumulates randomly.
GPTQ: round items one at a time, and after each rounding, adjust the remaining items to compensate for the error you just introduced. The total stays much closer to correct.
AWQ: notice that the “staff costs” line is multiplied by headcount (a big number) while the “stationery” line is multiplied by 1. Round stationery aggressively; keep staff costs precise.
4. Tiny example#
package main
import (
"fmt"
"math"
)
// quantInt4 rounds to symmetric 4-bit levels (-7..7) using one scale for the whole slice.
func quantInt4(w []float64) []float64 {
var absmax float64
for _, v := range w {
absmax = math.Max(absmax, math.Abs(v))
}
s := absmax / 7
out := make([]float64, len(w))
for i, v := range w {
out[i] = math.Round(v/s) * s
}
return out
}
func dot(a, b []float64) (s float64) {
for i := range a {
s += a[i] * b[i]
}
return s
}
func main() {
// One weight column, and the activations it sees
w := []float64{0.5, -0.3, 0.8, 0.1}
x := []float64{10.0, 0.1, 0.1, 0.1} // the FIRST input is much larger
fmt.Printf("exact : %.3f\n", dot(w, x)) // 5.0 - 0.03 + 0.08 + 0.01 = 5.06
// --- Naive symmetric INT4 ---
naive := quantInt4(w)
fmt.Printf("naive : %.3f → %.3f\n", naive, dot(naive, x))
// --- AWQ-style: scale up the important channel before quantizing ---
scale := []float64{4, 1, 1, 1} // protect channel 0
scaled := make([]float64, len(w))
for i := range w {
scaled[i] = w[i] * scale[i]
}
awq := quantInt4(scaled)
for i := range awq {
awq[i] /= scale[i] // undo the scaling
}
fmt.Printf("awq : %.3f → %.3f\n", awq, dot(awq, x))
}Output:
exact : 5.060
naive : [0.457 -0.343 0.800 0.114] → 4.629
awq : [0.500 -0.286 0.857 0.000] → 5.057The AWQ-style version has worse weight error on channels 1-3 but much better output error, because it protected the channel that mattered. Error: naive 0.43, AWQ 0.003 — 100x better on the metric that counts.
5. Technical explanation#
GPTQ#
Based on Optimal Brain Quantization / OBS. For a layer with weights W and calibration
activations X:
Objective: minimize ||W X − Ŵ X||²
Process, column by column (in some order):
1. Quantize column i: ŵ_i = quant(w_i)
2. Compute the error: e = w_i − ŵ_i
3. UPDATE all remaining unquantized columns to compensate:
W[:, i+1:] -= e · H⁻¹[i, i+1:] / H⁻¹[i,i]
where H = 2 X Xᵀ is the Hessian of the objective
4. Move to column i+1The Hessian inverse is computed once per layer via Cholesky decomposition. The “lazy batch update” and “act-order” (quantizing columns in order of decreasing importance) refinements make it fast and accurate.
Cost: ~10-60 minutes for a 70B model on one GPU
Data: 128-1024 calibration samples
Result: INT4 with ~1-2% quality loss (vs ~5-15% for naive RTN)act_order=True (also called desc_act) improves quality noticeably but can slow the kernel
(the weight permutation breaks the natural access order). Modern kernels handle it; older ones
were slower. Check your engine.
AWQ#
Based on the observation that ~0.1-1% of weight channels are salient, identified by activation magnitude (not weight magnitude).
Process:
1. Run calibration data; record per-channel activation magnitude |X|.
2. For each layer, search for a per-channel scale s that minimizes output error:
W' = W · diag(s), X' = X · diag(1/s)
(mathematically equivalent: W'X' = WX)
3. Quantize W' — the scaled-up salient channels now use more of the INT4 range.
4. Fold diag(1/s) into the PREVIOUS layer's output (or into the layernorm),
so there's no runtime cost.The scale search is a simple grid search over s = |X|^α for α ∈ [0, 1], minimizing measured
output error. Simple and effective.
Cost: ~5-20 minutes for a 70B model
Data: 128 calibration samples (fewer than GPTQ needs)
Result: INT4 with ~1-2% quality loss, often slightly better than GPTQ
Bonus: no weight reordering → simpler, faster kernelsComparison#
| GPTQ | AWQ | |
|---|---|---|
| Principle | error compensation via Hessian | protect salient channels via scaling |
| Calibration samples | 128-1024 | 128 |
| Quantization time (70B) | 20-60 min | 5-20 min |
| Quality at INT4 g128 | good | good, often slightly better |
| Robustness to calibration data | moderate | better (fewer samples needed) |
| Kernel complexity | act_order complicates it | simpler |
| Generalization to new domains | slightly worse | slightly better |
Practical guidance: try AWQ first (faster, simpler kernels, robust). Use GPTQ if AWQ’s quality is insufficient or if the pre-quantized checkpoints you need are GPTQ.
In practice the difference is small and both are far better than round-to-nearest. The bigger lever is group size: g128 → g32 helps more than GPTQ vs AWQ.
The kernels#
Quantized weight formats need matching kernels:
Marlin fast W4A16 GEMM, up to 4x over FP16 at small batch.
Requires a specific weight permutation, done at load time.
Machete Hopper-optimized successor.
exllamav2 fast kernels for the EXL2 format.
Naive dequant-then-cuBLAS: only ~1.3-1.5x. Avoid.The kernel matters as much as the quantization method. A model quantized with AWQ but served through a naive dequant path gets a fraction of the possible speedup. Check what your engine uses.
6-9. Under the hood, performance, production, mistakes#
Under the hood: the packed format:
INT4 weights, group size 128, per-group FP16 scale and zero-point:
qweight: [in_features/8, out_features] int32 (8 × 4-bit packed)
scales: [in_features/128, out_features] fp16
zeros: [in_features/128, out_features/8] int32
g_idx: [in_features] int32 (only with act_order)
Bits per weight = 4 + 16/128 (scale) + 4/128 (zero) ≈ 4.16Performance:
70B, one H100 equivalent, batch 1 decode:
BF16: 24 tok/s
INT4 naive dequant: 33 tok/s (1.4x)
INT4 Marlin: 72 tok/s (3.0x)
Batch 64:
BF16: 1,530 tok/s
INT4 Marlin: 2,900 tok/s (1.9x — less benefit, weights are amortized)Production:
- Use pre-quantized checkpoints where available; otherwise quantize offline and store.
- Verify your engine uses a fast kernel (Marlin/Machete), not naive dequant.
- Group size 128 is the default; 64 or 32 if you need quality and can afford ~2-6% more memory.
- Validate at Level 3 (long generations). INT4’s cost shows there.
- Keep calibration data representative and version it alongside the checkpoint.
Mistakes:
- Naive round-to-nearest INT4. 5-15% quality loss for no reason; use AWQ or GPTQ.
- Calibrating on the wrong distribution.
- Expecting prefill speedup. W4A16 doesn’t give it.
- Not checking which kernel is used.
- Per-tensor or per-channel scales at INT4. Use groups.
- Validating only on short benchmarks.
10. Hands-on exercise#
A. Quantize a model. Take a 7B model. Quantize it three ways: naive RTN INT4, GPTQ INT4 g128, AWQ INT4 g128. Measure perplexity for each. Quantify the benefit of the smart methods.
B. Group size. Quantize with g32, g64, g128, per-channel. Plot perplexity vs bits-per-weight (including scale overhead). Where’s the knee?
C. Reproduce the AWQ intuition. Extend the toy example in section 4 to a real weight matrix and real activations from a model. Compute output error with and without activation-aware scaling.
D. Kernel comparison. Serve the same AWQ model through a naive dequant path and through Marlin. Measure decode throughput at batch 1, 8, 64. Report the speedups.
E. Calibration sensitivity. Quantize with 32, 128, and 512 calibration samples, and with in-domain vs out-of-domain data. Build a 2×3 quality table.
11. Interview questions#
- Why does naive INT4 round-to-nearest lose so much quality?
- Explain GPTQ’s error-compensation mechanism.
- Explain AWQ’s core insight and how the scaling is made free at runtime.
- Compare GPTQ and AWQ. Which would you choose and why?
- What is
act_orderand what does it cost? - What is the effective bits-per-weight for INT4 g128?
- Why does the choice of kernel matter as much as the quantization method?
12. Further reading#
- [ESTABLISHED] Frantar et al., “GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers” (2022)
- [ESTABLISHED] Lin et al., “AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration” (2023)
- [ESTABLISHED] Frantar & Alistarh, “Marlin” kernel
- [REFERENCE]
autoawq,gptqmodel,llm-compressorrepositories - Next: 04 — Activation quantization and SmoothQuant