PidokuInfra

Weight-Only Quantization: GPTQ and AWQ

Intermediate Advanced 1h 30m Difficulty 4/5 Topic 03 of 14

Prerequisites III.12, 02


1. Problem → Why → Optimization#

PROBLEM   Naive round-to-nearest INT4 quantization loses too much quality.
WHY       Rounding error is uniform in magnitude, but weights are NOT uniformly
          important. Some weights matter far more than others for the layer's output.
OPTIMIZE  Choose the rounding to minimize OUTPUT error rather than weight error.

Two approaches, both post-training (no retraining required):

GPTQ  uses second-order information (the Hessian of the layer's output error)
      to decide rounding, compensating for each rounding decision in the
      remaining weights.

AWQ   identifies the ~1% of weight channels that see the largest activations
      and protects them by rescaling before quantization.

2. Why they exist#

The naive approach quantizes each weight independently:

w_hat = round(w / s) * s
error is uniform, up to s/2 per weight

But the layer’s output error is X · (W - Ŵ). A weight column that multiplies large activations contributes proportionally more error. Two insights follow:

  1. The error should be measured on the output, not the weights. (GPTQ)
  2. Weight importance depends on the activation magnitude it sees. (AWQ)

Both were discovered around the same time and remain the two dominant approaches.


3. Simple analogy#

Rounding a budget.

Naive: round every line item to the nearest £1,000. Total error accumulates randomly.

GPTQ: round items one at a time, and after each rounding, adjust the remaining items to compensate for the error you just introduced. The total stays much closer to correct.

AWQ: notice that the “staff costs” line is multiplied by headcount (a big number) while the “stationery” line is multiplied by 1. Round stationery aggressively; keep staff costs precise.


4. Tiny example#

Go
package main

import (
	"fmt"
	"math"
)

// quantInt4 rounds to symmetric 4-bit levels (-7..7) using one scale for the whole slice.
func quantInt4(w []float64) []float64 {
	var absmax float64
	for _, v := range w {
		absmax = math.Max(absmax, math.Abs(v))
	}
	s := absmax / 7
	out := make([]float64, len(w))
	for i, v := range w {
		out[i] = math.Round(v/s) * s
	}
	return out
}

func dot(a, b []float64) (s float64) {
	for i := range a {
		s += a[i] * b[i]
	}
	return s
}

func main() {
	// One weight column, and the activations it sees
	w := []float64{0.5, -0.3, 0.8, 0.1}
	x := []float64{10.0, 0.1, 0.1, 0.1}     // the FIRST input is much larger
	fmt.Printf("exact : %.3f\n", dot(w, x)) // 5.0 - 0.03 + 0.08 + 0.01 = 5.06

	// --- Naive symmetric INT4 ---
	naive := quantInt4(w)
	fmt.Printf("naive : %.3f → %.3f\n", naive, dot(naive, x))

	// --- AWQ-style: scale up the important channel before quantizing ---
	scale := []float64{4, 1, 1, 1} // protect channel 0
	scaled := make([]float64, len(w))
	for i := range w {
		scaled[i] = w[i] * scale[i]
	}
	awq := quantInt4(scaled)
	for i := range awq {
		awq[i] /= scale[i] // undo the scaling
	}
	fmt.Printf("awq   : %.3f → %.3f\n", awq, dot(awq, x))
}

Output:

exact : 5.060
naive : [0.457 -0.343 0.800 0.114] → 4.629
awq   : [0.500 -0.286 0.857 0.000] → 5.057

The AWQ-style version has worse weight error on channels 1-3 but much better output error, because it protected the channel that mattered. Error: naive 0.43, AWQ 0.003 — 100x better on the metric that counts.


5. Technical explanation#

GPTQ#

Based on Optimal Brain Quantization / OBS. For a layer with weights W and calibration activations X:

Objective: minimize ||W X − Ŵ X||²

Process, column by column (in some order):
  1. Quantize column i:  ŵ_i = quant(w_i)
  2. Compute the error:  e = w_i − ŵ_i
  3. UPDATE all remaining unquantized columns to compensate:
        W[:, i+1:] -= e · H⁻¹[i, i+1:] / H⁻¹[i,i]
     where H = 2 X Xᵀ is the Hessian of the objective
  4. Move to column i+1

The Hessian inverse is computed once per layer via Cholesky decomposition. The “lazy batch update” and “act-order” (quantizing columns in order of decreasing importance) refinements make it fast and accurate.

Cost:    ~10-60 minutes for a 70B model on one GPU
Data:    128-1024 calibration samples
Result:  INT4 with ~1-2% quality loss (vs ~5-15% for naive RTN)

act_order=True (also called desc_act) improves quality noticeably but can slow the kernel (the weight permutation breaks the natural access order). Modern kernels handle it; older ones were slower. Check your engine.

AWQ#

Based on the observation that ~0.1-1% of weight channels are salient, identified by activation magnitude (not weight magnitude).

Process:
  1. Run calibration data; record per-channel activation magnitude |X|.
  2. For each layer, search for a per-channel scale s that minimizes output error:
        W' = W · diag(s),   X' = X · diag(1/s)
     (mathematically equivalent: W'X' = WX)
  3. Quantize W' — the scaled-up salient channels now use more of the INT4 range.
  4. Fold diag(1/s) into the PREVIOUS layer's output (or into the layernorm),
     so there's no runtime cost.

The scale search is a simple grid search over s = |X|^α for α ∈ [0, 1], minimizing measured output error. Simple and effective.

Cost:    ~5-20 minutes for a 70B model
Data:    128 calibration samples (fewer than GPTQ needs)
Result:  INT4 with ~1-2% quality loss, often slightly better than GPTQ
Bonus:   no weight reordering → simpler, faster kernels

Comparison#

GPTQAWQ
Principleerror compensation via Hessianprotect salient channels via scaling
Calibration samples128-1024128
Quantization time (70B)20-60 min5-20 min
Quality at INT4 g128goodgood, often slightly better
Robustness to calibration datamoderatebetter (fewer samples needed)
Kernel complexityact_order complicates itsimpler
Generalization to new domainsslightly worseslightly better

Practical guidance: try AWQ first (faster, simpler kernels, robust). Use GPTQ if AWQ’s quality is insufficient or if the pre-quantized checkpoints you need are GPTQ.

In practice the difference is small and both are far better than round-to-nearest. The bigger lever is group size: g128 → g32 helps more than GPTQ vs AWQ.

The kernels#

Quantized weight formats need matching kernels:

Marlin       fast W4A16 GEMM, up to 4x over FP16 at small batch.
             Requires a specific weight permutation, done at load time.
Machete      Hopper-optimized successor.
exllamav2    fast kernels for the EXL2 format.
Naive dequant-then-cuBLAS:  only ~1.3-1.5x. Avoid.

The kernel matters as much as the quantization method. A model quantized with AWQ but served through a naive dequant path gets a fraction of the possible speedup. Check what your engine uses.


6-9. Under the hood, performance, production, mistakes#

Under the hood: the packed format:

INT4 weights, group size 128, per-group FP16 scale and zero-point:

  qweight:  [in_features/8, out_features]  int32   (8 × 4-bit packed)
  scales:   [in_features/128, out_features] fp16
  zeros:    [in_features/128, out_features/8] int32
  g_idx:    [in_features] int32              (only with act_order)

Bits per weight = 4 + 16/128 (scale) + 4/128 (zero) ≈ 4.16

Performance:

70B, one H100 equivalent, batch 1 decode:
  BF16:              24 tok/s
  INT4 naive dequant: 33 tok/s   (1.4x)
  INT4 Marlin:        72 tok/s   (3.0x)
  
Batch 64:
  BF16:              1,530 tok/s
  INT4 Marlin:       2,900 tok/s (1.9x — less benefit, weights are amortized)

Production:

  • Use pre-quantized checkpoints where available; otherwise quantize offline and store.
  • Verify your engine uses a fast kernel (Marlin/Machete), not naive dequant.
  • Group size 128 is the default; 64 or 32 if you need quality and can afford ~2-6% more memory.
  • Validate at Level 3 (long generations). INT4’s cost shows there.
  • Keep calibration data representative and version it alongside the checkpoint.

Mistakes:

  • Naive round-to-nearest INT4. 5-15% quality loss for no reason; use AWQ or GPTQ.
  • Calibrating on the wrong distribution.
  • Expecting prefill speedup. W4A16 doesn’t give it.
  • Not checking which kernel is used.
  • Per-tensor or per-channel scales at INT4. Use groups.
  • Validating only on short benchmarks.

10. Hands-on exercise#

A. Quantize a model. Take a 7B model. Quantize it three ways: naive RTN INT4, GPTQ INT4 g128, AWQ INT4 g128. Measure perplexity for each. Quantify the benefit of the smart methods.

B. Group size. Quantize with g32, g64, g128, per-channel. Plot perplexity vs bits-per-weight (including scale overhead). Where’s the knee?

C. Reproduce the AWQ intuition. Extend the toy example in section 4 to a real weight matrix and real activations from a model. Compute output error with and without activation-aware scaling.

D. Kernel comparison. Serve the same AWQ model through a naive dequant path and through Marlin. Measure decode throughput at batch 1, 8, 64. Report the speedups.

E. Calibration sensitivity. Quantize with 32, 128, and 512 calibration samples, and with in-domain vs out-of-domain data. Build a 2×3 quality table.


11. Interview questions#

  1. Why does naive INT4 round-to-nearest lose so much quality?
  2. Explain GPTQ’s error-compensation mechanism.
  3. Explain AWQ’s core insight and how the scaling is made free at runtime.
  4. Compare GPTQ and AWQ. Which would you choose and why?
  5. What is act_order and what does it cost?
  6. What is the effective bits-per-weight for INT4 g128?
  7. Why does the choice of kernel matter as much as the quantization method?

12. Further reading#

  • [ESTABLISHED] Frantar et al., “GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers” (2022)
  • [ESTABLISHED] Lin et al., “AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration” (2023)
  • [ESTABLISHED] Frantar & Alistarh, “Marlin” kernel
  • [REFERENCE] autoawq, gptqmodel, llm-compressor repositories
  • Next: 04 — Activation quantization and SmoothQuant

↑↓ navigate↵ openesc close