1. Problem → Why → Optimization#
PROBLEM Even at 8 bits, a 70B model is 70 GB — too large for one GPU, and
decode still reads 70 GB per step.
WHY Weight bytes dominate decode bandwidth at small-to-moderate batch.
OPTIMIZE Go to 4 bits (or fewer) for weights, keeping activations at FP16.
W4A16: 4x less weight traffic, 2.5-3.5x faster decode.
TRADE-OFFS
✓ 4x memory reduction → fits on smaller/fewer GPUs, more KV cache
✓ 2.5-3.5x decode speedup at small batch
✗ NO prefill speedup (compute is still FP16); can be slightly slower
✗ 1-3% quality loss, more on hard reasoning tasks
✗ Needs a good kernel (Marlin/Machete) or you get half the benefit
WHEN TO USE
Decode-dominated workloads, memory-constrained deployments, latency-critical
small-batch serving, consumer/edge hardware.
WHEN NOT TO USE
Prefill-heavy workloads (RAG, summarization) — no benefit, extra risk.
High-batch throughput serving — weights are already amortized.
Tasks where 1-3% quality matters (competitive reasoning, code correctness).2. Why it works at all#
Neural network weights are heavily over-parameterized and remarkably robust to precision reduction. Empirically:
Bits Typical perplexity increase (7B model, good method, g128)
16 baseline
8 +0.01 - 0.05
6 +0.05 - 0.15
4 +0.15 - 0.50
3 +0.5 - 1.5
2 +3 - 15 (usually unusable without special methods)The cliff is between 3 and 2 bits. Four bits sits comfortably above it — which is why INT4 is where the industry settled for weight-only quantization.
3. Simple analogy#
JPEG compression. A photo at 100% quality and at 85% quality are visually indistinguishable at 5x smaller. At 40% you see artifacts. At 10% it’s unrecognizable.
INT4 is around the 85% setting: the difference is measurable but usually not noticeable. INT2 is the 10% setting.
And, as with JPEG, what you’re compressing matters: fine detail (attention projections, reasoning-critical layers) suffers more than smooth areas (FFN bulk).
4. Tiny example — the memory and speed arithmetic#
Llama-3-70B:
Precision Weight bytes Fits on Decode floor (H100, b=1)
BF16 141 GB 2× H100 (TP=2) 42 ms (with TP=2: 21 ms)
FP8 71 GB 1× H100 21 ms
INT4 35 GB 1× H100, or 10.5 ms
1× A100-40GB (tight)
On a single H100 (80 GB):
BF16: doesn't fit
FP8: 71 GB weights, 9 GB budget → ~2-8 concurrent sequences. Cramped.
INT4: 35 GB weights, 45 GB KV budget → ~140 sequences at 1k context.INT4 turns “needs two GPUs and can barely batch” into “one GPU with real concurrency.” That’s the practical argument, and it’s often more valuable than the raw speed.
5. Technical explanation#
The formats#
INT4 symmetric, group 128:
4 bits per weight + FP16 scale per 128 + INT4 zero per 128
= 4 + 16/128 + 4/128 = 4.16 bits/weight (26% of FP16)
INT4 group 32:
4 + 16/32 + 4/32 = 4.63 bits/weight (29% of FP16)
noticeably better quality
NF4 (NormalFloat4, from QLoRA):
4-bit levels placed at the quantiles of a normal distribution rather than
uniformly. Better for Gaussian-distributed weights.
Used by bitsandbytes.
EXL2 / mixed-bit:
Different bits per layer based on measured sensitivity.
e.g. 6 bits for attention, 3 bits for FFN, averaging 4.
Best quality per bit; more complex.
AWQ / GPTQ INT4:
See file 03. The standard production choices.The kernel is half the story#
Approach Decode speedup vs FP16 (batch 1)
────────────────────────────────────────────────────────────
Dequantize to FP16, cuBLAS 1.3-1.5x ← wastes most of the benefit
Fused dequant in a simple kernel 2.0-2.5x
Marlin / Machete 3.0-3.7x ← what you wantMarlin’s tricks:
- Weights pre-permuted at load so unpacking is branch-free and coalesced.
- Dequantization fused into the tensor-core operand pipeline.
- Async copy and double buffering.
- Careful shared-memory layout to avoid bank conflicts.
Always check which kernel your engine uses. vLLM will log it; look for MarlinLinearMethod
or similar rather than a generic dequant path.
Where INT4 hurts#
Not uniformly. Empirically:
Robust to INT4: Sensitive to INT4:
general knowledge multi-step arithmetic
fluent generation code generation correctness
summarization long-chain reasoning
classification instruction following at the margins
short answers structured output adherenceThe pattern: tasks with long dependency chains suffer most, because small per-token errors compound (Section III.10). A 1% per-token quality reduction over 500 reasoning tokens is not a 1% reduction in answer quality.
This is why L3 validation (long generations) is non-negotiable for INT4.
Below 4 bits#
INT3 [EMERGING] Quality loss is significant (0.5-1.5 ppl).
Viable with careful methods for some models.
INT2 [RESEARCH] Requires special techniques (learned codebooks,
extreme fine-tuning). Not production-ready for general use.
Binary/ternary [RESEARCH] BitNet-style models are TRAINED in low precision
rather than quantized after. Different approach; promising
but requires training from scratch.
Mixed-bit [EMERGING] Per-layer or per-channel bit allocation based on
sensitivity. Best quality/bit. Complex kernels.
Vector/codebook [EMERGING] AQLM, QuIP#: quantize groups of weights jointly
quantization to entries in a learned codebook. Excellent quality at 2-3
bits, but slow kernels (codebook lookup is not GEMM-friendly).The honest state of the field: 4 bits is production-standard; 3 is possible; 2 is research.
6-9. Under the hood, performance, production, mistakes#
Under the hood — the packed layout:
Storage for a (4096, 11008) weight matrix, INT4 g128:
qweight: (512, 11008) int32 [4096/8 = 512, 8 weights per int32]
scales: (32, 11008) fp16 [4096/128 = 32 groups]
qzeros: (32, 1376) int32 [zeros also packed 8 per int32]
Total: 512×11008×4 + 32×11008×2 + 32×1376×4 = 22.5 + 0.70 + 0.18 MB = 23.4 MB
vs FP16: 4096×11008×2 = 90.2 MB
Ratio: 0.26 ✓Performance (70B, 1× H100, real measurements vary):
batch 1 batch 8 batch 32 batch 128
BF16 (TP=2) 47 tok/s 360 1,320 4,100
FP8 47 370 1,400 4,600
INT4 Marlin 92 640 1,900 4,900
Speedup vs BF16:
1.96x 1.78x 1.44x 1.20xThe benefit shrinks with batch size, because batching already amortizes the weight read. At batch 128, KV and activation traffic dominate and INT4 helps much less.
Production:
- Use INT4 for: single-GPU deployment of large models, latency-critical low-batch serving, memory-constrained environments.
- Don’t use INT4 for: high-throughput batch serving, prefill-heavy workloads.
- Verify the kernel. Naive dequant loses half the benefit.
- Validate at L3. Long generations, reasoning tasks, your domain.
- Consider g32 over g128 if quality is marginal — 3% more memory for meaningfully better quality.
- Keep the FP16 or FP8 checkpoint deployable for rollback and for comparison.
Mistakes:
- Using INT4 for a prefill-heavy workload. No gain, real risk.
- Expecting the batch-1 speedup at batch 128.
- Naive dequantization kernels.
- Per-tensor or per-channel scales. Use groups.
- Validating with perplexity only. INT4’s cost is in reasoning, not perplexity.
- Quantizing embeddings and LM head to INT4.
- Going below 4 bits in production without extraordinary validation.
10. Hands-on exercise#
A. The quality curve. Quantize a model to 8, 6, 4, 3, and 2 bits (using a method that supports them, e.g. EXL2 or GPTQ). Plot perplexity vs bits. Find the cliff.
B. Batch dependence. Measure INT4 vs FP16 decode throughput at batch 1, 8, 32, 128, 256. Plot the speedup ratio vs batch size. Explain the trend using the byte breakdown.
C. Kernel comparison. Serve the same INT4 model through a naive dequant path and through Marlin. Report both. Confirm the ~2x kernel difference.
D. Task sensitivity. Evaluate FP16 and INT4 versions on: a knowledge benchmark, a math benchmark, a code benchmark, and long-form generation quality. Which degrades most? Does it match the table in section 5?
E. Group size. Compare g32, g64, g128 on quality and memory. Compute quality per byte for each.
11. Interview questions#
- Why does INT4 weight-only quantization help decode but not prefill?
- What is the effective bits-per-weight for INT4 g128, including overhead?
- Why does the INT4 speedup shrink as batch size grows?
- Which tasks degrade most under INT4, and why?
- What does a good INT4 kernel do that a naive one doesn’t?
- When would you choose g32 over g128?
- What is the current honest state of sub-4-bit quantization?
12. Further reading#
- [ESTABLISHED] Frantar et al., “GPTQ”; Lin et al., “AWQ”
- [ESTABLISHED] Dettmers et al., “QLoRA” (2023) — NF4
- [ESTABLISHED] Marlin kernel (IST-DASLab)
- [EMERGING] Egiazarian et al., “AQLM” (2024); Tseng et al., “QuIP#” (2024)
- [RESEARCH] Ma et al., “BitNet b1.58” (2024)
- Next: 07 — Kernel and operator fusion