1. What is it?#
Before a model can serve anything, tens to hundreds of gigabytes must travel from wherever they are stored into GPU memory. That journey is the dominant term in cold-start time, and cold-start time is what makes LLM autoscaling hard (Section I.10).
Object storage (S3/GCS) → local disk → page cache → host RAM → GPU HBM
~0.1-10 GB/s 3-14 GB/s memcpy PCIe 25 GB/s2. Why does it exist as a topic?#
Because the naive path is 10-20x slower than the good path, and the difference is the gap between a 12-minute and a 45-second cold start. That gap determines whether you can autoscale at all, which determines how much idle capacity you must pay for, which is often a six-figure annual line item.
3. Simple analogy#
Moving a library into a reading room. The books exist in a warehouse across the country (object storage). You can ship them each time (slow), keep a local branch copy (disk cache), or keep them already on the shelves (warm replica). The engineering question is how much shelf space you’re willing to pay for to avoid the shipping delay.
4. Tiny example#
Measure the storage tiers on your machine:
# Sequential read from disk, bypassing page cache
dd if=/path/to/model.safetensors of=/dev/null bs=1M count=8192 iflag=direct
# → GB/s from the device itself
# Same file, second time (now in page cache)
dd if=/path/to/model.safetensors of=/dev/null bs=1M count=8192
# → GB/s from RAM, typically 5-20x faster
# Drop caches to compare fairly (needs root; don't do this in prod)
sync; echo 3 > /proc/sys/vm/drop_cachesTypical numbers:
Network object storage (single stream) 100-300 MB/s
Network object storage (32 parallel) 2-10 GB/s
Network filesystem (NFS/EFS) 0.2-3 GB/s
Local NVMe SSD 3-7 GB/s
Local NVMe RAID / Gen5 7-25 GB/s
Page cache (RAM) 10-50 GB/sCold-start arithmetic for a 140 GB model:
Single-stream S3 at 200 MB/s 700 s (11.7 min) ← the naive path
32-stream S3 at 5 GB/s 28 s
Local NVMe at 5 GB/s 28 s
Page cache at 20 GB/s 7 s
+ H2D over PCIe Gen4 at 20 GB/s 7 s35x difference between the worst and best path, using the same hardware. This is one of the highest-leverage, lowest-glamour optimizations in the field.
5. Technical explanation#
The model file formats#
| Format | Properties | Loading behavior |
|---|---|---|
.pt / .bin (pickle) | Python pickle; arbitrary code execution risk | must deserialize; slow; single-threaded |
.safetensors | JSON header + raw tensor bytes; no code execution | mmap-able, zero-copy, parallelizable |
| GGUF | llama.cpp format; embeds quantization + metadata | mmap-able |
| TensorRT engine | precompiled, hardware-specific | fast load, but must be built per GPU/config |
Use safetensors. It is safe (no pickle), fast (mmap), and shardable. The header tells you every tensor’s byte offset, so you can read shards in parallel and copy directly.
mmap and the page cache#
from safetensors.torch import load_file
sd = load_file("model.safetensors", device="cpu") # mmap-backedmmap maps the file into the address space without reading it. Pages fault in on access. Two
consequences:
- Second load is nearly free if the page cache still holds the file. Restarting a server on the same host is fast; the first start after a node boot is slow.
- RSS looks huge but is shared and reclaimable. Don’t panic at the number.
Parallel and direct paths#
The fast loading recipe:
1. Shard the model into N files (safetensors sharding is standard).
2. Read shards in parallel (thread pool; each thread does large sequential reads).
3. Copy to pinned host buffers.
4. cudaMemcpyAsync to GPU on multiple streams, overlapping with reads.
5. For tensor parallelism, each rank reads only ITS shard — N-way parallel by construction.Point 5 is important: with TP=8, each rank needs 1/8 of the weights, so eight processes read different shards concurrently and cold start drops ~8x if your storage can supply the aggregate bandwidth.
GPUDirect Storage (GDS) goes further: NVMe → GPU HBM via DMA, skipping host RAM entirely.
Requires supported hardware and nvidia-fs. Where available it gives another 1.5-2x and frees
host memory.
Where the rest of cold start goes#
Container image pull (if not cached) 30-180 s ← use small images, pre-pull, lazy loading
Python import torch + CUDA init 5-20 s
Weight read + transfer 30-600 s ← the topic of this file
Quantization at load time (if any) 10-120 s ← precompute instead!
CUDA graph capture / warmup 10-60 s
torch.compile (if used, uncached) 60-600 s ← cache the compilation artifacts
──────────────────────────────────────────────────Two of those lines are pure waste in production and are commonly overlooked: quantizing at load time (do it offline once and store the quantized checkpoint) and recompiling every start (persist the compile cache in the image or a volume).
6. Under the hood#
# Watch I/O during a model load
iostat -x 1
# %util near 100 with low MB/s → small random reads; you're not doing large sequential I/O
# Which files is the process reading?
sudo lsof -p <pid> | grep safetensors
# Page cache contents for a file
vmtouch -v /models/model.safetensors # from the vmtouch package
vmtouch -t /models/ # deliberately pre-load into page cachevmtouch -t is a legitimate production trick: after a node boots, pre-warm the page cache with
the models you expect to serve, so the first real load is RAM-speed.
7. Performance implications#
Cold start dominates your ability to react to load:
Traffic spike duration: 3 minutes
Cold start: 10 minutes
→ Autoscaling contributes NOTHING to this spike.
→ You must pre-provision.
Cold start: 45 seconds
→ Autoscaling can respond to sustained increases, though not to sharp spikes.The fix is layered:
- Make cold start fast (this file).
- Keep warm standby replicas (costs money, works instantly).
- Predictive scaling on traffic patterns (works for daily cycles).
- Overflow to a different model or provider during the gap.
8. Production implications#
- Cache weights on local NVMe. Pull from object storage once per node, not once per pod. A DaemonSet or an init container that populates a hostPath cache is standard practice.
- Pre-pull container images and keep them small. Do not bake 140 GB of weights into the image — image layers are not efficient for that, and every push/pull pays for it.
- Store quantized checkpoints, don’t quantize at startup.
- Persist
torch.compile/ TensorRT artifacts keyed by (model, GPU arch, shapes, version). - Measure and alarm on cold-start duration. It silently regresses.
- For multi-model platforms, model load time drives your placement and eviction policy (Section XII.04). An LRU cache of loaded models with a load cost of 8 minutes behaves very differently from one with a load cost of 30 seconds.
9. Common mistakes#
Loading from object storage on every pod start. Costs 10 minutes and money per start.
Single-threaded loading. One stream to S3 gets ~200 MB/s. Parallelism is free performance.
Using pickle-based .bin files. Slower, and a genuine security risk (arbitrary code
execution on load) if the weights are not fully trusted.
Baking weights into container images. Slow pulls, huge registries, no sharing between models that differ slightly.
Quantizing or compiling at startup. Move it offline.
Ignoring the page cache. Restarting on a warm node should be fast; if it isn’t, something is dropping caches or the file isn’t being mmap’d.
Forgetting host RAM. Loading a 140 GB model may need 140 GB of host RAM transiently unless you stream shard-by-shard into the GPU.
10. Hands-on exercise#
A. Measure your storage tiers. Use dd (with and without iflag=direct) and fio to
measure sequential read bandwidth from your disk and from page cache. Record in numbers.md.
B. Time a real model load, phase by phase. Instrument: process start → import complete → weights read → on GPU → first token. Where does the time go? Draw a bar chart.
C. Parallelize. Load a sharded safetensors model with 1, 2, 4, 8 reader threads. Plot load time vs threads. Where does it saturate, and what is the limiting resource?
D. Warm vs cold. Drop caches (on a test machine), time a load. Then load again immediately. Report the ratio.
11. Interview questions#
- Break down cold-start time for a 70B model. Which phase dominates and why?
- Why is safetensors preferred over pickle-based formats? Give both a security and a performance reason.
- How does tensor parallelism change model loading time?
- What is GPUDirect Storage and when is it worth the complexity?
- Your autoscaler adds replicas but latency doesn’t improve for 10 minutes. Explain and give three mitigations.
- Why shouldn’t you bake model weights into a container image?
12. Further reading#
- [REFERENCE] safetensors documentation and format spec
- [REFERENCE] NVIDIA GPUDirect Storage documentation
- [REFERENCE]
fiofor storage benchmarking;vmtouchfor page cache control - Next: 08 — Networking fundamentals