1. What is it?#
A container is not a virtual machine. It is a normal Linux process with three things applied:
namespaces → what the process can SEE (pids, mounts, network, users)
cgroups → what the process can USE (cpu, memory, io, devices)
layered fs → where its files come FROM (overlayfs)For GPU workloads, add a fourth: a device plugin / container runtime hook that injects the GPU device nodes and driver libraries.
2. Why does it exist as a topic?#
Because containers lie to your process about the machine, and inference stacks make decisions
based on what they see. os.cpu_count() returns the host’s core count; OpenMP spawns that many
threads; the cgroup quota is 4 CPUs; you get catastrophic throttling. This single mismatch is
one of the most common performance bugs in containerized ML.
3. Simple analogy#
A serviced office inside a large building. Namespaces are the frosted glass: you see only your own room and believe it’s the whole company. cgroups are the metered utilities: you may see the building’s 500 kW supply, but your room’s breaker trips at 2 kW.
The bug is when you plan your work based on what you can see rather than what you’re allowed to use.
4. Tiny example#
docker run --rm --cpus=2 python:3.11 python -c "
import os
print('os.cpu_count():', os.cpu_count())
print('affinity :', len(os.sched_getaffinity(0)))
print('cgroup quota :', open('/sys/fs/cgroup/cpu.max').read().strip())
"
Output on a 64-core host:
os.cpu_count(): 64 ← the lie
affinity : 64 ← also the lie (quota != affinity)
cgroup quota : 200000 100000 ← the truth: 2.0 CPUsNow the damage:
docker run --rm --cpus=2 pytorch/pytorch python -c "
import torch, time
print('torch threads:', torch.get_num_threads()) # often 64!
x = torch.randn(2000,2000)
t=time.time(); [x@x for _ in range(20)]; print('time', time.time()-t)
"
64 threads sharing a 2-CPU quota: constant throttling, context switching, and often slower than 2 threads. Fix:
docker run --rm --cpus=2 -e OMP_NUM_THREADS=2 -e MKL_NUM_THREADS=2 ...
Set thread counts from the quota, explicitly, in your entrypoint. A small helper:
#!/bin/sh
# entrypoint.sh
QUOTA=$(cat /sys/fs/cgroup/cpu.max 2>/dev/null | awk '{if ($1=="max") print 0; else print int($1/$2)}')
[ "${QUOTA:-0}" -gt 0 ] && export OMP_NUM_THREADS=$QUOTA MKL_NUM_THREADS=$QUOTA
exec "$@"
5. Technical explanation#
Namespaces#
| Namespace | Isolates | Inference relevance |
|---|---|---|
| PID | process IDs | your process is PID 1; signal handling differs |
| Mount | filesystem view | model volumes, /dev/shm size |
| Network | interfaces, ports | pod networking, NCCL between pods |
| IPC | shared memory, semaphores | /dev/shm size matters a lot |
| UTS | hostname | — |
| User | uid mapping | rootless containers, device access |
| Cgroup | cgroup view | — |
/dev/shm deserves attention. Docker defaults it to 64 MB. PyTorch DataLoader workers, NCCL
shared-memory transports, and multi-process engines all use it. Symptom: Bus error,
No space left on device on /dev/shm, or NCCL init failures.
docker run --shm-size=16g ...
# Kubernetes: an emptyDir with medium: Memory mounted at /dev/shm
cgroups v2 controllers#
/sys/fs/cgroup/
├── cpu.max "quota period" in µs — e.g. "400000 100000" = 4 CPUs
├── cpu.stat usage_usec, nr_throttled, throttled_usec ← THE diagnostic
├── cpu.weight relative share (1-10000), used when no quota
├── cpuset.cpus explicit CPU pinning (guaranteed QoS in K8s)
├── memory.max hard limit — exceeding = OOM kill
├── memory.high soft limit — throttles reclaim instead of killing
├── memory.current current usage
├── memory.stat detailed breakdown
├── io.max block I/O limits
└── pids.max process count limit
The CPU throttling mechanism (why it hurts so much)#
CFS bandwidth control gives you quota microseconds of CPU per period (default 100 ms).
Consume it early and you are frozen until the next period.
quota = 200ms per 100ms period (2 CPUs)
With 8 threads all runnable:
t=0-25ms: 8 threads × 25 ms = 200 ms consumed. Quota exhausted.
t=25-100ms: ALL THREADS FROZEN. ← 75 ms stall
t=100ms: period resets
Average CPU used: 2.0 (as configured)
p99 latency added: 75 msConfigured correctly on average, catastrophic on tail latency. For an inference engine loop this is a 75 ms ITL spike, repeatedly.
Mitigations, in order of preference:
- Set thread counts to match the quota (fewer threads → smoother consumption).
- Use
cpu.weight/ K8s requests without limits, so you get shares rather than a hard quota. - Use Guaranteed QoS with integer CPUs and the CPU Manager
staticpolicy → you get a dedicated cpuset with no quota throttling at all. This is the right answer for latency- critical inference pods.
GPUs in containers#
docker run --gpus all ... # requires the NVIDIA Container Toolkit
The toolkit’s runtime hook injects /dev/nvidia* device nodes and bind-mounts driver libraries
into the container. The container image must contain a CUDA runtime compatible with the
host driver — the driver stays on the host, the toolkit and userspace libraries come from
the image.
Host: NVIDIA driver 550.x → supports CUDA up to 12.4
Container: CUDA runtime 12.1 → OK (backward compatible)
Container: CUDA runtime 12.6 → FAILS unless forward-compat packages installedIn Kubernetes, the NVIDIA device plugin advertises nvidia.com/gpu as a schedulable resource
(Section XII.03).
6. Under the hood#
# Is my container being throttled? (run this INSIDE the container)
cat /sys/fs/cgroup/cpu.stat
# nr_throttled: 4821 ← nonzero and climbing = you have a problem
# throttled_usec: 91234567
# Memory pressure
cat /sys/fs/cgroup/memory.events # low, high, max, oom, oom_kill counters
# What device nodes do I have?
ls -l /dev/nvidia*
# Effective limits summary
for f in cpu.max memory.max pids.max; do echo "$f: $(cat /sys/fs/cgroup/$f)"; done
Export nr_throttled and throttled_usec as Prometheus metrics. A dashboard panel showing
throttling correlated with p99 ITL saves hours of confused debugging.
7. Performance implications#
| Misconfiguration | Impact |
|---|---|
| Threads > CPU quota | 2-10x slowdown from throttling + switching |
| CPU limit on latency-critical pod | 10-100 ms periodic stalls |
/dev/shm too small | NCCL failures, dataloader crashes |
| Memory limit too tight | OOM kill (137) with no traceback |
| Non-guaranteed QoS | no cpuset; NUMA-random placement; throttling |
| Huge container images | slow cold start (Section II.07) |
8. Production implications#
- Use Guaranteed QoS for inference pods: requests == limits, integer CPUs, and enable the
kubelet CPU Manager
staticpolicy plus the Topology Managersingle-numa-nodepolicy. You then get dedicated cores, NUMA-local memory, and NUMA-local GPUs. - Set every thread-count env var in the image entrypoint, derived from the quota.
- Size
/dev/shmgenerously (8-16 GB) for multi-process engines. - Set memory limits with headroom for the page cache used by mmap’d weights. Note that page
cache counts against
memory.maxin cgroup v2 but is reclaimable. - Pin the CUDA/driver compatibility matrix in your image build and test it in CI.
- Alert on
nr_throttled.
9. Common mistakes#
Leaving OMP_NUM_THREADS unset. The single most common containerized-ML performance bug.
Setting CPU limits on inference pods. Causes exactly the throttling described above. Use requests, or Guaranteed QoS.
Default 64 MB /dev/shm. Breaks NCCL and multiprocessing in confusing ways.
Assuming the container sees the GPU topology correctly. It sees the GPUs assigned to it,
renumbered. CUDA_VISIBLE_DEVICES inside a container refers to the assigned set.
Baking weights into images. Covered in file 07; also makes images too large to pull quickly.
Ignoring cgroup v1 vs v2 differences in tooling and paths.
10. Hands-on exercise#
A. Demonstrate the lie. Run the example in section 4 on your machine. Then write the entrypoint helper that derives thread counts from the quota, and prove it fixes the slowdown.
B. Induce throttling. Run a CPU-heavy container with --cpus=1 and 8 threads. Watch
nr_throttled and measure per-operation latency percentiles. Then set threads to 1 and compare
p99.
C. Break /dev/shm. Run a multi-process PyTorch job with the default 64 MB shm and observe
the failure. Fix with --shm-size.
D. Guaranteed QoS. On a Kubernetes cluster (kind/minikube is fine for the config part),
create a Guaranteed pod and a Burstable pod. Inspect cpuset.cpus and cpu.max inside each.
11. Interview questions#
- What are the three mechanisms that make a container, and what does each do?
- Why does
os.cpu_count()mislead inside a container, and what should you use? - Explain CFS bandwidth throttling and why it hurts tail latency more than average latency.
- Why do we prefer Guaranteed QoS for inference pods?
- What is
/dev/shmused for in an inference stack and what happens if it’s too small? - How does a container get access to a GPU, and what compatibility constraint applies?
12. Further reading#
- [REFERENCE] Kernel docs: cgroup-v2, namespaces(7)
- [REFERENCE] NVIDIA Container Toolkit documentation
- [REFERENCE] Kubernetes CPU Manager and Topology Manager docs
- Next: 13 — Syscalls and profiling basics