Below the API

Containers, Namespaces, and cgroups

Foundations Intermediate 1h Difficulty 3/5

Prerequisites 05, 06, 10


1. What is it?#

A container is not a virtual machine. It is a normal Linux process with three things applied:

namespaces  →  what the process can SEE   (pids, mounts, network, users)
cgroups     →  what the process can USE   (cpu, memory, io, devices)
layered fs  →  where its files come FROM  (overlayfs)

For GPU workloads, add a fourth: a device plugin / container runtime hook that injects the GPU device nodes and driver libraries.


2. Why does it exist as a topic?#

Because containers lie to your process about the machine, and inference stacks make decisions based on what they see. os.cpu_count() returns the host’s core count; OpenMP spawns that many threads; the cgroup quota is 4 CPUs; you get catastrophic throttling. This single mismatch is one of the most common performance bugs in containerized ML.


3. Simple analogy#

A serviced office inside a large building. Namespaces are the frosted glass: you see only your own room and believe it’s the whole company. cgroups are the metered utilities: you may see the building’s 500 kW supply, but your room’s breaker trips at 2 kW.

The bug is when you plan your work based on what you can see rather than what you’re allowed to use.


4. Tiny example#

docker run --rm --cpus=2 python:3.11 python -c "
import os
print('os.cpu_count():', os.cpu_count())
print('affinity     :', len(os.sched_getaffinity(0)))
print('cgroup quota :', open('/sys/fs/cgroup/cpu.max').read().strip())
"

Output on a 64-core host:

os.cpu_count(): 64          ← the lie
affinity     : 64           ← also the lie (quota != affinity)
cgroup quota : 200000 100000  ← the truth: 2.0 CPUs

Now the damage:

docker run --rm --cpus=2 pytorch/pytorch python -c "
import torch, time
print('torch threads:', torch.get_num_threads())   # often 64!
x = torch.randn(2000,2000)
t=time.time(); [x@x for _ in range(20)]; print('time', time.time()-t)
"

64 threads sharing a 2-CPU quota: constant throttling, context switching, and often slower than 2 threads. Fix:

docker run --rm --cpus=2 -e OMP_NUM_THREADS=2 -e MKL_NUM_THREADS=2 ...

Set thread counts from the quota, explicitly, in your entrypoint. A small helper:

#!/bin/sh
# entrypoint.sh
QUOTA=$(cat /sys/fs/cgroup/cpu.max 2>/dev/null | awk '{if ($1=="max") print 0; else print int($1/$2)}')
[ "${QUOTA:-0}" -gt 0 ] && export OMP_NUM_THREADS=$QUOTA MKL_NUM_THREADS=$QUOTA
exec "$@"

5. Technical explanation#

Namespaces#

NamespaceIsolatesInference relevance
PIDprocess IDsyour process is PID 1; signal handling differs
Mountfilesystem viewmodel volumes, /dev/shm size
Networkinterfaces, portspod networking, NCCL between pods
IPCshared memory, semaphores/dev/shm size matters a lot
UTShostname—
Useruid mappingrootless containers, device access
Cgroupcgroup view—

/dev/shm deserves attention. Docker defaults it to 64 MB. PyTorch DataLoader workers, NCCL shared-memory transports, and multi-process engines all use it. Symptom: Bus error, No space left on device on /dev/shm, or NCCL init failures.

docker run --shm-size=16g ...
# Kubernetes: an emptyDir with medium: Memory mounted at /dev/shm

cgroups v2 controllers#

/sys/fs/cgroup/
├── cpu.max              "quota period" in µs — e.g. "400000 100000" = 4 CPUs
├── cpu.stat             usage_usec, nr_throttled, throttled_usec  ← THE diagnostic
├── cpu.weight           relative share (1-10000), used when no quota
├── cpuset.cpus          explicit CPU pinning (guaranteed QoS in K8s)
├── memory.max           hard limit — exceeding = OOM kill
├── memory.high          soft limit — throttles reclaim instead of killing
├── memory.current       current usage
├── memory.stat          detailed breakdown
├── io.max               block I/O limits
└── pids.max             process count limit

The CPU throttling mechanism (why it hurts so much)#

CFS bandwidth control gives you quota microseconds of CPU per period (default 100 ms). Consume it early and you are frozen until the next period.

quota = 200ms per 100ms period (2 CPUs)

With 8 threads all runnable:
  t=0-25ms:   8 threads × 25 ms = 200 ms consumed. Quota exhausted.
  t=25-100ms: ALL THREADS FROZEN.                  ← 75 ms stall
  t=100ms:    period resets

Average CPU used: 2.0 (as configured)
p99 latency added: 75 ms

Configured correctly on average, catastrophic on tail latency. For an inference engine loop this is a 75 ms ITL spike, repeatedly.

Mitigations, in order of preference:

  1. Set thread counts to match the quota (fewer threads → smoother consumption).
  2. Use cpu.weight / K8s requests without limits, so you get shares rather than a hard quota.
  3. Use Guaranteed QoS with integer CPUs and the CPU Manager static policy → you get a dedicated cpuset with no quota throttling at all. This is the right answer for latency- critical inference pods.

GPUs in containers#

docker run --gpus all ...          # requires the NVIDIA Container Toolkit

The toolkit’s runtime hook injects /dev/nvidia* device nodes and bind-mounts driver libraries into the container. The container image must contain a CUDA runtime compatible with the host driver — the driver stays on the host, the toolkit and userspace libraries come from the image.

Host:      NVIDIA driver 550.x  →  supports CUDA up to 12.4
Container: CUDA runtime 12.1     →  OK (backward compatible)
Container: CUDA runtime 12.6     →  FAILS unless forward-compat packages installed

In Kubernetes, the NVIDIA device plugin advertises nvidia.com/gpu as a schedulable resource (Section XII.03).


6. Under the hood#

# Is my container being throttled? (run this INSIDE the container)
cat /sys/fs/cgroup/cpu.stat
# nr_throttled: 4821      ← nonzero and climbing = you have a problem
# throttled_usec: 91234567

# Memory pressure
cat /sys/fs/cgroup/memory.events    # low, high, max, oom, oom_kill counters

# What device nodes do I have?
ls -l /dev/nvidia*

# Effective limits summary
for f in cpu.max memory.max pids.max; do echo "$f: $(cat /sys/fs/cgroup/$f)"; done

Export nr_throttled and throttled_usec as Prometheus metrics. A dashboard panel showing throttling correlated with p99 ITL saves hours of confused debugging.


7. Performance implications#

MisconfigurationImpact
Threads > CPU quota2-10x slowdown from throttling + switching
CPU limit on latency-critical pod10-100 ms periodic stalls
/dev/shm too smallNCCL failures, dataloader crashes
Memory limit too tightOOM kill (137) with no traceback
Non-guaranteed QoSno cpuset; NUMA-random placement; throttling
Huge container imagesslow cold start (Section II.07)

8. Production implications#

  • Use Guaranteed QoS for inference pods: requests == limits, integer CPUs, and enable the kubelet CPU Manager static policy plus the Topology Manager single-numa-node policy. You then get dedicated cores, NUMA-local memory, and NUMA-local GPUs.
  • Set every thread-count env var in the image entrypoint, derived from the quota.
  • Size /dev/shm generously (8-16 GB) for multi-process engines.
  • Set memory limits with headroom for the page cache used by mmap’d weights. Note that page cache counts against memory.max in cgroup v2 but is reclaimable.
  • Pin the CUDA/driver compatibility matrix in your image build and test it in CI.
  • Alert on nr_throttled.

9. Common mistakes#

Leaving OMP_NUM_THREADS unset. The single most common containerized-ML performance bug.

Setting CPU limits on inference pods. Causes exactly the throttling described above. Use requests, or Guaranteed QoS.

Default 64 MB /dev/shm. Breaks NCCL and multiprocessing in confusing ways.

Assuming the container sees the GPU topology correctly. It sees the GPUs assigned to it, renumbered. CUDA_VISIBLE_DEVICES inside a container refers to the assigned set.

Baking weights into images. Covered in file 07; also makes images too large to pull quickly.

Ignoring cgroup v1 vs v2 differences in tooling and paths.


10. Hands-on exercise#

A. Demonstrate the lie. Run the example in section 4 on your machine. Then write the entrypoint helper that derives thread counts from the quota, and prove it fixes the slowdown.

B. Induce throttling. Run a CPU-heavy container with --cpus=1 and 8 threads. Watch nr_throttled and measure per-operation latency percentiles. Then set threads to 1 and compare p99.

C. Break /dev/shm. Run a multi-process PyTorch job with the default 64 MB shm and observe the failure. Fix with --shm-size.

D. Guaranteed QoS. On a Kubernetes cluster (kind/minikube is fine for the config part), create a Guaranteed pod and a Burstable pod. Inspect cpuset.cpus and cpu.max inside each.


11. Interview questions#

  1. What are the three mechanisms that make a container, and what does each do?
  2. Why does os.cpu_count() mislead inside a container, and what should you use?
  3. Explain CFS bandwidth throttling and why it hurts tail latency more than average latency.
  4. Why do we prefer Guaranteed QoS for inference pods?
  5. What is /dev/shm used for in an inference stack and what happens if it’s too small?
  6. How does a container get access to a GPU, and what compatibility constraint applies?

12. Further reading#

  • [REFERENCE] Kernel docs: cgroup-v2, namespaces(7)
  • [REFERENCE] NVIDIA Container Toolkit documentation
  • [REFERENCE] Kubernetes CPU Manager and Topology Manager docs
  • Next: 13 — Syscalls and profiling basics

↑↓ navigate ↵ open