Below the API

Threads, Processes, and Context Switching

Foundations Beginner 1h Difficulty 2/5

Prerequisites 01


1. What is it?#

  • Process — an isolated address space with its own memory, file descriptors, and at least one thread. Isolation is strong; communication is expensive.
  • Thread — an execution context (registers, stack, program counter) sharing the process’s address space. Communication is free; isolation is nonexistent.
  • Context switch — the kernel saving one thread’s state and restoring another’s. Costs 1-10 µs directly, plus cache/TLB pollution that can cost much more.

For inference: the engine is typically one process per GPU (or per TP rank), with threads for HTTP handling, tokenization, and the engine loop — and Python’s GIL complicating all of it.


2. Why does it exist?#

Because you have more work than cores, and because work has different isolation and communication needs. Processes for fault isolation and for escaping the GIL; threads for cheap sharing of the big data structures (like a KV cache manager) that you cannot afford to copy.


3. Simple analogy#

A workshop. A process is a separate workshop with its own tools and locked door — safe, but lending a tool means walking it across town. A thread is a colleague in your workshop — instant sharing, and also instant ability to knock over your work.

A context switch is being interrupted mid-task: you must write down where you were, and when you come back your workbench has been rearranged (cold caches).


4. Tiny example#

Two threads of CPU-bound work, in Go:

// threads.go — CPU-bound work on one goroutine, then on two.
package main

import (
	"fmt"
	"sync"
	"time"
)

//go:noinline
func cpuWork(n int) (x int) {
	for i := 0; i < n; i++ {
		x += i & 3
	}
	return x
}

func main() {
	const n = 2_000_000_000

	// Serial
	t0 := time.Now()
	cpuWork(n)
	cpuWork(n)
	serial := time.Since(t0)

	// Parallel: goroutines are scheduled onto real OS threads, one per core
	t0 = time.Now()
	var wg sync.WaitGroup
	for i := 0; i < 2; i++ {
		wg.Add(1)
		go func() { defer wg.Done(); cpuWork(n) }()
	}
	wg.Wait()
	parallel := time.Since(t0)

	fmt.Printf("serial %v   2 goroutines %v   speedup %.2fx\n",
		serial.Round(time.Millisecond), parallel.Round(time.Millisecond), float64(serial)/float64(parallel))
}

In Go the speedup is ≈ 2.0x: the runtime schedules goroutines onto real OS threads, one per core, and they genuinely run at the same time.

Write the same program with Python’s threading module and, on CPython ≤3.12, the speedup is ≈ 1.0x (sometimes worse). Python threads do not run Python bytecode in parallel. Only one thread holds the Global Interpreter Lock (GIL) at a time. This matters to you even as a Go engineer, because most inference engines (vLLM, TGI, SGLang) are Python processes, and the GIL shapes their architecture.

Inside such an engine, this still runs in parallel:

import torch
# GIL is RELEASED during this call — the C++ code runs in parallel
y = big_tensor @ big_matrix

Rule: the GIL is released around C extension calls, I/O, and CUDA launches. So Python threads are fine for an inference server’s I/O and for overlapping GPU work, and useless for CPU-bound Python logic. That single fact explains most of the threading design in vLLM, TGI, and friends.


5. Technical explanation#

Costs#

Thread creation        ~10-50 µs
Process creation(fork) ~100 µs - 1 ms
Context switch         1-10 µs direct
  + cache pollution    up to 100s of µs indirect (refilling L1/L2, TLB)
Mutex uncontended      ~20 ns
Mutex contended        ~1-10 µs (may involve a syscall + sleep)
Atomic increment       ~5-20 ns (worse under contention)
Pipe/socket IPC        ~5-20 µs round trip
Shared memory          ~free after setup

The thread model of an inference server#

A typical vLLM-style server:

Process: API server (Python, asyncio)
  ├─ event loop thread          HTTP, SSE streaming
  ├─ tokenizer thread pool      releases GIL (Rust tokenizers)
  └─ IPC to engine process

Process: Engine (one per TP rank)
  ├─ main loop thread           scheduler + kernel launches
  ├─ CUDA driver threads        internal
  └─ NCCL threads               collective communication

Why separate processes? Two reasons: to escape the GIL (the engine loop must not be blocked by HTTP parsing), and because tensor-parallel ranks must be separate processes anyway (each needs its own CUDA context and its own NCCL rank).

Oversubscription — the classic self-inflicted wound#

32-core machine, running:
  PyTorch intra-op threads:    32
  OpenMP (via MKL):            32
  Tokenizer pool:              16
  HTTP workers:                 8
  ────────────────────────────────
  88 runnable threads on 32 cores → constant context switching

Symptom: high system CPU time, high cs in vmstat, throughput lower than with fewer threads. Fix: set them explicitly.

export OMP_NUM_THREADS=8
export MKL_NUM_THREADS=8
export TOKENIZERS_PARALLELISM=false   # when you already parallelize at the request level
python -c "import torch; torch.set_num_threads(8)"

Async vs threads for the API layer#

An inference server’s HTTP layer is I/O-bound and long-lived (streaming responses last seconds-to-minutes). Thread-per-connection would need thousands of threads. asyncio (or Rust’s tokio, or Go goroutines) handles this with one or a few threads. This is why almost every modern inference server’s front end is async.


6. Under the hood#

Watch it happen:

# Context switches per second (cs column)
vmstat 1

# Per-process voluntary/involuntary switches
pidstat -w -p $(pgrep -f vllm) 1

# Thread list with CPU usage
top -H -p $(pgrep -f vllm)

# What are threads waiting on?
cat /proc/<pid>/task/*/stack   2>/dev/null   # kernel stacks
py-spy dump --pid <pid>                       # Python stacks of ALL threads

py-spy dump is the single most useful command for diagnosing a stuck or slow Python inference server. Learn it now; you will use it in Section X.

Interpreting: voluntary switches mean the thread blocked (I/O, lock, sleep) — usually fine. Involuntary switches mean the scheduler preempted it — high counts mean oversubscription.


7. Performance implications#

  • The engine loop is latency-critical. If it gets descheduled for 5 ms, every request’s ITL suffers by 5 ms. Give it a dedicated core; consider SCHED_FIFO in extreme cases (file 10).
  • GIL contention adds jitter. A CPU-heavy Python callback (e.g. a custom stopping criterion) holds the GIL and delays the engine loop.
  • Lock contention on the scheduler’s data structures shows up at large batch. Keep critical sections short.
  • Thread pools for tokenization should be sized to the CPU budget, not to os.cpu_count() (which lies inside containers — see file 12).

8. Production implications#

  • One engine process per GPU (or per TP group). Do not try to serve two models from one Python process expecting parallelism.
  • Set all the thread-count environment variables explicitly in your container image. Relying on defaults inside containers is a reliable way to oversubscribe.
  • Pin the engine thread to a NUMA-local core (file 03) on multi-socket hosts.
  • Use py-spy in production. It attaches without restarting and without instrumenting. Include it in your image.
  • Watch involuntary context switches as a health metric; a spike means CPU contention that will show up as ITL jitter.

9. Common mistakes#

Using Python threads for CPU-bound work. The GIL makes it pointless. Use processes, or push the work into C/Rust.

Leaving OMP_NUM_THREADS unset in a container. OpenMP sees the host’s core count, spawns 128 threads inside a 4-CPU cgroup, and you get pathological throttling.

Blocking the asyncio event loop. One synchronous time.sleep() or a heavy JSON dump in a coroutine stalls every streaming connection on that loop.

Creating threads per request. At 1,000 req/s that’s 1,000 thread creations/sec plus scheduler pressure. Use pools.

Assuming os.cpu_count() reflects your quota. It doesn’t in containers. Read the cgroup (file 12).


10. Hands-on exercise#

A. Scaling, and where it stops. Run the example in section 4 with 1, 2, 4, … goroutines up to twice your core count. Plot the speedup. Where does it flatten, and why? Then set GOMAXPROCS=1 and run it again: you have just reproduced what the GIL does to a Python engine.

B. Oversubscription. Run a CPU-heavy workload with OMP_NUM_THREADS = 1, 4, 8, 16, 32, 64 on an N-core machine. Plot throughput. Find the peak. Watch vmstat 1’s cs column at each setting.

C. Profile a real server’s threads. Start any Python inference server under load. Run top -H, pidstat -w, and py-spy dump. Identify: which thread is the engine loop, which is handling HTTP, and where CPU is actually going.

D. Measure context-switch cost. Write a ping-pong benchmark between two threads using a condition variable, 1M round trips. Compute µs per switch.


11. Interview questions#

  1. What is the GIL and when does it not block parallelism?
  2. Why do inference servers use separate processes for the API layer and the engine?
  3. What is oversubscription and how does it manifest?
  4. What is the difference between voluntary and involuntary context switches, diagnostically?
  5. Why is the API layer of an inference server usually async rather than thread-per-request?
  6. How would you debug a Python inference server that has become unresponsive?

12. Further reading#

↑↓ navigate ↵ open