PidokuInfra

Backpressure, Timeouts, and Retries

Intermediate Advanced 1h 15m Difficulty 4/5 Topic 05 of 14

Prerequisites I.10, 03


1. What is it?#

The three mechanisms that determine whether your service degrades gracefully or collapses.

BACKPRESSURE  signalling upstream that you're full, so it slows down
TIMEOUTS      giving up on work that has taken too long
RETRIES       trying again after a failure — and the discipline not to

For LLM serving, all three need different settings and different reasoning than for ordinary services, because requests are long-lived, expensive, and non-idempotent.


2. Why it exists#

Because the default behavior of a naive system under overload is not “slower” — it’s collapse, and often a collapse it cannot recover from.

Overload → queue grows → latency grows → clients time out → clients retry
        → load increases → queue grows more → ...

And critically, in LLM serving:
  the original request is STILL RUNNING, holding a KV slot, generating tokens
  nobody will read.

The system is now doing 3x the work for 0x the value. Arrival rate returning to normal doesn’t help, because the system is busy with abandoned work. This is a metastable failure.


3. Simple analogy#

A ticket office with a queue out the door.

Without backpressure: everyone joins the queue. After an hour, people at the back give up and rejoin at the front (retry), making the queue longer. The clerks serve people who left twenty minutes ago.

With backpressure: a sign at the door says “queue full, back at 3pm.” The queue stays short, served people are happy, and turned-away people know to come back rather than waiting three hours for nothing.


4. Tiny example — the retry storm#

Go
// retries.go — an overloaded server, with and without retries and cancellation.
package main

import (
	"container/heap"
	"fmt"
	"math"
	"math/rand"
)

type event struct {
	t       float64
	attempt int
}
type events []event

func (h events) Len() int           { return len(h) }
func (h events) Less(i, j int) bool { return h[i].t < h[j].t }
func (h events) Swap(i, j int)      { h[i], h[j] = h[j], h[i] }
func (h *events) Push(x any)        { *h = append(*h, x.(event)) }
func (h *events) Pop() any          { old := *h; x := old[len(old)-1]; *h = old[:len(old)-1]; return x }

type floats []float64

func (h floats) Len() int           { return len(h) }
func (h floats) Less(i, j int) bool { return h[i] < h[j] }
func (h floats) Swap(i, j int)      { h[i], h[j] = h[j], h[i] }
func (h *floats) Push(x any)        { *h = append(*h, x.(float64)) }
func (h *floats) Pop() any          { old := *h; x := old[len(old)-1]; *h = old[:len(old)-1]; return x }

func simulate(arrivalRate float64, retries, cancellation bool, simTime float64) (completed int, wasted float64) {
	rng := rand.New(rand.NewSource(1))
	slots := make(floats, 32) // when each of 32 slots becomes free
	var q events
	for t := rng.ExpFloat64() / arrivalRate; t < simTime; t += rng.ExpFloat64() / arrivalRate {
		heap.Push(&q, event{t, 0}) // the original requests
	}
	for q.Len() > 0 {
		ev := heap.Pop(&q).(event)
		if ev.t >= simTime {
			break
		}
		deadline := ev.t + 30 // the client gives up after 30 s
		start := math.Max(ev.t, slots[0])
		service := math.Exp(1.0 + 0.8*rng.NormFloat64())

		var busyUntil float64
		switch {
		case start+service <= deadline: // finished while the client was still listening
			busyUntil = start + service
			completed++
		case cancellation: // server notices the client left and aborts early
			busyUntil = math.Max(start, deadline)
			wasted += busyUntil - start
		default: // server runs to completion for nobody
			busyUntil = start + service
			wasted += service
		}
		slots[0] = busyUntil
		heap.Fix(&slots, 0)

		if start+service > deadline && retries && ev.attempt < 2 {
			heap.Push(&q, event{deadline, ev.attempt + 1}) // the client tries again
		}
	}
	return completed, wasted
}

func main() {
	for _, c := range []struct {
		label           string
		retries, cancel bool
	}{{"no retries, no cancel", false, false}, {"retries, no cancel", true, false}, {"retries + cancel", true, true}} {
		done, wasted := simulate(12, c.retries, c.cancel, 300)
		fmt.Printf("%-25s completed=%5d wasted_gpu_sec=%8.0f\n", c.label, done, wasted)
	}
}

The output:

no retries, no cancel     completed=  645 wasted_gpu_sec=   11599
retries, no cancel        completed=  639 wasted_gpu_sec=   30218   ← WORSE
retries + cancel          completed= 1906 wasted_gpu_sec=    6684   ← best

Retries without cancellation make things worse. Retries with cancellation help. The combination is what matters, and most systems implement only the first half.


5. Technical explanation#

Backpressure — the mechanisms#

1. REJECT (503 + Retry-After)
   The primary mechanism. Fast, explicit, honest.

2. BOUNDED QUEUES at every layer
   TCP backlog, API queue, engine waiting queue. Each bounded.

3. FLOW CONTROL on streams
   If a client isn't reading, stop generating. TCP backpressure propagates
   naturally IF you check it.

4. RATE LIMITING (token bucket, per tenant)
   Prevents one client from consuming the fleet.

5. CONCURRENCY LIMITS per tenant
   Simpler than rate limiting and directly bounds slot occupancy.

Mechanism 3 is LLM-specific and often missed. If a client stops reading the stream, the socket buffer fills, writes block, and — if you’re not checking — you either block the event loop or buffer unboundedly in userspace. Check is_disconnected() and handle write backpressure.

Timeouts — you need several#

TIMEOUT               TYPICAL VALUE          PURPOSE
connect               1-5 s                  network reachability
TLS handshake         2-5 s                  
queue wait            5-30 s                 don't queue forever
time to first token   30-60 s                catch prefill problems
inter-token stall     10-30 s                catch engine hangs   ← LLM-specific
total generation      5-15 min               absolute bound
idle connection       > max generation       don't kill live streams

The inter-token stall timeout is the important one. A total-request timeout of 10 minutes is useless for detecting a wedged engine; a “no token in 30 seconds” timeout catches it in 30 seconds.

Go
var ErrEngineStalled = errors.New("engine stalled")

// streamWithStallTimeout forwards tokens, but fails if the engine goes quiet.
// The timeout is per TOKEN, not per request: a long answer is fine, a silent engine is not.
func streamWithStallTimeout(ctx context.Context, in <-chan string, out chan<- string, stall time.Duration) error {
	timer := time.NewTimer(stall)
	defer timer.Stop()
	for {
		select {
		case tok, ok := <-in:
			if !ok {
				return nil // generation finished normally
			}
			out <- tok
			timer.Reset(stall)
		case <-timer.C:
			return fmt.Errorf("%w: no token in %v", ErrEngineStalled, stall)
		case <-ctx.Done():
			return ctx.Err() // client went away
		}
	}
}

Retries — the discipline#

RULE 1: Never retry a request that may still be running.
        Cancel first, confirm cancellation, then retry.

RULE 2: Never retry after streaming has started.
        The client has partial output. A retry duplicates it.

RULE 3: Retry only idempotent-ish failures:
        connection refused, 503 with Retry-After, 429.
        NOT: timeouts (the work may be in flight), 500 (may have partial effect).

RULE 4: Budget retries.
        Cap retries at ~10% of total requests, globally. A retry budget
        prevents a partial outage from becoming a total one.

RULE 5: Exponential backoff with jitter. Always jitter.
        Without jitter, retries synchronize into waves.

RULE 6: Respect Retry-After.

Retry budget implementation:

Go
// RetryBudget allows retries only while they are a small fraction of total traffic.
type RetryBudget struct {
	mu                sync.Mutex
	ratio             float64       // e.g. 0.1: retries may add at most 10% load
	window            time.Duration // e.g. one minute
	requests, retries []time.Time
}

func (b *RetryBudget) RecordRequest(now time.Time) {
	b.mu.Lock()
	defer b.mu.Unlock()
	b.requests = append(b.requests, now)
}

func (b *RetryBudget) AllowRetry(now time.Time) bool {
	b.mu.Lock()
	defer b.mu.Unlock()
	cutoff := now.Add(-b.window)
	expire := func(ts []time.Time) []time.Time {
		for len(ts) > 0 && ts[0].Before(cutoff) {
			ts = ts[1:]
		}
		return ts
	}
	b.requests, b.retries = expire(b.requests), expire(b.retries)
	if float64(len(b.retries)) >= b.ratio*float64(max(len(b.requests), 10)) {
		return false // budget spent: fail now rather than add load
	}
	b.retries = append(b.retries, now)
	return true
}

This one class prevents most retry storms. It is 15 lines and almost nobody implements it.

Cancellation propagation#

Client disconnects
   → the ASGI server sets the disconnect flag
   → the streaming handler notices (checks between tokens)
   → calls engine.abort(request_id)
   → the scheduler removes the request from `running`
   → the KV manager frees its blocks
   → the slot is immediately available

Every link in that chain must exist. If any is missing, capacity leaks. Test it explicitly: start a generation, kill the client, and verify the KV blocks are freed within one step.


6-9. Under the hood, performance, production, mistakes#

Under the hood — measuring abandonment:

Instrument:
  requests_aborted_total{reason="client_disconnect"}
  tokens_generated_after_disconnect_total     ← pure waste; should be ~0
  abort_latency_seconds                       ← time from disconnect to slot freed

In a chat product, client_disconnect is typically 10-30% of requests. If tokens_generated_after_disconnect is large, cancellation isn’t working.

Performance: the capacity impact of cancellation:

20% of requests abandoned at 50% of their generation
→ 10% of all generated tokens are wasted
→ fixing it is a 10-11% capacity increase, for a day of work

Production:

  • Implement cancellation first. It’s the highest return-on-effort item in this file.
  • Set a stall timeout, not just a total timeout.
  • Implement a retry budget.
  • Return Retry-After on 429 and 503, and honor it in your own clients.
  • Bound every queue.
  • Test the overload path. Load-test past capacity and verify you reject cleanly rather than collapsing. Most teams never do this.
  • Alert on rejection rate, not just on latency — rejections are the healthy failure mode.

Mistakes:

  • Retries without cancellation. Actively harmful.
  • Retrying on timeout. The original may still be running.
  • Retrying after streaming started.
  • No jitter. Synchronized retry waves.
  • No retry budget. Partial outage → total outage.
  • Idle timeout shorter than generation time. Streams die at 60 seconds.
  • No stall timeout. A wedged engine holds slots forever.
  • Never testing the overload path.

10. Hands-on exercise#

A. Reproduce the storm. Run the simulation in section 4. Vary the arrival rate to find where retries-without-cancellation cause collapse. Add cancellation and re-run.

B. Implement cancellation end to end. In a real server, wire up disconnect detection → abort → KV free. Verify by killing a client mid-generation and checking that the KV block count recovers within one step.

C. Stall timeout. Implement the stall-timeout wrapper. Simulate an engine hang and verify it fires in 30 seconds rather than at the total timeout.

D. Retry budget. Implement the RetryBudget class. Simulate a partial outage (30% of requests failing) with and without the budget. Compare total load on the backend.

E. Overload test. Load-test your service at 2x capacity. Does it reject cleanly? What’s the p99 latency of served requests? Does it recover when load drops?


11. Interview questions#

  1. What is a metastable failure and how do retries cause one?
  2. Why are retries harmful without cancellation in LLM serving?
  3. What timeouts does an LLM service need? Why is a stall timeout necessary?
  4. What is a retry budget and why does it matter?
  5. Why should you never retry after streaming has started?
  6. How does cancellation propagate from a disconnected client to freed KV blocks?
  7. What is the healthy failure mode under overload, and how do you verify you have it?

12. Further reading#

  • [FUNDAMENTAL] Google SRE Book, “Handling Overload,” “Addressing Cascading Failures”
  • [FUNDAMENTAL] Bronson et al., “Metastable Failures in Distributed Systems” (HotOS 2021)
  • [ESTABLISHED] AWS Builders’ Library: “Timeouts, retries, and backoff with jitter”
  • Next: 06 — Load balancing and routing

↑↓ navigate↵ openesc close