Below the API

Project 12 — Inference Gateway

Advanced 12h Difficulty 4/5

Prerequisites Project 06, VIII.05-06, XI.01, XI.08, XI.10, XII.06-08

Build the front door: the component that decides who gets in, where they go, and what happens when a backend dies.


1. What you build#

A reverse proxy for OpenAI-compatible backends that does authentication, token-based rate limiting, load-aware and prefix-aware routing, timeouts, bounded retries, circuit breaking, streaming passthrough, usage metering per tenant, and metrics. Backends are several instances of your Project 08 server (or vLLM).

Write it in Go, Rust, or async Python — whichever you would defend in a design review.

Diagram — Request pipeline#

flowchart LR
  C["Client"] --> AU["Auth<br/>key to tenant"]
  AU --> RL["Rate limit<br/>RPM + TPM buckets"]
  RL -->|"over quota"| X429["429"]
  RL --> RT["Router<br/>prefix affinity + load bound"]
  RT --> CB{"Circuit breaker<br/>closed?"}
  CB -->|"no: pick another"| RT
  CB -->|"yes"| B1["Backend A"]
  CB -->|"yes"| B2["Backend B"]
  B1 -->|"SSE passthrough"| C
  B2 -->|"SSE passthrough"| C
  B1 -.-> MT["Usage metering + metrics"]
  B2 -.-> MT

  class AU,RL,RT,CB queue
  class X429 warn
  class B1,B2 compute
  class MT io
  class C neutral

2. Why it matters#

Engines optimize one GPU. The gateway determines whether a fleet of them behaves like a service. Its decisions are worth real money: sending a request to the replica that already holds its prefix in KV cache can cut TTFT several-fold, and a retry policy without a budget can turn one slow backend into a full outage.

For someone coming from DevOps/platform work, this is the project where existing skills pay off most directly — and where LLM traffic breaks familiar assumptions (requests last minutes, cost varies 1000× between requests, round-robin is actively harmful).


3. Read first#


4. Spec#

Auth          API key -> tenant ; unknown key -> 401
Rate limits   per tenant: requests/min AND tokens/min (token bucket)
              reserve prompt_tokens + max_tokens up front, refund unused on completion
              concurrent-stream cap per tenant
Routing       model name -> backend pool
              policies: round_robin | least_in_flight | least_queue (scrape backend metrics)
                        | prefix_affinity (consistent hash of first N prompt tokens,
                          with load-based spill-over)
Resilience    connect timeout, TTFT timeout, idle (inter-token) timeout, total deadline
              retry ONLY before the first byte is sent to the client; retry budget ≤ 10%
              circuit breaker per backend (closed / open / half-open)
              active health checks on /health/ready
Streaming     pass SSE through unbuffered; propagate client disconnect upstream
Metering      per tenant: prompt/completion tokens, requests, errors (from `usage`)
Observability Prometheus metrics + a request ID header propagated to backends

5. Milestones#

  1. Plain proxy with streaming passthrough. Verify TTFT through the gateway is within 1-2 ms of direct.
  2. Auth and tenants from a config file.
  3. Rate limiting. Token bucket on tokens/min. Return 429 with Retry-After and remaining-quota headers.
  4. Routing policies. Implement all four behind one interface.
  5. Routing experiment. 4 backends; a workload where 80% of requests share one of 20 long system prompts. Compare policies on TTFT p50/p95 and backend prefix-cache hit rate.
  6. Failure injection. Kill a backend mid-stream; make one backend 10× slower; return 500s from one. Verify breaker, retries, and that the retry budget holds.
  7. Disconnect propagation. Client leaves → upstream request cancelled → backend slot freed. Confirm in backend metrics.
  8. Usage report per tenant.

6. Starter skeleton#

type Backend struct {
    URL       string
    inFlight  atomic.Int64
    breaker   *Breaker
    queueLen  atomic.Int64 // scraped from backend /metrics
}

// Prefix affinity with bounded load: stick to the hashed backend unless it is overloaded.
func (p *Pool) pick(prefixKey uint64) *Backend {
    order := p.ring.Lookup(prefixKey, len(p.backends)) // consistent-hash preference order
    avg := p.avgInFlight()
    for _, b := range order {
        if b.breaker.Allow() && float64(b.inFlight.Load()) <= 1.25*avg+1 {
            return b
        }
    }
    return p.leastInFlight()
}

type Bucket struct{ mu sync.Mutex; tokens, rate, burst float64; last time.Time }

func (b *Bucket) Take(n float64) (ok bool, retryAfter time.Duration) {
    b.mu.Lock(); defer b.mu.Unlock()
    now := time.Now()
    b.tokens = math.Min(b.burst, b.tokens+b.rate*now.Sub(b.last).Seconds())
    b.last = now
    if b.tokens >= n { b.tokens -= n; return true, 0 }
    return false, time.Duration((n-b.tokens)/b.rate*1e9)
}
func (b *Bucket) Refund(n float64) { b.mu.Lock(); b.tokens = math.Min(b.burst, b.tokens+n); b.mu.Unlock() }

7. What to measure#

MeasurementExpectation to write down first
Added latency of the gateway (p50/p99)< 1-2 ms; streaming unbuffered
TTFT p95: round-robin vs least-in-flight vs prefix-affinityAffinity wins on shared-prefix load
Backend prefix-cache hit rate per policyRound-robin ≈ 1/N of affinity
Load imbalance under affinity (max/mean in-flight)Bounded by the spill-over rule
Error rate seen by clients when one backend diesBrief blip, then zero
Total upstream requests / client requests during a brownout≤ 1.1 with a retry budget
Fairness: noisy tenant at 10× quota vs quiet tenant’s p95Quiet tenant unaffected

8. Done when#

  • Streaming passes through with negligible added TTFT.
  • A tenant exceeding tokens/min gets 429s; others are unaffected.
  • You have a table comparing routing policies on TTFT and cache hit rate.
  • Chaos tests pass: backend kill, slow backend, error backend.
  • No retry ever happens after bytes were streamed to the client.
  • Per-tenant usage matches the sum of backend usage fields.

9. Common pitfalls#

Round-robin for LLMs. Requests differ in cost by orders of magnitude and it scatters prefixes across replicas, destroying cache hits.

Rate limiting on requests only. One request can be 100 tokens or 100,000.

Retrying after the stream started. The client receives duplicated or spliced output.

Retries without a budget. A slow backend triggers retries, which add load, which slows more backends.

A single total timeout. A healthy 3-minute generation and a hung connection look the same. Use TTFT and inter-token idle timeouts.

Buffering the response (default in many HTTP libraries and proxies).

Pure consistent hashing with no load bound. One popular prefix overloads one replica.

Not cancelling upstream on client disconnect. You pay for tokens nobody reads.


10. Stretch goals#

  • Fallback cascade: primary model → smaller model on overload or timeout (XII.09).
  • Priority tiers with load shedding: drop batch-tier traffic first.
  • Weighted canary routing by percentage with sticky sessions (VIII.10).
  • Replace hash affinity with a real prefix index fed by backend cache events — the design behind KV-cache-aware schedulers such as llm-d and the Gateway API Inference Extension.
  • Cost attribution: tokens × per-model price per tenant (XI.03).

11. Interview questions this project answers#

  1. Why is round-robin a bad load balancer for LLM inference?
  2. How do you rate-limit when request cost varies 1000×?
  3. When is a retry safe for a streaming response?
  4. How does prefix-aware routing improve TTFT, and what stops it from creating hot spots?
  5. Which timeouts does an LLM gateway need, and why more than one?

12. Next#

Project 13 — Multi-GPU inference

↑↓ navigate ↵ open