1. What a gateway does#
The single front door for all inference traffic. Everything common to every request lives here, so it doesn’t have to live in every engine.
CLIENT ──► GATEWAY ──► ROUTER ──► ENGINE
GATEWAY RESPONSIBILITIES
TLS termination
authentication and authorization
rate limiting and quota enforcement
request validation (limits, schema)
model alias resolution
chat template application
tokenization
request/response logging and metrics
streaming (SSE) framing
detokenization
error normalization
usage accounting for billingThe gateway is the place where the platform’s policy lives. Engines should know nothing about tenants, quotas, or billing.
2. Why a separate component#
IF YOU PUT THIS IN THE ENGINE
✗ every engine process runs auth, rate limiting, and billing logic
✗ the GIL contends with the engine loop (Section VIII.01)
✗ policy changes require redeploying models (minutes each)
✗ N copies of the quota state to keep consistent
✗ engines become platform-specific rather than swappable
WITH A GATEWAY
✓ policy changes deploy in seconds
✓ engines stay generic and swappable
✓ one place for auth, billing, and observability
✓ scales independently (CPU-bound, cheap to replicate)3. The request pipeline#
1. TLS terminate
2. Parse HTTP, decode JSON (orjson/msgspec — 3-10x faster)
3. Authenticate (JWT verify, or API key lookup)
4. Authorize (may this key call this model?)
5. Validate (limits, parameter ranges)
6. Resolve alias ("chat-large" → model+version)
7. Apply chat template (from the registry)
8. Tokenize (thread pool; Rust tokenizer)
9. Rate limit / quota check (tokens, concurrency, requests)
10. Route (Section XII.06)
11. Forward to the engine (gRPC or HTTP)
12. Stream response back
- detokenize incrementally (Section V.02)
- frame as SSE
- detect client disconnect → abort (Section V.04)
13. Record usage, emit metrics and logs
14. Release the concurrency slot, refund unused quotaSteps 8 and 12 are the CPU-heavy ones. At 5,000 output tokens/sec across all streams, the per-token detokenize + SSE frame + write costs 0.1-0.5 of a core.
4. Sizing and scaling the gateway#
PER-REQUEST CPU COST
parse + validate 50-200 µs
auth (cached) 10-30 µs
auth (uncached, JWT verify) 100-500 µs
chat template 10-100 µs
tokenize (per 1k tokens) ~50 µs
routing decision 10-50 µs
PER-TOKEN CPU COST (streaming)
detokenize + SSE frame + write 20-100 µs
SIZING
at 100 req/s with 2k prompts and 400 output tokens:
per-request: 100 × 500 µs = 0.05 core
tokenize: 100 × 100 µs = 0.01 core
per-token: 100 × 400 × 50 µs = 2.0 cores ← dominates
→ ~3 cores plus headroom. Small.
at 1,000 req/s: ~25 cores. Run 4-6 gateway pods.The gateway is cheap relative to the GPUs, which is why putting work here rather than in the engine is almost always right.
Scale it independently. Gateway pods are stateless (with cached auth and routing tables) and scale in seconds — unlike engines.
5. Auth and authorization#
API KEYS
hash them; store the hash. Never log the key.
cache the lookup (key_hash → tenant, scopes) with a short TTL.
→ the cache is what keeps auth off the critical path when the auth
store is slow or down (Section XI.05: fail-closed with a cache
that rides out brief outages)
JWT
verify the signature locally with a cached JWKS.
✓ no per-request network call
✗ revocation is hard (short expiry + a revocation list)
SCOPES
what can this key do?
models: [chat-large, chat-small]
capabilities: [streaming, structured_output]
max_context: 32768
tier: standard
→ enforce at the gateway, so engines don't need to knowCapability scoping is under-used and valuable: restricting which models and features a key can access limits both cost exposure and blast radius.
6. Streaming through the gateway#
This is where most gateway implementations get it wrong.
func (g *Gateway) chat(w http.ResponseWriter, r *http.Request) {
rc, err := g.preamble(r) // auth, validate, tokenize
if err != nil {
writeError(w, err)
return
}
ctx := r.Context() // cancelled when the client disconnects
if !rc.Body.Stream {
result, err := g.engine.Generate(ctx, rc)
if err != nil {
writeError(w, err)
return
}
g.recordUsage(rc, result.OutputTokens)
json.NewEncoder(w).Encode(toOpenAI(result))
return
}
w.Header().Set("Content-Type", "text/event-stream")
w.Header().Set("Cache-Control", "no-cache")
w.Header().Set("X-Accel-Buffering", "no")
w.Header().Set("X-Model-Version", rc.ModelVersion)
flusher := w.(http.Flusher)
send := func(v any) {
b, _ := json.Marshal(v)
fmt.Fprintf(w, "data: %s\n\n", b)
flusher.Flush()
}
detok := NewIncrementalDetokenizer(rc.Tokenizer)
stop := NewStopSequenceMatcher(rc.Body.Stop)
nOut := 0
defer func() { // runs on every exit path, including disconnects
g.recordUsage(rc, nOut)
g.releaseQuota(rc, nOut) // refund unused
}()
tokens, errs := g.engine.GenerateStream(ctx, rc) // the engine aborts when ctx is cancelled
for tokens != nil {
select {
case <-ctx.Done(): // 1. disconnect — CRITICAL (Section V.14)
return
case err := <-errs:
// in-band error — we already sent 200 (Section V.14)
send(map[string]any{"error": map[string]string{"message": err.Error(), "type": "server_error"}})
tokens = nil
case id, ok := <-tokens:
if !ok {
tokens = nil
break
}
delta := detok.Add(id) // 2. incremental detokenization (Section V.02)
if delta == "" {
continue // incomplete UTF-8
}
emit, done := stop.Feed(delta) // 3. stop sequence hold-back
if emit != "" {
nOut++
send(chunk(emit))
}
if done {
tokens = nil
}
}
}
send(finalChunk(usage(rc, nOut)))
fmt.Fprint(w, "data: [DONE]\n\n")
flusher.Flush()
}Five things this gets right that implementations commonly get wrong:
- Disconnect detection and abort propagation.
- Incremental detokenization (no mojibake).
- Stop-sequence hold-back.
- In-band error reporting after streaming has started.
- Quota refund in
finally.
7. Deployment shape#
┌──────────────────────────┐
Internet ──────►│ L7 LB / Ingress │ TLS, WAF, DDoS
│ (proxy_buffering OFF!) │
└────────────┬─────────────┘
│
┌────────────▼─────────────┐
│ GATEWAY (4-8 pods) │ stateless, CPU-only
│ auth, quota, tokenize │ scales in seconds
└────────────┬─────────────┘
│
┌────────────▼─────────────┐
│ ROUTER (2-4 pods) │ routing table + load state
└────────────┬─────────────┘
│
┌────────────▼─────────────┐
│ ENGINES │ GPU
└──────────────────────────┘
Sometimes the gateway and router are one component. That's fine at small
scale; separate them when the routing logic grows.Every layer above the engines must have buffering disabled (Section II.08). Test the whole path from a real client.
8. Build vs adopt#
GENERIC API GATEWAYS (Kong, Envoy, APISIX, cloud gateways)
✓ auth, TLS, basic rate limiting, observability
✗ don't understand tokens (can't rate limit by tokens)
✗ don't tokenize or apply chat templates
✗ don't understand streaming semantics for LLMs
✗ can't do prefix-aware routing
→ USE FOR: TLS, WAF, coarse rate limiting at the edge
→ NOT SUFFICIENT as the inference gateway
LLM-SPECIFIC GATEWAYS (LiteLLM, Portkey, Kong AI Gateway, and others)
✓ multi-provider abstraction, token-aware limiting, caching
✗ maturity and performance vary; evaluate carefully
→ worth evaluating; may cover 80% of your needs
BUILD
✓ full control over routing, quota semantics, and observability
✗ real ongoing work
→ the usual answer for a serious platform, often layered behind a
generic gateway that handles TLS and edge concernsA common and sensible architecture: a generic gateway at the edge (TLS, WAF, IP rate limits) plus a custom inference gateway behind it (tokens, quotas, routing, streaming).
9. Production implications#
- The gateway owns policy; engines own inference. Keep the separation.
- Disable buffering at every layer above the engine.
- Cache auth lookups so the auth store isn’t in the critical path.
- Rate limit by tokens and concurrency at the gateway (Section XI.10).
- Implement disconnect propagation. Highest-value item in the streaming path.
- Use fast JSON (orjson/msgspec) and Rust tokenizers.
- Emit
X-Model-Versionso client-side issues can be correlated with deployments. - Scale the gateway independently. It’s cheap and fast to scale.
- Record usage in a
finallyblock. Otherwise aborted requests aren’t billed or accounted.
10. Common mistakes#
Putting policy in the engine. Slow to change, contends with the engine loop.
Proxy buffering enabled somewhere. Destroys streaming (Section X, case 8).
Per-request auth store lookups. Latency and a hard dependency.
Request-based rate limiting. (Section XI.10.)
No disconnect detection. 10-30% capacity waste.
Naive detokenization. Mojibake for non-Latin text.
No stop-sequence hold-back. Stop strings flash in the output.
Slow JSON parsing. Measurable at high QPS with large prompts.
Not recording usage on the error path.
11. Hands-on exercise#
A. Build the gateway. Implement the pipeline from section 3, with the streaming handler from
section 6. Test with the real openai client library. This is Project 12.
B. Test the streaming correctness. Verify: multi-byte UTF-8 across token boundaries, stop sequences spanning tokens, in-band errors, disconnect abort, and usage recording on all paths.
C. Measure the CPU cost. Load-test at 500 req/s with 2k prompts. Profile the gateway (Section X.09). Where does CPU go? How many cores per 1,000 req/s?
D. Auth caching. Implement the auth cache with TTL. Kill the auth store and verify requests continue for the TTL duration. Measure the latency difference cached vs uncached.
E. End-to-end streaming test. Deploy behind a real ingress. Measure client-side TTFT. Deliberately enable buffering somewhere and confirm your test catches it.
F. Quota refund. Verify that a request with max_tokens: 4096 that generates 100 tokens
refunds the difference, and that an aborted request refunds correctly.
12. Interview questions#
- What belongs in the gateway rather than the engine, and why?
- Walk through the request pipeline in order.
- What’s the CPU cost of a gateway, and what dominates it?
- How do you handle an error after streaming has started?
- Why cache auth lookups, and what’s the failure mode without it?
- Why is a generic API gateway insufficient for LLM traffic?
- What must the gateway do in a
finallyblock, and why?
13. Further reading#
- [REFERENCE] OpenAI API specification — the schema to implement
- [REFERENCE] Envoy, Kong documentation for the general patterns
- [REFERENCE] LiteLLM and similar LLM gateway projects — for comparison
- [REFERENCE] MDN Server-Sent Events
- Next: 09 — Cascades and fallbacks