Below the API

APIs: HTTP, gRPC, Streaming

Intermediate Intermediate 1h Difficulty 2/5

Prerequisites II.08, V.14


1. What is it?#

The contract between clients and your inference service: protocol, schema, streaming semantics, errors, and limits.

For LLMs, the industry has converged on the OpenAI-compatible API as the de facto standard. Deviating from it costs you every existing client library, framework integration, and evaluation harness.


2. Why the convergence matters#

If your API is OpenAI-compatible:
  ✓ works with openai-python, LangChain, LlamaIndex, Instructor, every eval harness
  ✓ clients can switch providers by changing a base_url
  ✓ your users already know the schema

If it isn't:
  ✗ every integration is bespoke
  ✗ you maintain client libraries
  ✗ evaluation tooling doesn't work

Implement the OpenAI schema even if you add extensions. The cost of compatibility is a day; the cost of incompatibility is permanent.


3. Simple analogy#

Electrical sockets. A better plug design doesn’t help if nothing plugs into it. The standard is worth more than the improvement.


4. The essential endpoints#

POST /v1/chat/completions       the main one
POST /v1/completions            legacy text completion
POST /v1/embeddings             if you serve embedding models
GET  /v1/models                 model discovery
GET  /health, /health/ready     liveness and readiness (different!)
GET  /metrics                   Prometheus
POST /tokenize, /detokenize     useful extensions

Request:

{
  "model": "llama-3-70b-instruct",
  "messages": [
    {"role": "system", "content": "You are helpful."},
    {"role": "user", "content": "Explain gravity"}
  ],
  "max_tokens": 500,
  "temperature": 0.7,
  "top_p": 0.9,
  "stream": true,
  "stop": ["\n\n"],
  "seed": 42,
  "response_format": {"type": "json_schema", "json_schema": {...}}
}

Streaming response chunks:

data: {"id":"...","object":"chat.completion.chunk","created":1234,"model":"...",
       "choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}

data: {"...","choices":[{"index":0,"delta":{"content":"Gravity"},"finish_reason":null}]}

data: {"...","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],
       "usage":{"prompt_tokens":12,"completion_tokens":143,"total_tokens":155}}

data: [DONE]

Note: usage in the final chunk is an extension many providers now support (stream_options: {"include_usage": true}). Clients need token counts for cost tracking.


5. Technical explanation#

Protocol choice#

                    HTTP/1.1 + SSE      HTTP/2          gRPC
Public API          ✓ THE STANDARD      ok              needs grpc-web
Internal service    ok                  ✓               ✓ best
Multiplexing        1 req/connection    many            many
Per-message overhead ~80 bytes text     ~20 binary      ~10 protobuf
Browser support     native EventSource  native          no
Middlebox friendly  ✓ (it's just HTTP)  mostly          often not
Bidirectional       no                  no (in practice) ✓

Practical answer: SSE over HTTP/1.1 for the public API, gRPC between gateway and engines. The 80 bytes of SSE framing per token is negligible (5,000 tok/s × 80 B = 400 KB/s), and the compatibility is worth everything.

Validation — what to enforce at the edge#

max input tokens           protects activation memory and prefill time
max output tokens          protects KV cache and slot occupancy
max total tokens           prompt + max_tokens ≤ model's context limit
temperature ∈ [0, 2]
top_p ∈ (0, 1]
top_k ≥ 1 or absent
n (completions) ≤ small    n=100 multiplies your cost by 100
stop sequences: ≤ 4, each ≤ 64 chars
message count and total size

Every one of these is a capacity control, not just hygiene. A client sending max_tokens: 100000, n: 20 can occupy a large fraction of your fleet.

Reject with a clear 400, naming the limit and the observed value.

Errors#

400  invalid request (bad params, too long)
401  unauthenticated
403  unauthorized for this model
404  unknown model
408  request timeout
413  payload too large
422  semantic validation failure
429  rate limited        ← include Retry-After
499  client closed request (nginx convention; useful in logs)
500  internal error
503  overloaded — no capacity   ← include Retry-After
504  upstream timeout

429 vs 503 matters. 429 means “you’re over your quota”; 503 means “we’re over capacity.” Clients should back off differently. Section VIII.05.

Mid-stream errors must go in-band (Section V.14):

data: {"error":{"message":"...","type":"server_error","code":"engine_crash"}}

data: [DONE]

Extensions worth having#

Beyond the OpenAI schema, useful additions:

  "priority": 0-9                  scheduling priority (multi-tenant)
  "ignore_eos": bool               for benchmarking
  "min_tokens": int                force a minimum length
  "repetition_penalty": float      not in the OpenAI schema, widely supported
  "min_p": float                   better than top_p (Section V.13)
  "guided_json" / "guided_regex"   structured output
  "logit_bias": {token_id: bias}   in the OpenAI schema
  "echo": bool                     return the prompt too
  "lora_request": name             which adapter to use (multi-LoRA)

Put extensions in a namespaced field or document them clearly, so clients know what’s portable.

Health checks — two different things#

GET /health        LIVENESS: is the process alive? Return 200 always if it
                   responds. Used to decide whether to RESTART the pod.

GET /health/ready  READINESS: is the model loaded, warmed up, and able to serve?
                   Returns 503 during model load and warmup.
                   Used to decide whether to SEND TRAFFIC.

Conflating them is a common and damaging mistake: if liveness fails during a 5-minute model load, Kubernetes kills the pod and it never starts. Set initialDelaySeconds generously on liveness, or use a startup probe.

Health checks must be cheap. Never run a real generation on every probe — at a 5-second probe interval across 100 replicas, that’s real GPU capacity spent on health checks.


6-9. Under the hood, performance, production, mistakes#

Under the hood — per-request CPU cost in the API layer:

TLS + HTTP parse         50-200 µs
JSON decode              20-500 µs (scales with message size)
Validation                5-20 µs
Chat template             10-100 µs
Tokenization             50 µs - 5 ms (scales with prompt length)
Per-token: JSON encode + SSE frame + write   20-100 µs

At 5,000 output tokens/sec across all streams, the per-token cost alone is 0.1-0.5 of a core. Use a fast JSON library (orjson, msgspec) and consider batching token emissions.

Production:

  • Be OpenAI-compatible. Test with the actual openai client library in CI.
  • Enforce all limits at the edge, with clear errors.
  • Separate liveness and readiness probes, with a startup probe for model loading.
  • Return usage in the final streaming chunk.
  • Include Retry-After on 429 and 503.
  • Log the request shape (input tokens, max_tokens, model, tenant) for every request — it’s the basis of all capacity analysis.
  • Version your API (/v1/) and don’t break it.

Mistakes:

  • Inventing your own schema. Loses the ecosystem.
  • Conflating liveness and readiness. Restart loops during model load.
  • Expensive health checks.
  • No limits on n or max_tokens. One client can consume the fleet.
  • Returning 500 for overload instead of 503. Clients retry immediately and make it worse.
  • Not handling mid-stream errors.
  • Slow JSON parsing on large messages.

10. Hands-on exercise#

A. Build an OpenAI-compatible endpoint. Implement /v1/chat/completions with streaming and non-streaming modes. Test it with the real openai Python client — that’s the compatibility test that matters.

B. Validation suite. Write tests for every limit in section 5. Verify each returns a clear 400 with a useful message.

C. Probe design. Implement liveness, readiness, and startup probes. Simulate a 5-minute model load and verify Kubernetes-style probing doesn’t kill the pod.

D. Measure the API cost. Profile the API layer at 500 req/s with 2,000-token prompts. How many cores does it need? What’s the biggest cost — parsing, tokenizing, or streaming?

E. Error semantics. Implement 429 and 503 with Retry-After. Write a client that backs off correctly for each and verify the behavior differs.


11. Interview questions#

  1. Why implement the OpenAI-compatible API rather than your own?
  2. When would you use gRPC instead of HTTP+SSE?
  3. What’s the difference between liveness and readiness, and what breaks if you conflate them?
  4. Which request parameters would you validate at the edge, and why is each a capacity control?
  5. How do you return an error after streaming has started?
  6. What’s the difference between 429 and 503, and how should a client treat each?
  7. What per-request information would you log, and what would you use it for?

12. Further reading#

  • [REFERENCE] OpenAI API reference — the de facto schema
  • [REFERENCE] vLLM’s OpenAI-compatible server implementation
  • [REFERENCE] Kubernetes probe documentation
  • [FUNDAMENTAL] Google SRE Book, chapter on load shedding (for error semantics)
  • Next: 03 — Queues, scheduling, admission control

↑↓ navigate ↵ open