1. The decision#
You need to choose, and the honest answer is that for most teams, most of the time, the decision matters less than how well you tune whatever you choose. A well-tuned vLLM beats a badly-tuned TensorRT-LLM by more than TRT-LLM beats vLLM when both are tuned.
That said, there are real distinctions. Here’s how to make the choice defensibly.
2. The decision tree#
START
├─ Do you serve models other than LLMs (embeddings, vision, rankers, classifiers)
│ from the same platform?
│ └─ YES → TRITON as the platform layer, with per-model backends
│ (tensorrtllm or vllm for LLMs, ORT for small models)
│
├─ Is your fleet non-NVIDIA (AMD, Intel, Gaudi, Inferentia, TPU)?
│ ├─ AMD/Intel/Gaudi → vLLM (ROCm and other backends); verify your model
│ ├─ AWS Inferentia → vLLM with the Neuron backend, or the Neuron SDK
│ └─ TPU → JAX/MaxText, or vLLM's TPU backend
│
├─ Is your workload dominated by SHARED PREFIXES?
│ (agentic loops, few-shot, multi-sample, document QA, tree search)
│ └─ YES → SGLANG. Its RadixAttention + router are built for this.
│
├─ Are you serving ONE OR TWO stable models at very high volume,
│ on NVIDIA, and is 15-25% throughput worth operational complexity?
│ └─ YES → TENSORRT-LLM (deployed via Triton)
│
├─ Are you running TGI today?
│ └─ It was archived in March 2026. Plan a migration to vLLM or SGLang
│ (section 6); do not start new work on it.
│
└─ Otherwise (the common case)
└─ VLLM. Broadest model support, simplest deployment, strong performance,
largest community. This is the correct default.“vLLM unless you have a specific reason” is a defensible position and the one most teams should take.
The layer above the engine#
The tree above picks an engine. Once you run more than a handful of replicas you also need a layer that routes, scales and shares KV cache across them. As of October 2026 the open-source options are:
llm-d Kubernetes-native distributed inference (CNCF sandbox
since March 2026). Cache- and load-aware routing through
the Gateway API Inference Extension, P/D disaggregation.
NVIDIA Dynamo 1.0 released in 2026. Router, KV block manager and P/D
disaggregation over vLLM, SGLang or TensorRT-LLM.
vLLM production-stack Helm-based reference deployment with a router.
and AIBrix Control plane for vLLM fleets.
KServe, Ray Serve General model-serving platforms that host LLM engines.Choose the engine first; add one of these when routing and autoscaling become the problem (Section XII). They are not alternatives to the engine.
A criterion this section learned the hard way#
TGI was a reasonable recommendation in this curriculum until Hugging Face archived it. Add maintenance trajectory to any evaluation: commit and release cadence, who funds the work, how quickly new model architectures land. An engine that stops moving becomes a migration.
3. The evaluation you should actually run#
Do not choose on published benchmarks. Run this:
1. DEFINE YOUR WORKLOAD
- prompt length distribution (p50, p95, p99) from real traffic
- output length distribution
- arrival pattern (rate, burstiness)
- prefix sharing rate
- required features (structured output, multi-LoRA, long context)
2. DEFINE YOUR SLOs
- p95 TTFT, p95 ITL, availability
3. BUILD ONE BENCHMARK HARNESS that works against any OpenAI-compatible endpoint.
Replay your real length distribution and arrival pattern.
4. FOR EACH CANDIDATE:
a. Deploy it (record how long this took — it's data)
b. Tune it using the procedure in Section VIII.04
c. Measure: throughput at your SLO, cost per million tokens,
p95/p99 TTFT and ITL, and the maximum sustainable arrival rate
d. Verify output quality matches your reference
e. Note operational friction: config complexity, docs, failure modes
5. COMPARE ON COST PER MILLION TOKENS AT YOUR SLO.
Not on peak throughput. Not on batch-1 latency. On the number that
determines your bill while meeting your commitments.Step 4a and 4e are real criteria. A system that’s 15% faster and takes three engineer-weeks to operate is not obviously better.
4. The criteria, weighted#
CRITERION WEIGHT WHY
Model support (yours, now) HIGH if it doesn't run your model, nothing else matters
Cost per M tokens at SLO HIGH the actual objective
Feature coverage HIGH structured output, LoRA, long context — do you need them?
Operational simplicity HIGH you'll live with this for years
Community / support MEDIUM how fast do bugs get fixed?
Peak throughput MEDIUM only matters at your operating point
Batch-1 latency VARIES critical for some products, irrelevant for others
Hardware flexibility VARIES matters if your fleet is heterogeneous or may change
Upgrade cadence MEDIUM fast-moving = new features, but also churnModel support is the top criterion and the most often overlooked at decision time. A new architecture appears in vLLM within days and in TensorRT-LLM within months. If your model strategy involves adopting new models quickly, that difference dominates everything else.
5. Common situations and answers#
"We're a startup serving one open model to our product."
→ vLLM. Deploy in an afternoon, tune in a day, revisit in a year.
"We serve 40 fine-tunes of one base model."
→ vLLM with multi-LoRA. This is a 50x cost decision (Section VIII.09).
"We run an agent platform with long tool-definition prefixes."
→ SGLang. Its whole design targets this.
"We're an inference provider with one flagship model at massive scale."
→ TensorRT-LLM on Triton. The 15-25% matters at your volume, and your
model set is stable enough to justify the rebuild treadmill.
"We serve LLMs, embeddings, rerankers, and a vision model."
→ Triton as the platform, with the appropriate backend per model.
Or: separate systems per model type, which is often simpler.
"We're on AMD MI300."
→ vLLM (ROCm support). Verify your specific model and features.
"We need on-device / browser inference."
→ ONNX Runtime, llama.cpp, MLC, or the platform's native runtime.
Different world; nothing in this list applies.
"We can't decide."
→ vLLM. You can migrate later; the API is compatible.6. The migration question#
Because you will reconsider:
WHAT MAKES MIGRATION EASY:
✓ Everything speaks the OpenAI-compatible API → clients don't change
✓ Your benchmark harness is engine-agnostic
✓ Your model artifacts are standard (safetensors + config.json)
✓ Your metrics have consistent names across engines (normalize at collection)
WHAT MAKES IT HARD:
✗ Engine-specific quantization formats (a TRT engine isn't portable)
✗ Engine-specific features in your product (SGLang's DSL, engine-specific params)
✗ Tuning knowledge that doesn't transfer
✗ Operational tooling built around one engine's metricsDesign for migration: keep the OpenAI API as your internal contract, keep model artifacts in a standard format, and normalize metrics names. Then the choice is reversible, which lowers the stakes considerably.
7. What NOT to do#
✗ Build your own engine.
Unless you have a genuinely novel requirement and a team to maintain it,
you will spend a year reimplementing continuous batching and paged
attention, badly. The existing systems represent hundreds of
engineer-years.
✗ Choose on a vendor benchmark.
Benchmarks are run by the party that wins them, on workloads that
favor them.
✗ Choose on peak throughput.
Your operating point is not peak throughput; it's throughput subject
to your latency SLO.
✗ Ignore operational cost.
A year of operating a complex system costs more than 20% of GPU spend
for most teams.
✗ Optimize the choice before tuning what you have.
The tuning delta usually exceeds the engine delta.The first item deserves emphasis. “We’ll write our own inference server” is the most expensive decision in this curriculum, and it is made surprisingly often. The legitimate reasons are: a genuinely novel architecture nobody supports, a hardware platform nobody targets, or you are an inference infrastructure company. Otherwise, contribute to an existing engine instead — you get the feature and the maintenance is shared.
8. Hands-on exercise#
A. Build the harness. Write a benchmark that works against any OpenAI-compatible endpoint, replays a configurable length distribution and arrival pattern, and reports throughput, TTFT/ITL percentiles, and cost per million tokens. This is the most reusable artifact in this section.
B. Evaluate two candidates. Deploy two engines for the same model. Tune both properly. Run your harness against each. Produce a one-page comparison with a recommendation and its justification.
C. Measure the tuning delta. For one engine, measure throughput with default settings and with proper tuning. Compare that delta to the delta between engines. Which is larger?
D. Migration test. Take a client application built against one engine and point it at another. What broke? Fix it and record what made migration hard.
E. Write the decision doc. For your actual situation, write the engine selection document: requirements, candidates, evaluation results, decision, and the conditions under which you’d revisit it. This is a real deliverable in a real job.
9. Interview questions#
- How would you choose a serving stack? Walk me through your process.
- What’s the top criterion, and why do people overlook it?
- When would you choose TensorRT-LLM over vLLM? When not?
- Why is peak throughput the wrong comparison metric?
- What makes migration between engines easy or hard?
- When is building your own inference engine justified?
- Your team wants to switch engines for a claimed 20% improvement. What do you ask them?
10. Further reading#
- [REFERENCE] Each system’s documentation and its own benchmark methodology
- [REFERENCE] MLPerf Inference — the closest thing to a neutral comparison
- [ESTABLISHED] vLLM
benchmarks/benchmark_serving.py— a good harness to start from - Next: Section IX — Distributed Inference