Below the API

TGI, Triton, and ONNX Runtime

Intermediate Advanced 1h 15m Difficulty 3/5

Prerequisites 11, 12


1. What is it?#

Three more serving systems, each designed for a different problem. Understanding what each optimized for is more useful than a feature checklist.

TGI (Text Generation Inference)   Hugging Face's LLM server. Rust router +
                                  Python/Rust engine. Production-hardened,
                                  tightly integrated with the HF ecosystem.

NVIDIA Triton Inference Server    A general-purpose model server (any framework,
                                  any modality). Not LLM-specific; hosts
                                  TensorRT-LLM as a backend.

ONNX Runtime                      A cross-platform inference runtime. Strong on
                                  CPU, edge, and non-NVIDIA hardware. The right
                                  answer for small models everywhere.

2. TGI — Text Generation Inference#

warning

Status, October 2026: TGI is no longer developed. Hugging Face put it into maintenance mode in December 2025 and archived the repository in March 2026; its documentation now points users to vLLM or SGLang. Do not start a new deployment on it. This section stays because the design — a compiled router in front of Python model workers — is worth understanding, and because you will meet TGI in existing systems and have to migrate them.

Design#

┌─────────────────────────────────────┐
│ ROUTER (Rust)                        │  ← the distinguishing choice
│   HTTP, validation, tokenization     │
│   batching decisions, queue          │
│   streaming, per-request state       │
└──────────────┬──────────────────────┘
               │ gRPC
┌──────────────▼──────────────────────┐
│ SHARD(s) (Python/Rust)               │
│   model execution, one per TP rank    │
│   paged attention, CUDA graphs        │
└─────────────────────────────────────┘

The Rust router is the notable decision. Everything CPU-bound and latency-sensitive — HTTP parsing, tokenization, batching logic, streaming — is outside Python. Consequences:

✓ no GIL contention with the model process
✓ very low per-request CPU overhead
✓ predictable latency under high request rates
✗ extending the router requires Rust

Strengths#

✓ Production hardening: it's been running HF's inference endpoints for years
✓ Excellent observability out of the box (OpenTelemetry, detailed metrics)
✓ Tight HF Hub integration — model IDs just work
✓ Good guided/structured generation support
✓ Solid multi-backend support (also runs on AMD, Intel, Gaudi, Neuron)
✓ Low CPU overhead at high QPS

Where it fits#

TGI was the pragmatic choice when you wanted a production-hardened server with excellent operational properties and did not need peak throughput. Today it fits one situation: you already run it, it works, and you are planning the move off it. The migration is mostly mechanical — both it and its successors speak the OpenAI-compatible API — but re-validate sampling defaults, chat templates, stop sequences and your dashboards (metric names differ), following the checklist in 14.

The lesson its history teaches is in Section 14’s criteria: an engine’s maintenance trajectory matters as much as its benchmark numbers.


3. Triton Inference Server#

note

NVIDIA now presents Triton as part of its Dynamo platform (you will see “Dynamo-Triton”). Dynamo itself — an open-source, LLM-specific distributed serving layer with a router, KV-cache management and prefill/decode disaggregation over engines such as vLLM, SGLang and TensorRT-LLM — reached version 1.0 in 2026. Triton remains the general-purpose, any-framework model server described here; Dynamo is the answer to “many GPUs, one LLM service” (XIII.06).

What it actually is#

A general model server, not an LLM server. It hosts models from any framework through a backend plugin architecture.

Backends:
  tensorrtllm      ← how you run TensorRT-LLM in production
  vllm             ← yes, you can run vLLM under Triton
  python           arbitrary Python models
  onnxruntime
  pytorch (libtorch), tensorflow
  dali             (data loading/preprocessing on GPU)
  fil              (forest/tree models)

The features that matter#

MODEL ENSEMBLES / BUSINESS LOGIC SCRIPTING
  Chain models together server-side:
    tokenizer (python backend) → LLM (trtllm backend) → detokenizer
  or:
    image preprocessor (DALI) → vision encoder → LLM
  One request, one round trip, all on the server.

DYNAMIC BATCHING
  For non-autoregressive models this is the correct batching strategy
  (Section V.08). Triton's implementation is mature and well-tuned.

MODEL REPOSITORY + VERSIONING
  Filesystem/S3 layout with explicit versions, load/unload policies,
  and a model control API.

CONCURRENT MODEL EXECUTION
  Multiple models and multiple instances per model on one GPU.

METRICS
  Comprehensive Prometheus metrics per model and per version.

Where it fits#

USE TRITON WHEN:
  - You serve heterogeneous models (LLM + embeddings + vision + rankers)
  - You want server-side pipelines (ensembles)
  - You're running TensorRT-LLM (this is the standard deployment path)
  - You need a mature model repository and versioning story

DON'T USE TRITON WHEN:
  - You serve only LLMs and want the simplest path (vLLM/SGLang are simpler)
  - Your team doesn't want to learn `config.pbtxt`

The operational surface is real. Model repositories, config.pbtxt files, ensemble definitions, and instance groups are all things to learn and maintain. Worth it for a heterogeneous platform; overhead for a single LLM.


4. ONNX Runtime#

What it is#

A cross-platform inference runtime with a graph optimizer and pluggable “execution providers.”

Execution providers:
  CPU (with oneDNN/MLAS)     ← very good; the main reason to use ORT
  CUDA, TensorRT
  DirectML (Windows)
  CoreML (Apple)
  OpenVINO (Intel)
  ROCm (AMD)
  QNN (Qualcomm), NNAPI (Android)
  WebGPU / WASM (browser)

Where it fits for inference engineering#

STRONG FOR:
  ✓ Small models on CPU (embeddings, classifiers, rerankers, small LLMs)
  ✓ Edge and on-device deployment
  ✓ Non-NVIDIA hardware
  ✓ Browser/WASM deployment
  ✓ Deterministic, portable deployment artifacts

WEAK FOR:
  ✗ Large LLM serving — no continuous batching, no paged KV cache
    (ONNX Runtime GenAI adds LLM features but is behind vLLM/TRT-LLM)
  ✗ Cutting-edge model architectures (ONNX export often lags)

The realistic use in an LLM platform: ORT serves the supporting models — the embedding model for RAG, the reranker, the classifier that routes requests, the safety filter. Those are small, often CPU-appropriate, and don’t need LLM-specific machinery.

Running them on CPU with ORT instead of on GPU with a heavyweight server frees GPUs for the LLM and is often cheaper.


5. The comparison table#

The TGI column describes the last released version; it is frozen there.

                    vLLM      SGLang    TGI       TRT-LLM   Triton    ORT
LLM continuous batch  ✓✓       ✓✓        ✓✓        ✓✓        via bkend  ✗
Paged KV              ✓✓       ✓✓        ✓✓        ✓✓        via bkend  ✗
Prefix caching        ✓✓       ✓✓✓       ✓         ✓✓        via bkend  ✗
Peak throughput       ✓✓       ✓✓        ✓✓        ✓✓✓       —          ✗
Batch-1 latency       ✓        ✓✓        ✓✓        ✓✓✓       —          ✓
Model support speed   ✓✓✓      ✓✓        ✓✓        ✓         —          ✓
Ease of deployment    ✓✓✓      ✓✓        ✓✓✓       ✓         ✓          ✓✓
Structured output     ✓✓       ✓✓✓       ✓✓        ✓✓        via bkend  ✗
Multi-modal models    ✓✓       ✓✓        ✓✓        ✓✓        ✓✓✓        ✓
Non-LLM models        ✗        ✗         ✗         ✗         ✓✓✓        ✓✓✓
Non-NVIDIA hardware   ✓        ✓         ✓✓        ✗         ✓✓         ✓✓✓
CPU inference         ✓        ✗         ✓         ✗         ✓✓         ✓✓✓
Observability         ✓✓       ✓✓        ✓✓✓       ✓✓        ✓✓✓        ✓
Operational surface   small    small     small     large     large      small

(Ratings are directional and drift with releases. Re-evaluate for your version.)


6. What each system teaches#

Beyond “which to use,” each embodies a lesson:

vLLM      → memory management is the bottleneck; solve it with paging
SGLang    → workload structure is exploitable; represent it explicitly
TGI       → keep the CPU-bound path out of Python
TRT-LLM   → giving up flexibility buys real performance
Triton    → a platform serving many model types needs a different abstraction
ORT       → portability and CPU efficiency are their own value

Reading more than one of these systems’ source is one of the best uses of your time in this curriculum. They solved the same problem with different priorities, and the differences are instructive.


7-9. Production, mistakes, and choosing#

Production considerations for each:

TGI:      archived (March 2026): no new models, features or fixes. Plan a
          migration; do not adopt
Triton:   budget time for config.pbtxt and the model repository; excellent
          for heterogeneous platforms; the standard TRT-LLM deployment path
ORT:      export is the hard part (opset support, dynamic shapes); verify
          numerics after export; excellent for the supporting-model tier

Mistakes:

  • Using Triton for a single LLM and paying the operational overhead for nothing.
  • Using ORT for large LLM serving. It lacks the LLM-specific machinery.
  • Running embedding and reranker models on GPUs with a heavyweight server when ORT on CPU would be cheaper and free the GPU.
  • Assuming feature parity. Check your specific model and feature against the specific version.
  • Choosing on benchmarks alone. Operational fit matters more over a year.

10. Hands-on exercise#

A. Deploy three ways. Serve the same small model with vLLM, TGI, and Triton+vLLM-backend. Compare: setup time, configuration complexity, throughput, p95 latency, and the metrics each exposes.

B. Triton ensemble. Build a Triton ensemble: a Python-backend tokenizer, a model, and a Python-backend detokenizer. Measure the latency versus doing tokenization client-side.

C. ORT for supporting models. Export an embedding model to ONNX and serve it with ORT on CPU. Compare throughput and cost per million embeddings against running it on GPU.

D. TGI’s router. Load-test TGI at high request rates (many short requests) and compare CPU usage of the router against vLLM’s Python API layer at the same rate. Quantify the Rust advantage.

E. Read one. Pick TGI’s Rust router or Triton’s core scheduling. Read enough to explain one design decision it made differently from vLLM, and why.


11. Interview questions#

  1. Why did TGI put its router in Rust? What does that buy?
  2. What is Triton and when would you choose it over vLLM?
  3. What is a Triton ensemble and what problem does it solve?
  4. Where does ONNX Runtime fit in an LLM serving platform?
  5. Why is dynamic batching the right choice for an embedding model but not an LLM?
  6. Compare the operational surface of vLLM and Triton.
  7. What lesson does each of these systems embody about serving design?

12. Further reading#

  • [REFERENCE] Hugging Face TGI documentation and source
  • [REFERENCE] NVIDIA Triton Inference Server documentation, especially the architecture and ensemble sections
  • [REFERENCE] ONNX Runtime documentation and ONNX Runtime GenAI
  • Next: 14 — Choosing a serving stack

↑↓ navigate ↵ open