1. What is it?#
Serving more than one model from shared infrastructure. Three distinct problems, often confused:
1. MANY BASE MODELS different architectures/sizes on a shared fleet
2. MANY FINE-TUNES one base model, many LoRA adapters
3. MODEL VERSIONS the same model, different versions (file 10)They have different solutions. Conflating them leads to bad architecture.
2. Problem → Why → Optimization#
PROBLEM You have 50 models. Dedicating a GPU to each means 50 GPUs, most
of them idle most of the time.
WHY Traffic per model is small and bursty; GPU memory is the constraint;
model loading is slow.
OPTIMIZE Depends on which of the three problems you have.3. The decision tree#
Are the models fine-tunes of ONE base?
├─ YES → MULTI-LORA. One base in memory, adapters swapped per request.
│ Serve hundreds of variants at the cost of one model. ← the big win
│
└─ NO, different base models
├─ All small (< 3B) and low traffic?
│ → Co-locate several per GPU with MPS or MIG.
│
├─ A few hot, many cold?
│ → Dedicated replicas for hot models;
│ an LRU-cached shared pool for cold ones.
│
├─ All large and all hot?
│ → Dedicated replicas. There is no trick; you need the GPUs.
│
└─ Very bursty, latency-tolerant?
→ Serverless / scale-to-zero with fast loading.4. Multi-LoRA — the highest-value case#
What LoRA is, from a serving perspective#
Base weight: W (d × k), frozen, shared
LoRA adapter: ΔW = B·A where A is (r × k), B is (d × r), r is typically 8-64
Forward: y = xW + x(BA) · (α/r)
↑ ↑
shared per-request, tiny
Adapter size for r=16 on a 7B model: ~40 MB
Base model: 14 GB
→ 100 adapters = 4 GB. You can serve 100 fine-tunes for 29% more memory
than serving one.This is one of the best deals in inference engineering.
How serving works#
Batch containing requests for 4 different adapters:
base GEMM: X @ W ← one big GEMM, all requests together
LoRA GEMM: grouped/segmented GEMM
request 0,3 → adapter A
request 1 → adapter B
request 2,4 → adapter C
→ a "grouped GEMM" where each group uses different weights
y = base_out + lora_out * scalingThe LoRA part is a grouped GEMM (Section IV.03) — many small matmuls with different weights, launched together. Punica and S-LoRA introduced the kernels (SGMV — segmented gather matrix vector) that make this efficient.
Overhead of multi-LoRA vs single model:
r=16, 4 adapters in a batch: 5-12% slower
r=64, 16 adapters: 15-25% slowerSmall overhead, enormous flexibility.
Adapter management#
Adapters live in a tiered cache:
GPU memory: the N most active (instant)
Host memory: the next M (copy over PCIe: 40 MB / 20 GB/s = 2 ms)
Disk/object storage: the rest (load on demand)
Because adapters are tiny, moving them is cheap. This is the key difference
from base models: a 40 MB adapter loads in milliseconds, a 14 GB model in
minutes.5. Different base models on shared GPUs#
Option A: co-location with MPS#
Several model processes on one GPU, with the MPS daemon so their kernels
run concurrently rather than time-sliced.
✓ flexible memory division
✓ better utilization than time-slicing for small models
✗ no isolation — one process's OOM or fault can affect others
✗ contention for SMs; unpredictable latencyOption B: MIG partitioning#
Hardware partitions: 1g.10gb, 2g.20gb, 3g.40gb, etc.
✓ hard isolation (separate SMs, L2 slice, memory)
✓ predictable performance
✗ fixed partition sizes; reconfiguration requires draining the whole GPU
✗ a 1g slice has 1/7 the SMs — often too little for a useful model
✗ can't burst beyond your partitionUse MIG for multi-tenant isolation requirements; use MPS for efficiency with cooperative workloads.
Option C: model swapping (LRU pool)#
A pool of GPUs, each hosting whichever model is currently needed.
On a request for an unloaded model: evict the LRU model, load the requested one.
Viability = load_time vs request_interarrival_time
load 60 s, requests every 5 min for that model → 20% of time spent loading. Bad.
load 10 s, requests every 5 min → 3%. Acceptable.
load 60 s, requests every 2 hours → fine, but users see 60 s TTFT.
→ Only viable with FAST LOADING (Section II.07). Otherwise you thrash.Mitigations: keep a “hot set” pinned; route cold-model requests to a dedicated swap pool so they don’t disturb the hot ones; queue requests for a loading model rather than starting another load.
6. Under the hood — the routing layer#
A multi-model platform needs a control plane (Section XII):
Registry: model_id → {weights_uri, config, replicas, adapters}
Router: request.model → which replica pool
Placement: which models on which GPUs
Autoscaler: per-model replica countsThe routing decision for multi-LoRA is simple (any replica with the base model loaded, plus adapter affinity). For multi-base-model it’s a bin-packing problem, and it’s where most of the complexity lives (Section XII.04).
7. Performance#
Approach Models/GPU Overhead Isolation Switch cost
Dedicated 1 0% perfect n/a
Multi-LoRA (same base) 100+ 5-25% none ~0 (ms)
MPS co-location 2-6 small 10-30% poor n/a
MIG up to 7 0-10% hard reconfiguration
LRU swapping 1 at a time 0% perfect load time (s-min)Multi-LoRA is in a different league for the case it covers. If your 50 models are fine-tunes of one base, this is the answer and nothing else comes close.
8. Production implications#
- Design for multi-LoRA if you serve fine-tunes. It changes the economics by ~50x.
- Cap
max_lorasper batch — the grouped GEMM overhead grows with the number of distinct adapters in a batch. 4-8 is a reasonable ceiling; route to concentrate adapters. - Adapter-aware routing: send requests for the same adapter to the same replica, so the adapter stays GPU-resident and batches together.
- For different base models, prefer dedicated replicas for hot models and pooling only for the long tail.
- Measure per-model traffic before designing. Usually 5 models are 95% of traffic.
- Model swapping needs fast loading. Fix loading first (Section II.07).
- Check CUDA graph compatibility — multi-LoRA can disable graph capture in some engine versions (Section VII.08).
9. Common mistakes#
Dedicating a GPU per fine-tune. 50x more expensive than multi-LoRA.
Unbounded distinct adapters per batch. The grouped GEMM overhead grows; cap it.
Model swapping with slow loading. Thrashing.
MIG for large models. The partitions are too small.
Co-locating without MPS. Time-slicing gives poor utilization.
Ignoring the traffic distribution. Designing a general solution when 5 models are 95% of traffic.
Not measuring the multi-LoRA overhead for your rank and adapter count.
10. Hands-on exercise#
A. Multi-LoRA. Serve one base model with 4+ LoRA adapters (vLLM supports this). Measure: throughput vs single-model, memory usage, and the overhead as you increase the number of distinct adapters in a batch. Plot overhead vs adapter count.
B. Adapter affinity. Simulate routing with and without adapter affinity across 4 replicas and 20 adapters. Compare the number of distinct adapters per batch and the resulting overhead.
C. Swapping economics. For a set of models with measured load times and request rates, compute whether LRU swapping is viable. Simulate it and measure the fraction of time spent loading.
D. MPS. Run two small model servers on one GPU with and without MPS. Compare aggregate throughput and latency variance.
E. Traffic analysis. For a real (or synthetic) multi-model workload, plot the traffic distribution across models. What fraction of models account for 95% of traffic? Design the placement accordingly.
11. Interview questions#
- What are the three distinct multi-model problems and how do their solutions differ?
- How does multi-LoRA work, and what does it cost?
- What is a grouped GEMM and why does multi-LoRA need one?
- When is model swapping viable? Give the arithmetic.
- Compare MPS and MIG for co-locating models.
- Why does adapter-aware routing matter?
- You have 50 fine-tuned variants of one 7B model. Design the serving architecture.
12. Further reading#
- [ESTABLISHED] Hu et al., “LoRA” (2021)
- [ESTABLISHED] Chen et al., “Punica: Multi-Tenant LoRA Serving” (2023)
- [ESTABLISHED] Sheng et al., “S-LoRA: Serving Thousands of Concurrent LoRA Adapters” (2023)
- [REFERENCE] vLLM multi-LoRA documentation; NVIDIA MPS and MIG documentation
- Next: 10 — Versioning, canary, A/B