PidokuInfra

Cost Per Token

Advanced 1h 15m Difficulty 4/5 Topic 03 of 11

Prerequisites I.11, 02


1. Why this is the metric#

Because it’s the number that connects every engineering decision to the business, and because it is the only fair way to compare configurations, models, and providers.

"We improved throughput 30%"      → so what? at what latency? on what hardware?
"Cost per million output tokens
 fell from $0.84 to $0.61"        → unambiguous, comparable, actionable

2. The full model#

                    total_infrastructure_cost_per_hour
cost_per_token = ────────────────────────────────────────
                      tokens_produced_per_hour

Expanding the numerator:

  GPU compute (on-demand / reserved / spot)          70-85%
  CPU hosts (gateway, routers, tokenization)          3-8%
  Storage (weights, caches, logs)                     2-5%
  Network (egress, inter-AZ, load balancers)          2-8%
  Observability (metrics, logs, traces)               1-4%
  Control plane (registry, scheduler, K8s)            1-3%

Expanding the denominator:

  tokens/hour = 3600 × batch_size / step_time_seconds
              × utilization
              × (1 − waste_fraction)

Both the overhead in the numerator and the waste in the denominator are routinely omitted, which is why internal cost estimates are typically 40-70% below actual.


3. The worked calculation#

CONFIGURATION
  Llama-3-70B, FP8, 8×H100 node, TP=8
  Node cost: $22/hour (reserved 1-year pricing; on-demand would be ~$32)

STEP 1 — raw throughput at the operating point
  batch 96, context 4k, measured step time 11.8 ms
  output tokens/sec = 96 / 0.0118 = 8,136 tok/s

STEP 2 — raw cost
  $22/hour ÷ 3600 = $0.00611/sec
  $0.00611 / 8136 tok/s = $7.51e-7 per token
  → $0.751 per million output tokens

STEP 3 — utilization
  Average over 24h (including the 3 a.m. trough): 58%
  → $0.751 / 0.58 = $1.295

STEP 4 — prefill
  Measured: prefill is 31% of GPU time for this workload
  The above counted only decode tokens. Attribute prefill cost:
  → $1.295 / (1 - 0.31) = $1.877 per million OUTPUT tokens
     (or: split it, charging input tokens for prefill — see section 5)

STEP 5 — waste
  Aborted generations: 14% of tokens generated for disconnected clients
  Padding waste (CUDA graph batch rounding): ~6%
  → $1.877 / (1 - 0.20) = $2.346

STEP 6 — redundancy and headroom
  Fleet is sized at 2.26× peak demand (Section XI.02)
  Only 1/2.26 of the fleet's capacity is "used" on average...
  ...but that's already in the utilization figure. Don't double-count.
  Instead: add N+1 redundancy explicitly: × 1.07
  → $2.510

STEP 7 — non-GPU infrastructure
  CPU hosts, storage, network, observability: +18% of GPU cost
  → $2.962

FINAL: ~$2.96 per million output tokens

Compare to step 2’s $0.75. The “real” cost is 4x the naive calculation, and every factor in between is legitimate.


4. Where the 4x goes, and which parts you can fix#

FACTOR                    MULTIPLIER   FIXABLE?
Raw hardware cost           1.00x      hardware choice, reserved pricing
Utilization (58%)           1.72x      YES — traffic consolidation, better
                                       scheduling, workload mixing
Prefill attribution         1.45x      YES — prefix caching
Waste (aborts + padding)    1.25x      YES — cancellation, better graph sizes
Redundancy                  1.07x      partly — larger fleets need less
                                       proportional redundancy
Non-GPU infrastructure      1.18x      partly — efficiency of the platform
                          ────────
                            3.95x

Three of the six are substantially fixable, and together they’re 3.1x.

If you fix them:
  utilization 58% → 78% (consolidation, mixing batch work into troughs)
  prefill 31% → 12% (prefix caching at 65% hit rate)
  waste 20% → 7% (cancellation propagation, tuned graph sizes)
  
  new multiplier: 1.28 × 1.14 × 1.075 × 1.07 × 1.18 = 2.01x
  → $0.751 × 2.01 = $1.51 per million output tokens

  A 49% reduction with no new hardware and no model change.

5. Input vs output token pricing#

Providers charge differently for input and output tokens, typically 3-5x. Here’s why, derived:

Prefill: compute-bound, processes S tokens in one pass
  cost per input token ∝ 2P / prefill_throughput

Decode: memory-bound, one token per pass
  cost per output token ∝ (weight_bytes + KV_bytes) / bandwidth

For Llama-3-70B FP8 on 8×H100:
  prefill throughput:  ~50,000 tok/s per node
  decode throughput:   ~8,100 tok/s per node
  
  ratio: 6.2x

  → an output token costs ~6x an input token in GPU time
  → typical pricing of 3-5x is roughly consistent, modulated by
    prefix caching (which makes input tokens cheaper still) and
    market factors

To attribute cost properly, measure the prefill/decode time split and allocate accordingly:

input_cost_per_token  = (node_cost/hour × prefill_time_fraction)
                      ÷ (input_tokens_per_hour)
output_cost_per_token = (node_cost/hour × decode_time_fraction)
                      ÷ (output_tokens_per_hour)

This is also how you price your own service and how you attribute cost to tenants.


6. Per-tenant attribution#

In a shared cluster, you need to know who is spending what.

NAIVE: cost per request. Wrong — requests vary 1000x.

BETTER: weighted tokens
  tenant_cost = (input_tokens × input_cost) + (output_tokens × output_cost)

BEST: account for actual resource occupancy
  tenant_cost = input_tokens × input_cost
              + output_tokens × output_cost
              + kv_block_seconds × kv_cost      ← long-context requests
                                                  occupy memory

That third term matters: a request with a 100k-token context holds 25x the KV of a 4k request for its entire duration, denying that memory to others. Charging only for tokens under-charges long-context users substantially.

kv_cost per block-second =
    (node_cost/hour ÷ 3600) × (fraction of node cost attributable to memory)
  ÷ (total KV blocks on the node)

Emit kv_block_seconds per request. Few systems do, and it’s the missing term in most attribution models.


7. Comparing to buying an API#

The build-vs-buy calculation, done honestly:

SELF-HOSTED
  cost per M output tokens: $2.96 (from section 3)
  plus: engineering time (2-4 FTE for a serious platform)
        = $500k-1M/year amortized over your volume
  plus: procurement lead time, capacity risk

API PROVIDER
  cost per M output tokens: published price (varies)
  no engineering, no capacity risk, instant scaling
  but: data residency, latency (their region), model choice,
       rate limits, no customization

CROSSOVER
  self_hosted_total = (volume × $2.96/1e6) + engineering_cost
  api_total         = volume × api_price/1e6
  
  breakeven volume = engineering_cost / ((api_price - 2.96) / 1e6)

Compute it with your real numbers. The answer depends heavily on your volume, and providers' prices fall over time while your engineering cost doesn’t.

The honest caveat: providers achieve utilization you probably won’t, because they pool traffic from thousands of customers. Their marginal cost is genuinely lower than yours at moderate scale.


8. Production implications#

  • Track cost per million tokens as a first-class metric, per model and per tenant, trended daily.
  • Decompose it into the six factors from section 4, so you know which to attack.
  • Attribute to tenants, including KV occupancy. Publish it internally — visibility changes behavior.
  • Include the non-GPU costs. They’re 15-25% and get forgotten.
  • Report input and output costs separately. They differ by 5x.
  • Set a cost target alongside the SLO. “p95 TTFT < 800 ms at < $1.50/M output tokens” is a complete engineering objective.
  • Recompute after every optimization. It’s how you demonstrate value.

9. Common mistakes#

Using only the raw GPU cost. 4x too low.

Ignoring utilization. The single largest factor.

Not attributing prefill cost. Under-prices input-heavy workloads.

Ignoring aborted generations. 10-30% pure waste.

Charging per request in a multi-tenant system.

Not charging for KV occupancy. Under-charges long-context users.

Comparing your cost to an API price without including engineering cost.

Computing it once. It changes with every configuration and traffic shift.


10. Hands-on exercise#

A. Compute yours. For a system you run, work through all seven steps of section 3. What’s your actual cost per million output tokens? How far is it from the naive figure?

B. Decompose. Build the factor table from section 4 for your system. Which factor is largest? Which is most fixable?

C. Fix one. Pick the largest fixable factor and improve it. Measure the cost change. Express it as annual dollars.

D. Input vs output. Measure your prefill/decode time split and compute separate input and output token costs. What’s the ratio? Does it match your pricing (if you charge)?

E. Tenant attribution. Implement per-tenant cost attribution including kv_block_seconds. Which tenant is most expensive per token? Is it who you expected?

F. Build vs buy. Compute the crossover volume for your situation. Be honest about engineering cost.


11. Interview questions#

  1. Derive cost per token from infrastructure cost and throughput.
  2. Name the six factors between raw hardware cost and actual cost. Which are fixable?
  3. Why do output tokens cost more than input tokens? Give the ratio and the derivation.
  4. How would you attribute cost to tenants in a shared cluster?
  5. Why should you charge for KV occupancy, not just tokens?
  6. Walk through a build-vs-buy calculation honestly.
  7. Your cost per token is 4x your estimate. Where did the estimate go wrong?

12. Further reading#

  • [ESTABLISHED] Published LLM API pricing — compare input/output ratios across providers
  • [ESTABLISHED] Cloud GPU pricing (on-demand, reserved, spot) for your regions
  • [FUNDAMENTAL] FinOps principles for cloud cost attribution
  • Next: 04 — Autoscaling in production

↑↓ navigate↵ openesc close