PidokuInfra

Tokenization and Its Consequences

Basic Beginner 1h Difficulty 2/5 Topic 02 of 15

Prerequisites III.06


1. What is it?#

Converting text to integers the model can process, and back again.

"Hello, world!"  →  [9906, 11, 1917, 0]  →  model  →  [2028, 374, ...]  →  "This is..."

Modern LLMs use subword tokenization: common words are single tokens, rare words split into pieces, and any string can be represented (no unknown tokens).


2. Why does it exist?#

Two bad alternatives frame the choice:

Character-level:  vocabulary ~256, sequences 5x longer
                  → 5x more decode steps, 5x more KV cache. Fatal.
Word-level:       vocabulary ~1M+, huge embedding table, and unknown words break it.

Subword tokenization sits in between: ~32k-256k vocabulary, ~4 characters per token for English, and complete coverage.

For inference specifically: tokens are your unit of work, your unit of cost, and your unit of memory. Everything is denominated in them, so how text maps to them directly determines what users pay and how much capacity you need.


3. Simple analogy#

Compression by dictionary. A shorthand dictionary maps common phrases to single symbols. Frequent phrases get short codes; unusual ones get spelled out letter by letter. Efficient for the language the dictionary was built for, and inefficient for anything else — which is exactly why tokenizers trained mostly on English are wasteful for Thai, Telugu, or minified JSON.


4. Tiny example#

Go
// tokens.go — ask a running server how it tokenizes text.
// Works with vLLM (POST /tokenize); for llama.cpp's server send {"content": s} instead.
package main

import (
	"bytes"
	"encoding/json"
	"fmt"
	"net/http"
	"os"
	"unicode/utf8"
)

func countTokens(base, model, s string) (int, error) {
	body, _ := json.Marshal(map[string]any{"model": model, "prompt": s, "add_special_tokens": false})
	resp, err := http.Post(base+"/tokenize", "application/json", bytes.NewReader(body))
	if err != nil {
		return 0, err
	}
	defer resp.Body.Close()
	var out struct {
		Tokens []int `json:"tokens"`
	}
	err = json.NewDecoder(resp.Body).Decode(&out)
	return len(out.Tokens), err
}

func main() {
	base, model := "http://localhost:8000", "meta-llama/Meta-Llama-3-8B-Instruct"
	if len(os.Args) > 2 {
		base, model = os.Args[1], os.Args[2]
	}
	for _, s := range []string{
		"Hello world",
		"supercalifragilisticexpialidocious",
		"def fibonacci(n):",
		"こんにちは世界",
		"🎉🎊🎈",
		"1234567890",
	} {
		n, err := countTokens(base, model, s)
		if err != nil || n == 0 {
			fmt.Println("tokenize failed:", err)
			return
		}
		chars := utf8.RuneCountInString(s)
		fmt.Printf("%3d chars → %3d tokens  (%.1f c/t)  %s\n", chars, n, float64(chars)/float64(n), s)
	}
}

Representative output:

 11 chars →   2 tokens  (5.5 c/t)  Hello world
 34 chars →   9 tokens  (3.8 c/t)  supercalifragilisticexpialido
 17 chars →   6 tokens  (2.8 c/t)  def fibonacci(n):
  7 chars →   5 tokens  (1.4 c/t)  こんにちは世界
  3 chars →   9 tokens  (0.3 c/t)  🎉🎊🎈
 10 chars →   4 tokens  (2.5 c/t)  1234567890

Look at the spread: 5.5 to 0.3 characters per token. Japanese costs ~4x more tokens per character than English. Emoji cost ~3 tokens each. This is a real fairness and cost issue: the same message in different languages costs users very different amounts.


5. Technical explanation#

BPE (Byte-Pair Encoding), the dominant algorithm#

Training:

1. Start with a vocabulary of all 256 bytes.
2. Count all adjacent pairs in the training corpus.
3. Merge the most frequent pair into a new token.
4. Repeat until the vocabulary reaches the target size.

Encoding: apply the learned merges greedily in the order they were learned.

"lower"  →  l o w e r
merge (e,r)→er:   l o w er
merge (l,o)→lo:   lo w er
merge (lo,w)→low: low er
→ ["low", "er"]

Variants: byte-level BPE (GPT-2 onward) operates on bytes so nothing is unrepresentable; SentencePiece/Unigram (Llama 1/2, T5) uses a probabilistic model instead of greedy merges; WordPiece (BERT) is BPE with a likelihood-based merge criterion.

Special tokens and chat templates#

<|begin_of_text|>  <|eot_id|>  <|start_header_id|>  <|end_header_id|>

Chat models expect a specific format:

<|begin_of_text|><|start_header_id|>system<|end_header_id|>

You are helpful.<|eot_id|><|start_header_id|>user<|end_header_id|>

Hello<|eot_id|><|start_header_id|>assistant<|end_header_id|>

Getting the template wrong degrades quality dramatically with no error. The model was trained on this exact format; deviating puts it out of distribution. Always use the tokenizer’s apply_chat_template, never hand-build the string.

Incremental detokenization — a real bug source#

You cannot simply decode each token independently and concatenate:

Go
// WRONG
for _, id := range generated {
	send(string(vocab[id])) // a token is BYTES, not a string: this splits multi-byte characters
}

Two problems:

  1. Multi-byte characters span tokens. A single Japanese character or emoji may be split across 2-4 tokens; decoding each alone produces replacement characters (�).
  2. Leading spaces. Many tokenizers encode " world" as one token; decoding it alone may or may not include the space depending on the decoder’s settings.

The correct approach maintains state:

Go
// Correct pattern: buffer bytes, emit only complete UTF-8 characters.
var pending []byte
for _, id := range generated {
	pending = append(pending, vocab[id]...) // vocab[id] is the token's raw bytes
	n := 0
	for n < len(pending) && utf8.FullRune(pending[n:]) { // is a whole character available?
		_, size := utf8.DecodeRune(pending[n:])
		n += size
	}
	if n > 0 {
		send(string(pending[:n]))
		pending = pending[n:] // an incomplete tail (if any) waits for the next token
	}
}

Production engines implement a more efficient incremental decoder, but the principle is the same: decode the accumulated sequence, emit the delta, and hold back incomplete characters.

Stop sequences#

Users specify stop strings ("\n\n", “”). But the model emits tokens, and a stop string may span token boundaries or appear mid-token. Correct implementation:

1. Accumulate decoded text.
2. Check for the stop string in the accumulated text.
3. When found, truncate the output at the stop position — which may be mid-token.
4. Also handle: a partial stop string at the end (don't emit it yet; it might complete).

That last point is subtle: if the stop string is “END” and you’ve just generated “EN”, you must not stream “EN” to the user yet, because the next token might make it a stop sequence that should be removed.


6. Under the hood#

Tokenization performance:

Pure Python BPE:              ~0.1-1 MB/s        ← unusable at scale
HuggingFace `tokenizers` (Rust): ~10-100 MB/s
tiktoken (Rust):              ~50-200 MB/s

For a 100k-token prompt (~400 KB of text), that’s 4 seconds in Python vs 4 ms in Rust. Always use the fast tokenizer (use_fast=True, which is the default for most models).

Tokenization is CPU work on the critical path of TTFT. At high QPS with long prompts, it can consume a full core per few hundred requests/sec. Parallelize it in a thread pool (the Rust tokenizer releases the GIL).


7. Performance implications#

  • Token count determines everything: prefill FLOPs, KV memory, decode steps, cost.
  • Language affects cost by 2-5x for the same semantic content.
  • Code and structured data tokenize poorly — JSON with lots of punctuation and whitespace can be 2 characters per token.
  • Tokenization latency matters at long prompts: 100k tokens is a few ms in Rust, seconds in Python.
  • Vocabulary size affects the model: bigger vocab = fewer tokens per text (cheaper decode) but a bigger embedding table and LM head (more memory, more FLOPs per step). Modern models have moved to 128k-256k vocabularies partly for this reason.

8. Production implications#

  • Bill and rate-limit by tokens, not characters or requests. Count them server-side; clients cannot be trusted.
  • Provide a token-counting endpoint so clients can estimate cost before sending.
  • Cap max_tokens and input length in tokens.
  • Use apply_chat_template. Never hand-construct prompts.
  • Test detokenization with non-Latin scripts and emoji in CI. This bug ships constantly.
  • Handle partial stop sequences correctly or you’ll leak stop strings into user output.
  • Consider the multilingual cost asymmetry in pricing if you serve global users.

9. Common mistakes#

Decoding tokens independently. Mojibake for non-English text.

Hand-building chat prompts. Silent quality degradation.

Estimating tokens as chars/4. Wrong by 2-5x for non-English and code.

Using a slow Python tokenizer. Seconds of TTFT on long prompts.

Not accounting for special tokens in length limits. The template adds 20-40 tokens.

Streaming a partial stop sequence. The user sees “END” flash before it’s removed.

Assuming the tokenizer matches the model. Using the wrong tokenizer produces valid-looking garbage.


10. Hands-on exercise#

A. Measure the language tax. Take the same paragraph translated into 8 languages. Tokenize each with a model’s tokenizer. Build a table of tokens per character. What’s the ratio between the most and least efficient?

B. Break detokenization. Generate text containing emoji and Japanese, decoding token by token independently. Observe the mojibake. Then implement the incremental decoder from section 5 and verify it’s fixed.

C. Tokenizer benchmark. Time tokenization of a 100k-character document with the fast and slow tokenizers. Report the ratio.

D. Chat template. Print tok.apply_chat_template(messages, tokenize=False) for a conversation. Count the overhead tokens. How much does the template cost for a 10-turn conversation?

E. Stop sequence edge case. Implement stop-sequence handling and write a test where the stop string spans two tokens and where a partial match occurs at the end of generation.


11. Interview questions#

  1. Why do LLMs use subword tokenization rather than characters or words?
  2. Explain BPE training and encoding.
  3. Why can’t you decode generated tokens one at a time and concatenate?
  4. Why does the same content cost more in Japanese than in English?
  5. What is a chat template and what happens if you get it wrong?
  6. How do you correctly handle a stop sequence that spans token boundaries?
  7. What are the tradeoffs of a larger vocabulary?

12. Further reading#

  • [FUNDAMENTAL] Sennrich et al., “Neural Machine Translation of Rare Words with Subword Units” (BPE, 2015)
  • [FUNDAMENTAL] Karpathy, “Let’s build the GPT Tokenizer” (video) and minbpe
  • [REFERENCE] HuggingFace tokenizers library docs; tiktoken
  • Next: 03 — Prefill vs decode

↑↓ navigate↵ openesc close