Below the API

Strings, Bytes and Runes

Basic Beginner 45 min Difficulty 2/5

Prerequisites 01

The idea in one minute#

A Go string is a two-word header — pointer and length — over a run of bytes that never change. It is not an array of characters. By convention those bytes are UTF-8, in which one character (a rune, a Unicode code point) takes one to four bytes. So len(s) counts bytes, s[i] is a byte, and for range s walks runes.

Because strings are immutable, slicing and passing them is free and safe. Because they are immutable, building them piece by piece is where the cost hides.

An analogy#

A printed page. Anyone can be handed a note saying “page 12, starting at character 40, for 15 characters” and read it; nobody can change the ink, so sharing is safe. To alter a word you must print a new page. A []byte is the same text written in pencil: editable, and therefore not safe to share blindly.

A picture#

flowchart TB
  subgraph H["Headers"]
    S["s = 'héllo, 世界'<br/>ptr, len 14"]
    SUB["sub = s[0:6]<br/>ptr, len 6  (no copy)"]
    B["b = []byte(s)<br/>ptr, len 14, cap 14  (copy)"]
  end
  subgraph MEM["Read-only bytes"]
    direction LR
    C0["h<br/>1 byte"] --- C1["é<br/>2 bytes"] --- C2["l l o , space<br/>5 bytes"] --- C3["世<br/>3 bytes"] --- C4["界<br/>3 bytes"]
  end
  HEAP[("New mutable array<br/>14 bytes")]
  S --> C0
  SUB --> C0
  B --> HEAP
  class S,SUB,B queue
  class C0,C1,C2,C3,C4 memory
  class HEAP compute

How it really works#

Three views of text#

TypeIslen countsMutableUse for
stringImmutable bytesBytesNoText you pass around, map keys
[]byteA slice of bytesBytesYesI/O, building, parsing
[]runeA slice of int32 code pointsCharactersYesPer-character editing (4 bytes each)
s := "héllo"
len(s)                       // 6: é is two bytes
s[1]                         // 0xC3: the first byte of é, not a character
for i, r := range s { }      // i is the byte offset, r is the rune: (0,'h') (1,'é') (3,'l') ...
utf8.RuneCountInString(s)    // 5

Invalid UTF-8 is allowed in a string; ranging over it yields U+FFFD for each bad byte.

What conversions cost#

OperationAllocates?
s[i:j]No. A new header over the same bytes
[]byte(s), string(b)Yes: a copy, because one side is mutable and the other must not change
s1 + s2Yes: a new string of the combined length
[]rune(s)Yes: decodes and stores 4 bytes per character
string(b) used only for a comparison or map lookup: m[string(b)]No. The compiler recognizes these and skips the copy
strconv.Itoa, fmt.SprintfYes. strconv.AppendInt(buf, n, 10) appends into your buffer instead

A substring keeps the whole original alive, exactly as a slice does: holding a 20-byte slice of a 10 MB string pins 10 MB. strings.Clone(sub) breaks the link.

Building strings#

s += x in a loop copies everything accumulated so far on every iteration: quadratic time and one allocation per step. Use:

  • strings.Builder — an append-only []byte that hands over its buffer as a string without a final copy. Call Grow(n) if you know the size.
  • bytes.Buffer — when you also need to read from it.
  • strings.Join(parts, sep) — computes the total size and allocates once.
  • fmt.Appendf(buf, ...), strconv.AppendX — format into an existing byte slice.

The standard toolbox#

strings and bytes mirror each other (Contains, Cut, Fields, Split, TrimSpace, HasPrefix, EqualFold, Index, …). strings.Cut(s, sep) is the idiomatic way to split once. unicode/utf8 decodes and validates; unicode classifies runes; strconv converts numbers.

Why this matters for AI code#

Tokenizers work on bytes. Byte-pair encoding (VI.04) starts from the 256 byte values and merges pairs, so any text — any language, emoji, invalid UTF-8 — is representable without an “unknown” token. A streaming LLM response can split a multi-byte character across two chunks: a server must buffer until utf8.FullRune says the character is complete, or the client sees garbage. And prompt assembly is string building in a hot path — exactly where += hurts.

Interning#

The unique package (Go 1.23) turns equal values into one canonical handle: unique.Make("gpt-large"). Comparing handles is a pointer comparison, and repeated strings — model names, tenant IDs, label values — are stored once.

Code#

// strings.go — bytes vs runes, what shares, what copies, and the cost of building.
package main

import (
	"fmt"
	"strings"
	"testing"
	"unicode/utf8"
	"unsafe"
)

func main() {
	s := "héllo, 世界"
	fmt.Printf("%q: %d bytes, %d runes, header %d bytes\n",
		s, len(s), utf8.RuneCountInString(s), unsafe.Sizeof(s))

	fmt.Print("bytes: ")
	for i := 0; i < len(s); i++ {
		fmt.Printf("%02x ", s[i])
	}
	fmt.Print("\nrunes: ")
	for i, r := range s {
		fmt.Printf("%d:%c(%d) ", i, r, utf8.RuneLen(r))
	}
	fmt.Println()

	// Substring: same bytes. Conversion: new bytes.
	sub := s[:6]
	b := []byte(s)
	fmt.Printf("\ns   data at %p\nsub data at %p  (shared)\nb   data at %p  (copied)\n",
		unsafe.StringData(s), unsafe.StringData(sub), unsafe.SliceData(b))

	// A character split across two chunks, as in a token stream.
	chunk1, chunk2 := []byte(s)[:9], []byte(s)[9:]
	fmt.Printf("\nchunk boundary inside a character: %q + %q\n", chunk1, chunk2)
	fmt.Println("is the tail of chunk 1 a complete rune?", utf8.FullRune(chunk1[8:]))

	// Building.
	parts := make([]string, 2000)
	for i := range parts {
		parts[i] = "tok"
	}
	bench := func(name string, f func() string) {
		r := testing.Benchmark(func(b *testing.B) {
			for i := 0; i < b.N; i++ {
				_ = f()
			}
		})
		fmt.Printf("%-22s %9d ns/op %6.0f allocs/op\n", name, r.NsPerOp(), testing.AllocsPerRun(5, func() { _ = f() }))
	}
	fmt.Println()
	bench("s += part", func() string {
		out := ""
		for _, p := range parts {
			out += p
		}
		return out
	})
	bench("strings.Builder", func() string {
		var sb strings.Builder
		for _, p := range parts {
			sb.WriteString(p)
		}
		return sb.String()
	})
	bench("Builder with Grow", func() string {
		var sb strings.Builder
		sb.Grow(len(parts) * 3)
		for _, p := range parts {
			sb.WriteString(p)
		}
		return sb.String()
	})
	bench("strings.Join", func() string { return strings.Join(parts, "") })

	// Map lookup with a []byte key: no allocation.
	m := map[string]int{"hello": 1}
	key := []byte("hello")
	fmt.Printf("\nm[string(key)] allocations: %.0f\n", testing.AllocsPerRun(100, func() { _ = m[string(key)] }))
}

Remember this#

  • A string is (pointer, length) over immutable bytes. len and indexing are in bytes.
  • range over a string yields runes; a rune is 1–4 bytes in UTF-8.
  • Slicing a string is free; converting to or from []byte copies.
  • Build with strings.Builder or Join, never += in a loop.

Try it#

  1. Run strings.go. How many times slower is += than Builder at 2,000 parts? At 20,000?
  2. Write reverse(s string) string that is correct for non-ASCII text.
  3. Write a StreamDecoder that accepts arbitrary byte chunks and emits only complete runes, holding back an incomplete tail. Test it with every split point of a multi-byte string.

Check yourself#

  1. What does len("héllo") return, and why?
  2. Why does []byte(s) have to copy?
  3. Why do tokenizers operate on bytes rather than characters?

↑↓ navigate ↵ open