The idea in one minute#
A Go string is a two-word header — pointer and length — over a run of bytes that never
change. It is not an array of characters. By convention those bytes are UTF-8, in which one
character (a rune, a Unicode code point) takes one to four bytes. So len(s) counts bytes,
s[i] is a byte, and for range s walks runes.
Because strings are immutable, slicing and passing them is free and safe. Because they are immutable, building them piece by piece is where the cost hides.
An analogy#
A printed page. Anyone can be handed a note saying “page 12, starting at character 40, for 15
characters” and read it; nobody can change the ink, so sharing is safe. To alter a word you
must print a new page. A []byte is the same text written in pencil: editable, and therefore
not safe to share blindly.
A picture#
flowchart TB
subgraph H["Headers"]
S["s = 'héllo, 世界'<br/>ptr, len 14"]
SUB["sub = s[0:6]<br/>ptr, len 6 (no copy)"]
B["b = []byte(s)<br/>ptr, len 14, cap 14 (copy)"]
end
subgraph MEM["Read-only bytes"]
direction LR
C0["h<br/>1 byte"] --- C1["é<br/>2 bytes"] --- C2["l l o , space<br/>5 bytes"] --- C3["世<br/>3 bytes"] --- C4["界<br/>3 bytes"]
end
HEAP[("New mutable array<br/>14 bytes")]
S --> C0
SUB --> C0
B --> HEAP
class S,SUB,B queue
class C0,C1,C2,C3,C4 memory
class HEAP computeHow it really works#
Three views of text#
| Type | Is | len counts | Mutable | Use for |
|---|---|---|---|---|
string | Immutable bytes | Bytes | No | Text you pass around, map keys |
[]byte | A slice of bytes | Bytes | Yes | I/O, building, parsing |
[]rune | A slice of int32 code points | Characters | Yes | Per-character editing (4 bytes each) |
s := "héllo"
len(s) // 6: é is two bytes
s[1] // 0xC3: the first byte of é, not a character
for i, r := range s { } // i is the byte offset, r is the rune: (0,'h') (1,'é') (3,'l') ...
utf8.RuneCountInString(s) // 5
Invalid UTF-8 is allowed in a string; ranging over it yields U+FFFD for each bad byte.
What conversions cost#
| Operation | Allocates? |
|---|---|
s[i:j] | No. A new header over the same bytes |
[]byte(s), string(b) | Yes: a copy, because one side is mutable and the other must not change |
s1 + s2 | Yes: a new string of the combined length |
[]rune(s) | Yes: decodes and stores 4 bytes per character |
string(b) used only for a comparison or map lookup: m[string(b)] | No. The compiler recognizes these and skips the copy |
strconv.Itoa, fmt.Sprintf | Yes. strconv.AppendInt(buf, n, 10) appends into your buffer instead |
A substring keeps the whole original alive, exactly as a slice does: holding a 20-byte
slice of a 10 MB string pins 10 MB. strings.Clone(sub) breaks the link.
Building strings#
s += x in a loop copies everything accumulated so far on every iteration: quadratic time and
one allocation per step. Use:
strings.Builder— an append-only[]bytethat hands over its buffer as a string without a final copy. CallGrow(n)if you know the size.bytes.Buffer— when you also need to read from it.strings.Join(parts, sep)— computes the total size and allocates once.fmt.Appendf(buf, ...),strconv.AppendX— format into an existing byte slice.
The standard toolbox#
strings and bytes mirror each other (Contains, Cut, Fields, Split, TrimSpace,
HasPrefix, EqualFold, Index, …). strings.Cut(s, sep) is the idiomatic way to split
once. unicode/utf8 decodes and validates; unicode classifies runes; strconv converts
numbers.
Why this matters for AI code#
Tokenizers work on bytes. Byte-pair encoding (VI.04) starts from the 256 byte values and
merges pairs, so any text — any language, emoji, invalid UTF-8 — is representable without an
“unknown” token. A streaming LLM response can split a multi-byte character across two chunks:
a server must buffer until utf8.FullRune says the character is complete, or the client sees
garbage. And prompt assembly is string building in a hot path — exactly where += hurts.
Interning#
The unique package (Go 1.23) turns equal values into one canonical handle:
unique.Make("gpt-large"). Comparing handles is a pointer comparison, and repeated strings —
model names, tenant IDs, label values — are stored once.
Code#
// strings.go — bytes vs runes, what shares, what copies, and the cost of building.
package main
import (
"fmt"
"strings"
"testing"
"unicode/utf8"
"unsafe"
)
func main() {
s := "héllo, 世界"
fmt.Printf("%q: %d bytes, %d runes, header %d bytes\n",
s, len(s), utf8.RuneCountInString(s), unsafe.Sizeof(s))
fmt.Print("bytes: ")
for i := 0; i < len(s); i++ {
fmt.Printf("%02x ", s[i])
}
fmt.Print("\nrunes: ")
for i, r := range s {
fmt.Printf("%d:%c(%d) ", i, r, utf8.RuneLen(r))
}
fmt.Println()
// Substring: same bytes. Conversion: new bytes.
sub := s[:6]
b := []byte(s)
fmt.Printf("\ns data at %p\nsub data at %p (shared)\nb data at %p (copied)\n",
unsafe.StringData(s), unsafe.StringData(sub), unsafe.SliceData(b))
// A character split across two chunks, as in a token stream.
chunk1, chunk2 := []byte(s)[:9], []byte(s)[9:]
fmt.Printf("\nchunk boundary inside a character: %q + %q\n", chunk1, chunk2)
fmt.Println("is the tail of chunk 1 a complete rune?", utf8.FullRune(chunk1[8:]))
// Building.
parts := make([]string, 2000)
for i := range parts {
parts[i] = "tok"
}
bench := func(name string, f func() string) {
r := testing.Benchmark(func(b *testing.B) {
for i := 0; i < b.N; i++ {
_ = f()
}
})
fmt.Printf("%-22s %9d ns/op %6.0f allocs/op\n", name, r.NsPerOp(), testing.AllocsPerRun(5, func() { _ = f() }))
}
fmt.Println()
bench("s += part", func() string {
out := ""
for _, p := range parts {
out += p
}
return out
})
bench("strings.Builder", func() string {
var sb strings.Builder
for _, p := range parts {
sb.WriteString(p)
}
return sb.String()
})
bench("Builder with Grow", func() string {
var sb strings.Builder
sb.Grow(len(parts) * 3)
for _, p := range parts {
sb.WriteString(p)
}
return sb.String()
})
bench("strings.Join", func() string { return strings.Join(parts, "") })
// Map lookup with a []byte key: no allocation.
m := map[string]int{"hello": 1}
key := []byte("hello")
fmt.Printf("\nm[string(key)] allocations: %.0f\n", testing.AllocsPerRun(100, func() { _ = m[string(key)] }))
}
Remember this#
- A string is (pointer, length) over immutable bytes.
lenand indexing are in bytes. rangeover a string yields runes; a rune is 1–4 bytes in UTF-8.- Slicing a string is free; converting to or from
[]bytecopies. - Build with
strings.BuilderorJoin, never+=in a loop.
Try it#
- Run
strings.go. How many times slower is+=thanBuilderat 2,000 parts? At 20,000? - Write
reverse(s string) stringthat is correct for non-ASCII text. - Write a
StreamDecoderthat accepts arbitrary byte chunks and emits only complete runes, holding back an incomplete tail. Test it with every split point of a multi-byte string.
Check yourself#
- What does
len("héllo")return, and why? - Why does
[]byte(s)have to copy? - Why do tokenizers operate on bytes rather than characters?