The idea in one minute#
There are two memories — the host’s and the device’s — and a pointer into one is meaningless in the other. You must allocate on the right side and copy between them explicitly. The four things to know are: device allocation, pageable vs pinned host memory, unified memory, and why GPU programs pool their memory instead of allocating as they go.
A picture#
flowchart LR
subgraph H["Host"]
PG[("Pageable memory<br/>ordinary make / malloc")]
PIN[("Pinned memory<br/>page-locked")]
PG -->|"hidden staging copy"| PIN
end
subgraph D["Device"]
POOL[("Memory pool<br/>one big allocation")]
T1["tensor"] --- POOL
T2["tensor"] --- POOL
end
PIN <-->|"DMA over PCIe<br/>fast, can be async"| POOL
class PG warn
class PIN,POOL memory
class T1,T2 computeHow it really works#
Device allocation#
cudaMalloc asks the driver for device memory and returns a device pointer. It is slow —
tens of microseconds to milliseconds — and each process sees its own separate allocations.
There is no swap: when device memory is full, the next allocation fails with an out-of-memory
error.
Pageable vs pinned host memory#
Ordinary host memory is pageable: the operating system may move it or write it to disk at any time. The GPU cannot safely read memory that might move, so for a pageable source the driver first copies your data into an internal pinned buffer, then transfers that. Two copies.
Pinned (page-locked) memory is allocated with a promise that it will stay put
(cudaMallocHost). The GPU reads it directly:
| Host memory | Copies | Typical speed | Can be asynchronous |
|---|---|---|---|
| Pageable | 2 | ~6–12 GB/s | No |
| Pinned | 1 | ~20–50 GB/s | Yes |
Pinned memory is not free: it cannot be swapped, so pinning too much starves the rest of the system. Use it for the buffers you actually transfer.
For Go programs this has a specific consequence: a Go slice lives in pageable memory managed
by the garbage collector. For high-rate transfers, allocate a pinned buffer on the C side, and
wrap it as a Go slice with unsafe.Slice so Go code can fill it without another copy.
Unified memory#
cudaMallocManaged returns one pointer valid on both sides. The system migrates pages on
demand: touch it on the host and the page moves to the host; touch it in a kernel and it moves
to the device.
It is convenient and good for prototypes. The cost is hidden page faults and migrations at unpredictable moments, so performance-sensitive code usually manages memory explicitly.
Why everyone uses a memory pool#
Because cudaMalloc is slow and device memory fragments, serious GPU programs allocate one
large region at start-up and hand out pieces themselves. PyTorch does this (which is why
nvidia-smi shows more memory “used” than your tensors need), and LLM servers go further:
they reserve most of the GPU up front and divide it into fixed-size blocks.
Fixed-size blocks have a wonderful property: any free block can satisfy any request, so the pool cannot fragment.
Code#
A fixed-size block allocator — the core data structure of every LLM server’s memory manager.
// pool.go — why GPU servers pre-allocate and hand out fixed-size blocks.
package main
import (
"errors"
"fmt"
)
type BlockPool struct {
blockBytes int
free []int // stack of free block indices
total int
}
func NewBlockPool(deviceBytes, blockBytes int) *BlockPool {
n := deviceBytes / blockBytes
p := &BlockPool{blockBytes: blockBytes, total: n, free: make([]int, n)}
for i := range p.free {
p.free[i] = n - 1 - i
}
return p
}
var ErrOOM = errors.New("out of device memory")
// Alloc returns enough blocks for `bytes`. O(blocks), no searching, no fragmentation.
func (p *BlockPool) Alloc(bytes int) ([]int, error) {
need := (bytes + p.blockBytes - 1) / p.blockBytes
if need > len(p.free) {
return nil, ErrOOM
}
got := append([]int{}, p.free[len(p.free)-need:]...)
p.free = p.free[:len(p.free)-need]
return got, nil
}
func (p *BlockPool) Free(blocks []int) { p.free = append(p.free, blocks...) }
func (p *BlockPool) Used() float64 { return 1 - float64(len(p.free))/float64(p.total) }
func main() {
const MB = 1 << 20
pool := NewBlockPool(1024*MB, 2*MB) // a 1 GB "device", 2 MB blocks
a, _ := pool.Alloc(300 * MB)
b, _ := pool.Alloc(300 * MB)
c, _ := pool.Alloc(300 * MB)
fmt.Printf("after 3 allocs: %.0f%% used\n", pool.Used()*100)
pool.Free(a)
pool.Free(c) // free memory is now in two separate places...
d, err := pool.Alloc(500 * MB)
fmt.Printf("500 MB after freeing a and c: %d blocks, err=%v\n", len(d), err)
// ...yet a 500 MB request succeeds: blocks need not be neighbours.
_, err = pool.Alloc(500 * MB)
fmt.Println("another 500 MB:", err)
_ = b
}
A contiguous allocator in the same situation would fail the 500 MB request: 600 MB free, but in two 300 MB holes. The price of block allocation is that data is no longer contiguous, so the code using it needs a table from logical position to block — the same idea as an operating system’s page table.
Remember this#
- Host and device memory are separate address spaces. Copy explicitly.
- Pinned host memory makes transfers 2–4x faster and allows them to be asynchronous.
- Unified memory is convenient but hides page-migration costs.
- Allocate device memory once, in a pool. Fixed-size blocks never fragment.
Try it#
- Run
pool.go. Then write aContiguousPoolwith the same interface that hands out one continuous range per allocation, and replay the same sequence. Show that it fails. - Add a
refcountper block so two owners can share a block, freeing it only when both release. (This is how servers share an identical prompt prefix between requests.) - Why does
nvidia-smireport high memory use for a framework process even when it is idle?
Check yourself#
- What is pinned memory and why is it faster to transfer?
- Why do GPU programs avoid calling
cudaMallocin their hot path? - What does a fixed-size block pool give up in exchange for never fragmenting?