Below the API

Host and Device Memory

Intermediate Intermediate 50 min Difficulty 3/5

Prerequisites 02, II.05

The idea in one minute#

There are two memories — the host’s and the device’s — and a pointer into one is meaningless in the other. You must allocate on the right side and copy between them explicitly. The four things to know are: device allocation, pageable vs pinned host memory, unified memory, and why GPU programs pool their memory instead of allocating as they go.

A picture#

flowchart LR
  subgraph H["Host"]
    PG[("Pageable memory<br/>ordinary make / malloc")]
    PIN[("Pinned memory<br/>page-locked")]
    PG -->|"hidden staging copy"| PIN
  end
  subgraph D["Device"]
    POOL[("Memory pool<br/>one big allocation")]
    T1["tensor"] --- POOL
    T2["tensor"] --- POOL
  end
  PIN <-->|"DMA over PCIe<br/>fast, can be async"| POOL
  class PG warn
  class PIN,POOL memory
  class T1,T2 compute

How it really works#

Device allocation#

cudaMalloc asks the driver for device memory and returns a device pointer. It is slow — tens of microseconds to milliseconds — and each process sees its own separate allocations. There is no swap: when device memory is full, the next allocation fails with an out-of-memory error.

Pageable vs pinned host memory#

Ordinary host memory is pageable: the operating system may move it or write it to disk at any time. The GPU cannot safely read memory that might move, so for a pageable source the driver first copies your data into an internal pinned buffer, then transfers that. Two copies.

Pinned (page-locked) memory is allocated with a promise that it will stay put (cudaMallocHost). The GPU reads it directly:

Host memoryCopiesTypical speedCan be asynchronous
Pageable2~6–12 GB/sNo
Pinned1~20–50 GB/sYes

Pinned memory is not free: it cannot be swapped, so pinning too much starves the rest of the system. Use it for the buffers you actually transfer.

For Go programs this has a specific consequence: a Go slice lives in pageable memory managed by the garbage collector. For high-rate transfers, allocate a pinned buffer on the C side, and wrap it as a Go slice with unsafe.Slice so Go code can fill it without another copy.

Unified memory#

cudaMallocManaged returns one pointer valid on both sides. The system migrates pages on demand: touch it on the host and the page moves to the host; touch it in a kernel and it moves to the device.

It is convenient and good for prototypes. The cost is hidden page faults and migrations at unpredictable moments, so performance-sensitive code usually manages memory explicitly.

Why everyone uses a memory pool#

Because cudaMalloc is slow and device memory fragments, serious GPU programs allocate one large region at start-up and hand out pieces themselves. PyTorch does this (which is why nvidia-smi shows more memory “used” than your tensors need), and LLM servers go further: they reserve most of the GPU up front and divide it into fixed-size blocks.

Fixed-size blocks have a wonderful property: any free block can satisfy any request, so the pool cannot fragment.

Code#

A fixed-size block allocator — the core data structure of every LLM server’s memory manager.

// pool.go — why GPU servers pre-allocate and hand out fixed-size blocks.
package main

import (
	"errors"
	"fmt"
)

type BlockPool struct {
	blockBytes int
	free       []int // stack of free block indices
	total      int
}

func NewBlockPool(deviceBytes, blockBytes int) *BlockPool {
	n := deviceBytes / blockBytes
	p := &BlockPool{blockBytes: blockBytes, total: n, free: make([]int, n)}
	for i := range p.free {
		p.free[i] = n - 1 - i
	}
	return p
}

var ErrOOM = errors.New("out of device memory")

// Alloc returns enough blocks for `bytes`. O(blocks), no searching, no fragmentation.
func (p *BlockPool) Alloc(bytes int) ([]int, error) {
	need := (bytes + p.blockBytes - 1) / p.blockBytes
	if need > len(p.free) {
		return nil, ErrOOM
	}
	got := append([]int{}, p.free[len(p.free)-need:]...)
	p.free = p.free[:len(p.free)-need]
	return got, nil
}

func (p *BlockPool) Free(blocks []int) { p.free = append(p.free, blocks...) }

func (p *BlockPool) Used() float64 { return 1 - float64(len(p.free))/float64(p.total) }

func main() {
	const MB = 1 << 20
	pool := NewBlockPool(1024*MB, 2*MB) // a 1 GB "device", 2 MB blocks

	a, _ := pool.Alloc(300 * MB)
	b, _ := pool.Alloc(300 * MB)
	c, _ := pool.Alloc(300 * MB)
	fmt.Printf("after 3 allocs: %.0f%% used\n", pool.Used()*100)

	pool.Free(a)
	pool.Free(c) // free memory is now in two separate places...
	d, err := pool.Alloc(500 * MB)
	fmt.Printf("500 MB after freeing a and c: %d blocks, err=%v\n", len(d), err)
	// ...yet a 500 MB request succeeds: blocks need not be neighbours.

	_, err = pool.Alloc(500 * MB)
	fmt.Println("another 500 MB:", err)
	_ = b
}

A contiguous allocator in the same situation would fail the 500 MB request: 600 MB free, but in two 300 MB holes. The price of block allocation is that data is no longer contiguous, so the code using it needs a table from logical position to block — the same idea as an operating system’s page table.

Remember this#

  • Host and device memory are separate address spaces. Copy explicitly.
  • Pinned host memory makes transfers 2–4x faster and allows them to be asynchronous.
  • Unified memory is convenient but hides page-migration costs.
  • Allocate device memory once, in a pool. Fixed-size blocks never fragment.

Try it#

  1. Run pool.go. Then write a ContiguousPool with the same interface that hands out one continuous range per allocation, and replay the same sequence. Show that it fails.
  2. Add a refcount per block so two owners can share a block, freeing it only when both release. (This is how servers share an identical prompt prefix between requests.)
  3. Why does nvidia-smi report high memory use for a framework process even when it is idle?

Check yourself#

  1. What is pinned memory and why is it faster to transfer?
  2. Why do GPU programs avoid calling cudaMalloc in their hot path?
  3. What does a fixed-size block pool give up in exchange for never fragmenting?

↑↓ navigate ↵ open