Below the API

Profiling and eBPF

Basic Intermediate 50 min Difficulty 3/5

Prerequisites I.02, 02

The idea in one minute#

A profile answers “which code is using the resource?”. A CPU profiler interrupts the program about a hundred times a second and records the call stack; stacks that appear often are where time goes. Continuous profiling leaves that running in production at low overhead, so you can compare this week against last week or one version against another.

eBPF lets small, verified programs run inside the Linux kernel. It is how modern tools profile every process on a machine and produce traces and metrics without changing or even restarting the application.

An analogy#

To find out how office staff spend their day you could ask everyone to log every task (instrumentation: accurate, intrusive). Or you could walk through the office at random moments and note what each person is doing (sampling). After a few hundred walks you know where the hours go, and nobody had to change how they work.

A picture#

flowchart TB
  P["Running process"] -->|"interrupt 100 times per second"| S["Capture the call stack"]
  S --> AGG["Count identical stacks"]
  AGG --> FG["Flame graph<br/>width = share of samples"]
  FG --> DIFF["Diff two profiles<br/>before and after a deploy"]
  subgraph K["Linux kernel"]
    EB["eBPF program<br/>attached to a timer, syscall or function"]
  end
  EB -->|"stacks from every process,<br/>no code changes"| AGG
  class P compute
  class S,AGG queue
  class FG,DIFF memory
  class EB io

How it really works#

Kinds of profile#

ProfileSamplesAnswers
CPUStacks that were on-CPUWhat burns CPU?
Heap / allocationsStacks that allocated memoryWhat allocates, what is retained?
GoroutineAll goroutine stacksWhat is everyone waiting on?
Block / mutexStacks that waitedWhere is the contention?
Off-CPU / wall clockStacks that were not runningWhy is it slow while the CPU is idle?

Profiling Go#

Add one import and the profiles are served over HTTP:

import _ "net/http/pprof"   // registers /debug/pprof/* on the default mux
go tool pprof -http=:0 http://localhost:6060/debug/pprof/profile?seconds=30   # CPU
go tool pprof -http=:0 http://localhost:6060/debug/pprof/heap
go test -bench . -cpuprofile cpu.out && go tool pprof -http=:0 cpu.out

Reading a flame graph#

  • Each box is a function; the box below is its caller.
  • Width is the share of samples — the only thing that matters. Left-to-right order is alphabetical, not time.
  • Look for wide plateaus at the top: functions that are themselves expensive.
  • A diff flame graph colours what grew and what shrank between two profiles — the fastest way to find a regression.

Continuous profiling#

Collect a short profile from every process every few seconds, tag it with service and version, and store it. Overhead is typically a percent or two of CPU. What it buys:

  • “CPU per request rose 12% in v1.8.2” — diff the two versions.
  • “Which function costs the most across the whole fleet?” — the answer is often a serializer, a logger or a regular expression, not business logic.
  • With span IDs attached to samples, a slow span links straight to the code that ran in it.

Tools: Grafana Pyroscope, Parca, Polar Signals, and the profilers built into the large commercial platforms. OpenTelemetry’s Profiles signal entered public alpha in March 2026; it defines a common format based on pprof and ships an eBPF profiling agent (donated by Elastic) that runs as part of the Collector.

eBPF in one page#

An eBPF program is a small piece of bytecode loaded into the kernel, checked by a verifier (it must terminate and cannot touch arbitrary memory), and attached to a hook:

HookFires onUsed for
Timer / perf eventA fixed frequencyCPU profiling of every process
kprobe / tracepointA kernel function or eventSyscalls, scheduling, disk and network I/O
uprobeA function in a user program or libraryTLS reads/writes, HTTP handlers, GPU runtime calls
Socket / TC / XDPNetwork packetsFlow metrics, service maps

Programs write results to maps that a user-space agent reads.

What it gives observability:

  • Zero-code instrumentation. OpenTelemetry eBPF Instrumentation (OBI, derived from Grafana Beyla) watches HTTP, gRPC and SQL calls and emits spans and RED metrics for any language — useful for services you cannot or will not modify.
  • Whole-machine profiling, including the kernel and native libraries.
  • Network observability without sidecars (Cilium Hubble, Pixie, Coroot).

Its limits: you get what can be seen from outside — protocol-level spans, not your business attributes. It needs a recent kernel and elevated privileges. Encrypted traffic needs uprobes on the TLS library. Use it for breadth, and SDK instrumentation for depth.

Profiling and GPUs#

A CPU profiler shows the host side of an inference server: tokenization, Python or Go overhead, scheduling — and stacks that sit in a driver call waiting for the GPU. What happens on the device needs GPU tools (Nsight Systems, the PyTorch profiler, CUPTI). eBPF tools can attach uprobes to the CUDA runtime to time kernel launches and memory copies per process; this is an active and fast-moving area (V.02, VI.02).

Code#

A program with an obvious hotspot that profiles itself and prints where the samples went.

// hot.go — profile a program from inside and print the top of the CPU profile.
package main

import (
	"fmt"
	"os"
	"os/exec"
	"regexp"
	"runtime/pprof"
	"strings"
)

var re = regexp.MustCompile(`^[a-z]+_[0-9]+$`)

// slowValidate compiles nothing new, but regex matching is far costlier than it looks.
func slowValidate(keys []string) int {
	n := 0
	for _, k := range keys {
		if re.MatchString(k) {
			n++
		}
	}
	return n
}

func fastValidate(keys []string) int {
	n := 0
	for _, k := range keys {
		if i := strings.IndexByte(k, '_'); i > 0 && i < len(k)-1 {
			n++
		}
	}
	return n
}

func main() {
	keys := make([]string, 200000)
	for i := range keys {
		keys[i] = fmt.Sprintf("tenant_%d", i)
	}

	f, err := os.CreateTemp("", "cpu-*.prof")
	if err != nil {
		panic(err)
	}
	defer os.Remove(f.Name())
	pprof.StartCPUProfile(f)
	total := 0
	for i := 0; i < 20; i++ {
		total += slowValidate(keys) + fastValidate(keys)
	}
	pprof.StopCPUProfile()
	f.Close()
	fmt.Println("validated:", total)

	// Ask the Go toolchain to summarize the profile: flat = time in the function itself.
	out, err := exec.Command("go", "tool", "pprof", "-top", "-nodecount=8", f.Name()).CombinedOutput()
	if err != nil {
		fmt.Println("could not run go tool pprof:", err)
		return
	}
	fmt.Println(string(out))
}

Both functions do the same job. The profile shows almost all samples under slowValidate and the regexp package — which is the kind of finding continuous profiling turns up fleet-wide.

Remember this#

  • A profile is aggregated stack samples. Width in a flame graph is share of the resource.
  • Continuous profiling makes “what changed between versions?” a diff.
  • eBPF observes from the kernel: every process, no code changes, protocol-level detail only.
  • CPU profiles show the host side of GPU workloads; device time needs GPU-aware tools.

Try it#

  1. Run hot.go. Then add import _ "net/http/pprof", serve on :6060, and open the flame graph with go tool pprof -http=:0.
  2. Take a heap profile of a program that builds a large slice. Find the allocating line.
  3. Write down three things an eBPF-generated span cannot contain that an SDK span can.

Check yourself#

  1. What does the width of a box in a flame graph mean?
  2. Why is continuous profiling cheap enough to leave on?
  3. What does the eBPF verifier guarantee?

↑↓ navigate ↵ open