The idea in one minute#
A data race is two goroutines accessing the same variable at the same time, at least one
writing, with nothing ordering them. The Go memory model is the specification of what
“ordering them” means: a read is guaranteed to see a write only if the write happens before
the read through a chain of synchronization — a channel operation, a mutex, an atomic, a
sync.Once, a goroutine start.
Without that chain, the compiler and CPU are free to reorder, cache in a register, or tear the
access. There is no such thing as a benign data race. The race detector (-race) finds
them at run time, and it should be on in every test run.
An analogy#
Two people editing the same paper document by post, with no agreement about whose copy is current. Each works from whatever arrived last. The result is not “the latest edit wins” — it is an unpredictable mix, and sometimes a page of one version stapled to a page of another. A synchronization point is a meeting where one hands the document to the other: after it, both agree on what the document says.
A picture#
flowchart LR
subgraph G1["goroutine A"]
A1["x = 42"] --> A2["ch <- struct{}{}"]
end
subgraph G2["goroutine B"]
B1["<-ch"] --> B2["print(x) sees 42"]
end
A2 -->|"a send happens before<br/>the matching receive completes"| B1
subgraph R["No edge: a data race"]
C1["goroutine C: y = 1"]
C2["goroutine D: print(y)<br/>may see 0, may see 1,<br/>may never see the write"]
end
class A1,A2,B1,B2 compute
class C1,C2 warnHow it really works#
Happens-before#
Within one goroutine, statements happen in program order. Between goroutines, only these events create an ordering:
| Event | Guarantee |
|---|---|
go f() | The go statement happens before f starts |
| Channel send | Happens before the corresponding receive completes |
| Unbuffered receive | Happens before the corresponding send completes |
close(ch) | Happens before a receive that returns because the channel is closed |
mu.Unlock() | Happens before the next mu.Lock() returns |
once.Do(f) | The single run of f happens before any Do returns |
| Atomic operations | Behave as if executed in one total order; a Load sees the latest Store |
wg.Done() / wg.Wait() | The final Done happens before Wait returns |
| Goroutine exit | No guarantee. Nothing is ordered after a goroutine merely finishing |
If a write and a read of the same variable are not connected by such a chain, they are concurrent, and the program has a data race.
What actually goes wrong#
- Lost updates.
n++is load, add, store. Two goroutines interleave and one increment vanishes (IV.04’s first counter). - Stale reads, forever. A loop
for !done { }may be compiled to readdoneonce into a register. The flag set by another goroutine is never seen; the loop never ends. - Torn values. An interface, slice or string is several words. A reader can observe the pointer of one value with the length of another — and then index past the end of memory. This is how a data race becomes memory corruption in a “memory-safe” language.
- Broken invariants. A map being written while read can crash the runtime
(
concurrent map read and map write). - Reordering. The compiler and the CPU may reorder writes that have no dependency. On weakly ordered CPUs such as ARM64 — which is what many GPU servers and laptops now are — far more reorderings are visible than on x86. Code that “worked” on one may fail on the other.
The race detector#
go test -race ./...
go run -race .
go build -race -o app-race .It instruments every memory access and tracks, per memory location, which goroutines touched it under which synchronization. When it observes two unsynchronized accesses it prints both stack traces and where the goroutines were created.
- It reports only races that actually occur in that run: no false positives, but coverage depends on your tests and load.
- Cost: roughly 5–10× CPU and memory. Run it in CI, in integration tests, and for a canary under real traffic if you can afford it — not usually fleet-wide.
- A report is always a real bug. Fix it; do not argue with it.
Frequent races and their fixes#
| Race | Fix |
|---|---|
| A shared counter or flag | atomic type |
| A map written from several goroutines | Mutex or sync.Map |
| Appending to a shared slice | Mutex; or give each goroutine its own index: results[i] = ... |
Lazy initialization with an if x == nil check | sync.Once / sync.OnceValue |
| Reading a config while another goroutine reloads it | atomic.Pointer to an immutable value |
A test goroutine calling t.Fatal after the test ended | Wait for goroutines before returning |
Sharing rand.Rand, bytes.Buffer, a tokenizer with internal state | One per goroutine, or a lock |
Writing to different elements of a slice from different goroutines is not a race:
results[i] for distinct i are distinct variables. That makes “preallocate, each worker
fills its own slot” the simplest correct way to collect parallel results.
Double-checked locking#
if instance == nil { // unsynchronized read: a data race
mu.Lock()
if instance == nil { instance = build() }
mu.Unlock()
}
return instanceThis pattern is wrong in Go: the first read is not ordered after the write, so a goroutine may
see a non-nil pointer to an object whose fields are not yet visible. Use sync.Once, which
does the fast-path check with an atomic and is correct.
Things that are safe#
- Reading from many goroutines data that nobody writes any more — provided it was published through a synchronization event (passed on a channel, stored before the goroutines started).
- Immutable values: build completely, then share.
- Confinement: data only one goroutine ever touches needs no synchronization.
Code#
// races.go — a stale read that never ends (shown safely), lost updates, and correct publication.
package main
import (
"fmt"
"sync"
"sync/atomic"
"time"
)
type Config struct {
Model string
Timeout time.Duration
}
func main() {
// 1. Publication through a synchronization event: always correct.
var cfg *Config
ready := make(chan struct{})
go func() {
cfg = &Config{"large", time.Second} // written before the close...
close(ready)
}()
<-ready // ...which happens before this receive returns
fmt.Println("published via channel:", cfg.Model)
// 2. Each goroutine writes its own slice element: no race, no lock.
results := make([]int, 8)
var wg sync.WaitGroup
for i := range results {
wg.Add(1)
go func() {
defer wg.Done()
results[i] = i * i
}()
}
wg.Wait() // the Done calls happen before Wait returns, so reading results is safe
fmt.Println("disjoint slice elements:", results)
// 3. A flag: the atomic version is guaranteed to be seen.
var stop atomic.Bool
spins := 0
go func() {
time.Sleep(5 * time.Millisecond)
stop.Store(true)
}()
for !stop.Load() {
spins++
}
fmt.Println("atomic flag: the loop ended after", spins, "spins")
fmt.Println(" (with a plain bool the compiler may hoist the read out of the loop: an endless loop)")
// 4. Lost updates on a shared int versus an atomic.
var plain int64
var safe atomic.Int64
for g := 0; g < 8; g++ {
wg.Add(1)
go func() {
defer wg.Done()
for i := 0; i < 100000; i++ {
plain++ // DATA RACE: run with -race to see the report
safe.Add(1)
}
}()
}
wg.Wait()
fmt.Printf("8 x 100,000 increments: plain=%d atomic=%d\n", plain, safe.Load())
// 5. Lazy initialization done right.
loads := 0
model := sync.OnceValue(func() *Config {
loads++
return &Config{"lazy", time.Second}
})
for g := 0; g < 16; g++ {
wg.Add(1)
go func() { defer wg.Done(); _ = model() }()
}
wg.Wait()
fmt.Println("OnceValue: 16 goroutines asked, the loader ran", loads, "time")
}Run it once normally, then with go run -race races.go and read the report for case 4.
Remember this#
- A data race is unsynchronized concurrent access with at least one write. Its behaviour is undefined, not merely “sometimes stale”.
- Ordering between goroutines comes only from channel operations, locks, atomics,
Once,WaitGroupand goroutine start. - Run tests with
-race. Every report is a real bug. - Distinct slice elements, immutable data, and confinement are safe without locks.
Try it#
- Run
races.gowith-race. Which lines does the report name? - Replace the atomic flag in case 3 with a plain
booland build with optimizations. Does the loop end? (Try it on ARM and on x86 if you can.) - Write the broken double-checked lock and a test that hammers it with
-race.
Check yourself#
- Give three events that establish a happens-before edge between goroutines.
- Why is a race on an interface or slice value more dangerous than one on an
int? - Why is writing
results[i]from goroutineinot a data race?