Part of the series: Production-Grade Concurrent AI Systems in Go
→ Full code for this post: github.com/madmmas/go-concurrent-ai-systems/tree/part-03 → Diff from Part 2: compare/part-02...part-03 → Run it: go run ./cmd/news-processor -articles=10 inside arc-1-foundations/part-03-race-conditions
Part 2 ended with a working pipeline and a data race. Ten goroutines appended to the same slice simultaneously, and the race detector flagged it:
WARNING: DATA RACE
Write at 0x00c0001b4000 by goroutine 8:
runtime.growslice(...)
Read at 0x00c0001b4000 by goroutine 12:
main.processAll(...)
The pipeline usually returned the right results anyway — which is what makes data races dangerous. They are not crashes. They are silent corruption that manifests unpredictably, only under certain timing conditions, and often only in production.
Part 3 fixes it. Then it shows what happens when the fix is applied in the wrong place.
The Fix: sync.Mutex
A mutex (mutual exclusion) ensures only one goroutine at a time can execute the code section it protects. The fix is three lines:
var mu sync.Mutex
// One goroutine at a time can append
mu.Lock()
results = append(results, result)
mu.Unlock()
The data race is gone. The pipeline produces consistent results regardless of scheduling order.
The full ProcessAll with the fix applied:
func (p *SafeProcessor) ProcessAll(articles []model.Article) ([]model.AIResult, time.Duration) {
start := time.Now()
var (
wg sync.WaitGroup
mu sync.Mutex
results = make([]model.AIResult, 0, len(articles))
)
for _, article := range articles {
wg.Add(1)
go func(a model.Article) {
defer wg.Done()
// AI work happens OUTSIDE the lock.
result := p.processArticle(a)
// Only the append is protected.
mu.Lock()
results = append(results, result)
mu.Unlock()
}(article)
}
wg.Wait()
return results, time.Since(start)
}
Run it with the race detector and nothing fires:
go test ./internal/... -race
# ok — no races reported
The Wrong Fix
Part 3 also demonstrates BadLockProcessor — what happens when the mutex wraps the LLM calls instead of just the append:
// BAD: lock wraps the entire LLM processing
mu.Lock()
result := p.processArticle(a) // LLM calls happen here
results = append(results, result)
mu.Unlock()
This is still correct for data races. But it serialises the pipeline entirely. Every goroutine must wait for the lock before calling the LLM, so all three LLM calls per article happen one at a time — the same as Part 1's sequential loop.
The benchmark makes this concrete:
BenchmarkBadLock — 10 articles: ~3.2s (re-serialised, like Part 1)
BenchmarkSafeLock — 10 articles: ~0.14s (mutex only on append, full concurrency)
The lesson: lock the minimum critical section. In this pipeline that is the append — the actual shared-memory write. Not the LLM call. Not the result construction. Just the append.
A Second Race You Won't See Coming
While verifying this part's code, a second race condition turned up — one that has nothing to do with the results slice.
Our simulator's LLMClient holds a *rand.Rand internally, used to pick a random latency for each call:
func (c *LLMClient) Call(task string, articleID int) {
spread := int64(c.cfg.MaxLatency - c.cfg.MinLatency)
latency := c.cfg.MinLatency + time.Duration(c.rng.Int63n(spread))
// ...
}
math/rand's *Rand type is not safe for concurrent use on its own. In Parts 1 and 2, this never mattered — each test or run created its own simulator, so nothing shared it across goroutines. But the moment several goroutines hold a reference to the same LLMClient and call Call() concurrently — exactly what's happening in this part's ProcessAll — they're all reading and advancing the same RNG state at once:
WARNING: DATA RACE
Read at 0x00c000080000 by goroutine 7:
math/rand.(*rngSource).Uint64()
...
simulator.(*LLMClient).Call()
simulator/llm.go:28
The fix is the same idea as the results slice, applied to the simulator itself — a mutex around just the RNG read:
func (c *LLMClient) Call(task string, articleID int) {
c.mu.Lock()
spread := int64(c.cfg.MaxLatency - c.cfg.MinLatency)
latency := c.cfg.MinLatency + time.Duration(c.rng.Int63n(spread))
c.mu.Unlock()
time.Sleep(latency) // outside the lock — no goroutine blocks another's sleep
}
This is worth sitting with for a second, because it's a slightly different lesson than the results-slice race. That race was a bug in our pipeline code. This one is a property of a dependency — math/rand's default source — that happens to be safe when used from one goroutine and unsafe the moment you share it across many. Any shared client, cache, or connection pool in a concurrent system deserves the same question: is this safe to call from multiple goroutines at once, or have I just never tested it that way?
sync.RWMutex: When Reads Outnumber Writes
sync.Mutex serialises all access — whether a goroutine is reading or writing. Only one goroutine at a time, regardless of what it is doing with the protected data.
That is conservative. Consider a results cache: once an article URL is processed, subsequent articles with the same URL can be served from cache without another LLM call. The cache is read by every goroutine on every article. It is written only when a new URL is first encountered.
Concurrent reads of the cache are safe — no goroutine modifies shared state during a read. Only the write requires exclusive access. sync.Mutex would still serialise all reads, causing goroutines to queue up behind each other even when they are doing nothing that conflicts.
sync.RWMutex is the right tool:
type CachedProcessor struct {
llm *simulator.LLMClient
mu sync.RWMutex
cache map[string]model.AIResult // keyed by article URL
}
The read path acquires a read lock — multiple goroutines can hold it simultaneously:
// Multiple goroutines can hold RLock at the same time.
// No goroutine can write to the cache while any reader holds RLock.
p.mu.RLock()
cached, hit := p.cache[a.URL]
p.mu.RUnlock()
The write path acquires an exclusive lock — all readers and other writers are blocked:
// Exclusive write lock — held only for the map write, not the LLM call.
p.mu.Lock()
p.cache[a.URL] = result
p.mu.Unlock()
The LLM call that produces the result happens between the read lock (cache miss confirmed) and the write lock (cache populated). No lock is held during the actual work.
What the Cache Demonstrates
With 9 articles across 3 unique URLs (each URL repeated 3 times), the output shows:
Processing article 1 (https://news.example.com/article/1)...
Processing article 2 (https://news.example.com/article/2)...
Processing article 3 (https://news.example.com/article/3)...
...all 9 goroutines check the cache simultaneously, all miss...
...3 goroutines write their results to cache...
Processed : 9 articles
Cache size: 3 unique URLs
Duration : 119ms
All 9 goroutines launch simultaneously. All 9 check the cache in the same moment, all holding RLock concurrently. All 9 miss, because the cache starts empty. All 9 call the LLM. All 9 then compete for the write lock to populate the cache.
This reveals a subtlety: the first time an article URL appears, the cache is cold, and the LLM is called regardless. The cache benefit is realised on the second and subsequent batches, or when you pre-warm the cache before processing begins. In a real news platform, the same wire-service article gets scraped by multiple sources — the cache pays for itself quickly.
The three tests confirm the correct behaviour:
// 10 articles, 3 unique URLs → exactly 3 LLM calls made, 3 cache entries
TestRWMutex_CacheHitReducesLLMCalls
// 50 goroutines reading the cache concurrently — race detector must pass
TestRWMutex_NoRaceUnderConcurrentReads
// All unique URLs → cache grows to len(articles), no entries lost
TestRWMutex_AllUniqueURLs
Three Mutex Patterns, One Benchmark
The benchmark file compares all three approaches directly:
BenchmarkBadLock — over-locked, throughput ≈ sequential
BenchmarkSafeLock — minimum critical section, full concurrency
BenchmarkRWMutex_AllUnique — RWMutex overhead vs Mutex (no cache benefit yet)
BenchmarkRWMutex_WithCacheHits — RWMutex benefit when reads dominate
Run them:
go test ./benchmarks/... -bench=. -benchmem -benchtime=2s -run='^$'
BenchmarkRWMutex_AllUnique is slower than BenchmarkSafeLock — when every URL is unique, all accesses are writes, and RWMutex has more overhead than plain Mutex for write-heavy workloads. BenchmarkRWMutex_WithCacheHits shows the reversal when reads dominate: concurrent RLock holders do not block each other, so throughput rises as cache hit rate increases.
Rule of thumb: reach for RWMutex when reads are frequent and concurrent reads are safe. If you are writing as often as reading, plain Mutex is cheaper.
When to Use Which
| Pattern | When |
|---|---|
sync.Mutex | Any shared state that goroutines both read and write. Default choice. |
sync.RWMutex | Read-heavy shared state where concurrent reads are safe. Benchmark before switching — overhead only pays off when read concurrency is high. |
Both are covered in Part 3. The remaining sync primitives — sync.Once for one-time initialisation, sync.Map for concurrent maps, and the full atomic API — appear in Arc 3.
What's Next
The mutex works. But the results slice is still shared state that every goroutine touches. Shared state — even when correctly protected — is the source of the bugs in Parts 2 and 3.
Part 4 shows what happens when the lock itself causes the problem: deadlocks. Three classic deadlock patterns, the Go runtime's detection output, and two rules that prevent all of them.
See you in Part 4.
This is Part 3 of the series "Production-Grade Concurrent AI Systems in Go." Read Part 2 — Goroutines and WaitGroups or continue to Part 4 — Deadlocks: When Goroutines Wait Forever.