Part of the series: Production-Grade Concurrent AI Systems in Go
→ Full code for this post: github.com/madmmas/go-concurrent-ai-systems/tree/part-15 → Diff from Part 14: compare/part-14...part-15 → Run it: go run ./cmd/news-processor -articles=20 -workers=3 -error-rate=0.5 inside arc-2-production/part-15-circuit-breaker
Part 13 added retries. Part 14 added rate limiting. These two together handle the routine failure modes of an LLM pipeline under normal operating conditions: occasional 429s from bursting over the rate limit, occasional 503s from transient provider hiccups. In both cases the provider recovers within seconds, and a retry after a short backoff succeeds.
But providers also fail for longer stretches. An incident that takes 10 minutes to resolve means 10 minutes of your workers making calls that time out, retrying, timing out again, and exhausting their retry budgets — all while holding worker slots that could be doing something else, burning through timeout budget on connections that will never complete, and potentially adding load to a provider that is already struggling to recover.
The circuit breaker pattern exists specifically for this scenario. When the failure rate crosses a threshold, the circuit breaker opens and all subsequent calls are rejected immediately — no network round trip, no timeout wait, no retry. The error comes back in microseconds instead of seconds. When the cooldown period elapses, one probe call is allowed through. If it succeeds, the circuit closes and normal operation resumes. If it fails, the circuit reopens for another cooldown cycle.
This is how you protect a pipeline from a broken provider without the pipeline itself breaking.
The Three States
The circuit breaker is a state machine with three states:
threshold failures
CLOSED ─────────────────────→ OPEN
↑ │
│ probe succeeds │ cooldown elapses
└──────────────── HALF-OPEN ←─┘
│
│ probe fails
└──────────→ OPEN
Closed is normal operation. Every call is allowed through. Failures are counted. When consecutive failures hit the threshold, the circuit transitions to Open.
Open is the fail-fast state. No calls are allowed through — Allow() returns ErrCircuitOpen immediately. The circuit stays Open for the cooldown duration. This gives the provider time to recover without the pipeline hammering it with failing requests.
Half-Open is the recovery probe. After the cooldown elapses, one call is permitted. If it succeeds, the circuit transitions to Closed and normal operation resumes. If it fails, the circuit reopens for another full cooldown cycle.
The Implementation
Three methods drive the state machine. Allow is called before every LLM call:
func (cb *CircuitBreaker) Allow() error {
cb.mu.Lock()
defer cb.mu.Unlock()
switch cb.state {
case stateClosed:
return nil
case stateOpen:
if time.Since(cb.openedAt) >= cb.cooldown {
fmt.Println("[circuit-breaker] cooldown elapsed → HALF-OPEN")
cb.state = stateHalfOpen
return nil // allow the probe
}
return ErrCircuitOpen
case stateHalfOpen:
return nil // allow one probe
}
return nil
}
The Open → Half-Open transition happens lazily inside Allow, not on a timer goroutine. When the cooldown has elapsed and a call arrives, Allow transitions the state and permits the call through. No background goroutine means one less thing that can leak or need cleanup.
RecordFailure and RecordSuccess update the state after each call:
func (cb *CircuitBreaker) RecordFailure() {
cb.mu.Lock()
defer cb.mu.Unlock()
cb.failures++
if cb.state == stateHalfOpen || cb.failures >= cb.threshold {
cb.state = stateOpen
cb.openedAt = time.Now()
fmt.Printf("[circuit-breaker] %d failures → OPEN (cooldown %v)\n",
cb.failures, cb.cooldown)
}
}
func (cb *CircuitBreaker) RecordSuccess() {
cb.mu.Lock()
defer cb.mu.Unlock()
if cb.state == stateHalfOpen {
fmt.Println("[circuit-breaker] probe succeeded → CLOSED")
}
cb.state = stateClosed
cb.failures = 0
}
The Half-Open path is strict: one failure in Half-Open reopens immediately, regardless of how many successes preceded it. This reflects the conservative position — the provider just came back from an outage, and one failure means it may not have fully recovered.
Wiring It Into the Pool
The circuit breaker wraps every LLM call in processArticle:
func (p *CBPool) processArticle(ctx context.Context, article model.Article) model.AIResult {
result := model.AIResult{ArticleID: article.ID}
// Check BEFORE the call — fail fast if open
if err := p.cb.Allow(); err != nil {
fmt.Printf("[article %d] circuit OPEN — rejected\n", article.ID)
result.Err = err
return result
}
if err := p.llm.Call(ctx, "Summarisation", article.ID); err != nil {
p.cb.RecordFailure()
result.Err = err
return result
}
p.cb.RecordSuccess()
result.Summary = "AI-generated summary"
if err := p.llm.Call(ctx, "Sentiment Analysis", article.ID); err != nil {
p.cb.RecordFailure()
result.Err = err
return result
}
p.cb.RecordSuccess()
result.Sentiment = "Positive"
return result
}
The placement of Allow() and RecordFailure()/RecordSuccess() matters. Allow() gates the call — if the circuit is open, the function returns without touching the network. RecordFailure() is called on every error the LLM returns, which is what drives the threshold counter. RecordSuccess() resets the counter and, in Half-Open, closes the circuit.
One detail worth noticing: RecordSuccess() is called after the first successful call, not after the article is fully processed. That means a successful summarisation followed by a failed sentiment analysis records one success and one failure in the same article. This is intentional — the circuit breaker tracks the health of individual calls, not articles.
What the Output Shows
Healthy provider — circuit stays closed throughout:
Circuit breaker: threshold=3 cooldown=800ms error-rate=0%
[article 3] processed (circuit: CLOSED)
[article 1] processed (circuit: CLOSED)
[article 2] processed (circuit: CLOSED)
...
Succeeded:8 | Failed:0 (circuit-open:0) | 1.352s
Unhealthy provider — watch the failure cascade, circuit opening, and fast rejection:
Circuit breaker: threshold=3 cooldown=800ms error-rate=50%
[2] Sentiment Analysis → 503 server error
[3] Summarisation → 503 server error
[4] Sentiment Analysis → 503 server error
[circuit-breaker] 3 failures → OPEN (cooldown 800ms)
[article 12] circuit OPEN — rejected
[article 13] circuit OPEN — rejected
[article 14] circuit OPEN — rejected
[article 15] circuit OPEN — rejected
[article 16] circuit OPEN — rejected
[article 17] circuit OPEN — rejected
[article 18] circuit OPEN — rejected
[article 19] circuit OPEN — rejected
[article 20] circuit OPEN — rejected
Succeeded:1 | Failed:19 (circuit-open:9) | 747ms
Three things stand out in this output.
First, the rejected articles return in microseconds. Articles 12 through 20 are processed and their results written to the results channel before article 9's in-flight LLM call even completes. There is no timeout wait, no retry budget consumed, no network round trip. ErrCircuitOpen is returned immediately from Allow().
Second, the total duration is 747ms for 20 articles. Without a circuit breaker, 20 articles with a 500ms timeout and 50% error rate would take seconds — each failing call waits up to 500ms before returning an error. The circuit breaker absorbs all of that into a single short burst of real calls followed by immediate rejections.
Third, the output distinguishes between two kinds of failure: Failed:19 (circuit-open:9) means 10 articles failed from actual LLM errors and 9 were rejected by the open circuit. These are meaningfully different. The 10 real errors contributed to opening the circuit. The 9 circuit-open rejections happened because the circuit was already protecting the pipeline.
Choosing Threshold and Cooldown
Two parameters drive the circuit breaker's behaviour. Getting them roughly right matters more than getting them precisely right.
Threshold — how many consecutive failures before opening. A low threshold (2–3) reacts quickly but trips on brief glitches that would have resolved with one retry. A high threshold (10+) gives the provider more benefit of the doubt but means more wasted calls before protection kicks in. For LLM providers with stable P99 latency and rare failure events, 3–5 is a reasonable starting point.
Cooldown — how long to stay open before probing. This should be longer than your provider's typical recovery time for transient incidents, but shorter than how long you are willing to reject new work. If your provider's status page shows incidents typically resolving in 2–3 minutes, a cooldown of 30–60 seconds is reasonable. Too short and you are probing a provider that has not finished recovering. Too long and you are rejecting work that could have succeeded minutes ago.
Both values should come from observing your specific provider's behaviour under real failure conditions, not from defaults in a tutorial.
The Relationship with Retries and Rate Limiting
The three resilience patterns from Parts 13, 14, and 15 each operate at a different scope:
| Pattern | Scope | Handles |
|---|---|---|
| Rate limiter (Part 14) | Every call | Prevents overloading the provider |
| Retry (Part 13) | Per article | Recovers from transient call failures |
| Circuit breaker (Part 15) | Provider level | Detects sustained outage, stops all calls |
They compose in a specific order. Rate limiting is the outermost layer — it caps what the pipeline sends before retries or circuit breakers are involved. Inside that, retries give individual articles multiple chances before giving up. The circuit breaker sits alongside retries, monitoring the aggregate failure signal across all articles: when enough individual failures stack up, it short-circuits everything.
In a production pipeline you would wire them together: the rate limiter gates Acquire before every call, the retry loop calls processWithRetry, and the circuit breaker sits inside processWithRetry, checked on every attempt. A rate limit 429 triggers a retry. Enough retries trigger the circuit breaker. The circuit breaker stops retries from firing at all.
What's Next
Parts 13 through 15 built the resilience layer for outbound LLM calls. The pipeline now handles transient failures, sustained outages, and rate pressure without falling over.
Part 16 turns inward and looks at what happens between stages when one part of the pipeline is slower than another. If the scrape stage produces articles faster than the embed stage can consume them, something has to give. Part 16 introduces backpressure — the mechanism that propagates the slow signal upstream so the producer slows down instead of the queue growing without bound.
See you in Part 16.
This is Part 15 of the series "Production-Grade Concurrent AI Systems in Go." Read Part 14 — Rate Limiting or continue to Part 16 — Backpressure: Letting the Queue Push Back.