Write-Behind Caching in Go: Async Flush Mechanics, Ordering Hazards, and the Durability Boundary
# Write-Behind Caching in Go: Async Flush Mechanics, Ordering Hazards, and the Durability Boundary
Write-through caching is safe but slow—every write blocks on two stores. Write-behind (also called write-back) pushes the database write off the critical path: you acknowledge to the caller after the cache write, then flush to the database asynchronously. The latency win is real, often 5–20× on write-heavy workloads where MongoDB round-trips dominate. But the durability contract is now probabilistic, and most implementations break silently under the failure modes that actually occur in production.
This article covers the mechanics you need to build a production-grade write-behind layer in Go: flush worker design, ordering invariants, failure handling, and where to draw the durability boundary given your SLA.
## What Changes in the Write Path
In write-through, the call graph is synchronous:
caller → cache.Set(k, v) → db.Upsert(k, v) → ack
In write-behind:
caller → cache.Set(k, v) → dirty-queue.Enqueue(k, v) → ack
↓ (async)
flush-worker → db.Upsert(k, v)
The caller sees lower latency. The database sees batched, coalesced writes. The system absorbs burst write load without proportional DB connection pressure. What you lose: any window between the cache write and the flush is a durability gap. A crash or eviction during that window loses data unless you build explicit recovery.
## Dirty Queue Design
The dirty queue is the central structure. Its semantics determine everything downstream.
**Keyed coalescing** is the first decision. If the same key is written three times before a flush, you only need to write the latest value to the database. A naive FIFO queue produces three DB writes. A keyed map collapses them to one but requires a lock or concurrent map with careful memory ordering.
type DirtyEntry struct {
Value any
DirtyAt time.Time
SeqNum uint64 // monotonic, per-key
}
type DirtyQueue struct {
mu sync.Mutex
entries map[string]*DirtyEntry
seq atomic.Uint64
}
func (q *DirtyQueue) Mark(key string, value any) {
seq := q.seq.Add(1)
q.mu.Lock()
q.entries[key] = &DirtyEntry{Value: value, DirtyAt: time.Now(), SeqNum: seq}
q.mu.Unlock()
}
func (q *DirtyQueue) Drain(limit int) map[string]*DirtyEntry {
q.mu.Lock()
defer q.mu.Unlock()
out := make(map[string]*DirtyEntry, min(limit, len(q.entries)))
count := 0
for k, v := range q.entries {
if count >= limit {
break
}
out[k] = v
delete(q.entries, k)
count++
}
return out
}
The sequence number matters for a subtle reason: if a flush worker drains a key and then a concurrent `Mark` call re-dirtifies it before the DB write completes, you need to know whether the in-flight value is still current. Compare `SeqNum` at drain time against the sequence at write completion; if a newer write arrived during the flush, re-enqueue rather than considering the key clean.
## Flush Worker Mechanics
A single flush goroutine avoids lock contention on the drain path but becomes a bottleneck under high write velocity. A pool of workers requires partitioning by key hash to preserve per-key ordering—two workers racing to flush the same key against MongoDB can produce last-write-wins anomalies that violate your update semantics.
func (c *Cache) flushWorker(ctx context.Context, partitionKeys <-chan string) {
ticker := time.NewTicker(c.flushInterval)
defer ticker.Stop()
for {
select {
case <-ticker.C:
batch := c.dirty.Drain(c.batchSize)
if len(batch) == 0 {
continue
}
c.flushBatch(ctx, batch)
case <-ctx.Done():
// Drain remaining before exit
final := c.dirty.Drain(math.MaxInt)
c.flushBatch(context.Background(), final)
return
}
}
}
The `context.Done()` drain is not optional. When your service receives SIGTERM, Kubernetes gives you a grace period (typically 30s). If you skip the final flush, every dirty key in the queue is lost. Most implementations miss this.
## Ordering Hazards
Ordering breaks in at least three ways in write-behind systems:
**Cross-key ordering.** If write A to key `user:1` causally precedes write B to key `order:99` (e.g., a balance deduction before an order creation), and your flush batches them independently, the database may reflect the order before the balance deduction. A reader hitting the DB directly during the flush window sees an inconsistent state. Mitigation: coerce related keys into the same flush batch by grouping them on a correlation ID, or accept that read-your-writes consistency requires reading from the cache layer, not the DB.
**Retry reordering.** When a DB write fails and you retry with exponential backoff, writes that entered the dirty queue after the failed key may flush before the retry succeeds. If the second write depends on the first (a foreign key, a version check), the retry produces a constraint error. Mitigation: per-key retry queues with head-of-line semantics—a failing key blocks its own subsequent writes but not unrelated keys.
**Eviction before flush.** Redis under memory pressure evicts keys. If your dirty-queue map lives in-process but the cache value lives in Redis, an eviction loses the value you intended to write. You must either store the value in the dirty queue itself (not just the key reference) or pin dirty keys with a Redis `PERSIST` or TTL extension. Storing the full value in the dirty queue increases heap pressure—a real tradeoff at high write volume.
## Flush Failure and the Durability Boundary
A failed flush leaves a key in a limbo state: the cache holds the new value, the database holds a stale value. Your retry policy determines the exposure window. Three strategies:
1. **Requeue on failure.** Re-mark the key as dirty with its current value. Simple, but if the failure is persistent (DB down), the dirty queue grows unbounded. Add a max-retry cap with a dead-letter channel that pages on-call.
2. **Circuit breaker on the flush path.** Wrap the DB write in a circuit breaker. When the breaker opens, stop draining the dirty queue entirely rather than accumulating retry failures. Resume drain when the breaker half-opens. This bounds DB connection exhaustion during an outage.
3. **WAL-backed dirty queue.** Write the dirty entry to an append-only log (a local file, Redis Stream, or SQS queue) before acknowledging the cache write. The flush worker reads from the WAL. Crash recovery replays unacknowledged entries. This is the only approach that provides durable write-behind; everything else accepts data loss on hard failure. The operational cost is a second write per cache mutation—evaluate whether the latency savings still justify write-behind over write-through at that point.
## Batching Strategy and DB Pressure
Write-behind's value proposition is batch efficiency. MongoDB's `bulkWrite` with `ordered: false` processes independent writes in parallel server-side. A flush of 200 coalesced writes over one `bulkWrite` call is fundamentally cheaper than 200 individual upserts—fewer round-trips, fewer index writes if updates are coalesced, lower oplog pressure.
Batch size is a tuning parameter with a ceiling. A batch of 2000 documents saturates the MongoDB write path and causes latency spikes for concurrent reads. In practice, 50–200 documents per flush at 100–500ms intervals is a reasonable starting range. Measure oplog lag and secondary replication latency as you increase batch size—they reveal when the primary is saturated before connection count does.
## Observability Requirements
Write-behind adds invisible lag to your write path. Without instrumentation, a growing dirty queue looks like healthy batching until it becomes a data loss event. Track:
* **Dirty queue depth** (gauge): sustained growth indicates flush worker can't keep up.
* **Flush latency** (histogram): p99 flush duration against the DB; spikes predict backpressure.
* **Keys requeued on failure** (counter): a non-zero rate is an active durability gap.
* **Eviction-before-flush events** (counter): requires Redis keyspace notification hooks on eviction of dirty keys—`notify-keyspace-events` with `Ke` flag.
* **Coalescing ratio** : `(marks / db-writes)`. A ratio near 1.0 means you're getting no coalescing benefit; reconsider whether write-behind is justified.
## Decision Framework
Before adopting write-behind caching:
**Accept write-behind when:**
* Write volume is high and bursty; DB write latency is the dominant bottleneck.
* Data loss of the flush window is tolerable under SLA (e.g., counters, analytics events, session metadata where eventual consistency is acceptable).
* You can afford WAL infrastructure for durability or can define an explicit loss boundary in your runbook.
**Reject write-behind when:**
* Reads hit the DB directly (not the cache); you cannot reason about the consistency window across all readers.
* The data is financial, transactional, or carries regulatory durability requirements without WAL.
* Your cache eviction policy is aggressive; you will lose dirty values under memory pressure without value-in-queue storage.
* You lack observability tooling to detect a silently growing dirty queue.
Write-behind is not a drop-in latency optimization. It is a durability tradeoff that requires explicit contracts, failure mode handling, and ongoing observability. Implemented with those constraints in view, it is one of the highest-leverage write optimizations available to a Go backend service under sustained write pressure.