System design interview Β· Infrastructure

Design a distributed cache

Everyone's favorite interview topic β€” and everyone's favorite production footgun. A cache doesn't make your system fast. It makes your system fast on average, and the average hides the stampedes.

🎯 The takeaway, first

Cache-aside is the default; write-through when you can't afford stale reads. But the real design isn't the happy path β€” it's the invalidation story and the thundering-herd plan. Think of the cache like a whiteboard next to the filing cabinet: reading the whiteboard is fast, but someone has to erase it when the files change, and if the whiteboard gets wiped, a hundred people rushing the cabinet at once will jam the door. If you can't say what happens when the cache is cold, you don't have a cache design β€” you have a hope.

🚫 Misconception, busted

"Adding a cache always makes things faster." A cold cache makes p99 worse than no cache (every request pays cache-lookup plus DB). A hot key can melt a single shard while the rest sit idle. And a stale-cache bug serves wrong data at lightning speed β€” confidently. The cache is a liability you manage, not a turbo button you bolt on.

Requirements

Say these out loud before drawing a single box.

Functional

  • GET / SET / DELETE with TTL and LRU eviction
  • Scale out: add nodes without rehashing everything
  • Key namespacing (svc:entity:id)
  • Stats: hit rate, evictions, per-key hotness

Non-functional

  • p99 < 5ms, 1M+ ops/s
  • 99.99% availability β€” survive a node loss without avalanche
  • Hit rate β‰₯ 95% at steady state
  • Bounded staleness: define it, don't discover it
  • No single key may take down a shard

Back-of-the-envelope math

Assumptions labeled.

WhatAssumptionMath
Read load1M reads/s, 95% hit rate50K DB reads/s on miss β€” your DB must handle that
Memory100M keys Γ— ~1 KBβ‰ˆ 100 GB RAM β†’ e.g. 8 nodes Γ— 16 GB + replicas
Consistent hashing~1000 virtual nodes per physical nodeadding a node moves only β‰ˆ 1/N of keys
Thundering herd1 hot key expires; 10K concurrent requestsnaive: 10K DB queries on one row; singleflight: 1
TTL jitterbase TTL 300s Β± 10% randomexpirations spread over 60s instead of one cliff β€” herd prevention for free
Hot keyone key = 20% of reads on one shardshard saturates while others idle β†’ replicate hot keys or add local L1
Go deeper: why does adding a node only move 1/N of keys?

Consistent hashing places both nodes and keys on a ring; each key belongs to the first node clockwise from it. Adding a node only steals the key ranges immediately counter-clockwise of its positions. With plain modulo hashing (key % N), changing N reshuffles nearly every key β€” a full cache flush in disguise. Virtual nodes (each physical node claims ~1000 ring positions) smooth out the distribution so no node gets a lumpy share. Draw the ring in the interview; it's the expected picture.

Architecture

The centerpiece: the client library owns the ring.

flowchart LR
    APP[App servers] --> CL[Cache client lib
consistent hash ring] CL --> S1[Shard 1
primary + replica] CL --> S2[Shard 2
primary + replica] CL --> S3[Shard 3
primary + replica] CL --> SN[Shard N ...] S1 --> DB[(Primary DB)] S2 --> DB S3 --> DB SN --> DB APP -->|cache-aside miss| DB APP -->|singleflight| SF[In-flight request
dedup map]

The client library hashes the key and talks directly to the owning shard β€” no proxy hop. Each shard has a replica for failover. On a miss, the app (not the cache) loads from the DB β€” that's what makes it cache-aside. The singleflight map sits in the app layer: 100 concurrent GETs for the same missing key collapse into one DB query, and 99 waiters share the result.

Component deep-dives

Cache-aside read + write (the default)

sequenceDiagram
    participant A as App
    participant C as Cache shard
    participant D as DB
    A->>C: GET user:42
    alt hit
        C-->>A: value (p99 < 5ms)
    else miss
        C-->>A: nil
        A->>D: SELECT ... (singleflight dedups)
        D-->>A: row
        A->>C: SET user:42 row TTL 300s+jitter
        A-->>A: return row
    end
    Note over A,D: write path: UPDATE DB first, then DEL cache key.
Delete (not update) β€” the next read repopulates. A
crashed-in-between DEL just means one stale window of TTL.

Write-through (when stale is not an option)

sequenceDiagram
    participant A as App
    participant C as Cache shard
    participant D as DB
    A->>C: SET user:42 newval
    C->>D: write-through (sync)
    D-->>C: ack
    C-->>A: ack
    Note over A,D: Every write pays DB latency. Reads never stale.
Use for: balances, inventory, anything money-shaped.

Thundering herd, defused by singleflight

sequenceDiagram
    participant R1 as Req 1..100
    participant A as App (singleflight map)
    participant D as DB
    R1->>A: GET hot:key (expired, 100 concurrent)
    A->>A: key already in-flight? 99 wait on the promise
    A->>D: exactly ONE query
    D-->>A: row
    A-->>R1: all 100 served from the one result

API + data model

GET    /cache/{key}            β†’ 200 {value} | 404
SET    /cache/{key} {value, ttl_s}
DELETE /cache/{key}
EXPIRE /cache/{key} {ttl_s}
key format:  {service}:{entity}:{id}      e.g.  web:user:42
value:       { v: 3, data: {...} }         -- versioned! readers ignore v < expected
metadata:    TTL = 300s + rand(Β±30s)       -- jittered, always
Go deeper: why version the value?

Rolling deploys mean old and new code share the cache. A version field lets new code ignore (and overwrite) values written by old code instead of deserializing garbage. It's also your escape hatch for a bad serialization deploy: bump the version, old entries become invisible, no flush needed.

Trade-offs

DecisionOption AOption BPick
Write strategyCache-aside (fast writes, brief staleness)Write-through (never stale, slow writes)Aside by default; through for money-shaped data
ExpiryTTL onlyExplicit invalidationBoth β€” TTL as the backstop, DEL on write as the primary
EngineRedis (data structures, pub/sub)Memcached (simpler, multithreaded)Redis β€” pub/sub gives you cross-instance invalidation fan-out
HashingClient-side ringProxy (twemproxy/Envoy)Client-side β€” one fewer hop at p99; proxy when clients are heterogeneous

Failure modes

Thundering herd: hot key expires under loadSingleflight + jittered TTLs + probabilistic early refresh (one request refreshes at 90% TTL)
Hot key melts one shardDetect via per-key counters; replicate the key to N shards; small local L1 in the app
Cold restart (fleet deploy wipes cache)Rolling restarts; warm from replicas; never flush the whole ring at once
Cache penetration: attackers request keys that never existBloom filter of known keys in front; short negative caching ("null" with 10s TTL)
Node dies mid-writeReplica promotion; client retries on the replica; 1/N of keys briefly miss β€” acceptable

What I'd actually build

Opinionated. Steal this for the interview.

Stack: Redis Cluster, client-side consistent hashing in the app, singleflight in the request path, TTL = 300s Β± jitter always. Cache-aside with delete-on-write; write-through only for balances/inventory. Redis pub/sub invalidation channel so every app instance drops its local L1 on writes. Versioned values. Dashboards: hit rate, eviction rate, per-shard hotness, singleflight collapses/sec. Load-test the cold-start path quarterly β€” that's the one that pages you at 3am.

Interview tips

πŸ”¬ Interactive: cache lab

Two demos. First: run GET/SET against cache-aside and write-through side by side. Then: fire a thundering herd and watch singleflight defuse it.

Demo A β€” write strategies, side by side

Cache-aside

client→cache→DB
hits 0 Β· misses 0 Β· DB value v1

Write-through

client→cache→DB
hits 0 Β· misses 0 Β· DB value v1

Try: SET on aside, then GET β€” note the invalidation. Then SET on write-through and compare the write path cost.

Demo B β€” thundering herd

The key hot:key just expired. 100 requests arrive at once.


clients100 reqs
cacheMISS
DBidle
0DB queries
0requests served
–slowest request
0cache hits
v2026.10.03-01