π― The takeaway, first
Cache-aside is the default; write-through when you can't afford stale reads. But the real design isn't the happy path β it's the invalidation story and the thundering-herd plan. Think of the cache like a whiteboard next to the filing cabinet: reading the whiteboard is fast, but someone has to erase it when the files change, and if the whiteboard gets wiped, a hundred people rushing the cabinet at once will jam the door. If you can't say what happens when the cache is cold, you don't have a cache design β you have a hope.
π« Misconception, busted
"Adding a cache always makes things faster." A cold cache makes p99 worse than no cache (every request pays cache-lookup plus DB). A hot key can melt a single shard while the rest sit idle. And a stale-cache bug serves wrong data at lightning speed β confidently. The cache is a liability you manage, not a turbo button you bolt on.
Requirements
Say these out loud before drawing a single box.
Functional
GET / SET / DELETEwith TTL and LRU eviction- Scale out: add nodes without rehashing everything
- Key namespacing (
svc:entity:id) - Stats: hit rate, evictions, per-key hotness
Non-functional
- p99 < 5ms, 1M+ ops/s
- 99.99% availability β survive a node loss without avalanche
- Hit rate β₯ 95% at steady state
- Bounded staleness: define it, don't discover it
- No single key may take down a shard
Back-of-the-envelope math
Assumptions labeled.
| What | Assumption | Math |
|---|---|---|
| Read load | 1M reads/s, 95% hit rate | 50K DB reads/s on miss β your DB must handle that |
| Memory | 100M keys Γ ~1 KB | β 100 GB RAM β e.g. 8 nodes Γ 16 GB + replicas |
| Consistent hashing | ~1000 virtual nodes per physical node | adding a node moves only β 1/N of keys |
| Thundering herd | 1 hot key expires; 10K concurrent requests | naive: 10K DB queries on one row; singleflight: 1 |
| TTL jitter | base TTL 300s Β± 10% random | expirations spread over 60s instead of one cliff β herd prevention for free |
| Hot key | one key = 20% of reads on one shard | shard saturates while others idle β replicate hot keys or add local L1 |
Go deeper: why does adding a node only move 1/N of keys?
Consistent hashing places both nodes and keys on a ring; each key belongs to the first node clockwise from it. Adding a node only steals the key ranges immediately counter-clockwise of its positions. With plain modulo hashing (key % N), changing N reshuffles nearly every key β a full cache flush in disguise. Virtual nodes (each physical node claims ~1000 ring positions) smooth out the distribution so no node gets a lumpy share. Draw the ring in the interview; it's the expected picture.
Architecture
The centerpiece: the client library owns the ring.
flowchart LR
APP[App servers] --> CL[Cache client lib
consistent hash ring]
CL --> S1[Shard 1
primary + replica]
CL --> S2[Shard 2
primary + replica]
CL --> S3[Shard 3
primary + replica]
CL --> SN[Shard N ...]
S1 --> DB[(Primary DB)]
S2 --> DB
S3 --> DB
SN --> DB
APP -->|cache-aside miss| DB
APP -->|singleflight| SF[In-flight request
dedup map]
The client library hashes the key and talks directly to the owning shard β no proxy hop. Each shard has a replica for failover. On a miss, the app (not the cache) loads from the DB β that's what makes it cache-aside. The singleflight map sits in the app layer: 100 concurrent GETs for the same missing key collapse into one DB query, and 99 waiters share the result.
Component deep-dives
Cache-aside read + write (the default)
sequenceDiagram
participant A as App
participant C as Cache shard
participant D as DB
A->>C: GET user:42
alt hit
C-->>A: value (p99 < 5ms)
else miss
C-->>A: nil
A->>D: SELECT ... (singleflight dedups)
D-->>A: row
A->>C: SET user:42 row TTL 300s+jitter
A-->>A: return row
end
Note over A,D: write path: UPDATE DB first, then DEL cache key.
Delete (not update) β the next read repopulates. A
crashed-in-between DEL just means one stale window of TTL.
Write-through (when stale is not an option)
sequenceDiagram
participant A as App
participant C as Cache shard
participant D as DB
A->>C: SET user:42 newval
C->>D: write-through (sync)
D-->>C: ack
C-->>A: ack
Note over A,D: Every write pays DB latency. Reads never stale.
Use for: balances, inventory, anything money-shaped.
Thundering herd, defused by singleflight
sequenceDiagram
participant R1 as Req 1..100
participant A as App (singleflight map)
participant D as DB
R1->>A: GET hot:key (expired, 100 concurrent)
A->>A: key already in-flight? 99 wait on the promise
A->>D: exactly ONE query
D-->>A: row
A-->>R1: all 100 served from the one result
API + data model
GET /cache/{key} β 200 {value} | 404
SET /cache/{key} {value, ttl_s}
DELETE /cache/{key}
EXPIRE /cache/{key} {ttl_s}
key format: {service}:{entity}:{id} e.g. web:user:42
value: { v: 3, data: {...} } -- versioned! readers ignore v < expected
metadata: TTL = 300s + rand(Β±30s) -- jittered, always
Go deeper: why version the value?
Rolling deploys mean old and new code share the cache. A version field lets new code ignore (and overwrite) values written by old code instead of deserializing garbage. It's also your escape hatch for a bad serialization deploy: bump the version, old entries become invisible, no flush needed.
Trade-offs
| Decision | Option A | Option B | Pick |
|---|---|---|---|
| Write strategy | Cache-aside (fast writes, brief staleness) | Write-through (never stale, slow writes) | Aside by default; through for money-shaped data |
| Expiry | TTL only | Explicit invalidation | Both β TTL as the backstop, DEL on write as the primary |
| Engine | Redis (data structures, pub/sub) | Memcached (simpler, multithreaded) | Redis β pub/sub gives you cross-instance invalidation fan-out |
| Hashing | Client-side ring | Proxy (twemproxy/Envoy) | Client-side β one fewer hop at p99; proxy when clients are heterogeneous |
Failure modes
What I'd actually build
Opinionated. Steal this for the interview.
Stack: Redis Cluster, client-side consistent hashing in the app, singleflight in the request path, TTL = 300s Β± jitter always. Cache-aside with delete-on-write; write-through only for balances/inventory. Redis pub/sub invalidation channel so every app instance drops its local L1 on writes. Versioned values. Dashboards: hit rate, eviction rate, per-shard hotness, singleflight collapses/sec. Load-test the cold-start path quarterly β that's the one that pages you at 3am.
Interview tips
- Draw the ring. Consistent hashing with virtual nodes is the expected picture β draw it before they ask.
- Say "thundering herd" early. Then defuse it with singleflight. It's the highest-value 60 seconds in this interview.
- TTL jitter is a free point. One sentence β "jitter every TTL so nothing expires in a cliff" β signals production experience.
- Name what you won't cache. "Money-shaped data goes write-through or bypasses the cache" shows judgment.
- Close with the cold start. Ask them: "want me to walk through a full fleet restart?" β it reframes you as the operator, not the theorist.
π¬ Interactive: cache lab
Two demos. First: run GET/SET against cache-aside and write-through side by side. Then: fire a thundering herd and watch singleflight defuse it.
Demo A β write strategies, side by side
Cache-aside
Write-through
Try: SET on aside, then GET β note the invalidation. Then SET on write-through and compare the write path cost.
Demo B β thundering herd
The key hot:key just expired. 100 requests arrive at once.