3. Design a Distributed Cache
Partition cached data, survive node loss and keep cache misses from overwhelming the source of truth.
The brief
Design a distributed cache for a read-heavy service with a separate durable database. Specify get, set and delete behavior, key ownership, expiry and eviction. Explain how clients find a key during membership changes and how the database is protected when popular entries disappear.
- Keep 100 million entries averaging 1 KB of value data; peak demand is 500,000 reads/second.
- Five percent of keys account for 80% of requests, and values may be stale for up to 30 seconds.
- Cache nodes can restart or be replaced during normal traffic. The database cannot sustain the full miss load.
Constraints
- Cache-hit latency≤ 5 milliseconds
- Declare p95 cache-hit server latency in one region.
- Freshness contract
- Explain invalidation, expiry and the circumstances under which a stale value may be returned.
- Database protection
- Bound concurrent fills and load on the source when nodes or hot keys disappear.
What to cover
- 01
Contract and key placement
Define cache APIs, namespace isolation and client routing.
- 02
Memory and replication
Estimate capacity including overhead and any replicas; explain eviction.
- 03
Membership change
Trace node loss and replacement, including misses and any moved keys.
- 04
Hot-key recovery
Walk through mass expiry and compare duplicate fills, request coalescing and stale serving.
Worked designs
Explore the architecture and decisions, then build on an example with Coach.
Review rubric
AI feedback uses these criteria. Scores are practice feedback.
Cache contract
Expiry, invalidation and durable ownership are distinct.
Placement and membership
Routing and node changes have coherent behavior.
Capacity and load protection
Memory, hot keys and miss admission are quantified.
Recovery choices
An outage walkthrough explains availability and freshness tradeoffs.
Discussion
Share an approach, ask a question, or tag @Coach.
Loading discussion…