2. Design a Distributed Rate Limiter
Enforce tenant quotas across servers while making burst behavior and outage policy explicit.
Start with a template. Work through each step. Ask Coach when you need a second opinion.
Company tags are community-reported. Counts on cards show how many people reported that design.
Enforce tenant quotas across servers while making burst behavior and outage policy explicit.
Partition cached data, survive node loss and keep cache misses from overwhelming the source of truth.
Route API traffic with authentication, tenant limits and safe configuration rollout.
Reconcile desired workloads, schedule containers and recover from failed nodes safely.
Safely converge 40,000 gateways while some sites are disconnected and old controllers can return.
Survive a regional outage without confusing a health check with safe recovery.
A read-heavy service sends every request to a database that is slower and more expensive per read than memory, and the same few rows are asked for over and over.
A limit enforced per server is not a limit: ten servers each allowing a hundred requests a minute allow a thousand. And a counter per fixed window lets twice the limit through across a window boundary.
Millions of future actions — reminders, retries, expiries — must fire near their due time. Scanning everything every minute does not scale, and an in-process timer dies with the process.
Collect 5 million samples a second from 500,000 hosts, store them as time series, chart them and page people.
Run 10,000 jobs a second within two seconds of their time, at least once, with retries, fairness between tenants and exactly-once effects.
Every goal on 25 million open apps within two seconds: SSE or WebSockets, pub/sub to the right server, resume by seq, hot topics, reconnect storms and push.
Serve a product page a million times a second: a cheap query, replicas, Redis and the edge, with hot keys, stampedes and invalidation handled.
A live TV vote at a million writes a second: spread by key, buffer in Kafka, aggregate the hot counter in two stages, shed what can wait.
Files up to 50 GB up and down without a byte through the API: presigned URLs, resumable multipart, S3 events, scanning and CloudFront.
Accept in milliseconds, work in the background: leases and heartbeats, retries and dead letters, progress by SSE and webhooks, fairness across tenants.
An in-memory data store like Redis, from one event loop to Redis Cluster to a durable platform.
Workflows as code that survive any crash: event histories and replay, activities retried under timeouts, durable timers, sharded history.
A coordination service like ZooKeeper: znodes, sessions and watches, a ZAB quorum of five with observers, and the recipes for elections and locks.
An S3-inspired object store: bytes before metadata, a partitioned namespace with cache coherence, then repair, heat, lifecycle and safe change. Three interview boards with explicit assumptions and AWS source boundaries.
Build a queue around visibility leases and explicit acknowledgement, then add replicated partitions, FIFO groups, tenant fairness and bounded redrive: three self-contained interview boards.
Deny requests by caller address, range or URL at a global edge, with lists from governments, threat feeds and detection reaching 300 proxies in seconds.
Turn 10 pushes a second into job graphs on ephemeral runners, with fair queues per pool, leases fenced by attempt and live logs.
Ten terabytes of logs an hour from host agents through Kafka into tiered OpenSearch and S3, with regex search, live tail and exceptions grouped into issues.
Billions of files, hundreds of petabytes, one strongly consistent tree: a namespace partitioned by directory, chunk servers, replication, repair, erasure coding. One-hour boards for junior, senior and staff, with the theory behind them.
Thousands of GPUs shared by many teams: gang admission, topology-aware placement, quotas with lending, checkpointed preemption and failure recovery.
A 500 GB model on 1,000 GPU servers in minutes: chunks and hashes, a topology-aware swarm, signed manifests, bandwidth budgets and waves.
Detect, classify and repair failing GPU nodes (Xid errors, ECC, NVLink and EFA faults, stragglers, silent corruption) without draining the fleet.