2. Design a Distributed Rate Limiter
Enforce tenant quotas across servers while making burst behavior and outage policy explicit.
The brief
Design a rate limiter used by API servers in two regions. Explain the client-visible quota contract, how simultaneous requests consume allowance, and what happens when shared state is unavailable. Choose and justify an algorithm without assuming all limits need identical consistency.
- Serve 100,000 checks/second across 10,000 tenants, with one tenant producing 20% of checks.
- Each tenant has a sustained request limit and a separately configured burst allowance.
- Clients retry after rejection. Regions can lose contact for five minutes while still receiving traffic.
Constraints
- Check latency≤ 10 milliseconds
- Declare p95 additional server-side latency during normal operation.
- Quota contract
- Specify the maximum overshoot allowed across simultaneous requests and regional partitions.
- State outage policy
- State which requests fail open or closed and how recovered allowance avoids a second full burst.
What to cover
- 01
Algorithm and contract
Define tokens or counters, time semantics, rejection responses and retry guidance.
- 02
State and atomicity
Show key ownership, atomic updates and tenant isolation across API servers.
- 03
Capacity and hot keys
Estimate state size and load; handle the hot tenant and expired counters.
- 04
Partition walkthrough
Trace a regional partition and recovery, comparing accuracy with availability.
Worked designs
Explore the architecture and decisions, then build on an example with Coach.
Review rubric
AI feedback uses these criteria. Scores are practice feedback.
Quota semantics
Burst, time and retry behavior are precise.
Concurrent enforcement
Atomic updates and any regional overshoot are justified.
Capacity and isolation
Hot tenants and state lifetime fit the stated load.
Outage policy
Failure and recovery respect an explicit availability tradeoff.
Discussion
Share an approach, ask a question, or tag @Coach.
Loading discussion…