13. Keep a regional ingress service alive
Survive a regional outage without confusing a health check with safe recovery.
The brief
Design ingress and regional failover for a business API. Clients cache addresses, requests may be retried and a recovered region can contain stale application state. State precisely what “failover complete” means.
- Normal traffic is 60,000 requests/second, split evenly across two regions; bursts reach 90,000 requests/second.
- A third failure domain is available for coordination but cannot host the full serving workload.
- Some clients cache DNS for ten minutes. A payment-like write has an idempotency key and cannot be processed twice.
Constraints
- Surviving-region capacity≥ 90,000 requests/second
- Declare the peak request rate the surviving serving region must sustain.
- Service recovery objective≤ 120 seconds
- Declare recovery time for clients honoring the routing contract; explain separately how ten-minute DNS caches behave.
- Write safety
- Routing retries and regional recovery must not duplicate a committed write.
- Cached addresses
- Explain the user-visible degradation for stale-address clients rather than promising a DNS change instantly reaches them.
What to cover
- 01
Routing and failure detection
Show request routing, health signals and who may withdraw or restore a region.
- 02
State and retry semantics
Describe replicated state, idempotency scope, partitions and failback fencing.
- 03
Load and recovery budget
Calculate headroom and break recovery time into detection, decision, routing and application readiness.
- 04
Outage exercise
Trace total region loss and a false-positive health failure; state abort and failback gates.
Worked designs
No worked design has been published for this brief yet. You can start an attempt and share your approach in the discussion.
Review rubric
AI feedback uses these criteria. Scores are practice feedback.
Routing contract and failure model
Makes DNS/cache behavior, partial failures and routing ownership explicit.
Write consistency under failover
Explains idempotency and ownership across partitions and recovery without assuming exactly-once networks.
Recovery and capacity evidence
Budget supports the claimed service recovery and full surviving-region peak load.
Safe exercise and failback
Observable gates and rollback limit false-positive failover and premature failback.
Discussion
Share an approach, ask a question, or tag @Coach.
Loading discussion…