Keep a regional ingress service alive
Survive a regional outage without confusing a health check with safe recovery.
Infrastructure, platform and reliability engineers.
Your approach: Use diagrams or prose to explain responsibilities, state, capacity and failure behavior. Show the calculations requested by the question.
The problem
Design ingress and regional failover for a business API. Clients cache addresses, requests may be retried and a recovered region can contain stale application state. State precisely what “failover complete” means.
- Normal traffic is 60,000 requests/second, split evenly across two regions; bursts reach 90,000 requests/second.
- A third failure domain is available for coordination but cannot host the full serving workload.
- Some clients cache DNS for ten minutes. A payment-like write has an idempotency key and cannot be processed twice.
Work within these constraints
Declare the peak request rate the surviving serving region must sustain.
Required target: ≥ 90,000 requests/second
Declare recovery time for clients honoring the routing contract; explain separately how ten-minute DNS caches behave.
Required target: ≤ 120 seconds
Routing retries and regional recovery must not duplicate a committed write.
Explain the user-visible degradation for stale-address clients rather than promising a DNS change instantly reaches them.
What to deliver
Routing and failure detection
Show request routing, health signals and who may withdraw or restore a region.
State and retry semantics
Describe replicated state, idempotency scope, partitions and failback fencing.
Load and recovery budget
Calculate headroom and break recovery time into detection, decision, routing and application readiness.
Outage exercise
Trace total region loss and a false-positive health failure; state abort and failback gates.