Survive a regional outage without confusing a health check with safe recovery.
Infrastructure, platform and reliability engineers.
Your approach: Use diagrams or prose to explain responsibilities, state, capacity and failure behavior. Show the calculations requested by the question.
Design ingress and regional failover for a business API. Clients cache addresses, requests may be retried and a recovered region can contain stale application state. State precisely what “failover complete” means.
Declare the peak request rate the surviving serving region must sustain.
Required target: ≥ 90,000 requests/second
Declare recovery time for clients honoring the routing contract; explain separately how ten-minute DNS caches behave.
Required target: ≤ 120 seconds
Routing retries and regional recovery must not duplicate a committed write.
Explain the user-visible degradation for stale-address clients rather than promising a DNS change instantly reaches them.
Show request routing, health signals and who may withdraw or restore a region.
Describe replicated state, idempotency scope, partitions and failback fencing.
Calculate headroom and break recovery time into detection, decision, routing and application readiness.
Trace total region loss and a false-positive health failure; state abort and failback gates.