Official · Infrastructure · advanced
Keep a regional ingress service alive
Survive a regional outage without confusing a health check with safe recovery.
The question
Design ingress and regional failover for a business API. Clients cache addresses, requests may be retried and a recovered region can contain stale application state. State precisely what “failover complete” means.
- Normal traffic is 60,000 requests/second, split evenly across two regions; bursts reach 90,000 requests/second.
- A third failure domain is available for coordination but cannot host the full serving workload.
- Some clients cache DNS for ten minutes. A payment-like write has an idempotency key and cannot be processed twice.
What to cover
Routing and failure detection
Show request routing, health signals and who may withdraw or restore a region.
State and retry semantics
Describe replicated state, idempotency scope, partitions and failback fencing.
Load and recovery budget
Calculate headroom and break recovery time into detection, decision, routing and application readiness.
Outage exercise
Trace total region loss and a false-positive health failure; state abort and failback gates.
How your practice is reviewed
- Routing contract and failure model (25 points): Makes DNS/cache behavior, partial failures and routing ownership explicit.
- Write consistency under failover (30 points): Explains idempotency and ownership across partitions and recovery without assuming exactly-once networks.
- Recovery and capacity evidence (25 points): Budget supports the claimed service recovery and full surviving-region peak load.
- Safe exercise and failback (20 points): Observable gates and rollback limit false-positive failover and premature failback.
Discuss & learn
Ask the community, or tag @Coach for a contextual AI answer.