System Design AI
← Question bank

Official · Infrastructure · advanced

Keep a regional ingress service alive

Survive a regional outage without confusing a health check with safe recovery.

Practice this questionDiscuss with @Coach

The question

Design ingress and regional failover for a business API. Clients cache addresses, requests may be retried and a recovered region can contain stale application state. State precisely what “failover complete” means.

What to cover

Routing and failure detection

Show request routing, health signals and who may withdraw or restore a region.

State and retry semantics

Describe replicated state, idempotency scope, partitions and failback fencing.

Load and recovery budget

Calculate headroom and break recovery time into detection, decision, routing and application readiness.

Outage exercise

Trace total region loss and a false-positive health failure; state abort and failback gates.

How your practice is reviewed
  • Routing contract and failure model (25 points): Makes DNS/cache behavior, partial failures and routing ownership explicit.
  • Write consistency under failover (30 points): Explains idempotency and ownership across partitions and recovery without assuming exactly-once networks.
  • Recovery and capacity evidence (25 points): Budget supports the claimed service recovery and full surviving-region peak load.
  • Safe exercise and failback (20 points): Observable gates and rollback limit false-positive failover and premature failback.

Discuss & learn

Ask the community, or tag @Coach for a contextual AI answer.