System Design AI
All designs

Infrastructure & SREadvanced

13. Keep a regional ingress service alive

Survive a regional outage without confusing a health check with safe recovery.

The brief

Design ingress and regional failover for a business API. Clients cache addresses, requests may be retried and a recovered region can contain stale application state. State precisely what “failover complete” means.

  • Normal traffic is 60,000 requests/second, split evenly across two regions; bursts reach 90,000 requests/second.
  • A third failure domain is available for coordination but cannot host the full serving workload.
  • Some clients cache DNS for ten minutes. A payment-like write has an idempotency key and cannot be processed twice.

Constraints

Surviving-region capacity≥ 90,000 requests/second
Declare the peak request rate the surviving serving region must sustain.
Service recovery objective≤ 120 seconds
Declare recovery time for clients honoring the routing contract; explain separately how ten-minute DNS caches behave.
Write safety
Routing retries and regional recovery must not duplicate a committed write.
Cached addresses
Explain the user-visible degradation for stale-address clients rather than promising a DNS change instantly reaches them.

What to cover

  1. 01

    Routing and failure detection

    Show request routing, health signals and who may withdraw or restore a region.

  2. 02

    State and retry semantics

    Describe replicated state, idempotency scope, partitions and failback fencing.

  3. 03

    Load and recovery budget

    Calculate headroom and break recovery time into detection, decision, routing and application readiness.

  4. 04

    Outage exercise

    Trace total region loss and a false-positive health failure; state abort and failback gates.

Worked designs

No worked design has been published for this brief yet. You can start an attempt and share your approach in the discussion.

Review rubric

AI feedback uses these criteria. Scores are practice feedback.

Routing contract and failure model

Makes DNS/cache behavior, partial failures and routing ownership explicit.

25points

Write consistency under failover

Explains idempotency and ownership across partitions and recovery without assuming exactly-once networks.

30points

Recovery and capacity evidence

Budget supports the claimed service recovery and full surviving-region peak load.

25points

Safe exercise and failback

Observable gates and rollback limit false-positive failover and premature failback.

20points

Discussion

Share an approach, ask a question, or tag @Coach.

Share your interview experience

Published under your alias as a community report. Leave out interviewer names and confidential material.

Loading discussion…