System Design AI
Design brief
Infrastructure & SREAdvanced50 min suggested

Keep a regional ingress service alive

Survive a regional outage without confusing a health check with safe recovery.

Infrastructure, platform and reliability engineers.

Your approach: Use diagrams or prose to explain responsibilities, state, capacity and failure behavior. Show the calculations requested by the question.

The problem

Design ingress and regional failover for a business API. Clients cache addresses, requests may be retried and a recovered region can contain stale application state. State precisely what “failover complete” means.

  • Normal traffic is 60,000 requests/second, split evenly across two regions; bursts reach 90,000 requests/second.
  • A third failure domain is available for coordination but cannot host the full serving workload.
  • Some clients cache DNS for ten minutes. A payment-like write has an idempotency key and cannot be processed twice.

Work within these constraints

Surviving-region capacity

Declare the peak request rate the surviving serving region must sustain.

Required target: ≥ 90,000 requests/second

Service recovery objective

Declare recovery time for clients honoring the routing contract; explain separately how ten-minute DNS caches behave.

Required target: ≤ 120 seconds

Write safety

Routing retries and regional recovery must not duplicate a committed write.

Cached addresses

Explain the user-visible degradation for stale-address clients rather than promising a DNS change instantly reaches them.

What to deliver

1

Routing and failure detection

Show request routing, health signals and who may withdraw or restore a region.

2

State and retry semantics

Describe replicated state, idempotency scope, partitions and failback fencing.

3

Load and recovery budget

Calculate headroom and break recovery time into detection, decision, routing and application readiness.

4

Outage exercise

Trace total region loss and a false-positive health failure; state abort and failback gates.