12. Roll out configuration to a disconnected fleet
Safely converge 40,000 gateways while some sites are disconnected and old controllers can return.
The brief
Design a tenant-aware configuration controller for a managed gateway fleet. Explain the journey from an approved change to an observed running version, including an interrupted rollout. Identify the responsibilities of management and request-serving components; separate visual planes only if that helps your explanation.
- 40,000 gateways across 800 tenants and four regions poll or receive changes. About 10% of sites can be disconnected for six hours.
- A gateway already serving traffic must keep serving its last known-good configuration while management is unavailable.
- Operators need canaries, rollback and a reliable answer to “which version is actually running?” Two controllers can temporarily believe they own a rollout.
Constraints
- Connected fleet convergence≤ 300 seconds
- Declare a target for 95% of connected gateways to observe an approved change, excluding intentionally paused canaries.
- Fleet sizing≥ 40,000 gateways
- Size the system for at least the initial 40,000 gateways; show the request-rate and state-retention calculation.
- Management outage
- Serving the last known-good configuration must not require a live management request.
- Tenant isolation
- A tenant must not read or apply another tenant’s configuration, including through retries and shared queues.
What to cover
- 01
Architecture and ownership
Diagram or describe approval, distribution, acknowledgement and serving paths, naming the state owner at each step.
- 02
Version and reconciliation protocol
Define desired versus observed state, version ordering, stale-leader fencing and reconnect behavior.
- 03
Capacity calculation
Estimate steady-state and reconnect-burst load with stated polling or push assumptions.
- 04
Bad rollout walkthrough
Walk through a faulty canary, a disconnected site and two active controllers; define measurable rollback/stop gates.
- 05
Tradeoff and operations
Compare two distribution approaches and state what you would monitor to detect a stuck rollout.
Worked designs
No worked design has been published for this brief yet. You can start an attempt and share your approach in the discussion.
Review rubric
AI feedback uses these criteria. Scores are practice feedback.
Reconciliation and version safety
Desired/observed state, monotonic versions or equivalent safeguards, idempotency and reconnect behavior form a coherent protocol.
Isolation and serving continuity
Tenant authentication/authorization and continued serving under management failure are addressed end to end.
Quantified capacity
Calculations expose polling, reconnect bursts and retention assumptions; claimed convergence is supported rather than merely repeated.
Rollout decisions and tradeoffs
Canary gates, rollback limits, ownership and an explicitly rejected alternative are operationally credible.
Discussion
Share an approach, ask a question, or tag @Coach.
Loading discussion…