Safely converge 40,000 gateways while some sites are disconnected and old controllers can return.
Infrastructure, platform and reliability engineers.
Your approach: Use diagrams or prose to explain responsibilities, state, capacity and failure behavior. Show the calculations requested by the question.
Design a tenant-aware configuration controller for a managed gateway fleet. Explain the journey from an approved change to an observed running version, including an interrupted rollout. Identify the responsibilities of management and request-serving components; separate visual planes only if that helps your explanation.
Declare a target for 95% of connected gateways to observe an approved change, excluding intentionally paused canaries.
Required target: ≤ 300 seconds
Size the system for at least the initial 40,000 gateways; show the request-rate and state-retention calculation.
Required target: ≥ 40,000 gateways
Serving the last known-good configuration must not require a live management request.
A tenant must not read or apply another tenant’s configuration, including through retries and shared queues.
Diagram or describe approval, distribution, acknowledgement and serving paths, naming the state owner at each step.
Define desired versus observed state, version ordering, stale-leader fencing and reconnect behavior.
Estimate steady-state and reconnect-burst load with stated polling or push assumptions.
Walk through a faulty canary, a disconnected site and two active controllers; define measurable rollback/stop gates.
Compare two distribution approaches and state what you would monitor to detect a stuck rollout.