Design a Container Orchestrator
Reconcile desired workloads, schedule containers and recover from failed nodes safely.
Infrastructure, platform and reliability engineers.
Your approach: Use diagrams or prose to explain responsibilities, state, capacity and failure behavior. Show the calculations requested by the question.
The problem
Design a container orchestrator inspired by Kubernetes. Define desired and observed workload state, scheduling, reconciliation and service discovery. Explain how running workloads behave when management is unavailable, and trace recovery when a node is partitioned but its containers may still be running.
- Manage 10,000 nodes and 200,000 running containers across three zones.
- Workloads declare CPU, memory, replica count and placement constraints; deploys happen continuously.
- Nodes can partition for ten minutes, and controllers can restart or briefly compete for ownership.
Work within these constraints
Declare supported node count and estimate status-update load.
Required target: ≥ 10,000 nodes
Reconciliation must be repeatable and must not allow competing controllers to corrupt desired state.
Distinguish rescheduling a stateless replica from safely moving a stateful workload with exclusive storage.
What to deliver
Resource and state model
Define workload identity, desired/observed state, ownership and versioning.
Scheduling and reconciliation
Show admission, placement, node agents, retries and discovery updates.
Capacity and fairness
Estimate control traffic and explain resource accounting, unschedulable work and tenant isolation.
Partition and rollout walkthrough
Trace controller loss, an uncertain node and a bad rollout; compare recovery speed with duplicate execution risk.