11. Design a Container Orchestrator
Reconcile desired workloads, schedule containers and recover from failed nodes safely.
The brief
Design a container orchestrator inspired by Kubernetes. Define desired and observed workload state, scheduling, reconciliation and service discovery. Explain how running workloads behave when management is unavailable, and trace recovery when a node is partitioned but its containers may still be running.
- Manage 10,000 nodes and 200,000 running containers across three zones.
- Workloads declare CPU, memory, replica count and placement constraints; deploys happen continuously.
- Nodes can partition for ten minutes, and controllers can restart or briefly compete for ownership.
Constraints
- Node capacity≥ 10,000 nodes
- Declare supported node count and estimate status-update load.
- Convergent ownership
- Reconciliation must be repeatable and must not allow competing controllers to corrupt desired state.
- Partition behavior
- Distinguish rescheduling a stateless replica from safely moving a stateful workload with exclusive storage.
What to cover
- 01
Resource and state model
Define workload identity, desired/observed state, ownership and versioning.
- 02
Scheduling and reconciliation
Show admission, placement, node agents, retries and discovery updates.
- 03
Capacity and fairness
Estimate control traffic and explain resource accounting, unschedulable work and tenant isolation.
- 04
Partition and rollout walkthrough
Trace controller loss, an uncertain node and a bad rollout; compare recovery speed with duplicate execution risk.
Worked designs
Explore the architecture and decisions, then build on an example with Coach.
Review rubric
AI feedback uses these criteria. Scores are practice feedback.
Resource state and ownership
Desired state, observed status and concurrency form a coherent model.
Convergence and placement
Scheduling and repeated reconciliation respect resources and constraints.
Failure semantics
Partitions, stateful fencing and rollout rollback are explicit.
Operating capacity
Status load, fairness and observability support the fleet size.
Discussion
Share an approach, ask a question, or tag @Coach.
Loading discussion…