System Design AI
All designs

Classic designInfrastructure & SREadvanced

11. Design a Container Orchestrator

Reconcile desired workloads, schedule containers and recover from failed nodes safely.

The brief

Design a container orchestrator inspired by Kubernetes. Define desired and observed workload state, scheduling, reconciliation and service discovery. Explain how running workloads behave when management is unavailable, and trace recovery when a node is partitioned but its containers may still be running.

  • Manage 10,000 nodes and 200,000 running containers across three zones.
  • Workloads declare CPU, memory, replica count and placement constraints; deploys happen continuously.
  • Nodes can partition for ten minutes, and controllers can restart or briefly compete for ownership.

Constraints

Node capacity≥ 10,000 nodes
Declare supported node count and estimate status-update load.
Convergent ownership
Reconciliation must be repeatable and must not allow competing controllers to corrupt desired state.
Partition behavior
Distinguish rescheduling a stateless replica from safely moving a stateful workload with exclusive storage.

What to cover

  1. 01

    Resource and state model

    Define workload identity, desired/observed state, ownership and versioning.

  2. 02

    Scheduling and reconciliation

    Show admission, placement, node agents, retries and discovery updates.

  3. 03

    Capacity and fairness

    Estimate control traffic and explain resource accounting, unschedulable work and tenant isolation.

  4. 04

    Partition and rollout walkthrough

    Trace controller loss, an uncertain node and a bad rollout; compare recovery speed with duplicate execution risk.

Worked designs

Explore the architecture and decisions, then build on an example with Coach.

Review rubric

AI feedback uses these criteria. Scores are practice feedback.

Resource state and ownership

Desired state, observed status and concurrency form a coherent model.

30points

Convergence and placement

Scheduling and repeated reconciliation respect resources and constraints.

25points

Failure semantics

Partitions, stateful fencing and rollout rollback are explicit.

30points

Operating capacity

Status load, fairness and observability support the fleet size.

15points

Discussion

Share an approach, ask a question, or tag @Coach.

Loading discussion…