System Design AI
Design brief
Infrastructure & SREAdvanced55 min suggested

Design a Container Orchestrator

Reconcile desired workloads, schedule containers and recover from failed nodes safely.

Infrastructure, platform and reliability engineers.

Your approach: Use diagrams or prose to explain responsibilities, state, capacity and failure behavior. Show the calculations requested by the question.

The problem

Design a container orchestrator inspired by Kubernetes. Define desired and observed workload state, scheduling, reconciliation and service discovery. Explain how running workloads behave when management is unavailable, and trace recovery when a node is partitioned but its containers may still be running.

  • Manage 10,000 nodes and 200,000 running containers across three zones.
  • Workloads declare CPU, memory, replica count and placement constraints; deploys happen continuously.
  • Nodes can partition for ten minutes, and controllers can restart or briefly compete for ownership.

Work within these constraints

Node capacity

Declare supported node count and estimate status-update load.

Required target: ≥ 10,000 nodes

Convergent ownership

Reconciliation must be repeatable and must not allow competing controllers to corrupt desired state.

Partition behavior

Distinguish rescheduling a stateless replica from safely moving a stateful workload with exclusive storage.

What to deliver

1

Resource and state model

Define workload identity, desired/observed state, ownership and versioning.

2

Scheduling and reconciliation

Show admission, placement, node agents, retries and discovery updates.

3

Capacity and fairness

Estimate control traffic and explain resource accounting, unschedulable work and tenant isolation.

4

Partition and rollout walkthrough

Trace controller loss, an uncertain node and a bad rollout; compare recovery speed with duplicate execution risk.