11. Design a Container Orchestrator
Reconcile desired workloads, schedule containers and recover from failed nodes safely.
Pick a system. Work through the problem. Compare your approach.
Company tags are community-reported. Counts on cards show how many people reported that design.
Reconcile desired workloads, schedule containers and recover from failed nodes safely.
Millions of future actions — reminders, retries, expiries — must fire near their due time. Scanning everything every minute does not scale, and an in-process timer dies with the process.
Run 10,000 jobs a second within two seconds of their time, at least once, with retries, fairness between tenants and exactly-once effects.
Thousands of GPUs shared by many teams: gang admission, topology-aware placement, quotas with lending, checkpointed preemption and failure recovery.