102. Design a GPU Cluster Scheduler for Training Jobs
Thousands of GPUs shared by many teams: gang admission, topology-aware placement, quotas with lending, checkpointed preemption and failure recovery.
Pick a system. Work through the problem. Compare your approach.
Company tags are community-reported. Counts on cards show how many people reported that design.
Thousands of GPUs shared by many teams: gang admission, topology-aware placement, quotas with lending, checkpointed preemption and failure recovery.
Train 1 B to 400 B models on 8 to 10,240 H100s: DDP, ZeRO and FSDP, tensor and pipeline parallel over NVLink and EFA, checkpoints, hot spares, MFU and goodput.
Open-weight LLMs from 8B to 405B served to many products over a streaming OpenAI-style API on AWS GPUs.
Detect, classify and repair failing GPU nodes (Xid errors, ECC, NVLink and EFA faults, stragglers, silent corruption) without draining the fleet.