102. Design a GPU Cluster Scheduler for Training Jobs
Thousands of GPUs shared by many teams: gang admission, topology-aware placement, quotas with lending, checkpointed preemption and failure recovery.
ClassicMedium
Pick a system. Work through the problem. Compare your approach.
Company tags are community-reported. Counts on cards show how many people reported that design.
Thousands of GPUs shared by many teams: gang admission, topology-aware placement, quotas with lending, checkpointed preemption and failure recovery.
Detect, classify and repair failing GPU nodes (Xid errors, ECC, NVLink and EFA faults, stragglers, silent corruption) without draining the fleet.