102. Design a GPU Cluster Scheduler for Training Jobs
Thousands of GPUs shared by many teams: gang admission, topology-aware placement, quotas with lending, checkpointed preemption and failure recovery.
ClassicMedium
Pick a system. Work through the problem. Compare your approach.
Company tags are community-reported. Counts on cards show how many people reported that design.
Thousands of GPUs shared by many teams: gang admission, topology-aware placement, quotas with lending, checkpointed preemption and failure recovery.
A 500 GB model on 1,000 GPU servers in minutes: chunks and hashes, a topology-aware swarm, signed manifests, bandwidth budgets and waves.
Detect, classify and repair failing GPU nodes (Xid errors, ECC, NVLink and EFA faults, stragglers, silent corruption) without draining the fleet.