105. Design a Distributed Training Platform
Train 1 B to 400 B models on 8 to 10,240 H100s: DDP, ZeRO and FSDP, tensor and pipeline parallel over NVLink and EFA, checkpoints, hot spares, MFU and goodput.
Start with a template. Work through each step. Ask Coach when you need a second opinion.
Company tags are community-reported. Counts on cards show how many people reported that design.
Train 1 B to 400 B models on 8 to 10,240 H100s: DDP, ZeRO and FSDP, tensor and pipeline parallel over NVLink and EFA, checkpoints, hot spares, MFU and goodput.
Save the 1 TB state of a 70 B model on 10,000 GPUs often enough that a failure costs minutes, without stalling training, and restore fast.
Pretraining a base LLM on a fixed budget: sizing by 6ND and scaling laws, the data factory, tokenizer, stable runs, failures and evals.
A 70B teacher distilled into an 8B student: transfer sets, logit and on-policy distillation, per-slice gates, escalation routing, a refresh loop.
Teacher models write training data for smaller students: conditioned and evolved prompts, verified answers, dedup, decontamination, versioned datasets with lineage.