104. Design a Feature Store
Define a feature once, train on point-in-time correct history and serve 200 fresh values in under 10 ms.
Pick a system. Work through the problem. Compare your approach.
Company tags are community-reported. Counts on cards show how many people reported that design.
Define a feature once, train on point-in-time correct history and serve 200 fresh values in under 10 ms.
Train 1 B to 400 B models on 8 to 10,240 H100s: DDP, ZeRO and FSDP, tensor and pipeline parallel over NVLink and EFA, checkpoints, hot spares, MFU and goodput.
From registry to traffic for hundreds of models: gates, shadow and canary with automatic rollback, packed CPU and GPU pools, LLM serving, cells and audit.
Save the 1 TB state of a 70 B model on 10,000 GPUs often enough that a failure costs minutes, without stalling training, and restore fast.
Know within hours when hundreds of production models see broken inputs, drift or falling quality, before the labels arrive.
Route images, text, audio and model replies to thousands of annotators, buy quality with gold and consensus, cut cost with models, and ship versioned datasets.
Pretraining a base LLM on a fixed budget: sizing by 6ND and scaling laws, the data factory, tokenizer, stable runs, failures and evals.
GPU batch generation behind a shared cache, sandboxed pass@k, calibrated judges, paired statistics, release gates, contamination checks.
Experiment tracking and a model registry: non-blocking logging, chunked metric curves, content-addressed artifacts, lineage and a promotion gate. One-hour boards for junior, senior and staff, with the theory behind them.
A tuning service like Vizier: random and Bayesian search, ASHA and population-based training, trials on a shared GPU cluster.