19. Ship a cheaper model without hiding regressions
Build a release decision from noisy evaluations, customer slices and production feedback.
Start with a template. Work through each step. Ask Coach when you need a second opinion.
Company tags are community-reported. Counts on cards show how many people reported that design.
Build a release decision from noisy evaluations, customer slices and production feedback.
Define a feature once, train on point-in-time correct history and serve 200 fresh values in under 10 ms.
Train 1 B to 400 B models on 8 to 10,240 H100s: DDP, ZeRO and FSDP, tensor and pipeline parallel over NVLink and EFA, checkpoints, hot spares, MFU and goodput.
From registry to traffic for hundreds of models: gates, shadow and canary with automatic rollback, packed CPU and GPU pools, LLM serving, cells and audit.
Choose 20 videos out of 10 million for 100 million daily users: implicit labels, two-tower retrieval, a multi-task ranker, bias and exploration.
Save the 1 TB state of a 70 B model on 10,000 GPUs often enough that a failure costs minutes, without stalling training, and restore fast.
Know within hours when hundreds of production models see broken inputs, drift or falling quality, before the labels arrive.
Route images, text, audio and model replies to thousands of annotators, buy quality with gold and consensus, cut cost with models, and ship versioned datasets.
Pretraining a base LLM on a fixed budget: sizing by 6ND and scaling laws, the data factory, tokenizer, stable runs, failures and evals.
Preference tuning as a weekly loop: rater and AI labels, Bradley–Terry reward models, DPO and PPO with a KL leash, vLLM rollouts, reward-hacking checks.
Tenants upload examples and get a tuned, gated model on the same API: LoRA and QLoRA arithmetic, chat templates, Kueue fair sharing, eval gates, multi-LoRA serving.
GPU batch generation behind a shared cache, sandboxed pass@k, calibrated judges, paired statistics, release gates, contamination checks.
Experiment tracking and a model registry: non-blocking logging, chunked metric curves, content-addressed artifacts, lineage and a promotion gate. One-hour boards for junior, senior and staff, with the theory behind them.
A tuning service like Vizier: random and Bayesian search, ASHA and population-based training, trials on a shared GPU cluster.
A 70B teacher distilled into an 8B student: transfer sets, logit and on-policy distillation, per-slice gates, escalation routing, a refresh loop.
Every new LLM version shrunk to FP8, INT8 or INT4 with a smaller KV cache, gated against its BF16 parent and benchmarked on the serving GPU.
Teacher models write training data for smaller students: conditioned and evolved prompts, verified answers, dedup, decontamination, versioned datasets with lineage.
Search by photo over a billion images: contrastive embeddings from engagement pairs, object crops, IVF-PQ with re-scoring, versioned indexes.
Predict the chance a person clicks an ad, calibrated for the auction, at 10 B requests a day: sampling and its correction, DCN-V2, delayed clicks.
Find and act on posts that break the rules in text, images, video and live: hash matching, a calibrated multimodal model, review ranked by expected harm, prevalence: one-hour boards for junior, senior and staff, with the theory.
Pick 12 homes a guest could book instead, from listing embeddings learned on browsing sessions, filtered by dates and party size, then ranked.
Suggest people a member knows out of a billion: bounded friends of friends, affiliations and contacts, two ranking heads, GNN embeddings, privacy and abuse.
Order each home feed so the time is worth it: candidate sources, a multi-task ranker and value model, integrity re-ranking and feedback loops.
Rank upcoming events when every event is new and expires: geo and time candidates, live features, calibrated ranking, a two-sided market.
Find the right videos for a typed query among billions: BM25 and a dual encoder over text, speech and frames, LambdaMART on debiased clicks, human raters.
Blur every face and licence plate in billions of street panoramas: tiles for tiny faces, recall-first detection, a batch pipeline that fails closed.