109. Design an Embedding Generation and Indexing Pipeline
Embed a billion chunks on GPUs, keep a k-NN index fresh through CDC with versioned writes and provable deletes, and migrate models blue-green.
Start with a template. Work through each step. Ask Coach when you need a second opinion.
Company tags are community-reported. Counts on cards show how many people reported that design.
Embed a billion chunks on GPUs, keep a k-NN index fresh through CDC with versioned writes and provable deletes, and migrate models blue-green.
Choose 20 videos out of 10 million for 100 million daily users: implicit labels, two-tower retrieval, a multi-task ranker, bias and exploration.
Save the 1 TB state of a 70 B model on 10,000 GPUs often enough that a failure costs minutes, without stalling training, and restore fast.
Know within hours when hundreds of production models see broken inputs, drift or falling quality, before the labels arrive.
Route images, text, audio and model replies to thousands of annotators, buy quality with gold and consensus, cut cost with models, and ship versioned datasets.
Petabytes of crawled pages into trillions of clean, deduplicated, tokenized and versioned training tokens, with opt-outs honoured.
Millions of LLM prompts in JSONL files, one result per custom_id within 24 hours at about half the online price, on GPUs that come and go.
Pretraining a base LLM on a fixed budget: sizing by 6ND and scaling laws, the data factory, tokenizer, stable runs, failures and evals.
Detect, classify and repair failing GPU nodes (Xid errors, ECC, NVLink and EFA faults, stragglers, silent corruption) without draining the fleet.
One API in front of every model: token quotas, routing and fallbacks, caching, cost attribution, masked logs and guardrails, streamed without buffering.
Preference tuning as a weekly loop: rater and AI labels, Bradley–Terry reward models, DPO and PPO with a KL leash, vLLM rollouts, reward-hacking checks.
Tenants upload examples and get a tuned, gated model on the same API: LoRA and QLoRA arithmetic, chat templates, Kueue fair sharing, eval gates, multi-LoRA serving.
GPU batch generation behind a shared cache, sandboxed pass@k, calibrated judges, paired statistics, release gates, contamination checks.
Experiment tracking and a model registry: non-blocking logging, chunked metric curves, content-addressed artifacts, lineage and a promotion gate. One-hour boards for junior, senior and staff, with the theory behind them.
A tuning service like Vizier: random and Bayesian search, ASHA and population-based training, trials on a shared GPU cluster.
Serve 10,000 LoRA adapters for 2,000 tenants on one shared base model, with mixed-adapter batches, tiered adapter caches, affinity routing and base upgrades.
Exact, semantic and prefix caching in front of LLMs: scoped keys, distances with an error budget, tenant isolation, invalidation and measured savings.
A 70B teacher distilled into an 8B student: transfer sets, logit and on-policy distillation, per-slice gates, escalation routing, a refresh loop.
Every new LLM version shrunk to FP8, INT8 or INT4 with a smaller KV cache, gated against its BF16 parent and benchmarked on the serving GPU.
Teacher models write training data for smaller students: conditioned and evolved prompts, verified answers, dedup, decontamination, versioned datasets with lineage.
Billions of image-text pairs from Common Crawl: polite fetching, CLIP scoring, dedup, recaptioning, WebDataset shards and takedowns.
Search by photo over a billion images: contrastive embeddings from engagement pairs, object crops, IVF-PQ with re-scoring, versioned indexes.
Answers from private documents with citations: chunking, BM25 + vector retrieval, reranking, permission filters, grounding checks and evaluation.
Next words within 100 ms of a keystroke: conditional language modelling, a distilled student on draft-pinned KV caches, private n-grams, DP and canary audits.
Translation for 20 to 240 languages: one multilingual encoder-decoder, mined and back-translated data, distilled students, COMET and MQM, quality-aware serving: one-hour boards for junior, senior and staff, with the theory.
Prompt to four images in seconds: latent diffusion, guidance, a distilled few-step student, recaptioned data, safety on both sides, watermarks and C2PA.
Captions read aloud as alt text: encoder, bridge and decoder, CLIP-filtered data, beam search, hallucination-aware evaluation, lanes on one GPU pool.
Faces of people who do not exist: a style-based GAN, truncation, latent-space edits, selfie inversion, provenance and memorisation audits.
Predict the chance a person clicks an ad, calibrated for the auction, at 10 B requests a day: sampling and its correction, DCN-V2, delayed clicks.
Find and act on posts that break the rules in text, images, video and live: hash matching, a calibrated multimodal model, review ranked by expected harm, prevalence: one-hour boards for junior, senior and staff, with the theory.