121. Design an LLM Evaluation Platform
GPU batch generation behind a shared cache, sandboxed pass@k, calibrated judges, paired statistics, release gates, contamination checks.
Pick a system. Work through the problem. Compare your approach.
Company tags are community-reported. Counts on cards show how many people reported that design.
GPU batch generation behind a shared cache, sandboxed pass@k, calibrated judges, paired statistics, release gates, contamination checks.
Experiment tracking and a model registry: non-blocking logging, chunked metric curves, content-addressed artifacts, lineage and a promotion gate. One-hour boards for junior, senior and staff, with the theory behind them.
A tuning service like Vizier: random and Bayesian search, ASHA and population-based training, trials on a shared GPU cluster.
Serve 10,000 LoRA adapters for 2,000 tenants on one shared base model, with mixed-adapter batches, tiered adapter caches, affinity routing and base upgrades.
Exact, semantic and prefix caching in front of LLMs: scoped keys, distances with an error budget, tenant isolation, invalidation and measured savings.
A 70B teacher distilled into an 8B student: transfer sets, logit and on-policy distillation, per-slice gates, escalation routing, a refresh loop.
Every new LLM version shrunk to FP8, INT8 or INT4 with a smaller KV cache, gated against its BF16 parent and benchmarked on the serving GPU.
Teacher models write training data for smaller students: conditioned and evolved prompts, verified answers, dedup, decontamination, versioned datasets with lineage.
Billions of image-text pairs from Common Crawl: polite fetching, CLIP scoring, dedup, recaptioning, WebDataset shards and takedowns.
Search by photo over a billion images: contrastive embeddings from engagement pairs, object crops, IVF-PQ with re-scoring, versioned indexes.
Answers from private documents with citations: chunking, BM25 + vector retrieval, reranking, permission filters, grounding checks and evaluation.
Next words within 100 ms of a keystroke: conditional language modelling, a distilled student on draft-pinned KV caches, private n-grams, DP and canary audits.
Translation for 20 to 240 languages: one multilingual encoder-decoder, mined and back-translated data, distilled students, COMET and MQM, quality-aware serving: one-hour boards for junior, senior and staff, with the theory.
Prompt to four images in seconds: latent diffusion, guidance, a distilled few-step student, recaptioned data, safety on both sides, watermarks and C2PA.
Captions read aloud as alt text: encoder, bridge and decoder, CLIP-filtered data, beam search, hallucination-aware evaluation, lanes on one GPU pool.
Faces of people who do not exist: a style-based GAN, truncation, latent-space edits, selfie inversion, provenance and memorisation audits.
Predict the chance a person clicks an ad, calibrated for the auction, at 10 B requests a day: sampling and its correction, DCN-V2, delayed clicks.
Find and act on posts that break the rules in text, images, video and live: hash matching, a calibrated multimodal model, review ranked by expected harm, prevalence: one-hour boards for junior, senior and staff, with the theory.
Pick 12 homes a guest could book instead, from listing embeddings learned on browsing sessions, filtered by dates and party size, then ranked.
Suggest people a member knows out of a billion: bounded friends of friends, affiliations and contacts, two ranking heads, GNN embeddings, privacy and abuse.
Order each home feed so the time is worth it: candidate sources, a multi-task ranker and value model, integrity re-ranking and feedback loops.
Rank upcoming events when every event is new and expires: geo and time candidates, live features, calibrated ranking, a two-sided market.
Find the right videos for a typed query among billions: BM25 and a dual encoder over text, speech and frames, LambdaMART on debiased clicks, human raters.
Selfies in, professional headshots out: per-order LoRA on SDXL, an identity-encoder preview, face-match ranking, GPU pools, consent and deletion.
Images at 2K–4K without paying for every pixel: latent diffusion, an SR cascade with noise augmentation, tiled decoding, distilled drafts, cost per megapixel: one-hour boards for junior, senior and staff, with the theory.
Short clips from a prompt: latent video diffusion, recaptioned data, spacetime attention, cascades, multi-GPU jobs, previews, fair queues, provenance.
Blur every face and licence plate in billions of street panoramas: tiles for tiny faces, recall-first detection, a batch pipeline that fails closed.
The URL shortener with a payload: the same id generation and cache-first read path, but the value is kilobytes of text, so it moves out of the database and into object storage.