Building blocks, complete architectures, and variations worth comparing. Search by the problem you want to solve or a component you want to understand.
Billions of image-text pairs from Common Crawl: polite fetching, CLIP scoring, dedup, recaptioning, WebDataset shards and takedowns: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infradata-pipelinemultimodalvision
System Design AI·Junior → Staff0 bookmarks0 comments
Exact, semantic and prefix caching in front of LLMs: scoped keys, distances with an error budget, tenant isolation, invalidation and measured savings: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infrallmcachingembeddings
System Design AI·Junior → Staff0 bookmarks0 comments
Serve 10,000 LoRA adapters for 2,000 tenants on one shared base model, with mixed-adapter batches, tiered adapter caches, affinity routing and base upgrades: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infrallmlorainference
System Design AI·Junior → Staff0 bookmarks0 comments
A tuning service like Vizier: random and Bayesian search, ASHA and population-based training, trials on a shared GPU cluster: one-hour boards for junior, senior and staff, with the theory behind them.
ai-inframlopshyperparameter-tuningoptimization
System Design AI·Junior → Staff0 bookmarks0 comments
Experiment tracking and a model registry: non-blocking logging, chunked metric curves, content-addressed artifacts, lineage and a promotion gate. One-hour boards for junior, senior and staff, with the theory behind them.
ai-inframlopsexperiment-trackingmodel-registry
System Design AI·Junior → Staff0 bookmarks0 comments
GPU batch generation behind a shared cache, sandboxed pass@k, calibrated judges, paired statistics, release gates, contamination checks: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infrallmevaluationmlops
System Design AI·Junior → Staff0 bookmarks0 comments
One API in front of every model: token quotas, routing and fallbacks, caching, cost attribution, masked logs and guardrails, streamed without buffering: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infrallmgatewayrate-limiting
System Design AI·Junior → Staff0 bookmarks0 comments
Detect, classify and repair failing GPU nodes (Xid errors, ECC, NVLink and EFA faults, stragglers, silent corruption) without draining the fleet: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infragpureliabilityobservability
System Design AI·Junior → Staff0 bookmarks0 comments
Pretraining a base LLM on a fixed budget: sizing by 6ND and scaling laws, the data factory, tokenizer, stable runs, failures and evals: one-hour boards for junior, senior and staff, with the theory behind them.
genaillmtrainingai-infra
System Design AI·Junior → Staff0 bookmarks0 comments
Millions of LLM prompts in JSONL files, one result per custom_id within 24 hours at about half the online price, on GPUs that come and go: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infrallmbatchinference
System Design AI·Junior → Staff0 bookmarks0 comments
Petabytes of crawled pages into trillions of clean, deduplicated, tokenized and versioned training tokens, with opt-outs honoured: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infradata-pipelinellmdeduplication
System Design AI·Junior → Staff0 bookmarks0 comments
Route images, text, audio and model replies to thousands of annotators, buy quality with gold and consensus, cut cost with models, and ship versioned datasets: one-hour boards for junior, senior and staff, with the theory behind them.
ai-inframl-platformlabelinghuman-in-the-loop
System Design AI·Junior → Staff0 bookmarks0 comments
Know within hours when hundreds of production models see broken inputs, drift or falling quality, before the labels arrive: one-hour boards for junior, senior and staff, with the theory behind them.
ai-inframlopsmonitoringdrift
System Design AI·Junior → Staff0 bookmarks0 comments
Save the 1 TB state of a 70 B model on 10,000 GPUs often enough that a failure costs minutes, without stalling training, and restore fast: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infratrainingcheckpointingstorage
System Design AI·Junior → Staff0 bookmarks0 comments
Embed a billion chunks on GPUs, keep a k-NN index fresh through CDC with versioned writes and provable deletes, and migrate models blue-green: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infraembeddingsvector-searchbatch-inference
System Design AI·Junior → Staff0 bookmarks0 comments
Open-weight LLMs from 8B to 405B served to many products over a streaming OpenAI-style API on AWS GPUs: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infrallmgpuinference
System Design AI·Junior → Staff0 bookmarks0 comments
From registry to traffic for hundreds of models: gates, shadow and canary with automatic rollback, packed CPU and GPU pools, LLM serving, cells and audit: one-hour boards for junior, senior and staff, with the theory behind them.
ai-inframlopsservingdeployment
System Design AI·Junior → Staff0 bookmarks0 comments
Train 1 B to 400 B models on 8 to 10,240 H100s: DDP, ZeRO and FSDP, tensor and pipeline parallel over NVLink and EFA, checkpoints, hot spares, MFU and goodput: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infratraininggpuparallelism
System Design AI·Junior → Staff0 bookmarks0 comments
Define a feature once, train on point-in-time correct history and serve 200 fresh values in under 10 ms: one-hour boards for junior, senior and staff, with the theory behind them.
ai-inframl-platformfeaturesstreaming
System Design AI·Junior → Staff0 bookmarks0 comments
A 500 GB model on 1,000 GPU servers in minutes: chunks and hashes, a topology-aware swarm, signed manifests, bandwidth budgets and waves: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infrainfrastructurep2pdistribution
System Design AI·Junior → Staff0 bookmarks0 comments
Thousands of GPUs shared by many teams: gang admission, topology-aware placement, quotas with lending, checkpointed preemption and failure recovery: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infragpuschedulingkubernetes
System Design AI·Junior → Staff0 bookmarks0 comments