Building blocks, complete architectures, and variations worth comparing. Search by the problem you want to solve or a component you want to understand.
Next words within 100 ms of a keystroke: conditional language modelling, a distilled student on draft-pinned KV caches, private n-grams, DP and canary audits: one-hour boards for junior, senior and staff, with the theory behind them.
genaillmautocompletelatency
System Design AI·Junior → Staff0 bookmarks0 comments
Answers from private documents with citations: chunking, BM25 + vector retrieval, reranking, permission filters, grounding checks and evaluation: one-hour boards for junior, senior and staff, with the theory behind them.
genairagretrievalllm
System Design AI·Junior → Staff0 bookmarks0 comments
Teacher models write training data for smaller students: conditioned and evolved prompts, verified answers, dedup, decontamination, versioned datasets with lineage: one-hour boards for junior, senior and staff, with the theory behind them.
genaillmsynthetic-datatraining
System Design AI·Junior → Staff0 bookmarks0 comments
Every new LLM version shrunk to FP8, INT8 or INT4 with a smaller KV cache, gated against its BF16 parent and benchmarked on the serving GPU: one-hour boards for junior, senior and staff, with the theory behind them.
genaillmquantizationinference
System Design AI·Junior → Staff0 bookmarks0 comments
A 70B teacher distilled into an 8B student: transfer sets, logit and on-policy distillation, per-slice gates, escalation routing, a refresh loop: one-hour boards for junior, senior and staff, with the theory behind them.
genaillmdistillationtraining
System Design AI·Junior → Staff0 bookmarks0 comments
Exact, semantic and prefix caching in front of LLMs: scoped keys, distances with an error budget, tenant isolation, invalidation and measured savings: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infrallmcachingembeddings
System Design AI·Junior → Staff0 bookmarks0 comments
Serve 10,000 LoRA adapters for 2,000 tenants on one shared base model, with mixed-adapter batches, tiered adapter caches, affinity routing and base upgrades: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infrallmlorainference
System Design AI·Junior → Staff0 bookmarks0 comments
GPU batch generation behind a shared cache, sandboxed pass@k, calibrated judges, paired statistics, release gates, contamination checks: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infrallmevaluationmlops
System Design AI·Junior → Staff0 bookmarks0 comments
Tenants upload examples and get a tuned, gated model on the same API: LoRA and QLoRA arithmetic, chat templates, Kueue fair sharing, eval gates, multi-LoRA serving: one-hour boards for junior, senior and staff, with the theory behind them.
genaillmfine-tuninglora
System Design AI·Junior → Staff0 bookmarks0 comments
Preference tuning as a weekly loop: rater and AI labels, Bradley–Terry reward models, DPO and PPO with a KL leash, vLLM rollouts, reward-hacking checks: one-hour boards for junior, senior and staff, with the theory behind them.
genaillmrlhfalignment
System Design AI·Junior → Staff0 bookmarks0 comments
One API in front of every model: token quotas, routing and fallbacks, caching, cost attribution, masked logs and guardrails, streamed without buffering: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infrallmgatewayrate-limiting
System Design AI·Junior → Staff0 bookmarks0 comments
Pretraining a base LLM on a fixed budget: sizing by 6ND and scaling laws, the data factory, tokenizer, stable runs, failures and evals: one-hour boards for junior, senior and staff, with the theory behind them.
genaillmtrainingai-infra
System Design AI·Junior → Staff0 bookmarks0 comments
Millions of LLM prompts in JSONL files, one result per custom_id within 24 hours at about half the online price, on GPUs that come and go: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infrallmbatchinference
System Design AI·Junior → Staff0 bookmarks0 comments
Petabytes of crawled pages into trillions of clean, deduplicated, tokenized and versioned training tokens, with opt-outs honoured: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infradata-pipelinellmdeduplication
System Design AI·Junior → Staff0 bookmarks0 comments
Open-weight LLMs from 8B to 405B served to many products over a streaming OpenAI-style API on AWS GPUs: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infrallmgpuinference
System Design AI·Junior → Staff0 bookmarks0 comments
A chat assistant on models we train and serve: next-token framing, three training stages, resumable SSE streams, prefix caching, token quotas, safety, evaluation: one-hour boards for junior, senior and staff, with the theory behind them.
genaillmai-mlstreaming
System Design AI·Junior → Staff0 bookmarks0 comments