91. Design a CI/CD System like GitHub Actions
Turn 10 pushes a second into job graphs on ephemeral runners, with fair queues per pool, leases fenced by attempt and live logs.
Pick a system. Work through the problem. Compare your approach.
Company tags are community-reported. Counts on cards show how many people reported that design.
Turn 10 pushes a second into job graphs on ephemeral runners, with fair queues per pool, leases fenced by attempt and live logs.
Real-time rated chess with server-validated moves, lag-compensated clocks, pairing by rating and a million watchers on one game.
Rank 200 million players in real time with sharded Redis sorted sets, count every game exactly once, and close a season fairly.
Count 4 million moving drivers per map cell from a million pings a second, and serve the density to a million viewers as cached map tiles.
Design a team chat app like Slack: channels and DMs in real time, threads, unread counts and search. One-hour boards for junior, senior and staff, with the theory behind them.
Find restaurants that deliver to you, check out without double charges, time couriers to the kitchen and track them live.
Ten terabytes of logs an hour from host agents through Kafka into tiered OpenSearch and S3, with regex search, live tail and exceptions grouped into issues.
Record every money movement as balanced, append-only entries, exactly once, with holds, payouts and reconciliation.
One set of tags across Jira issues, Confluence pages and Bitbucket pull requests: batch renders, tag pages that never leak, suggestions, popular tags.
Billions of files, hundreds of petabytes, one strongly consistent tree: a namespace partitioned by directory, chunk servers, replication, repair, erasure coding. One-hour boards for junior, senior and staff, with the theory behind them.
Signed HTTPS callbacks for a billion events a day, at least once, retried for three days, with no endpoint able to slow another.
Thousands of GPUs shared by many teams: gang admission, topology-aware placement, quotas with lending, checkpointed preemption and failure recovery.
A 500 GB model on 1,000 GPU servers in minutes: chunks and hashes, a topology-aware swarm, signed manifests, bandwidth budgets and waves.
Define a feature once, train on point-in-time correct history and serve 200 fresh values in under 10 ms.
Train 1 B to 400 B models on 8 to 10,240 H100s: DDP, ZeRO and FSDP, tensor and pipeline parallel over NVLink and EFA, checkpoints, hot spares, MFU and goodput.
A chat assistant on models we train and serve: next-token framing, three training stages, resumable SSE streams, prefix caching, token quotas, safety, evaluation.
From registry to traffic for hundreds of models: gates, shadow and canary with automatic rollback, packed CPU and GPU pools, LLM serving, cells and audit.
Open-weight LLMs from 8B to 405B served to many products over a streaming OpenAI-style API on AWS GPUs.
Embed a billion chunks on GPUs, keep a k-NN index fresh through CDC with versioned writes and provable deletes, and migrate models blue-green.
Choose 20 videos out of 10 million for 100 million daily users: implicit labels, two-tower retrieval, a multi-task ranker, bias and exploration.
Save the 1 TB state of a 70 B model on 10,000 GPUs often enough that a failure costs minutes, without stalling training, and restore fast.
Know within hours when hundreds of production models see broken inputs, drift or falling quality, before the labels arrive.
Route images, text, audio and model replies to thousands of annotators, buy quality with gold and consensus, cut cost with models, and ship versioned datasets.
Petabytes of crawled pages into trillions of clean, deduplicated, tokenized and versioned training tokens, with opt-outs honoured.
Millions of LLM prompts in JSONL files, one result per custom_id within 24 hours at about half the online price, on GPUs that come and go.
Pretraining a base LLM on a fixed budget: sizing by 6ND and scaling laws, the data factory, tokenizer, stable runs, failures and evals.
Detect, classify and repair failing GPU nodes (Xid errors, ECC, NVLink and EFA faults, stragglers, silent corruption) without draining the fleet.
One API in front of every model: token quotas, routing and fallbacks, caching, cost attribution, masked logs and guardrails, streamed without buffering.
Preference tuning as a weekly loop: rater and AI labels, Bradley–Terry reward models, DPO and PPO with a KL leash, vLLM rollouts, reward-hacking checks.
Tenants upload examples and get a tuned, gated model on the same API: LoRA and QLoRA arithmetic, chat templates, Kueue fair sharing, eval gates, multi-LoRA serving.