28. Aggregate events that arrive late or twice
Keep tenant usage totals correct across duplicate delivery, late events, outages and replay.
Start with a template. Work through each step. Ask Coach when you need a second opinion.
Company tags are community-reported. Counts on cards show how many people reported that design.
Keep tenant usage totals correct across duplicate delivery, late events, outages and replay.
Finding the nearby members of a set that is constantly moving. A table scan with a distance function is hopeless at any scale, and a stale position is worse than none — it sends a car to someone who left ten minutes ago.
Count every ad click once, fast enough to chart live and exactly enough to bill. One-hour interview boards for junior, senior and staff: requirements, data layer, low-level design and what goes wrong at every component.
The classes and code behind the ad click aggregator: signed click tokens, a redirect that never waits on the log, and a stream counter that de-duplicates and handles late clicks. Tests run all of it.
The K most-viewed videos for the last hour, day, month and all time from 700,000 views a second, exactly and in milliseconds.
Ten billion pages in five days, politely, without losing progress: fetchers and parsers, domain locks, bandwidth math, deduplication and a crawl that stays fresh. One-hour boards for junior, senior and staff.
Search a billion documents in 200 ms: inverted indexes, analysers, Lucene segments, shards and replicas, query then fetch with BM25, fed by CDC.
A partitioned, replicated append-only log: a million messages a second, order per key, nothing acknowledged ever lost, a week of replayable history.
A leaderless wide-column store: a token ring with virtual nodes, consistency tuned per query, an LSM write path and repair that keeps replicas converged.
A key-value store like DynamoDB, from one durable node to Paxos-replicated partitions to global tables run for thousands of tenants.
A relational database like PostgreSQL: WAL and MVCC on one server, quorum replication and failover, then a sharded fleet.
A stateful stream processor: dataflow graphs, keyed state in RocksDB, event time and watermarks, windows, barrier checkpoints, exactly-once into Kafka.
Radius, k-nearest and map-box search over places that rarely move and drivers that ping every few seconds: R-trees, geohash, H3 and S2 cells, Redis GEO and OpenSearch. One-hour boards for junior, senior and staff, with the theory.
Unique visitors, top pages, seen URLs and p99 latency over 10 B events a day with HyperLogLog, Count-Min, Bloom filters and t-digests.
Nearest-neighbour search over a billion embeddings: HNSW, IVF-PQ and disk graphs, filters, segments, sharding, and hybrid search with re-ranking.
The database under a metrics platform: a WAL and head block, compressed chunks, a label index, rollups, compaction and blocks in S3.
Every committed change in Aurora and DynamoDB, read from the log and delivered in order per row to search, caches, Redshift and an S3 lake.
Count 4 million moving drivers per map cell from a million pings a second, and serve the density to a million viewers as cached map tiles.
Embed a billion chunks on GPUs, keep a k-NN index fresh through CDC with versioned writes and provable deletes, and migrate models blue-green.
Petabytes of crawled pages into trillions of clean, deduplicated, tokenized and versioned training tokens, with opt-outs honoured.
Billions of image-text pairs from Common Crawl: polite fetching, CLIP scoring, dedup, recaptioning, WebDataset shards and takedowns.