42. Design a Metrics Monitoring Platform
Collect 5 million samples a second from 500,000 hosts, store them as time series, chart them and page people.
Start with a template. Work through each step. Ask Coach when you need a second opinion.
Company tags are community-reported. Counts on cards show how many people reported that design.
Collect 5 million samples a second from 500,000 hosts, store them as time series, chart them and page people.
The database under a metrics platform: a WAL and head block, compressed chunks, a label index, rollups, compaction and blocks in S3.
Ten terabytes of logs an hour from host agents through Kafka into tiered OpenSearch and S3, with regex search, live tail and exceptions grouped into issues.
Know within hours when hundreds of production models see broken inputs, drift or falling quality, before the labels arrive.
Detect, classify and repair failing GPU nodes (Xid errors, ECC, NVLink and EFA faults, stragglers, silent corruption) without draining the fleet.