42. Design a Metrics Monitoring Platform
Collect 5 million samples a second from 500,000 hosts, store them as time series, chart them and page people.
ClassicMedium
Pick a system. Work through the problem. Compare your approach.
Company tags are community-reported. Counts on cards show how many people reported that design.
Collect 5 million samples a second from 500,000 hosts, store them as time series, chart them and page people.
Ten terabytes of logs an hour from host agents through Kafka into tiered OpenSearch and S3, with regex search, live tail and exceptions grouped into issues.
Detect, classify and repair failing GPU nodes (Xid errors, ECC, NVLink and EFA faults, stragglers, silent corruption) without draining the fleet.