97. Design a Distributed Logging and Error Tracking System
Ten terabytes of logs an hour from host agents through Kafka into tiered OpenSearch and S3, with regex search, live tail and exceptions grouped into issues.
Pick a system. Work through the problem. Compare your approach.
Company tags are community-reported. Counts on cards show how many people reported that design.
Ten terabytes of logs an hour from host agents through Kafka into tiered OpenSearch and S3, with regex search, live tail and exceptions grouped into issues.
Billions of files, hundreds of petabytes, one strongly consistent tree: a namespace partitioned by directory, chunk servers, replication, repair, erasure coding. One-hour boards for junior, senior and staff, with the theory behind them.
Signed HTTPS callbacks for a billion events a day, at least once, retried for three days, with no endpoint able to slow another.