Keep tenant usage totals correct across duplicate delivery, late events, outages and replay.
Data engineers building storage, streaming and processing systems.
Your approach: Use diagrams or prose for data contracts, state ownership, throughput, ordering and recovery where the brief requires them.
Design a pipeline that publishes per-tenant, five-minute usage counts and unit totals. Events carry a tenant ID, stable event ID, event time, integer usage units and schema version. Define when a result is provisional or final, and explain how replay repairs a bad deployment without double counting.
Declare peak supported events per second and show partition and storage assumptions.
Required target: ≥ 100,000 events/second
Declare p95 time from durable acceptance to a visible provisional total, excluding recovery periods.
Required target: ≤ 60 seconds
Declare how many days of accepted raw events remain replayable, with a storage estimate.
Required target: ≥ 7 days
Define tenant-scoped deduplication, event-time windows, late corrections and what final means; handle conflicting duplicate payloads and events outside the correction window.
Explain durable acknowledgement, checkpoint/output coordination and how a backfill is verified and published while live traffic continues.
Show acceptance, durable storage, validation, partitioning, aggregation and serving. Identify each state owner.
Define event identity, time windows, deduplication retention, provisional/final results and late correction behavior. Trace a duplicate across a restart.
Estimate peak ingress, seven-day raw storage and processing state. Address the hot tenant and explain catch-up capacity after a six-hour outage.
Walk through repairing a bad aggregation version while new events arrive. Specify validation, atomic publication or version selection, rollback and useful alerts.