28. Aggregate events that arrive late or twice
Keep tenant usage totals correct across duplicate delivery, late events, outages and replay.
The brief
Design a pipeline that publishes per-tenant, five-minute usage counts and unit totals. Events carry a tenant ID, stable event ID, event time, integer usage units and schema version. Define when a result is provisional or final, and explain how replay repairs a bad deployment without double counting.
- Ingest averages 20,000 events/second and peaks at 100,000; each event is about 600 bytes. There are 2,000 tenants, and one produces 35% of traffic.
- Sources retry and deliver out of order. Events can arrive two hours late; clock errors and incompatible schemas also occur.
- Dashboards accept provisional totals. Raw events are retained for seven days, and a six-hour processor outage must be recoverable.
Constraints
- Peak ingest capacity≥ 100,000 events/second
- Declare peak supported events per second and show partition and storage assumptions.
- Normal-operation freshness≤ 60 seconds
- Declare p95 time from durable acceptance to a visible provisional total, excluding recovery periods.
- Replay retention≥ 7 days
- Declare how many days of accepted raw events remain replayable, with a storage estimate.
- Late and duplicate events
- Define tenant-scoped deduplication, event-time windows, late corrections and what final means; handle conflicting duplicate payloads and events outside the correction window.
- Safe recovery and replay
- Explain durable acknowledgement, checkpoint/output coordination and how a backfill is verified and published while live traffic continues.
What to cover
- 01
Data path and ownership
Show acceptance, durable storage, validation, partitioning, aggregation and serving. Identify each state owner.
- 02
Correctness contract
Define event identity, time windows, deduplication retention, provisional/final results and late correction behavior. Trace a duplicate across a restart.
- 03
Capacity and latency
Estimate peak ingress, seven-day raw storage and processing state. Address the hot tenant and explain catch-up capacity after a six-hour outage.
- 04
Replay and release plan
Walk through repairing a bad aggregation version while new events arrive. Specify validation, atomic publication or version selection, rollback and useful alerts.
Worked designs
No worked design has been published for this brief yet. You can start an attempt and share your approach in the discussion.
Review rubric
AI feedback uses these criteria. Scores are practice feedback.
Result correctness
Identity, windows, late corrections and state lifetime form a consistent contract through retries and restarts.
Capacity and freshness
Calculations cover peak load, skew, retention and recovery headroom; latency claims state their assumptions.
Recovery and replay
Replay avoids double counting, verifies results and safely publishes a version alongside live processing.
Operational choices
Schema failures, lag and bad deployments have visible handling; alternatives and limits are explained.
Discussion
Share an approach, ask a question, or tag @Coach.
Loading discussion…