System Design AI
All designs

Data engineeringintermediate

28. Aggregate events that arrive late or twice

Keep tenant usage totals correct across duplicate delivery, late events, outages and replay.

The brief

Design a pipeline that publishes per-tenant, five-minute usage counts and unit totals. Events carry a tenant ID, stable event ID, event time, integer usage units and schema version. Define when a result is provisional or final, and explain how replay repairs a bad deployment without double counting.

  • Ingest averages 20,000 events/second and peaks at 100,000; each event is about 600 bytes. There are 2,000 tenants, and one produces 35% of traffic.
  • Sources retry and deliver out of order. Events can arrive two hours late; clock errors and incompatible schemas also occur.
  • Dashboards accept provisional totals. Raw events are retained for seven days, and a six-hour processor outage must be recoverable.

Constraints

Peak ingest capacity≥ 100,000 events/second
Declare peak supported events per second and show partition and storage assumptions.
Normal-operation freshness≤ 60 seconds
Declare p95 time from durable acceptance to a visible provisional total, excluding recovery periods.
Replay retention≥ 7 days
Declare how many days of accepted raw events remain replayable, with a storage estimate.
Late and duplicate events
Define tenant-scoped deduplication, event-time windows, late corrections and what final means; handle conflicting duplicate payloads and events outside the correction window.
Safe recovery and replay
Explain durable acknowledgement, checkpoint/output coordination and how a backfill is verified and published while live traffic continues.

What to cover

  1. 01

    Data path and ownership

    Show acceptance, durable storage, validation, partitioning, aggregation and serving. Identify each state owner.

  2. 02

    Correctness contract

    Define event identity, time windows, deduplication retention, provisional/final results and late correction behavior. Trace a duplicate across a restart.

  3. 03

    Capacity and latency

    Estimate peak ingress, seven-day raw storage and processing state. Address the hot tenant and explain catch-up capacity after a six-hour outage.

  4. 04

    Replay and release plan

    Walk through repairing a bad aggregation version while new events arrive. Specify validation, atomic publication or version selection, rollback and useful alerts.

Worked designs

No worked design has been published for this brief yet. You can start an attempt and share your approach in the discussion.

Review rubric

AI feedback uses these criteria. Scores are practice feedback.

Result correctness

Identity, windows, late corrections and state lifetime form a consistent contract through retries and restarts.

35points

Capacity and freshness

Calculations cover peak load, skew, retention and recovery headroom; latency claims state their assumptions.

25points

Recovery and replay

Replay avoids double counting, verifies results and safely publishes a version alongside live processing.

25points

Operational choices

Schema failures, lag and bad deployments have visible handling; alternatives and limits are explained.

15points

Discussion

Share an approach, ask a question, or tag @Coach.

Share your interview experience

Published under your alias as a community report. Leave out interviewer names and confidential material.

Loading discussion…