Millions of future actions — reminders, retries, expiries — must fire near their due time. Scanning everything every minute does not scale, and an in-process timer dies with the process.
Infrastructure, platform and reliability engineers.
Your approach: Use diagrams or prose to explain responsibilities, state, capacity and failure behavior. Show the calculations requested by the question.
Delay Queue Scheduler. Millions of future actions — reminders, retries, expiries — must fire near their due time. Scanning everything every minute does not scale, and an in-process timer dies with the process. Work from the scoping questions below. State assumptions for any unspecified load, guarantee or target, then trace your design end to end. Explain one difficult case and a credible alternative; the worked example is a reference, not a required implementation.
Resolve the scoping questions for Delay Queue Scheduler. Separate stated behavior from assumptions, and identify what is outside your design.
Declare relevant volume, latency, freshness, quality or cost targets with units. Show calculations or an evaluation plan that can test them; unspecified targets are your assumptions, not hidden pass criteria.
Explain how your guarantees hold in a difficult case relevant to this subject. Address: What happens if a worker fails during execution? How many actions become due together, and how is recovery paced?
Identify users, required behavior and exclusions. Answer: Which actions run later and who chooses their due times? How late may a scheduled action run?
Define the information owned by the system and the inputs, outputs and errors at its boundaries. Resolve: Can an action be canceled or rescheduled after acceptance? What happens if a worker fails during execution?
Estimate the dominant workload and resource demand with units and explicit assumptions. For a learned system, also state how quality is measured and what data is available.
Draw or describe the responsibilities needed for Delay Queue Scheduler. Trace a representative request, event or job from its input to a visible result; identify durable state owners.
Walk through a difficult case step by step, including detection and recovery. Consider: How many actions become due together, and how is recovery paced?
Compare a credible alternative using your chosen workload and guarantees. Explain a remaining risk, a signal to watch and when you would change the design.