System Design AI
All designs

Machine learningintermediate

19. Ship a cheaper model without hiding regressions

Build a release decision from noisy evaluations, customer slices and production feedback.

The brief

Design the evaluation and rollout system for replacing a support assistant model with a cheaper candidate. The candidate improves average benchmark scores but may fail disproportionately on rare, high-cost cases. Describe how the release decision is made and reversed.

  • The product handles 100,000 conversations/day. Human review capacity is 300 cases/day.
  • Existing ratings are sparse and skewed toward dissatisfied users; an LLM judge sometimes prefers verbose answers.
  • The candidate costs 35% less per call, but tool errors can trigger expensive manual escalations.

Constraints

Human review budget≤ 300 cases/day
Declare the daily human-reviewed sample, including calibration and critical slices.
Evaluation integrity
Separate development examples from held-out release evidence; record model, prompt, tool and dataset versions.
Critical error gate
Overall average improvements cannot waive the defined critical-safety or incorrect-tool-action gate.
Release reversal
Identify the rollback signal, decision owner and how to attribute behavior to a version after rollback.

What to cover

  1. 01

    Metrics and slices

    Define task-success, tool correctness, escalation cost and customer slices, including at least one rare critical case.

  2. 02

    Judging and calibration

    Describe human labels, judge disagreement, uncertainty and protection against style/verbosity bias.

  3. 03

    Release experiment

    Specify offline gates, traffic assignment, guardrails, sample limitations and go/no-go ownership.

  4. 04

    Feedback and cost loop

    Explain failure triage, dataset updates without holdout leakage, and total cost including escalations.

Worked designs

No worked design has been published for this brief yet. You can start an attempt and share your approach in the discussion.

Review rubric

AI feedback uses these criteria. Scores are practice feedback.

Representative release evidence

Metrics and slices reflect real task outcomes; versioned holdouts and sampling expose rare failures.

30points

Judge calibration and uncertainty

Uses human comparison, disagreement and uncertainty without presenting automated scores as ground truth.

30points

Experiment and rollback design

Defines attributable cohorts, guardrails, decision owners and reversible rollout under limited evidence.

25points

End-to-end economics

Quantifies the tradeoff between model savings, escalations and the constrained review budget.

15points

Discussion

Share an approach, ask a question, or tag @Coach.

Share your interview experience

Published under your alias as a community report. Leave out interviewer names and confidential material.

Loading discussion…