Official · AI & data · intermediate
Ship a cheaper model without hiding regressions
Build a release decision from noisy evaluations, customer slices and production feedback.
The question
Design the evaluation and rollout system for replacing a support assistant model with a cheaper candidate. The candidate improves average benchmark scores but may fail disproportionately on rare, high-cost cases. Describe how the release decision is made and reversed.
- The product handles 100,000 conversations/day. Human review capacity is 300 cases/day.
- Existing ratings are sparse and skewed toward dissatisfied users; an LLM judge sometimes prefers verbose answers.
- The candidate costs 35% less per call, but tool errors can trigger expensive manual escalations.
What to cover
Define task-success, tool correctness, escalation cost and customer slices, including at least one rare critical case.
Describe human labels, judge disagreement, uncertainty and protection against style/verbosity bias.
Specify offline gates, traffic assignment, guardrails, sample limitations and go/no-go ownership.
Explain failure triage, dataset updates without holdout leakage, and total cost including escalations.
How your practice is reviewed
- Representative release evidence (30 points): Metrics and slices reflect real task outcomes; versioned holdouts and sampling expose rare failures.
- Judge calibration and uncertainty (30 points): Uses human comparison, disagreement and uncertainty without presenting automated scores as ground truth.
- Experiment and rollback design (25 points): Defines attributable cohorts, guardrails, decision owners and reversible rollout under limited evidence.
- End-to-end economics (15 points): Quantifies the tradeoff between model savings, escalations and the constrained review budget.
Discuss & learn
Ask the community, or tag @Coach for a contextual AI answer.