19. Ship a cheaper model without hiding regressions
Build a release decision from noisy evaluations, customer slices and production feedback.
The brief
Design the evaluation and rollout system for replacing a support assistant model with a cheaper candidate. The candidate improves average benchmark scores but may fail disproportionately on rare, high-cost cases. Describe how the release decision is made and reversed.
- The product handles 100,000 conversations/day. Human review capacity is 300 cases/day.
- Existing ratings are sparse and skewed toward dissatisfied users; an LLM judge sometimes prefers verbose answers.
- The candidate costs 35% less per call, but tool errors can trigger expensive manual escalations.
Constraints
- Human review budget≤ 300 cases/day
- Declare the daily human-reviewed sample, including calibration and critical slices.
- Evaluation integrity
- Separate development examples from held-out release evidence; record model, prompt, tool and dataset versions.
- Critical error gate
- Overall average improvements cannot waive the defined critical-safety or incorrect-tool-action gate.
- Release reversal
- Identify the rollback signal, decision owner and how to attribute behavior to a version after rollback.
What to cover
- 01
Metrics and slices
Define task-success, tool correctness, escalation cost and customer slices, including at least one rare critical case.
- 02
Judging and calibration
Describe human labels, judge disagreement, uncertainty and protection against style/verbosity bias.
- 03
Release experiment
Specify offline gates, traffic assignment, guardrails, sample limitations and go/no-go ownership.
- 04
Feedback and cost loop
Explain failure triage, dataset updates without holdout leakage, and total cost including escalations.
Worked designs
No worked design has been published for this brief yet. You can start an attempt and share your approach in the discussion.
Review rubric
AI feedback uses these criteria. Scores are practice feedback.
Representative release evidence
Metrics and slices reflect real task outcomes; versioned holdouts and sampling expose rare failures.
Judge calibration and uncertainty
Uses human comparison, disagreement and uncertainty without presenting automated scores as ground truth.
Experiment and rollback design
Defines attributable cohorts, guardrails, decision owners and reversible rollout under limited evidence.
End-to-end economics
Quantifies the tradeoff between model savings, escalations and the constrained review budget.
Discussion
Share an approach, ask a question, or tag @Coach.
Loading discussion…