Ship a cheaper model without hiding regressions
Build a release decision from noisy evaluations, customer slices and production feedback.
Machine-learning engineers designing data, training and model systems.
Your approach: Explain the data, learning or model lifecycle in the brief. Support relevant quality and operating targets with evidence.
The problem
Design the evaluation and rollout system for replacing a support assistant model with a cheaper candidate. The candidate improves average benchmark scores but may fail disproportionately on rare, high-cost cases. Describe how the release decision is made and reversed.
- The product handles 100,000 conversations/day. Human review capacity is 300 cases/day.
- Existing ratings are sparse and skewed toward dissatisfied users; an LLM judge sometimes prefers verbose answers.
- The candidate costs 35% less per call, but tool errors can trigger expensive manual escalations.
Work within these constraints
Declare the daily human-reviewed sample, including calibration and critical slices.
Required target: ≤ 300 cases/day
Separate development examples from held-out release evidence; record model, prompt, tool and dataset versions.
Overall average improvements cannot waive the defined critical-safety or incorrect-tool-action gate.
Identify the rollback signal, decision owner and how to attribute behavior to a version after rollback.
What to deliver
Metrics and slices
Define task-success, tool correctness, escalation cost and customer slices, including at least one rare critical case.
Judging and calibration
Describe human labels, judge disagreement, uncertainty and protection against style/verbosity bias.
Release experiment
Specify offline gates, traffic assignment, guardrails, sample limitations and go/no-go ownership.
Feedback and cost loop
Explain failure triage, dataset updates without holdout leakage, and total cost including escalations.