Build a release decision from noisy evaluations, customer slices and production feedback.
AI, machine-learning and data practitioners.
Your approach: Use prose or diagrams to explain data and model behavior, evaluation evidence and lifecycle controls requested in the question.
Design the evaluation and rollout system for replacing a support assistant model with a cheaper candidate. The candidate improves average benchmark scores but may fail disproportionately on rare, high-cost cases. Describe how the release decision is made and reversed.
Declare the daily human-reviewed sample, including calibration and critical slices.
Required target: ≤ 300 cases/day
Separate development examples from held-out release evidence; record model, prompt, tool and dataset versions.
Overall average improvements cannot waive the defined critical-safety or incorrect-tool-action gate.
Identify the rollback signal, decision owner and how to attribute behavior to a version after rollback.
Define task-success, tool correctness, escalation cost and customer slices, including at least one rare critical case.
Describe human labels, judge disagreement, uncertainty and protection against style/verbosity bias.
Specify offline gates, traffic assignment, guardrails, sample limitations and go/no-go ownership.
Explain failure triage, dataset updates without holdout leakage, and total cost including escalations.