107. Design an ML Model Serving and Deployment Platform
From registry to traffic for hundreds of models: gates, shadow and canary with automatic rollback, packed CPU and GPU pools, LLM serving, cells and audit.
Start with a template. Work through each step. Ask Coach when you need a second opinion.
Company tags are community-reported. Counts on cards show how many people reported that design.
From registry to traffic for hundreds of models: gates, shadow and canary with automatic rollback, packed CPU and GPU pools, LLM serving, cells and audit.
Know within hours when hundreds of production models see broken inputs, drift or falling quality, before the labels arrive.
GPU batch generation behind a shared cache, sandboxed pass@k, calibrated judges, paired statistics, release gates, contamination checks.
Experiment tracking and a model registry: non-blocking logging, chunked metric curves, content-addressed artifacts, lineage and a promotion gate. One-hour boards for junior, senior and staff, with the theory behind them.
A tuning service like Vizier: random and Bayesian search, ASHA and population-based training, trials on a shared GPU cluster.