System Design AI
All designs

Classic designMachine learningintermediate

122. Design an Experiment Tracking and Model Registry Service

Experiment tracking and a model registry: non-blocking logging, chunked metric curves, content-addressed artifacts, lineage and a promotion gate. One-hour boards for junior, senior and staff, with the theory behind them.

The brief

Design an Experiment Tracking and Model Registry Service. Experiment tracking and a model registry: non-blocking logging, chunked metric curves, content-addressed artifacts, lineage and a promotion gate. One-hour boards for junior, senior and staff, with the theory behind them. Work from the scoping questions below. State assumptions for any unspecified load, guarantee or target, then trace your design end to end. Explain one difficult case and a credible alternative; the worked example is a reference, not a required implementation.

  • Set the scope: How much is logged? How fresh must charts be?
  • Define the contract: How long are runs? What are artifacts?
  • Test the boundaries: What gates production? Who logs from a 512-GPU job? Self-hosted or SaaS?

Constraints

Explicit scope and guarantees
Resolve the scoping questions for an Experiment Tracking and Model Registry Service. Separate stated behavior from assumptions, and identify what is outside your design.
Supported operating targets
Declare relevant volume, latency, freshness, quality or cost targets with units. Show calculations or an evaluation plan that can test them; unspecified targets are your assumptions, not hidden pass criteria.
Failure and boundary behavior
Explain how your guarantees hold in a difficult case relevant to this subject. Address: Who logs from a 512-GPU job? Self-hosted or SaaS?

What to cover

  1. 01

    Scope and behavior contract

    Identify users, required behavior and exclusions. Answer: How much is logged? How fresh must charts be?

  2. 02

    State and interfaces

    Define the information owned by the system and the inputs, outputs and errors at its boundaries. Resolve: How long are runs? What are artifacts?

  3. 03

    Capacity and operating targets

    Estimate the dominant workload and resource demand with units and explicit assumptions. For a learned system, also state how quality is measured and what data is available.

  4. 04

    Architecture and central flow

    Draw or describe the responsibilities needed for an Experiment Tracking and Model Registry Service. Trace a representative request, event or job from its input to a visible result; identify durable state owners.

  5. 05

    Failure and boundary walkthrough

    Walk through a difficult case step by step, including detection and recovery. Consider: What gates production? Who logs from a 512-GPU job? Self-hosted or SaaS?

  6. 06

    Tradeoffs and operations

    Compare a credible alternative using your chosen workload and guarantees. Explain a remaining risk, a signal to watch and when you would change the design.

Worked designs

Explore the architecture and decisions, then build on an example with Coach.

Review rubric

AI feedback uses these criteria. Scores are practice feedback.

Scope and contracts

The scoping questions have explicit, consistent answers.

25points

End-to-end design

State ownership and the central flow satisfy the chosen scope.

30points

Operating evidence

Calculations or evaluations support the declared targets.

20points

Boundaries and tradeoffs

A difficult case and an alternative are traced concretely.

25points

Discussion

Share an approach, ask a question, or tag @Coach.

Loading discussion…