System Design AI
Design brief
AI systemsAdvanced55 min suggested

Keep enterprise AI search fresh and private

Handle document changes, ACL revocation and deletion across a retrieval pipeline.

Engineers designing AI products and the systems that serve them.

Your approach: Explain the AI behavior and system boundaries in the brief, with evidence for the requested quality, privacy, latency and cost.

The problem

Design an internal question-answering product over enterprise documents. Focus on the lifecycle from ingestion to a cited answer, especially when permission or source content changes. Explain a useful degraded response when freshness or authorization cannot be established.

  • Ten million documents change at up to 200 updates/second; the product serves 400 questions/second.
  • Permissions exist at both workspace and document level. A user can lose access while a retrieval or generation request is running.
  • Documents may contain instructions aimed at the assistant. Answers must distinguish source evidence from model assumptions.

Work within these constraints

Content freshness target

Declare the p95 delay from an accepted source update to retrieval reflecting it under normal operation.

Required target: ≤ 300 seconds

Access revocation target

Declare the maximum intended stale-access window and specify where authorization is rechecked.

Required target: ≤ 60 seconds

Deletion lifecycle

Track removal across source cache, chunk store, embeddings and answer caches; describe backup retention separately.

Grounding and untrusted documents

Source text cannot override system instructions or authorization; unsupported answers need an explicit abstention or uncertainty path.

What to deliver

1

Data and query paths

Show ingestion, versioned chunks/indexes, retrieval, authorization checks and cited answer assembly.

2

Change and deletion protocol

Trace an ACL revoke and a deleted document during an in-flight question; explain cache invalidation and reconciliation.

3

Quality and security evaluation

Define test slices, relevance/grounding/leakage measures, a baseline and launch thresholds.

4

Latency, cost and failure budget

Size ingest/query load, estimate per-answer cost assumptions and choose a degraded behavior when dependencies fail.