83. Design a Vector Database
Nearest-neighbour search over a billion embeddings: HNSW, IVF-PQ and disk graphs, filters, segments, sharding, and hybrid search with re-ranking.
Start with a template. Work through each step. Ask Coach when you need a second opinion.
Company tags are community-reported. Counts on cards show how many people reported that design.
Nearest-neighbour search over a billion embeddings: HNSW, IVF-PQ and disk graphs, filters, segments, sharding, and hybrid search with re-ranking.
Embed a billion chunks on GPUs, keep a k-NN index fresh through CDC with versioned writes and provable deletes, and migrate models blue-green.
Exact, semantic and prefix caching in front of LLMs: scoped keys, distances with an error budget, tenant isolation, invalidation and measured savings.
Search by photo over a billion images: contrastive embeddings from engagement pairs, object crops, IVF-PQ with re-scoring, versioned indexes.
Pick 12 homes a guest could book instead, from listing embeddings learned on browsing sessions, filtered by dates and party size, then ranked.