System Design AI

Preparation

What systems are built from

The services that appear in nearly every design, grouped by what they are for, AWS name first. Each group is really one decision, written the way you would say it in the interview. Every board in the library uses this vocabulary.

Traffic in

The choice. Terminate at the edge or at the origin; a plain load balancer or a gateway that authenticates and throttles; L7 (HTTP, routing by path) or L4 (TCP, raw speed).

  • Route 53

    also Cloudflare DNS

    DNS with health checks and latency routing: how a person reaches the nearest healthy region, and how a region fails over.

  • CloudFront

    also Cloudflare · Akamai · Fastly

    CDN: static files and cacheable responses served from the edge; the origin sees a fraction of the traffic.

  • ALB

    also Nginx · HAProxy · Envoy

    Layer 7 load balancer: TLS ends here, routes by path and host, drops instances that fail health checks.

  • NLB

    Layer 4: millions of connections, static IPs, microsecond latency; nothing about HTTP.

  • API Gateway

    also Kong · Apigee

    Authentication, throttling, request shaping and WebSockets in front of services; adds latency and cost per request.

  • WAF

    Rate limits per IP, bot control, the OWASP rules; the flood never reaches the service.

  • Envoy + Istio

    also Linkerd · App Mesh

    Service mesh: mTLS, retries, timeouts and telemetry between services, without code in each one.

Compute

The choice. Containers you run (predictable at steady load) or serverless (zero ops, scales to zero, cold starts and a 15-minute limit).

  • EC2

    Virtual machines. The default when nothing else fits, and what everything else runs on.

  • ECS / EKS

    also Kubernetes · Nomad

    Container orchestration: stateless services, autoscaling, rolling deploys. EKS is Kubernetes; ECS is simpler.

  • Fargate

    Containers without servers to manage; pay per task.

  • Lambda

    also Cloud Functions

    Functions on events: glue, small jobs, change feeds. Cold starts and 15 minutes are the limits.

OLTP: rows, transactions, point reads and writes

The choice. SQL with joins and transactions, or a key-value store that scales by key and does neither. Strong consistency or eventual. Scale up one writer, or scale out and give up joins.

  • Aurora / RDS (PostgreSQL, MySQL)

    also Cloud SQL · Postgres · MySQL

    Relational, ACID, joins, indexes. Aurora replicates storage across 3 AZs and fails over in about 30 s. One writer is the limit.

  • DynamoDB

    also Cassandra · ScyllaDB · Bigtable

    Key-value at any scale: design the table around the queries (PK, SK, GSIs), conditional writes, TTL, Streams, global tables. No joins.

  • MongoDB / DocumentDB

    Documents with flexible schema and secondary indexes; good for objects that are read whole.

  • Spanner / CockroachDB

    also Aurora DSQL · YugabyteDB

    Distributed SQL: transactions across shards and regions, at the cost of latency per commit.

  • Vitess / TiDB

    Sharded MySQL when you have MySQL and outgrew one writer.

OLAP: scan millions of rows and aggregate

The choice. OLTP vs OLAP is row vs columnar: serve by key from one, analyse by scan from the other, and never make one do the other's job. Warehouse (load, then query) or lakehouse (query the files where they are). Batch analytics or sub-second real-time OLAP.

  • Redshift

    also Snowflake · BigQuery · Databricks

    The warehouse: columnar, joins across billions of rows, for analysts and finance. Fed by CDC and batch loads; never on the serving path.

  • Athena

    also Trino · Presto · BigQuery external tables

    SQL over files in S3, paid by bytes scanned: recounts, ad-hoc questions, the nightly batch.

  • ClickHouse / Druid / Pinot

    also Rockset

    Real-time OLAP: sub-second aggregates over fresh events for dashboards with many dimensions. You run it.

  • DuckDB

    OLAP in one process over local Parquet: analysis on a laptop or inside a job.

Lake and file formats

The choice. Keep the raw record cheap and forever, readable by anything, and let every store be rebuilt from it. Columnar files for scans; row formats for the wire.

  • S3

    also GCS · MinIO · HDFS

    Object storage: eleven nines of durability, versioning, lifecycle to Glacier. The archive everything else is derived from.

  • Parquet / ORC

    Columnar files: a scan reads the three columns it needs, compressed 5–10x.

  • Avro / Protobuf

    Row formats with schemas for the wire and the log; a schema registry keeps producers and consumers compatible.

  • Iceberg / Hudi / Delta Lake

    Table formats over files: transactions, schema evolution and time travel on a lake.

  • Glue Data Catalog

    also Hive Metastore

    The table definitions over S3 that Athena, Spark and Redshift Spectrum share.

Cache

The choice. Cache-aside (the service fills it on a miss) or write-through (every write updates it). Expire by TTL or invalidate on change. Redis for data structures, Memcached for plain keys.

  • ElastiCache (Redis)

    also Valkey · MemoryDB · Redis

    In-memory data structures: cache, sorted sets for leaderboards and ranges, rate limiters, locks, pub/sub. Replica in another AZ.

  • Memcached

    Plain key-value cache, multithreaded, nothing else. Simpler than Redis when that is all you need.

  • DAX

    A read cache in front of DynamoDB that speaks DynamoDB's API.

  • In-process cache

    also Caffeine · Guava

    The fastest cache is the one with no network hop; fine when a little staleness per machine is acceptable.

Streams and queues

The choice. A log (replayable, ordered per key, many readers) or a queue (delete on read, one consumer group, delays and retries on a timer). Kafka by default; SQS only when a message must wait. Push (SNS, EventBridge) or pull (Kafka, SQS). At-least-once everywhere; exactly-once by idempotent consumers.

  • Kafka (MSK)

    also Redpanda · Pulsar · Confluent

    The log: partitions keyed for order, replication factor 3, retention for days, consumer groups with offsets. Replay is the superpower.

  • Kinesis Data Streams

    Kafka-shaped streaming on AWS with shards instead of partitions; less to run, less control.

  • SQS

    also RabbitMQ · ActiveMQ

    A queue: visibility timeout, delay queues, dead-letter queues. The tool when a message must be delayed or retried on a schedule.

  • SNS / EventBridge

    also Google Pub/Sub · NATS

    Fan-out and an event bus with rules: one event, many subscribers, filtered by content.

  • Kafka Connect

    Sources and sinks without code: S3 sink, JDBC, Debezium. The archiver in most designs.

Change data capture

The choice. CDC (follow the database's commit log) or dual writes (write to two places and pray) or the outbox pattern (write the event in the same transaction, publish it after). CDC keeps caches, search and warehouses in sync with one writer.

  • Debezium

    Reads Postgres, MySQL or MongoDB commit logs into Kafka topics, one per table. The standard way to feed everything downstream from the source of truth.

  • DynamoDB Streams

    The change feed of a DynamoDB table, consumed by Lambda or Kinesis: cache invalidation, projections, triggers.

  • AWS DMS

    Managed replication and CDC between databases, including into S3 and Redshift.

  • Postgres logical replication

    The commit log as a stream, which is what Debezium reads.

  • Flink CDC

    CDC connectors inside a Flink job, so the stream processor reads tables directly.

Stream processing

The choice. Event time (when it happened) or processing time (when we saw it); watermarks and allowed lateness decide when a window is done. Stateful (keyed state, checkpoints) or stateless. One streaming path (kappa) or stream plus batch (lambda).

  • Flink (Managed Service for Apache Flink)

    Event time, watermarks, keyed state in RocksDB, checkpoints to S3, exactly-once with idempotent sinks. The reference stream processor.

  • Kafka Streams / ksqlDB

    Stream processing as a library inside your service, or as SQL over topics; no cluster of its own.

  • Spark Structured Streaming

    Micro-batches over Spark: one engine for stream and batch, with more latency.

  • Materialize / RisingWave

    Streaming SQL databases: incrementally maintained views over streams.

Batch, ETL and orchestration

The choice. Batch (simple, exact, reproducible, late) or stream (fresh, complex, approximate until reconciled). An orchestrator that owns the workflow, or choreography through events. Long transactions as sagas with compensation.

  • Spark (EMR)

    also Databricks · Glue jobs

    Distributed batch over the lake: the daily close, backfills, feature building.

  • Glue

    Managed ETL and the catalog; crawlers that discover schemas in S3.

  • dbt

    SQL transformations in the warehouse, versioned and tested.

  • Airflow

    also Dagster · Prefect · MWAA

    The scheduler of DAGs: when the daily close runs, what it waits for, who is paged when it is late.

  • Step Functions

    also Temporal · Cadence

    Durable workflows with retries and compensation: multi-step business processes that must finish.

  • EventBridge Scheduler

    also cron

    Run this at 06:00 UTC. The timer behind every batch.

  • Celery / Sidekiq

    also BullMQ · Resque

    Background jobs from a web app: send the email, resize the image, later.

Time series and metrics

The choice. Pull (the collector scrapes) or push (services send). Keep raw for days, downsample for months.

  • Prometheus + Grafana

    also Amazon Managed Prometheus · VictoriaMetrics

    Metrics by scraping, PromQL, alerting rules, dashboards. The lingua franca of service metrics.

  • Timestream

    also InfluxDB · TimescaleDB

    Time-series databases: devices, sensors, financial ticks, with retention tiers built in.

Graph, vectors and geo

The choice. A graph database when queries walk many hops; adjacency tables in SQL when they walk one. Approximate nearest neighbours for vectors at scale, exact when small. Geohashes or cells for "near me".

  • Neptune

    also Neo4j · JanusGraph · TigerGraph

    Graph: friends of friends, fraud rings, dependencies.

  • pgvector / OpenSearch k-NN

    also Pinecone · Milvus · Weaviate

    Vector search for embeddings: semantic search, recommendations, RAG.

  • Redis GEO / PostGIS / S2, H3

    Geo queries: nearby drivers, delivery zones, cells that shard the map.

Coordination

The choice. Leader election and locks are for the few places that need exactly one actor; design most things so they do not. Leases with expiry beat locks that can be held forever.

  • ZooKeeper

    also etcd · Consul

    Consensus (Raft, ZAB): configuration, leader election, service discovery. Kafka used to depend on it.

  • DynamoDB conditional writes / Redis leases

    A lock with a TTL in a store you already have: good enough for one scheduler at a time.

  • Raft / Paxos

    The algorithms behind every consistent replicated log; name them, do not implement them.

Identity and secrets

The choice. Sessions (revocable, a store) or JWTs (stateless, expire only). Delegated sign-in through OAuth/OIDC over passwords of your own. Keys rotate; every token names its key.

  • Cognito

    also Auth0 · Okta · Keycloak · Firebase Auth

    Sign-in, OAuth/OIDC, tokens; nothing about passwords to build yourself.

  • IAM

    Roles between services, least privilege; the reason a leaked instance cannot read every bucket.

  • Secrets Manager / KMS

    also Vault

    Secrets with rotation, and the keys that encrypt everything at rest.

Observability and on-call

The choice. Metrics (cheap, aggregate), logs (expensive, exact), traces (a request across services). SLOs with error budgets over alerting on every blip.

  • CloudWatch

    also Datadog · New Relic

    Metrics, logs, alarms and dashboards for everything on AWS; the alarms page.

  • OpenTelemetry + X-Ray

    also Jaeger · Zipkin · Honeycomb

    Distributed tracing: one request across the ALB, the service and the database, with the slow hop named.

  • OpenSearch (logs)

    also ELK · Loki · Splunk

    Structured logs you can search when the metrics say something is wrong.

  • PagerDuty

    also Opsgenie

    Who is paged, escalation, the runbook link.

  • Sentry

    Exceptions with stack traces, grouped, from services and clients.

Delivery

The choice. Blue-green (two fleets, switch) or canary (a slice first, watch, widen). Infrastructure as code so an environment can be rebuilt.

  • Terraform / CloudFormation

    also CDK · Pulumi

    Infrastructure as code: reviewed, versioned, reproducible.

  • GitHub Actions / CodePipeline

    also Jenkins · ArgoCD

    Build, test, deploy; ArgoCD for GitOps into Kubernetes.

  • LaunchDarkly

    also Unleash · AppConfig

    Feature flags: ship dark, turn on for 1%, turn off without a deploy.

APIs and protocols

The choice. Synchronous request-response or asynchronous events; push (WebSockets, SSE) or pull (polling); text (JSON) or binary (Protobuf) on the wire.

  • REST

    Resources over HTTP; the default for public APIs. Idempotency keys on writes.

  • gRPC

    Binary, typed, fast, streaming: between services.

  • GraphQL

    One query shaped by the client; watch the N+1s.

  • WebSockets / SSE / long polling

    Server push to browsers: chat, live dashboards, notifications. API Gateway can hold the sockets.

Building blocks everyone knows

The choice. Algorithms, not services; name the one that fits and say its trade-off.

  • Consistent hashing

    Add a node and move 1/N of the keys, not all of them. Caches, shards, Cassandra's ring.

  • Bloom filter

    "Definitely not here" in a few bits per key: skip the disk read, dedupe cheaply. False positives, never false negatives.

  • HyperLogLog

    Count distinct in 12 KB with 1% error: unique visitors, unique clicks.

  • Token bucket rate limiter

    N requests per second with bursts, in Redis or at the gateway.

  • Snowflake / ULID ids

    Time-ordered unique ids without a coordinator.

  • Quorum (W + R > N)

    How replicated stores trade consistency for availability, one knob at a time.

  • CRDTs

    Data types that merge without coordination: counters, sets, collaborative text.

  • Merkle trees

    Find which of a million blocks differ by comparing hashes: anti-entropy, Git, blockchains.

Third-party essentials

The choice. Buy the parts that are not your product.

  • Stripe

    also Adyen · Braintree

    Payments, with idempotency keys and webhooks you must handle twice.

  • SES / SNS / FCM / APNs

    also Twilio · SendGrid

    Email, SMS and push notifications; each has its own retry and rate rules.

  • Maps APIs

    also Google Maps · Mapbox

    Geocoding, routing, tiles.

  • S3 + CloudFront + MediaConvert

    also Mux · Cloudinary

    Upload to S3, transcode or resize, serve from the edge: every media pipeline.

See them used: the designs.