Preparation
What systems are built from
The services that appear in nearly every design, grouped by what they are for, AWS name first. Each group is really one decision, written the way you would say it in the interview. Every board in the library uses this vocabulary.
Traffic in
The choice. Terminate at the edge or at the origin; a plain load balancer or a gateway that authenticates and throttles; L7 (HTTP, routing by path) or L4 (TCP, raw speed).
Route 53
also Cloudflare DNS
DNS with health checks and latency routing: how a person reaches the nearest healthy region, and how a region fails over.
CloudFront
also Cloudflare · Akamai · Fastly
CDN: static files and cacheable responses served from the edge; the origin sees a fraction of the traffic.
ALB
also Nginx · HAProxy · Envoy
Layer 7 load balancer: TLS ends here, routes by path and host, drops instances that fail health checks.
NLB
Layer 4: millions of connections, static IPs, microsecond latency; nothing about HTTP.
API Gateway
also Kong · Apigee
Authentication, throttling, request shaping and WebSockets in front of services; adds latency and cost per request.
WAF
Rate limits per IP, bot control, the OWASP rules; the flood never reaches the service.
Envoy + Istio
also Linkerd · App Mesh
Service mesh: mTLS, retries, timeouts and telemetry between services, without code in each one.
Compute
The choice. Containers you run (predictable at steady load) or serverless (zero ops, scales to zero, cold starts and a 15-minute limit).
EC2
Virtual machines. The default when nothing else fits, and what everything else runs on.
ECS / EKS
also Kubernetes · Nomad
Container orchestration: stateless services, autoscaling, rolling deploys. EKS is Kubernetes; ECS is simpler.
Fargate
Containers without servers to manage; pay per task.
Lambda
also Cloud Functions
Functions on events: glue, small jobs, change feeds. Cold starts and 15 minutes are the limits.
OLTP: rows, transactions, point reads and writes
The choice. SQL with joins and transactions, or a key-value store that scales by key and does neither. Strong consistency or eventual. Scale up one writer, or scale out and give up joins.
Aurora / RDS (PostgreSQL, MySQL)
also Cloud SQL · Postgres · MySQL
Relational, ACID, joins, indexes. Aurora replicates storage across 3 AZs and fails over in about 30 s. One writer is the limit.
DynamoDB
also Cassandra · ScyllaDB · Bigtable
Key-value at any scale: design the table around the queries (PK, SK, GSIs), conditional writes, TTL, Streams, global tables. No joins.
MongoDB / DocumentDB
Documents with flexible schema and secondary indexes; good for objects that are read whole.
Spanner / CockroachDB
also Aurora DSQL · YugabyteDB
Distributed SQL: transactions across shards and regions, at the cost of latency per commit.
Vitess / TiDB
Sharded MySQL when you have MySQL and outgrew one writer.
OLAP: scan millions of rows and aggregate
The choice. OLTP vs OLAP is row vs columnar: serve by key from one, analyse by scan from the other, and never make one do the other's job. Warehouse (load, then query) or lakehouse (query the files where they are). Batch analytics or sub-second real-time OLAP.
Redshift
also Snowflake · BigQuery · Databricks
The warehouse: columnar, joins across billions of rows, for analysts and finance. Fed by CDC and batch loads; never on the serving path.
Athena
also Trino · Presto · BigQuery external tables
SQL over files in S3, paid by bytes scanned: recounts, ad-hoc questions, the nightly batch.
ClickHouse / Druid / Pinot
also Rockset
Real-time OLAP: sub-second aggregates over fresh events for dashboards with many dimensions. You run it.
DuckDB
OLAP in one process over local Parquet: analysis on a laptop or inside a job.
Lake and file formats
The choice. Keep the raw record cheap and forever, readable by anything, and let every store be rebuilt from it. Columnar files for scans; row formats for the wire.
S3
also GCS · MinIO · HDFS
Object storage: eleven nines of durability, versioning, lifecycle to Glacier. The archive everything else is derived from.
Parquet / ORC
Columnar files: a scan reads the three columns it needs, compressed 5–10x.
Avro / Protobuf
Row formats with schemas for the wire and the log; a schema registry keeps producers and consumers compatible.
Iceberg / Hudi / Delta Lake
Table formats over files: transactions, schema evolution and time travel on a lake.
Glue Data Catalog
also Hive Metastore
The table definitions over S3 that Athena, Spark and Redshift Spectrum share.
Cache
The choice. Cache-aside (the service fills it on a miss) or write-through (every write updates it). Expire by TTL or invalidate on change. Redis for data structures, Memcached for plain keys.
ElastiCache (Redis)
also Valkey · MemoryDB · Redis
In-memory data structures: cache, sorted sets for leaderboards and ranges, rate limiters, locks, pub/sub. Replica in another AZ.
Memcached
Plain key-value cache, multithreaded, nothing else. Simpler than Redis when that is all you need.
DAX
A read cache in front of DynamoDB that speaks DynamoDB's API.
In-process cache
also Caffeine · Guava
The fastest cache is the one with no network hop; fine when a little staleness per machine is acceptable.
Search
The choice. An inverted index that lags the source of truth by seconds, or a database `LIKE` that does not scale. Relevance ranking vs exact match.
OpenSearch
also Elasticsearch · Solr
Full-text search, aggregations, geo queries, k-NN vectors. Fed by CDC; eventually consistent with the source.
Algolia
also Typesense · Meilisearch
Search as a service: instant results, typo tolerance, no cluster to run.
Postgres full-text search
Good enough for a small corpus without another system to operate.
Streams and queues
The choice. A log (replayable, ordered per key, many readers) or a queue (delete on read, one consumer group, delays and retries on a timer). Kafka by default; SQS only when a message must wait. Push (SNS, EventBridge) or pull (Kafka, SQS). At-least-once everywhere; exactly-once by idempotent consumers.
Kafka (MSK)
also Redpanda · Pulsar · Confluent
The log: partitions keyed for order, replication factor 3, retention for days, consumer groups with offsets. Replay is the superpower.
Kinesis Data Streams
Kafka-shaped streaming on AWS with shards instead of partitions; less to run, less control.
SQS
also RabbitMQ · ActiveMQ
A queue: visibility timeout, delay queues, dead-letter queues. The tool when a message must be delayed or retried on a schedule.
SNS / EventBridge
also Google Pub/Sub · NATS
Fan-out and an event bus with rules: one event, many subscribers, filtered by content.
Kafka Connect
Sources and sinks without code: S3 sink, JDBC, Debezium. The archiver in most designs.
Change data capture
The choice. CDC (follow the database's commit log) or dual writes (write to two places and pray) or the outbox pattern (write the event in the same transaction, publish it after). CDC keeps caches, search and warehouses in sync with one writer.
Debezium
Reads Postgres, MySQL or MongoDB commit logs into Kafka topics, one per table. The standard way to feed everything downstream from the source of truth.
DynamoDB Streams
The change feed of a DynamoDB table, consumed by Lambda or Kinesis: cache invalidation, projections, triggers.
AWS DMS
Managed replication and CDC between databases, including into S3 and Redshift.
Postgres logical replication
The commit log as a stream, which is what Debezium reads.
Flink CDC
CDC connectors inside a Flink job, so the stream processor reads tables directly.
Stream processing
The choice. Event time (when it happened) or processing time (when we saw it); watermarks and allowed lateness decide when a window is done. Stateful (keyed state, checkpoints) or stateless. One streaming path (kappa) or stream plus batch (lambda).
Flink (Managed Service for Apache Flink)
Event time, watermarks, keyed state in RocksDB, checkpoints to S3, exactly-once with idempotent sinks. The reference stream processor.
Kafka Streams / ksqlDB
Stream processing as a library inside your service, or as SQL over topics; no cluster of its own.
Spark Structured Streaming
Micro-batches over Spark: one engine for stream and batch, with more latency.
Materialize / RisingWave
Streaming SQL databases: incrementally maintained views over streams.
Batch, ETL and orchestration
The choice. Batch (simple, exact, reproducible, late) or stream (fresh, complex, approximate until reconciled). An orchestrator that owns the workflow, or choreography through events. Long transactions as sagas with compensation.
Spark (EMR)
also Databricks · Glue jobs
Distributed batch over the lake: the daily close, backfills, feature building.
Glue
Managed ETL and the catalog; crawlers that discover schemas in S3.
dbt
SQL transformations in the warehouse, versioned and tested.
Airflow
also Dagster · Prefect · MWAA
The scheduler of DAGs: when the daily close runs, what it waits for, who is paged when it is late.
Step Functions
also Temporal · Cadence
Durable workflows with retries and compensation: multi-step business processes that must finish.
EventBridge Scheduler
also cron
Run this at 06:00 UTC. The timer behind every batch.
Celery / Sidekiq
also BullMQ · Resque
Background jobs from a web app: send the email, resize the image, later.
Time series and metrics
The choice. Pull (the collector scrapes) or push (services send). Keep raw for days, downsample for months.
Prometheus + Grafana
also Amazon Managed Prometheus · VictoriaMetrics
Metrics by scraping, PromQL, alerting rules, dashboards. The lingua franca of service metrics.
Timestream
also InfluxDB · TimescaleDB
Time-series databases: devices, sensors, financial ticks, with retention tiers built in.
Graph, vectors and geo
The choice. A graph database when queries walk many hops; adjacency tables in SQL when they walk one. Approximate nearest neighbours for vectors at scale, exact when small. Geohashes or cells for "near me".
Neptune
also Neo4j · JanusGraph · TigerGraph
Graph: friends of friends, fraud rings, dependencies.
pgvector / OpenSearch k-NN
also Pinecone · Milvus · Weaviate
Vector search for embeddings: semantic search, recommendations, RAG.
Redis GEO / PostGIS / S2, H3
Geo queries: nearby drivers, delivery zones, cells that shard the map.
Coordination
The choice. Leader election and locks are for the few places that need exactly one actor; design most things so they do not. Leases with expiry beat locks that can be held forever.
ZooKeeper
also etcd · Consul
Consensus (Raft, ZAB): configuration, leader election, service discovery. Kafka used to depend on it.
DynamoDB conditional writes / Redis leases
A lock with a TTL in a store you already have: good enough for one scheduler at a time.
Raft / Paxos
The algorithms behind every consistent replicated log; name them, do not implement them.
Identity and secrets
The choice. Sessions (revocable, a store) or JWTs (stateless, expire only). Delegated sign-in through OAuth/OIDC over passwords of your own. Keys rotate; every token names its key.
Cognito
also Auth0 · Okta · Keycloak · Firebase Auth
Sign-in, OAuth/OIDC, tokens; nothing about passwords to build yourself.
IAM
Roles between services, least privilege; the reason a leaked instance cannot read every bucket.
Secrets Manager / KMS
also Vault
Secrets with rotation, and the keys that encrypt everything at rest.
Observability and on-call
The choice. Metrics (cheap, aggregate), logs (expensive, exact), traces (a request across services). SLOs with error budgets over alerting on every blip.
CloudWatch
also Datadog · New Relic
Metrics, logs, alarms and dashboards for everything on AWS; the alarms page.
OpenTelemetry + X-Ray
also Jaeger · Zipkin · Honeycomb
Distributed tracing: one request across the ALB, the service and the database, with the slow hop named.
OpenSearch (logs)
also ELK · Loki · Splunk
Structured logs you can search when the metrics say something is wrong.
PagerDuty
also Opsgenie
Who is paged, escalation, the runbook link.
Sentry
Exceptions with stack traces, grouped, from services and clients.
Delivery
The choice. Blue-green (two fleets, switch) or canary (a slice first, watch, widen). Infrastructure as code so an environment can be rebuilt.
Terraform / CloudFormation
also CDK · Pulumi
Infrastructure as code: reviewed, versioned, reproducible.
GitHub Actions / CodePipeline
also Jenkins · ArgoCD
Build, test, deploy; ArgoCD for GitOps into Kubernetes.
LaunchDarkly
also Unleash · AppConfig
Feature flags: ship dark, turn on for 1%, turn off without a deploy.
APIs and protocols
The choice. Synchronous request-response or asynchronous events; push (WebSockets, SSE) or pull (polling); text (JSON) or binary (Protobuf) on the wire.
REST
Resources over HTTP; the default for public APIs. Idempotency keys on writes.
gRPC
Binary, typed, fast, streaming: between services.
GraphQL
One query shaped by the client; watch the N+1s.
WebSockets / SSE / long polling
Server push to browsers: chat, live dashboards, notifications. API Gateway can hold the sockets.
Building blocks everyone knows
The choice. Algorithms, not services; name the one that fits and say its trade-off.
Consistent hashing
Add a node and move 1/N of the keys, not all of them. Caches, shards, Cassandra's ring.
Bloom filter
"Definitely not here" in a few bits per key: skip the disk read, dedupe cheaply. False positives, never false negatives.
HyperLogLog
Count distinct in 12 KB with 1% error: unique visitors, unique clicks.
Token bucket rate limiter
N requests per second with bursts, in Redis or at the gateway.
Snowflake / ULID ids
Time-ordered unique ids without a coordinator.
Quorum (W + R > N)
How replicated stores trade consistency for availability, one knob at a time.
CRDTs
Data types that merge without coordination: counters, sets, collaborative text.
Merkle trees
Find which of a million blocks differ by comparing hashes: anti-entropy, Git, blockchains.
Third-party essentials
The choice. Buy the parts that are not your product.
Stripe
also Adyen · Braintree
Payments, with idempotency keys and webhooks you must handle twice.
SES / SNS / FCM / APNs
also Twilio · SendGrid
Email, SMS and push notifications; each has its own retry and rate rules.
Maps APIs
also Google Maps · Mapbox
Geocoding, routing, tiles.
S3 + CloudFront + MediaConvert
also Mux · Cloudinary
Upload to S3, transcode or resize, serve from the edge: every media pipeline.
See them used: the designs.