Capstone & System Design Architectures
Two things live here: the seven Day-14 system designs, and the capstone the learner assembles across the course.
Part A — The system design method
Applied identically to all seven systems and to the capstone.
flowchart TD
R[1 Requirements<br/>functional + non-functional] --> T[2 Traffic & scale estimation]
T --> A[3 API design]
A --> C[4 Compute]
C --> D[5 Database]
D --> CA[6 Caching]
CA --> M[7 Messaging]
M --> S[8 Storage]
S --> SEC[9 Security]
SEC --> SC[10 Scaling]
SC --> F[11 Failure handling]
F --> O[12 Observability]
O --> CO[13 Cost]Interview time allocation (45 minutes): requirements + estimation 8 min · API + high-level architecture 10 min · the 2–3 decisions that actually matter 15 min · failure + scaling 8 min · cost + wrap 4 min. The course teaches this budget explicitly, because the most common senior-interview failure is spending 25 minutes on requirements.
Part B — The seven systems
Each ships: requirements, numbers, the four-diagram set, the 2–3 load-bearing decisions, failure analysis against the seven questions, and a cost envelope.
System 1 — High-scale URL shortener
Why first: smallest surface, forces key-design and read-scaling reasoning immediately.
| Requirement | Value |
|---|---|
| Writes | 1,000/s peak |
| Reads | 100,000/s peak (100:1 skew) |
| Redirect latency | p99 < 50ms globally |
| Retention | 5 years, ~15B URLs |
flowchart TB
U[User] --> CF[CloudFront]
CF -->|cache miss| AG[API Gateway]
AG --> RD[Redirect Service · Lambda]
RD --> DX[(DAX / ElastiCache)]
DX --> DDB[(DynamoDB<br/>PK = shortCode)]
W[Create API] --> WS[Write Service · Fargate]
WS --> DDB
WS --> CTR[(Counter / ID allocator)]The decisions that matter: ID generation (counter-range allocation vs hash-with-collision-retry vs Snowflake — and why random-looking keys are required for DynamoDB partition spread); caching layers (CloudFront does most of the work — the redirect is the ideal cacheable response); DynamoDB over relational because the access pattern is exactly one key lookup forever.
Failure analysis highlight: a cache-miss storm on a viral link, and why the answer is CloudFront + DAX rather than more DynamoDB capacity.
System 2 — Order processing
This is S17, the course's backbone scenario, presented as a full design. See 10 — Production Scenario Plan §3 for the numbers and the architecture. Day 14 adds the interview presentation, the consistency-boundary discussion, and the cost attack.
The decisions that matter: where the durability boundary sits (DB commit, not payment success); the transactional outbox; independent queues per consumer rather than one shared queue; idempotency keyed on (orderId, operation).
System 3 — Notification platform
| Requirement | Value |
|---|---|
| Volume | 10M notifications/day, 80% within a 4-hour window |
| Channels | email, SMS, push, in-app |
| Requirement | per-channel failure isolation; user preferences; no duplicate sends |
| Priority | transactional (seconds) vs marketing (minutes) |
flowchart TB
SRC[Producing services] --> EB[EventBridge]
EB --> NS[Notification Service]
NS --> PREF[(Preferences · DynamoDB)]
NS --> TPL[Template render]
NS --> QT[(SQS transactional)]
NS --> QM[(SQS marketing)]
QT --> WE[Email Worker] --> SES[SES]
QT --> WP[Push Worker] --> SNSM[SNS mobile push]
QM --> WS2[SMS Worker] --> PIN[SNS SMS / Pinpoint]
WE --> DDQ[(Dedup store · DynamoDB TTL)]
QT -.-> DLQ[(DLQ per channel)]The decisions that matter: a queue per channel and per priority (failure isolation is the requirement, and one slow channel must not block another); dedupe with a TTL'd key rather than exactly-once delivery claims; rate limiting per provider, with backpressure rather than dropping.
System 4 — File upload/download platform
| Requirement | Value |
|---|---|
| File size | up to 5 GB |
| Volume | 50,000 uploads/day, 2M downloads/day |
| Requirement | virus scan, metadata search, expiring share links, global delivery |
flowchart TB
C[Client] -->|1 request upload| API[Upload API]
API -->|2 presigned multipart URLs| C
C -->|3 PUT parts directly| S3[(S3 raw)]
S3 -->|4 ObjectCreated| EB[EventBridge]
EB --> SCAN[Scan Lambda] --> S3C[(S3 clean)]
SCAN --> META[(DynamoDB metadata)]
S3C --> CF[CloudFront signed URLs] --> C
EB -.->|infected| QUAR[(Quarantine + alert)]The decisions that matter: the application never touches the bytes (presigned + multipart); a two-bucket clean/raw model so an unscanned object is never servable; signed CloudFront URLs rather than presigned S3 URLs for downloads (cache + control); a lifecycle rule for abandoned multipart parts, because that bill is invisible otherwise.
System 5 — Payment / event processing
| Requirement | Value |
|---|---|
| Volume | 5,000 payments/min peak |
| Requirement | never double-charge; never lose a payment intent; full audit trail; PCI-conscious |
| External | provider with p99 2.5s and periodic outages |
flowchart TB
API[Payment API] -->|idempotency-key header| IDEM[(DynamoDB idempotency<br/>conditional write)]
API --> DB[(Aurora payments + outbox)]
OUT[Outbox relay] --> SQS[(SQS payment-intents FIFO<br/>group = accountId)]
SQS --> W[Payment Worker]
W --> PROV[Provider]
W --> DB
PROV -.->|webhook| WH[Webhook handler] --> DB
SQS -.-> DLQ[(DLQ)] --> OPS[Manual review queue]
W --> SFN[Step Functions saga<br/>compensate on partial failure]The decisions that matter: idempotency enforced at the API edge and at the worker (defence in depth, with the two-store failure window discussed honestly); FIFO grouped by account so one account's ordering is guaranteed without serializing the platform; the saga for multi-step settlement; the DLQ routes to humans, not to retry.
This system is where the course's hardest question lives: what exactly is guaranteed, and to whom, at each of the five hops?
System 6 — Subscription / billing
| Requirement | Value |
|---|---|
| Subscribers | 2M |
| Billing runs | monthly, with daily proration events |
| Requirement | correct money, idempotent retries, dunning, invoicing, tax |
flowchart TB
SCH[EventBridge Scheduler] --> SFN[Step Functions · billing run]
SFN --> SEG[Segment accounts into batches]
SEG --> Q[(SQS billing-jobs)]
Q --> BW[Billing Worker] --> INV[(Aurora invoices)]
BW --> PAY[Payment system · System 5]
BW --> S3[(S3 invoice PDFs)]
PAY -.->|failed| DUN[Dunning state machine]
INV --> NOTIF[Notification platform · System 3]The decisions that matter: a batched, resumable billing run rather than a cron that must succeed atomically (2M accounts cannot be one transaction); idempotency per (accountId, billingPeriod); separating "invoice generated" from "payment collected" as distinct state machines; why a scheduled Step Function beats a Lambda with a 15-minute ceiling.
System 7 — Real-time data processing
| Requirement | Value |
|---|---|
| Ingest | 500,000 events/s peak |
| Latency | aggregates visible within 30s |
| Retention | raw 7 days hot, 2 years in S3 |
flowchart TB
P[Producers] --> KDS[(Kinesis Data Streams<br/>N shards)]
KDS --> AGG[Aggregation consumers<br/>Lambda / Flink]
AGG --> DDB[(DynamoDB · live aggregates)]
KDS --> FH[Firehose] --> S3[(S3 raw, partitioned)]
S3 --> ATH[Athena / Glue]
DDB --> API[Query API] --> DASH[Dashboards]
KDS -.->|iterator age alarm| OPS[On-call]The decisions that matter: shard count from throughput and partition-key cardinality (the hot-shard trap); Kinesis over SQS because multiple independent consumers need the same ordered data; the hot/cold split (DynamoDB for serving, S3+Athena for history); why iterator age is the only health metric that matters.
Part C — The capstone
Production-Grade Order Processing Platform
Built incrementally as the Mini Assignments of Days 4–13, assembled and drilled on Day 14. Budget 4–6 hours outside the daily schedule.
Architecture
flowchart TB
C[Client] --> CF[CloudFront]
CF --> AG[API Gateway<br/>JWT authorizer · throttling]
AG --> ALB[ALB] --> OS[Order Service · ECS Fargate<br/>private subnets, 2 AZs]
OS --> RDS[(Aurora PostgreSQL<br/>Multi-AZ · orders + outbox)]
OS --> RED[(ElastiCache Redis<br/>catalogue cache)]
OS --> S3[(S3 · order documents)]
REL[Outbox relay] -.-> RDS
REL --> EB[EventBridge · order events]
EB --> QP[(SQS payments)] --> PW[Payment Worker]
EB --> QI[(SQS inventory)] --> IW[Inventory Worker]
EB --> QN[(SQS notifications)] --> NW[Notification Worker · Lambda]
PW --> IDEM[(DynamoDB idempotency)]
QP -.-> DLQP[(payments DLQ)]
QI -.-> DLQI[(inventory DLQ)]
QN -.-> DLQN[(notifications DLQ)]
SM[Secrets Manager] --> OS
KMS[KMS] -.-> RDS & S3 & QP
OS & PW & IW & NW --> CW[CloudWatch Logs · Metrics · X-Ray]
CW --> AL[Alarms]What every component is there to prove
| Component | Demonstrates | Built on day |
|---|---|---|
| API Gateway + JWT authorizer | Authentication, throttling at the edge | 5, 11 |
| ALB + ECS Fargate, 2 AZs | Horizontal scale, zero-downtime deploy, AZ tolerance | 4 |
| Aurora Multi-AZ + Secrets Manager | Managed relational data, failover behaviour, no credentials in code | 6 |
| Outbox + relay | Atomic "save and publish" | 8, 9 |
| EventBridge | Routing and replay without producer changes | 9 |
| Three SQS queues + three DLQs | Failure isolation, retry, poison handling | 8 |
| DynamoDB idempotency table | Correctness under at-least-once delivery | 7, 8 |
| ElastiCache | Read scaling with TTL jitter | 7 |
| S3 + presigned URLs | Large objects out of the application path | 10 |
| KMS on DB, bucket and queues | Encryption at rest, key policy reasoning | 11 |
| Least-privilege task roles | Four distinct roles, none reusable by another service | 11 |
| CloudWatch + X-Ray + alarms | Trace an order across five async hops | 13 |
Assembly schedule
| Day | Mini Assignment contributes |
|---|---|
| 4 | Order Service containerized, on Fargate, behind an ALB, autoscaling |
| 5 | Notification Worker as a Lambda; API Gateway front door |
| 6 | Aurora persistence, Secrets Manager, Flyway migration |
| 7 | Redis catalogue cache; DynamoDB idempotency table |
| 8 | SQS + DLQ + idempotent Payment Worker |
| 9 | EventBridge routing; Inventory Worker; replay proven |
| 10 | S3 order documents via presigned URLs; CloudFront |
| 11 | Four least-privilege roles; KMS; JWT auth |
| 12 | Autoscaling policies; timeouts, retries with jitter, circuit breaker on the provider call |
| 13 | Structured logs, correlation IDs across async hops, X-Ray, the alarm set |
| 14 | Assembly, the failure drill, cost review, teardown |
The Day 14 failure drill (the capstone's real exam)
The learner runs five injected failures and must show the system ends in a correct state, with the evidence:
- Kill a Fargate task mid-request. → ALB retries/drains; no data loss; verified by the client's continuous request log.
- Deliver a duplicate
OrderPlacedevent. → one payment row; the conditional write rejected the second. - Make the payment provider return 500 for 5 minutes. → messages retried with backoff, circuit breaker opens, queue age rises and recovers, nothing enters the DLQ prematurely.
- Poison one message. → exactly
maxReceiveCountattempts, then DLQ, then alarm. - Force an Aurora failover. → bounded error window, app reconnects, the outbox is not double-published.
Then the cost review: read the actual bill for the capstone's lifetime, identify the top three line items, and propose a 30% reduction with the trade-off named.
Completion criteria
The capstone is complete when the learner can, in 10 minutes and without notes: draw the architecture, explain why every component exists, name the failure mode each one mitigates, state where the consistency boundary is, and give the monthly cost at 50k orders/min with its three largest lines.