Learning/AWS Backend Developer/14 — Capstone & System Design Architectures

Capstone & System Design Architectures

Two things live here: the seven Day-14 system designs, and the capstone the learner assembles across the course.

Part A — The system design method

Applied identically to all seven systems and to the capstone.

flowchart TD
    R[1 Requirements<br/>functional + non-functional] --> T[2 Traffic & scale estimation]
    T --> A[3 API design]
    A --> C[4 Compute]
    C --> D[5 Database]
    D --> CA[6 Caching]
    CA --> M[7 Messaging]
    M --> S[8 Storage]
    S --> SEC[9 Security]
    SEC --> SC[10 Scaling]
    SC --> F[11 Failure handling]
    F --> O[12 Observability]
    O --> CO[13 Cost]

Interview time allocation (45 minutes): requirements + estimation 8 min · API + high-level architecture 10 min · the 2–3 decisions that actually matter 15 min · failure + scaling 8 min · cost + wrap 4 min. The course teaches this budget explicitly, because the most common senior-interview failure is spending 25 minutes on requirements.

Part B — The seven systems

Each ships: requirements, numbers, the four-diagram set, the 2–3 load-bearing decisions, failure analysis against the seven questions, and a cost envelope.

System 1 — High-scale URL shortener

Why first: smallest surface, forces key-design and read-scaling reasoning immediately.

RequirementValue
Writes1,000/s peak
Reads100,000/s peak (100:1 skew)
Redirect latencyp99 < 50ms globally
Retention5 years, ~15B URLs
flowchart TB
    U[User] --> CF[CloudFront]
    CF -->|cache miss| AG[API Gateway]
    AG --> RD[Redirect Service · Lambda]
    RD --> DX[(DAX / ElastiCache)]
    DX --> DDB[(DynamoDB<br/>PK = shortCode)]
    W[Create API] --> WS[Write Service · Fargate]
    WS --> DDB
    WS --> CTR[(Counter / ID allocator)]

The decisions that matter: ID generation (counter-range allocation vs hash-with-collision-retry vs Snowflake — and why random-looking keys are required for DynamoDB partition spread); caching layers (CloudFront does most of the work — the redirect is the ideal cacheable response); DynamoDB over relational because the access pattern is exactly one key lookup forever.

Failure analysis highlight: a cache-miss storm on a viral link, and why the answer is CloudFront + DAX rather than more DynamoDB capacity.

System 2 — Order processing

This is S17, the course's backbone scenario, presented as a full design. See 10 — Production Scenario Plan §3 for the numbers and the architecture. Day 14 adds the interview presentation, the consistency-boundary discussion, and the cost attack.

The decisions that matter: where the durability boundary sits (DB commit, not payment success); the transactional outbox; independent queues per consumer rather than one shared queue; idempotency keyed on (orderId, operation).

System 3 — Notification platform

RequirementValue
Volume10M notifications/day, 80% within a 4-hour window
Channelsemail, SMS, push, in-app
Requirementper-channel failure isolation; user preferences; no duplicate sends
Prioritytransactional (seconds) vs marketing (minutes)
flowchart TB
    SRC[Producing services] --> EB[EventBridge]
    EB --> NS[Notification Service]
    NS --> PREF[(Preferences · DynamoDB)]
    NS --> TPL[Template render]
    NS --> QT[(SQS transactional)]
    NS --> QM[(SQS marketing)]
    QT --> WE[Email Worker] --> SES[SES]
    QT --> WP[Push Worker] --> SNSM[SNS mobile push]
    QM --> WS2[SMS Worker] --> PIN[SNS SMS / Pinpoint]
    WE --> DDQ[(Dedup store · DynamoDB TTL)]
    QT -.-> DLQ[(DLQ per channel)]

The decisions that matter: a queue per channel and per priority (failure isolation is the requirement, and one slow channel must not block another); dedupe with a TTL'd key rather than exactly-once delivery claims; rate limiting per provider, with backpressure rather than dropping.

System 4 — File upload/download platform

RequirementValue
File sizeup to 5 GB
Volume50,000 uploads/day, 2M downloads/day
Requirementvirus scan, metadata search, expiring share links, global delivery
flowchart TB
    C[Client] -->|1 request upload| API[Upload API]
    API -->|2 presigned multipart URLs| C
    C -->|3 PUT parts directly| S3[(S3 raw)]
    S3 -->|4 ObjectCreated| EB[EventBridge]
    EB --> SCAN[Scan Lambda] --> S3C[(S3 clean)]
    SCAN --> META[(DynamoDB metadata)]
    S3C --> CF[CloudFront signed URLs] --> C
    EB -.->|infected| QUAR[(Quarantine + alert)]

The decisions that matter: the application never touches the bytes (presigned + multipart); a two-bucket clean/raw model so an unscanned object is never servable; signed CloudFront URLs rather than presigned S3 URLs for downloads (cache + control); a lifecycle rule for abandoned multipart parts, because that bill is invisible otherwise.

System 5 — Payment / event processing

RequirementValue
Volume5,000 payments/min peak
Requirementnever double-charge; never lose a payment intent; full audit trail; PCI-conscious
Externalprovider with p99 2.5s and periodic outages
flowchart TB
    API[Payment API] -->|idempotency-key header| IDEM[(DynamoDB idempotency<br/>conditional write)]
    API --> DB[(Aurora payments + outbox)]
    OUT[Outbox relay] --> SQS[(SQS payment-intents FIFO<br/>group = accountId)]
    SQS --> W[Payment Worker]
    W --> PROV[Provider]
    W --> DB
    PROV -.->|webhook| WH[Webhook handler] --> DB
    SQS -.-> DLQ[(DLQ)] --> OPS[Manual review queue]
    W --> SFN[Step Functions saga<br/>compensate on partial failure]

The decisions that matter: idempotency enforced at the API edge and at the worker (defence in depth, with the two-store failure window discussed honestly); FIFO grouped by account so one account's ordering is guaranteed without serializing the platform; the saga for multi-step settlement; the DLQ routes to humans, not to retry.

This system is where the course's hardest question lives: what exactly is guaranteed, and to whom, at each of the five hops?

System 6 — Subscription / billing

RequirementValue
Subscribers2M
Billing runsmonthly, with daily proration events
Requirementcorrect money, idempotent retries, dunning, invoicing, tax
flowchart TB
    SCH[EventBridge Scheduler] --> SFN[Step Functions · billing run]
    SFN --> SEG[Segment accounts into batches]
    SEG --> Q[(SQS billing-jobs)]
    Q --> BW[Billing Worker] --> INV[(Aurora invoices)]
    BW --> PAY[Payment system · System 5]
    BW --> S3[(S3 invoice PDFs)]
    PAY -.->|failed| DUN[Dunning state machine]
    INV --> NOTIF[Notification platform · System 3]

The decisions that matter: a batched, resumable billing run rather than a cron that must succeed atomically (2M accounts cannot be one transaction); idempotency per (accountId, billingPeriod); separating "invoice generated" from "payment collected" as distinct state machines; why a scheduled Step Function beats a Lambda with a 15-minute ceiling.

System 7 — Real-time data processing

RequirementValue
Ingest500,000 events/s peak
Latencyaggregates visible within 30s
Retentionraw 7 days hot, 2 years in S3
flowchart TB
    P[Producers] --> KDS[(Kinesis Data Streams<br/>N shards)]
    KDS --> AGG[Aggregation consumers<br/>Lambda / Flink]
    AGG --> DDB[(DynamoDB · live aggregates)]
    KDS --> FH[Firehose] --> S3[(S3 raw, partitioned)]
    S3 --> ATH[Athena / Glue]
    DDB --> API[Query API] --> DASH[Dashboards]
    KDS -.->|iterator age alarm| OPS[On-call]

The decisions that matter: shard count from throughput and partition-key cardinality (the hot-shard trap); Kinesis over SQS because multiple independent consumers need the same ordered data; the hot/cold split (DynamoDB for serving, S3+Athena for history); why iterator age is the only health metric that matters.

Part C — The capstone

Production-Grade Order Processing Platform

Built incrementally as the Mini Assignments of Days 4–13, assembled and drilled on Day 14. Budget 4–6 hours outside the daily schedule.

Architecture

flowchart TB
    C[Client] --> CF[CloudFront]
    CF --> AG[API Gateway<br/>JWT authorizer · throttling]
    AG --> ALB[ALB] --> OS[Order Service · ECS Fargate<br/>private subnets, 2 AZs]
    OS --> RDS[(Aurora PostgreSQL<br/>Multi-AZ · orders + outbox)]
    OS --> RED[(ElastiCache Redis<br/>catalogue cache)]
    OS --> S3[(S3 · order documents)]
    REL[Outbox relay] -.-> RDS
    REL --> EB[EventBridge · order events]
    EB --> QP[(SQS payments)] --> PW[Payment Worker]
    EB --> QI[(SQS inventory)] --> IW[Inventory Worker]
    EB --> QN[(SQS notifications)] --> NW[Notification Worker · Lambda]
    PW --> IDEM[(DynamoDB idempotency)]
    QP -.-> DLQP[(payments DLQ)]
    QI -.-> DLQI[(inventory DLQ)]
    QN -.-> DLQN[(notifications DLQ)]
    SM[Secrets Manager] --> OS
    KMS[KMS] -.-> RDS & S3 & QP
    OS & PW & IW & NW --> CW[CloudWatch Logs · Metrics · X-Ray]
    CW --> AL[Alarms]

What every component is there to prove

ComponentDemonstratesBuilt on day
API Gateway + JWT authorizerAuthentication, throttling at the edge5, 11
ALB + ECS Fargate, 2 AZsHorizontal scale, zero-downtime deploy, AZ tolerance4
Aurora Multi-AZ + Secrets ManagerManaged relational data, failover behaviour, no credentials in code6
Outbox + relayAtomic "save and publish"8, 9
EventBridgeRouting and replay without producer changes9
Three SQS queues + three DLQsFailure isolation, retry, poison handling8
DynamoDB idempotency tableCorrectness under at-least-once delivery7, 8
ElastiCacheRead scaling with TTL jitter7
S3 + presigned URLsLarge objects out of the application path10
KMS on DB, bucket and queuesEncryption at rest, key policy reasoning11
Least-privilege task rolesFour distinct roles, none reusable by another service11
CloudWatch + X-Ray + alarmsTrace an order across five async hops13

Assembly schedule

DayMini Assignment contributes
4Order Service containerized, on Fargate, behind an ALB, autoscaling
5Notification Worker as a Lambda; API Gateway front door
6Aurora persistence, Secrets Manager, Flyway migration
7Redis catalogue cache; DynamoDB idempotency table
8SQS + DLQ + idempotent Payment Worker
9EventBridge routing; Inventory Worker; replay proven
10S3 order documents via presigned URLs; CloudFront
11Four least-privilege roles; KMS; JWT auth
12Autoscaling policies; timeouts, retries with jitter, circuit breaker on the provider call
13Structured logs, correlation IDs across async hops, X-Ray, the alarm set
14Assembly, the failure drill, cost review, teardown

The Day 14 failure drill (the capstone's real exam)

The learner runs five injected failures and must show the system ends in a correct state, with the evidence:

  1. Kill a Fargate task mid-request. → ALB retries/drains; no data loss; verified by the client's continuous request log.
  2. Deliver a duplicate OrderPlaced event. → one payment row; the conditional write rejected the second.
  3. Make the payment provider return 500 for 5 minutes. → messages retried with backoff, circuit breaker opens, queue age rises and recovers, nothing enters the DLQ prematurely.
  4. Poison one message. → exactly maxReceiveCount attempts, then DLQ, then alarm.
  5. Force an Aurora failover. → bounded error window, app reconnects, the outbox is not double-published.

Then the cost review: read the actual bill for the capstone's lifetime, identify the top three line items, and propose a 30% reduction with the trade-off named.

Completion criteria

The capstone is complete when the learner can, in 10 minutes and without notes: draw the architecture, explain why every component exists, name the failure mode each one mitigates, state where the consistency boundary is, and give the monthly cost at 50k orders/min with its three largest lines.