Learning/AWS Backend Developer/10 — Production Scenario Plan

Production Scenario Plan

Every major topic carries a realistic corporate scenario with real numbers. Scenarios are how the course converts "here is a service" into "here is a decision you will have to defend."

1. Scenario format

Each scenario is written to this shape:

Context        the company, the product, the constraint that makes it interesting
The numbers    traffic, data size, growth, latency budget, cost ceiling
The ask        what you are being asked to decide or fix
Naive approach what a competent engineer without AWS depth would do
What breaks    at what scale, and how it manifests
Better design  with the reasoning, not just the answer
Trade-offs     what you gave up
Cost           actual arithmetic, monthly
The interview  how you would present this in 5 minutes

Numbers are internally consistent and defensible. Where a figure is a current AWS limit or price, it is dated and flagged as subject to change.

2. The scenario catalogue

Days 1–3 · Fundamentals

#ScenarioDayThe decision
S1A startup's single EC2 box serves 500 RPS and the CTO asks "what happens if that AZ dies?"1Why Region/AZ topology is an application concern
S2A developer's laptop credentials are found in a public GitHub repo1Roles vs keys; the blast radius of an account
S3"The app can't reach the database" — six possible causes, one symptom2The connectivity debugging order
S4The monthly bill shows $1,400 of NAT Gateway charges for a service that only talks to S32Gateway endpoints; cost as a design signal
S5A t3.micro fleet is fine for a week, then latency triples every afternoon3Burstable credits; instance family selection
S6CloudWatch Logs is the third-largest line item on the bill3Retention, log level, sampling

Days 4–7 · Core backend

#ScenarioDayThe decision
S7Every deploy drops ~2% of requests for 30 seconds4Deregistration delay + graceful shutdown + health check tuning
S8Black Friday: 8× traffic arriving over 4 minutes; autoscaling "works" but users see errors4Scale-out lag, pre-scaling, load shedding
S9A Java Lambda behind API Gateway shows a 3.2s p99 and a 180ms p505Cold starts; provisioned concurrency vs Fargate; the cost crossover
S10A Lambda triggered by S3 writes to the same bucket. The bill is $11,000 in 6 hours.5, 12Recursive invocation; the prefix/suffix filter; the spend guardrail
S11200 Fargate tasks × 20-connection pools meet a database that allows 500 connections6Pool math; RDS Proxy; the deploy that doubles both
S12A Multi-AZ failover takes 40 seconds of application errors; the team expected zero6What Multi-AZ actually promises; JVM DNS TTL; retry on reconnect
S13Reporting queries on a read replica return orders that "don't exist yet"6Replica lag; read-your-writes; routing reads deliberately
S14A DynamoDB table is throttled at 40% of provisioned capacity7Hot partition; key design; adaptive capacity's limits
S15A cache node restarts at 09:00 and the database falls over at 09:00:037Stampede; TTL jitter; request coalescing; warmup
S16A "simple" Scan in a nightly job costs $900/month7Query vs scan; GSI design; the access-pattern audit

Days 8–11 · Advanced backend

#ScenarioDayThe decision
S17E-commerce: 50,000 orders/minute. Payments, inventory and notification must all happen.8The canonical async decomposition — full treatment
S18A customer is charged twice. The logs show one order and two payment messages.8At-least-once delivery; idempotency keys; conditional writes
S19A worker takes 90 seconds; visibility timeout is 30 seconds. The queue never drains and the DLQ fills.8Visibility timeout math; heartbeating; maxReceiveCount
S20A FIFO queue "isn't scaling" — throughput is stuck at ~10 msg/s8Message group ID as the parallelism unit
S21A new consumer is added and the producer team has to deploy9Coupling direction; SNS/EventBridge fan-out; contract ownership
S22A Kinesis consumer's iterator age climbs steadily for 6 hours9Shard/consumer math; poison record; enhanced fan-out
S23An EventBridge rule silently matches nothing after a producer adds a field9Pattern brittleness; schema registry; target DLQs
S24Users upload 4 GB video files through the application tier; the app tier OOMs10Presigned URLs + multipart; getting the app out of the data path
S25CloudFront hit rate is 0%; origin cost is 20× expected10Cache key composition; Vary; cookie/query forwarding
S26Data transfer is 35% of the monthly bill10, 12Cross-AZ, NAT, egress, and where each hides
S27The service works in dev and gets AccessDenied in prod — same code, same policy text11Resource-policy vs identity-policy interaction; permission boundary; SCP
S28A leaked task role: what could an attacker actually do?11Least privilege scoped by condition keys; blast radius
S29A KMS-encrypted queue starts throttling at peak11KMS request costs and quotas; data key caching

Days 12–14 · Production and senior

#ScenarioDayThe decision
S30A downstream slows to 800ms. Within 90 seconds the whole platform is down.12Retry storms, metastable failure, circuit breakers, load shedding
S31The board asks for "multi-Region." What does it actually cost and buy?12RTO/RPO honesty; the four DR tiers; when to say no
S32A $40,000 surprise bill — five plausible causes, one true one12Cost forensics with Cost Explorer and CUR
S33"The API got slow at 14:20." You have metrics, logs and traces. Go.13The full investigation, end to end
S34A 6-year-old monolith on-prem must be on AWS in 9 months13The 7 Rs, sequencing, strangler fig, the risk register
S35A team wants to move a 400 GB Postgres table to DynamoDB13Access-pattern audit; dual-write; backfill; rollback; when to refuse
S36–S42The seven Day 14 system designs14See 14 — Capstone & System Designs

Total: 42 scenarios, at least two per teaching day.

3. The flagship scenario, expanded

S17 — E-commerce at 50,000 orders/minute is the course's recurring backbone. It is introduced on Day 8 and revisited on Days 9, 11, 12, 13 and 14, each time with the lens of that day.

The numbers (fixed, reused throughout)

MetricValue
Peak order rate50,000/min ≈ 833 orders/sec
Average order rate6,000/min ≈ 100/sec
Peak-to-average ratio8.3×
Order payload~4 KB
Payment call latencyp50 400ms, p99 2.5s, external provider
Inventory checkp50 15ms against DynamoDB
Notificationemail + push, no latency requirement
Order API latency budgetp99 < 300ms
Durability requirementAn accepted order is never lost
Correctness requirementA customer is never charged twice

Why the naive design fails

flowchart LR
    C[Client] --> API[Order API]
    API --> DB[(Orders DB)]
    API --> PAY[Payment provider<br/>p99 2.5s]
    API --> INV[Inventory]
    API --> NOTIF[Email/Push]

At 833 orders/sec with a p99 of 2.5s on the payment hop, the API's own p99 can never be 300ms; worse, the service's threads are held hostage by an external provider, so a provider slowdown becomes a full outage. Order durability is also tied to the payment provider's availability, which is the wrong coupling entirely.

The design the course builds toward

flowchart TB
    C[Client] --> CF[CloudFront] --> AG[API Gateway]
    AG --> OS[Order Service · ECS Fargate]
    OS -->|1 write + outbox, one txn| DB[(Aurora orders)]
    OS -->|2 return 202 + orderId| C
    REL[Outbox relay] -.-> DB
    REL --> EB[EventBridge · OrderPlaced]
    EB --> QP[(SQS payments)] --> PW[Payment Worker]
    EB --> QI[(SQS inventory)] --> IW[Inventory Worker]
    EB --> QN[(SQS notifications)] --> NW[Notification Worker]
    PW --> IDEM[(DynamoDB idempotency)]
    PW --> PROV[Payment provider]
    QP -.->|after N attempts| DLQP[(payments DLQ)]
    PW --> CW[CloudWatch]

The questions each day asks of it

DayLens
8Visibility timeouts, DLQ thresholds, idempotency key design, the outbox
9Why EventBridge rather than SNS here; replay after a consumer bug
11What each worker's role may do; encryption of the queues; what a leaked worker role exposes
12Scaling each worker independently; what happens when the payment provider is down for 20 minutes; retry budget; DR posture
13Tracing an order across five async hops; the alarm set that would have caught a stuck payment queue
14The full 13-step design, presented in 45 interview minutes

Cost envelope (illustrative, us-east-1, prices dated at publication)

ComponentMonthly estimate
ECS Fargate — API (peak-scaled, avg 12 tasks × 1 vCPU/2 GB)~$360
ECS Fargate — 3 worker services (avg 9 tasks total)~$270
Aurora PostgreSQL (2 × r6g.large, Multi-AZ)~$500
DynamoDB idempotency (on-demand, ~260M writes/mo)~$330
SQS (~780M requests/mo across 3 queues, batched 10×)~$32
EventBridge (260M events)~$260
ALB + API Gateway (260M requests)~$290
CloudWatch Logs (retention 14d, sampled)~$180
Data transfer + NAT~$220
Total≈ $2,440/month

The course then asks the senior question: which three of these lines would you attack first, and what would you give up? (Answers: EventBridge → SNS saves ~$230 but loses routing/replay; batching SQS harder; log sampling; and questioning whether the idempotency table needs on-demand.)

4. Scenario quality rules

  1. Numbers must be arithmetically consistent. If peak is 833/sec and the worker handles 40/sec, the design needs ≥21 workers and the text says so.
  2. The naive approach must be genuinely plausible — a straw man teaches nothing.
  3. Failure must be described as it is experienced, not as a category: "the queue depth chart goes up and to the right and does not come back down" beats "the consumer could not keep up."
  4. Every scenario names its trade-off. No scenario resolves into a free win.
  5. Prices and limits are dated and marked as verifiable against current AWS docs.