Production Scenario Plan
Every major topic carries a realistic corporate scenario with real numbers. Scenarios are how the course converts "here is a service" into "here is a decision you will have to defend."
1. Scenario format
Each scenario is written to this shape:
Context the company, the product, the constraint that makes it interesting
The numbers traffic, data size, growth, latency budget, cost ceiling
The ask what you are being asked to decide or fix
Naive approach what a competent engineer without AWS depth would do
What breaks at what scale, and how it manifests
Better design with the reasoning, not just the answer
Trade-offs what you gave up
Cost actual arithmetic, monthly
The interview how you would present this in 5 minutesNumbers are internally consistent and defensible. Where a figure is a current AWS limit or price, it is dated and flagged as subject to change.
2. The scenario catalogue
Days 1–3 · Fundamentals
| # | Scenario | Day | The decision |
|---|---|---|---|
| S1 | A startup's single EC2 box serves 500 RPS and the CTO asks "what happens if that AZ dies?" | 1 | Why Region/AZ topology is an application concern |
| S2 | A developer's laptop credentials are found in a public GitHub repo | 1 | Roles vs keys; the blast radius of an account |
| S3 | "The app can't reach the database" — six possible causes, one symptom | 2 | The connectivity debugging order |
| S4 | The monthly bill shows $1,400 of NAT Gateway charges for a service that only talks to S3 | 2 | Gateway endpoints; cost as a design signal |
| S5 | A t3.micro fleet is fine for a week, then latency triples every afternoon | 3 | Burstable credits; instance family selection |
| S6 | CloudWatch Logs is the third-largest line item on the bill | 3 | Retention, log level, sampling |
Days 4–7 · Core backend
| # | Scenario | Day | The decision |
|---|---|---|---|
| S7 | Every deploy drops ~2% of requests for 30 seconds | 4 | Deregistration delay + graceful shutdown + health check tuning |
| S8 | Black Friday: 8× traffic arriving over 4 minutes; autoscaling "works" but users see errors | 4 | Scale-out lag, pre-scaling, load shedding |
| S9 | A Java Lambda behind API Gateway shows a 3.2s p99 and a 180ms p50 | 5 | Cold starts; provisioned concurrency vs Fargate; the cost crossover |
| S10 | A Lambda triggered by S3 writes to the same bucket. The bill is $11,000 in 6 hours. | 5, 12 | Recursive invocation; the prefix/suffix filter; the spend guardrail |
| S11 | 200 Fargate tasks × 20-connection pools meet a database that allows 500 connections | 6 | Pool math; RDS Proxy; the deploy that doubles both |
| S12 | A Multi-AZ failover takes 40 seconds of application errors; the team expected zero | 6 | What Multi-AZ actually promises; JVM DNS TTL; retry on reconnect |
| S13 | Reporting queries on a read replica return orders that "don't exist yet" | 6 | Replica lag; read-your-writes; routing reads deliberately |
| S14 | A DynamoDB table is throttled at 40% of provisioned capacity | 7 | Hot partition; key design; adaptive capacity's limits |
| S15 | A cache node restarts at 09:00 and the database falls over at 09:00:03 | 7 | Stampede; TTL jitter; request coalescing; warmup |
| S16 | A "simple" Scan in a nightly job costs $900/month | 7 | Query vs scan; GSI design; the access-pattern audit |
Days 8–11 · Advanced backend
| # | Scenario | Day | The decision |
|---|---|---|---|
| S17 | E-commerce: 50,000 orders/minute. Payments, inventory and notification must all happen. | 8 | The canonical async decomposition — full treatment |
| S18 | A customer is charged twice. The logs show one order and two payment messages. | 8 | At-least-once delivery; idempotency keys; conditional writes |
| S19 | A worker takes 90 seconds; visibility timeout is 30 seconds. The queue never drains and the DLQ fills. | 8 | Visibility timeout math; heartbeating; maxReceiveCount |
| S20 | A FIFO queue "isn't scaling" — throughput is stuck at ~10 msg/s | 8 | Message group ID as the parallelism unit |
| S21 | A new consumer is added and the producer team has to deploy | 9 | Coupling direction; SNS/EventBridge fan-out; contract ownership |
| S22 | A Kinesis consumer's iterator age climbs steadily for 6 hours | 9 | Shard/consumer math; poison record; enhanced fan-out |
| S23 | An EventBridge rule silently matches nothing after a producer adds a field | 9 | Pattern brittleness; schema registry; target DLQs |
| S24 | Users upload 4 GB video files through the application tier; the app tier OOMs | 10 | Presigned URLs + multipart; getting the app out of the data path |
| S25 | CloudFront hit rate is 0%; origin cost is 20× expected | 10 | Cache key composition; Vary; cookie/query forwarding |
| S26 | Data transfer is 35% of the monthly bill | 10, 12 | Cross-AZ, NAT, egress, and where each hides |
| S27 | The service works in dev and gets AccessDenied in prod — same code, same policy text | 11 | Resource-policy vs identity-policy interaction; permission boundary; SCP |
| S28 | A leaked task role: what could an attacker actually do? | 11 | Least privilege scoped by condition keys; blast radius |
| S29 | A KMS-encrypted queue starts throttling at peak | 11 | KMS request costs and quotas; data key caching |
Days 12–14 · Production and senior
| # | Scenario | Day | The decision |
|---|---|---|---|
| S30 | A downstream slows to 800ms. Within 90 seconds the whole platform is down. | 12 | Retry storms, metastable failure, circuit breakers, load shedding |
| S31 | The board asks for "multi-Region." What does it actually cost and buy? | 12 | RTO/RPO honesty; the four DR tiers; when to say no |
| S32 | A $40,000 surprise bill — five plausible causes, one true one | 12 | Cost forensics with Cost Explorer and CUR |
| S33 | "The API got slow at 14:20." You have metrics, logs and traces. Go. | 13 | The full investigation, end to end |
| S34 | A 6-year-old monolith on-prem must be on AWS in 9 months | 13 | The 7 Rs, sequencing, strangler fig, the risk register |
| S35 | A team wants to move a 400 GB Postgres table to DynamoDB | 13 | Access-pattern audit; dual-write; backfill; rollback; when to refuse |
| S36–S42 | The seven Day 14 system designs | 14 | See 14 — Capstone & System Designs |
Total: 42 scenarios, at least two per teaching day.
3. The flagship scenario, expanded
S17 — E-commerce at 50,000 orders/minute is the course's recurring backbone. It is introduced on Day 8 and revisited on Days 9, 11, 12, 13 and 14, each time with the lens of that day.
The numbers (fixed, reused throughout)
| Metric | Value |
|---|---|
| Peak order rate | 50,000/min ≈ 833 orders/sec |
| Average order rate | 6,000/min ≈ 100/sec |
| Peak-to-average ratio | 8.3× |
| Order payload | ~4 KB |
| Payment call latency | p50 400ms, p99 2.5s, external provider |
| Inventory check | p50 15ms against DynamoDB |
| Notification | email + push, no latency requirement |
| Order API latency budget | p99 < 300ms |
| Durability requirement | An accepted order is never lost |
| Correctness requirement | A customer is never charged twice |
Why the naive design fails
flowchart LR
C[Client] --> API[Order API]
API --> DB[(Orders DB)]
API --> PAY[Payment provider<br/>p99 2.5s]
API --> INV[Inventory]
API --> NOTIF[Email/Push]At 833 orders/sec with a p99 of 2.5s on the payment hop, the API's own p99 can never be 300ms; worse, the service's threads are held hostage by an external provider, so a provider slowdown becomes a full outage. Order durability is also tied to the payment provider's availability, which is the wrong coupling entirely.
The design the course builds toward
flowchart TB
C[Client] --> CF[CloudFront] --> AG[API Gateway]
AG --> OS[Order Service · ECS Fargate]
OS -->|1 write + outbox, one txn| DB[(Aurora orders)]
OS -->|2 return 202 + orderId| C
REL[Outbox relay] -.-> DB
REL --> EB[EventBridge · OrderPlaced]
EB --> QP[(SQS payments)] --> PW[Payment Worker]
EB --> QI[(SQS inventory)] --> IW[Inventory Worker]
EB --> QN[(SQS notifications)] --> NW[Notification Worker]
PW --> IDEM[(DynamoDB idempotency)]
PW --> PROV[Payment provider]
QP -.->|after N attempts| DLQP[(payments DLQ)]
PW --> CW[CloudWatch]The questions each day asks of it
| Day | Lens |
|---|---|
| 8 | Visibility timeouts, DLQ thresholds, idempotency key design, the outbox |
| 9 | Why EventBridge rather than SNS here; replay after a consumer bug |
| 11 | What each worker's role may do; encryption of the queues; what a leaked worker role exposes |
| 12 | Scaling each worker independently; what happens when the payment provider is down for 20 minutes; retry budget; DR posture |
| 13 | Tracing an order across five async hops; the alarm set that would have caught a stuck payment queue |
| 14 | The full 13-step design, presented in 45 interview minutes |
Cost envelope (illustrative, us-east-1, prices dated at publication)
| Component | Monthly estimate |
|---|---|
| ECS Fargate — API (peak-scaled, avg 12 tasks × 1 vCPU/2 GB) | ~$360 |
| ECS Fargate — 3 worker services (avg 9 tasks total) | ~$270 |
| Aurora PostgreSQL (2 × r6g.large, Multi-AZ) | ~$500 |
| DynamoDB idempotency (on-demand, ~260M writes/mo) | ~$330 |
| SQS (~780M requests/mo across 3 queues, batched 10×) | ~$32 |
| EventBridge (260M events) | ~$260 |
| ALB + API Gateway (260M requests) | ~$290 |
| CloudWatch Logs (retention 14d, sampled) | ~$180 |
| Data transfer + NAT | ~$220 |
| Total | ≈ $2,440/month |
The course then asks the senior question: which three of these lines would you attack first, and what would you give up? (Answers: EventBridge → SNS saves ~$230 but loses routing/replay; batching SQS harder; log sampling; and questioning whether the idempotency table needs on-demand.)
4. Scenario quality rules
- Numbers must be arithmetically consistent. If peak is 833/sec and the worker handles 40/sec, the design needs ≥21 workers and the text says so.
- The naive approach must be genuinely plausible — a straw man teaches nothing.
- Failure must be described as it is experienced, not as a category: "the queue depth chart goes up and to the right and does not come back down" beats "the consumer could not keep up."
- Every scenario names its trade-off. No scenario resolves into a free win.
- Prices and limits are dated and marked as verifiable against current AWS docs.