Production Troubleshooting & Failure Plan
10 troubleshooting scenarios (symptom → root cause, with the investigation shown) and 10 architecture failure scenarios (design-level, "what breaks and what should have been there").
1. The investigation method
Taught on Day 13 and then applied identically in all ten scenarios:
flowchart TD
S[Symptom + when it started] --> B{Is it everything<br/>or one thing?}
B -->|Everything| G[Look at shared dependencies:<br/>DB, cache, auth, network, deploy]
B -->|One thing| N[Narrow by dimension:<br/>endpoint, AZ, task, customer, key]
G --> T[Correlate with the change log:<br/>deploy, config, traffic, AWS event]
N --> T
T --> M[Pick the ONE signal that<br/>discriminates between hypotheses]
M --> C{Confirmed?}
C -->|No| M
C -->|Yes| F[Fix · then ask what<br/>should have alarmed first]The rule the course repeats: form two hypotheses, then choose the single metric or query that tells them apart. Reading dashboards top-to-bottom is not investigation.
The signal ladder
| Question | Signal |
|---|---|
| Is it broken, and since when? | Metrics (RED: rate, errors, duration) |
| Where in the path? | Traces / service map |
| Why, exactly? | Logs, filtered by trace ID |
| Is it AWS or us? | Service Health Dashboard, service-side metrics, CloudTrail |
| Did someone change something? | CloudTrail, deploy history, Config timeline |
2. Ten troubleshooting scenarios (Day 13)
Each is written as: symptom → hypotheses → discriminating signal (with the exact query/command) → root cause → fix → what alarm should have caught it.
| # | Symptom | Discriminating signal | Root cause taught |
|---|---|---|---|
| T1 | API p99 jumps from 120ms to 2.1s at 14:20; p50 unchanged | X-Ray service map, then per-subsegment latency | A GSI query fell back to a scan after a deploy changed a filter; only the long tail of large customers hit it |
| T2 | ALB returns 504s at ~3% under load; targets show healthy | Target response time vs ALB idle timeout; target 5xx count = 0 | App is slower than the ALB idle timeout; no server-side timeout so requests pile up. 502 vs 504 distinction is the key insight |
| T3 | SQS ApproximateAgeOfOldestMessage climbing for 40 min | Messages-received vs deleted rate; consumer error log | Visibility timeout < processing time → infinite reprocessing; work is being done repeatedly and never acknowledged |
| T4 | DynamoDB throttling at 38% of provisioned capacity | ThrottledRequests by partition via CloudWatch Contributor Insights | Hot partition — partition key is status, four values, one dominant |
| T5 | Connection is not available, request timed out after 30000ms | HikariCP pending-threads metric vs RDS DatabaseConnections | Pool exhaustion from a slow query plus a deploy that doubled task count; DB itself is idle |
| T6 | Lambda Throttles spiking; some customers charged twice | Concurrency metric vs account limit; async retry behaviour | Reserved concurrency too low; async invocations retried twice by the service while the handler was non-idempotent |
| T7 | 502 immediately after every deploy, clears after 2 min | Target group deregistration events + container exit codes | No graceful shutdown; tasks killed mid-request while still registered |
| T8 | CloudFront hit rate 4%; origin bill 20× budget | CloudFront cache statistics + X-Cache header on a sample request | Cache key includes a per-user cookie and a cache-busting query param forwarded by policy |
| T9 | Data transfer cost triples with no traffic change | Cost Explorer by usage type, then VPC Flow Logs | Service moved to a second AZ; chatty service-to-service calls now cross AZ on every hop |
| T10 | AccessDenied in prod only, identical code and identity policy | AssumeRole session in CloudTrail + the full denial context | The bucket's resource policy lacks the prod account; SCP also denies outside an allowed Region |
Each scenario includes the actual CLI/Logs Insights command used, for example:
aws logs start-query \
--log-group-name /ecs/order-service \
--start-time $START --end-time $END \
--query-string 'fields @timestamp, traceId, durationMs
| filter durationMs > 1000
| stats count() by bin(1m)'3. Ten architecture failure scenarios (Days 4–12, one per relevant day)
These are not "something broke, debug it" — they are "this design has a failure mode; find it before production does."
| # | Architecture | What fails | What should have been there |
|---|---|---|---|
| A1 | Single-AZ ECS service + single-AZ RDS | AZ loss = total outage | Multi-AZ subnets, Multi-AZ DB, and a test that proves failover |
| A2 | Synchronous chain: API → Service B → Service C → external provider | B's availability = product of all four | Async boundary at the external hop; timeouts; circuit breaker |
| A3 | Worker with no DLQ and infinite retries | One poison message blocks the queue forever | DLQ + maxReceiveCount + a DLQ-depth alarm |
| A4 | Non-idempotent consumer on a standard queue | Duplicate side effects under normal operation | Idempotency key with conditional write |
| A5 | "Write to DB, then publish to SQS" | Crash between the two = silent data divergence | Transactional outbox (or DynamoDB Streams / CDC) |
| A6 | Read-heavy service reading from a replica, writes to primary | Users don't see their own writes | Read-your-writes routing, or session-consistent reads |
| A7 | Cache with uniform TTL across 2M keys | Synchronized expiry → origin stampede | TTL jitter, request coalescing, stale-while-revalidate |
| A8 | Autoscaling on CPU for an I/O-bound service | Never scales; queues and latency grow while CPU sits at 30% | Scale on RPS, concurrency, or backlog-per-worker |
| A9 | Aggressive client retries with no jitter or budget | Retry storm converts a brownout into an outage that survives the original cause (metastable failure) | Exponential backoff + jitter, retry budget, load shedding, circuit breaker |
| A10 | Multi-AZ described to the board as "disaster recovery" | Region-level event or a logical failure (bad deploy, dropped table) destroys both AZs' copies | Explicit RTO/RPO, cross-Region backups, restore drills, and honest language |
4. Failure-first questions (applied everywhere)
Every architecture the course presents answers these seven, explicitly:
1. What can fail?
2. What happens when it fails — from the user's point of view?
3. How does recovery happen — automatic or human?
4. How quickly?
5. Can the operation happen twice?
6. Can data be lost?
7. Can data be duplicated?The course treats questions 5–7 as the ones that separate a 5-YOE answer from a 9-YOE answer, because they are about correctness, not availability.
5. Alarms worth having (the Day 13 deliverable)
The course ships an opinionated starter alarm set, with justification for each and an explicit "do not alarm on this" list.
| Layer | Alarm | Threshold philosophy |
|---|---|---|
| API | 5xx rate, p99 latency | Symptom-based, user-visible, pages |
| ALB/Target | Unhealthy host count > 0 for 2 min | Structural |
| ECS | Running vs desired count mismatch | Catches capacity and crash loops |
| SQS | ApproximateAgeOfOldestMessage | The single best queue alarm — depth alone lies |
| SQS | DLQ depth > 0 | Always. A message in a DLQ is a bug. |
| Lambda | Throttles > 0, error rate, duration p99 vs timeout | Duration approaching timeout is a leading indicator |
| DynamoDB | ThrottledRequests, SystemErrors | Leading indicator of key design problems |
| RDS | Connections vs max, replica lag, CPU, free storage | Storage is the one that takes you down silently |
| Cache | Evictions, hit rate drop | Hit rate falling is a cost and latency event |
| Cost | Budget forecast > threshold | Yes, cost is an alarm |
Do not alarm on: raw CPU (unless it's the actual constraint), individual error log lines, queue depth without age, or any metric nobody will act on at 3am. The course states the rule directly: an alarm that does not change what a human does is a notification, and notifications belong in a dashboard.