Learning/AWS Backend Developer/11 — Production Troubleshooting & Failure Plan

Production Troubleshooting & Failure Plan

10 troubleshooting scenarios (symptom → root cause, with the investigation shown) and 10 architecture failure scenarios (design-level, "what breaks and what should have been there").

1. The investigation method

Taught on Day 13 and then applied identically in all ten scenarios:

flowchart TD
    S[Symptom + when it started] --> B{Is it everything<br/>or one thing?}
    B -->|Everything| G[Look at shared dependencies:<br/>DB, cache, auth, network, deploy]
    B -->|One thing| N[Narrow by dimension:<br/>endpoint, AZ, task, customer, key]
    G --> T[Correlate with the change log:<br/>deploy, config, traffic, AWS event]
    N --> T
    T --> M[Pick the ONE signal that<br/>discriminates between hypotheses]
    M --> C{Confirmed?}
    C -->|No| M
    C -->|Yes| F[Fix · then ask what<br/>should have alarmed first]

The rule the course repeats: form two hypotheses, then choose the single metric or query that tells them apart. Reading dashboards top-to-bottom is not investigation.

The signal ladder

QuestionSignal
Is it broken, and since when?Metrics (RED: rate, errors, duration)
Where in the path?Traces / service map
Why, exactly?Logs, filtered by trace ID
Is it AWS or us?Service Health Dashboard, service-side metrics, CloudTrail
Did someone change something?CloudTrail, deploy history, Config timeline

2. Ten troubleshooting scenarios (Day 13)

Each is written as: symptom → hypotheses → discriminating signal (with the exact query/command) → root cause → fix → what alarm should have caught it.

#SymptomDiscriminating signalRoot cause taught
T1API p99 jumps from 120ms to 2.1s at 14:20; p50 unchangedX-Ray service map, then per-subsegment latencyA GSI query fell back to a scan after a deploy changed a filter; only the long tail of large customers hit it
T2ALB returns 504s at ~3% under load; targets show healthyTarget response time vs ALB idle timeout; target 5xx count = 0App is slower than the ALB idle timeout; no server-side timeout so requests pile up. 502 vs 504 distinction is the key insight
T3SQS ApproximateAgeOfOldestMessage climbing for 40 minMessages-received vs deleted rate; consumer error logVisibility timeout < processing time → infinite reprocessing; work is being done repeatedly and never acknowledged
T4DynamoDB throttling at 38% of provisioned capacityThrottledRequests by partition via CloudWatch Contributor InsightsHot partition — partition key is status, four values, one dominant
T5Connection is not available, request timed out after 30000msHikariCP pending-threads metric vs RDS DatabaseConnectionsPool exhaustion from a slow query plus a deploy that doubled task count; DB itself is idle
T6Lambda Throttles spiking; some customers charged twiceConcurrency metric vs account limit; async retry behaviourReserved concurrency too low; async invocations retried twice by the service while the handler was non-idempotent
T7502 immediately after every deploy, clears after 2 minTarget group deregistration events + container exit codesNo graceful shutdown; tasks killed mid-request while still registered
T8CloudFront hit rate 4%; origin bill 20× budgetCloudFront cache statistics + X-Cache header on a sample requestCache key includes a per-user cookie and a cache-busting query param forwarded by policy
T9Data transfer cost triples with no traffic changeCost Explorer by usage type, then VPC Flow LogsService moved to a second AZ; chatty service-to-service calls now cross AZ on every hop
T10AccessDenied in prod only, identical code and identity policyAssumeRole session in CloudTrail + the full denial contextThe bucket's resource policy lacks the prod account; SCP also denies outside an allowed Region

Each scenario includes the actual CLI/Logs Insights command used, for example:

aws logs start-query \
  --log-group-name /ecs/order-service \
  --start-time $START --end-time $END \
  --query-string 'fields @timestamp, traceId, durationMs
                  | filter durationMs > 1000
                  | stats count() by bin(1m)'

3. Ten architecture failure scenarios (Days 4–12, one per relevant day)

These are not "something broke, debug it" — they are "this design has a failure mode; find it before production does."

#ArchitectureWhat failsWhat should have been there
A1Single-AZ ECS service + single-AZ RDSAZ loss = total outageMulti-AZ subnets, Multi-AZ DB, and a test that proves failover
A2Synchronous chain: API → Service B → Service C → external providerB's availability = product of all fourAsync boundary at the external hop; timeouts; circuit breaker
A3Worker with no DLQ and infinite retriesOne poison message blocks the queue foreverDLQ + maxReceiveCount + a DLQ-depth alarm
A4Non-idempotent consumer on a standard queueDuplicate side effects under normal operationIdempotency key with conditional write
A5"Write to DB, then publish to SQS"Crash between the two = silent data divergenceTransactional outbox (or DynamoDB Streams / CDC)
A6Read-heavy service reading from a replica, writes to primaryUsers don't see their own writesRead-your-writes routing, or session-consistent reads
A7Cache with uniform TTL across 2M keysSynchronized expiry → origin stampedeTTL jitter, request coalescing, stale-while-revalidate
A8Autoscaling on CPU for an I/O-bound serviceNever scales; queues and latency grow while CPU sits at 30%Scale on RPS, concurrency, or backlog-per-worker
A9Aggressive client retries with no jitter or budgetRetry storm converts a brownout into an outage that survives the original cause (metastable failure)Exponential backoff + jitter, retry budget, load shedding, circuit breaker
A10Multi-AZ described to the board as "disaster recovery"Region-level event or a logical failure (bad deploy, dropped table) destroys both AZs' copiesExplicit RTO/RPO, cross-Region backups, restore drills, and honest language

4. Failure-first questions (applied everywhere)

Every architecture the course presents answers these seven, explicitly:

1. What can fail?
2. What happens when it fails — from the user's point of view?
3. How does recovery happen — automatic or human?
4. How quickly?
5. Can the operation happen twice?
6. Can data be lost?
7. Can data be duplicated?

The course treats questions 5–7 as the ones that separate a 5-YOE answer from a 9-YOE answer, because they are about correctness, not availability.

5. Alarms worth having (the Day 13 deliverable)

The course ships an opinionated starter alarm set, with justification for each and an explicit "do not alarm on this" list.

LayerAlarmThreshold philosophy
API5xx rate, p99 latencySymptom-based, user-visible, pages
ALB/TargetUnhealthy host count > 0 for 2 minStructural
ECSRunning vs desired count mismatchCatches capacity and crash loops
SQSApproximateAgeOfOldestMessageThe single best queue alarm — depth alone lies
SQSDLQ depth > 0Always. A message in a DLQ is a bug.
LambdaThrottles > 0, error rate, duration p99 vs timeoutDuration approaching timeout is a leading indicator
DynamoDBThrottledRequests, SystemErrorsLeading indicator of key design problems
RDSConnections vs max, replica lag, CPU, free storageStorage is the one that takes you down silently
CacheEvictions, hit rate dropHit rate falling is a cost and latency event
CostBudget forecast > thresholdYes, cost is an alarm

Do not alarm on: raw CPU (unless it's the actual constraint), individual error log lines, queue depth without age, or any metric nobody will act on at 3am. The course states the rule directly: an alarm that does not change what a human does is a notification, and notifications belong in a dashboard.