Learning/AWS Backend Developer/04 — The 14-Day Curriculum

The 14-Day Curriculum

Every day is time-boxed block by block. Totals are held to 3h 00m – 4h 00m. The 14-day total is 48h 35m, inside the 42–56 hour budget.

The arc

flowchart LR
    subgraph P1["Days 1-3 · Fundamentals"]
        D1[1 Cloud, Account, IAM] --> D2[2 Networking/VPC] --> D3[3 Compute, Storage, Observability]
    end
    subgraph P2["Days 4-7 · Core Backend AWS"]
        D4[4 Containers, ALB, Scaling] --> D5[5 Serverless + API Gateway] --> D6[6 Relational Data] --> D7[7 NoSQL + Caching]
    end
    subgraph P3["Days 8-11 · Advanced Backend AWS"]
        D8[8 Queues + Idempotency] --> D9[9 Events + Streams] --> D10[10 Storage + Edge] --> D11[11 Security + IAM Depth]
    end
    subgraph P4["Days 12-14 · Production & Senior"]
        D12[12 Scalability, Reliability, Cost] --> D13[13 Observability + Debugging + Migration] --> D14[14 System Design + Interview]
    end
    D3 --> D4
    D7 --> D8
    D11 --> D12

Phase outcomes

AfterThe learner can say
Days 1–3"I understand what AWS is and how my backend application communicates with AWS."
Days 4–7"I can build backend applications using AWS services."
Days 8–11"I understand how AWS-backed applications behave in production."
Days 12–13"I can reason about scalability, reliability, security, cost and failure."
Day 14"Given a backend problem I can select services, design the architecture, explain trade-offs, identify failure modes, troubleshoot it, and discuss it at a senior interview."

PHASE 1 — DAYS 1–3 · AWS FUNDAMENTALS

Day 1 — Cloud, the AWS Account, and Identity

Outcome: you can explain what AWS is, reach it three different ways (Console, CLI, SDK), and you understand that every single AWS call is an authenticated, authorized API call — including the ones the Console makes for you.

Prerequisites: none beyond the baseline.

BlockMinutesContent
1.120Why cloud exists. The capacity/procurement problem a company with its own servers actually has. IaaS / PaaS / SaaS as a responsibility boundary, not a marketing tier. The shared responsibility model — and what it means for you, the developer.
1.225AWS global infrastructure. Region → Availability Zone → data centre. Edge locations. Global vs regional vs zonal services, and why s3.amazonaws.com and sqs.eu-west-1.amazonaws.com look different. What an AZ failure looks like vs a Region failure. Choosing a Region (latency, price, data residency, service availability).
1.320The AWS account. The account as the primary blast radius and billing boundary. Root user: what it can do, why you lock it with MFA immediately, and why you never use it again. Multi-account patterns named only (deferred to Day 11). Lab 0: secure the root user, create a billing alarm, create an admin IAM user/role.
1.445IAM fundamentals. Principal, action, resource, condition. Users, groups, roles, policies. Identity-based vs resource-based policies (introduced, deepened Day 11). The crucial idea: a role is not a user — it is a set of permissions anything trustworthy can temporarily wear. Reading a policy document line by line.
1.530How an AWS API call actually works. The request lifecycle: SDK builds request → credential provider chain resolves credentials → SigV4 signing → HTTPS to a regional endpoint → IAM authorization → service executes → response. Why this single diagram explains 80% of "access denied" tickets.
1.635Interacting with AWS. Console vs CLI v2 vs SDK v2 — same API underneath. aws configure, named profiles, ~/.aws/config. First Java + SDK v2 program: StsClient.getCallerIdentity() — "who does AWS think I am?"
1.720Lab 1 — S3 upload/download (Console, then CLI, then Java SDK v2, the same object three ways).
1.815Common mistakes · interview questions · day-end revision · mini assignment.
3h 30m

Optional Deep Dive: SigV4 signature computation step by step; AWS Organizations and SCPs.

Day 2 — Networking for Backend Developers

Outcome: you can draw your application's VPC, and when a connection fails you can say which of the six things between the client and the database rejected it.

Prerequisites: Day 1.

BlockMinutesContent
2.125From laptop to application. The pre-VPC mental model. Why "just give the server a public IP" stops working. Public vs private address space, restated in AWS terms.
2.230CIDR and subnets in practice. Sizing a VPC (10.0.0.0/16), carving subnets per AZ, the 5 reserved addresses per subnet, and why you leave room to grow. The public/private/isolated three-tier layout and why three.
2.335Routing. Route tables as the thing that actually makes a subnet "public". Internet Gateway. NAT Gateway — what it does, why it costs real money, and the per-AZ-vs-shared decision. Egress-only IGW (IPv6) mentioned.
2.435Security Groups vs NACLs. Stateful vs stateless, the single most-tested distinction in this area. SG-referencing-SG as the idiomatic pattern (app-sg allows alb-sg on 8080; db-sg allows app-sg on 5432). Default-deny reasoning.
2.525Reaching AWS services privately. Why a Lambda in a private subnet calling S3 is a surprisingly interesting question. Gateway endpoints (S3, DynamoDB — free) vs Interface endpoints / PrivateLink (ENI, hourly + data cost). The NAT-Gateway-bill-reduction angle.
2.620DNS in AWS. Route 53 public vs private hosted zones, VPC DNS resolution, record types you'll actually create, alias records vs CNAME.
2.730Lab 2 — build the three-tier VPC, launch a host in the private subnet, prove it can reach the Internet through NAT and cannot be reached from it.
2.820The connectivity debugging checklist (route table → NACL → SG → listener → health check → application) · failure scenarios · interview questions · revision · mini assignment.
3h 40m

Optional Deep Dive: VPC peering, Transit Gateway, Direct Connect, IPv6-only subnets, VPC Flow Logs analysis.

Day 3 — Compute, Storage and Seeing What Happened

Outcome: your Spring Boot application is running on an EC2 instance in a private subnet, reading and writing S3 through an instance role, with its logs and metrics in CloudWatch.

Prerequisites: Days 1–2.

BlockMinutesContent
3.130EC2 fundamentals. Instance = virtual machine; AMI = the image; instance families and what the letters mean (c/m/r/t, and burstable credits as a production trap); EBS volume types (gp3 vs io2) and the "your disk has IOPS limits" surprise; user data; instance metadata service (IMDSv2) — and how it delivers role credentials.
3.225How EC2 gets credentials. The credential provider chain revisited, now concrete: instance profile → IMDS → temporary STS credentials → rotated automatically. Why this is the moment hardcoded access keys stop being necessary forever.
3.340S3 fundamentals. Bucket, key, object, metadata. Flat namespace vs the "folders" the Console shows you. Durability vs availability. Strong read-after-write consistency (and what that replaced). Storage classes introduced (deep-dived Day 10). Bucket policies, Block Public Access, and the classic public-bucket incident.
3.435Lab 3 — deploy Spring Boot to EC2. Build the jar, ship it, run it as a systemd service, front it with nothing yet, reach it from a bastion/SSM Session Manager. Lab 4 — the app reads/writes S3 via its instance role (zero credentials in code or config).
3.535CloudWatch basics. The three pillars as AWS implements them: Logs (groups, streams, retention — and the retention default that costs you money), Metrics (namespace, dimension, statistic, period), Alarms (threshold, evaluation periods, missing-data treatment). The CloudWatch agent vs the SDK's log appender. Shipping structured JSON logs from Spring Boot.
3.620Scaling, previewed. Vertical vs horizontal; why "make the instance bigger" ends; what makes an app safe to run in more than one copy. Sets up Day 4.
3.725Cost of what you just built · cleanup · interview questions · revision · Phase 1 checkpoint quiz.
3h 30m

End of Phase 1 checkpoint: "I understand what AWS is and how my backend application communicates with AWS."

PHASE 2 — DAYS 4–7 · CORE BACKEND AWS

Day 4 — Running Services: Scaling Groups, Load Balancers, Containers

Outcome: the same Spring Boot app runs as multiple containers on ECS Fargate behind an ALB, scaling on a metric, with zero-downtime deploys.

Prerequisites: Days 1–3.

BlockMinutesContent
4.130Auto Scaling Groups. Launch templates, desired/min/max, health checks (EC2 vs ELB), scaling policies (target tracking vs step vs scheduled), cooldowns, and why scaling on CPU is usually the wrong signal for a request-driven service.
4.240Load balancing. ALB anatomy: listener → rule → target group → target. Health checks (the parameters that cause 80% of "my deploy never goes healthy"). Connection draining/deregistration delay. Sticky sessions and why needing them is a design smell. ALB vs NLB vs GWLB — with NLB's use cases stated precisely. Internal request lifecycle of an ALB request, including where TLS terminates and what X-Forwarded-For does.
4.330Containers, briefly. Image, layer, registry, ECR. Building a Spring Boot image well (layered jars, JVM flags for containers, MALLOC_ARENA_MAX, why the JVM used to ignore cgroup limits). Just enough Docker, no more.
4.445ECS. Cluster, task definition, task, service. Fargate vs EC2 launch type. Task role vs execution role — the distinction people get wrong constantly. Service discovery options. Rolling deploys, circuit breaker, blue/green with CodeDeploy named. Task-level autoscaling.
4.525EC2 vs ECS vs Fargate vs EKS vs Lambda — the first full decision matrix, on 10 axes. Deliberately placed before Lambda is taught, so Day 5 answers a question you already have.
4.630Lab 5 — containerize and deploy to ECS Fargate behind an ALB, with autoscaling and a rolling deploy watched in real time.
4.720Failure scenarios (task thrashing, health-check flapping, capacity errors) · cost · interview questions · revision · mini assignment.
3h 40m

Optional Deep Dive: EKS in depth, ECS on EC2 capacity providers, Spot strategies, App Runner.

Day 5 — Serverless Compute and the API Layer

Outcome: you can explain a Lambda cold start to an interviewer in terms of what AWS actually does, and you can put an authenticated, throttled API in front of both Lambda and ECS.

Prerequisites: Days 1–4.

BlockMinutesContent
5.140Lambda internals. Not "Lambda runs your function" but: invoke → service → execution environment (Firecracker microVM) → cold start (download code, start runtime, run static init) vs warm reuse → handler → response → freeze → reuse or reap. Init-phase vs invoke-phase billing. Where the 10 GB /tmp and the 15-minute ceiling come from. Memory as the CPU dial.
5.230Lambda operationally. Concurrency: account limit, reserved, provisioned. Throttling and 429s. Sync vs async vs poll-based invocation and the completely different retry behaviour of each. Event source mappings. Lambda in a VPC (and how Hyperplane ENIs removed the old cold-start penalty). Java on Lambda specifically: cold start reality, SnapStart, and when the answer is "use Fargate instead".
5.340API Gateway. REST API vs HTTP API vs WebSocket API — capability and cost trade-off table. Stages, routes, integrations (proxy vs non-proxy), request/response mapping. Authorizers: IAM, Cognito, Lambda authorizer, JWT. Usage plans, API keys, throttling and burst. Caching. Request lifecycle diagram end to end.
5.425API Gateway vs ALB. When a load balancer is the right front door and when it isn't; cost crossover; the "we put API Gateway in front of everything" anti-pattern.
5.535Lab 6 — Lambda from Java (handler, SDK v2 client reuse across invocations, env config, cold-start measurement). Lab 7 — API Gateway HTTP API → Lambda, with a JWT authorizer and throttling proven with a load generator.
5.620Failure scenarios (throttling storms, retry amplification into a downstream, poison async events) · cost modelling: a request/month break-even between Lambda and Fargate · interview questions · revision.
3h 10m

Optional Deep Dive: Lambda extensions and layers, Step Functions Express, SnapStart internals, GraalVM native images.

Day 6 — Relational Data in Production

Outcome: your Spring Boot app talks to Aurora PostgreSQL over a correctly-sized pool, with credentials it never sees in config, and you can describe exactly what happens to in-flight transactions during a failover.

Prerequisites: Days 1–4.

BlockMinutesContent
6.130RDS fundamentals. What "managed" removes and what it doesn't. Engines, instance classes, storage autoscaling, parameter groups, option groups, maintenance windows, minor/major version upgrades. Backups: automated backups + retention vs manual snapshots; PITR and what the recovery window really is.
6.235Multi-AZ and read replicas — the distinction that matters. Multi-AZ instance vs Multi-AZ cluster. Synchronous standby vs asynchronous replica. What failover does to your connections (they break — plan for it), DNS-based endpoint switching and client DNS caching (the JVM networkaddress.cache.ttl trap). Replica lag and read-your-own-writes.
6.330Aurora. The decoupled storage layer and why that changes replica cost, failover speed and read scaling. Writer/reader endpoints, custom endpoints. Aurora Serverless v2 and when it fits. Global Database, named for Day 12 DR. Aurora vs RDS decision criteria.
6.435The application side. HikariCP sizing (and why maxPoolSize = 100 per task is how you take down a database with a deploy). Connection lifetime vs failover. RDS Proxy: the problem it solves (connection storms from Lambda/Fargate scale-out), what it costs, when it's unnecessary. Timeouts at every layer. Schema migrations with Flyway/Liquibase against a live service, expand/contract.
6.525Secrets. Secrets Manager vs SSM Parameter Store (SecureString) — cost, rotation, size, cross-account. IAM database authentication as a third option. Spring Boot wiring so the password exists only in memory. Never a credential in application.yml, an env var baked into an image, or a Git repo.
6.630Lab 8 — Spring Boot + Aurora PostgreSQL in private subnets, secret from Secrets Manager, Flyway migration, then trigger a failover and observe what the app does.
6.720Failure scenarios · performance (slow query log, Performance Insights) · cost (the instance is on 24/7 — the biggest line item so far) · interview questions · revision.
3h 25m

Optional Deep Dive: DMS, logical replication, Babelfish, Aurora Limitless.

Day 7 — DynamoDB and Caching

Outcome: you can model a real access pattern in DynamoDB without reaching for a relational instinct, predict what it costs, and know which reads belong in Redis instead.

Prerequisites: Days 1–4, 6 (the relational contrast is the teaching device).

BlockMinutesContent
7.135DynamoDB from first principles. Start with the problem: a table that must serve predictable single-digit-millisecond reads at any size. Hashing the partition key → partitions; the sort key → ordered range within a partition. Items, attributes, the 400 KB item limit. Why there are no JOINs and why that is the point, not a limitation.
7.240Modeling. Access-pattern-first design: list the queries before the schema. Composite keys, begins_with, sparse indexes. GSI (own capacity, eventually consistent, can re-partition your data) vs LSI (same partition, strongly consistent, 10 GB item-collection limit, must be created with the table). Single-table design — taught honestly, including when not to do it. Worked model: orders by customer, by status, by date.
7.330Operational DynamoDB. On-demand vs provisioned + autoscaling, RCU/WCU arithmetic you can do in your head. Hot partitions and adaptive capacity. Strong vs eventual reads. Conditional writes and optimistic locking. TransactWriteItems — what it really guarantees and what it costs (2×). BatchGet/BatchWrite and partial failures. TTL. Streams (→ Day 9). PITR and backups. DAX, briefly.
7.430ElastiCache & caching strategy. Redis vs Memcached vs Valkey; cluster mode, replicas, failover. The strategies: cache-aside (default), write-through, write-behind, read-through. TTL and jitter. The three failure modes with names: stampede/dogpile, hot key, thundering herd on restart. Cache invalidation honestly discussed. Redis as more than a cache (rate limiter, distributed lock — with the lock caveats stated).
7.525Decision matrices: RDS vs DynamoDB; DynamoDB vs Redis; when "both" is correct. Includes the migration question (RDBMS → DynamoDB) previewed for Day 13.
7.630Lab 9 — Spring Boot + DynamoDB Enhanced Client (bean mapping, a GSI query, a conditional write that prevents a double-update). Lab 10 — add ElastiCache Redis cache-aside and measure the hit rate and latency change.
7.720Failure scenarios · cost (the on-demand bill surprise; scan vs query) · common mistakes · interview questions · revision.
3h 30m

End of Phase 2 checkpoint: "I can build backend applications using AWS services."

Optional Deep Dive: DynamoDB Global Tables, NoSQL Workbench, DAX in depth, ElastiCache Serverless.

PHASE 3 — DAYS 8–11 · ADVANCED BACKEND AWS

Day 8 — Queues, Delivery Semantics and Idempotency

Outcome: you can build a consumer that is correct when — not if — the same message arrives twice, and you can defend every timeout number you chose.

Prerequisites: Days 1–7.

BlockMinutesContent
8.125Why a queue. The engineering problems it solves: load levelling, decoupling, retry with durability, spiky traffic, slow downstreams. What it costs you: ordering, immediate errors, simple tracing, and your ability to say "it's done" when the API returns.
8.245SQS internals. Distributed storage across AZs; why standard queues can reorder and duplicate (it falls out of the architecture, it's not a bug). The full lifecycle: send → store → receive → visibility timeout → delete-or-reappear. Long polling vs short polling and the empty-receive cost. Message attributes, size limits and the extended-client (S3) pattern. Delay queues. Redrive policy, DLQs, and redrive-back.
8.330Getting the numbers right. Visibility timeout vs processing time vs Lambda timeout vs maxReceiveCount — a worked calculation with a table of what breaks at each wrong value. Heartbeating (ChangeMessageVisibility) for long jobs.
8.430FIFO queues. Message group ID as the unit of ordering and parallelism (the thing everyone misunderstands). Deduplication ID and the 5-minute window. Exactly-once processing vs exactly-once delivery. Throughput limits, high-throughput mode. FIFO's real cost: a hot message group is a serial bottleneck.
8.535Idempotency, properly. Natural idempotency vs enforced idempotency. Implementations: conditional write on DynamoDB with a dedupe key, unique constraint in Postgres, idempotency-key table with state machine and TTL. The transactional outbox pattern for "write to DB and publish" atomicity. Poison messages.
8.625SNS + fan-out. Topics, subscriptions, filter policies, delivery retry policies, SNS DLQ. The canonical SNS → multiple SQS fan-out and why you almost never subscribe a service directly to SNS. FIFO topics. SQS vs SNS vs "both".
8.730Lab 11 — Spring Boot producer + consumer over SQS with a DLQ, a deliberately duplicated message, an idempotent consumer proving exactly-once effect, and consumer autoscaling on queue depth (ApproximateNumberOfMessagesVisible / backlog-per-task).
8.820Failure scenarios · monitoring a queue (the 4 metrics that matter, incl. ApproximateAgeOfOldestMessage) · interview traps · revision.
4h 00m

Day 9 — Events, Streams and Orchestration

Outcome: you can tell the difference between a queue, a topic, an event bus and a stream in one sentence each, and pick correctly under interview pressure.

Prerequisites: Day 8.

BlockMinutesContent
9.125Commands vs events. "Do this" vs "this happened". Why the distinction determines your coupling direction. Choreography vs orchestration, with the failure-visibility trade-off made explicit.
9.235EventBridge. Buses, rules, content-based pattern matching, targets, input transformers, archive and replay, schema registry. DLQ and retry on targets. Scheduler. Pipes as the "source → filter → enrich → target" glue. SQS vs SNS vs EventBridge decision matrix — routing intelligence vs throughput vs latency vs cost.
9.340Streams. The log abstraction: ordered, replayable, multi-consumer, retained. Kinesis Data Streams — shards, partition keys, sequence numbers, iterator age (the one metric that tells you everything), enhanced fan-out, on-demand vs provisioned, resharding. DynamoDB Streams and the CDC pattern. MSK / Kafka — what it gives you that Kinesis doesn't (ecosystem, compacted topics, consumer groups, unbounded retention) and what it costs you (operational surface, even on MSK Serverless).
9.425Queue vs stream vs bus — the comparison this whole day exists to earn. Includes SQS vs Kafka, the interview question where "they're interchangeable" is the trap answer.
9.525Step Functions. Standard vs Express. State types, error handling/retry/catch, the saga pattern for distributed transactions, callback patterns with task tokens. When a state machine beats a chain of queues — and when it's ceremony.
9.630Lab 12 — event-driven architecture. Order service publishes to EventBridge; rules fan out to a payment consumer (SQS + Lambda) and an analytics consumer (Kinesis → Firehose → S3); one consumer fails and is replayed from the archive.
9.720Failure scenarios (silent rule mismatch, iterator age climbing, poison record blocking a shard) · cost comparison at 1M events/day and 1B events/day · interview questions · revision.
3h 20m

Optional Deep Dive: Kafka Connect, Flink / Managed Service for Apache Flink, Firehose transformations, event sourcing in depth.

Day 10 — Storage at Scale and the Edge

Outcome: you can design a file-upload/download platform that handles 5 GB files, expiring links, and a global audience — and know what each piece costs.

Prerequisites: Days 1–5 (Day 3 S3 intro assumed).

BlockMinutesContent
10.135S3 deep. Storage classes and the real decision rule (access frequency × retrieval latency tolerance × minimum-duration charges); Intelligent-Tiering. Lifecycle policies. Versioning + MFA delete; delete markers and the "my bucket keeps growing" bill. Object Lock. Replication (CRR/SRR) and RTC. Request-rate scaling and the prefix myth, corrected. Event notifications → SQS/SNS/Lambda/EventBridge.
10.235Moving bytes correctly. Multipart upload: why, thresholds, part sizes, resumption, the abandoned-parts bill and the lifecycle rule that fixes it. Transfer Manager in SDK v2. Presigned URLs — the pattern that keeps large uploads out of your application entirely, expiry, scoping, and the checks you lose by doing so (and how to get them back with S3 event validation). Byte-range GETs.
10.325Object vs block vs file. S3 vs EBS vs EFS: the actual decision, with the "we used EFS as a database" anti-pattern. EBS snapshots, volume resizing, multi-attach. EFS performance and throughput modes, and its role for containers.
10.430CloudFront. POPs, origins (S3 with OAC, ALB, API Gateway), cache behaviours and path patterns, cache key composition (the #1 cause of 0% hit rate), TTLs and invalidation vs versioned object keys, signed URLs/cookies, origin failover, compression. CloudFront vs S3 direct and why "just make the bucket public" is the wrong answer three separate ways. CloudFront Functions vs Lambda@Edge.
10.520Route 53 for applications. Routing policies (simple, weighted, latency, failover, geolocation) as deployment tools: blue/green, canary, regional failover. Health checks. TTL as your rollback speed limit.
10.630Lab 13 — file upload/download platform. Presigned PUT from a browser-ish client → S3 event → Lambda validates and writes metadata to DynamoDB → CloudFront serves downloads with signed URLs. Multipart for a large file.
10.720Failure scenarios · data transfer cost, done properly (the egress/NAT/cross-AZ lesson that saves the most money of any hour in the course) · interview questions · revision.
3h 15m

Optional Deep Dive: S3 Access Points, Mountpoint for S3, S3 Tables, Global Accelerator, Storage Gateway.

Day 11 — Security and IAM at Developer Depth

Outcome: you can read a policy evaluation like a compiler error, hand your service exactly the permissions it needs, and say where every piece of data is encrypted and with whose key.

Prerequisites: Days 1–10.

BlockMinutesContent
11.140Policy evaluation, properly. The decision flow: explicit deny → SCP → resource policy → identity policy → permission boundary → session policy → implicit deny. Policy structure, Principal, conditions, variables, NotAction/NotResource hazards. The reasoning chain: who am I → what role am I assuming → what action → which resource → which policies apply → allow or deny? Reading real denial messages and mapping them back to the flow.
11.230Roles, STS and temporary credentials. AssumeRole, trust policy vs permission policy, ExternalId, session duration and chaining limits. Service roles: EC2 instance profile, ECS task role vs execution role, Lambda execution role. Cross-account access patterns. OIDC federation for CI (GitHub Actions without long-lived keys). IRSA named for EKS.
11.325Least privilege in practice. Starting broad and narrowing with Access Analyzer and CloudTrail evidence. Resource-level vs action-level scoping. Condition keys that do real work (aws:SourceVpce, aws:PrincipalOrgID, s3:prefix). Permission boundaries. The honest trade-off: least privilege vs deployment velocity.
11.435Encryption. At rest vs in transit, stated per service. KMS: CMK vs AWS-managed vs owned, envelope encryption explained mechanically, data keys, key policies (the resource policy that can lock you out), grants, rotation, multi-Region keys, and the request-cost/throttling reality. Client-side vs server-side. TLS everywhere, ACM certificate issuance and renewal, where TLS terminates in each architecture from Days 4–5.
11.525Application identity. Cognito user pools vs identity pools, the OAuth2/OIDC flows a backend developer actually implements, JWT validation in Spring Security, token lifetime and revocation. Alternatives named honestly (your own IdP, Auth0, Entra). Machine-to-machine auth with client credentials.
11.620Perimeter and detection. WAF rule groups, rate-based rules, the managed rule sets worth turning on. Shield Standard vs Advanced. GuardDuty, Security Hub, CloudTrail (management vs data events) and Config — what each answers, from a developer's point of view.
11.725Lab 14 — least privilege end to end. Take the Day 8 service, replace its permissive policy with a minimal one derived from CloudTrail, add a KMS-encrypted queue and bucket, rotate a secret, and prove a cross-account read works through AssumeRole.
11.820"What happens if this credential leaks?" exercise across every architecture built so far · interview questions · revision.
3h 40m

End of Phase 3 checkpoint: "I understand how AWS-backed applications behave in production."

PHASE 4 — DAYS 12–14 · PRODUCTION ARCHITECTURE AND SENIOR DEPTH

Day 12 — Scalability, Reliability and Cost as Design Inputs

Outcome: given a load number and a budget, you can produce an architecture and defend both the scaling story and the bill.

Prerequisites: Days 1–11.

BlockMinutesContent
12.135Scaling patterns. Vertical vs horizontal; statelessness as the enabling constraint; where state actually goes (DB, cache, object store, queue). Scaling the tiers in order — compute, then database (replicas → caching → partitioning → sharding), then the queue, then the cache. Choosing scaling signals (RPS, queue backlog per worker, concurrency) over CPU. Warm-up, scale-out lag, and why autoscaling never saves you from a 10-second spike.
12.240Reliability patterns, with AWS specifics. Timeouts (at every layer, with actual numbers), retries with exponential backoff and jitter, retry budgets, and the retry-storm/metastable-failure discussion that separates senior answers from junior ones. Circuit breaker, bulkhead, load shedding, graceful degradation, backpressure. Idempotency revisited as a system property. Multi-AZ by default; single points of failure hiding in "managed" services.
12.330Disaster recovery. RTO/RPO defined and then used. Backup-restore vs pilot light vs warm standby vs active-active, with AWS building blocks and honest cost per tier. Multi-Region: what actually becomes hard (data, consistency, failover decision-making, testing). Route 53 + Global Database + DynamoDB Global Tables. Why most companies should not go multi-Region, and how to say that in an interview without sounding unambitious.
12.425Failure-first analysis as a habit. The seven questions applied to every architecture built in this course: What can fail? What happens? How does it recover? How fast? Can it happen twice? Can data be lost? Can data be duplicated? Blast-radius and failure-domain thinking (cell-based architecture, shuffle sharding — introduced, not exhaustive).
12.540Cost engineering for developers. The bill's real shape: data transfer (egress, cross-AZ, NAT), NAT Gateway, idle RDS/ALB/NAT hourly charges, CloudWatch Logs ingestion + retention, DynamoDB capacity mode, Lambda GB-seconds, S3 requests + storage class minimums, ELB LCUs. "How does a developer accidentally create a $40,000 bill?" — five true-to-life scenarios with arithmetic (recursive Lambda on S3 events, chatty cross-AZ traffic, DEBUG logging at 20k RPS, unbounded scan, egress through NAT for S3). Cost-aware design choices that cost nothing to make. Tagging, Cost Explorer, Budgets.
12.620Architecture exercises (5 scenarios, load and budget given) · interview questions · revision.
3h 10m

Day 13 — Observability, Production Debugging and Migration

Outcome: given "the API got slow at 14:20," you have a procedure — not a hunch — and you can move a system from where it is to where it should be without a big-bang cutover.

Prerequisites: Days 1–12.

BlockMinutesContent
13.130Instrumenting an application properly. Structured JSON logging from Spring Boot, correlation/trace IDs propagated across HTTP and SQS/EventBridge (the hop where most people lose the thread), log levels and sampling, what never goes in a log (PII, secrets, tokens). Custom metrics: EMF (Embedded Metric Format) vs PutMetricData vs Micrometer/CloudWatch. RED and USE method, applied. Lab 15 starts here.
13.230CloudWatch as a tool, not a dashboard. Logs Insights query language, with the 8 queries you will actually reuse. Metric math, anomaly detection, composite alarms, alarms that don't page you at 3am for nothing. Dashboards worth having. Log retention and cost. Container Insights and Lambda Insights.
13.325Distributed tracing. X-Ray concepts (segments, subsegments, sampling, service map), instrumenting Spring Boot, tracing through async hops, ADOT/OpenTelemetry and the portability argument. Reading a service map to localize latency in 30 seconds. Lab 15 concludes: find an injected regression using telemetry only.
13.450Ten production troubleshooting scenarios, worked. Each one: symptom → hypotheses → the exact metric/log/trace/CLI command that discriminates → root cause → fix → prevention. (Latency spike; 504s under load; SQS backlog growing; DynamoDB throttling; RDS connection exhaustion; Lambda throttles and duplicate side-effects; ALB 502s after deploy; CloudFront 0% hit rate; cross-AZ cost spike; an IAM denial that only happens in production.) See 11 — Troubleshooting Plan.
13.545Migration. Six playbooks — monolith → microservices (strangler fig), on-prem → AWS (the 7 Rs, chosen not recited), EC2 → ECS, ECS/EC2 → Lambda, RDBMS → DynamoDB (the hardest, with the access-pattern audit), synchronous → event-driven (outbox + dual-write + backfill). For each: strategy, sequencing, data migration, dual-run and verification, rollback plan, risks, and the trade-off that makes it worth doing — or not.
13.620Architecture failure exercises · interview questions · revision.
3h 20m

End of Phase 4a checkpoint: "I can reason about scalability, reliability, security, cost and failure."

Day 14 — AWS System Design and Interview Readiness

Outcome: a repeatable method for any "design X on AWS" question, rehearsed on seven systems, plus the revision material for the night before.

Prerequisites: Days 1–13.

BlockMinutesContent
14.125The method. A 13-step framework applied identically every time: requirements → traffic/scale estimation → API design → compute → database → caching → messaging → storage → security → scaling → failure handling → observability → cost. How to run it out loud in 45 interview minutes, where to spend time, and how to signal seniority (state assumptions, name trade-offs, choose, then stress-test your own choice).
14.22h 00mSeven systems, ~17 min each. (1) High-scale URL shortener · (2) Order processing · (3) Notification platform · (4) File upload/download · (5) Payment/event processing · (6) Subscription/billing · (7) Real-time data processing. Each ships a full architecture, the numbers, the two or three decisions that actually matter, the failure analysis, and the cost envelope. See 14 — Capstone & System Designs.
14.325Interview traps. Nine statements that sound right and are wrong or incomplete (exactly-once SQS, Lambda is always cheaper, Multi-AZ is DR, private subnet means unreachable, …) with the corrected mental model for each.
14.420Question banks by YOE — how to use them, and a timed mock: 6 questions at your level. See 12 — Interview Plan.
14.515The revision system. 30-min / 1-hour / 3-hour / night-before / morning-of plans and when each is the right one. See 13 — Cheatsheet Structure.
14.6—Capstone — built across Days 4–13 as the Mini Assignments accumulate, assembled and reviewed here (Lab 16). Budget 4–6 hours outside the 14-day schedule; it is explicitly not counted in the 48h 35m.
3h 25m

Time budget summary

DayDuration
13h 30m
23h 40m
33h 30m
43h 40m
53h 10m
63h 25m
73h 30m
84h 00m
93h 20m
103h 15m
113h 40m
123h 10m
133h 20m
143h 25m
Total48h 35m

Plus, outside the budget and clearly labelled as such: the optional Foundation Module (4h 30m), the capstone assembly (4–6h), and every Optional Deep Dive.

Per-day file template

Every day file follows this exact heading structure, so a reader can navigate Day 11 the same way they navigated Day 1:

# Day X — Topic
## Learning Objectives
## Prerequisites
## Why Does This Exist?
## Beginner Explanation
## Core Concepts
## Architecture
## How It Works Internally
## Code
## Hands-on
## Production Considerations
## Failure Scenarios
## Security
## Performance
## Cost
## Alternatives
## Trade-offs
## Common Mistakes
## Interview Questions
## Senior-Level Questions
## Day-End Revision
## Mini Assignment

Sections that would be empty for a given day are omitted rather than padded — but the order never changes.