Hands-on Lab Plan
17 labs (Lab 0 – Lab 16). Every lab is a real AWS resource the learner creates, verifies and then deletes. No screenshot-following; each lab states what it proves.
Mandatory lab structure
Every lab file uses exactly this shape:
## Objective what you will be able to do afterwards
## Prerequisites prior labs, tools, IAM permissions needed
## Architecture Mermaid diagram of what you are about to build
## Steps numbered, each with the reason it exists
## Commands copy-pasteable AWS CLI v2 / Maven
## Code Java 17 + Spring Boot 3 + AWS SDK v2
## Expected Output what success actually looks like on screen
## Verification how you prove it works (not "it didn't error")
## Cleanup every billable resource, in dependency order
## Common Errors symptom → cause → fix table
## Production Relevance what changes when this is real trafficTwo rules that are never broken:
- Cleanup is not optional. Any lab creating a NAT Gateway, RDS/Aurora instance, ALB, ElastiCache node or Interface Endpoint states its hourly cost up front and its teardown at the end.
- No long-lived access keys. After Lab 0, every lab authenticates through a role. An access key pasted into
application.ymlis treated as a defect, not a shortcut.
The labs
| # | Lab | Day | Time | Builds | Est. cost if cleaned up |
|---|---|---|---|---|---|
| 0 | Account hardening | 1 | 20m | MFA on root, admin IAM role, billing alarm, CLI profile | $0 |
| 1 | S3 three ways | 1 | 20m | One object uploaded via Console, CLI, and Java SDK v2 | ~$0 |
| 2 | Build the VPC | 2 | 30m | 3-tier VPC, 2 AZs, IGW, NAT, SGs; private host proven reachable outward, unreachable inward | ~$1.10 (NAT, 1 day) |
| 3 | Spring Boot on EC2 | 3 | 20m | Jar deployed, systemd service, reachable via SSM Session Manager | ~$0.10 |
| 4 | Spring Boot + S3 via instance role | 3 | 15m | Zero credentials anywhere in the app | ~$0 |
| 5 | ECS Fargate behind an ALB | 4 | 30m | ECR image, task def, service, target group, autoscaling, rolling deploy watched live | ~$0.80 |
| 6 | Lambda in Java | 5 | 20m | Handler, client reuse, cold-start measured warm vs cold | ~$0 |
| 7 | API Gateway → Lambda | 5 | 25m | HTTP API, JWT authorizer, throttling proven with a load generator | ~$0 |
| 8 | Spring Boot + Aurora PostgreSQL | 6 | 30m | Private subnets, Secrets Manager, HikariCP, Flyway, manual failover observed | ~$1.50 |
| 9 | DynamoDB with the Enhanced Client | 7 | 20m | Table + GSI, bean mapping, a conditional write that blocks a double-update | ~$0 |
| 10 | ElastiCache cache-aside | 7 | 20m | Redis, TTL with jitter, measured hit rate and latency delta | ~$0.40 |
| 11 | SQS producer/consumer with a DLQ | 8 | 30m | Duplicate delivered on purpose, idempotent consumer, DLQ redrive, scale on backlog | ~$0 |
| 12 | Event-driven fan-out | 9 | 30m | EventBridge → (SQS+Lambda) and (Kinesis→Firehose→S3); one consumer fails and is replayed from archive | ~$0.20 |
| 13 | File upload/download platform | 10 | 30m | Presigned PUT → S3 event → Lambda → DynamoDB metadata; CloudFront signed-URL download; multipart for a large file | ~$0.10 |
| 14 | Least privilege end to end | 11 | 25m | CloudTrail-derived minimal policy, KMS-encrypted queue + bucket, secret rotation, cross-account AssumeRole | ~$0.10 |
| 15 | Observability and a real investigation | 13 | 30m | Structured logs, EMF custom metrics, X-Ray traces across an SQS hop, then find an injected latency regression using only telemetry | ~$0.30 |
| 16 | Capstone: production order platform | 14 | 4–6h | Everything, assembled | ~$3–6 |
Total in-course lab time: ~6h 15m (already counted inside each day's block budget). Total lab spend if every cleanup is followed: roughly $8–12, plus the capstone.
What each lab proves (the verification, not the steps)
| Lab | The verification that matters |
|---|---|
| 0 | aws sts get-caller-identity returns a role session, not a root or long-lived user identity |
| 1 | The same object is byte-identical whichever of the three paths created it — proving Console, CLI and SDK are one API |
| 2 | From the private host: curl https://checkip.amazonaws.com succeeds; from your laptop, nc to that host times out at the SG, not the NACL — and you can say which |
| 3 | The app answers /actuator/health through a port-forwarded SSM session, with no public IP and no SSH key |
| 4 | Removing every credential from config and rebooting the instance still lists the bucket |
| 5 | During a rolling deploy, a continuous curl loop records zero failed requests |
| 6 | Cold invocation vs warm invocation durations differ by the init time you can point to in the log's Init Duration |
| 7 | The 11th request in a second returns 429 with the throttle headers, and an unsigned request returns 401 |
| 8 | During failover the app throws for N seconds, then recovers — and you can state why N was what it was (DNS TTL + pool eviction) |
| 9 | Two concurrent conditional writes: exactly one succeeds, the other gets ConditionalCheckFailedException |
| 10 | Cache hit rate > 90% on the second pass, p99 drops by an order of magnitude, and a deliberately stampeding key is visible in the metrics |
| 11 | The same message delivered twice produces one row in the outcome table; a poison message lands in the DLQ after exactly maxReceiveCount attempts |
| 12 | Replay from the EventBridge archive reprocesses only the failed consumer's events, not everyone's |
| 13 | The presigned URL works before expiry and returns 403 after; a 1 GB upload resumes after being interrupted |
| 14 | The narrowed policy still passes every happy-path test, and aws s3 ls on an out-of-scope bucket is denied |
| 15 | You locate the injected latency to a specific subsegment using the service map, before reading any application code |
| 16 | The full failure drill: kill a task, drop a message, throttle the DB, leak a duplicate — the system still ends in a correct state |
Cost-control rules applied to every lab
| Resource | Rule |
|---|---|
| NAT Gateway | Created in Lab 2, deleted at end of Day 2, recreated only for labs that need private egress (and each such lab says so) |
| RDS/Aurora | Single smallest instance; deleted same day; final snapshot skipped deliberately (stated) |
| ALB | One, shared by Labs 5 and 15; deleted at end of Day 4 unless the learner is continuing to the capstone |
| ElastiCache | cache.t4g.micro, single node, deleted same day |
| Interface VPC endpoints | Demonstrated, then deleted in the same lab (hourly charge per AZ per endpoint) |
| CloudWatch Logs | Every log group gets an explicit retention (7 days) at creation — never left at "Never expire" |
| S3 | Lifecycle rule to abort incomplete multipart uploads after 1 day, set in Lab 13 |
| Everything | Tagged Project=aws-course so a single Cost Explorer filter and a single cleanup script can find it all |
A cleanup.sh per day and a master teardown checklist ship with the course. The final instruction of Day 14 is to run the master teardown and confirm a $0 forecast.
Labs mapped to the required lab list
The spec named 14 labs; all are covered, several merged or expanded:
| Spec lab | Covered by |
|---|---|
| S3 upload/download | Lab 1, Lab 13 |
| IAM role and permissions | Lab 0, Lab 4, Lab 14 |
| Deploy Spring Boot application | Lab 3, Lab 5 |
| Spring Boot + S3 | Lab 4 |
| SQS send/receive | Lab 11 |
| Asynchronous processing | Lab 11, Lab 12 |
| Lambda | Lab 6 |
| API Gateway + Lambda | Lab 7 |
| Spring Boot + RDS | Lab 8 |
| DynamoDB | Lab 9 |
| Redis/ElastiCache | Lab 10 |
| Event-driven architecture | Lab 12 |
| CloudWatch monitoring | Lab 15 |
| Complete production-style architecture | Lab 16 (capstone) |
| (added) Networking | Lab 2 |
| (added) Storage/CDN at scale | Lab 13 |
| (added) Security/least privilege | Lab 14 |