Day 3 — Compute, Storage and Seeing What Happened
3h 30m · Phase 1 of 4 · Curriculum
Learning Objectives
By the end of today you can:
- Choose an EC2 instance family and size for a JVM workload, and explain what the letters and numbers mean.
- Explain how an application running on EC2 obtains AWS credentials without any credential ever being stored, tracing it to the metadata service.
- Describe S3's data model precisely — and correct three common misconceptions about it.
- Deploy a Spring Boot application to a private-subnet EC2 instance and have it read and write S3 with zero credentials in code or config.
- Ship structured logs and custom metrics to CloudWatch, and create an alarm that would actually be worth waking up for.
- Explain why "make the instance bigger" eventually stops working — setting up Day 4.
The sentence you should be able to say tonight: "My application has no credentials; it has a role, and the platform hands it expiring keys."
Prerequisites
- Day 1 — IAM roles and policies, the credential provider chain, the CLI.
- Day 2 — subnets, Security Groups, SSM access to a private instance.
- The VPC from Lab 2. If you deleted it, rebuild it (the script in Day 2 takes ~5 minutes).
Why Does This Exist?
You have a network. You have an identity. You still have nowhere to run code and nowhere to put bytes.
Those two needs are the oldest in computing, and AWS's answers to them — EC2 and S3 — are its two oldest services. Almost everything else in the catalogue is a more specialized version of one of them, or glue between them.
But there is a third need that is less obvious and just as fundamental: you cannot operate what you cannot see. A process on a server you cannot SSH into, in a subnet with no inbound access, is a black box. When it misbehaves at 3am you need its logs, its metrics, and something that tells you before the customer does. That is CloudWatch, and the reason it appears on Day 3 rather than Day 13 is simple: from the first thing you deploy, you should be able to see what it did.
Beginner Explanation
EC2 in one sentence
EC2 gives you a virtual machine. You choose how much CPU and memory it has, which operating system image it boots, which subnet it sits in, and which Security Group guards it. Then it is a Linux box, and everything you know about Linux boxes applies.
That is both the appeal and the cost. You get complete control — and you own the operating system, the patching, the process supervision, the log rotation and the capacity planning. Day 4 shows what it looks like to give some of that back.
S3 in one sentence
S3 stores objects — blobs of bytes with a name and some metadata — and hands them back over HTTP.
flowchart LR
subgraph B["Bucket: acme-orders-prod (globally unique, in one Region)"]
O1["Key: documents/2026/01/inv-8812.pdf<br/>Body: 240 KB · Metadata: content-type, tags"]
O2["Key: exports/daily.csv"]
endThere is no filesystem here. No directories, no appending to an existing object, no partial overwrite, no rename. An object is created whole and replaced whole. What you do get is essentially unlimited capacity, extreme durability, and an HTTP interface.
CloudWatch in one sentence
CloudWatch is where AWS puts the numbers and the text your systems produce, plus a way to be told when a number crosses a line.
| Pillar | What it is | Example |
|---|---|---|
| Logs | Text your application emitted | "orderId":"o-88","status":"PAID" |
| Metrics | Numeric time series | CPUUtilization = 34% at 14:05 |
| Alarms | A rule over a metric, with an action | "CPU > 80% for 5 minutes → notify me" |
Core Concepts
1. EC2: the pieces
flowchart TB
AMI["AMI<br/>the boot image<br/>(Region-scoped)"] --> INST
TYPE["Instance type<br/>m7g.large<br/>(vCPU, memory, network)"] --> INST
SUBNET["Subnet<br/>(which AZ, which tier)"] --> INST
SG["Security Group<br/>(who may connect)"] --> INST
PROFILE["Instance profile<br/>(which IAM role)"] --> INST
UD["User data<br/>(first-boot script)"] --> INST
INST["EC2 instance"] --> EBS["EBS volume(s)<br/>network-attached block storage<br/>(AZ-bound)"]
INST -.-> IS["Instance store<br/>(local NVMe, ephemeral)"]Instance types — reading the name
m7g.large
│││ └──── size: nano < micro < small < medium < large < xlarge < 2xlarge < ...
││└────── processor: g = AWS Graviton (ARM), i = Intel, a = AMD, (none) = varies
│└─────── generation: 7 — newer is usually cheaper per unit of work
└──────── family| Family | Optimized for | Typical backend use |
|---|---|---|
| t | Burstable, cheap baseline | Dev boxes, low-traffic services — with a caveat, below |
| m | Balanced CPU:memory (1:4) | The default for a general JVM service |
| c | Compute (1:2) | CPU-bound work: encoding, compression, heavy serialization |
| r | Memory (1:8) | Caches, in-memory data, large JVM heaps |
| i / d | Local NVMe storage | Databases you operate yourself |
| g / p / inf | GPU / accelerators | Out of scope for this course |
Graviton (
gsuffix) is usually the right default for JVM workloads. Java runs on ARM without source changes, and Graviton instances are typically 20%+ cheaper for comparable performance. The only real friction is native dependencies and container base images, both of which are usually solved now.
The burstable trap
t3/t4g instances have a baseline CPU allocation (as low as 5–40% of a vCPU depending on size) and earn CPU credits while running below it. Bursting above baseline spends credits. When credits run out:
- T2 /
standardmode: performance is throttled hard to baseline. Your service slows to a crawl and nothing in your application metrics explains it. - T3 /
unlimitedmode (the default for T3): it keeps bursting and charges you extra per vCPU-hour. Your service stays fast and the bill quietly grows.
Either way the symptom is mysterious: "it was fine for a week, then every afternoon latency tripled." That is S5, one of today's failure scenarios. Watch the CPUCreditBalance metric, or use a non-burstable family for anything with sustained load.
Storage: EBS vs instance store
| EBS | Instance store | |
|---|---|---|
| What | Network-attached block device | Physical NVMe on the host |
| Lifetime | Independent of the instance; survives stop/start | Lost on stop, terminate, or host failure |
| AZ | Bound to one AZ | Bound to the host |
| Snapshots | Yes, to S3 | No |
| Performance | Very good; limited by volume type and instance's EBS bandwidth | Extremely fast |
| Use | Root volume, databases, anything you care about | Scratch, caches, shuffle space |
EBS volume types — the two that matter:
| Type | Baseline | Notes |
|---|---|---|
| gp3 | 3,000 IOPS and 125 MB/s included at any size, independently provisionable higher | The correct default. Decoupling size from performance is the whole point. |
| gp2 | 3 IOPS per GB (min 100, burst to 3,000) | Legacy. A small gp2 volume is slow because it is small. |
| io2 / io2 Block Express | Provisioned IOPS, higher durability | Databases with genuine IOPS requirements |
| st1 / sc1 | Throughput-optimized HDD | Big sequential reads, logs, cheap cold data |
⚠️ "My disk has IOPS limits" is a surprise for people coming from physical servers. A JVM writing verbose logs to a small gp2 root volume can saturate it and stall the application — with no CPU or memory pressure to explain it. Use gp3, and keep logs off the root volume (or ship them straight to CloudWatch, as we do today).
User data
A script that runs on first boot, as root. This is how an instance goes from a blank AMI to a running service without anyone logging in:
#!/bin/bash
dnf install -y java-17-amazon-corretto-headless
aws s3 cp s3://my-artifacts/order-service.jar /opt/app/app.jar
systemctl enable --now order-serviceKeep it short and idempotent. Anything complex belongs in the AMI (built with Packer/Image Builder) or in a container image (Day 4).
2. How EC2 gets credentials — the Instance Metadata Service
This is the concept that makes Day 1's "never use access keys for workloads" actionable.
Every EC2 instance can reach a link-local address, 169.254.169.254, which serves metadata about itself. It is not routed anywhere; it is answered by the hypervisor. When an instance profile (a wrapper around an IAM role) is attached, that service also vends temporary credentials for the role — and rotates them automatically before they expire.
sequenceDiagram
autonumber
participant APP as Your app (SDK v2)
participant CH as DefaultCredentialsProvider
participant IMDS as IMDSv2 at 169.254.169.254
participant STS as AWS STS (behind the scenes)
participant S3 as S3
APP->>CH: need credentials
Note over CH: steps 1-5 of the chain find nothing:<br/>no system properties, no env vars,<br/>no profile file, not a container
CH->>IMDS: PUT /latest/api/token<br/>X-aws-ec2-metadata-token-ttl-seconds: 21600
IMDS-->>CH: session token
CH->>IMDS: GET /latest/meta-data/iam/security-credentials/<br/>X-aws-ec2-metadata-token: ...
IMDS-->>CH: role name
CH->>IMDS: GET /latest/meta-data/iam/security-credentials/{role}
IMDS-->>CH: AccessKeyId, SecretAccessKey, Token, Expiration
Note over STS: credentials were minted by STS for the role<br/>and are cached/rotated by the instance
CH-->>APP: temporary credentials
APP->>S3: signed request
Note over CH: SDK refreshes automatically<br/>before ExpirationWhy this is the right answer, in one table:
| Access key in config | Instance role via IMDS | |
|---|---|---|
| Lifetime | Until someone deletes it | Hours; auto-rotated |
| Stored where | A file, an env var, a repo, a wiki | Nowhere |
| Rotation | Manual, and therefore never | Automatic |
| If leaked | Valid until noticed | Expires shortly |
| Scoped to | Whatever was convenient | One role, one workload |
IMDSv2 is the session-oriented version, and it is the default for new instances. The PUT-then-GET dance exists specifically to defeat SSRF: a tricked application that fetches an attacker-supplied URL cannot mint a token, because that requires a PUT with a custom header — something a naive "fetch this URL" bug will not do. The hop limit (default 1) additionally prevents the response from being forwarded to a container or another host.
The Day 1 credential chain is not trivia. Steps 5 and 6 — container credentials and IMDS — are what make every production example in the rest of this course credential-free. ECS (Day 4) and Lambda (Day 5) use the same idea with a different endpoint.
Try it yourself in today's lab:
TOKEN=$(curl -sX PUT "http://169.254.169.254/latest/api/token" \
-H "X-aws-ec2-metadata-token-ttl-seconds: 600")
curl -s -H "X-aws-ec2-metadata-token: $TOKEN" \
http://169.254.169.254/latest/meta-data/iam/security-credentials/3. S3: the data model, and three corrections
Bucket — a container, in one Region, with a globally unique name (3–63 characters, DNS-compatible). Key — the object's full name within the bucket. Up to 1,024 UTF-8 bytes. Object — the bytes (0 bytes to 5 TB), plus system metadata (size, ETag, storage class, encryption) and optional user metadata.
Correction 1: there are no folders
The Console shows folders. The API has only a flat key space. documents/2026/inv.pdf is a single key that happens to contain slashes; documents/ is not an entity. "Listing a folder" is really ListObjectsV2 with prefix=documents/2026/ and delimiter=/.
Consequences that matter:
- There is no atomic "rename a folder". Renaming means copy + delete, per object.
- Listing a prefix with millions of objects is paginated and slow; if you need to query, keep an index elsewhere (DynamoDB — Day 7 — or S3 Inventory).
- Deleting a "folder" is a bulk delete of every key under a prefix.
Correction 2: durability is not availability
| Property | S3 Standard design target | Means |
|---|---|---|
| Durability | 99.999999999% (11 nines) | Your bytes are very unlikely to be lost |
| Availability | 99.99% SLA-backed | The service may be briefly unreachable |
Eleven nines is about redundancy across devices and AZs. It says nothing about whether you can read the object at 14:03 on a bad day — and nothing at all about you deleting it yourself. Durability does not protect you from your own DeleteObject call. Versioning does (Day 10).
Correction 3: consistency is strong now
S3 provides strong read-after-write consistency for PUTs of new objects, overwrites, and DELETEs, for all requests. Older material describes eventual consistency for overwrites — that changed in December 2020. Interviewers sometimes still ask the old question; know both the current behaviour and that it changed.
Defaults you get for free (and should still be deliberate about)
| Default (for buckets created recently) | Meaning |
|---|---|
| Block Public Access: on | Bucket policies and ACLs cannot make objects public |
| Encryption at rest: on (SSE-S3, AES-256) | Every object encrypted server-side; KMS is opt-in (Day 11) |
| ACLs disabled (Object Ownership: bucket owner enforced) | Access is controlled by policies, not per-object ACLs. This is a big improvement — object ACLs were the cause of many public-data incidents. |
| Versioning: off | You must enable it (Day 10) |
Request-rate scaling: S3 scales to at least 3,500 write and 5,500 read requests per second per partitioned prefix, and partitions automatically as load grows. Older guidance about randomizing key prefixes for performance is obsolete for most workloads.
4. CloudWatch
Logs
flowchart LR
APP[Application] -->|agent / SDK / task driver| LG["Log group<br/>/aws/ec2/order-service"]
LG --> LS1["Log stream: i-0abc... "]
LG --> LS2["Log stream: i-0def..."]
LG --> RET["Retention setting<br/>⚠️ default: Never expire"]
LG --> INS[Logs Insights queries]| Concept | Meaning |
|---|---|
| Log group | A named container, usually one per application. Retention and encryption are set here. |
| Log stream | A sequence of events from one source — one instance, one container, one Lambda environment |
| Retention | ⚠️ Defaults to "Never expire." Set it at creation, every time. This single default is responsible for a great deal of surprise CloudWatch spend. |
Structure your logs as JSON. Unstructured text is greppable; JSON is queryable:
# Unstructured — you can only grep
2026-01-14 09:12:04 INFO Processed order o-8812 in 240ms for customer c-51
# Structured — you can aggregate, filter, and chart
{"ts":"2026-01-14T09:12:04Z","level":"INFO","msg":"order processed",
"orderId":"o-8812","customerId":"c-51","durationMs":240,"traceId":"1-65a..."}With the second form:
fields @timestamp, orderId, durationMs
| filter durationMs > 1000
| stats count(), avg(durationMs), pct(durationMs, 99) by bin(5m)Day 13 builds an entire investigative practice on this. Today you simply make sure the logs are structured, because retrofitting it later means touching every log statement in the codebase.
Metrics
| Concept | Meaning |
|---|---|
| Namespace | A grouping — AWS/EC2, AWS/S3, or your own Acme/OrderService |
| Metric name | CPUUtilization, OrdersProcessed |
| Dimensions | Key-value pairs that make a metric specific: InstanceId=i-0abc. A distinct combination of dimensions is a distinct metric, and is billed as one. |
| Statistic | How points are aggregated in a period: Average, Sum, Maximum, p99 |
| Period | 60s standard; 1s high-resolution (costs more) |
⚠️ Two EC2 metric facts that catch everyone:
- Basic monitoring publishes at 5-minute granularity. Detailed monitoring (1-minute) is a paid opt-in. If your alarm evaluates 1-minute periods on a basic-monitoring instance, it will sit in
INSUFFICIENT_DATA. - Memory and disk usage are not EC2 metrics. The hypervisor cannot see inside your OS.
CPUUtilization, network and EBS metrics come free; memory, swap and filesystem usage require the CloudWatch agent running inside the instance. Every engineer new to AWS looks for a memory graph and does not find one.
Alarms
An alarm watches one metric (or a metric-math expression) and has three states: OK, ALARM, INSUFFICIENT_DATA.
The parameters that decide whether it is useful or noise:
| Parameter | Guidance |
|---|---|
| Threshold + comparison | Set from observed behaviour, not from a round number you like |
| Period × Evaluation periods | The real detection latency. 60s × 5 means up to five minutes before it fires. |
| Datapoints to alarm | "3 out of 5" is far more robust than "5 out of 5" against a single flappy sample |
| Treat missing data | notBreaching, breaching, ignore, missing. Choose deliberately — an instance that dies stops publishing, and missing means your alarm never fires for the worst possible failure |
| Action | SNS topic → email/chat/pager. An alarm with no action is a dashboard widget. |
The rule of thumb this course repeats: alarm on symptoms your users feel (error rate, latency, queue age), not on causes (CPU). Cause-alarms fire when nothing is wrong and stay silent when something is.
Architecture
What you build today
flowchart TB
DEV[Developer] -->|SSM Session Manager| INST
subgraph VPC["VPC 10.0.0.0/16"]
subgraph PRIV["private-a 10.0.16.0/20"]
INST["EC2 t3.micro<br/>Amazon Linux 2023<br/>Spring Boot on :8080<br/>systemd + CloudWatch agent"]
PROF["Instance profile →<br/>course-app-role"]
end
subgraph PUB["public-a"]
NAT[NAT Gateway]
end
GWE[S3 gateway endpoint]
end
INST --- PROF
INST -->|"S3 API, via endpoint<br/>(no NAT charge)"| S3[(S3 bucket)]
INST -->|"logs + metrics"| CW[CloudWatch]
INST -->|"SSM, other AWS APIs"| NAT
PROF -.->|"temporary credentials<br/>via IMDSv2"| INST
CW --> ALARM[Alarm → SNS → email]Note what is absent: no public IP, no SSH key, no inbound Security Group rule, no credentials file. Everything the instance can do, it can do because of the role attached to it.
How It Works Internally
S3 PutObject, end to end
sequenceDiagram
autonumber
participant APP as Application
participant SDK as SDK v2
participant EP as s3.eu-west-1.amazonaws.com
participant FE as S3 front-end fleet
participant IAM as IAM / bucket policy
participant ST as Storage layer (multi-AZ)
APP->>SDK: putObject(bucket, key, body)
SDK->>SDK: SHA-256 of body, canonical request, SigV4 sign
SDK->>EP: PUT /key (HTTPS, chunked or single)
EP->>FE: route to a front-end host
FE->>IAM: is this principal allowed s3:PutObject on this ARN?<br/>(identity policy ∩ bucket policy ∩ endpoint policy)
IAM-->>FE: Allow
FE->>ST: write, redundantly, across multiple devices<br/>in multiple AZs (for Standard)
ST-->>FE: durable
FE-->>SDK: 200 OK · ETag · x-amz-server-side-encryption
SDK-->>APP: PutObjectResponse
Note over ST: Only now is the object readable —<br/>strong read-after-write consistencyWhat to take from this:
- The 200 means durable. S3 does not acknowledge until the object is redundantly stored. There is no "written but not yet safe" window to worry about.
- The ETag is usually the MD5 of the body for single-part uploads — useful for integrity checks. For multipart uploads it is not an MD5 of the whole object (Day 10 explains the
-Nsuffix). - Authorization is an intersection. Your identity policy must allow it, and the bucket policy must not deny it, and any VPC endpoint policy must allow it. A 403 can come from any of the three. Day 11 makes this precise.
ListBucketandGetObjectare different actions on different ARNs — the bucket and the objects. Granting one does not grant the other, and this trips up nearly everyone once.
How an instance becomes a running service
flowchart TD
A[RunInstances API call] --> B[EC2 selects a host in the chosen AZ]
B --> C[EBS root volume created from the AMI snapshot]
C --> D[Instance boots · ENI attached in your subnet · SG applied]
D --> E["cloud-init runs user data as root"]
E --> F[systemd starts your unit]
F --> G["App starts · SDK asks IMDSv2 for credentials"]
G --> H[Service healthy]
D -.-> I["Instance profile association<br/>makes credentials available at 169.254.169.254"]
I --> GThe gap between step A and step H — typically 60–120 seconds for a JVM app — is why autoscaling cannot rescue you from a sudden spike, and why Day 4 cares so much about it.
Code
The Spring Boot service
pom.xml (additions to Day 1's file):
<parent>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-parent</artifactId>
<version>3.3.4</version>
</parent>
<dependencies>
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-web</artifactId>
</dependency>
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-actuator</artifactId>
</dependency>
<dependency>
<groupId>software.amazon.awssdk</groupId>
<artifactId>s3</artifactId> <!-- version from the BOM -->
</dependency>
<!-- structured JSON logging -->
<dependency>
<groupId>net.logstash.logback</groupId>
<artifactId>logstash-logback-encoder</artifactId>
<version>8.0</version>
</dependency>
</dependencies>The S3 client bean — created once, for the life of the process:
package com.acme.config;
import org.springframework.boot.context.properties.ConfigurationProperties;
import org.springframework.context.annotation.Bean;
import org.springframework.context.annotation.Configuration;
import software.amazon.awssdk.auth.credentials.DefaultCredentialsProvider;
import software.amazon.awssdk.core.retry.RetryMode;
import software.amazon.awssdk.regions.Region;
import software.amazon.awssdk.services.s3.S3Client;
import java.time.Duration;
@Configuration
public class AwsConfig {
@Bean
S3Client s3Client(AppProperties props) {
return S3Client.builder()
.region(Region.of(props.region()))
// On EC2 this resolves via IMDSv2 to the instance role.
// Locally it resolves to your ~/.aws profile. Same code.
.credentialsProvider(DefaultCredentialsProvider.create())
.overrideConfiguration(c -> c
// Total budget for the call INCLUDING retries.
.apiCallTimeout(Duration.ofSeconds(10))
// Budget for ONE attempt. Confusing these two is why
// "we set a 2s timeout" turns into a 10s p99.
.apiCallAttemptTimeout(Duration.ofSeconds(3))
.retryStrategy(RetryMode.STANDARD))
.build();
}
}
@ConfigurationProperties(prefix = "app")
record AppProperties(String region, String bucket) {}application.yml — note what is not in it:
app:
region: eu-west-1
bucket: ${COURSE_BUCKET} # injected by the environment, never hardcoded
# There is no access key here. There never will be.
management:
endpoints:
web:
exposure:
include: health,info,metrics
endpoint:
health:
probes:
enabled: true
server:
shutdown: graceful # matters on Day 4; start the habit now
spring:
lifecycle:
timeout-per-shutdown-phase: 25sThe controller and service:
package com.acme.documents;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import org.springframework.web.bind.annotation.*;
import software.amazon.awssdk.core.ResponseBytes;
import software.amazon.awssdk.core.sync.RequestBody;
import software.amazon.awssdk.services.s3.S3Client;
import software.amazon.awssdk.services.s3.model.*;
import java.nio.charset.StandardCharsets;
import java.util.UUID;
@RestController
@RequestMapping("/documents")
public class DocumentController {
private static final Logger log = LoggerFactory.getLogger(DocumentController.class);
private final S3Client s3;
private final String bucket;
DocumentController(S3Client s3, AppProperties props) {
this.s3 = s3;
this.bucket = props.bucket();
}
@PostMapping
public DocumentRef create(@RequestBody String body) {
String id = UUID.randomUUID().toString();
String key = "documents/" + id + ".txt";
long start = System.nanoTime();
s3.putObject(PutObjectRequest.builder()
.bucket(bucket).key(key).contentType("text/plain").build(),
RequestBody.fromString(body, StandardCharsets.UTF_8));
long ms = (System.nanoTime() - start) / 1_000_000;
// Structured fields, not an interpolated sentence — this is what makes
// Day 13's Logs Insights queries possible.
log.atInfo()
.addKeyValue("event", "document.created")
.addKeyValue("documentId", id)
.addKeyValue("bytes", body.length())
.addKeyValue("durationMs", ms)
.log("document stored");
return new DocumentRef(id, key);
}
@GetMapping("/{id}")
public String read(@PathVariable String id) {
try {
ResponseBytes<GetObjectResponse> obj = s3.getObjectAsBytes(
GetObjectRequest.builder().bucket(bucket)
.key("documents/" + id + ".txt").build());
return obj.asUtf8String();
} catch (NoSuchKeyException e) {
// A domain outcome, not an error worth alerting on.
throw new DocumentNotFound(id);
} catch (S3Exception e) {
// 403 here means the instance role lacks s3:GetObject on this ARN.
// Log the ARN so the fix is obvious at 3am.
log.atError()
.addKeyValue("event", "s3.error")
.addKeyValue("status", e.statusCode())
.addKeyValue("bucket", bucket)
.addKeyValue("documentId", id)
.log(e.awsErrorDetails().errorMessage());
throw e;
}
}
}
record DocumentRef(String id, String key) {}logback-spring.xml — JSON to stdout, which the CloudWatch agent picks up:
<configuration>
<appender name="JSON" class="ch.qos.logback.core.ConsoleAppender">
<encoder class="net.logstash.logback.encoder.LogstashEncoder">
<includeMdcKeyName>traceId</includeMdcKeyName>
<customFields>{"service":"order-service"}</customFields>
</encoder>
</appender>
<root level="INFO">
<appender-ref ref="JSON"/>
</root>
</configuration>The IAM policy for the instance role
Deliberately narrow, from the first deployment:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "ObjectsInOnePrefixOnly",
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:PutObject"],
"Resource": "arn:aws:s3:::aws-course-123456789012-day3/documents/*"
},
{
"Sid": "ListOnlyThatPrefix",
"Effect": "Allow",
"Action": "s3:ListBucket",
"Resource": "arn:aws:s3:::aws-course-123456789012-day3",
"Condition": { "StringLike": { "s3:prefix": "documents/*" } }
},
{
"Sid": "WriteOwnLogs",
"Effect": "Allow",
"Action": ["logs:CreateLogStream", "logs:PutLogEvents", "logs:DescribeLogStreams"],
"Resource": "arn:aws:logs:eu-west-1:123456789012:log-group:/aws/ec2/order-service:*"
},
{
"Sid": "PublishAppMetrics",
"Effect": "Allow",
"Action": "cloudwatch:PutMetricData",
"Resource": "*",
"Condition": { "StringEquals": { "cloudwatch:namespace": "Acme/OrderService" } }
}
]
}Three things worth noticing, because they generalize:
s3:ListBuckettargets the bucket ARN;s3:GetObjecttargets the object ARN. Different resources, different statements.cloudwatch:PutMetricDatadoes not support resource-level permissions, soResourcemust be*— but a condition on the namespace restores the scoping. Recognizing "this action can't be resource-scoped, so scope it with a condition" is a genuinely senior IAM move.- Nothing grants
s3:DeleteObject. The application has no business deleting, so it cannot.
systemd unit
# /etc/systemd/system/order-service.service
[Unit]
Description=Order Service
After=network-online.target
Wants=network-online.target
[Service]
User=appuser
Environment=COURSE_BUCKET=aws-course-123456789012-day3
Environment=JAVA_TOOL_OPTIONS=-XX:MaxRAMPercentage=75
ExecStart=/usr/bin/java -jar /opt/app/app.jar
SuccessExitStatus=143
Restart=on-failure
RestartSec=5
StandardOutput=journal
StandardError=journal
[Install]
WantedBy=multi-user.target
MaxRAMPercentagerather than a fixed-Xmx: the same artifact then sizes itself correctly on at3.microand anm7g.xlarge, and — importantly for Day 4 — inside a container with a memory limit.
Hands-on
Lab 3 — Deploy Spring Boot to a private EC2 instance (20 min)
Objective: a JVM service running in a private subnet, reachable by you through SSM, with no SSH key and no inbound rule.
Prerequisites: the Day 2 VPC (NAT Gateway running), Lab 0 CLI profile.
⚠️ Billable:
t3.micro(~$0.01/hr) plus the NAT Gateway from Day 2. Under $0.50 for the session.
Steps:
- Create the S3 bucket for the app (same pattern as Lab 1, with
-day3). - Create the IAM role
course-app-rolewith: the policy above, plus the managed policyAmazonSSMManagedInstanceCore. Create an instance profile and add the role to it. - Build the jar:
mvn -q clean package. - Upload it to an artifacts prefix in the bucket.
- Launch a
t3.microinprivate-awithapp-sg, no public IP, the instance profile, and this user data:
#!/bin/bash
set -euxo pipefail
dnf install -y java-17-amazon-corretto-headless amazon-cloudwatch-agent
useradd -r -s /sbin/nologin appuser || true
mkdir -p /opt/app && chown appuser /opt/app
aws s3 cp s3://BUCKET/artifacts/app.jar /opt/app/app.jar --region eu-west-1
# ... write the systemd unit and the agent config here ...
systemctl daemon-reload
systemctl enable --now order-service
systemctl enable --now amazon-cloudwatch-agent- Connect and port-forward, so you can reach the app from your laptop without any inbound rule at all:
aws ssm start-session --target "$INSTANCE_ID" \
--document-name AWS-StartPortForwardingSession \
--parameters '{"portNumber":["8080"],"localPortNumber":["8080"]}'Verification:
curl -s localhost:8080/actuator/health
# {"status":"UP","groups":["liveness","readiness"]}Then confirm the security posture:
aws ec2 describe-instances --instance-ids "$INSTANCE_ID" \
--query 'Reservations[].Instances[].{Public:PublicIpAddress,Private:PrivateIpAddress,Key:KeyName}'
# Public: null Key: null ← the point of the labCommon errors:
| Symptom | Cause | Fix |
|---|---|---|
| Instance never appears in SSM | No egress (NAT down / route missing), or no instance profile | Check the route table, then the profile attachment |
cloud-init finished but no jar | Role lacks s3:GetObject on the artifacts prefix | The policy above scopes to documents/* — add the artifacts prefix, deliberately |
| Service fails to start, no logs | Look at journalctl -u order-service -n 100 | Usually a missing env var or a Java version mismatch |
Port forward connects, curl refuses | App bound to 127.0.0.1, or still starting | ss -lntp; wait for the JVM |
Cleanup: terminate the instance. Keep the bucket and role for Lab 4.
Lab 4 — S3 via the instance role, with zero credentials (15 min)
Objective: prove the credential chain resolves to the instance role, and that removing every trace of a credential changes nothing.
Steps:
- On the instance, confirm there is no credentials file:
ls -la ~/.aws 2>&1 # No such file or directory
env | grep -i AWS_ # nothing but maybe AWS_REGION- Ask the metadata service what identity the instance has:
TOKEN=$(curl -sX PUT "http://169.254.169.254/latest/api/token" \
-H "X-aws-ec2-metadata-token-ttl-seconds: 600")
ROLE=$(curl -s -H "X-aws-ec2-metadata-token: $TOKEN" \
http://169.254.169.254/latest/meta-data/iam/security-credentials/)
echo "Role: $ROLE"
curl -s -H "X-aws-ec2-metadata-token: $TOKEN" \
"http://169.254.169.254/latest/meta-data/iam/security-credentials/${ROLE}" \
| python3 -c 'import json,sys; d=json.load(sys.stdin); print("Expires:", d["Expiration"])'- Confirm the SDK agrees:
aws sts get-caller-identity
# Arn: arn:aws:sts::123456789012:assumed-role/course-app-role/i-0abc...
# Note "assumed-role" and "sts" — not an IAM user. These are temporary credentials.- Exercise the application end to end:
curl -s -XPOST localhost:8080/documents -H 'Content-Type: text/plain' -d 'first document'
# {"id":"3f2a...","key":"documents/3f2a....txt"}
curl -s localhost:8080/documents/3f2a...
# first document- Prove least privilege is real. The role has no
DeleteObject:
aws s3api delete-object --bucket "$BUCKET" --key documents/3f2a....txt
# An error occurred (AccessDenied) ...- Prove the credentials expire and are rotated. Note the
Expirationfrom step 2, and check again later — the value moves without anyone doing anything.
Verification — the one that matters: reboot the instance. The application comes back and still works, and there is still no credential anywhere on disk.
Cleanup:
aws ec2 terminate-instances --instance-ids "$INSTANCE_ID"
aws s3 rm "s3://${BUCKET}" --recursive && aws s3api delete-bucket --bucket "$BUCKET"
# And, if you are stopping here for the day, the NAT Gateway + EIP from Day 2.Production relevance: this is exactly how every AWS-hosted application should obtain credentials. ECS (Day 4) and Lambda (Day 5) use the same mechanism through a different endpoint. If you take one habit from Day 3, take this one.
CloudWatch: logs, a custom metric, and one good alarm (part of block 3.5)
Agent configuration — ship the journal to a log group with a retention set from the start:
{
"logs": {
"logs_collected": {
"files": {
"collect_list": [{
"file_path": "/var/log/messages",
"log_group_name": "/aws/ec2/order-service",
"log_stream_name": "{instance_id}",
"retention_in_days": 7
}]
}
},
"force_flush_interval": 5
},
"metrics": {
"namespace": "Acme/OrderService",
"append_dimensions": { "InstanceId": "${aws:InstanceId}" },
"metrics_collected": {
"mem": { "measurement": ["mem_used_percent"], "metrics_collection_interval": 60 },
"disk": { "measurement": ["used_percent"], "resources": ["/"], "metrics_collection_interval": 60 }
}
}
}Memory and disk appear here because, as noted, they are not EC2 metrics. Without the agent there is no memory graph.
Set retention explicitly — the single highest-value CloudWatch command:
aws logs put-retention-policy --log-group-name /aws/ec2/order-service --retention-in-days 7A query worth keeping:
fields @timestamp, documentId, durationMs
| filter event = "document.created"
| stats count() as n, avg(durationMs) as avg, pct(durationMs, 99) as p99 by bin(5m)
| sort @timestamp descAn alarm worth having. Not CPU — a symptom:
# SNS topic first, so the alarm has somewhere to go
TOPIC=$(aws sns create-topic --name course-alerts --query TopicArn --output text)
aws sns subscribe --topic-arn "$TOPIC" --protocol email --notification-endpoint you@example.com
# Alarm on p99 latency of the application's own metric
aws cloudwatch put-metric-alarm \
--alarm-name order-service-p99-latency \
--namespace Acme/OrderService --metric-name DocumentWriteLatency \
--extended-statistic p99 --period 60 \
--evaluation-periods 5 --datapoints-to-alarm 3 \
--threshold 1000 --comparison-operator GreaterThanThreshold \
--treat-missing-data notBreaching \
--alarm-actions "$TOPIC"Read the flags as a sentence: "If, in 3 of any 5 consecutive minutes, p99 write latency exceeds one second, tell me — and don't fire merely because no requests arrived."
Production Considerations
| Concern | Production practice |
|---|---|
| Don't deploy by copying a jar | Bake an AMI (Image Builder/Packer) or build a container image. User data that downloads artifacts is fine for learning, fragile in production. |
| Never a single instance | An instance is a single point of failure in a single AZ. Day 4 fixes this with an ASG across AZs. |
| Log retention always | Set it at log-group creation. 7–30 days hot, archive to S3 if you need more. |
| Install the agent | Otherwise you are blind to memory and disk, which are how JVM services actually die. |
| gp3, not gp2 | And keep logs off the root volume. |
| Graviton where you can | Meaningful savings for JVM workloads with no code change. |
| Patching is yours | On EC2, guest OS patching is your responsibility (shared responsibility model, Day 1). SSM Patch Manager, or move up the ladder to containers/serverless. |
| Tag everything | Project, Environment, Owner. Cost allocation and cleanup both depend on it. |
| S3 naming | Include account and environment. Bucket names are global and permanent-ish; test-bucket was taken in 2006. |
Failure Scenarios
S5 — "Fine for a week, then latency triples every afternoon"
A service runs on t3.micro instances. Response times are excellent for six days, then degrade badly each afternoon, recovering overnight.
| Question | Answer |
|---|---|
| What can fail? | CPU credit exhaustion on a burstable instance |
| What happens? | In standard mode, CPU is throttled to baseline (a fraction of a vCPU) and everything queues. In unlimited mode, performance holds and the bill grows instead. |
| Recovery? | Credits accrue overnight when load is low — which is exactly why it "fixes itself" and returns the next day |
| How quickly? | Hours; it is a slow accumulation |
| Data lost / duplicated? | No — but requests time out, and retrying clients make it worse |
The discriminating signal: CPUCreditBalance in the AWS/EC2 namespace trending to zero, while CPUUtilization flatlines at a suspiciously round number. Application metrics show nothing, because nothing is wrong with the application.
The fix: use a non-burstable family (m7g) for sustained load, or accept unlimited mode knowingly and alarm on CPUSurplusCreditsCharged.
The lesson that generalizes: AWS resources have limits that are invisible from inside the operating system. The instance reports plenty of idle CPU while being throttled. This same shape recurs with EBS IOPS, network bandwidth, and DynamoDB capacity (Day 7).
S6 — CloudWatch Logs is the third-largest line item
A team enables DEBUG logging to diagnose an issue and forgets to turn it off. Log groups were created with the default retention.
Arithmetic, at 2,000 requests/second with 3 KB of logs per request:
2,000 × 3 KB = 6 MB/s
≈ 518 GB/day
Ingestion 518 × 30 × $0.50 ≈ $7,770/month
Storage, accumulating forever at ~$0.03/GB-month, growing by ~15 TB/month| Question | Answer |
|---|---|
| What can fail? | Nothing functionally. The failure is the invoice. |
| What happens? | Ingestion charges dominate; storage grows without bound because retention is Never expire |
| Recovery? | Set retention (applies to existing data too), reduce log level, sample high-volume events |
| How quickly? | Retention change is immediate; the ingestion charge is already spent |
| Data lost? | Only what retention now expires — which is the point |
Prevention, all of it cheap: set retention at creation; make log level runtime-configurable (Spring Boot Actuator's /actuator/loggers does this without a redeploy); sample debug-level events; never log request/response bodies at scale; and put a CloudWatch Logs cost alarm in place.
A1 (from the failure plan) — single-AZ everything
Today's architecture has one instance in one AZ. Run the seven questions on it:
| What can fail? | The instance, the host, the AZ |
| What happens? | Total outage. No health check, no replacement, no second copy. |
| How does recovery happen? | A human notices and launches another instance |
| How quickly? | However long it takes someone to notice, plus boot time |
| Can the operation happen twice? | N/A |
| Can data be lost? | Anything on the EBS root volume without a snapshot — but note that today's design already avoids this by keeping all state in S3 |
| Can data be duplicated? | N/A |
Which is exactly why Day 4 exists. The one design decision that survives from today is keeping state out of the instance: because documents live in S3, replacing the instance loses nothing. That property — statelessness — is what makes horizontal scaling possible tomorrow.
Security
| Question | Today's answer |
|---|---|
| Who can access this? | The instance: no inbound path at all. You, via SSM, which is IAM-authorized and logged in CloudTrail — strictly better than an SSH key, which is a file someone can copy. |
| What credentials are used? | Temporary, from the instance role, via IMDSv2. Nothing persists on disk. |
| Where is data encrypted? | S3 at rest by default (SSE-S3); in transit by TLS on every API call. The EBS root volume should be encrypted too — it is a launch option, on by default in many accounts, and worth verifying. |
| What if credentials leak? | IMDS credentials expire within hours. An attacker who exfiltrates them gets a short window, scoped to a role that can read and write one prefix of one bucket and cannot delete anything. That is what least privilege buys you — not prevention, but a small, expiring blast radius. |
Two controls specific to today:
- IMDSv2 required, hop limit 1. Set
HttpTokens=requiredat launch. This is the mitigation for SSRF-based credential theft, and it is the specific control that would have prevented several well-known cloud breaches. - SSM instead of SSH. No port 22 open, no key material to manage or rotate, every session authorized by IAM and recorded. In production, enable session logging to S3/CloudWatch.
Performance
| Lever | What to know |
|---|---|
| Instance family | Match the shape of the workload: c for CPU-bound, r for heap-hungry, m when unsure. Wrong family is the most common and most fixable performance problem. |
| EBS | gp3 gives 3,000 IOPS baseline at any size. The instance also has an EBS bandwidth ceiling — a big volume on a small instance is still slow. |
| JVM in a small instance | Use MaxRAMPercentage, not a fixed heap. Leave headroom for metaspace, threads and native memory — a 1 GB instance with a 900 MB heap gets OOM-killed by the kernel, which looks like a crash with no Java stack trace. |
| S3 throughput | Very high, but every operation is a network round trip (~10–30 ms). Batch where you can; stream large objects rather than buffering them (Day 10's multipart). |
| S3 request rate | ≥3,500 writes and ≥5,500 reads per second per partitioned prefix, scaling automatically. Rarely your bottleneck. |
| Boot time | 60–120 s from RunInstances to a serving JVM. This number is why autoscaling lags demand (Day 4, Day 12). |
Cost
Today's build (eu-west-1, illustrative — verify current pricing):
| Resource | Rate | 3 hours | Left running 1 month |
|---|---|---|---|
t3.micro | ~$0.0114/hr | ~$0.03 | ~$8.30 |
| EBS gp3 8 GB root | ~$0.08/GB-mo | ~$0.003 | ~$0.64 |
| S3 storage (a few MB) | ~$0.023/GB-mo | ~$0 | ~$0 |
| S3 requests | ~$0.005/1k PUT | ~$0 | ~$0 |
| CloudWatch Logs ingestion | ~$0.50/GB | ~$0 | depends entirely on log volume |
| CloudWatch custom metrics | ~$0.30/metric-mo | — | ~$0.90 for 3 metrics |
| CloudWatch alarm | ~$0.10/alarm-mo | — | ~$0.10 |
| NAT Gateway (from Day 2) | ~$0.045/hr | ~$0.14 | ~$32 |
Three cost lessons from today:
- The NAT Gateway still dominates. A
t3.microcosts less than a third of what the NAT Gateway costs to sit idle beside it. Cost intuition from on-premises — where compute is the expensive thing — does not transfer. - CloudWatch custom metrics are priced per unique dimension combination. A metric dimensioned by
InstanceIdacross 200 instances is 200 metrics, ~$60/month, for one number. Dimension by service and environment; use EMF or percentile statistics rather than high-cardinality dimensions. Day 13 covers this properly. - Log retention is the setting that costs the most to forget. One command, at creation, every time.
Alternatives
| Instead of | You could | Trade-off |
|---|---|---|
| EC2 | ECS/Fargate (Day 4) | No OS to own, faster scaling, per-task isolation. Less control; containerization required. |
| EC2 | Lambda (Day 5) | No servers at all, per-request billing. Execution limits, cold starts, a different programming model. |
| EC2 | Elastic Beanstalk / App Runner | Faster to a running app; an abstraction you will eventually fight |
| S3 | EBS | A filesystem, low latency, one instance. No HTTP interface, no 11-nines durability, AZ-bound. |
| S3 | EFS | Shared POSIX filesystem across instances. More expensive, slower, but sometimes genuinely required. |
| CloudWatch | Datadog, Grafana Cloud, ELK | Better UX and cross-cloud; another vendor, another bill, and CloudWatch is still where AWS-native metrics originate |
| CloudWatch agent | OpenTelemetry Collector / ADOT | Portable and vendor-neutral; more to configure. Day 13 returns to this. |
| SSM Session Manager | SSH with a bastion | Familiar; you now manage keys, a bastion host, and an inbound rule |
Trade-offs
| Decision | Gain | Cost |
|---|---|---|
| EC2 over a managed platform | Total control, any runtime, any process | You own patching, supervision, capacity, scaling |
| State in S3 rather than on disk | Instances become disposable → horizontal scaling becomes possible | Every access is a network call, ~10–30 ms; no partial writes |
| Instance role over access keys | Nothing to leak, automatic rotation | You must understand trust policies (Day 11) |
| Structured JSON logs | Queryable, aggregatable, chartable | Slightly larger payloads; harder to read raw with tail |
| Narrow IAM policy from day one | Small blast radius | You will hit AccessDenied during development — which is the system working |
| Burstable instances | Very cheap baseline | Throttling or surprise charges under sustained load |
| Detailed monitoring (1-min metrics) | Faster detection | ~$2.10/instance/month; unnecessary for most instances |
Common Mistakes
- Putting application state on the instance. Local files, in-memory sessions, a local upload directory. It works until there are two instances, and then it fails in ways that look random.
- Expecting a memory metric for free. Install the agent, or stay blind to the most common JVM failure mode.
- Leaving log retention at "Never expire." The most expensive default in AWS.
- Using
t3for sustained production load without watching credit balance. - A fixed
-Xmxin a container or a resized instance. UseMaxRAMPercentage. - Granting
s3:*on*"to get it working." Then never narrowing it. Start narrow; widen on evidence. - Confusing
s3:ListBucket(bucket ARN) withs3:GetObject(object ARN). Everyone does this once. - Believing 11 nines of durability protects you from yourself. It does not. Versioning does.
- Alarming on CPU. It fires when nothing is wrong and stays quiet when something is. Alarm on latency, errors, and queue age.
- Treating "missing data" as OK by accident. If an instance dies and stops publishing, a
missing-treated alarm never fires — for the worst failure you have.
Interview Questions
2–3 YOE
- What is an AMI, and what is the relationship between an AMI, an instance and an EBS volume?
- How does an application on EC2 get AWS credentials without any credentials being stored on it?
- What is the difference between S3 and EBS?
- What are the three things CloudWatch gives you?
- Why is there no memory metric for an EC2 instance by default?
- What does S3 durability of "11 nines" mean, and what does it not protect you from?
4–5 YOE
- Your service is slow every afternoon and CPU utilization sits flat at 20%. What is your first hypothesis?
- Why does a small gp2 volume perform badly, and what changed with gp3?
- Explain IMDSv2's
PUT-then-GETflow and the attack it is designed to prevent. - Your application gets
AccessDeniedonListObjectsV2butGetObjectworks. What is wrong? - You need a CloudWatch alarm that fires when your service is failing but not when it is merely idle. Which settings matter, and how do you set them?
Senior-Level Questions
6–7 YOE
- Design the logging and metrics strategy for a 40-service platform where CloudWatch costs must stay under $2,000/month. What do you log, at what level, with what retention, and what do you deliberately not collect?
- A team wants to store user uploads on the instance's EBS volume "for speed" and sync to S3 nightly. Make the argument against it — and describe the one situation where they might be right.
- How would you bootstrap 200 instances so that a new deployment takes under 90 seconds and requires no human in the loop?
- Your instance role needs
cloudwatch:PutMetricData, which does not support resource-level permissions. How do you scope it anyway? - What would make you choose EC2 over ECS/Fargate in 2026, and how would you defend that in a design review?
8–10 YOE
- You inherit a fleet where every instance has an admin-equivalent role "because narrowing it kept breaking things." Describe how you would get to least privilege without stopping delivery, and how long it would take.
- An application is being throttled by a limit that is invisible from inside the OS. Name three distinct AWS resources where this happens, and describe how you would detect each before it becomes an incident.
- Argue both sides of "all state in S3" for a latency-sensitive service. When is the network round trip unacceptable, and what would you do instead?
- Your CloudWatch bill is dominated by custom metrics. Explain how that happens, and describe a redesign that keeps the observability and removes the cost.
Day-End Revision
The five sentences
- EC2 gives you a virtual machine and, with it, ownership of the OS, patching and capacity.
- An instance profile lets the SDK obtain temporary, auto-rotating credentials from IMDSv2 — so no workload ever needs an access key.
- S3 stores immutable objects in a flat key space; durability is not availability, folders do not exist, and reads are strongly consistent.
- CloudWatch gives you logs, metrics and alarms — but not memory or disk unless you install the agent, and not 1-minute EC2 metrics unless you pay for detailed monitoring.
- Keeping state out of the instance is what makes the instance disposable, and disposability is the prerequisite for everything on Day 4.
The diagram to redraw from memory: the IMDSv2 credential sequence — SDK → chain → token → role → temporary credentials → signed call.
The numbers
| gp3 baseline | 3,000 IOPS / 125 MB/s at any size |
| S3 object size | 0 B – 5 TB (single PUT up to 5 GB) |
| S3 request rate per prefix | ≥3,500 write / ≥5,500 read per second |
| EC2 basic monitoring | 5-minute periods (detailed = 1-minute, paid) |
| CloudWatch Logs ingestion | ~$0.50/GB |
| Custom metric | ~$0.30 per metric per month |
| JVM boot on t3.micro | ~60–120 s to serving |
Today's trap: "It's on EC2, so I'll put the access key in an environment variable." Attach a role. This is the last day anyone should consider the alternative.
Tomorrow needs: the instance role concept (ECS task roles are the same idea), the statelessness property you just built, the Security Groups from Day 2, and the observation that one instance in one AZ is not an architecture.
Phase 1 Checkpoint
Before Day 4, you should be able to answer all of these without notes. If any fail, the day to revisit is named.
| # | Question | Revisit |
|---|---|---|
| 1 | What happens between s3.putObject() and the object existing? | Day 1 |
| 2 | Why does the same code behave differently on your laptop and on EC2? | Day 1, Day 3 |
| 3 | What makes a subnet public? | Day 2 |
| 4 | Your app can reach S3 but not an external API. What do you check? | Day 2 |
| 5 | Stateful vs stateless firewall — which is which, and what is the symptom of getting it wrong? | Day 2 |
| 6 | How does an EC2 instance get credentials, and how long do they last? | Day 3 |
| 7 | Why is there no memory metric by default? | Day 3 |
| 8 | Name three things that cost money in what you built, in order of size. | Days 2–3 |
| 9 | What is the single point of failure in today's architecture? | Day 3 |
| 10 | Say the Phase 1 sentence. | — |
The Phase 1 sentence: "I understand what AWS is and how my backend application communicates with AWS."
Mini Assignment
Time: 45 minutes. This becomes the first service in your capstone repository.
- In
day-03/, take the Spring Boot service from today and add:GET /documents— list document IDs under the prefix, paginated (ListObjectsV2with a continuation token);DELETE /documents/{id}— which must fail with a clear 403-derived error, because the role has nos3:DeleteObject. Handle it gracefully and log it as a policy problem, not a bug;- a custom CloudWatch metric
DocumentWriteLatencypublished viaPutMetricDatain theAcme/OrderServicenamespace; - a custom Actuator
HealthIndicatorthat reportsDOWNif S3 is unreachable — and think about whether that is the right behaviour.
- Deploy it, generate ~50 documents, and write a Logs Insights query returning p50/p99 write latency in 5-minute bins.
- Create the p99 latency alarm and trigger it deliberately (add a
Thread.sleepbehind a query parameter, or write a 10 MB object). - In
day-03/NOTES.md:- Should the health check report
DOWNwhen S3 is unreachable? Argue both sides, then decide. (This is a genuine senior question — it decides whether a dependency outage becomes a total outage when the load balancer removes every instance. Day 4 returns to it.) - What exactly would be lost if this instance were terminated right now?
- Which three resources from Days 1–3 are costing you money as you write this?
- Should the health check report
- Run the full teardown and confirm with Cost Explorer that tomorrow's forecast is near zero.
- Commit.
Success criterion: the alarm fired and you received the notification; your NOTES answer to the health-check question names the failure mode of both choices.
AWS Documentation
- EC2 User Guide · Instance types · Burstable performance
- EBS volume types
- Instance Metadata Service (IMDSv2)
- IAM roles for EC2
- S3 User Guide · Consistency model · Best practices for request rates
- CloudWatch User Guide · Logs Insights query syntax · Alarms and missing data
- CloudWatch agent configuration
- SSM Session Manager
Previous: Day 2 — Networking for Backend Developers · Next: Day 4 — Running Services: Scaling Groups, Load Balancers, Containers (Phase 3)