Day 4 — Running Services: Scaling Groups, Load Balancers, Containers
3h 40m · Phase 2 of 4 · Curriculum
Learning Objectives
By the end of today you can:
- Configure an Auto Scaling Group and explain why scaling on CPU is usually the wrong signal for a request-driven service.
- Trace an ALB request from the client to your container and back, and say where TLS terminates.
- Diagnose
502vs503vs504from a load balancer without reading application logs. - Package a Spring Boot application as a container image that behaves correctly under a memory limit.
- Explain the difference between an ECS task role and an execution role — and why confusing them produces two completely different errors.
- Run a rolling deployment with zero failed requests, and explain each setting that makes that true.
- Choose between EC2, ECS/Fargate, EKS and Lambda with reasons, before you have been taught Lambda.
The sentence you should be able to say tonight: "Zero-downtime deployment is not a feature you enable; it is four settings that agree with each other."
Prerequisites
- Day 1 — IAM roles, the credential provider chain.
- Day 2 — public/private subnets, Security Groups referencing Security Groups.
- Day 3 — EC2, instance roles via IMDS, statelessness, CloudWatch metrics and alarms.
- Docker Desktop (or Podman/colima) installed and running.
- The Spring Boot service from Day 3's Mini Assignment.
Why Does This Exist?
Yesterday you deployed a service to one EC2 instance in one Availability Zone. Run the seven failure questions on it and the answers are all bad: the instance dies, the service is gone, and a human has to notice.
So you want more than one copy. That immediately creates three new problems you did not have:
- Who decides how many copies there are? Not a human at 3am. → Auto Scaling
- How does a client reach "the service" rather than a specific instance? → Load balancing
- How do you guarantee that copy #7, launched at 04:12 on a Tuesday, is byte-identical to copy #1? → Containers
That third one deserves a moment. Yesterday's instance was configured by a user data script that installed a JDK, downloaded a jar, and wrote a systemd unit. It worked. But it is a recipe, and recipes drift: the base AMI gets a new package version, a dnf install pulls a different minor release, the artifact in S3 is overwritten. Six weeks later, instance #7 is subtly not the same as instance #1, and you have a bug that reproduces on one host out of twelve.
A container image is not a recipe. It is the result of the recipe, frozen and content-addressed. Every copy is the same bytes. That is the whole reason containers took over server-side deployment, and it matters far more than any efficiency argument.
Beginner Explanation
Horizontal scaling, and what it demands
flowchart LR
subgraph V["Vertical: make it bigger"]
V1["t3.micro"] --> V2["m7g.xlarge"] --> V3["m7g.16xlarge"] --> V4["...and then?"]
end
subgraph H["Horizontal: make more of it"]
H1[copy] & H2[copy] & H3[copy] & H4[copy] --> LB[one address]
endVertical scaling is easier and has a ceiling: eventually there is no larger instance, and the whole time you have a single point of failure paying for peak capacity 24/7.
Horizontal scaling has no ceiling, survives losing a copy, and lets capacity follow demand. Its price is a constraint on your application: any copy must be able to serve any request. That is statelessness, and Day 3 already got you there by putting documents in S3 instead of on the instance's disk.
The three pieces
flowchart TB
C[Clients] --> ALB["Load Balancer<br/>one stable address,<br/>spreads traffic,<br/>removes sick copies"]
ALB --> T1[copy] & T2[copy] & T3[copy]
ASG["Auto Scaling<br/>decides HOW MANY copies,<br/>replaces dead ones"] -.manages.-> T1 & T2 & T3
IMG["Container image<br/>guarantees every copy<br/>is identical"] -.-> T1 & T2 & T3Each answers a different question: how many, how to reach them, and are they the same. You need all three, and they are independent — you can have an ASG of EC2 instances with no containers, or containers with no autoscaling. Production usually wants all three.
Core Concepts
1. Auto Scaling Groups
An ASG maintains a desired count of instances across the subnets you give it, replacing any that fail a health check.
flowchart TB
LT["Launch template<br/>AMI · instance type · SGs ·<br/>instance profile · user data"] --> ASG
ASG["Auto Scaling Group<br/>min=2 desired=4 max=20<br/>subnets: private-a, private-b, private-c"]
ASG --> I1[instance] & I2[instance] & I3[instance] & I4[instance]
HC{"Health check<br/>EC2 or ELB"} -.monitors.-> I1 & I2 & I3 & I4
HC -->|unhealthy| TERM["Terminate + launch replacement"]
POL["Scaling policy<br/>target tracking on a metric"] --> ASG| Setting | What it does | The trap |
|---|---|---|
| Launch template | The blueprint for every instance (launch configurations are legacy — do not use them) | Versioned; an ASG pins a version, so "I updated the template" may change nothing |
| min / desired / max | The bounds and the current target | max too low silently caps you at peak; an alarm on desired == max is worth having |
| Subnets | One per AZ; the ASG balances across them | List only one subnet and you have built a single-AZ architecture with extra steps |
| Health check type | EC2 (is the VM running?) or ELB (is the app responding?) | Default is EC2. A hung JVM passes an EC2 check forever. Always use ELB for a service behind a load balancer. |
| Health check grace period | Seconds after launch before health checks count | Shorter than your JVM boot time → an infinite launch/terminate loop. This is the most common ASG misconfiguration. |
| Termination policy | Who dies when scaling in | Defaults are sensible; OldestInstance is useful during migrations |
| Instance refresh | Rolling replacement of every instance | How you deploy a new AMI without a load balancer swap |
| Lifecycle hooks | Pause on launch/terminate to run something | Used for graceful drain, warm-up, registration |
Scaling policies — and why CPU is the wrong signal
| Policy | How it works | Use it when |
|---|---|---|
| Target tracking | "Keep metric X at value Y." AWS manages the alarms. | Default choice. Almost always what you want. |
| Step scaling | Explicit alarm thresholds → explicit capacity changes | You need asymmetric or non-linear behaviour |
| Simple scaling | Legacy; one adjustment with a cooldown | Don't |
| Scheduled | Capacity by clock | Known traffic shapes — market open, a nightly batch, a televised event |
| Predictive | ML forecast from history, pre-scales | Strongly cyclical traffic; combine with target tracking as a floor |
Now the important part. CPU utilization is a poor scaling signal for most request-driven backends, and understanding why is a senior-level idea:
- A service whose time goes on waiting — for a database, an HTTP call, a queue — can be fully saturated at 20% CPU. Its thread pool is exhausted, latency is terrible, and CPU-based scaling never fires.
- CPU is a lagging, indirect proxy. By the time it moves, your latency already moved.
- The JVM makes it worse: JIT compilation and GC produce CPU spikes unrelated to load.
Better signals, in rough order of preference:
| Signal | Why it is better |
|---|---|
ALBRequestCountPerTarget | Directly ties capacity to demand. "Keep each target at 1,000 requests/minute." |
| Concurrency / in-flight requests | Measures the actual constrained resource for a thread-per-request service |
| Queue backlog per worker | The correct signal for consumers (Day 8) |
| p99 latency | A symptom, not a cause — better as an alarm than a scaling trigger |
| CPU | Correct only when your service is genuinely CPU-bound |
Scale-out lag — the number that ruins everything
gantt
title From traffic spike to serving capacity
dateFormat ss
axisFormat %S s
section Detection
Metric published (60s period) :a1, 00, 60s
Alarm evaluation (2 datapoints) :a2, after a1, 60s
section Launch
EC2 launch + boot :b1, after a2, 45s
JVM start + warm-up :b2, after b1, 60s
section Registration
ALB health checks pass (2 × 30s) :c1, after b2, 60sThat is roughly 4–5 minutes from "traffic arrived" to "a new instance is serving." A spike that arrives over 90 seconds will not be met by autoscaling — it will be met by whatever headroom you already had.
Three consequences worth internalizing now:
- Autoscaling handles trends, not spikes. Run enough headroom to absorb the spike, and let autoscaling handle the trend.
- Every second you shave off boot time is capacity you get earlier. This is a real argument for containers (seconds, not minutes) and for Lambda (Day 5).
- Scale out fast, scale in slowly. Aggressive scale-in plus a bumpy metric produces thrashing. Target tracking builds in asymmetric behaviour for exactly this reason.
2. Load balancing
flowchart TB
C[Client] -->|":443"| L["Listener<br/>HTTPS :443<br/>ACM certificate"]
L --> R1{"Rule priority 10<br/>path /api/orders/*"}
L --> R2{"Rule priority 20<br/>host admin.acme.com"}
L --> RD["Default rule"]
R1 --> TG1["Target group: orders<br/>protocol HTTP :8080<br/>health check /actuator/health"]
R2 --> TG2["Target group: admin"]
RD --> TG1
TG1 --> T1["Target<br/>10.0.16.7:8080"] & T2["Target<br/>10.0.32.9:8080"]Listener → Rule → Target group → Target. Memorize that chain; every ALB problem lives at one of the four links.
| Concept | Detail |
|---|---|
| Listener | A port + protocol. HTTPS listeners hold the ACM certificate — this is where TLS terminates. |
| Rule | Evaluated by priority. Conditions can match host header, path, HTTP header, query string, source IP, or method. |
| Target group | A set of targets plus the health check and the routing algorithm. Targets can be instance IDs, IP addresses, a Lambda function, or (for NLB) an ALB. |
| Target | One instance or task on one port. Registered and deregistered automatically by the ASG or ECS. |
Health checks — the settings that cause most deployment failures
| Parameter | Default | What goes wrong |
|---|---|---|
| Path | / | Hitting / runs your whole app. Use a dedicated endpoint. |
| Interval | 30s | Long interval = slow detection of a dead target |
| Timeout | 5s | A slow-but-alive app fails a 5s timeout under load and gets removed — making the load worse |
| Healthy threshold | 5 | 5 × 30s = 150 s before a new target receives traffic. Compounds with boot time. |
| Unhealthy threshold | 2 | 2 × 30s = 60 s to remove a dead target |
| Success codes | 200 | Spring Actuator returns 503 when a component is DOWN — which may be exactly what you want, or may take your whole fleet out |
The health check design question — the one raised in Day 3's Mini Assignment, now answered properly:
Should
/actuator/healthreportDOWNwhen a downstream dependency (S3, the database) is unreachable?
Usually no — and this matters enormously. If every instance reports unhealthy because a shared dependency is degraded, the load balancer removes every target and returns 503 to everyone. A partial outage (some requests fail) becomes a total outage (no requests are even attempted). Worse, the ASG then terminates and relaunches the entire fleet, and the new instances fail too.
The discipline is to separate two questions:
| Probe | Question | Should it check dependencies? |
|---|---|---|
| Liveness | Is this process broken and in need of a restart? | No |
| Readiness (what the LB uses) | Can this instance serve traffic right now? | Only for dependencies without which it can serve nothing, and even then think hard |
Spring Boot gives you both: /actuator/health/liveness and /actuator/health/readiness, with management.endpoint.health.probes.enabled=true. Point the load balancer at readiness, keep external dependencies out of it, and alarm on the dependency separately.
Deregistration and draining
When a target is removed — a deploy, a scale-in, a failed health check — the ALB stops sending it new requests and waits deregistration delay (default 300 s) for in-flight requests to finish.
This only works if your application cooperates:
sequenceDiagram
autonumber
participant ECS as ECS / ASG
participant ALB
participant APP as Your app
ECS->>ALB: deregister target
ALB->>ALB: stop new requests; keep draining
ECS->>APP: SIGTERM
Note over APP: Spring graceful shutdown:<br/>stop accepting, finish in-flight,<br/>then exit
APP-->>ECS: exit 0
ALB->>ALB: drained
ECS->>APP: SIGKILL after stopTimeout<br/>(if still alive)Four settings must agree, and this is the answer to "how do I get zero-downtime deploys":
ALB deregistration delay ≥ longest in-flight request (e.g. 30 s)
ECS stopTimeout / ASG hook ≥ deregistration delay + margin (e.g. 45 s)
Spring graceful shutdown enabled, timeout < stopTimeout (e.g. 25 s)
Health check grace period ≥ JVM boot + warm-up (e.g. 90 s)Get one wrong and you drop requests on every deployment — a ~2% error blip for 30 seconds that everyone blames on "the network." That is scenario S7.
The 300-second default deregistration delay is usually too long and makes deploys crawl. Set it to your real p99 request duration plus headroom, typically 20–60 s.
Sticky sessions
The ALB can pin a client to a target with a cookie. It works, and needing it is a design smell: it means your service is not stateless, which breaks scale-in, deployments, and AZ failure. If you need shared session state, put it in Redis (Day 7). Reach for stickiness only for legacy applications you cannot change.
ALB vs NLB vs GWLB
| ALB | NLB | GWLB | |
|---|---|---|---|
| Layer | 7 (HTTP/HTTPS) | 4 (TCP/UDP/TLS) | 3 (packet) |
| Routing on | Host, path, headers, method, query | Port only | n/a |
| Latency added | Low (single-digit ms) | Very low (~100 µs) | n/a |
| Static IP | No — use a Route 53 alias | Yes, one Elastic IP per AZ | n/a |
| Preserves client source IP | No (use X-Forwarded-For) | Yes | Yes |
| WebSockets / HTTP/2 / gRPC | Yes | Pass-through | n/a |
| Lambda target | Yes | No | No |
| Cross-zone LB | Always on, free | Off by default, charged when on | — |
| Use for | Almost every HTTP backend | Non-HTTP, extreme latency needs, static IPs, very high connection rates | Inline security appliances |
Default to ALB for HTTP services. Choose NLB when you need a static IP (a partner is allowlisting you), a non-HTTP protocol, source IP preservation, or microsecond-level overhead.
The ALB's own IP addresses change as it scales its nodes. This is why you point DNS at it with a Route 53 alias record, never a hardcoded IP.
3. Containers, exactly as much as you need
| Term | Meaning |
|---|---|
| Image | An immutable, layered filesystem plus metadata (entrypoint, env, ports). Content-addressed by digest. |
| Layer | One filesystem diff. Cached and shared between images — which is why layer order determines build and pull speed. |
| Container | A running instance of an image, isolated by kernel namespaces and cgroups |
| Registry | Where images live. On AWS: ECR |
| Digest vs tag | sha256:abc… is immutable; :latest is a mutable pointer. Deploy by digest or by an immutable tag. |
A Spring Boot image that behaves correctly
Two things matter more than everything else: layer for cache efficiency, and let the JVM see the container's memory limit.
# ---- build stage ----
FROM maven:3.9-eclipse-temurin-17 AS build
WORKDIR /build
# Dependencies first: this layer is cached until pom.xml changes.
COPY pom.xml .
RUN mvn -q -B dependency:go-offline
COPY src ./src
RUN mvn -q -B clean package -DskipTests
# Spring Boot's layertools splits the fat jar into layers that change
# at different rates: dependencies (rarely) vs your classes (every commit).
RUN java -Djarmode=layertools -jar target/*.jar extract --destination /build/layers
# ---- runtime stage ----
FROM amazoncorretto:17-alpine
RUN addgroup -S app && adduser -S app -G app
WORKDIR /app
# Order matters: least-frequently-changed layer first.
COPY --from=build /build/layers/dependencies/ ./
COPY --from=build /build/layers/spring-boot-loader/ ./
COPY --from=build /build/layers/snapshot-dependencies/ ./
COPY --from=build /build/layers/application/ ./
USER app
EXPOSE 8080
# MaxRAMPercentage, not -Xmx: the JVM reads the cgroup limit, so the same
# image is correctly sized at 512 MB and at 8 GB.
ENV JAVA_TOOL_OPTIONS="-XX:MaxRAMPercentage=70 -XX:+ExitOnOutOfMemoryError"
ENTRYPOINT ["java","org.springframework.boot.loader.launch.JarLauncher"]Why MaxRAMPercentage and not -Xmx: since Java 10 the JVM is container-aware (UseContainerSupport, on by default) and reads the cgroup memory limit rather than the host's total memory. Before that, a JVM in a 512 MB container would see a 64 GB host, size its heap accordingly, and get OOM-killed by the kernel — a crash with no Java stack trace and no OutOfMemoryError, just exit code 137. If you ever see 137, this is what happened.
70% rather than 100% leaves room for metaspace, thread stacks, code cache, and direct buffers — all of which live outside the heap and all of which count against the container limit.
ECR
aws ecr create-repository --repository-name order-service \
--image-scanning-configuration scanOnPush=true \
--image-tag-mutability IMMUTABLE # a tag can never be repointed
aws ecr get-login-password | docker login --username AWS \
--password-stdin "${ACCOUNT}.dkr.ecr.${REGION}.amazonaws.com"
docker build -t order-service:1.0.0 .
docker tag order-service:1.0.0 "${ACCOUNT}.dkr.ecr.${REGION}.amazonaws.com/order-service:1.0.0"
docker push "${ACCOUNT}.dkr.ecr.${REGION}.amazonaws.com/order-service:1.0.0"Add a lifecycle policy immediately — untagged image layers accumulate and are billed per GB:
{"rules":[{
"rulePriority": 1,
"description": "Expire untagged images after 7 days",
"selection": {"tagStatus":"untagged","countType":"sinceImagePushed","countUnit":"days","countNumber":7},
"action": {"type":"expire"}
}]}4. ECS
flowchart TB
CL["Cluster<br/>a logical grouping + capacity"] --> SVC
TD["Task definition (family:revision)<br/>containers · cpu/memory · roles ·<br/>network mode · log config"] --> SVC
SVC["Service<br/>desired count · deployment config ·<br/>load balancer registration · autoscaling"]
SVC --> T1["Task (awsvpc: its own ENI,<br/>its own private IP, its own SG)"]
SVC --> T2[Task]
SVC --> T3[Task]
TG[ALB target group] -.registers.-> T1 & T2 & T3| Concept | Meaning |
|---|---|
| Cluster | A namespace plus capacity. With Fargate it is almost nothing; with EC2 launch type it holds the container instances. |
| Task definition | The immutable, versioned spec. Editing it creates a new revision; services point at a revision. |
| Task | One running instantiation — one or more containers scheduled together, sharing a network namespace |
| Service | The controller that keeps N tasks running, registers them with a load balancer, and performs deployments |
Fargate vs EC2 launch type
| Fargate | EC2 launch type | |
|---|---|---|
| You manage | Nothing below the container | AMI, patching, the ASG, capacity, bin-packing |
| Billing | Per vCPU-second and GB-second of the task | Per instance-hour, used or not |
| Start time | ~30–60 s | Seconds (if capacity exists), minutes if the ASG must scale |
| Density | One task per micro-VM | Many tasks per instance |
| GPU / privileged / host networking | No | Yes |
| Best for | Almost everything. Start here. | Very high steady density, GPUs, special kernel needs, deep cost optimization with Spot |
Default to Fargate. Move to EC2 launch type when you have measured a cost or capability reason.
Task role vs execution role — get this right
This distinction produces two completely different failures, and knowing which is which saves an hour every time.
flowchart TB
subgraph TASK["A running task"]
AGENT["ECS agent / Fargate infrastructure"]
APP["Your application container"]
end
EXEC["EXECUTION ROLE<br/>used by the platform, BEFORE your code runs:<br/>• pull the image from ECR<br/>• create log streams, put log events<br/>• read secrets injected into the task definition"] --> AGENT
TASKROLE["TASK ROLE<br/>used by YOUR CODE at runtime:<br/>• s3:PutObject<br/>• sqs:SendMessage<br/>• dynamodb:GetItem"] --> APP| Symptom | Which role | Why |
|---|---|---|
Task stops immediately: CannotPullContainerError | Execution role | It could not read from ECR |
| Task runs but no logs appear in CloudWatch | Execution role | It could not create the log stream |
Task starts, your code gets AccessDenied on S3 | Task role | Your application's own permissions |
| Secret injection fails at startup | Execution role | The platform fetches those, not your code |
Your application receives task-role credentials through the same credential-provider mechanism as Day 3's instance role — step 5 of the chain instead of step 6. ECS sets AWS_CONTAINER_CREDENTIALS_RELATIVE_URI, and the SDK fetches from 169.254.170.2. Your code does not change at all between EC2 and ECS. That is the payoff of never hardcoding credentials.
Deployments
stateDiagram-v2
[*] --> Steady: N tasks, revision 1
Steady --> InProgress: new revision registered
InProgress --> Starting: launch new tasks (up to maximumPercent)
Starting --> HealthCheck: register with target group
HealthCheck --> Draining: new tasks healthy → deregister old
Draining --> Steady: old tasks stopped (after deregistration delay)
HealthCheck --> RollingBack: circuit breaker trips
RollingBack --> Steady: revision 1 restored| Setting | Meaning | Typical |
|---|---|---|
minimumHealthyPercent | Floor on running capacity during a deploy | 100 (never lose capacity) |
maximumPercent | Ceiling, allowing extra tasks temporarily | 200 (start all new before stopping old) |
| Deployment circuit breaker | Detects a failing deployment and rolls back automatically | Enable it. Always. |
stopTimeout | SIGTERM → SIGKILL grace | 45 s (default 30, max 120) |
With min=100 / max=200, ECS starts new tasks, waits for them to pass health checks, registers them, deregisters the old ones, waits out the deregistration delay, then stops them. Combined with graceful shutdown, that is a genuinely zero-downtime deploy.
For stronger guarantees — instant rollback, traffic shifting, pre-traffic validation hooks — use blue/green via CodeDeploy, which runs two target groups and shifts the listener between them.
Service discovery
| Option | Mechanism | Use |
|---|---|---|
| ALB | Everything goes through the load balancer | Public-facing, and honestly fine for internal too |
| ECS Service Connect | A managed sidecar proxy; logical service names, per-request metrics, retries | Service-to-service inside ECS. The current best answer. |
| Cloud Map / DNS | Route 53 private hosted zone records per task | Simple, but DNS caching makes failover slow |
| Internal ALB | A second, non-internet-facing ALB | Simple and costs an extra ALB |
Task autoscaling
Application Auto Scaling, with target tracking:
aws application-autoscaling register-scalable-target \
--service-namespace ecs --scalable-dimension ecs:service:DesiredCount \
--resource-id service/course-cluster/order-service --min-capacity 2 --max-capacity 20
aws application-autoscaling put-scaling-policy \
--service-namespace ecs --scalable-dimension ecs:service:DesiredCount \
--resource-id service/course-cluster/order-service \
--policy-name track-requests --policy-type TargetTrackingScaling \
--target-tracking-scaling-policy-configuration '{
"TargetValue": 1000.0,
"PredefinedMetricSpecification": {
"PredefinedMetricType": "ALBRequestCountPerTarget",
"ResourceLabel": "app/course-alb/abc123/targetgroup/orders-tg/def456"
},
"ScaleOutCooldown": 60,
"ScaleInCooldown": 300
}'Read the asymmetry: scale out after 60 s, scale in after 300 s. Adding capacity quickly is cheap insurance; removing it quickly risks thrashing.
5. The compute decision matrix
This is deliberately placed before you learn Lambda, so that Day 5 answers a question you already have.
| EC2 | ECS on EC2 | ECS Fargate | EKS | Lambda (Day 5) | |
|---|---|---|---|---|---|
| Unit of deployment | AMI / jar | Container | Container | Container | Function / zip / image |
| You patch the OS | Yes | Yes (the hosts) | No | Yes (nodes) or no (Fargate) | No |
| Scale granularity | Instance | Task on existing capacity | Task | Pod | Request |
| Scale-out latency | 2–5 min | Seconds–minutes | ~30–60 s | Seconds–minutes | ~0–1 s warm, 0.2–3 s cold |
| Billing | Instance-hour | Instance-hour | vCPU-s + GB-s per task | Instance-hour + $0.10/hr cluster | Request + GB-ms |
| Idle cost | Full | Full | Zero when no tasks | Control plane + nodes | Zero |
| Max run time | Unbounded | Unbounded | Unbounded | Unbounded | 15 min |
| Statefulness | Anything | Anything | Ephemeral only | StatefulSets | Stateless |
| Long-lived connections | Yes | Yes | Yes | Yes | Awkward |
| Operational surface | Large | Medium-large | Small | Largest | Smallest |
| Ecosystem portability | Low | Medium | Medium | Highest | Lowest |
| Best at | Legacy, special kernels, GPUs | High density, Spot optimization | The default for services | Multi-cloud, a platform team, an existing K8s investment | Event-driven, spiky, glue |
How to actually decide:
flowchart TD
A{"Event-driven or very spiky,<br/>each unit < 15 min?"} -->|yes| L[Lambda]
A -->|no| B{"Needs a specific kernel,<br/>GPU, or licence tied to a host?"}
B -->|yes| E[EC2 or ECS on EC2]
B -->|no| C{"Does the org already run<br/>Kubernetes with a platform team?"}
C -->|yes| K[EKS]
C -->|no| D{"Sustained high density where<br/>per-task billing costs more<br/>than instances you'd bin-pack?"}
D -->|yes| EC2T[ECS on EC2 + Spot]
D -->|no| F[ECS Fargate]Say this in an interview, and mean it: "Fargate unless there is a reason. EKS is a platform decision, not a service decision — if you are choosing it for one service, you are choosing it for the wrong reason."
Architecture
High-level
flowchart TB
U((Users)) --> R53[Route 53 alias] --> ALB[Public ALB]
ALB --> SVC["ECS Service · order-service<br/>desired 4, min 2, max 20"]
SVC --> S3[(S3)]
SVC --> CW[CloudWatch]Detailed
flowchart TB
U((Users)) --> IGW[Internet Gateway]
subgraph VPC["VPC 10.0.0.0/16"]
subgraph PUBA["public-a"]
ALBA[ALB node]
NATA[NAT GW]
end
subgraph PUBB["public-b"]
ALBB[ALB node]
end
subgraph PRIA["private-a 10.0.16.0/20"]
T1["Task ENI 10.0.16.7<br/>app-sg"]
T2["Task ENI 10.0.16.8"]
end
subgraph PRIB["private-b 10.0.32.0/20"]
T3["Task ENI 10.0.32.11"]
T4["Task ENI 10.0.32.12"]
end
GWE[S3 gateway endpoint]
end
IGW --> ALBA & ALBB
ALBA --> T1 & T3
ALBB --> T2 & T4
T1 & T2 & T3 & T4 --> GWE --> S3[(S3)]
T1 & T2 & T3 & T4 -->|"ECR pull, logs"| NATA
ECR[(ECR)] -.-> NATA
TASKROLE["task role"] -.-> T1 & T2 & T3 & T4
EXECROLE["execution role"] -.->|"image pull, log stream"| T1Each task has its own ENI and its own private IP (awsvpc network mode, mandatory on Fargate). That means Security Groups apply per task — and it means each task consumes a subnet IP address. This is exactly why Day 2 sized the private subnets at /20: 4,091 addresses, not 251.
Request flow
sequenceDiagram
autonumber
participant C as Client
participant R as Route 53
participant A as ALB
participant T as Task (Fargate)
participant S as S3
C->>R: resolve api.acme.com
R-->>C: alias → ALB addresses
C->>A: TLS ClientHello :443
Note over A: TLS TERMINATES HERE.<br/>ACM certificate on the listener.
C->>A: GET /api/documents/123
A->>A: rule match → orders-tg<br/>algorithm picks a healthy target
A->>T: HTTP/1.1 GET :8080<br/>X-Forwarded-For: <client IP><br/>X-Forwarded-Proto: https
Note over T: The app sees plain HTTP.<br/>Client IP only via XFF.
T->>S: GetObject (task role credentials)
S-->>T: bytes
T-->>A: 200
A-->>C: 200 over the existing TLS connectionTwo consequences developers hit constantly:
request.isSecure()is false inside the app, and generated absolute URLs come out ashttp://. Fix:server.forward-headers-strategy=frameworkin Spring Boot, so it honoursX-Forwarded-Proto.- The remote address is the ALB node, not the client. Rate limiting or geo-logging on remote address will bucket your entire user base together. Use the leftmost
X-Forwarded-Forentry — and only trust it because the ALB set it.
Failure flow
flowchart TB
C[Client] --> ALB[ALB]
ALB --> T1["Task A ✅"]
ALB -->|"health check fails ×2"| T2["Task B ❌ hung"]
ALB -.->|"removed from rotation<br/>after 60 s"| T2
T2 -.->|"ECS sees unhealthy,<br/>stops and replaces"| T5["Task B' (new)"]
ALB --> T3["Task C ✅ in AZ b"]
AZ["AZ a impaired"] -.-> T1
ALB -->|"all AZ-a targets unhealthy →<br/>all traffic to AZ b"| T3
ALB -->|"NO healthy targets anywhere"| E503["HTTP 503<br/>Service Unavailable"]How It Works Internally
An ALB request, and where each status code comes from
flowchart TD
REQ[Request arrives at an ALB node] --> LIS{"Listener match<br/>on port?"}
LIS -->|no| DROP[Connection refused]
LIS -->|yes| TLS["TLS handshake<br/>(HTTPS listener, ACM cert)"]
TLS --> RULE{"Rule match by priority"}
RULE --> TG[Target group selected]
TG --> HEALTHY{"Any healthy targets?"}
HEALTHY -->|no| E503["503 Service Unavailable"]
HEALTHY -->|yes| PICK["Pick a target<br/>(round robin, or LOR,<br/>or weighted random)"]
PICK --> CONN{"Connection to target<br/>established?"}
CONN -->|refused / reset / malformed| E502["502 Bad Gateway"]
CONN -->|yes| WAIT{"Response within<br/>idle timeout (60 s)?"}
WAIT -->|no| E504["504 Gateway Timeout"]
WAIT -->|yes| OK[Response relayed to client]This diagram is a diagnostic tool. Memorize the three codes:
| Code | The ALB is telling you | Look at |
|---|---|---|
| 503 | "I have no healthy targets." | Health checks, task/instance count, the ASG or ECS service, whether your readiness probe checks a shared dependency |
| 502 | "A target answered, and the answer was broken." | The container crashed mid-request, closed the connection, or returned a malformed response. Check exit codes (137 = OOM-killed), stack traces, and whether draining is misconfigured. |
| 504 | "A target accepted the request and never answered in time." | Your application is slow, not dead. Missing server-side timeouts, a blocked thread pool, a hung downstream. Raising the ALB idle timeout hides it; fixing the timeout chain solves it. |
The single most useful ALB metric pairing:
HTTPCode_ELB_5XX_Count(the ALB's own errors — 502/503/504) versusHTTPCode_Target_5XX_Count(your application's own 500s). If the first is high and the second is zero, your code never ran. That fact alone redirects an investigation.
ECS task placement and deployment
sequenceDiagram
autonumber
participant API as UpdateService
participant SCH as ECS scheduler
participant FG as Fargate infrastructure
participant ECR
participant TG as Target group
participant OLD as Old tasks
API->>SCH: desired = revision 2
SCH->>SCH: honour minimumHealthyPercent / maximumPercent
SCH->>FG: provision micro-VM, attach ENI in your subnet
Note over FG: uses the EXECUTION ROLE
FG->>ECR: pull image by digest
FG->>FG: start container; app begins boot
FG->>TG: register target
TG->>FG: health checks... (healthyThreshold × interval)
TG-->>SCH: target healthy
SCH->>TG: deregister an old target
TG->>TG: drain for deregistrationDelay
SCH->>OLD: SIGTERM
Note over OLD: graceful shutdown finishes in-flight work
OLD-->>SCH: exited
SCH->>SCH: repeat until all tasks are revision 2Where the time goes, for a JVM service:
provision micro-VM + ENI ~10-20 s
image pull (cached layers) ~5-15 s
JVM start + Spring context ~20-40 s
health checks (2 × 15 s) ~30 s
─────────
first new task serving traffic ~65-105 s
× N tasks, batched the full deployThis is the number to quote when someone asks "why does a deploy take six minutes?" — and the number to attack if it matters (smaller images, spring-context-indexer, fewer health-check datapoints, AOT/CDS).
Code
Production-shaped application.yml
server:
port: 8080
shutdown: graceful
# Honour X-Forwarded-* from the ALB so isSecure() and generated URLs are right.
forward-headers-strategy: framework
tomcat:
# Cap the thread pool deliberately: it is your concurrency limit and
# therefore the thing your autoscaling signal should track.
threads:
max: 100
connection-timeout: 5s
spring:
lifecycle:
# Must be LESS than ECS stopTimeout, or SIGKILL interrupts the drain.
timeout-per-shutdown-phase: 25s
management:
endpoints.web.exposure.include: health,info,metrics,prometheus
endpoint.health:
probes.enabled: true # gives /health/liveness and /health/readiness
group:
readiness:
# Deliberately NOT including S3 or the database.
# A shared-dependency outage must not take every task out of rotation.
include: ping
liveness:
include: pingGraceful shutdown that actually drains
package com.acme.lifecycle;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import org.springframework.boot.availability.AvailabilityChangeEvent;
import org.springframework.boot.availability.ReadinessState;
import org.springframework.context.ApplicationListener;
import org.springframework.context.event.ContextClosedEvent;
import org.springframework.stereotype.Component;
/**
* On SIGTERM, Spring closes the context. We flip readiness to REFUSING_TRAFFIC
* first and pause briefly, so the ALB's next health check fails BEFORE the
* connector stops accepting. Without this pause the ALB can still be sending
* requests to a socket that is already closing -> a burst of 502s per deploy.
*/
@Component
class DrainOnShutdown implements ApplicationListener<ContextClosedEvent> {
private static final Logger log = LoggerFactory.getLogger(DrainOnShutdown.class);
private final org.springframework.context.ApplicationEventPublisher publisher;
DrainOnShutdown(org.springframework.context.ApplicationEventPublisher publisher) {
this.publisher = publisher;
}
@Override
public void onApplicationEvent(ContextClosedEvent event) {
log.info("SIGTERM received: refusing traffic, draining");
AvailabilityChangeEvent.publish(publisher, this, ReadinessState.REFUSING_TRAFFIC);
try {
// ~1 health-check interval. Tune to your interval, not to a magic number.
Thread.sleep(java.time.Duration.ofSeconds(15).toMillis());
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}
}A readiness probe that is honest about its dependencies
package com.acme.health;
import org.springframework.boot.actuate.health.Health;
import org.springframework.boot.actuate.health.HealthIndicator;
import org.springframework.stereotype.Component;
import software.amazon.awssdk.services.s3.S3Client;
import software.amazon.awssdk.services.s3.model.HeadBucketRequest;
/**
* Registered as a health indicator, but DELIBERATELY EXCLUDED from the
* readiness group in application.yml.
*
* Why: if S3 is degraded, every task reports DOWN, the ALB removes every
* target, and a partial failure becomes a total outage — plus ECS starts
* replacing a fleet whose replacements will also fail.
*
* So: surface it on /actuator/health for humans and alarms, and keep it out
* of the signal the load balancer acts on.
*/
@Component("s3")
class S3HealthIndicator implements HealthIndicator {
private final S3Client s3;
private final String bucket;
S3HealthIndicator(S3Client s3, com.acme.config.AppProperties props) {
this.s3 = s3;
this.bucket = props.bucket();
}
@Override
public Health health() {
try {
s3.headBucket(HeadBucketRequest.builder().bucket(bucket).build());
return Health.up().withDetail("bucket", bucket).build();
} catch (Exception e) {
return Health.down().withDetail("bucket", bucket)
.withDetail("error", e.getClass().getSimpleName()).build();
}
}
}Task definition — with the two roles visibly distinct
{
"family": "order-service",
"networkMode": "awsvpc",
"requiresCompatibilities": ["FARGATE"],
"cpu": "1024",
"memory": "2048",
"runtimePlatform": { "cpuArchitecture": "ARM64", "operatingSystemFamily": "LINUX" },
"executionRoleArn": "arn:aws:iam::123456789012:role/ecsTaskExecutionRole",
"taskRoleArn": "arn:aws:iam::123456789012:role/order-service-task-role",
"containerDefinitions": [{
"name": "app",
"image": "123456789012.dkr.ecr.eu-west-1.amazonaws.com/order-service@sha256:9f2c...",
"portMappings": [{ "containerPort": 8080, "protocol": "tcp" }],
"essential": true,
"environment": [
{ "name": "COURSE_BUCKET", "value": "aws-course-123456789012-day4" },
{ "name": "SPRING_PROFILES_ACTIVE", "value": "prod" }
],
"stopTimeout": 45,
"healthCheck": {
"command": ["CMD-SHELL", "curl -fsS http://localhost:8080/actuator/health/liveness || exit 1"],
"interval": 15, "timeout": 5, "retries": 3, "startPeriod": 60
},
"logConfiguration": {
"logDriver": "awslogs",
"options": {
"awslogs-group": "/ecs/order-service",
"awslogs-region": "eu-west-1",
"awslogs-stream-prefix": "app"
}
}
}]
}Notes that generalize:
- Image by digest, not
:latest. A mutable tag makes "which code is running?" unanswerable, and makes a rollback a guess. - ARM64 + Graviton, for the same 20%-ish saving as Day 3, with no code change.
stopTimeout: 45exceeds the 25 s Spring drain plus the 15 s readiness pause. The chain is consistent by construction.startPeriod: 60stops the container health check from failing during JVM boot.
Rolling deploy with zero dropped requests
aws ecs create-service \
--cluster course-cluster --service-name order-service \
--task-definition order-service:7 --desired-count 4 --launch-type FARGATE \
--network-configuration "awsvpcConfiguration={subnets=[$PRI_A,$PRI_B],securityGroups=[$APP_SG],assignPublicIp=DISABLED}" \
--load-balancers "targetGroupArn=$TG_ARN,containerName=app,containerPort=8080" \
--health-check-grace-period-seconds 90 \
--deployment-configuration '{
"minimumHealthyPercent": 100,
"maximumPercent": 200,
"deploymentCircuitBreaker": { "enable": true, "rollback": true }
}'
# Deregistration delay tuned to real request duration, not the 300 s default.
aws elbv2 modify-target-group-attributes --target-group-arn "$TG_ARN" \
--attributes Key=deregistration_delay.timeout_seconds,Value=30 \
Key=load_balancing.algorithm.type,Value=least_outstanding_requests
least_outstanding_requestsbeats round robin whenever request costs vary — it stops a slow request from being queued behind another slow request on the same target. For heterogeneous workloads it is a free latency win.
Hands-on
Lab 5 — Containerize and deploy to ECS Fargate behind an ALB (30 min)
Objective: the Day 3 service running as 4 Fargate tasks across 2 AZs behind an ALB, autoscaling on requests per target, with a rolling deploy that drops zero requests.
Prerequisites: Day 2 VPC (NAT Gateway running, for the ECR pull), Docker running locally, the Day 3 jar and bucket.
⚠️ Billable: ALB (
$0.0225/hr + LCUs), Fargate ($0.04/vCPU-hr + ~$0.0044/GB-hr), NAT Gateway, ECR storage. A 2-hour lab is well under $1. The ALB is the one to remember to delete — ~$16/month idle.
Architecture: the Detailed diagram above.
Steps:
- Create the ECR repository with
scanOnPushandIMMUTABLEtags; add the untagged-expiry lifecycle policy. - Build the layered image from the Dockerfile above; push by tag
1.0.0; record the digest. - Create the execution role (
AmazonECSTaskExecutionRolePolicy) and the task role (the narrow S3 + CloudWatch policy from Day 3). Two roles, deliberately. - Create the ALB in the public subnets with
alb-sg; create an HTTP listener on 80 (HTTPS needs a domain and an ACM certificate — noted as the production requirement, skipped here). - Create the target group: target type
ip(required forawsvpc), port 8080, health check path/actuator/health/readiness, interval 15 s, healthy threshold 2. - Register the task definition above. Create the service with the
create-servicecall above. - Attach the
ALBRequestCountPerTargettarget-tracking policy (target 1000).
Verification — three tests, each proving one thing:
ALB_DNS=$(aws elbv2 describe-load-balancers --names course-alb \
--query 'LoadBalancers[0].DNSName' --output text)
# (a) It serves, and tasks are spread across AZs
curl -s "http://${ALB_DNS}/actuator/health"
aws ecs describe-tasks --cluster course-cluster \
--tasks $(aws ecs list-tasks --cluster course-cluster --query 'taskArns[]' --output text) \
--query 'tasks[].availabilityZone' --output text
# → expect BOTH eu-west-1a and eu-west-1b# (b) THE ZERO-DOWNTIME TEST — the point of the lab.
# Start a continuous request loop, then deploy, and count failures.
( while true; do
code=$(curl -s -o /dev/null -w '%{http_code}' --max-time 5 "http://${ALB_DNS}/actuator/health")
[ "$code" = "200" ] || echo "$(date +%T) FAILED: $code"
sleep 0.2
done ) > /tmp/deploy-watch.log 2>&1 &
WATCH=$!
# Force a new deployment (same image — we are testing the mechanics)
aws ecs update-service --cluster course-cluster --service order-service --force-new-deployment
aws ecs wait services-stable --cluster course-cluster --services order-service
kill $WATCH
echo "Failures during deploy: $(wc -l < /tmp/deploy-watch.log)"
# ✅ Success criterion: 0If it is not zero, do not move on — work the chain: is minimumHealthyPercent 100? Is graceful shutdown on? Is stopTimeout > the drain? Is the deregistration delay ≥ your longest request? That debugging is the lab.
# (c) Autoscaling responds to load, not to CPU
hey -z 3m -c 200 "http://${ALB_DNS}/actuator/health" &
watch -n 15 'aws ecs describe-services --cluster course-cluster \
--services order-service --query "services[0].{desired:desiredCount,running:runningCount}"'
# Watch desiredCount climb. Note how long it takes — that is scale-out lag, measured.Then break it deliberately, so you recognize each signature later:
# 502: make the container die mid-request
aws ecs stop-task --cluster course-cluster --task "$SOME_TASK_ARN" # while (b)'s loop runs
# 503: remove every healthy target
aws ecs update-service --cluster course-cluster --service order-service --desired-count 0
curl -si "http://${ALB_DNS}/" | head -1 # → HTTP/1.1 503 Service Unavailable
aws ecs update-service --cluster course-cluster --service order-service --desired-count 4
# 504: be slow rather than dead (add a sleep endpoint, or drop the idle timeout)
aws elbv2 modify-load-balancer-attributes --load-balancer-arn "$ALB_ARN" \
--attributes Key=idle_timeout.timeout_seconds,Value=2
curl -si "http://${ALB_DNS}/slow?ms=5000" | head -1 # → HTTP/1.1 504 Gateway TimeoutCleanup — in dependency order:
aws ecs update-service --cluster course-cluster --service order-service --desired-count 0
aws ecs delete-service --cluster course-cluster --service order-service --force
aws elbv2 delete-listener --listener-arn "$LISTENER_ARN"
aws elbv2 delete-load-balancer --load-balancer-arn "$ALB_ARN" # ← the expensive one
aws elbv2 delete-target-group --target-group-arn "$TG_ARN"
aws ecs delete-cluster --cluster course-cluster
aws ecr delete-repository --repository-name order-service --force
# Keep the VPC for Day 5/6. Delete the NAT Gateway + EIP if you are stopping today.Common errors:
| Symptom | Cause | Fix |
|---|---|---|
CannotPullContainerError | Execution role lacks ECR permissions, or no route to ECR (no NAT, no endpoints) | Check the role first, then the route |
| Task starts, no logs in CloudWatch | Execution role lacks logs:CreateLogStream / PutLogEvents | Add them; the log group must exist or be creatable |
Task runs, app gets AccessDenied on S3 | Task role is wrong — a different role entirely | Two roles. Read the error to know which. |
Target stuck unhealthy, task looks fine | Health check path/port wrong, app-sg doesn't allow alb-sg on 8080, or grace period < boot time | Curl the path from inside the task |
| Exit code 137 | OOM-killed by the kernel. Heap + non-heap exceeded the task memory limit. | Lower MaxRAMPercentage, or raise task memory |
| Exit code 143 | SIGTERM, i.e. a normal stop. Not an error. | Set SuccessExitStatus=143 where relevant |
RESOURCE:ENI placement failure | The private subnet ran out of IP addresses | This is Day 2's subnet-sizing lesson arriving in production |
| Deploy hangs at "in progress" for 10+ min | New tasks never go healthy; circuit breaker not enabled | Enable the circuit breaker so it rolls back instead of hanging |
Production relevance: this is the shape of a real ECS service. Days 6–13 add a database, a queue, security and observability to exactly this deployment.
Production Considerations
| Concern | Practice |
|---|---|
| Deployment safety | Circuit breaker with rollback, always. Blue/green via CodeDeploy when the change is risky. |
| Immutable images | Deploy by digest. IMMUTABLE tags in ECR. :latest in production is how you lose the ability to answer "what is running?" |
| Two roles, two purposes | Execution role shared and boring; task role per service and narrow. Never reuse one task role across services — it silently grants every service the union of all permissions. |
| Subnet capacity | Each task consumes an IP. desired × 2 during deploys, plus ALB nodes, plus headroom. |
| Capacity providers | FARGATE_SPOT for interruptible work (workers, batch) at a large discount; keep a FARGATE base for the serving tier. |
| Right-sizing | Fargate bills for what you request, not what you use. Over-requesting CPU/memory is pure waste; Container Insights shows the real numbers. |
| Image size | Affects pull time on every task launch, which affects scale-out lag. Alpine/distroless + layering. |
| Log volume | The awslogs driver ships everything to CloudWatch. Day 3's retention and cost lessons apply immediately, multiplied by task count. |
| Health check design | Readiness excludes shared dependencies. Liveness is about this process. Alarm on dependencies separately. |
| ALB access logs | Off by default. Turn them on to S3 — they are the only record of what the LB saw, and indispensable for 502/504 forensics. |
Failure Scenarios
S7 — "Every deploy drops ~2% of requests for 30 seconds"
A team deploys 6 times a day. Each deploy produces a brief spike of 502s. Everyone has learned to ignore it.
| Question | Answer |
|---|---|
| What can fail? | The handoff between "task is being stopped" and "ALB stopped sending it traffic" |
| What happens? | The container receives SIGTERM and closes its listening socket while the ALB is still routing to it. In-flight requests are cut; new ones hit a closed socket → 502. |
| Recovery? | Automatic — the deploy completes and errors stop. Which is precisely why it never gets fixed. |
| How quickly? | 20–60 s per deploy, 6 times a day, every day |
| Can the operation happen twice? | Yes — clients retry the failed requests, so non-idempotent operations can double. This is the part nobody notices. |
| Can data be lost? | Yes: requests killed mid-write with no retry |
The four-part fix, which must be consistent:
1. server.shutdown=graceful app stops accepting, finishes in-flight
2. spring.lifecycle.timeout=25s bounded drain
3. ECS stopTimeout=45s SIGKILL arrives AFTER the drain, not during
4. deregistration_delay=30s ALB stops routing BEFORE SIGTERM lands
+ the readiness pause in DrainOnShutdown the health check fails firstThe senior observation: a "2% blip on deploy" is not cosmetic. Multiply by deploy frequency and request volume and it is thousands of failed operations a week, some of them retried into duplicates. Zero is achievable, and the test in Lab 5(b) proves it.
S8 — Black Friday: 8× traffic over 4 minutes
Autoscaling is configured correctly. Users still see errors for the first six minutes.
| Question | Answer |
|---|---|
| What can fail? | Nothing is misconfigured. The physics of scale-out lag. |
| What happens? | Metric publication (60 s) + alarm evaluation (60–120 s) + task launch (30–60 s) + health checks (30 s) ≈ 4 min before the first extra task serves. Existing tasks saturate, queue, and time out. |
| Recovery? | Automatic, once capacity catches up |
| How quickly? | 4–8 minutes of degradation |
| Can the operation happen twice? | Yes, and worse — clients retry timed-out requests, adding load exactly when there is least capacity. A retry storm (Day 12). |
| Can data be lost? | Yes, for non-retried requests |
What actually works, in order of effectiveness:
- Scheduled scaling. You knew about Black Friday. Pre-scale an hour ahead. This is the single highest-value action and it is free.
- Headroom. Target 50–60% utilization, not 80%. The unused capacity is the spike absorber.
- Faster launches. Smaller images, fewer health-check datapoints, shorter grace period, Java AOT/CDS. Each second is capacity earned earlier.
- Load shedding. Reject excess with
429fast rather than queueing everything into timeouts. Fast failure beats slow failure — it lets clients back off instead of piling on. - Queue the work. The structural answer: return
202and process asynchronously so the spike lands in a queue rather than in your thread pool. That is Day 8.
The interview-grade version of this answer: "Autoscaling is a cost optimization, not an availability mechanism. Availability under a spike comes from headroom, pre-scaling, and shedding."
A8 — Autoscaling on CPU for an I/O-bound service
A service spends 80% of its time waiting on a downstream API. Latency climbs, queues build, and the ASG never scales.
| Question | Answer |
|---|---|
| What can fail? | The scaling signal is uncorrelated with the constrained resource |
| What happens? | Tomcat's 100 threads are all blocked on socket reads. CPU sits at 25%. The CPU target-tracking policy is satisfied. New requests queue in the accept backlog and time out. |
| Recovery? | None automatic. Manual intervention. |
| Discriminating signal | tomcat_threads_busy at max while CPUUtilization is low; ALBRequestCountPerTarget climbing; TargetResponseTime climbing with no CPU movement |
| Fix | Scale on ALBRequestCountPerTarget or thread-pool saturation. Add timeouts and a circuit breaker so a slow downstream cannot hold your threads (Day 12). |
Security
| Question | Answer |
|---|---|
| Who can access this? | The ALB, from the Internet on 443 only. Tasks: only from alb-sg on 8080 — each task has its own ENI and its own Security Group membership, so this is enforced per task. |
| What credentials are used? | Two roles. The execution role (platform: ECR, logs, secrets) and the task role (your code). Both temporary, both rotated, neither stored. |
| Where is data encrypted? | In transit: client→ALB via TLS (ACM certificate on the listener). ALB→task is plain HTTP inside the VPC — acceptable for many threat models, but if you need end-to-end encryption, use an HTTPS target group or a service mesh. At rest: ECR images and S3 objects encrypted by default. |
| What if credentials leak? | The task role is narrow (one bucket prefix, no delete) and expires. The execution role can pull images and write logs — unpleasant but limited. This is why they are separate roles. |
ECS-specific controls worth knowing:
readonlyRootFilesystem: trueon the container, withtmpfsmounts where writes are needed. Blocks a large class of post-exploitation.- Non-root user in the image (the Dockerfile above does this). A container escape from root is far worse than from an unprivileged user.
- ECR image scanning on push, plus basic scanning or Inspector for continuous rescanning as new CVEs land against images you already shipped.
- Secrets via
secretsin the task definition (Secrets Manager / SSM), fetched by the execution role and injected as env vars — never baked into the image. Day 6 does this properly. assignPublicIp: DISABLED. A Fargate task in a public subnet with a public IP is directly reachable; there is no reason for it.
Performance
| Lever | Effect |
|---|---|
least_outstanding_requests | Real p99 improvement when request costs vary. Cheap and under-used. |
| Task CPU/memory sizing | Fargate CPU is allocated per task; 0.25 vCPU throttles a JVM badly. 1 vCPU / 2 GB is a sane floor for Spring Boot. |
| Thread pool vs task count | maxThreads × tasks is your concurrency ceiling. Raising threads on an I/O-bound service is often cheaper than adding tasks — until the downstream becomes the bottleneck. |
| Keep-alive | ALB reuses connections to targets. Ensure your app's keep-alive timeout exceeds the ALB's (60 s) or you get races that surface as 502s. |
| Image size and layering | Directly determines pull time, which is scale-out lag |
| JVM start time | CDS / AOT / spring-context-indexer / lazy init. Matters most for scale-out and deploy speed. |
| Graviton (ARM64) | Comparable or better throughput at ~20% lower cost for JVM workloads |
| Cross-AZ hops | ALB→task can cross AZs. It costs ~$0.01/GB each way and adds a millisecond or two. Usually the right trade for availability. |
Cost
Lab 5 for 2 hours (eu-west-1, illustrative — verify current pricing):
| Resource | Rate | 2 hours | Left running 1 month |
|---|---|---|---|
| ALB (hourly) | ~$0.0225/hr | $0.045 | ~$16.40 |
| ALB LCUs | ~$0.008/LCU-hr | ~$0.02 | varies with traffic |
| Fargate, 4 tasks × 1 vCPU / 2 GB | ~$0.04048/vCPU-hr + ~$0.004445/GB-hr | ~$0.40 | ~$146 |
| NAT Gateway (from Day 2) | ~$0.045/hr | $0.09 | ~$32 |
| ECR storage | ~$0.10/GB-mo | ~$0 | ~$0.03 |
| CloudWatch Logs | ~$0.50/GB ingested | depends | depends |
| Total | ~$0.56 | ~$195+ |
Fargate arithmetic worth memorizing. One task at 1 vCPU / 2 GB:
vCPU: 1 × $0.04048 = $0.04048/hr
memory: 2 × $0.004445 = $0.00889/hr
─────────────
~$0.0494/hr ≈ $36/month per taskSo "4 tasks minimum for HA" is ~$144/month before traffic. That number drives real decisions:
- Graviton (ARM64) cuts it by roughly 20%.
- Fargate Spot cuts it by ~70% for interruption-tolerant tasks — excellent for the Day 8 workers, wrong for the serving tier.
- Right-sizing. If Container Insights shows 15% CPU and 40% memory at 1 vCPU / 2 GB, you are paying for a request you do not use. Fargate bills the request, not the usage.
- Compute Savings Plans apply to Fargate: ~20% off for a 1-year commitment on a predictable baseline.
The EC2-vs-Fargate crossover: an m7g.xlarge (4 vCPU, 16 GB) is ~$0.163/hr ≈ $119/month and could host 4 of those tasks, versus ~$144 on Fargate — before you account for the ASG, patching, bin-packing and the capacity you must keep idle for headroom. Fargate is usually cheaper once operational cost is honest, and the crossover only favours EC2 at sustained high density.
Alternatives
| Instead of | You could | Trade-off |
|---|---|---|
| ECS Fargate | EC2 + ASG + jar (Day 3's model) | No Docker to learn; you own AMIs, patching and drift. Slower scale-out. |
| ECS Fargate | ECS on EC2 | Higher density, Spot flexibility, lower cost at scale; you own the hosts |
| ECS | EKS | Kubernetes ecosystem and portability; a control plane fee, upgrade treadmill, and a genuine platform-team requirement |
| ECS | App Runner | Fastest path from container to URL; much less control, fewer knobs |
| ECS | Lambda (Day 5) | Zero idle cost, per-request scaling; 15-minute ceiling, cold starts, stateless-only |
| ALB | NLB | Static IPs, source IP preservation, lower latency; no path/host routing, no Lambda targets |
| ALB | API Gateway (Day 5) | Built-in auth, throttling, keys, caching; higher per-request cost and a 30 s integration timeout |
| Rolling deploy | Blue/green (CodeDeploy) | Instant rollback, pre-traffic hooks, canary shifting; two target groups and more machinery |
| Target tracking | Scheduled scaling | Beats reactive scaling for known events; needs a human to know |
Trade-offs
| Decision | Gain | Cost |
|---|---|---|
| Containers over jars-on-instances | Identical copies, fast launches, no drift | A build pipeline, a registry, image hygiene, a new failure vocabulary |
| Fargate over EC2 | No hosts to own; per-task billing and isolation | Higher per-unit price at density; no GPUs or special kernels |
min=100 / max=200 | Never lose capacity mid-deploy | Temporarily double the task cost during a deploy |
| Short deregistration delay | Fast deploys | Cutting long-running requests if set below your true p99 |
| Readiness excludes dependencies | A dependency outage stays partial | A task that genuinely cannot work still receives traffic and returns errors |
| Scaling on requests, not CPU | Capacity tracks demand | You must pick and maintain a target value; CPU is the lazy default for a reason |
| Multi-AZ tasks | Survives an AZ | Cross-AZ data transfer charges and a millisecond or two |
| Autoscaling at all | Pay for what you need | Scale-out lag means it never saves you from a spike |
Common Mistakes
- Health check type
EC2instead ofELBon an ASG. A hung JVM stays "healthy" forever. - Health check grace period shorter than boot time. Infinite launch/terminate loop, and the cause is invisible unless you read ASG activity history.
- Confusing task role and execution role. Learn the four symptoms; it saves an hour each time.
:latestin production. You cannot answer "what is running?" and cannot roll back deterministically.- Readiness probe that checks the database. One shared dependency blip removes every target and turns a partial failure into a total one.
- Leaving deregistration delay at 300 s. Deploys crawl and nobody knows why.
- No graceful shutdown. The 2%-per-deploy error rate everyone learns to ignore.
- Fixed
-Xmxin a container. Exit code 137, no stack trace, three hours of confusion. - Scaling on CPU for an I/O-bound service. It will never scale when it needs to.
- Undersized private subnets.
RESOURCE:ENIplacement failures during a deploy, because a deploy transiently needs 2× the IPs. - Sticky sessions to "fix" a state problem. Now scale-in, deploys and AZ failure all break instead.
- No ALB access logs. The only forensic record of what the load balancer saw, off by default.
Interview Questions
2–3 YOE
- What is an Auto Scaling Group, and what does "desired capacity" mean?
- What is the difference between vertical and horizontal scaling?
- What are the four levels of an ALB's routing — listener to target?
- What is a container image, and why does it make deployments more reliable than a setup script?
- What is the difference between an ECS task and an ECS service?
- Where does TLS terminate in an ALB → ECS architecture, and what does the application see?
4–5 YOE
- Your ALB returns
503. Your application logs show nothing at all. What happened, and where do you look? - Your ALB returns
502only during deployments. Diagnose it and list the settings you would change. - Your task exits with code 137 and there is no Java stack trace. What happened?
- Why is CPU usually the wrong autoscaling signal for a request-driven service, and what would you use instead?
- Explain the difference between an ECS task role and an execution role, and give one symptom of getting each wrong.
- Your health check hits an endpoint that verifies the database. Argue why that is dangerous.
Senior-Level Questions
6–7 YOE
- Design the deployment strategy for a service where a bad deploy costs real money. What do you use, what do you measure, and what triggers an automatic rollback?
- Your service must handle an 8× spike arriving in 3 minutes, and the budget forbids running 8× capacity all day. What is your actual plan?
- A team wants to move from ECS to EKS. What questions do you ask, and under what conditions do you support it?
- How do you keep 40 ECS services' IAM permissions least-privilege without a central bottleneck?
- Fargate costs 30% more than the equivalent EC2 fleet on paper. Make the case for Fargate anyway — and then the case against.
8–10 YOE
- Your platform is "multi-AZ" and an AZ impairment still caused a full outage. Give three plausible causes at this layer (ALB, ASG/ECS, dependencies) and say how you would detect each before it happens.
- Explain how a health check can turn a 20% failure into a 100% outage, and design the probe strategy that prevents it. Where does this reasoning break down?
- You inherit a service that deploys 8 times a day with a known 2% error blip each time. Quantify the real business impact, then describe the fix and how you would prove it worked.
- When is horizontal scaling the wrong answer, and what would you do instead? Give two concrete situations.
- Design the compute strategy for a company with 60 services of mixed shapes: request/response, batch, event-driven, and one that needs GPUs. How many platforms do you run, and how do you defend that number?
Day-End Revision
The five sentences
- An ASG keeps N instances alive across AZs; use
ELBhealth checks and a grace period longer than your boot time. - An ALB routes listener → rule → target group → target, terminates TLS at the listener, and its three 5xx codes each name a different failure: 503 no healthy targets, 502 broken response, 504 too slow.
- Container images make every copy identical;
MaxRAMPercentagemakes the JVM respect the container's memory limit. - ECS has two roles — execution (platform: pull image, write logs) and task (your code's AWS permissions) — and confusing them produces two different errors.
- Zero-downtime deployment is four settings agreeing: graceful shutdown, drain timeout,
stopTimeout, deregistration delay.
The diagram to redraw from memory: the ALB request flow with the three 5xx branches.
The numbers
| ALB default idle timeout | 60 s |
| Target group deregistration delay default | 300 s (usually too long) |
| ALB health check defaults | 30 s interval, 5 s timeout, healthy 5, unhealthy 2 |
ECS stopTimeout | default 30 s, max 120 s |
| Fargate 1 vCPU / 2 GB | ~$0.049/hr ≈ ~$36/month |
| ALB idle | ~$16/month |
| Scale-out lag, JVM | ~4–5 minutes end to end |
| Exit code 137 / 143 | OOM-killed / SIGTERM |
Today's trap: "Autoscaling means we can handle any spike." Autoscaling handles trends. Spikes are handled by headroom, scheduled scaling, and load shedding.
Tomorrow needs: the ALB (API Gateway is compared against it), the container/task mental model (Lambda is compared against it), the two-role distinction (Lambda has an execution role), and today's scale-out lag number — Lambda's answer to it is the reason Day 5 exists.
Mini Assignment
Time: 45–60 minutes. Capstone contribution: the Order Service, containerized and autoscaling.
- In
day-04/, containerize the Day 3 service with the layered Dockerfile. Confirm that a source-only change rebuilds only the application layer (comparedocker historybefore and after). - Add to the service:
GET /slow?ms=N— sleeps, so you can produce 504s on demand;GET /crash— exits the JVM, so you can produce 502s on demand;- a Micrometer gauge exposing
tomcat_threads_busyso you can see thread-pool saturation.
- Deploy to ECS Fargate behind an ALB with 2 tasks across 2 AZs,
ALBRequestCountPerTargetautoscaling, and the circuit breaker enabled. - Produce, capture and explain each of
502,503and504— one screenshot orcurl -sioutput each, with the CloudWatch metric that confirms it (HTTPCode_ELB_5XX_CountvsHTTPCode_Target_5XX_Count). - Run the zero-downtime test from Lab 5(b). If it is not zero, fix it and document what was wrong.
- Measure scale-out lag: start a load test, record the timestamp of the first 429/latency rise and the timestamp of the first new task passing health checks. Report the delta.
- In
day-04/NOTES.md:- Your measured scale-out lag, and what you would change to halve it.
- Should the readiness probe check S3? Your answer from Day 3, revisited — and did it change?
- Monthly cost of this service at 2 tasks, at 20 tasks, and with Graviton + Spot for a worker tier.
- Which of EC2 / ECS-on-EC2 / Fargate / EKS / Lambda you would choose for this service, and the one fact that would change your mind.
- Tear down the ALB and cluster. Commit.
Success criterion: the deploy test reports 0 failures, and you can explain each of the four settings that made it zero.
AWS Documentation
- Auto Scaling Groups · Target tracking policies
- Application Load Balancer guide · Target groups and health checks · Troubleshooting 5xx
- ECS Developer Guide · Task definitions · Task IAM role · Execution role
- Fargate · Deployment circuit breaker · Service auto scaling
- ECR · Lifecycle policies
- Spring Boot container images · Graceful shutdown
Previous: Day 3 — Compute, Storage and Seeing What Happened · Next: Day 5 — Serverless Compute and the API Layer