Day 5 — Serverless Compute and the API Layer
3h 10m · Phase 2 of 4 · Curriculum
Learning Objectives
By the end of today you can:
- Describe a Lambda cold start in terms of what AWS actually does — microVM, init phase, invoke phase, freeze, reuse, reap — not "it warms up."
- Explain why memory is the CPU dial, and size a function accordingly.
- Predict what happens to a request when a Lambda is throttled, differently for synchronous, asynchronous and poll-based invocation.
- Choose between API Gateway REST, HTTP and WebSocket APIs, and between API Gateway and an ALB, with cost arithmetic.
- Wire an authorizer and a throttle, and prove both work.
- Say when the answer to "should this be a Lambda?" is no — and defend it.
The sentence you should be able to say tonight: "An asynchronous Lambda invocation will be retried by the service, which means a non-idempotent handler will produce duplicate side effects as a matter of routine."
Prerequisites
- Day 1 — roles, the credential chain.
- Day 2 — subnets, NAT, VPC endpoints.
- Day 3 — S3, CloudWatch, structured logs.
- Day 4 — the ALB, scale-out lag, the compute decision matrix. Today answers the question Day 4's matrix left open.
Why Does This Exist?
Day 4 left you with a measured number: roughly four to five minutes from a traffic spike to a new instance serving traffic. And a bill: ~$36/month per Fargate task, running whether or not anyone calls you.
Both are consequences of the same thing — you are renting a long-lived container and filling it with requests. That is the right model for a service handling steady traffic. It is a poor model for three very common shapes of work:
| Shape | Why containers fit badly |
|---|---|
| Spiky | 0 requests for an hour, then 5,000 in a minute. You pay for peak capacity all day, or you are too slow. |
| Event-driven glue | "When an object lands in S3, validate it." A container idling 23 hours a day to run for 40 minutes. |
| Rare but latency-tolerant | A nightly report, a webhook receiver, a scheduled reconciliation. |
Lambda inverts the model: you rent an execution of your code, per invocation, per millisecond. AWS owns the container, the scaling, the placement, the patching, and the idle time. Scale-out is measured in milliseconds to a couple of seconds, not minutes, and idle costs exactly zero.
You pay for that with a set of hard constraints — 15 minutes, no persistent state, a cold-start tax, and a stateless programming model — and the whole skill is knowing when the trade is worth it.
Then there is the second half of today. Once your compute is a function rather than a server, you need something in front of it that does what a web framework used to: routing, authentication, rate limiting, request validation, throttling. That is API Gateway — and it is useful in front of containers too.
Beginner Explanation
Lambda in one picture
flowchart LR
E1[HTTP request] --> L
E2[S3 object created] --> L
E3[SQS message] --> L
E4[Schedule: every 5 min] --> L
L["Lambda function<br/>your handler code"] --> R[Result / side effect]You upload a function. You tell AWS what should trigger it. AWS runs it — once per event, as many in parallel as needed — and bills you for the milliseconds it ran and the memory you configured.
There is no server to size, patch, scale or pay for while idle. There is also nowhere to keep state between invocations, no way to run longer than 15 minutes, and a small delay the first time a new copy starts.
API Gateway in one picture
flowchart LR
C[Client] --> AG["API Gateway<br/>• route<br/>• authenticate<br/>• validate<br/>• throttle<br/>• cache<br/>• log"]
AG --> L[Lambda]
AG --> ALB[ALB → ECS]
AG --> AWS[AWS service directly]
AG --> HTTP[Any HTTP endpoint]API Gateway is a managed front door. It does the cross-cutting HTTP concerns before your code runs, which is exactly the set of things you would otherwise implement in every service and get subtly wrong in some of them.
Core Concepts
1. The execution environment — what actually happens
This is the diagram Lambda is misunderstood without. It is one of the eleven mandatory internal-working diagrams.
flowchart TB
INV[Invocation arrives] --> WARM{"Is a warm execution<br/>environment free?"}
WARM -->|yes| HANDLER
WARM -->|"no — COLD START"| MV["1 · Create a Firecracker microVM<br/>(AWS-managed, ~ms)"]
MV --> DL["2 · Download your code<br/>(zip from S3-backed store, or image layers)"]
DL --> RT["3 · Start the runtime<br/>(JVM boot, Node bootstrap, ...)"]
RT --> INIT["4 · INIT PHASE<br/>run static initializers, constructors,<br/>Spring context, SDK client creation"]
INIT --> HANDLER
HANDLER["5 · INVOKE PHASE<br/>your handler method runs"]
HANDLER --> RESP[6 · Response returned]
RESP --> FREEZE["7 · Environment FROZEN<br/>memory preserved, no CPU,<br/>background threads stop"]
FREEZE -->|"next invocation<br/>within minutes"| THAW["Thawed → straight to step 5"]
FREEZE -->|"idle too long, or a deploy,<br/>or AWS reclaims it"| REAP["Environment destroyed"]
THAW --> HANDLEREverything that matters about Lambda falls out of this diagram.
Cold start vs warm invocation
| Cold | Warm | |
|---|---|---|
| Steps run | 1 → 6 | 5 → 6 only |
| Java + plain handler | ~300 ms – 1 s | single-digit ms |
| Java + Spring Boot | ~2 – 8 s | single-digit ms |
| Node/Python | ~100 – 400 ms | ms |
The cold-start tax is paid per execution environment, not per request. A function serving steady traffic on 10 warm environments pays it 10 times and then not again. A function receiving one request a minute may pay it on most requests.
Cold starts affect the tail, not the average. A function at 1% cold-start rate with a 3-second init has a fine p50 and a terrible p99. If your SLO is on p99 — and it should be — this is your problem.
The freeze is the subtlety that catches people
Between invocations the environment is frozen: memory is preserved, but the CPU is taken away. Consequences that produce real bugs:
- Background threads stop mid-execution. A
@Scheduledtask, an async metric flusher, a connection-pool reaper — none of them run between invocations. An async log or metric emitted just before you return may never be flushed. Flush before returning. - In-memory state survives into the next invocation of that environment. Useful for caching a config value or reusing an SDK client. Dangerous if you accidentally retain per-request state — one user's data leaking into another user's request is a real class of Lambda bug.
- Nothing is guaranteed to survive. The environment can be reaped at any time. Never treat in-memory state as durable.
Init phase — what you should and should not do there
The init phase is where you want expensive one-time work, because it runs once per environment rather than once per request:
public class DocumentHandler implements RequestHandler<S3Event, String> {
// STATIC: runs once per execution environment, during INIT.
// Every subsequent warm invocation reuses these. Creating an S3Client
// inside handleRequest() would add ~100-300ms of credential resolution
// and connection setup to EVERY invocation.
private static final S3Client S3 = S3Client.builder()
.region(Region.of(System.getenv("AWS_REGION")))
.httpClientBuilder(UrlConnectionHttpClient.builder()) // lighter than Apache for Lambda
.build();
@Override
public String handleRequest(S3Event event, Context context) { ... }
}Billing note, and a genuine "verify this" item: AWS has changed how the INIT phase is billed. Historically, init time for zip-packaged functions using managed runtimes was not billed; AWS moved to billing INIT duration for managed runtimes in 2025. Container-image functions and provisioned concurrency were always billed differently. Check current pricing documentation — but either way, init time is latency your users feel, so optimizing it is worth doing regardless of who pays for it.
Memory is the CPU dial
This is the single most valuable practical fact about Lambda, and it is not obvious from the name.
| Memory configured | Approximate CPU |
|---|---|
| 128 MB | a small fraction of a vCPU |
| 1,769 MB | ~1 full vCPU |
| 3,538 MB | ~2 vCPUs |
| 10,240 MB (max) | ~6 vCPUs |
CPU (and network, and I/O) scale linearly with configured memory. So:
Raising memory often lowers cost. A CPU-bound function at 512 MB taking 4,000 ms may take 900 ms at 2,048 MB. You are billed GB-ms:
0.5 × 4000 = 2,000vs2 × 900 = 1,800. Cheaper and 4× faster.
This is deeply counter-intuitive for anyone used to sizing servers, and it is why AWS Lambda Power Tuning (an open-source Step Functions tool that runs your function at a range of memory settings and charts cost vs duration) is worth running on any function that matters. The optimum is almost never the default.
For a JVM function, do not go below ~1,024 MB — a memory-starved JVM is a slow JVM, and you pay for the slowness.
The hard limits
| Limit | Value | Consequence |
|---|---|---|
| Timeout | 15 minutes max (default 3 s) | Anything longer needs Step Functions (Day 9) or a container |
| Memory | 128 MB – 10,240 MB | Also your CPU setting |
/tmp | 512 MB – 10,240 MB, ephemeral | Scratch space, shared within an environment |
| Sync payload | 6 MB request and response | Larger data goes via S3 with a pointer |
| Async payload | 256 KB | Smaller than you expect; catches people on SNS/EventBridge events |
| Deployment package | 50 MB zipped / 250 MB unzipped; 10 GB as a container image | Spring Boot fat jars get close to the zip limit |
| Environment variables | 4 KB total | Config beyond that goes to Parameter Store |
| Concurrent executions | 1,000 per Region by default (raisable) | Shared across every function in the Region — one runaway function starves the others |
That concurrency limit deserves emphasis: it is per account per Region, shared by all your functions. A misbehaving batch job can throttle your customer-facing API. Reserved concurrency exists to prevent exactly that.
2. Concurrency, throttling, and the three invocation models
Concurrency
Concurrency = the number of invocations executing simultaneously, not requests per second. By Little's Law:
concurrency ≈ requests_per_second × average_duration_seconds
100 rps × 0.2 s = 20 concurrent
100 rps × 2.0 s = 200 concurrentSlow functions consume concurrency. Halving duration halves concurrency need.
| Control | What it does | Cost |
|---|---|---|
| Account concurrency limit | Ceiling for the whole Region (default 1,000) | — |
| Reserved concurrency | Carves out a guaranteed slice for one function, and caps it. Also protects everything else from it. | Free, but the reservation is removed from the shared pool |
| Provisioned concurrency | Pre-initialized environments, kept warm. Eliminates cold starts for that many concurrent executions. | Charged per hour whether used or not — this is the setting that quietly removes Lambda's "pay only for use" property |
Reserved concurrency is dual-purpose and worth stating explicitly:
- Set it low on a function to protect the rest of your account from it (a batch job, an untrusted webhook).
- Set it at all on a critical function to guarantee it can always run.
- Set it to zero to disable a function instantly without deleting it — the emergency stop.
Lambda also applies a burst limit on how fast concurrency can ramp, after which it scales in steps. For a genuine cold-start-from-zero spike, you will be throttled during ramp-up. (Burst values have changed several times; check current documentation rather than memorizing a number.)
The three invocation models — and their completely different retry behaviour
This is the highest-value table on this page. Getting it wrong is how duplicate charges happen.
flowchart TB
subgraph SYNC["SYNCHRONOUS (RequestResponse)"]
S1["API Gateway, ALB,<br/>SDK Invoke, Function URL"] --> S2[Lambda]
S2 --> S3["Result returned to caller"]
S2 -.->|"error or throttle (429)"| S4["Caller sees it.<br/>LAMBDA DOES NOT RETRY.<br/>The CLIENT decides."]
end
subgraph ASYNC["ASYNCHRONOUS (Event)"]
A1["SNS, EventBridge, S3 events,<br/>Invoke with InvocationType=Event"] --> A2["Lambda internal queue"]
A2 --> A3[Lambda]
A3 -.->|error| A4["LAMBDA RETRIES: 2 more times<br/>(configurable 0-2), with backoff"]
A4 -.->|still failing| A5["On-failure destination / DLQ"]
A2 -.->|throttled| A6["Retried by the service<br/>for up to ~6 hours"]
end
subgraph POLL["POLL-BASED (Event source mapping)"]
P1["SQS, Kinesis, DynamoDB Streams,<br/>Kafka"] --> P2["Lambda-managed poller"]
P2 --> P3[Lambda]
P3 -.->|error| P4["THE SOURCE decides:<br/>SQS → visibility timeout, redrive, DLQ<br/>Kinesis → retry the shard, blocking it"]
end| Model | Who retries | Default retries | Duplicate risk |
|---|---|---|---|
| Synchronous | The caller | None by Lambda | Only if the client retries |
| Asynchronous | Lambda itself | 2 (3 attempts total), plus up to ~6 hours of retries on throttling | High, and it is normal operation |
| Poll-based | The event source | SQS: maxReceiveCount; Kinesis: retries the shard | High |
⚠️ The statement to remember: an asynchronously-invoked Lambda will run twice sometimes, with no failure anywhere that anyone reports. A retry after a timeout can mean the first attempt actually completed. If the handler charges a card, sends an email, or increments a counter, it must be idempotent. Day 8 is largely about how.
For async invocations, configure the failure path explicitly:
aws lambda put-function-event-invoke-config \
--function-name document-validator \
--maximum-retry-attempts 2 \
--maximum-event-age-in-seconds 3600 \
--destination-config '{
"OnFailure": {"Destination": "arn:aws:sqs:eu-west-1:123456789012:validator-dlq"},
"OnSuccess": {"Destination": "arn:aws:events:eu-west-1:123456789012:event-bus/default"}
}'On-failure destinations are strictly better than the older DLQ setting — they include the invocation context and the error, not just the original payload.
3. Lambda in a VPC
By default a Lambda runs in an AWS-managed network with Internet access and no access to your VPC. Attach it to your VPC and it gets an ENI in your subnets — and loses Internet access unless you provide NAT or VPC endpoints. Exactly the Day 2 rules, applied to a function.
flowchart LR
subgraph VPC["Your VPC"]
subgraph PRIV["private subnets"]
LENI["Lambda's Hyperplane ENI"]
end
NAT[NAT Gateway] --> IGW[IGW]
GWE[S3/DynamoDB gateway endpoint]
IEP[Interface endpoints:<br/>SQS, Secrets Manager, KMS]
end
LENI -->|"private DB, internal services"| RDS[(RDS)]
LENI -->|"S3, DynamoDB — free"| GWE
LENI -->|"other AWS services"| IEP
LENI -->|"public Internet — costs"| NATTwo important pieces of history, because stale advice is everywhere:
- VPC attachment used to add seconds to every cold start because an ENI was created per environment. Since 2019 Lambda uses shared Hyperplane ENIs created at function create/update time. VPC cold-start penalty is now negligible. Advice telling you to avoid VPC-attaching Lambdas for performance is obsolete.
- Lambda still consumes IP addresses from your subnets, and a very high-concurrency function in a small subnet can exhaust them. Day 2's sizing lesson again.
Attach to a VPC only when you need to reach something private. If your function only talks to S3, DynamoDB and SQS, leave it outside the VPC — it is simpler and avoids the NAT/endpoint question entirely.
4. Java on Lambda — the honest version
Java is the least natural fit for Lambda among the mainstream runtimes, and pretending otherwise leads people into bad architectures. The reason is the init phase: JVM startup plus classloading plus framework initialization is measured in seconds.
| Approach | Cold start | Verdict |
|---|---|---|
Plain Java handler, minimal deps, UrlConnectionHttpClient | ~300 ms – 1 s | Good. This is Java on Lambda done right. |
Spring Boot via spring-cloud-function / AWS_PROXY adapter | ~2 – 6 s | Works. Painful tail latency. Use SnapStart. |
| Full Spring Boot web app in a container image | ~4 – 10 s | You are running a server in a function. Use Fargate. |
| Quarkus / Micronaut (build-time DI) | ~500 ms – 1.5 s | Designed for this. Genuinely good. |
| GraalVM native image | ~100 – 300 ms | Excellent runtime, painful build and reflection story |
| SnapStart (Java) | ~200 – 800 ms | The pragmatic answer for Java. Use it. |
SnapStart, and its two real caveats
SnapStart runs your init phase once at version publish time, snapshots the memory state of the initialized microVM, and restores from that snapshot on invocation — skipping JVM boot and framework init entirely.
Two caveats that produce genuine bugs:
- Uniqueness. Anything generated during init is now shared by every restored environment:
Randomseeds, UUIDs cached at startup, cryptographic key material. Use theafterRestoreruntime hook (CRaCResource) to re-seed. - Connections. Network connections captured in the snapshot are dead on restore. Database connections, long-lived HTTP connections and SDK clients holding sockets must be re-established in
afterRestore.
import org.crac.Core;
import org.crac.Resource;
public class Handler implements RequestHandler<Map<String,Object>, String>, Resource {
private SecureRandom random;
public Handler() {
Core.getGlobalContext().register(this);
this.random = new SecureRandom();
}
@Override public void beforeCheckpoint(org.crac.Context<? extends Resource> c) {
// close anything socket-backed before the snapshot is taken
}
@Override public void afterRestore(org.crac.Context<? extends Resource> c) {
// CRITICAL: re-seed, or every restored environment shares a seed
this.random = new SecureRandom();
}
}The senior framing: "If your Java workload needs a framework-heavy container and sub-100 ms p99, Lambda is the wrong tool — use Fargate. If it is event-driven, spiky, or glue, use Lambda with a lean handler and SnapStart."
5. API Gateway
Three API types
| REST API | HTTP API | WebSocket API | |
|---|---|---|---|
| Cost per million requests | ~$3.50 | ~$1.00 | ~$1.00 + connection-minutes |
| Latency added | Higher | Lower | — |
| Authorizers | IAM, Cognito, Lambda (token + request) | IAM, JWT, Lambda | IAM, Lambda |
| API keys + usage plans | Yes | No | No |
| Request/response mapping templates (VTL) | Yes | Simple parameter mapping only | — |
| Request validation (schema) | Yes | No | — |
| Response caching | Yes (priced hourly) | No | — |
| WAF | Yes | Limited — verify current support | — |
| Private (VPC-only) endpoints | Yes | No | — |
| Canary deployments | Yes | No (use stages + weighted routing) | — |
| Integration timeout | 29 s (raisable on some configurations — verify) | 30 s | — |
| Bidirectional | No | No | Yes |
How to choose, in one line each:
- HTTP API by default. Cheaper, faster, sufficient for most backends.
- REST API when you need API keys/usage plans (monetization, per-partner quotas), schema validation, response caching, private endpoints, or mapping templates for a legacy contract.
- WebSocket API for genuinely bidirectional traffic — live dashboards, chat, notifications.
The 3.5× price gap is not a rounding error. At 500M requests/month: ~$500 vs ~$1,750. On the other hand, at 500M requests/month the response cache on a REST API might save you far more than that in compute. Do the arithmetic for your traffic.
The request pipeline
flowchart TB
REQ[Client request] --> RL{"Account/stage throttle<br/>exceeded?"}
RL -->|yes| E429["429 Too Many Requests"]
RL -->|no| ROUTE{"Route match?"}
ROUTE -->|no| E404[404 / 403 Forbidden]
ROUTE -->|yes| AUTH{"Authorizer"}
AUTH -->|deny| E401["401 / 403"]
AUTH -->|"allow (result cached, TTL)"| KEY{"API key + usage plan<br/>(REST only)"}
KEY -->|over quota| E429b["429"]
KEY -->|ok| VAL{"Request validation<br/>(REST only)"}
VAL -->|invalid| E400["400 Bad Request<br/>— your code never runs"]
VAL -->|valid| CACHE{"Cache hit?<br/>(REST only)"}
CACHE -->|hit| CACHED["Cached response<br/>— integration not invoked"]
CACHE -->|miss| MAP["Request mapping"]
MAP --> INT["Integration:<br/>Lambda / HTTP / ALB / AWS service"]
INT -->|"no response in 29-30 s"| E504["504"]
INT --> RMAP["Response mapping"] --> RESP[Response]The point of this diagram: throttling, authentication, key checking, validation and caching all happen before your code runs. Every one of them is work you would otherwise write, deploy and maintain in every service. That is API Gateway's actual value proposition — not routing.
Authorizers
| Type | How | Use |
|---|---|---|
| IAM (SigV4) | The caller signs with AWS credentials | Service-to-service, internal tools, SDK clients |
| Cognito user pool | Validates a Cognito-issued JWT | You use Cognito (Day 11) |
| JWT (HTTP API) | Validates any OIDC-compliant JWT against a JWKS issuer | The clean answer for a modern API with any standards-compliant IdP |
| Lambda authorizer | Your function returns an IAM policy or a simple allow/deny | Custom schemes, opaque tokens, per-request context enrichment |
Lambda authorizers cache by default (TTL up to 1 hour). The cache key matters enormously:
TOKENtype keys on the token → one cold authorizer call per token per TTL. Good.REQUESTtype can key on headers, paths and query strings → a poor cache key means an authorizer invocation on every request, which doubles your Lambda cost and adds its latency to every call.
Throttling
Two independent layers:
| Layer | Scope | Configure |
|---|---|---|
| Account-level | All APIs in the Region (default ~10,000 RPS with a burst allowance for REST) | Support request to raise |
| Stage / route | Per stage or per method | Steady-state rate + burst |
| Usage plan (REST only) | Per API key | Rate, burst, and a daily/weekly/monthly quota |
Throttling uses a token bucket: burst is the bucket size, rate the refill. Exceeding it returns 429 with Retry-After. Clients should back off — which is exactly the retry-with-jitter behaviour Day 12 covers.
Throttling is a protection mechanism, not a monetization feature. Its real job is to stop one client from consuming all your downstream capacity. Set it even if you never bill anyone.
6. API Gateway vs ALB
| API Gateway | ALB | |
|---|---|---|
| Pricing model | Per request (~$1–3.50/M) | Per hour (~$16/mo) + LCU |
| Cost at 1M req/mo | ~$1 | ~$17 |
| Cost at 500M req/mo | ~$500 | ~$50–150 |
| Idle cost | $0 | ~$16/month |
| Auth built in | Yes (IAM, JWT, Cognito, Lambda) | No (do it in your app, or OIDC on the listener) |
| Rate limiting / quotas | Yes | No |
| Request validation | Yes (REST) | No |
| Caching | Yes (REST) | No |
| Lambda target | Native | Yes |
| Timeout | 29–30 s | Configurable idle timeout, much longer |
| WebSockets | Yes (WebSocket API) | Yes (pass-through) |
| Long-lived / streaming | Poor | Good |
flowchart TD
A{"Target is Lambda?"} -->|yes| AG[API Gateway]
A -->|no, containers| B{"Need built-in auth,<br/>API keys, quotas,<br/>or request validation?"}
B -->|yes| AG
B -->|no| C{"Very high, steady<br/>request volume?"}
C -->|yes| ALB["ALB — the per-hour model wins"]
C -->|no| D{"Long-lived connections,<br/>streaming, or >30 s requests?"}
D -->|yes| ALB
D -->|no| E["Either. Prefer API Gateway<br/>for the managed features;<br/>ALB if you already have one."]The anti-pattern to name explicitly: API Gateway in front of an ALB in front of ECS, with no authorizer, no validation, no caching and no quotas. You have added a per-request charge, a 29-second timeout and a hop, in exchange for nothing. If you are not using API Gateway's features, you are paying for a proxy.
Architecture
High-level
flowchart LR
C[Client] --> AG[API Gateway HTTP API<br/>JWT authorizer · throttle]
AG --> L["Lambda: document-api"]
L --> S3[(S3)]
S3 -->|ObjectCreated| LV["Lambda: document-validator<br/>(async)"]
LV --> DDB[(metadata)]
LV -.->|"after 3 attempts"| DLQ[(SQS DLQ)]Detailed
flowchart TB
C[Client] -->|"Authorization: Bearer <JWT>"| AG
subgraph APIGW["API Gateway · HTTP API · stage: prod"]
AUTH["JWT authorizer<br/>issuer + audience validated"]
R1["POST /documents"]
R2["GET /documents/{id}"]
TH["Throttle: 100 rps, burst 200"]
end
AG --- AUTH
AG --- TH
R1 --> L1["Lambda document-api<br/>1024 MB · 10 s timeout<br/>SnapStart enabled<br/>reserved concurrency 50"]
R2 --> L1
L1 --> ROLE["Execution role:<br/>s3:PutObject/GetObject on one prefix<br/>logs:* on its own log group"]
L1 --> S3[(S3 bucket)]
S3 -->|"ObjectCreated:* (prefix filter)"| L2["Lambda document-validator<br/>ASYNC invocation<br/>maxRetry 2 · maxEventAge 1h"]
L2 --> DDB[(DynamoDB metadata)]
L2 -.->|"on failure"| DLQ[(SQS validator-dlq)]
L1 & L2 --> CW[CloudWatch Logs + Metrics]Request flow — the internal working of an API Gateway → Lambda call
sequenceDiagram
autonumber
participant C as Client
participant AG as API Gateway
participant JWKS as IdP JWKS endpoint
participant L as Lambda service
participant EE as Execution environment
participant S3
C->>AG: POST /documents Authorization: Bearer eyJ...
AG->>AG: stage throttle check (token bucket)
AG->>AG: route match POST /documents
AG->>JWKS: fetch signing keys (cached)
AG->>AG: verify signature, issuer, audience, exp
AG->>L: Invoke (RequestResponse) with the event JSON
L->>L: any warm environment free?
alt Cold start
L->>EE: create microVM, download code, start JVM
EE->>EE: INIT: static init, SDK clients<br/>(or SnapStart restore)
end
L->>EE: invoke handler(event, context)
EE->>S3: PutObject (execution-role credentials)
S3-->>EE: 200
EE-->>L: response payload (≤ 6 MB)
L-->>AG: 200 + body
AG->>AG: response mapping, access log, metrics
AG-->>C: 201 Created
Note over EE: environment FROZEN — CPU removed,<br/>memory retained, awaiting reuseFailure flow
flowchart TB
C[Client] --> AG[API Gateway]
AG -->|"over throttle"| E429["429 — your code never ran"]
AG -->|"bad JWT"| E401[401]
AG --> L[Lambda]
L -->|"account concurrency exhausted"| T["Throttled → 429 to API Gateway → 429 to client<br/>(sync: NO automatic retry)"]
L -->|"handler throws"| E502["502 — API Gateway cannot<br/>interpret the response"]
L -->|"runs > 29 s"| E504["504 Gateway Timeout<br/>(the Lambda may still be running!)"]
S3[(S3 event)] -.->|async| L2[validator]
L2 -->|"attempt 1 fails"| R1["retry after ~1 min"]
R1 -->|"attempt 2 fails"| R2["retry after ~2 min"]
R2 -->|"attempt 3 fails"| DLQ[(DLQ + alarm)]
L2 -.->|"⚠️ attempt 1 timed out but COMPLETED"| DUP["Duplicate side effect<br/>→ needs idempotency"]The most dangerous box in that diagram is
E504. API Gateway gives up at 29–30 seconds; the Lambda keeps running to its own timeout. The client sees a failure, retries, and you now have two concurrent executions doing the same work. Always set the Lambda timeout below the API Gateway integration timeout.
Code
A lean Java handler — the way Java on Lambda should look
<dependencies>
<dependency>
<groupId>com.amazonaws</groupId>
<artifactId>aws-lambda-java-core</artifactId>
<version>1.2.3</version>
</dependency>
<dependency>
<groupId>com.amazonaws</groupId>
<artifactId>aws-lambda-java-events</artifactId>
<version>3.14.0</version>
</dependency>
<dependency>
<groupId>software.amazon.awssdk</groupId>
<artifactId>s3</artifactId> <!-- BOM-managed -->
</dependency>
<dependency>
<!-- Lighter than apache-client: fewer classes to load during INIT,
which is a real cold-start saving. No connection pooling, which
is fine because a Lambda handles one request at a time. -->
<groupId>software.amazon.awssdk</groupId>
<artifactId>url-connection-client</artifactId>
</dependency>
</dependencies>package com.acme.lambda;
import com.amazonaws.services.lambda.runtime.Context;
import com.amazonaws.services.lambda.runtime.RequestHandler;
import com.amazonaws.services.lambda.runtime.events.APIGatewayV2HTTPEvent;
import com.amazonaws.services.lambda.runtime.events.APIGatewayV2HTTPResponse;
import software.amazon.awssdk.core.sync.RequestBody;
import software.amazon.awssdk.http.urlconnection.UrlConnectionHttpClient;
import software.amazon.awssdk.regions.Region;
import software.amazon.awssdk.services.s3.S3Client;
import software.amazon.awssdk.services.s3.model.PutObjectRequest;
import java.nio.charset.StandardCharsets;
import java.time.Duration;
import java.util.Map;
import java.util.UUID;
public class DocumentApiHandler
implements RequestHandler<APIGatewayV2HTTPEvent, APIGatewayV2HTTPResponse> {
// ---- INIT PHASE: runs once per execution environment ----
// Anything expensive belongs here, not in handleRequest.
private static final S3Client S3 = S3Client.builder()
.region(Region.of(System.getenv("AWS_REGION")))
.httpClientBuilder(UrlConnectionHttpClient.builder()
.connectionTimeout(Duration.ofSeconds(2))
.socketTimeout(Duration.ofSeconds(5)))
.overrideConfiguration(c -> c
// Keep the total SDK budget well inside the function timeout,
// so we fail with a useful error rather than being killed.
.apiCallTimeout(Duration.ofSeconds(6))
.apiCallAttemptTimeout(Duration.ofSeconds(2)))
.build();
private static final String BUCKET = System.getenv("COURSE_BUCKET");
@Override
public APIGatewayV2HTTPResponse handleRequest(APIGatewayV2HTTPEvent event, Context ctx) {
// The request id is the correlation id that ties this invocation to
// CloudWatch Logs, X-Ray, and the client's error report.
String requestId = ctx.getAwsRequestId();
try {
String body = event.getBody();
if (body == null || body.isBlank()) {
return json(400, "{\"error\":\"empty body\"}", requestId);
}
String id = UUID.randomUUID().toString();
S3.putObject(PutObjectRequest.builder()
.bucket(BUCKET)
.key("documents/" + id + ".txt")
.contentType("text/plain")
.build(),
RequestBody.fromString(body, StandardCharsets.UTF_8));
// Structured log line. CloudWatch Logs Insights can aggregate on these
// fields; Day 13 builds the queries.
System.out.printf(
"{\"event\":\"document.created\",\"documentId\":\"%s\",\"requestId\":\"%s\","
+ "\"remainingMs\":%d}%n",
id, requestId, ctx.getRemainingTimeInMillis());
return json(201, "{\"id\":\"" + id + "\"}", requestId);
} catch (Exception e) {
// Never let an exception escape a synchronous handler: API Gateway
// turns it into an opaque 502. Return a shaped error instead.
System.err.printf("{\"event\":\"document.create.failed\",\"requestId\":\"%s\","
+ "\"error\":\"%s\"}%n", requestId, e.getClass().getSimpleName());
return json(500, "{\"error\":\"internal\",\"requestId\":\"" + requestId + "\"}", requestId);
}
}
private static APIGatewayV2HTTPResponse json(int status, String body, String requestId) {
return APIGatewayV2HTTPResponse.builder()
.withStatusCode(status)
.withHeaders(Map.of("Content-Type", "application/json",
"X-Request-Id", requestId))
.withBody(body)
.build();
}
}Five deliberate choices in that file:
- Static SDK client. Created during init, reused across warm invocations.
UrlConnectionHttpClient. Fewer classes to load than the Apache client — measurable cold-start saving, and connection pooling buys nothing when a function handles one request at a time.- Timeouts inside the function timeout. So you fail with a logged error rather than being killed at the ceiling with no diagnostics.
getRemainingTimeInMillis()logged. The one Lambda-specific observability field: it tells you how close you ran to the timeout.- No exception escapes. A thrown exception from a sync handler becomes a 502 with no useful body.
An idempotent async handler
package com.acme.lambda;
import com.amazonaws.services.lambda.runtime.Context;
import com.amazonaws.services.lambda.runtime.RequestHandler;
import com.amazonaws.services.lambda.runtime.events.S3Event;
import software.amazon.awssdk.services.dynamodb.DynamoDbClient;
import software.amazon.awssdk.services.dynamodb.model.*;
import java.util.Map;
/**
* Invoked ASYNCHRONOUSLY by S3 events. Lambda will retry this up to 2 extra
* times on failure — and a "failure" can be a timeout on an attempt that
* actually completed. So this handler MUST be safe to run twice.
*
* The mechanism: a conditional write on DynamoDB keyed by the S3 object's
* versioned identity. The second attempt loses the race and exits cleanly.
* Day 7 covers conditional writes; Day 8 covers idempotency as a discipline.
*/
public class DocumentValidatorHandler implements RequestHandler<S3Event, Void> {
private static final DynamoDbClient DDB = DynamoDbClient.create();
private static final String TABLE = System.getenv("METADATA_TABLE");
@Override
public Void handleRequest(S3Event event, Context ctx) {
event.getRecords().forEach(record -> {
String bucket = record.getS3().getBucket().getName();
String key = record.getS3().getObject().getKey();
// eTag identifies this exact object version — a re-upload is genuinely
// new work, a retry of the same event is not.
String dedupeKey = bucket + "/" + key + "#" + record.getS3().getObject().geteTag();
try {
DDB.putItem(PutItemRequest.builder()
.tableName(TABLE)
.item(Map.of(
"pk", AttributeValue.fromS(dedupeKey),
"status", AttributeValue.fromS("VALIDATED"),
"requestId", AttributeValue.fromS(ctx.getAwsRequestId())))
// Succeeds only if nothing has claimed this key yet.
.conditionExpression("attribute_not_exists(pk)")
.build());
} catch (ConditionalCheckFailedException already) {
// NOT an error. A previous attempt already did this work.
System.out.printf("{\"event\":\"validation.duplicate.skipped\",\"key\":\"%s\"}%n", key);
return;
}
validate(bucket, key);
});
return null;
}
private void validate(String bucket, String key) { /* ... */ }
}Deploying, with the settings that matter
# Reserved concurrency: caps this function at 50 AND protects the other 950
# of the account pool from it.
aws lambda create-function \
--function-name document-api \
--runtime java17 --architectures arm64 \
--handler com.acme.lambda.DocumentApiHandler::handleRequest \
--zip-file fileb://target/document-api.jar \
--role arn:aws:iam::123456789012:role/document-api-role \
--memory-size 1024 \
--timeout 10 \
--environment "Variables={COURSE_BUCKET=$BUCKET}" \
--tracing-config Mode=Active
aws lambda put-function-concurrency \
--function-name document-api --reserved-concurrent-executions 50
# SnapStart applies to published VERSIONS, not $LATEST. This is the whole
# workflow difference: publish a version, alias it, point the API at the alias.
aws lambda update-function-configuration \
--function-name document-api --snap-start ApplyOn=PublishedVersions
VERSION=$(aws lambda publish-version --function-name document-api --query Version --output text)
aws lambda create-alias --function-name document-api --name prod --function-version "$VERSION"
# Log retention. Lambda creates the log group for you — with no retention.
aws logs put-retention-policy \
--log-group-name "/aws/lambda/document-api" --retention-in-days 14The HTTP API, with an authorizer and a throttle
API_ID=$(aws apigatewayv2 create-api \
--name course-api --protocol-type HTTP \
--query ApiId --output text)
# JWT authorizer — no code, no Lambda, no cold start.
AUTH_ID=$(aws apigatewayv2 create-authorizer \
--api-id "$API_ID" --name jwt-auth --authorizer-type JWT \
--identity-source '$request.header.Authorization' \
--jwt-configuration "Audience=my-api,Issuer=https://issuer.example.com/" \
--query AuthorizerId --output text)
INT_ID=$(aws apigatewayv2 create-integration \
--api-id "$API_ID" --integration-type AWS_PROXY \
--integration-uri "arn:aws:lambda:eu-west-1:123456789012:function:document-api:prod" \
--payload-format-version 2.0 \
--query IntegrationId --output text)
aws apigatewayv2 create-route --api-id "$API_ID" \
--route-key 'POST /documents' \
--target "integrations/${INT_ID}" \
--authorization-type JWT --authorizer-id "$AUTH_ID"
# Throttle + access logs. Both are off/unlimited by default.
aws apigatewayv2 create-stage --api-id "$API_ID" --stage-name prod --auto-deploy \
--default-route-settings 'ThrottlingRateLimit=100,ThrottlingBurstLimit=200' \
--access-log-settings "DestinationArn=arn:aws:logs:eu-west-1:123456789012:log-group:/aws/apigw/course-api,Format={\"requestId\":\"\$context.requestId\",\"status\":\"\$context.status\",\"latency\":\"\$context.responseLatency\",\"integrationLatency\":\"\$context.integration.latency\"}"
# API Gateway must be granted permission to invoke the function.
aws lambda add-permission --function-name document-api:prod \
--statement-id apigw --action lambda:InvokeFunction \
--principal apigateway.amazonaws.com \
--source-arn "arn:aws:execute-api:eu-west-1:123456789012:${API_ID}/*/*/documents"
$context.responseLatencyvs$context.integration.latencyin that access-log format is not decoration. The difference between them is time spent in API Gateway — authorizer, mapping, throttling — as opposed to in your function. When someone says "the API is slow", that subtraction tells you whose problem it is.
Hands-on
Lab 6 — Lambda in Java, with cold starts measured (20 min)
Objective: deploy the lean handler, measure cold vs warm, and see what memory sizing and SnapStart actually do.
Prerequisites: the S3 bucket from Day 3. No VPC needed — this function only talks to S3.
⚠️ Billable: essentially nothing at this scale; Lambda's free tier covers it. Log retention still matters.
Steps:
- Build the lean handler jar. Note its size — if it is over ~40 MB, you have dependencies you do not need.
- Create the execution role: a narrow S3 policy plus
AWSLambdaBasicExecutionRole. - Deploy at 512 MB. Invoke once, note
Init DurationandDuration. Invoke again immediately and note the absence ofInit Duration. - Redeploy at 1,024 and 2,048 MB, repeating the measurement.
- Enable SnapStart, publish a version, invoke the version cold.
Verification:
# Force a cold start by changing configuration (this discards warm environments)
aws lambda update-function-configuration --function-name document-api \
--environment "Variables={COURSE_BUCKET=$BUCKET,CACHE_BUST=$(date +%s)}" >/dev/null
aws lambda wait function-updated-v2 --function-name document-api
aws lambda invoke --function-name document-api \
--payload "$(printf '{"body":"hello","isBase64Encoded":false}' | base64)" \
--log-type Tail out.json --query LogResult --output text | base64 -d | grep REPORT
# REPORT ... Duration: 842.11 ms Billed Duration: 843 ms Memory Size: 1024 MB
# Max Memory Used: 187 MB Init Duration: 2104.55 ms ← COLD
# Immediately again — no Init Duration line
aws lambda invoke --function-name document-api --payload ... --log-type Tail out.json \
--query LogResult --output text | base64 -d | grep REPORT
# REPORT ... Duration: 61.30 ms Billed Duration: 62 ms Memory Size: 1024 MB
# Max Memory Used: 189 MB ← WARMRecord a table like this — building it yourself is the lab:
| Memory | Init Duration | Warm Duration | Max Memory Used | Cost per 1M invocations |
|---|---|---|---|---|
| 512 MB | ~3,100 ms | ~140 ms | 180 MB | |
| 1,024 MB | ~2,100 ms | ~61 ms | 187 MB | |
| 2,048 MB | ~1,500 ms | ~38 ms | 190 MB | |
| 1,024 + SnapStart | ~400 ms (restore) | ~61 ms | 187 MB |
Three conclusions to draw from your own numbers:
Max Memory Usedbarely moves with configured memory — so the extra memory bought you CPU, not memory.- There is a point past which extra memory stops paying for itself. Find it.
- SnapStart changes Java on Lambda from "questionable" to "fine."
Also observe the freeze, because it explains a class of bug:
// Add to the handler, then invoke several times in a row and read the logs.
private static int invocationCount = 0;
// ... inside handleRequest:
System.out.printf("{\"event\":\"env.reuse\",\"count\":%d}%n", ++invocationCount);The counter increments across invocations that share an environment and resets when a new one starts. That is in-memory state surviving between requests — useful for caching, dangerous if you leave request data in it.
Cleanup: set log retention; the function costs nothing idle. Delete it at the end of the day.
Lab 7 — API Gateway HTTP API with a JWT authorizer and proven throttling (25 min)
Objective: a real API front door — authenticated, throttled, logged — and proof that both controls work.
Steps:
- Create the HTTP API, JWT authorizer, integration, route and stage using the CLI above.
- For the IdP: use a Cognito user pool (Day 11 covers it properly) or any OIDC provider you can get a token from. The JWKS URL and audience are all API Gateway needs.
- Add the
lambda:InvokeFunctionpermission for API Gateway. - Set the stage throttle to 10 rps, burst 20 — deliberately low, so you can prove it.
- Turn on access logging with the format above.
Verification — three proofs:
ENDPOINT="https://${API_ID}.execute-api.eu-west-1.amazonaws.com/prod"
# (a) Unauthenticated is rejected BEFORE your code runs
curl -si -XPOST "$ENDPOINT/documents" -d 'hello' | head -1
# → HTTP/2 401
# Confirm in CloudWatch: the Lambda has ZERO invocations for this request.
# (b) Authenticated works
curl -si -XPOST "$ENDPOINT/documents" \
-H "Authorization: Bearer $TOKEN" -d 'hello' | head -1
# → HTTP/2 201
# (c) Throttling: 50 concurrent requests against a 10 rps / burst 20 limit
hey -n 200 -c 50 -m POST -H "Authorization: Bearer $TOKEN" -d 'hello' "$ENDPOINT/documents"
# Status code distribution:
# [201] 62 responses
# [429] 138 responses ← the throttle is real# Prove your code did not run for the throttled requests:
aws cloudwatch get-metric-statistics --namespace AWS/Lambda --metric-name Invocations \
--dimensions Name=FunctionName,Value=document-api \
--start-time "$(date -u -v-10M +%Y-%m-%dT%H:%M:%SZ)" \
--end-time "$(date -u +%Y-%m-%dT%H:%M:%SZ)" --period 60 --statistics Sum
# Sum ≈ 62, not 200.Then measure who is slow, using the access log:
fields @timestamp, status, latency, integrationLatency,
(latency - integrationLatency) as gatewayOverheadMs
| filter status = 201
| stats avg(integrationLatency), avg(gatewayOverheadMs), pct(latency, 99) by bin(1m)Cleanup:
aws apigatewayv2 delete-api --api-id "$API_ID"
aws lambda delete-function --function-name document-apiCommon errors:
| Symptom | Cause | Fix |
|---|---|---|
500 Internal Server Error, no Lambda logs | API Gateway lacks lambda:InvokeFunction | add-permission with the right source-arn |
502 Bad Gateway, Lambda logs show success | Handler returned a shape API Gateway cannot parse | With AWS_PROXY + payload format 2.0, return statusCode/body, or a plain object for auto-wrapping |
403 Forbidden with "Missing Authentication Token" | Route does not exist — the path or method is wrong | Check the route key exactly, including the leading / |
| Everything is a cold start | Function is invoked rarely; or every deploy discards environments | Expected. Provisioned concurrency or SnapStart if p99 matters. |
401 with a token that works elsewhere | Audience/issuer mismatch, or exp in the past | Decode the token and compare claim by claim |
504 after 29–30 s | Integration exceeded the API Gateway timeout | Lambda timeout must be lower, and long work belongs behind a queue |
Production relevance: this is a complete, production-shaped API front door — authentication, throttling, structured access logs, per-function concurrency limits — in about 25 minutes and with no servers.
Production Considerations
| Concern | Practice |
|---|---|
| Reserved concurrency on everything that matters | The account pool is shared. One runaway function throttling your payment API is a real incident, and reserved concurrency is the only structural prevention. |
| Lambda timeout < API Gateway timeout | Otherwise the client gets a 504 while the function keeps running — and retries create duplicates. |
| Idempotency for anything async | Non-negotiable. Async invocations retry by design. |
| On-failure destinations, and alarm on them | A DLQ nobody watches is a data-loss mechanism with extra steps. |
| Versions and aliases | Point API Gateway at an alias, not $LATEST. Aliases enable weighted traffic shifting for canaries and make rollback a single API call. |
| Log retention | Lambda creates log groups with no retention. Set it for every function. |
Max Memory Used in the REPORT line | Free right-sizing data on every invocation. Use it. |
| Power-tune the functions that matter | The default memory setting is almost never the cost optimum. |
| ARM64 (Graviton) | ~20% cheaper per GB-second, and Java runs unchanged. |
| Stay out of the VPC unless you need it | Simpler, no NAT/endpoint question, no subnet IP consumption. |
| Bundle only what you need | Every megabyte is cold-start time. Spring Boot Lambdas are large for a reason worth questioning. |
| Don't use Lambda as a web server | If you find yourself adding provisioned concurrency to every function to fix latency, you have chosen the wrong compute model. That is Fargate's job. |
Failure Scenarios
S9 — "p99 is 3.2 s, p50 is 180 ms"
A Java Lambda behind API Gateway. Average latency looks fine. The p99 is terrible and support tickets say the app is "sometimes very slow."
| Question | Answer |
|---|---|
| What can fail? | Nothing is broken. The distribution is bimodal: warm invocations at ~180 ms, cold starts at ~3 s. |
| What happens? | Traffic pattern means ~1–3% of requests land on a new environment. Those users wait 3 seconds. |
| Recovery? | N/A — it is steady-state behaviour |
| Discriminating signal | The Init Duration field in the REPORT lines; or a Logs Insights query counting how many invocations have one |
| Data lost / duplicated? | No |
fields @timestamp, @initDuration, @duration
| filter ispresent(@initDuration)
| stats count() as coldStarts, avg(@initDuration) as avgInitMsThe options, with honest trade-offs:
| Fix | Effect | Cost |
|---|---|---|
| SnapStart | Init ~3 s → ~400 ms | Free; requires versions; the uniqueness/connection caveats |
| Provisioned concurrency | Zero cold starts up to the provisioned level | Charged hourly, used or not — it removes Lambda's pay-per-use property |
| Raise memory | Faster init and faster invoke | Often cost-neutral or better |
| Trim dependencies | Less to classload | Engineering time |
| Quarkus / Micronaut / native | Init in hundreds of ms | A framework migration |
| Move to Fargate | No cold starts at all | Idle cost returns; Day 4's model |
The senior answer: "First I'd quantify it — what fraction of requests are cold, and does the p99 actually breach an SLO anyone agreed to? If it does, SnapStart plus memory tuning, measured. Provisioned concurrency is the paid answer, and if I need it on every function I'd re-examine whether this workload should be on Lambda at all."
S10 — The $11,000 Lambda bill in six hours
A Lambda is triggered by s3:ObjectCreated:* on a bucket. It processes the object and writes the result back to the same bucket.
flowchart LR
U[Upload] --> B[(bucket)]
B -->|ObjectCreated| L[Lambda]
L -->|"writes result to the SAME bucket"| B
B -->|"ObjectCreated again!"| L
L -.->|"∞"| L| Question | Answer |
|---|---|
| What can fail? | Nothing throws. The loop is working as configured. |
| What happens? | Each invocation creates an object that triggers another invocation. Growth is exponential until the concurrency limit caps the rate — at which point it runs flat out, 1,000 concurrent, indefinitely. |
| Recovery? | put-function-concurrency --reserved-concurrent-executions 0 — the emergency stop. Then remove the trigger, then delete the generated objects. |
| How quickly? | Detected by a billing alarm in hours; by a human in days. AWS has recursive-loop detection for some patterns now, but do not rely on it. |
| Data lost / duplicated? | Enormously duplicated |
The arithmetic, so it is not abstract — 1,000 concurrent invocations at 1 s and 1 GB, for 6 hours:
invocations: 1,000 × 3,600 × 6 = 21,600,000
GB-seconds: 21,600,000 × 1 = 21,600,000
compute: 21.6M × $0.0000166667 ≈ $360
requests: 21.6M / 1M × $0.20 ≈ $4.32
S3 PUTs: 21.6M / 1,000 × $0.005 ≈ $108
S3 GETs, storage, CloudWatch Logs ingestion at ~1 KB/invocation ≈ 21 GB × $0.50 ≈ $10That is ~$500 for compute — and the eye-watering totals in the well-known versions of this story come from larger memory settings, longer durations, expensive downstream calls (a paid API, a database write), and days rather than hours before anyone noticed.
Prevention, in order of value:
- Never write to the bucket that triggers you. Two buckets, or at minimum a prefix/suffix filter on the notification that cannot match your own output.
- Reserved concurrency on every event-driven function. A cap turns a runaway into a slow leak you have time to notice.
- A billing alarm and an anomaly detector. Day 1's Lab 0 exists for this.
- A duration + invocation-count alarm per function. An invocation count 100× its baseline is an incident regardless of cause.
A4 (from the failure plan) — duplicate side effects from async retries
A Lambda subscribed to an SNS topic sends a confirmation email. Customers occasionally receive two.
| Question | Answer |
|---|---|
| What can fail? | Nothing reports a failure. The function "succeeded" both times. |
| What happens? | Attempt 1 sent the email, then timed out before returning (or the response was lost). Lambda saw a failure and retried. Attempt 2 sent it again. |
| Recovery? | None — the emails are sent |
| Can the operation happen twice? | Yes. This is the default behaviour of async invocation, not an edge case. |
| Fix | Idempotency: a conditional write on a dedupe key before the side effect, keyed on something stable in the event (message id, not a timestamp) |
The framing that makes this click: retry is not "try again because it failed." Retry is "try again because I do not know whether it failed." Day 8 turns this into a discipline.
Security
| Question | Answer |
|---|---|
| Who can access this? | The API: only holders of a valid JWT from the configured issuer, within the throttle. The function: only API Gateway (via the resource policy on the function) and anyone with lambda:InvokeFunction. |
| What credentials are used? | The function's execution role — temporary, rotated, delivered by the Lambda runtime environment. Same principle as Day 3's instance role and Day 4's task role; different endpoint. |
| Where is data encrypted? | TLS to API Gateway and from API Gateway to Lambda. Environment variables are encrypted at rest with a KMS key (AWS-managed by default; use your own for sensitive values). S3 at rest by default. |
| What if the execution role leaks? | Short-lived and narrowly scoped. This is why per-function roles matter — a shared "lambda-role" with the union of every function's permissions turns any one function's compromise into all of them. |
Lambda-specific security practice:
- One role per function. Not per team, not per service. The cost of extra roles is zero; the cost of a shared role is a much larger blast radius.
- Environment variables are not secrets. They are visible to anyone with
lambda:GetFunctionConfiguration, and they show in the Console. Real secrets come from Secrets Manager or Parameter Store at runtime (Day 6), or are injected by the platform. - Function URLs bypass API Gateway entirely and support only
AWS_IAMorNONEauth.NONEmeans a public, unauthenticated endpoint with no throttle and no WAF. Use them knowingly or not at all. - A Lambda authorizer is code you own, and therefore an attack surface. A JWT authorizer is configuration AWS runs. Prefer configuration.
- Validate the event. An S3 event, an SNS message, a webhook body — all untrusted input. A function that trusts
event.bodyis as injectable as any web endpoint.
Performance
| Lever | Effect |
|---|---|
| Memory | The CPU dial. The biggest single lever, and often free or cheaper. Power-tune. |
| SnapStart | 4–8× cold-start reduction for Java. Nearly always worth it. |
| Static init | Move SDK clients, config parsing and any expensive setup into init; it runs once per environment. |
| Lighter HTTP client | UrlConnectionHttpClient loads far fewer classes than Apache. |
| Smaller package | Directly reduces download and classload time. |
| ARM64 | Cheaper, generally equal or better throughput. |
| Batch size (poll-based) | Amortizes overhead across messages (Day 8) |
| Stay out of the VPC | Modern Lambda has negligible VPC cold-start penalty, but staying out avoids NAT/endpoint latency entirely |
| Avoid the "fix" that isn't | Scheduled "warming" pings are a workaround, not a solution: they do not scale with concurrency, and provisioned concurrency or SnapStart is the supported mechanism. |
Cost
| Item | Rate (us-east-1, illustrative — verify) |
|---|---|
| Lambda requests | ~$0.20 per 1M |
| Lambda compute (x86) | ~$0.0000166667 per GB-second |
| Lambda compute (ARM64) | ~20% less |
| Provisioned concurrency | ~$0.0000041667 per GB-second provisioned, plus a lower per-invocation rate |
| Free tier | 1M requests + 400,000 GB-seconds per month |
| API Gateway HTTP API | ~$1.00 per million requests |
| API Gateway REST API | ~$3.50 per million requests |
| API Gateway REST cache | ~$0.02/hr (0.5 GB) up to much more for larger sizes |
The break-even every senior engineer should be able to do
A service handling N requests/month, each taking 200 ms at 1 GB.
Lambda + HTTP API:
requests: N/1M × ($0.20 + $1.00) = N × $0.0000012
compute: N × 0.2 s × 1 GB × $0.0000166667 = N × $0.00000333
total ≈ N × $0.0000045Fargate + ALB (2 tasks × 1 vCPU / 2 GB for HA, ~$36/task/month, ALB ~$17):
fixed: 2 × $36 + $17 = ~$89/month, regardless of N| Requests/month | Lambda + API GW | Fargate + ALB |
|---|---|---|
| 100,000 | ~$0.45 | ~$89 |
| 1,000,000 | ~$4.50 | ~$89 |
| 10,000,000 | ~$45 | ~$89 |
| 20,000,000 | ~$90 | ~$89 ← crossover |
| 100,000,000 | ~$450 | ~$120 (with autoscaling) |
Read the table properly, because the naive reading is wrong in both directions:
- Below ~10M requests/month, Lambda is dramatically cheaper. "Lambda is expensive" usually comes from someone who has never priced the idle Fargate task.
- Above ~20M, containers win on raw compute — and the gap widens.
- But: the crossover moves a long way based on duration. A 2-second function is 10× the compute cost and crosses over at ~2M requests. A 20 ms function crosses far later.
- And: this ignores operational cost. Two Fargate tasks come with a cluster, images, a deployment pipeline, patching and capacity decisions. For a small service, Lambda's total cost of ownership can win well past the compute crossover.
The interview-grade statement: "Lambda's cost advantage is idle time and operational overhead, not per-unit compute. It wins on spiky and low-volume workloads and loses on sustained high-throughput ones — and the crossover depends mostly on function duration."
Alternatives
| Instead of | You could | Trade-off |
|---|---|---|
| Lambda | Fargate (Day 4) | No cold starts, no 15-minute limit, long-lived connections; idle cost and scale-out lag return |
| Lambda | Step Functions (Day 9) | Orchestration, waits, retries and >15-minute workflows as configuration; a new paradigm and per-transition cost |
| Lambda | Fargate with a queue | For long-running async work, a container consuming SQS beats chained Lambdas |
| Java on Lambda | Node/Python | Far better cold starts for glue code. Genuinely worth considering for a 30-line function. |
| Java on Lambda | Quarkus / Micronaut / GraalVM native | Built for fast startup; a framework decision |
| API Gateway | ALB (Day 4) | Cheaper at high volume, longer timeouts, better streaming; no built-in auth/quotas/validation |
| API Gateway | Lambda Function URL | Simplest possible HTTP endpoint, no extra cost; only IAM or no auth, no throttling, no WAF |
| API Gateway | CloudFront + Lambda@Edge / Functions | Edge execution, global latency; very limited runtime |
| JWT authorizer | Lambda authorizer | Arbitrary logic; your code, your cold start, your bug |
| Provisioned concurrency | SnapStart | Free vs hourly. Try SnapStart first. |
Trade-offs
| Decision | Gain | Cost |
|---|---|---|
| Lambda over containers | Zero idle cost, per-request scaling, no patching | Cold starts, 15-minute ceiling, stateless only, harder local testing |
| Higher memory | Faster, often cheaper overall | Higher per-millisecond rate; must be measured, not guessed |
| Provisioned concurrency | Predictable latency | Hourly charge that removes the pay-per-use property |
| SnapStart | Large cold-start win, free | Version-publish workflow; uniqueness and connection caveats |
| Reserved concurrency | Isolation and a blast-radius cap | Reserved capacity is removed from the shared pool; a hard cap can throttle a legitimate spike |
| Async invocation | Caller decoupled, automatic retries | Duplicates become normal — idempotency is now mandatory |
| API Gateway over ALB | Auth, quotas, validation, caching for free | Per-request cost, 29–30 s ceiling |
| HTTP API over REST API | ~3.5× cheaper, lower latency | No API keys, no caching, no request validation, no private endpoints |
| VPC-attached Lambda | Reaches private resources | NAT or endpoints required for egress; consumes subnet IPs |
Common Mistakes
- Creating SDK clients inside the handler. Hundreds of milliseconds added to every invocation.
- Assuming async invocations run exactly once. They do not, by design.
- Lambda timeout ≥ API Gateway timeout. Client sees 504, function keeps running, retry duplicates the work.
- A shared execution role for every function. One compromise becomes all of them.
- Leaving memory at the default. Usually both slower and more expensive than the optimum.
- A Lambda writing to the bucket that triggers it. The $11,000 lesson.
- No reserved concurrency anywhere. The account pool is shared; one function can starve the rest.
- Forgetting log retention. Lambda creates log groups with none.
- Expecting background threads to run between invocations. The environment is frozen. Flush before returning.
- Leaking per-request state into static fields. It survives into the next request, possibly another user's.
- Running a full Spring Boot web app on Lambda and then adding provisioned concurrency to hide it. That is a container with extra steps.
- API Gateway in front of an ALB with none of its features enabled. Paying per request for a proxy.
- A
REQUEST-type Lambda authorizer with a poor cache key. Doubles invocations and adds latency to every request.
Interview Questions
2–3 YOE
- What is AWS Lambda, and what do you pay for?
- What is a cold start?
- What is the maximum time a Lambda can run?
- What does API Gateway do that your application would otherwise have to?
- What is the difference between synchronous and asynchronous Lambda invocation?
- Where does a Lambda get its AWS credentials from?
4–5 YOE
- Your function is slow. Why might increasing its memory make it cheaper?
- A Lambda is triggered by SNS and sends emails. Customers sometimes get two. Explain exactly how, with no error anywhere.
- When would you choose an ALB over API Gateway?
- Your API returns 504 after 29 seconds but CloudWatch shows the Lambda completing in 45 seconds. What is happening and what do you change?
- What is reserved concurrency, and name two different reasons to set it.
- Why is Java a harder fit for Lambda than Node, and what would you do about it?
Senior-Level Questions
6–7 YOE
- Design a file-processing pipeline where uploads are up to 5 GB and processing takes up to 40 minutes. Where does Lambda fit, and where does it not?
- Your p99 is dominated by cold starts and the SLO is 500 ms. Walk me through your options in the order you would try them, with costs.
- You have 200 Lambda functions and one shared account concurrency limit. How do you prevent one team from taking down another's API?
- When would you put API Gateway in front of a container service, and when is that a mistake?
- Design the authentication strategy for an API serving a mobile app, a web app, and three partner integrations with different quotas.
8–10 YOE
- A team proposes rewriting a steady 5,000-RPS service from Fargate to Lambda "to save money." Do the arithmetic out loud and give your answer.
- Explain precisely how an asynchronous Lambda architecture can lose data, and design one that cannot.
- You inherit a system of 40 chained Lambdas where each invokes the next asynchronously. What is wrong with it, what would you change first, and how would you migrate without a big-bang cutover?
- Where is the consistency boundary in an "API Gateway → Lambda → DynamoDB, plus an async Lambda on the stream" architecture, and what does a client observe inside it?
- Argue that serverless increases operational complexity rather than reducing it. Then argue the opposite. Which do you actually believe, and for which workloads?
Day-End Revision
The five sentences
- A cold start is: create a microVM, download code, start the runtime, run init — then the handler; afterwards the environment is frozen and reused, so init cost is per environment, not per request.
- Memory is the CPU dial; raising it often makes a function both faster and cheaper, and the default is rarely optimal.
- Synchronous invocations are not retried by Lambda; asynchronous ones are retried twice by default, which makes idempotency mandatory.
- Reserved concurrency both caps a function and protects the shared account pool; provisioned concurrency and SnapStart attack cold starts, one with an hourly bill and one for free.
- API Gateway does throttling, authentication, validation and caching before your code runs — and if you use none of them, an ALB is cheaper.
The diagram to redraw from memory: the execution environment lifecycle, with the freeze/reuse/reap branch.
The numbers
| Lambda timeout | max 15 min (default 3 s) |
| Lambda memory | 128 MB – 10,240 MB; ~1 vCPU at ~1,769 MB |
| Sync payload / async payload | 6 MB / 256 KB |
| Default account concurrency | 1,000 per Region, shared |
| Async retries | 2 (3 attempts), up to ~6 h on throttling |
| API Gateway integration timeout | 29 s (REST) / 30 s (HTTP) |
| HTTP API vs REST API | ~$1.00 vs ~$3.50 per million |
| Lambda compute | ~$0.0000166667 per GB-second |
| Lambda/Fargate crossover | roughly 10–20M requests/month at 200 ms |
Today's trap: "Lambda is always cheaper." It is cheaper on idle time and operations. At sustained high throughput, containers win — and the crossover depends mostly on function duration.
Tomorrow needs: the execution-role pattern (a Lambda connecting to a database has a credentials problem), the concurrency idea (200 concurrent Lambdas opening database connections is Day 6's most interesting failure), and today's timeout discipline.
Mini Assignment
Time: 45–60 minutes. Capstone contribution: the Notification Worker as a Lambda, and the API Gateway front door.
- In
day-05/, write two functions:document-api— the lean synchronous handler, behind an HTTP API with a JWT authorizer and a 50 rps throttle;document-validator— triggered asynchronously by S3ObjectCreated, idempotent via a DynamoDB conditional write, with an on-failure destination to an SQS DLQ.
- Power-tune
document-apiby hand. Deploy at 512 / 1,024 / 2,048 / 3,008 MB. For each, record init duration, warm duration,Max Memory Used, and the computed cost per million invocations. Produce a table and identify the optimum. - Enable SnapStart and re-measure the cold start. Report the ratio.
- Prove the duplicate problem, then prove your fix. Invoke
document-validatortwice with the identical event payload. Show that the business side effect happened once and that the second invocation loggedvalidation.duplicate.skipped. - Prove the DLQ works. Make the validator throw unconditionally; confirm exactly 3 attempts, then a message in the DLQ, then an alarm.
- In
day-05/NOTES.md:- Your power-tuning table and chosen memory setting, with reasoning.
- At what monthly request volume does this API become cheaper on Fargate + ALB? Show the arithmetic with your measured duration.
- What breaks if
document-validatoris not idempotent? Be specific about the sequence of events. - One thing in this design you would change if the requirement became "p99 under 300 ms".
- Tear down. Commit.
Success criterion: your cost crossover is computed from your own measured duration, and your duplicate test demonstrates exactly one side effect from two invocations.
AWS Documentation
- Lambda Developer Guide · Execution environment lifecycle · Quotas
- Lambda concurrency · Asynchronous invocation and retries · Recursive loop detection
- SnapStart for Java · SnapStart uniqueness and CRaC hooks
- Lambda with a VPC · Lambda pricing
- AWS Lambda Power Tuning
- API Gateway Developer Guide · Choosing between REST and HTTP APIs · Throttling · JWT authorizers
Previous: Day 4 — Running Services: Scaling Groups, Load Balancers, Containers · Next: Day 6 — Relational Data in Production