Day 2 — Networking for Backend Developers
3h 40m · Phase 1 of 4 · Curriculum
Learning Objectives
By the end of today you can:
- Read and write CIDR notation without a calculator, and carve a VPC into subnets across AZs.
- Explain what actually makes a subnet "public", and draw the traffic path for inbound and outbound.
- State the difference between a Security Group and a NACL precisely enough to debug with it.
- Explain why a service that only talks to S3 can still generate a large NAT Gateway bill — and fix it.
- Work through a connectivity failure in a fixed order instead of guessing.
- Describe how DNS resolution works inside a VPC and what Route 53 adds.
The sentence you should be able to say tonight: "A private IP is not routable on the Internet; everything about VPC design follows from that one fact."
Prerequisites
- Day 1 — accounts, Regions and AZs, IAM basics, a working CLI profile.
- From the baseline: IP addresses, ports, DNS at a high level, what a firewall rule is.
- If CIDR is unfamiliar, read Foundation § 4 first. Today is much harder without it.
Why Does This Exist?
Yesterday's mental model was fine:
flowchart LR
L[My laptop] --> I((Internet)) --> A[AWS] --> App[My application]It stops being fine the moment you have more than one thing. Consider a completely ordinary backend: a web tier, an application tier, and a database. Give all three public IP addresses and you have just published your database to the Internet. Every credential-stuffing bot on the planet will find it within hours — this is not hypothetical, it is what happens.
So you need a way to say:
- This component may be reached from the Internet.
- That component may be reached only by the first one.
- The database may be reached only by the application tier, and may not initiate connections outward at all.
- All of them need to reach AWS services and pull OS updates, without being reachable themselves.
That is a network design problem, and AWS's answer is the VPC — a logically isolated virtual network that you define, inside a Region, where every one of those sentences becomes a concrete configuration.
You are not learning networking to become a network engineer. You are learning exactly enough to answer one question when it matters: "my service cannot reach that thing — which of the six layers between them said no?"
Beginner Explanation
Private addresses, one more time
Your home router gives your laptop something like 192.168.1.47. Your colleague's router gives their laptop the same address. Both work. Neither can reach the other directly.
These are private addresses — three ranges reserved by RFC 1918 that are not routable on the public Internet:
10.0.0.0/8 10.0.0.0 – 10.255.255.255 16.7M addresses
172.16.0.0/12 172.16.0.0 – 172.31.255.255 1.0M addresses
192.168.0.0/16 192.168.0.0 – 192.168.255.255 65.5K addressesNo Internet router will forward a packet toward 10.0.3.47. That is not a bug; it is what makes the ranges reusable by every private network on Earth.
A VPC is your private network inside AWS. You choose its address range from those blocks. Everything you put in it gets a private address. Nothing in it is reachable from the Internet until you deliberately arrange for it.
flowchart LR
subgraph V["Your VPC · 10.0.0.0/16"]
A["app 10.0.1.5"]
B["db 10.0.3.9"]
end
A <-->|works, they are in the same network| B
A -.->|needs a translator or a front door| I((Internet))
I -.->|cannot reach 10.0.3.9 directly| BCIDR in ten minutes
10.0.0.0/16 means: the first 16 bits are fixed, the remaining 16 bits are yours to vary.
10.0.0.0/16
└┬────┘└┬─┘
fixed variable → 2^16 = 65,536 addresses (10.0.0.0 – 10.0.255.255)The rule: a /N has 2^(32−N) addresses. The bigger the number after the slash, the smaller the network.
| CIDR | Addresses | Usable in AWS | Typical use |
|---|---|---|---|
| /16 | 65,536 | 65,531 | A whole VPC |
| /20 | 4,096 | 4,091 | A generous subnet |
| /24 | 256 | 251 | A normal subnet |
| /26 | 64 | 59 | A small subnet |
| /28 | 16 | 11 | The smallest AWS allows |
AWS reserves 5 addresses in every subnet. In 10.0.1.0/24:
| Address | Reserved for |
|---|---|
10.0.1.0 | Network address |
10.0.1.1 | VPC router |
10.0.1.2 | AWS DNS (the "VPC+2" address) |
10.0.1.3 | Reserved for future use |
10.0.1.255 | Broadcast address (AWS does not support broadcast, but reserves it) |
So a /24 gives you 251 usable addresses, not 256. This matters the first time you try to scale to 260 tasks in one subnet.
Four facts to know without thinking:
| Question | Answer |
|---|---|
| How many /24s fit in a /16? | 256 |
Is 10.0.5.7 inside 10.0.4.0/22? | Yes — a /22 covers 10.0.4.0–10.0.7.255 |
Do 10.0.0.0/16 and 10.1.0.0/16 overlap? | No — which is why they can be peered |
Do 10.0.0.0/16 and 10.0.0.0/24 overlap? | Yes — the second is entirely inside the first |
That last pair is the one that bites in real life: two networks with overlapping CIDRs can never be connected. Companies discover this during acquisitions, and the fix is renumbering an entire environment. Pick your VPC ranges with that in mind.
Core Concepts
1. The VPC and its subnets
A VPC is a regional construct. A subnet lives in exactly one AZ. That single sentence is why multi-AZ architecture exists: to span AZs, you create a subnet per AZ and place resources in each.
flowchart TB
subgraph VPC["VPC · 10.0.0.0/16 · eu-west-1 (regional)"]
subgraph AZA["AZ eu-west-1a"]
PUBA["public-a<br/>10.0.0.0/24"]
PRIA["private-a<br/>10.0.10.0/24"]
DATA["data-a<br/>10.0.20.0/24"]
end
subgraph AZB["AZ eu-west-1b"]
PUBB["public-b<br/>10.0.1.0/24"]
PRIB["private-b<br/>10.0.11.0/24"]
DATB["data-b<br/>10.0.21.0/24"]
end
endWhy three tiers and not two?
| Tier | Contains | Reachable from Internet? | Can reach Internet? |
|---|---|---|---|
| Public | Load balancers, NAT Gateways, bastions | Yes (inbound) | Yes |
| Private | Your application — EC2, ECS tasks, Lambdas in VPC | No | Yes, outbound only, via NAT |
| Data / isolated | RDS, ElastiCache | No | No — no route out at all |
The third tier exists because a database has no legitimate reason to initiate an outbound Internet connection. If one ever does, something is badly wrong — and with no route, it simply cannot. That is defence in depth achieved by absence, which is the cheapest kind.
A sane address plan — leave room, you cannot resize a subnet later:
VPC 10.0.0.0/16 65,536 addresses
public-a 10.0.0.0/24 251 usable ┐
public-b 10.0.1.0/24 251 usable ├ small: only LBs/NAT live here
public-c 10.0.2.0/24 251 usable ┘
private-a 10.0.16.0/20 4,091 usable ┐
private-b 10.0.32.0/20 4,091 usable ├ large: this is where scale happens
private-c 10.0.48.0/20 4,091 usable ┘
data-a 10.0.80.0/24 251 usable ┐
data-b 10.0.81.0/24 251 usable ├ small: a handful of DB ENIs
data-c 10.0.82.0/24 251 usable ┘
10.0.96.0/19 … unallocated, deliberately⚠️ You can add CIDR blocks to a VPC later, but you cannot resize a subnet. Running out of addresses in the private tier during a traffic spike — because autoscaling could not place tasks — is a genuinely common production incident. Size the tier that scales, generously.
2. Routing — what actually makes a subnet public
Here is the part that surprises people: there is no "public" checkbox on a subnet. A subnet is public if, and only if, its route table has a route to an Internet Gateway.
Every subnet is associated with exactly one route table. Every route table starts with one route you cannot delete:
| Destination | Target |
|---|---|
10.0.0.0/16 | local |
That is what lets everything inside the VPC talk to everything else inside the VPC, across subnets and AZs, with no further configuration. (Security Groups still apply — routing and filtering are separate concerns.)
Public subnet route table:
| Destination | Target | Meaning |
|---|---|---|
10.0.0.0/16 | local | Stay inside the VPC |
0.0.0.0/0 | igw-abc123 | Everything else → the Internet Gateway |
Private subnet route table:
| Destination | Target | Meaning |
|---|---|---|
10.0.0.0/16 | local | Stay inside the VPC |
0.0.0.0/0 | nat-xyz789 | Everything else → the NAT Gateway in the public subnet |
Data subnet route table:
| Destination | Target |
|---|---|
10.0.0.0/16 | local |
No default route. Nothing leaves. Nothing arrives.
Routes are matched most-specific-first: a route for 10.0.20.0/24 wins over 10.0.0.0/16, which wins over 0.0.0.0/0.
Internet Gateway
An IGW is a horizontally-scaled, redundant, free component attached to a VPC. It does two things:
- Provides a target for
0.0.0.0/0routes. - Performs one-to-one NAT between a resource's private address and its public address.
That second point is the one people miss: an EC2 instance never sees its own public IP. Its operating system is configured with 10.0.0.42. The IGW rewrites addresses on the way through. Run ip addr on an instance with a public IP and you will find only the private one — which is why applications that need to know their public address have to ask the metadata service.
A resource is reachable from the Internet only if all of these are true:
- It has a public IPv4 address (or is behind something that does, like an ALB).
- It is in a subnet whose route table points
0.0.0.0/0at an IGW. - Its Security Group allows the inbound traffic.
- Its subnet's NACL allows the traffic both inbound and outbound.
Four independent conditions. When something is unreachable, it is one of these four, and you check them in order.
NAT Gateway
Your application needs to reach the outside world — call a payment provider, pull a container image, fetch an OS package — but must not be reachable from it. That asymmetry is exactly what Network Address Translation provides.
flowchart LR
subgraph PRIV["Private subnet 10.0.10.0/24"]
APP["app · 10.0.10.7"]
end
subgraph PUB["Public subnet 10.0.0.0/24"]
NAT["NAT Gateway<br/>Elastic IP 52.x.x.x"]
end
APP -->|"src 10.0.10.7 → dst 1.2.3.4"| NAT
NAT -->|"src 52.x.x.x → dst 1.2.3.4"| IGW[Internet Gateway] --> EXT((External service))
EXT -->|"reply to 52.x.x.x"| IGW --> NAT -->|"rewritten back to 10.0.10.7"| APPThe NAT Gateway records the outbound flow so it can rewrite the reply. Nothing on the outside can initiate a connection inward — there is no mapping for a conversation that hasn't started. That is the whole security property.
Two things to know immediately:
It is zonal. A NAT Gateway lives in one AZ. If that AZ fails, private subnets routed through it lose Internet egress. Production deployments put one NAT Gateway per AZ, with each AZ's private route table pointing at its own. This costs three times as much and is usually correct — and also avoids cross-AZ data charges on every outbound byte.
It is expensive, in a way that surprises people. Roughly (us-east-1, verify current pricing):
| Charge | Approx. |
|---|---|
| Per NAT Gateway hour | |
| Per GB processed | ~$0.045 |
Three NAT Gateways for high availability is ~$100/month before a single byte moves. And that per-GB charge applies to all traffic through it — including traffic to AWS services that never needed to leave the AWS network at all. Which brings us to endpoints.
3. Security Groups and NACLs
Two firewalls, at two different levels, with one critical difference.
flowchart TB
I((Internet)) --> NACL1["NACL · subnet boundary<br/>STATELESS · allow + deny · ordered rules"]
NACL1 --> SG1["Security Group · ENI boundary<br/>STATEFUL · allow only"]
SG1 --> R[Your resource]| Security Group | Network ACL | |
|---|---|---|
| Attached to | An ENI (instance, task, LB node, RDS) | A subnet |
| Stateful? | Yes — return traffic auto-allowed | No — you must allow return traffic explicitly |
| Rules | Allow only | Allow and Deny |
| Evaluation | All rules; any match allows | Numbered, lowest first, first match wins |
| Can reference | Other Security Groups, prefix lists, CIDRs | CIDRs only |
| Default (new custom) | Deny all inbound, allow all outbound | Deny all both ways |
| Default (VPC's default) | Allow from itself, allow all out | Allow all both ways |
The stateful/stateless difference, concretely. A client connects from port 54321 to your app on port 8080.
Security Group: allow inbound TCP 8080. Done. The reply from 8080 back to 54321 is automatically permitted because the SG remembers the connection.
NACL: allow inbound TCP 8080 and allow outbound TCP to ports 1024–65535 (the ephemeral port range), because the NACL has no memory and sees the reply as brand-new traffic in the other direction. Forget the second rule and connections hang — they establish and then appear to freeze. This is the single most common NACL bug, and its symptom (a hang, not a refusal) is why it wastes so much time.
The idiomatic pattern: Security Groups referencing Security Groups
Do not write CIDRs between your own tiers. Reference the group:
alb-sg inbound: 443 from 0.0.0.0/0
outbound: 8080 to app-sg
app-sg inbound: 8080 from alb-sg ← not a CIDR
outbound: 5432 to db-sg
443 to 0.0.0.0/0 (external APIs, AWS endpoints)
db-sg inbound: 5432 from app-sg ← not a CIDR
outbound: (none needed)Why this is better than CIDRs:
- It survives subnet changes, resizing and new AZs — nothing to update.
- It expresses intent: "the app tier may reach the database", not "10.0.16.0/20 may reach 10.0.80.0/24".
- It is self-documenting in the Console and in a review.
- It scales: adding a new app subnet requires no security change at all.
Practical guidance: use Security Groups for essentially all of your access control. Reach for NACLs only for coarse subnet-wide denies — blocking a hostile CIDR range, or enforcing "this subnet may never talk to the Internet" as a belt-and-braces control alongside the absent route. Most well-run AWS accounts leave NACLs at their defaults and do all real work in Security Groups.
4. Reaching AWS services privately
Here is a question that sounds trivial and is not: your application is in a private subnet and needs to call S3. How does the traffic get there?
By default, through the NAT Gateway, out to the Internet, and back into AWS via S3's public endpoint. Every byte pays the NAT per-GB charge. For a service that writes 5 TB of objects a month, that is ~$225/month in NAT processing alone for traffic that never needed to leave AWS.
VPC endpoints fix this. There are two kinds, and the difference is worth memorizing.
| Gateway endpoint | Interface endpoint (PrivateLink) | |
|---|---|---|
| Services | S3 and DynamoDB only | Most AWS services, plus third-party and your own |
| Mechanism | A route table entry to a prefix list | An ENI with a private IP in each subnet you choose |
| DNS | Unchanged (uses the public hostname, routed privately) | Private DNS overrides the service hostname to the ENI |
| Security | Endpoint policy | Endpoint policy + a Security Group on the ENI |
| Cost | Free | ~$0.01/hr per AZ + ~$0.01/GB processed |
| Works from on-prem (via VPN/DX)? | No | Yes |
flowchart LR
subgraph VPC["VPC"]
APP[App in private subnet]
ENI["Interface endpoint ENI<br/>10.0.10.200"]
end
APP -->|"S3, DynamoDB<br/>via route table"| GW[Gateway endpoint] --> S3[(S3)]
APP -->|"SQS, Secrets Manager,<br/>KMS, ECR, ..."| ENI --> SVC[AWS service]
APP -.->|"everything else:<br/>the expensive path"| NAT[NAT Gateway] --> INT((Internet))The rule of thumb: always add the S3 and DynamoDB gateway endpoints. They are free, they reduce your NAT bill, and they let you write an endpoint policy restricting which buckets are reachable from the VPC at all. Add interface endpoints when the traffic volume or the "must never traverse the Internet" requirement justifies the hourly cost — and do the arithmetic, because for a low-traffic service three interface endpoints across three AZs can cost more than the NAT processing they save.
5. DNS inside a VPC
Every VPC has a DNS resolver at VPC base + 2 (10.0.0.2 for 10.0.0.0/16), also reachable at the link-local address 169.254.169.253. It resolves:
- public Internet names, by recursion;
- AWS service endpoints (
s3.eu-west-1.amazonaws.com) — and returns the private endpoint address when a matching interface endpoint with private DNS exists; - names in any Route 53 private hosted zone associated with the VPC;
- internal hostnames for your resources.
Two VPC attributes control this, and both must be true for the common case:
| Attribute | Effect |
|---|---|
enableDnsSupport | The VPC resolver at base+2 works at all |
enableDnsHostnames | Instances with public IPs get public DNS names; required for interface endpoint private DNS to work |
A classic failure: you create an interface endpoint for Secrets Manager, and the application still goes out through NAT. Cause:
enableDnsHostnamesis false, so the private DNS override never took effect and the public hostname still resolves to a public address.
Route 53 adds:
| Record type | Use |
|---|---|
A / AAAA | Name → IP address |
CNAME | Name → another name. Cannot exist at a zone apex (example.com) |
| Alias | An AWS-specific record pointing at an AWS resource (ALB, CloudFront, S3 website, another Route 53 record). Works at the apex, resolves to the resource's current addresses, and is free to query |
Use alias records for AWS resources, always. An ALB's IP addresses change; an alias tracks them. A CNAME to an ALB works but cannot be used at the apex and costs per query.
Private hosted zones give you internal names (orders.internal.acme.com → 10.0.10.7) resolvable only from associated VPCs. Routing policies — weighted, latency, failover — turn DNS into a deployment and failover tool, which we use on Day 10 and Day 12.
Architecture
The reference VPC — the one every later day assumes
flowchart TB
U((Users)) --> R53[Route 53<br/>alias → ALB]
R53 --> IGW[Internet Gateway]
subgraph VPC["VPC 10.0.0.0/16 · eu-west-1"]
subgraph AZA["AZ eu-west-1a"]
PUBA["public-a 10.0.0.0/24"]
ALBA[ALB node]
NATA[NAT GW a]
PRIA["private-a 10.0.16.0/20"]
APPA[App tasks]
DATA["data-a 10.0.80.0/24"]
DBA[(RDS primary)]
end
subgraph AZB["AZ eu-west-1b"]
PUBB["public-b 10.0.1.0/24"]
ALBB[ALB node]
NATB[NAT GW b]
PRIB["private-b 10.0.32.0/20"]
APPB[App tasks]
DATB["data-b 10.0.81.0/24"]
DBB[(RDS standby)]
end
GWE[S3/DynamoDB<br/>gateway endpoint]
end
IGW --> ALBA & ALBB
ALBA --> APPA
ALBB --> APPB
APPA --> DBA
APPB --> DBA
DBA -.sync replication.-> DBB
APPA --> NATA --> IGW
APPB --> NATB --> IGW
APPA & APPB --> GWERead the properties off the picture:
- Losing
eu-west-1aentirely: the ALB stops sending traffic to AZ a, tasks in AZ b serve everything, RDS fails over to the standby. Degraded capacity, not an outage. - Nothing in
private-*ordata-*has a public address or an inbound path from the IGW. - Traffic to S3 and DynamoDB bypasses NAT entirely (and its per-GB charge).
- Each AZ's egress stays inside that AZ — no cross-AZ charges on outbound traffic.
Request flow, annotated
sequenceDiagram
autonumber
participant C as Client
participant R as Route 53
participant I as Internet Gateway
participant A as ALB (public subnet)
participant T as Task (private subnet)
participant D as RDS (data subnet)
C->>R: resolve api.acme.com
R-->>C: alias → ALB addresses
C->>I: TLS to ALB public IP :443
I->>A: 1:1 NAT to the ALB node's private address
Note over A: alb-sg allows 443 from 0.0.0.0/0
A->>T: HTTP :8080 to the task's private IP
Note over T: app-sg allows 8080 from alb-sg
T->>D: TCP :5432
Note over D: db-sg allows 5432 from app-sg
D-->>T: rows
T-->>A: 200
A-->>C: 200 (TLS terminated at the ALB)Notice that the client never learns any private address, and the database's Security Group names a group, not a range.
How It Works Internally
What happens to a packet leaving a private subnet
flowchart TD
P["Packet from 10.0.16.7 → 54.239.28.85:443"] --> SGO{"Security Group<br/>outbound rule?"}
SGO -->|no| DROP1[Dropped silently]
SGO -->|yes| NO{"NACL outbound<br/>rule, lowest number first"}
NO -->|deny| DROP2[Dropped silently]
NO -->|allow| RT{"Route table:<br/>longest prefix match"}
RT -->|"10.0.0.0/16 → local"| LOCAL[Stays in the VPC]
RT -->|"pl-xxxx S3 → gateway endpoint"| GWE[AWS backbone, no NAT charge]
RT -->|"0.0.0.0/0 → nat-xyz"| NAT[NAT Gateway rewrites source]
RT -->|no matching route| BLACKHOLE[Dropped: no route]
NAT --> IGWO[Internet Gateway] --> NET((Internet))
NET --> RET["Reply arrives"]
RET --> NACLI{"NACL inbound:<br/>ephemeral ports allowed?"}
NACLI -->|deny| HANG["Connection hangs —<br/>the classic NACL bug"]
NACLI -->|allow| SGI["Security Group: stateful,<br/>return traffic auto-allowed"]
SGI --> DONE[Delivered]Three things to take from this diagram:
- Drops are silent. Neither Security Groups nor NACLs send a rejection. You get a timeout, not a refusal. A connection that is refused quickly is usually your application not listening; a connection that hangs is usually a network rule.
- Routing happens after filtering on the way out, and the longest prefix wins. A gateway endpoint inserts a more-specific prefix-list route, which is precisely how it steals S3 traffic away from the NAT route.
- The return path is where stateless bites. Everything in the outbound direction can be perfect and the reply still gets dropped by an inbound NACL rule.
The 1:1 NAT performed by the Internet Gateway
flowchart LR
subgraph OS["Instance operating system"]
NIC["eth0: 10.0.0.42<br/>(this is all the OS sees)"]
end
NIC --> IGW["Internet Gateway<br/>rewrites 10.0.0.42 ⇄ 54.72.10.9"]
IGW --> NET((Internet))This is why curl ifconfig.me from an instance returns an address you cannot find anywhere in ip addr, and why code that needs its own public address must query the instance metadata service instead of inspecting its interfaces.
Code
Networking is configured, not programmed — so today's "code" is CLI. These commands are the ones you will actually reuse when investigating.
Build the reference VPC
REGION=eu-west-1
AZ_A=${REGION}a
AZ_B=${REGION}b
# 1. VPC
VPC_ID=$(aws ec2 create-vpc --cidr-block 10.0.0.0/16 \
--tag-specifications 'ResourceType=vpc,Tags=[{Key=Name,Value=course-vpc},{Key=Project,Value=aws-course}]' \
--query Vpc.VpcId --output text)
# DNS attributes — both must be on for private endpoint DNS to work later
aws ec2 modify-vpc-attribute --vpc-id "$VPC_ID" --enable-dns-support
aws ec2 modify-vpc-attribute --vpc-id "$VPC_ID" --enable-dns-hostnames
# 2. Subnets — note the deliberate sizing: small public, large private
PUB_A=$(aws ec2 create-subnet --vpc-id "$VPC_ID" --cidr-block 10.0.0.0/24 --availability-zone "$AZ_A" --query Subnet.SubnetId --output text)
PUB_B=$(aws ec2 create-subnet --vpc-id "$VPC_ID" --cidr-block 10.0.1.0/24 --availability-zone "$AZ_B" --query Subnet.SubnetId --output text)
PRI_A=$(aws ec2 create-subnet --vpc-id "$VPC_ID" --cidr-block 10.0.16.0/20 --availability-zone "$AZ_A" --query Subnet.SubnetId --output text)
PRI_B=$(aws ec2 create-subnet --vpc-id "$VPC_ID" --cidr-block 10.0.32.0/20 --availability-zone "$AZ_B" --query Subnet.SubnetId --output text)
DAT_A=$(aws ec2 create-subnet --vpc-id "$VPC_ID" --cidr-block 10.0.80.0/24 --availability-zone "$AZ_A" --query Subnet.SubnetId --output text)
DAT_B=$(aws ec2 create-subnet --vpc-id "$VPC_ID" --cidr-block 10.0.81.0/24 --availability-zone "$AZ_B" --query Subnet.SubnetId --output text)
# 3. Internet Gateway (free)
IGW_ID=$(aws ec2 create-internet-gateway --query InternetGateway.InternetGatewayId --output text)
aws ec2 attach-internet-gateway --vpc-id "$VPC_ID" --internet-gateway-id "$IGW_ID"
# 4. Public route table — THIS is what makes a subnet public
RT_PUB=$(aws ec2 create-route-table --vpc-id "$VPC_ID" --query RouteTable.RouteTableId --output text)
aws ec2 create-route --route-table-id "$RT_PUB" --destination-cidr-block 0.0.0.0/0 --gateway-id "$IGW_ID"
aws ec2 associate-route-table --route-table-id "$RT_PUB" --subnet-id "$PUB_A"
aws ec2 associate-route-table --route-table-id "$RT_PUB" --subnet-id "$PUB_B"
# 5. NAT Gateway in public-a ⚠️ BILLABLE from this moment (~$0.045/hr + $0.045/GB)
EIP_ALLOC=$(aws ec2 allocate-address --domain vpc --query AllocationId --output text)
NAT_ID=$(aws ec2 create-nat-gateway --subnet-id "$PUB_A" --allocation-id "$EIP_ALLOC" \
--query NatGateway.NatGatewayId --output text)
aws ec2 wait nat-gateway-available --nat-gateway-ids "$NAT_ID"
# 6. Private route table → NAT
# Production uses one NAT and one route table PER AZ. We use one to halve the lab cost.
RT_PRI=$(aws ec2 create-route-table --vpc-id "$VPC_ID" --query RouteTable.RouteTableId --output text)
aws ec2 create-route --route-table-id "$RT_PRI" --destination-cidr-block 0.0.0.0/0 --nat-gateway-id "$NAT_ID"
aws ec2 associate-route-table --route-table-id "$RT_PRI" --subnet-id "$PRI_A"
aws ec2 associate-route-table --route-table-id "$RT_PRI" --subnet-id "$PRI_B"
# 7. Data subnets: no route table association → they use the main table,
# which has only the local route. No path out. That is the point.
# 8. S3 gateway endpoint — free, and it takes S3 traffic off the NAT
aws ec2 create-vpc-endpoint --vpc-id "$VPC_ID" \
--service-name "com.amazonaws.${REGION}.s3" \
--route-table-ids "$RT_PRI"Security Groups that reference each other
ALB_SG=$(aws ec2 create-security-group --group-name alb-sg --description "ALB" --vpc-id "$VPC_ID" --query GroupId --output text)
APP_SG=$(aws ec2 create-security-group --group-name app-sg --description "App tier" --vpc-id "$VPC_ID" --query GroupId --output text)
DB_SG=$(aws ec2 create-security-group --group-name db-sg --description "Data tier" --vpc-id "$VPC_ID" --query GroupId --output text)
aws ec2 authorize-security-group-ingress --group-id "$ALB_SG" \
--protocol tcp --port 443 --cidr 0.0.0.0/0
# The important line: source is a GROUP, not a CIDR.
aws ec2 authorize-security-group-ingress --group-id "$APP_SG" \
--protocol tcp --port 8080 --source-group "$ALB_SG"
aws ec2 authorize-security-group-ingress --group-id "$DB_SG" \
--protocol tcp --port 5432 --source-group "$APP_SG"The investigation commands
# What route table applies to this subnet, and what is in it?
aws ec2 describe-route-tables \
--filters "Name=association.subnet-id,Values=${PRI_A}" \
--query 'RouteTables[].Routes[].{Dest:DestinationCidrBlock,Target:GatewayId,NAT:NatGatewayId,State:State}' \
--output table
# What can reach this instance?
aws ec2 describe-security-groups --group-ids "$APP_SG" \
--query 'SecurityGroups[].IpPermissions[].{Port:FromPort,Src:UserIdGroupPairs[].GroupId,Cidr:IpRanges[].CidrIp}'
# Does the NACL allow the return traffic?
aws ec2 describe-network-acls --filters "Name=association.subnet-id,Values=${PRI_A}" \
--query 'NetworkAcls[].Entries[].{Num:RuleNumber,Egress:Egress,Proto:Protocol,Ports:PortRange,Cidr:CidrBlock,Action:RuleAction}' \
--output table
# Let AWS do the analysis: Reachability Analyzer answers "can A reach B, and if not, what blocked it"
aws ec2 create-network-insights-path \
--source "$SOURCE_ENI" --destination "$DEST_ENI" --protocol tcp --destination-port 5432Reachability Analyzer is the tool most developers never discover. It statically analyses your route tables, Security Groups, NACLs and gateways and tells you the exact component that blocks a path — no packets sent, works even when nothing is running. It costs a small per-analysis fee and saves hours.
Hands-on
Lab 2 — Build the VPC and prove its asymmetry (30 min)
Objective: stand up the reference three-tier VPC, put a host in a private subnet, and prove it can reach the Internet while the Internet cannot reach it.
Prerequisites: Lab 0/1. Permissions: ec2:* on VPC resources, ssm:StartSession.
⚠️ Billable: the NAT Gateway (~$0.045/hr) and the Elastic IP. Total for a 2-hour lab: well under $1. Delete both at the end — an orphaned NAT Gateway is ~$32/month.
Architecture: the reference VPC above, with one t3.micro in private-a and no bastion — you will reach it with SSM Session Manager, which needs no inbound rule and no SSH key at all.
Steps:
- Run the CLI block above through step 8.
- Create an IAM role for the instance with the AWS managed policy
AmazonSSMManagedInstanceCore, and an instance profile from it. (This is the Day 3 instance-role mechanism arriving one day early — today just use it.) - Launch a
t3.microinprivate-a, withapp-sg, no public IP, the instance profile attached, and a recent Amazon Linux 2023 AMI. - Wait for it to register with SSM (a minute or two), then connect:
aws ssm start-session --target "$INSTANCE_ID"SSM works here because the instance reaches the SSM service outbound through the NAT Gateway. No inbound rule exists, and none is needed — this is the asymmetry in action. (In a NAT-free design you would add interface endpoints for
ssm,ssmmessagesandec2messagesinstead.)
Verification — three tests that each prove one thing:
# Inside the SSM session:
# (a) Outbound to the Internet works — via NAT
curl -s https://checkip.amazonaws.com
# → the NAT Gateway's Elastic IP, NOT the instance's address
# (b) The instance has no public address of its own
ip -4 addr show | grep inet
# → only 10.0.16.x
# (c) S3 traffic bypasses NAT entirely — it matched the endpoint's prefix-list route
aws s3 ls# From your laptop:
nc -vz -w 5 10.0.16.7 8080 # times out: the address is not routable from here
nc -vz -w 5 <any public IP you might hope for> 8080 # there isn't oneThen make the failure deliberate, so you can recognize it later:
# Remove the private subnet's default route and watch egress die
aws ec2 delete-route --route-table-id "$RT_PRI" --destination-cidr-block 0.0.0.0/0
# In the session: curl https://checkip.amazonaws.com → hangs, then times out
# `aws s3 ls` still works — the endpoint route is separate. That is the lesson.
aws ec2 create-route --route-table-id "$RT_PRI" --destination-cidr-block 0.0.0.0/0 --nat-gateway-id "$NAT_ID"Expected output: checkip returns the NAT's EIP; ip addr shows only a private address; your laptop times out; with the route deleted, general egress fails but S3 does not.
Cleanup — in this order (dependencies matter):
aws ec2 terminate-instances --instance-ids "$INSTANCE_ID"
aws ec2 wait instance-terminated --instance-ids "$INSTANCE_ID"
aws ec2 delete-nat-gateway --nat-gateway-id "$NAT_ID"
aws ec2 wait nat-gateway-deleted --nat-gateway-ids "$NAT_ID"
aws ec2 release-address --allocation-id "$EIP_ALLOC"
# Then: endpoints, route tables, subnets, IGW (detach first), security groups, VPC.Keep the VPC if you plan to continue straight into Day 3 — it is free once the NAT Gateway and Elastic IP are gone. The NAT Gateway and the Elastic IP are the only meaningful costs; delete those today regardless.
Common errors:
| Symptom | Cause | Fix |
|---|---|---|
InvalidSubnet.Conflict | Overlapping subnet CIDRs | Recheck your ranges |
| Instance never appears in SSM | No route out (NAT missing/not ready), or no instance profile, or the AMI has no SSM agent | Check all three, in that order |
curl hangs instead of failing fast | A network rule dropped it silently — route, SG or NACL | Reachability Analyzer |
DependencyViolation on delete | Something still uses the resource | Delete in dependency order |
NAT Gateway stuck in pending for minutes | Normal — it takes a few minutes | aws ec2 wait nat-gateway-available |
Production relevance: this is, structurally, the network every application in this course runs in. Days 4, 5 and 6 place ECS tasks, Lambdas and databases into exactly these tiers.
The connectivity debugging checklist
When something cannot reach something, do not guess. Work down this list in order; each step eliminates a layer.
flowchart TD
S[A cannot reach B] --> DNS{"Does the name<br/>resolve?"}
DNS -->|no| DNSFIX["DNS: enableDnsSupport,<br/>private hosted zone,<br/>endpoint private DNS"]
DNS -->|yes| RT{"Is there a route<br/>from A's subnet to B?"}
RT -->|no| RTFIX["Route table: missing default route,<br/>wrong table associated,<br/>blackhole route"]
RT -->|yes| NACL{"NACL: allowed<br/>OUT from A and<br/>IN to B — and the<br/>ephemeral return?"}
NACL -->|no| NACLFIX["Add the return rule.<br/>This is the hang, not the refusal."]
NACL -->|yes| SG{"Security Groups:<br/>A outbound, B inbound<br/>on the right port?"}
SG -->|no| SGFIX["Reference the source SG,<br/>check the port, check the protocol"]
SG -->|yes| LISTEN{"Is B actually<br/>listening on that port?"}
LISTEN -->|no| APPFIX["ss -lntp on B.<br/>Bound to 127.0.0.1<br/>instead of 0.0.0.0?"]
LISTEN -->|yes| HEALTH["Health checks, TLS,<br/>application-level auth<br/>→ Day 4"]The two symptoms and what they mean:
| Symptom | Almost always |
|---|---|
| Connection times out / hangs | Something dropped the packet silently: route, Security Group, or NACL |
| Connection refused immediately | The network is fine. Nothing is listening on that port — or it is bound to 127.0.0.1 |
That distinction alone resolves a large fraction of real incidents in seconds.
Production Considerations
| Concern | Production practice |
|---|---|
| AZ coverage | Subnets in at least 3 AZs where the Region offers them. Two is the minimum; three survives one AZ loss without halving capacity. |
| NAT per AZ | One NAT Gateway per AZ, each AZ's private route table pointing at its own. Costs ~3×, removes a single point of failure and cross-AZ data charges. |
| Address planning | Allocate VPC ranges centrally across the organization so nothing ever overlaps. You cannot peer overlapping networks, and renumbering is brutal. |
| Subnet sizing | Private tier generously (/20 or larger). Autoscaling that cannot place tasks because a subnet is out of addresses is a real outage. |
| Endpoints | S3 and DynamoDB gateway endpoints always. Interface endpoints where volume or compliance justifies them. |
| IaC | Nobody builds a real VPC by CLI. Terraform/CDK/CloudFormation — but knowing what the tool creates is exactly what today taught you. |
| Flow Logs | Enable to S3 or CloudWatch for the VPC. They are how you prove what was or was not attempted. Sample in high-traffic accounts; the ingestion cost is real. |
| Public IPv4 cost | AWS charges per public IPv4 address per hour (~$0.005). At scale this is a line item; it is also an argument for keeping things private. |
Failure Scenarios
S3 — "The app can't reach the database"
One symptom, six plausible causes. This scenario is run as an exercise: the learner is given the symptom and must name the discriminating check for each hypothesis.
| # | Cause | The check that confirms it |
|---|---|---|
| 1 | db-sg has no rule allowing app-sg on 5432 | describe-security-groups on db-sg |
| 2 | The database is in a subnet whose NACL denies the ephemeral return range | describe-network-acls; symptom is a hang |
| 3 | The app is in a subnet with no route to the data subnet — impossible within a VPC (the local route), so really: the database is in a different VPC | Compare VPC IDs |
| 4 | Wrong endpoint — the app has the old primary's address after a failover | Resolve the RDS endpoint name; compare |
| 5 | The database is not listening / still starting | RDS status; connection refused not timeout |
| 6 | Credentials or pg_hba-level rejection | An authentication error, not a network error — read the message |
The lesson: the error text distinguishes these. Timeout = layers 1–3. Refused = layer 5. Authentication failed = layer 6. Read the error before forming a theory.
S4 — $1,400 of NAT Gateway charges for a service that only talks to S3
An image-processing service in private subnets reads and writes ~15 TB/month to S3.
| Question | Answer |
|---|---|
| What can fail? | Nothing, technically — it works perfectly. The failure is economic. |
| What happens? | 15 TB × |
| Recovery? | Add an S3 gateway endpoint. Free. Takes two minutes. |
| How quickly? | Immediate — the more-specific prefix-list route wins as soon as it exists |
| Data lost / duplicated? | No |
The lesson, and it is a Day 1-through-14 lesson: in AWS, an architecture can be entirely correct and still be wrong. Cost is a design property. A route table entry that costs nothing eliminated 70% of this service's infrastructure bill.
How you would have caught it: Cost Explorer grouped by usage type, where NatGateway-Bytes appears as a distinct line. Day 12 makes this a routine.
A1 (preview) — Single-AZ everything
Covered fully on Day 4, but visible already: if your NAT Gateway, your tasks and your database are all in eu-west-1a, then eu-west-1a is your entire availability story. The subnets in AZ b cost nothing to create. Create them on day one, before there is anything to migrate.
Security
Applying the four questions to today's network:
| Question | Answer |
|---|---|
| Who can access this? | Inbound: only what a Security Group explicitly allows, only in public subnets, only through the IGW. The data tier has no inbound path from outside the VPC at all. |
| What credentials are used? | None at the network layer — this is pure network authorization. Which is exactly why it is defence in depth: it holds even when a credential leaks. |
| Where is data encrypted? | In transit: your responsibility. TLS terminates at the ALB (Day 4); traffic inside the VPC is not automatically encrypted, though AWS does encrypt some inter-AZ traffic at the physical layer. Encrypt app↔DB connections explicitly. |
| What if credentials leak? | Network controls limit what a leaked credential can reach from where. An endpoint policy on your S3 gateway endpoint can restrict VPC traffic to your own buckets — a genuinely useful control against exfiltration. |
Two security postures to know the names of:
- Least privilege at the network layer — Security Groups that reference groups, not ranges; no
0.0.0.0/0anywhere except the ALB's:443. - No path is better than a blocked path — the data tier's lack of a default route is stronger than any deny rule, because there is nothing to misconfigure.
Performance
| Factor | Impact |
|---|---|
| Cross-AZ latency | Single-digit milliseconds. Fine for a database call; expensive if a request makes 30 cross-AZ hops. Day 12 covers keeping call graphs AZ-local. |
| NAT Gateway throughput | Scales to tens of Gbps automatically, but it is a shared choke point and its connection tracking has limits. High-connection-churn workloads can exhaust NAT ports (ErrorPortAllocation in CloudWatch). |
| Instance network bandwidth | Varies by instance size; small instances are bandwidth-limited and burstable. A t3.micro cannot saturate a 10 Gbps link. |
| Endpoints | Gateway endpoints keep traffic on the AWS backbone — generally lower latency and more consistent than the public path. |
| DNS TTL | Client-side caching of AWS endpoint names is a correctness issue at failover time, not a performance one — but it is a networking behaviour. Day 6 has the JVM-specific trap. |
Cost
| Component | Approx. cost (us-east-1, verify current) |
|---|---|
| VPC, subnets, route tables, Security Groups, NACLs | Free |
| Internet Gateway | Free (you pay for the data transfer, not the gateway) |
| NAT Gateway | |
| Elastic IP attached to a running resource | Free; charged when idle |
| Public IPv4 address | |
| Gateway endpoint (S3, DynamoDB) | Free |
| Interface endpoint | ~$0.01/hr per AZ + ~$0.01/GB |
| Cross-AZ data transfer | ~$0.01/GB each way |
| Data transfer out to Internet | First 100 GB/mo free, then ~$0.09/GB |
| VPC Flow Logs | No charge for the feature; you pay CloudWatch Logs or S3 ingestion + storage |
Today's cost lesson: the NAT Gateway is the most commonly underestimated line item in a backend AWS bill. It charges by the hour whether or not it is used, three times over if you do it correctly for HA, and again per gigabyte. Two habits, learned today, prevent most of that pain:
- Add the free S3 and DynamoDB gateway endpoints to every private route table, always.
- Before adding an interface endpoint, do the arithmetic:
hours × AZs × $0.01vsGB × $0.045saved. For low-volume services the NAT is genuinely cheaper.
Alternatives
| Instead of | You could | Trade-off |
|---|---|---|
| A custom VPC | The default VPC (172.31.0.0/16, all subnets public) | Fine for a throwaway experiment. Every subnet is public, so nothing is protected by design. Never production. |
| NAT Gateway | NAT instance (an EC2 instance doing NAT) | Cheaper at tiny scale; you now own patching, HA and throughput. AWS's own guidance is to use the gateway. |
| NAT Gateway | Interface endpoints for everything | Eliminates NAT entirely if your service only talks to AWS. Hourly cost per endpoint per AZ; worth modelling. |
| Security Groups | NACLs for primary access control | NACLs are stateless and rule-ordered — harder to reason about, and cannot reference groups. Use them for coarse denies only. |
| Route 53 | Any DNS provider | External DNS works, but you lose alias records, health-check-driven failover, and private hosted zones. |
| Multi-AZ | Single AZ | Cheaper (one NAT, no cross-AZ transfer) and strictly worse. Acceptable for dev, never for production. |
Trade-offs
| Decision | Gain | Cost |
|---|---|---|
| Three tiers instead of two | The data tier physically cannot reach the Internet | More subnets, route tables, and mental overhead |
| NAT per AZ | No single point of failure; no cross-AZ egress charges | ~3× the NAT cost |
| Large private subnets | Room to scale | Address space consumed; harder to sub-allocate later |
| SG-referencing-SG | Self-documenting, survives topology change | Slightly less obvious to someone reading a CIDR-based diagram |
| Interface endpoints | Traffic never leaves AWS; on-prem reachable | Hourly cost per AZ per service; can exceed the NAT it replaces |
| Big VPC CIDR (/16) | Never run out | Consumes organizational address space; overlap risk in a merger |
| Private-only compute | Dramatically smaller attack surface | You need SSM/bastion to debug, and endpoints or NAT for egress |
Common Mistakes
- Believing a subnet has a "public" setting. A route to an IGW is what makes it public. Nothing else.
- Forgetting the NACL return rule. Stateless means the reply is separate traffic. The symptom is a hang, which is why it costs so much time.
- Writing CIDRs between your own tiers. Reference Security Groups.
- Undersizing private subnets. Autoscaling failing on address exhaustion during a traffic spike is a genuine outage mode.
- One NAT Gateway for a "multi-AZ" architecture. The other AZs lose egress when it fails, and pay cross-AZ charges meanwhile.
- Leaving S3 traffic on the NAT path. Free endpoint, immediate saving, no downside.
- Overlapping CIDRs across VPCs or with on-prem. Discovered during an acquisition, fixed by renumbering. Plan centrally.
- Putting a database in a public subnet "because it was easier." Then adding a Security Group and calling it safe. One misconfiguration away from exposure — versus zero routes away.
- Forgetting
enableDnsHostnamesand then wondering why the interface endpoint did nothing. - Deleting resources in the wrong order and fighting
DependencyViolationfor twenty minutes. Instances → NAT → EIP → endpoints → route tables → subnets → IGW → SGs → VPC.
Interview Questions
2–3 YOE
- What is a VPC, and what is a subnet?
- What makes a subnet public?
- What is the difference between a Security Group and a NACL?
- Why can't you reach an instance at
10.0.1.5from your laptop? - How many usable IP addresses are in a
/24subnet in AWS, and why isn't it 256? - What does a NAT Gateway do?
4–5 YOE
- Your application in a private subnet can reach S3 but not an external payment API. What do you check, and in what order?
- A connection to your database hangs instead of being refused. What does that tell you, and what does it rule out?
- Why is a Security Group that references another Security Group better than one that references a CIDR?
- A service writes 10 TB/month to S3 from a private subnet. Where is the money going, and how do you fix it?
- What is the difference between a gateway endpoint and an interface endpoint, and when would you pay for the second?
Senior-Level Questions
6–7 YOE
- Design the VPC layout for a platform with three environments and a requirement that production be network-isolated from everything else. How many VPCs, how many accounts, and why?
- Your company is acquiring another company whose VPCs use the same
10.0.0.0/16you do. What are your options, and what does each cost? - When would you choose to eliminate NAT Gateways entirely, and what would you need in place first?
- Walk me through how you would give an on-premises data centre private access to a service running in your VPC — and the reverse.
- How do you enforce, organization-wide, that no database is ever placed in a public subnet?
8–10 YOE
- Your architecture is "multi-AZ" but an AZ impairment caused a full outage anyway. Give three plausible causes involving networking, and say how you would detect each in advance.
- Argue for and against a single large VPC shared by 40 microservices, versus a VPC per service. Then decide, and name what would change your mind.
- Cross-AZ data transfer is 22% of your bill. Describe your investigation and the architectural changes you would consider — including the ones you would reject and why.
- What network-layer controls actually reduce the impact of a compromised application credential, and which are theatre?
Day-End Revision
The five sentences
- A VPC is your private network in one Region; a subnet lives in exactly one AZ, which is why multi-AZ means multi-subnet.
- A subnet is public if and only if its route table has
0.0.0.0/0→ Internet Gateway. - Security Groups are stateful, allow-only, and can reference other groups; NACLs are stateless, ordered, allow-and-deny, and need explicit return rules.
- NAT Gateway gives outbound-only Internet access to private subnets, is zonal, and charges hourly plus per gigabyte.
- Gateway endpoints for S3 and DynamoDB are free and take that traffic off the NAT path.
The diagram to redraw from memory: the reference three-tier VPC across two AZs, with route targets labelled.
The numbers
| Reserved IPs per subnet | 5 |
| Usable in a /24 | 251 |
| VPC CIDR size range | /16 to /28 |
| VPC DNS resolver | base + 2 |
| NAT Gateway | ~$0.045/hr + ~$0.045/GB |
| Cross-AZ transfer | ~$0.01/GB each way |
| Gateway endpoint | free |
Today's trap: "It's in a private subnet, so it's safe." Private means no route from an Internet Gateway. It is still reachable via VPN, peering, PrivateLink, a compromised bastion, or a load balancer someone attached — and it can still reach outward through NAT.
Tomorrow needs: subnets (you will place an EC2 instance in one), Security Groups (the instance needs one), and the SSM access pattern you used today.
Mini Assignment
Time: 40–50 minutes.
- In
day-02/, write a shell scriptvpc-audit.shthat, for a given VPC ID, prints:- every subnet with its CIDR, AZ, usable address count, and whether it is public (determined by inspecting its route table, not by its name);
- every route table with its routes and associated subnets, flagging any
blackholeroute; - every Security Group rule, marking any rule with source
0.0.0.0/0as ⚠️; - whether S3 and DynamoDB gateway endpoints exist.
- Run it against the VPC you built. Then run it against your account's default VPC and compare.
- In
day-02/NOTES.md:- Which subnets in the default VPC are public, and what does that imply about using it for anything real?
- You need to support 3,000 concurrent ECS tasks across 3 AZs, each task consuming one IP. What subnet sizes do you choose, and how much headroom do you leave?
- Your service transfers 8 TB/month to S3 and 200 GB/month to an external API. Compute the monthly NAT cost with and without an S3 gateway endpoint.
- Commit.
Success criterion: your script determines "public" from the route table, and your NOTES arithmetic is correct to within a few dollars.
AWS Documentation
- VPC User Guide
- VPC and subnet sizing
- Route tables
- NAT Gateways
- Security Groups · Network ACLs
- VPC endpoints
- DNS attributes for your VPC
- Reachability Analyzer
- Route 53 developer guide
Previous: Day 1 — Cloud, the AWS Account, and Identity · Next: Day 3 — Compute, Storage and Seeing What Happened