Diagram Plan
Every major concept gets at least one diagram. Diagrams are Mermaid, they are explained in prose immediately after, and they are never decorative.
1. Rules
- Mermaid by default. ASCII art only where Mermaid genuinely cannot express the idea (rare; a couple of layout sketches).
- A diagram must teach something the prose cannot. If the paragraph above it already says everything, the diagram is cut.
- Every box is a term the reader already knows — or is being defined right there. No unexplained boxes.
- Consistent vocabulary and shapes across all 14 days (§4).
- Every diagram is followed by an explanation, walking the reader through it in the order the data moves.
- Complex architectures ship as a set of four, never one crowded picture (§3).
- Technically correct. If TLS terminates at the ALB, the diagram shows it terminating at the ALB.
2. Diagram types and where each is used
| Type | Mermaid | Used for | Example day |
|---|---|---|---|
| Architecture / component | flowchart TB with subgraphs | System shape, VPC layout, service topology | D2, D4, D14 |
| Request lifecycle | sequenceDiagram | What happens in order, across components, with timing | D1 (signed API call), D5 (API GW → Lambda), D8 (SQS receive) |
| Data flow | flowchart LR | Where bytes go and who transforms them | D9, D10 |
| Internal working | flowchart TB + stateDiagram-v2 | What AWS does inside a service | D5 (execution environment), D8 (message states) |
| State machine | stateDiagram-v2 | Message visibility, order status, circuit breaker, deploy states | D8, D12, D14 |
| Decision tree | flowchart TD with diamonds | Service selection | D4, D7, D9, D12 |
| Failure flow | flowchart TB with a marked failure point and the recovery path | What breaks and what happens next | Every day's Failure Scenarios |
| Comparison | Table + a paired two-panel flowchart | Sync vs async, queue vs stream, SG vs NACL | D7, D8, D9 |
| Timeline / Gantt | gantt | Failover timelines, cold-start budget, deploy sequences | D6, D5, D13 |
3. The four-diagram rule for complex architecture
Any architecture with more than ~6 components is presented as four diagrams, in this order:
flowchart LR
A["1 · High-level<br/>boxes a stakeholder understands"] --> B["2 · Detailed<br/>subnets, roles, policies, endpoints"]
B --> C["3 · Request / data flow<br/>numbered sequence of one real operation"]
C --> D["4 · Failure flow<br/>the same path with a component removed"]This applies to: Day 4 (ECS+ALB), Day 5 (API GW+Lambda), Day 9 (event-driven), Day 10 (upload platform), Day 12 (multi-AZ/DR), Day 14 (all seven system designs + capstone).
4. Visual conventions (fixed for all 14 days)
| Element | Convention |
|---|---|
| Client / user | C[Client] — plain rectangle, leftmost or topmost |
| AWS managed service | plain rectangle with the service name, e.g. SQS[SQS orders-queue] |
| Datastore | cylinder: DB[(Aurora orders)] |
| Queue / topic / stream | cylinder with the primitive named: Q[(SQS)], S[(Kinesis)] |
| Your code | rectangle prefixed with the service role: OS[Order Service] |
| Network boundary | subgraph — VPC, then subnet tiers nested inside |
| Synchronous call | solid arrow --> |
| Asynchronous / eventual | dotted arrow -.-> |
| Failure point | X node or an edge label `--> |
| Recovery path | dashed arrow annotated with the mechanism (retry, DLQ, failover) |
| Numbered flows | edge labels `--> |
Direction: LR for pipelines and lifecycles, TB for architecture with network tiers.
5. Per-day diagram inventory
| Day | Diagrams (minimum) |
|---|---|
| 1 | Shared responsibility boundary · Region/AZ/edge hierarchy · the signed API call sequence (SDK → credential chain → SigV4 → endpoint → IAM → service) · IAM entity relationships · Console/CLI/SDK converging on one API |
| 2 | Laptop→Internet→AWS (the naive model) · VPC CIDR carve-up · public vs private subnet routing · SG (stateful) vs NACL (stateless) two-panel · NAT Gateway traffic flow · gateway vs interface endpoint · the connectivity debugging decision tree |
| 3 | EC2 + EBS + AMI relationship · IMDS credential delivery sequence · S3 bucket/key/object model · flat namespace vs "folders" · CloudWatch three-pillar flow · vertical vs horizontal scaling |
| 4 | ASG + launch template + health check loop · ALB listener→rule→target group→target · ALB request lifecycle sequence (incl. TLS termination and XFF) · deregistration/draining timeline · ECS cluster/service/task/task-def hierarchy · task role vs execution role · rolling deploy state diagram · compute decision tree (EC2/ECS/Fargate/EKS/Lambda) |
| 5 | Lambda execution environment lifecycle (cold vs warm, init vs invoke) · concurrency and throttling · the three invocation models with their different retry behaviour · API Gateway request lifecycle sequence · REST vs HTTP vs WebSocket capability map · ALB vs API Gateway decision tree |
| 6 | RDS Multi-AZ (sync standby) vs read replica (async) two-panel · failover timeline gantt (detection → promotion → DNS → client reconnect) · Aurora shared-storage architecture · connection pool exhaustion flow · Secrets Manager retrieval sequence · expand/contract migration flow |
| 7 | Hash(partition key) → partition placement · partition + sort key range read · GSI vs LSI two-panel · hot partition failure flow · RDS vs DynamoDB decision tree · cache-aside sequence · cache stampede failure flow · write-through vs cache-aside two-panel |
| 8 | Why a queue (before/after coupling) · SQS message state diagram (visible → in-flight → deleted/returned → DLQ) · visibility timeout timeline · long vs short polling · FIFO message-group parallelism · idempotency decision + conditional-write sequence · outbox pattern · SNS→SQS fan-out · SQS vs SNS two-panel |
| 9 | Command vs event · choreography vs orchestration two-panel · EventBridge rule matching and fan-out · Kinesis shard/partition-key/consumer model · iterator age failure flow · queue vs stream vs bus comparison · Step Functions saga with compensation |
| 10 | S3 storage class decision tree · multipart upload sequence · presigned URL sequence (who signs, who uploads, who validates) · S3 event → consumer fan-out · object vs block vs file two-panel · CloudFront POP/origin/cache-key flow · cache key composition failure · Route 53 routing policies as deployment tools · data transfer cost map (the diagram with dollar signs on the edges) |
| 11 | Policy evaluation flowchart (explicit deny → SCP → resource → identity → boundary → session → implicit deny) · AssumeRole/STS sequence · cross-account trust · the who am I → what action → which resource → allow or deny chain · envelope encryption sequence · TLS termination points across every architecture built so far · Cognito OIDC flow · WAF placement |
| 12 | Scaling the tiers in order · scaling signal comparison · retry storm / metastable failure flow · circuit breaker state diagram · bulkhead isolation · DR strategy ladder (RTO/RPO vs cost) · multi-AZ vs multi-Region · blast radius / failure domains · the anatomy of a surprise bill (five cost-leak flows) |
| 13 | Correlation ID propagation across an async hop · EMF metric pipeline · X-Ray service map (annotated) · the debugging decision tree (symptom → signal → component) · one worked investigation per troubleshooting scenario · strangler fig migration · dual-write + backfill + cutover timeline |
| 14 | Four diagrams × 7 systems (28) + four for the capstone · the 13-step design method as a flow |
Total planned diagrams: ~150.
6. Internal-working diagrams (mandatory set)
These ten must exist and must show what AWS actually does, not what the docs summary says:
| Operation | Day | Must show |
|---|---|---|
| Signed AWS API call | 1 | credential chain → SigV4 → endpoint → IAM eval → service |
| IAM authorization decision | 11 | full evaluation order with all policy types |
| Lambda invocation | 5 | microVM, init vs invoke phases, freeze/thaw/reap, billing boundaries |
| API Gateway request | 5 | authorizer → mapping → integration → response mapping → throttle/cache points |
| ALB request | 4 | listener rules → target selection → health-check state → draining |
| ECS task placement & deploy | 4 | scheduler → capacity → registration → health → draining old tasks |
| SQS consumption | 8 | replication, receive, visibility timer, delete, redrive |
| S3 upload (incl. multipart) | 10 | part boundaries, ETag assembly, durability replication, event emission |
| RDS connection + failover | 6 | pool → endpoint DNS → standby promotion → reconnect |
| DynamoDB request | 7 | key hashing → partition routing → replica quorum → consistency choice |
| CloudFront request | 10 | POP → cache key → origin shield/origin → cache fill → TTL |
7. Rendering and accuracy checks
Before any day is published:
- every Mermaid block is rendered and visually inspected — no truncated labels, no crossing edges that mislead;
- every arrow direction is checked against the actual protocol direction (a common error: drawing the consumer "receiving" as a push from SQS, when SQS is pull-based — the diagram must show the poll);
- every diagram has a one-line caption stating what to notice;
- diagrams containing numbers (costs, latencies, limits) have those numbers verified against current AWS documentation, with the retrieval date noted.