Interview Question Plan by Experience Level
Four question banks, ~285 questions total, plus the Top 100 with answers. Questions test reasoning, not recall — no multiple choice, no "which of these is a valid S3 storage class."
1. How questions are distributed
| Bank | Count | Character | Where they live |
|---|---|---|---|
| 2–3 YOE | 80 | What is it · how do I use it · basic implementation | End of each day, "Interview Questions" |
| 4–5 YOE | 75 | Integration · performance · service selection · failure handling | End of each day, "Interview Questions" |
| 6–7 YOE | 70 | Architecture · scalability · reliability · distributed systems | End of each day, "Senior-Level Questions" |
| 8–10 YOE | 60 | Trade-offs · failure domains · cost · security · migration · debugging | End of each day, "Senior-Level Questions" + Day 14 |
| Top 100 | 100 (drawn from the above, with model answers) | Rapid revision | Cheatsheet §10 |
| Traps | 9 | "Sounds right, is wrong" | Day 14 §14.3 |
Per teaching day: roughly 6 questions at 2–5 YOE and 5 at 6–10 YOE, always drawn from that day's material.
2. What each level is actually testing
flowchart LR
A["2-3 YOE<br/>Do you know<br/>what it is?"] --> B["4-5 YOE<br/>Can you use it<br/>correctly?"]
B --> C["6-7 YOE<br/>Can you design<br/>with it?"]
C --> D["8-10 YOE<br/>Can you defend<br/>the decision and<br/>the failure modes?"]The same topic, all four levels — SQS:
| Level | Question | What a good answer contains |
|---|---|---|
| 2–3 | What is SQS and when would you use it? | Durable queue; decoupling, buffering, retry; producer/consumer |
| 4–5 | Your consumer takes 90 seconds. Visibility timeout is 30. What happens and how do you fix it? | Message reappears, is processed repeatedly, eventually hits the DLQ; raise the timeout or heartbeat; mentions maxReceiveCount |
| 6–7 | Design order processing for 50k orders/min using SQS. How do payment, inventory and notification relate? | Fan-out over independent queues, independent scaling and failure isolation, the durability boundary, the API returning 202 |
| 8–10 | Your consumer is idempotent but the idempotency store is a different database from the business data. What can go wrong, and what would you change? | Non-atomic dedupe-then-write; crash window between the two; move the dedupe key into the same transaction/conditional write; or accept and reconcile; names the trade-off |
3. Coverage matrix — every day feeds every bank
| Day | 2–3 YOE | 4–5 YOE | 6–7 YOE | 8–10 YOE |
|---|---|---|---|---|
| 1 Cloud/IAM | 6 | 5 | 5 | 4 |
| 2 Networking | 6 | 5 | 5 | 4 |
| 3 Compute/Storage/CW | 6 | 5 | 5 | 4 |
| 4 ECS/ALB/Scaling | 6 | 6 | 5 | 4 |
| 5 Lambda/API GW | 6 | 6 | 5 | 5 |
| 6 RDS/Aurora | 6 | 6 | 5 | 5 |
| 7 DynamoDB/Cache | 6 | 6 | 5 | 5 |
| 8 SQS/SNS/Idempotency | 6 | 6 | 6 | 5 |
| 9 Events/Streams | 5 | 5 | 5 | 4 |
| 10 S3/CloudFront | 6 | 5 | 5 | 4 |
| 11 Security/IAM depth | 6 | 5 | 5 | 4 |
| 12 Scale/Reliability/Cost | 5 | 5 | 5 | 4 |
| 13 Observability/Migration | 5 | 5 | 5 | 4 |
| 14 System design | 5 | 5 | 4 | 4 |
| Total | 80 | 75 | 70 | 60 |
4. Sample questions at each level
2–3 YOE (10 of 80)
- What is the difference between a Region and an Availability Zone?
- What is an IAM role, and how is it different from an IAM user?
- What is S3 and what is it not good at?
- What is a security group?
- What does a load balancer do for a backend service?
- What is a Lambda cold start?
- What is the difference between SQS and SNS?
- What is a partition key in DynamoDB?
- How does your application get AWS credentials when running on EC2?
- What is CloudWatch used for?
4–5 YOE (10 of 75)
- When would you use S3 instead of EBS — and when is EBS the only right answer?
- Your ALB health check passes but users get 502s. Where do you look first?
- How do you size a connection pool for a service running 30 tasks against one database?
- Explain visibility timeout and what happens when you get it wrong in both directions.
- When is DynamoDB the wrong choice for a workload that "just needs key-value"?
- How do you keep a secret out of your application config and out of your image?
- Your Lambda gets throttled. Walk through what happens to a synchronous caller vs an asynchronous one.
- How would you cache a read-heavy endpoint, and what breaks on a cache miss storm?
- What is the difference between a read replica and a Multi-AZ standby, operationally?
- Why might a service that only talks to S3 still generate a large NAT Gateway bill?
6–7 YOE (10 of 70)
- Design a notification platform that sends email, SMS and push at 10M/day with per-channel failure isolation.
- You must guarantee an accepted order is never lost, but the payment provider is unreliable. Where is your durability boundary?
- How do you scale a service whose bottleneck is a single relational database? Give four options in the order you'd try them.
- Design idempotency for a payment API. What do you store, keyed on what, for how long?
- When would you choose Kinesis over SQS, and what do you lose?
- How do you do a zero-downtime schema change on a live service?
- Your event-driven system needs to replay three days of events for one consumer. How is that possible, and what did you have to build in advance?
- Design the caching strategy for a product catalogue with 5M items and heavy read skew.
- How do you structure IAM so 40 microservices in 3 accounts follow least privilege without blocking deploys?
- A service must serve users in Europe and the US with <100ms reads. Walk me through the options.
8–10 YOE (10 of 60)
- Your platform degrades and stays degraded after the triggering cause is fixed. Explain what happened and how the architecture caused it. (Metastable failure / retry storm.)
- You are asked to make the system multi-Region. What do you tell the executive sponsor about cost, complexity, and what it does not buy?
- A team proposes moving a 400 GB relational table to DynamoDB "for scale." Make the case both for and against, then decide.
- Where exactly is the consistency boundary in the order system you just designed, and what does a customer observe inside it?
- Your idempotency store and your business data live in different systems. Describe the failure window and three ways to close it.
- The AWS bill grew 3× while traffic grew 1.2×. Describe your investigation and the three most likely causes.
- You inherit a system with 200 Lambdas and no tracing. What do you do in the first two weeks, and what do you refuse to do?
- Design the failure domains for a platform that must survive an AZ loss, a bad deploy, and a poisoned message, without a human in the loop for the first two.
- When is "just use a bigger instance" the correct senior answer, and how do you defend it?
- What would you have to see to conclude that your event-driven architecture was a mistake and a synchronous call was better?
5. How the course teaches answering, not memorizing
Every 6–10 YOE question ships with a reasoning trace, not an answer paragraph:
What I'd clarify first the two or three questions that change the answer
The constraint that drives it what actually forces the decision
Options considered 3, with the real trade-off of each
What I'd choose and why a decision, stated plainly
What I'd give up explicitly
How I'd know I was wrong the metric or event that would change my mindThat last line is the one the course emphasizes most: a senior answer includes its own falsification condition.
The five-move senior pattern (taught on Day 14)
flowchart LR
A[State assumptions<br/>out loud] --> B[Name the constraint<br/>that dominates]
B --> C[Give 2-3 options,<br/>not a single answer]
C --> D[Choose, and say<br/>what you traded]
D --> E[Stress-test your<br/>own choice first]6. Common interview traps (Day 14)
Nine statements that sound correct and are not. Each gets the wrong mental model, the correct one, and the one-sentence version to say in an interview.
| Trap | Corrected in one line |
|---|---|
| "SQS guarantees exactly-once processing" | FIFO gives exactly-once delivery within a dedup window; exactly-once processing is something your consumer achieves through idempotency |
| "Lambda is always cheaper" | Cheaper at low/spiky volume; there is a crossover (roughly steady-state utilization above ~40–50%) where Fargate/EC2 wins, and the course does the arithmetic |
| "DynamoDB is always faster" | Predictable at any scale — which is different from faster; a well-indexed Postgres beats it on complex access patterns, and DynamoDB punishes unknown access patterns brutally |
| "SQS guarantees ordering" | Standard: best-effort, expect reordering. FIFO: ordered per message group, which is also your parallelism unit |
| "Kafka and SQS are interchangeable" | Queue (consume-and-delete, competing consumers) vs log (retained, replayable, multi-consumer, ordered per partition). Different primitives |
| "Multi-AZ means disaster recovery" | Multi-AZ is high availability within a Region. DR is about Region loss and logical failures — a bad deploy replicates to your standby instantly |
| "Private subnet means completely inaccessible" | It means no route from an IGW. Reachable via VPN, peering, PrivateLink, a misconfigured ALB, or a compromised bastion — and it can still reach the Internet outbound |
| "Store IAM user credentials in application config" | Use roles. Temporary, auto-rotated, no secret to leak. If you have an access key in a config file, that is the finding |
| "RDS solves database scaling" | It solves operations. Scaling writes is still your problem — caching, partitioning, sharding, or a different data model |
| (bonus) "More replicas always improve performance" | Replicas add read capacity and replication lag. Write throughput is unchanged; replica lag adds correctness hazards |
7. Mock interview formats (Day 14)
| Format | Duration | Use |
|---|---|---|
| Rapid fire | 15 min, 20 questions | Fundamentals recall at 2–5 YOE level |
| Deep dive | 20 min, one service | "Tell me everything about SQS" — tests depth ladder |
| System design | 45 min | One of the seven Day 14 systems, run end to end |
| Debugging | 20 min | One of the ten troubleshooting scenarios, interviewer plays the metrics |
| Trade-off defence | 15 min | Interviewer argues the opposite choice; candidate holds or updates with reasons |