Learning/AWS Backend Developer/Day 7 — DynamoDB and Caching

Day 7 — DynamoDB and Caching

3h 30m · Phase 2 of 4 · Curriculum

Learning Objectives

By the end of today you can:

  1. Explain DynamoDB's data model from first principles — hashing, partitions, and ordered ranges — rather than as a list of features.
  2. Design a table from access patterns, not from entities, and justify a GSI over an LSI (or neither).
  3. Compute RCU/WCU and on-demand cost in your head, and recognize when a query is a scan wearing a disguise.
  4. Diagnose a hot partition and fix it at the key-design level.
  5. Use a conditional write to make an operation safe under concurrency — the primitive Day 8's idempotency depends on.
  6. Implement cache-aside correctly, including TTL jitter, and name the three cache failure modes.
  7. Choose between RDS, DynamoDB and Redis for a given access pattern, and say when the answer is "two of them."

The sentence you should be able to say tonight: "DynamoDB gives me predictable latency at any scale in exchange for knowing my access patterns in advance — and that trade is the whole decision."

Prerequisites

  • Day 3 — CloudWatch metrics and alarms.
  • Day 5 — Lambda, and the conditional-write example in the idempotent handler.
  • Day 6 — essential. DynamoDB is taught today as a deliberate contrast with the relational model, and the contrast is the teaching device.
  • From the baseline: indexes, composite index column order, EXPLAIN.

Why Does This Exist?

Day 6 ended at a wall. A relational database has one writer. You can make it bigger, add replicas for reads, cache in front of it, and eventually you must partition your data across multiple databases and manage that yourself — at which point joins, transactions and foreign keys stop working across the boundary, and you have built a distributed database by hand, badly.

There is also a second, quieter wall: the performance of a relational query depends on the size of the data. A query that takes 3 ms on a million rows takes 200 ms on a billion, unless the index and the plan stay perfect forever. Your p99 degrades as you succeed.

DynamoDB is built around a different bargain:

Tell me how you will access the data, and I will give you single-digit-millisecond latency at any size, forever.

To deliver that it removes the things that make performance unpredictable:

RemovedWhy
JoinsA join's cost depends on data you cannot bound
Ad-hoc queriesAn arbitrary WHERE may require reading everything
A query plannerA planner is a system that sometimes chooses wrong
Unbounded result setsEvery read is bounded in size

That is not a limitation list; it is the mechanism. The absence of joins is why the latency is predictable. If you resent the constraints, you want a relational database — and that is a perfectly good answer.

Beginner Explanation

The one idea: hash, then range

flowchart TB
    K["Item key:<br/>customerId = c-51<br/>orderId = o-8812"] --> H["hash(c-51)"]
    H --> P{"Which partition?"}
    P --> P1["Partition 1"]
    P --> P2["Partition 2<br/>all items for c-51 live here<br/>o-1001<br/>o-4520<br/>o-8812 (sorted by sort key)<br/>o-9903"]
    P --> P3["Partition 3"]
    P2 --> R["Read one item: O(1)<br/>Read a RANGE of one customer's<br/>orders: O(result size)"]

That is genuinely the whole model:

  • The partition key is hashed. The hash decides which physical partition holds the item. This makes a single-item lookup a direct address computation, not a search — and it costs the same whether the table holds a thousand items or a trillion.
  • The sort key orders items within a partition. This makes "the 20 most recent orders for this customer" a contiguous read, stopping as soon as it has 20.

Compare with Day 6's relational picture:

RELATIONAL   index on (customer_id, created_at)
             - traverse a B-tree, seek to the customer, read in order
             - cost grows (slowly) with total data size
             - and you may add a different index tomorrow for a new query

DYNAMODB     partition key = customer_id, sort key = created_at
             - hash straight to the partition, read in order
             - cost is independent of total data size
             - but you CANNOT add a different access pattern tomorrow
               without designing for it (a GSI, or a rewrite)

A DynamoDB table is a composite index you cannot avoid using. If that sentence lands, the rest of today is detail.

Why there are no joins

Two items in a relational join may live anywhere. Joining them means finding both, which means a search whose cost depends on the data. DynamoDB's promise is incompatible with that.

So instead of joining at read time, you arrange the data so that things read together are stored together. That is the whole of "single-table design", and it is a genuine shift in thinking: in relational modelling you normalize and join; here you decide your queries first and store the answers adjacently.

Core Concepts

1. The building blocks

ConceptDetail
TableA collection of items. No schema beyond the key attributes.
ItemA row. Max 400 KB, including attribute names.
AttributeA field. Types: S, N, B, BOOL, NULL, L (list), M (map), SS/NS/BS (sets).
Partition key (PK / hash key)Required. Hashed to select a partition. Max 2,048 bytes.
Sort key (SK / range key)Optional. Orders items within a partition. Max 1,024 bytes.
Primary keyPK alone, or PK + SK. Must be unique.
Item collectionAll items sharing a PK. Relevant to the LSI size limit.

Partition limits — the source of the hot-partition problem

Limit per physical partitionValue
Storage10 GB
Read throughput3,000 RCU
Write throughput1,000 WCU

DynamoDB splits partitions automatically as data or throughput grows. But throughput is shared within a partition, so all the traffic for one partition key lands on one partition and is subject to those ceilings. A single, extremely popular key cannot exceed them no matter how much capacity you provision. That is scenario S14, and it is the single most common DynamoDB production problem.

2. Capacity, and arithmetic you should be able to do in your head

UnitBuys you
1 RCUone strongly consistent read of up to 4 KB/second — or two eventually consistent reads of up to 4 KB/s
1 WCUone write of up to 1 KB/second
Transactional read2 RCU per 4 KB
Transactional write2 WCU per 1 KB

Worked example — 500 reads/second of 6 KB items:

6 KB rounds up to 2 x 4 KB units
strongly consistent:   500 x 2      = 1,000 RCU
eventually consistent: 500 x 2 / 2  =   500 RCU   <- half the cost, for a read
                                                     that may be a few ms stale

And 200 writes/second of 3 KB items:

3 KB rounds up to 3 x 1 KB units
200 x 3 = 600 WCU

Eventually consistent reads are half price. That is not a rounding difference — it halves the read cost of your entire application. Use them by default and choose strong consistency deliberately where correctness needs it.

On-demand vs provisioned

On-demandProvisioned (+ autoscaling)
BillingPer requestPer capacity-unit-hour, reserved
Price (illustrative)~$1.25 per million write request units, ~$0.25 per million read request units~$0.00013 per WCU-hour, ~$0.000065 per RCU-hour
ScalingInstant, no configurationAutoscaling reacts in minutes
ThrottlingRare, but there are per-table ramp limitsWhen you exceed provisioned capacity
Best forUnknown, spiky, or new workloads; low duty cyclePredictable, sustained load

The break-even is roughly 15–20% sustained utilization of the equivalent provisioned capacity. Below that, on-demand wins. Above it, provisioned with autoscaling is substantially cheaper.

1,000 WCU steady for a month:
  provisioned:  1,000 x $0.00013 x 730 h            = about  $95
  on-demand:    1,000 x 3600 x 730 / 1M x $1.25     = about $3,285   (34x more)

Start on-demand — it removes a whole class of early mistakes. Move to provisioned once you have a month of real usage data. The switch is a single API call (limited to once every 24 hours).

3. Modelling: access patterns first

This is the part that cannot be shortcut. Write the queries before the schema. Every time.

Worked example — an order system

Step 1: enumerate access patterns. Not entities. Queries.

#Access patternFrequency
1Get one order by idVery high
2List a customer's orders, newest first, paginatedHigh
3List a customer's orders in a given statusMedium
4List all orders in status PENDING_PAYMENT (for a worker)Medium
5List all orders created on a given day (for ops)Low
6Get the line items of an orderHigh, always with #1

Step 2: design keys so that each pattern is a single Query or GetItem.

TABLE: orders          PK = pk (String)        SK = sk (String)

Pattern 1 and 6 - the order and its items are ONE contiguous read:
  pk = ORDER#o-8812    sk = META                 -> the order itself
  pk = ORDER#o-8812    sk = ITEM#001             |
  pk = ORDER#o-8812    sk = ITEM#002             |  Query pk=ORDER#o-8812
  pk = ORDER#o-8812    sk = ITEM#003             |  returns all of it, once

Pattern 2 - GSI1, newest first:
  GSI1PK = CUST#c-51   GSI1SK = 2026-01-14T09:12:04Z#o-8812
  -> Query GSI1 with pk=CUST#c-51, ScanIndexForward=false, Limit=20

Pattern 3 - GSI2, status within a customer:
  GSI2PK = CUST#c-51   GSI2SK = PAID#2026-01-14T09:12:04Z
  -> Query GSI2 with pk=CUST#c-51 AND begins_with(sk, "PAID#")

Pattern 4 - GSI3, SPARSE: the attribute exists ONLY while pending.
  GSI3PK = PENDING     GSI3SK = 2026-01-14T09:12:04Z
  -> when the order is paid, DELETE the GSI3PK attribute; the item leaves
     the index entirely. The index stays small forever.
     But see the hot-partition discussion: a single PK value is one partition.

Pattern 5 - GSI4, with a write-sharded key to avoid a daily hot partition:
  GSI4PK = DATE#2026-01-14#<0-9>    GSI4SK = <timestamp>#o-8812
  -> query all 10 shards in parallel and merge. Low frequency, so the
     extra complexity buys throughput where a single key would throttle.

Three techniques in that design worth naming:

  1. Prefixed keys (ORDER#, CUST#) let one table hold several entity types without collision — the basis of single-table design.
  2. Sparse indexes (pattern 4): an item appears in a GSI only if it has that GSI's key attributes. A "find the work queue" index that automatically shrinks to just the outstanding work is extremely efficient.
  3. Write sharding (pattern 5): append a suffix to spread a naturally hot key across N partitions, then fan out reads. The cost is read complexity; the benefit is throughput you could not otherwise have.

GSI vs LSI

Global Secondary IndexLocal Secondary Index
Partition keyCan differ from the table'sMust be the table's PK
Sort keyAny attributeAny attribute
ConsistencyEventually consistent onlyStrong consistency available
CapacityIts own, separate from the tableShares the table's
CreatedAny timeOnly with the table — never afterwards
DeletedAny timeNever
Limit20 per table (default, raisable)5 per table
Item collection sizeNo constraint10 GB limit per partition key across the table and all its LSIs

Choose a GSI almost always. LSIs impose that 10 GB item-collection limit — once a single partition key's items plus their LSI entries exceed 10 GB, writes to that key fail permanently, and you cannot drop the LSI to fix it. The only reasons to accept an LSI are a genuine need for strongly consistent reads on an alternate sort key, and a bounded item collection.

The GSI trap that takes down tables: a GSI has its own capacity, and if a GSI is throttled, DynamoDB throttles writes to the base table. It cannot let the index fall arbitrarily behind. So an under-provisioned GSI — or a GSI with a hot partition key — becomes an outage of the whole table's write path. Provision GSIs like you mean it, and watch their throttle metrics separately.

Single-table design, honestly

The canonical DynamoDB advice is "one table per application, not per entity." It is correct, for real reasons — one set of capacity to manage, related entities retrieved in one round trip, fewer moving parts — and it is oversold.

Single table whenMultiple tables when
Entities are queried togetherEntities have completely unrelated access patterns
Access patterns are well understood and stableAccess patterns are still being discovered
The team understands the modelThe team is new to DynamoDB and will get it wrong
One-round-trip retrieval of related items mattersDifferent entities need very different capacity, TTL or backup policies

The pragmatic position to take in an interview: "I'd use single-table design where entities are genuinely accessed together, and separate tables where they aren't. The overlapping-keys style is powerful and it costs real legibility — a team that can't read the model will produce a worse design than a slightly less optimal multi-table one."

Query vs Scan

QueryScan
ReadsOne partition, a range of sort keysEvery item in the table
CostProportional to results returnedProportional to table size
LatencyMillisecondsSeconds to hours
FiltersFilterExpression applies after the read — you pay for the filtered-out itemsSame

FilterExpression is not a WHERE clause. It is applied after items are read and after capacity is consumed. Scan with a filter returning 10 items from a 50 GB table costs the full 50 GB of reads. This is scenario S16, and it is how a "simple nightly job" costs $900 a month.

If you need a scan in a request path, your key design is wrong. Legitimate uses: one-off migrations, analytics exports (better: export to S3 and use Athena), and small configuration tables.

4. Conditional writes — the most important API on this page

A conditional write succeeds only if a condition on the item holds, evaluated atomically at the item level.

Java
// Create only if it does not exist. The atomic primitive behind idempotency.
.conditionExpression("attribute_not_exists(pk)")

// Optimistic locking - the DynamoDB equivalent of a version column.
.conditionExpression("#v = :expectedVersion")

// Business invariant enforced in the database, not in application logic.
.conditionExpression("#status = :pending AND #stock >= :qty")

If the condition fails you get ConditionalCheckFailedException. That is frequently not an error — it is the answer. "This message was already processed." "Someone else updated it first." "There is not enough stock."

This matters enormously because DynamoDB has no transactions across requests and no row locks. Conditional writes are how you get correctness under concurrency, and they are the mechanism for:

PatternCondition
Idempotency (Day 8)attribute_not_exists(pk)
Optimistic lockingversion = :expected
A distributed lock with an owner and expiryattribute_not_exists(lockedBy) OR expiresAt < :now
Enforcing an invariantstock >= :qty
State-machine transitionsstatus = :fromState

A conditional write that fails still consumes write capacity. A workload built on frequently-failing conditions pays for the failures.

Transactions

TransactWriteItems gives ACID across up to 100 items (and 4 MB) in one Region, all-or-nothing.

Cost2x the capacity of the same non-transactional writes
ConstraintNo two operations on the same item in one transaction
FailureTransactionCanceledException, with a per-item reason list — read the reasons, they tell you which condition failed

Use it where you genuinely need atomicity across items (debit one account, credit another). Do not use it as a reflex — most DynamoDB writes are single-item and therefore already atomic.

5. Consistency, and the rest of the operational surface

Read typeBehaviourCost
Eventually consistent (default)May not reflect a write from the last few milliseconds0.5 RCU per 4 KB
Strongly consistentReflects all prior successful writes1 RCU per 4 KB
TransactionalSerializable, cross-item2 RCU per 4 KB

GSIs are always eventually consistent. There is no strongly-consistent GSI read, ever. If a workflow writes an item and immediately queries a GSI for it, it will sometimes not be there. This is exactly Day 6's read-replica staleness problem in a different costume — and worth noticing that the shape of the problem generalizes across every distributed data store.

FeatureWhat to know
Batch operationsBatchGetItem up to 100 items/16 MB; BatchWriteItem up to 25 items/16 MB. Both can partially fail — you must retry UnprocessedItems with backoff. Not atomic.
TTLSet an epoch-seconds attribute; DynamoDB deletes the item, typically within 48 hours of expiry. Free, consumes no WCU. Ideal for idempotency keys and sessions.
StreamsAn ordered, 24-hour change log per item. KEYS_ONLY / NEW_IMAGE / OLD_IMAGE / NEW_AND_OLD_IMAGES. The basis of CDC (Day 9).
PITRContinuous backups, restore to any second in the last 35 days. Restores to a new table.
Adaptive capacityAutomatically isolates and boosts frequently accessed partitions. It mitigates uneven access; it does not raise the per-partition ceiling, so a single hot key still throttles.
DAXAn in-VPC, write-through cache in front of DynamoDB. Microsecond reads. Worth it only for read-heavy, repeated-key workloads.

6. Caching with ElastiCache

Why cache at all, when DynamoDB is already fast

ReasonDetail
CostA Redis GET is effectively free against a per-request DynamoDB read. At high read volume caching is a cost decision before it is a latency one.
LatencySub-millisecond vs single-digit milliseconds
Protecting a slower storeCaching in front of Aurora is where caching earns most of its keep
Computed resultsA cached aggregate that would cost 50 reads to recompute
Cross-cutting stateRate limiters, sessions, leaderboards, locks

Redis vs Memcached vs Valkey

Redis / ValkeyMemcached
Data structuresStrings, hashes, lists, sets, sorted sets, streams, HLLStrings only
Persistence / snapshotsYesNo
Replication + failoverYes, Multi-AZNo
Transactions, Lua, pub/subYesNo
Multi-threadedMostly single-threaded for command executionYes

Use Redis (or Valkey, which is API-compatible and cheaper) unless you have a specific reason not to. Memcached's multi-threading matters only for very high-throughput pure-cache workloads.

Because Redis executes commands on a single thread, one slow command blocks everything. KEYS * on a production cache is a self-inflicted outage. Use SCAN.

The four strategies

flowchart TB
    subgraph CA["Cache-aside (lazy loading) - the default"]
        A1[Read] --> A2{"in cache?"}
        A2 -->|hit| A3[return]
        A2 -->|miss| A4[read DB] --> A5[write cache] --> A3
        A6[Write] --> A7[write DB] --> A8["invalidate cache"]
    end
    subgraph WT["Write-through"]
        B1[Write] --> B2[write cache] --> B3[write DB]
        B4[Read] --> B5["always a hit for written data"]
    end
    subgraph WB["Write-behind"]
        C1[Write] --> C2[write cache] --> C3["return immediately"]
        C2 -.->|async, batched| C4[write DB]
    end
StrategyGainsCosts
Cache-asideOnly caches what is read; resilient to cache failureEvery miss pays DB latency; stale until invalidated or expired
Write-throughCache always current for written dataWrite latency includes both; caches data nobody may read
Write-behindFastest writes; absorbs burstsData loss if the cache dies before flushing. Rarely right for a system of record.
Read-throughClean application codeRequires cache-provider support

Default to cache-aside. It is simple, it degrades gracefully, and the failure mode (a miss) is harmless.

The three failure modes, by name

1. Cache stampede (dogpile) — S15. A popular key expires. A thousand concurrent requests all miss and all hit the database at once.

flowchart LR
    K["Hot key TTL expires"] --> M["1,000 concurrent misses"]
    M --> DB[("Database gets 1,000x<br/>its normal load, instantly")]
    DB --> T["Timeouts, then retries, then worse"]

Three mitigations, all cheap:

MitigationMechanism
TTL jitterttl = base +/- random(0, base * 0.2). Keys stop expiring simultaneously. Do this always — it is one line of code.
Request coalescingPer-key in-process lock so one thread refills and the rest wait for that result
Stale-while-revalidateServe the expired value and refresh in the background. Best UX, needs two TTLs.

2. Hot key. One key gets a disproportionate share of traffic, saturating the single Redis shard that owns it. Mitigations: a local in-process cache in front of Redis (short TTL), or replicating the value under N suffixed keys and reading a random one.

3. Thundering herd on restart. A cache node restarts empty. Every request misses, and the database — which has been comfortably serving 10% of read traffic — receives 100% of it, instantly. Mitigations: warm the cache before taking traffic; a replica so restarts do not empty everything; and always have a circuit breaker so a database overload does not become total failure.

Redis beyond caching

A rate limiter — correct because Redis commands are atomic:

Java
// INCR + EXPIRE on a time-bucketed key. Atomic, so no race between
// checking and incrementing.
String key = "rate:" + userId + ":" + (System.currentTimeMillis() / 60_000);
Long count = redis.opsForValue().increment(key);
if (count == 1L) redis.expire(key, Duration.ofMinutes(2));
if (count > LIMIT) throw new TooManyRequestsException();

A distributed lock — with the caveats stated, because this is where people go wrong:

Java
// SET key value NX PX ttl  - atomic acquire with an expiry
Boolean acquired = redis.opsForValue()
        .setIfAbsent(lockKey, ownerToken, Duration.ofSeconds(30));

A Redis lock is not a correctness guarantee. If your process pauses (GC, a slow syscall) past the TTL, the lock expires, someone else takes it, and now two holders believe they own it. Redis-based locking is fine for efficiency ("don't do this work twice, usually") and not for correctness ("this must never happen twice"). For correctness, use a fencing token — a monotonically increasing number checked by the resource itself — or, in AWS, a DynamoDB conditional write, which is atomic and durable. This is the honest version of a much-debated topic, and knowing it is a senior signal.

Architecture

What today adds

flowchart TB
    C[Client] --> ALB[ALB] --> APP[Order Service · ECS]
    subgraph DATA["Data tier"]
        RED[("ElastiCache Redis<br/>cluster mode enabled<br/>catalogue cache")]
        DDB[("DynamoDB<br/>orders + GSIs")]
        IDEM[("DynamoDB<br/>idempotency keys · TTL 24h")]
        AUR[("Aurora<br/>system of record")]
    end
    APP -->|"1 cache-aside read"| RED
    RED -.->|"miss"| AUR
    APP -->|"high-volume order reads/writes"| DDB
    APP -->|"conditional write before side effects"| IDEM
    DDB -->|"Streams to Day 9"| STR[(DynamoDB Streams)]

The read path, in detail

sequenceDiagram
    autonumber
    participant APP as Order Service
    participant R as Redis
    participant D as DynamoDB
    APP->>R: GET catalogue:sku-991
    alt hit (about 0.3 ms)
        R-->>APP: value
    else miss
        R-->>APP: nil
        APP->>APP: acquire per-key in-process lock<br/>(request coalescing, stops a stampede)
        APP->>D: GetItem pk=SKU#991 (eventually consistent, 0.5 RCU)
        D-->>APP: item (about 4 ms)
        APP->>R: SETEX catalogue:sku-991 with jittered TTL
        APP->>APP: release lock; waiters read the cached value
    end

Failure flow

flowchart TB
    APP[Application] --> R{"Redis reachable?"}
    R -->|"no, node failure"| FB["FALL BACK to DynamoDB.<br/>A cache outage must degrade<br/>latency and cost, NEVER<br/>availability."]
    FB --> HERD["but now 100% of reads hit the DB:<br/>thundering herd"]
    HERD --> CB["Circuit breaker + load shedding<br/>keep the DB alive (Day 12)"]
    R -->|yes| HIT[serve]
    APP --> D{"DynamoDB throttles?"}
    D -->|"ProvisionedThroughputExceeded"| RETRY["SDK retries with backoff<br/>(automatic in v2)"]
    RETRY -->|"persistent"| HOT["Hot partition or<br/>under-provisioned GSI:<br/>a KEY DESIGN problem,<br/>not a capacity problem"]
    D -->|"GSI throttled"| BASE["BASE TABLE WRITES<br/>are throttled too"]

How It Works Internally

A DynamoDB request

sequenceDiagram
    autonumber
    participant APP as Application
    participant EP as dynamodb.eu-west-1.amazonaws.com
    participant RT as Request router
    participant PART as Storage partition (3 replicas across AZs)
    APP->>EP: GetItem pk=ORDER#o-8812 (SigV4 signed)
    EP->>RT: authorize (IAM), then route
    RT->>RT: hash(pk) gives partition id and replica set
    alt Eventually consistent read (default)
        RT->>PART: read from ANY of the 3 replicas
        Note over PART: may miss a write from the<br/>last few milliseconds. 0.5 RCU.
    else Strongly consistent read
        RT->>PART: read from the LEADER replica
        Note over PART: reflects all prior writes. 1 RCU.
    end
    PART-->>APP: item + ConsumedCapacity

For a write:

sequenceDiagram
    autonumber
    participant APP
    participant RT as Request router
    participant L as Leader replica
    participant F as Follower replicas (2, other AZs)
    APP->>RT: PutItem with ConditionExpression
    RT->>L: route to the leader for this partition
    L->>L: evaluate the condition ATOMICALLY on the item
    alt condition fails
        L-->>APP: ConditionalCheckFailedException<br/>(capacity IS still consumed)
    else condition holds
        L->>F: replicate
        F-->>L: ack (durability quorum)
        L-->>APP: 200 + ConsumedCapacity
        L-->>APP: Streams record written
        L-->>APP: GSI update is asynchronous and<br/>consumes the GSI's own capacity
    end

Three things to take from this pair of diagrams:

  • Condition evaluation happens at the leader for that item. That is why it is atomic, and why it is a reliable concurrency primitive without locks.
  • GSI updates are asynchronous — hence eventual consistency on every GSI, always.
  • ConsumedCapacity is returned on every request. Ask for it (ReturnConsumedCapacity=TOTAL) and log it. It is free, precise cost telemetry, and almost nobody uses it.

How a hot partition happens

flowchart TB
    subgraph BAD["BAD: pk = status"]
        B1["hash(PENDING)"] --> BP1["Partition A<br/>95% of all traffic<br/>ceiling: 1,000 WCU"]
        B2["hash(SHIPPED)"] --> BP2["Partition B<br/>nearly idle"]
        BP1 --> THR["THROTTLED at about 1,000 WCU<br/>even with 20,000 provisioned:<br/>the other 19,000 are unreachable"]
    end
    subgraph GOOD["GOOD: pk = orderId"]
        G1["hash(o-8812)"] --> GP1[Partition 1]
        G2["hash(o-8813)"] --> GP2[Partition 2]
        G3["hash(o-9001)"] --> GP3[Partition N]
        GP1 & GP2 & GP3 --> OK["Even distribution.<br/>Scales to the table's full capacity."]
    end

The rule: a partition key must have high cardinality and even access. Low-cardinality attributes — status, country, tenant tier, true/false, a date — are never partition keys on a high-throughput table. If you need to query by one, use a sparse GSI (small, so throughput is manageable) or write sharding (PENDING#<0-9>), and fan out reads.

Code

Dependencies

<dependency>
  <groupId>software.amazon.awssdk</groupId>
  <artifactId>dynamodb-enhanced</artifactId>   <!-- BOM-managed -->
</dependency>
<dependency>
  <groupId>org.springframework.boot</groupId>
  <artifactId>spring-boot-starter-data-redis</artifactId>
</dependency>

An entity for the Enhanced Client

Java
package com.acme.orders;

import software.amazon.awssdk.enhanced.dynamodb.mapper.annotations.*;

import java.math.BigDecimal;
import java.time.Instant;

@DynamoDbBean
public class OrderItem {

    private String pk;          // ORDER#o-8812
    private String sk;          // META, or ITEM#001
    private String orderId;
    private String customerId;
    private String status;
    private BigDecimal amount;
    private Instant createdAt;
    private Long version;

    // GSI1: a customer's orders, newest first
    private String gsi1pk;      // CUST#c-51
    private String gsi1sk;      // 2026-01-14T09:12:04Z#o-8812

    // GSI3: SPARSE work queue. These attributes exist ONLY while pending;
    // clearing them removes the item from the index entirely.
    private String gsi3pk;      // PENDING#<shard>, or null
    private String gsi3sk;

    @DynamoDbPartitionKey                       public String getPk() { return pk; }
    @DynamoDbSortKey                            public String getSk() { return sk; }

    @DynamoDbSecondaryPartitionKey(indexNames = "gsi1")
    public String getGsi1pk() { return gsi1pk; }
    @DynamoDbSecondarySortKey(indexNames = "gsi1")
    public String getGsi1sk() { return gsi1sk; }

    @DynamoDbSecondaryPartitionKey(indexNames = "gsi3")
    public String getGsi3pk() { return gsi3pk; }
    @DynamoDbSecondarySortKey(indexNames = "gsi3")
    public String getGsi3sk() { return gsi3sk; }

    // Optimistic locking: the Enhanced Client turns this into a
    // ConditionExpression on every update. No locks, no lost updates.
    @DynamoDbVersionAttribute                   public Long getVersion() { return version; }

    // remaining getters/setters omitted for brevity
}

The repository — query, conditional write, and pagination done right

Java
package com.acme.orders;

import org.springframework.stereotype.Repository;
import software.amazon.awssdk.enhanced.dynamodb.*;
import software.amazon.awssdk.enhanced.dynamodb.model.*;
import software.amazon.awssdk.services.dynamodb.model.*;

import java.util.List;

@Repository
public class OrderRepository {

    private final DynamoDbTable<OrderItem> table;
    private final DynamoDbIndex<OrderItem> gsi1;
    private final software.amazon.awssdk.services.dynamodb.DynamoDbClient raw;

    OrderRepository(DynamoDbEnhancedClient enhanced,
                    software.amazon.awssdk.services.dynamodb.DynamoDbClient raw) {
        this.table = enhanced.table("orders", TableSchema.fromBean(OrderItem.class));
        this.gsi1 = table.index("gsi1");
        this.raw = raw;
    }

    /** Pattern 1 + 6: the order AND its line items, in ONE round trip. */
    public List<OrderItem> findOrderWithItems(String orderId) {
        return table.query(r -> r
                        .queryConditional(QueryConditional
                                .keyEqualTo(k -> k.partitionValue("ORDER#" + orderId)))
                        // Eventually consistent by default = half the RCU.
                        // Explicit here because it is a decision, not an accident.
                        .consistentRead(false))
                .items().stream().toList();
    }

    /**
     * Pattern 2: a customer's orders, newest first, paginated.
     *
     * The .stream().limit(n) mistake: DynamoDB's Limit bounds ITEMS EXAMINED
     * per request, and a client-side limit still pays for everything fetched.
     * Take ONE page and hand the client the LastEvaluatedKey.
     */
    public Page<OrderItem> findCustomerOrders(String customerId, int pageSize,
                                              java.util.Map<String, AttributeValue> exclusiveStartKey) {
        var request = QueryEnhancedRequest.builder()
                .queryConditional(QueryConditional
                        .keyEqualTo(k -> k.partitionValue("CUST#" + customerId)))
                .scanIndexForward(false)        // newest first
                .limit(pageSize)
                .exclusiveStartKey(exclusiveStartKey)
                .build();
        return gsi1.query(request).iterator().next();   // exactly one page
    }

    /**
     * A state transition that is safe under concurrency, without a lock.
     *
     * The condition enforces the state machine IN THE DATABASE. Two concurrent
     * attempts to pay the same order: exactly one succeeds.
     */
    public void markPaid(String orderId) {
        try {
            raw.updateItem(UpdateItemRequest.builder()
                    .tableName("orders")
                    .key(java.util.Map.of(
                            "pk", AttributeValue.fromS("ORDER#" + orderId),
                            "sk", AttributeValue.fromS("META")))
                    .updateExpression(
                            "SET #st = :paid, paidAt = :now REMOVE gsi3pk, gsi3sk")
                    // REMOVE takes the item out of the sparse work-queue GSI
                    .conditionExpression("#st = :pending")
                    .expressionAttributeNames(java.util.Map.of("#st", "status"))
                    .expressionAttributeValues(java.util.Map.of(
                            ":paid", AttributeValue.fromS("PAID"),
                            ":pending", AttributeValue.fromS("PENDING_PAYMENT"),
                            ":now", AttributeValue.fromS(java.time.Instant.now().toString())))
                    .returnConsumedCapacity(ReturnConsumedCapacity.TOTAL)
                    .build());

        } catch (ConditionalCheckFailedException e) {
            // NOT an error. Either it was already paid (a duplicate, fine) or
            // it is in a state that cannot transition to PAID (a real problem).
            // The caller decides; this is a domain outcome.
            throw new IllegalOrderStateException(orderId, e);
        }
    }
}

Cache-aside with jitter and coalescing

Java
package com.acme.catalogue;

import org.springframework.data.redis.core.StringRedisTemplate;
import org.springframework.stereotype.Service;

import java.time.Duration;
import java.util.Optional;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.locks.ReentrantLock;

@Service
public class CatalogueService {

    private static final Duration BASE_TTL = Duration.ofMinutes(10);
    private static final double JITTER = 0.2;            // plus or minus 20%
    private static final String NEGATIVE_MARKER = "__ABSENT__";

    private final StringRedisTemplate redis;
    private final ProductRepository repo;
    // Per-key locks: request coalescing. On a miss, ONE thread refills and the
    // others wait for its result, instead of 1,000 threads all hitting the DB.
    private final ConcurrentHashMap<String, ReentrantLock> locks = new ConcurrentHashMap<>();

    CatalogueService(StringRedisTemplate redis, ProductRepository repo) {
        this.redis = redis; this.repo = repo;
    }

    public Optional<Product> get(String sku) {
        String key = "catalogue:" + sku;

        try {
            String cached = redis.opsForValue().get(key);
            if (NEGATIVE_MARKER.equals(cached)) return Optional.empty();
            if (cached != null) return Optional.of(deserialize(cached));
        } catch (Exception e) {
            // A cache outage must degrade latency and cost, NEVER availability.
            // Fall through to the database.
            log.warn("cache read failed, falling through", e);
            return repo.findBySku(sku);
        }

        ReentrantLock lock = locks.computeIfAbsent(key, k -> new ReentrantLock());
        lock.lock();
        try {
            // Double-check: another thread may have filled it while we waited.
            String cached = redis.opsForValue().get(key);
            if (NEGATIVE_MARKER.equals(cached)) return Optional.empty();
            if (cached != null) return Optional.of(deserialize(cached));

            Optional<Product> fromDb = repo.findBySku(sku);

            try {
                if (fromDb.isPresent()) {
                    redis.opsForValue().set(key, serialize(fromDb.get()), jitteredTtl());
                } else {
                    // Negative caching with a SHORT ttl: stops a missing-SKU
                    // hammer from hitting the database on every request.
                    redis.opsForValue().set(key, NEGATIVE_MARKER, Duration.ofSeconds(30));
                }
            } catch (Exception ignored) {
                // A cache write failure is not fatal.
            }
            return fromDb;

        } finally {
            lock.unlock();
            locks.remove(key, lock);
        }
    }

    /**
     * TTL jitter. Without it, 2 million keys written during a deploy all expire
     * in the same second, ten minutes later, and the database falls over.
     * This is three lines and it prevents S15.
     */
    private Duration jitteredTtl() {
        long base = BASE_TTL.toSeconds();
        long delta = (long) (base * JITTER * (Math.random() * 2 - 1));
        return Duration.ofSeconds(base + delta);
    }

    /** Write path: invalidate rather than update, so we never cache a stale write. */
    public void update(Product p) {
        repo.save(p);
        try { redis.delete("catalogue:" + p.sku()); } catch (Exception ignored) { }
    }
}

Invalidate, don't update, on write. Updating the cache from the write path introduces a race: two concurrent writers can leave the cache holding the older value. Deleting is idempotent and always safe — the next read repopulates from the source of truth.

Creating the table

aws dynamodb create-table \
  --table-name orders \
  --billing-mode PAY_PER_REQUEST \
  --attribute-definitions \
      AttributeName=pk,AttributeType=S AttributeName=sk,AttributeType=S \
      AttributeName=gsi1pk,AttributeType=S AttributeName=gsi1sk,AttributeType=S \
  --key-schema AttributeName=pk,KeyType=HASH AttributeName=sk,KeyType=RANGE \
  --global-secondary-indexes '[{
      "IndexName":"gsi1",
      "KeySchema":[{"AttributeName":"gsi1pk","KeyType":"HASH"},
                   {"AttributeName":"gsi1sk","KeyType":"RANGE"}],
      "Projection":{"ProjectionType":"INCLUDE",
                    "NonKeyAttributes":["orderId","status","amount","createdAt"]}
  }]' \
  --stream-specification StreamEnabled=true,StreamViewType=NEW_AND_OLD_IMAGES \
  --sse-specification Enabled=true

aws dynamodb update-continuous-backups --table-name orders \
  --point-in-time-recovery-specification PointInTimeRecoveryEnabled=true

# The idempotency table for Day 8. TTL means it cleans itself up for free.
aws dynamodb create-table --table-name idempotency \
  --billing-mode PAY_PER_REQUEST \
  --attribute-definitions AttributeName=pk,AttributeType=S \
  --key-schema AttributeName=pk,KeyType=HASH
aws dynamodb update-time-to-live --table-name idempotency \
  --time-to-live-specification "Enabled=true,AttributeName=expiresAt"

ProjectionType is a real cost decision. ALL duplicates every attribute into the index — doubling storage and the write cost of every item. KEYS_ONLY is cheapest but forces a second GetItem per result. INCLUDE with exactly the attributes your query needs is usually right, and is the choice people skip.

Hands-on

Lab 9 — DynamoDB with the Enhanced Client (20 min)

Objective: model the order access patterns, query a GSI, and prove a conditional write prevents a double-update.

Billable: negligible at this scale; DynamoDB's free tier covers it. PITR adds a small storage charge.

Steps:

  1. Create the orders table and gsi1 as above. Create the idempotency table with TTL.
  2. Implement OrderItem and OrderRepository.
  3. Seed 200 orders across 10 customers with varied statuses and timestamps.
  4. Implement findOrderWithItems, findCustomerOrders (paginated), and markPaid.

Verification:

# (a) ONE Query returns the order and all its items - the single-table payoff
aws dynamodb query --table-name orders \
  --key-condition-expression "pk = :pk" \
  --expression-attribute-values '{":pk":{"S":"ORDER#o-8812"}}' \
  --return-consumed-capacity TOTAL \
  --query '{items: Count, capacity: ConsumedCapacity.CapacityUnits}'
# e.g. {"items": 4, "capacity": 0.5}   <- four items, half an RCU
# (b) THE CONDITIONAL WRITE TEST - the point of the lab.
# Five concurrent attempts to pay the same order. Exactly one must win.
for i in 1 2 3 4 5; do
  curl -s -XPOST "http://localhost:8080/orders/o-8812/pay" -o "/tmp/pay-$i.json" \
       -w "%{http_code}\n" &
done
wait
sort /tmp/pay-*.json | uniq -c
# Expect exactly ONE 200 and four 409-equivalents
# (ConditionalCheckFailedException becomes IllegalOrderStateException)

aws dynamodb get-item --table-name orders \
  --key '{"pk":{"S":"ORDER#o-8812"},"sk":{"S":"META"}}' \
  --query 'Item.{status:status.S,paidAt:paidAt.S}'
# Exactly one paidAt timestamp. The state machine held, with no locks.
# (c) SCAN vs QUERY - the cost lesson, measured
aws dynamodb scan --table-name orders \
  --filter-expression "#s = :s" \
  --expression-attribute-names '{"#s":"status"}' \
  --expression-attribute-values '{":s":{"S":"PENDING_PAYMENT"}}' \
  --return-consumed-capacity TOTAL \
  --query '{scanned: ScannedCount, returned: Count, capacity: ConsumedCapacity.CapacityUnits}'
# {"scanned": 850, "returned": 12, "capacity": 106.5}
# YOU PAID FOR 850 ITEMS TO GET 12.

aws dynamodb query --table-name orders --index-name gsi3 \
  --key-condition-expression "gsi3pk = :p" \
  --expression-attribute-values '{":p":{"S":"PENDING"}}' \
  --return-consumed-capacity TOTAL \
  --query '{scanned: ScannedCount, returned: Count, capacity: ConsumedCapacity.CapacityUnits}'
# {"scanned": 12, "returned": 12, "capacity": 1.5}
# Same answer, about 70x cheaper. THIS is why key design is the whole job.
# (d) Watch the sparse index shrink as work completes
aws dynamodb query --table-name orders --index-name gsi3 \
  --key-condition-expression "gsi3pk = :p" \
  --expression-attribute-values '{":p":{"S":"PENDING"}}' --select COUNT
# Pay some orders, then re-run. The count drops, because items removed the
# GSI3 attributes and therefore left the index. The work queue is self-cleaning.

Common errors:

SymptomCauseFix
ValidationException: schema mismatchAttributes in the key schema not declared in --attribute-definitionsDeclare only key attributes, and all of them
ResourceNotFoundException on a GSI queryIndex still CREATING, or the name is wrongdescribe-table, wait for ACTIVE
GSI query returns nothing for an item you just wroteGSIs are eventually consistent, alwaysRetry, or read the base table
ProvisionedThroughputExceededException on-demandTable-level ramp limits, or a hot partitionCheck key distribution first
ConditionalCheckFailedException you did not expectThe item is not in the state you assumedThat is the API doing its job — handle it as a domain case
Item write fails with a size error400 KB limit, including attribute namesStore the blob in S3, keep a pointer

Lab 10 — ElastiCache Redis cache-aside (20 min)

Objective: put a cache in front of a data store, measure the hit rate and latency change, and then cause a stampede on purpose and fix it.

Billable: cache.t4g.micro about $0.016/hr. Under $0.10 for the lab. Delete the cluster.

Steps:

  1. Create a Redis (or Valkey) cluster in the private subnets, with cache-sg allowing 6379 from app-sg only.
aws elasticache create-cache-cluster \
  --cache-cluster-id course-cache \
  --engine redis --cache-node-type cache.t4g.micro --num-cache-nodes 1 \
  --cache-subnet-group-name course-private-subnets \
  --security-group-ids "$CACHE_SG" \
  --transit-encryption-enabled
aws elasticache wait cache-cluster-available --cache-cluster-id course-cache
  1. Wire CatalogueService with a 10-minute TTL and no jitter yet — you will need the bug.
  2. Seed 1,000 products.

Verification:

# (a) Cold pass: all misses. Warm pass: hits.
hey -n 2000 -c 20 "http://localhost:8080/catalogue/sku-991"   # first
hey -n 2000 -c 20 "http://localhost:8080/catalogue/sku-991"   # second

# Compare the p99 between the two runs. Expect roughly an order of magnitude.
aws cloudwatch get-metric-statistics --namespace AWS/ElastiCache \
  --metric-name CacheHits --dimensions Name=CacheClusterId,Value=course-cache \
  --start-time "$(date -u -v-15M +%Y-%m-%dT%H:%M:%SZ)" \
  --end-time "$(date -u +%Y-%m-%dT%H:%M:%SZ)" --period 60 --statistics Sum
# (b) CAUSE A STAMPEDE. Warm 1,000 keys in one burst, then expire them all at
# the same instant - exactly what a deploy-time warm does with a fixed TTL.
for i in $(seq 1 1000); do curl -s "http://localhost:8080/catalogue/sku-$i" >/dev/null; done
redis-cli -h "$CACHE_ENDPOINT" --tls --scan --pattern 'catalogue:*' | \
  xargs -n 100 redis-cli -h "$CACHE_ENDPOINT" --tls DEL
# SCAN, not KEYS: Redis is single-threaded and KEYS blocks every other command.

hey -n 5000 -c 500 "http://localhost:8080/catalogue/sku-1"
# Watch the source store: a spike of identical concurrent reads.
# (c) Fix it, then re-measure:
#   1. enable TTL jitter (jitteredTtl())
#   2. enable per-key request coalescing (the ReentrantLock)
# Re-run (b). The source-store read spike should collapse to a handful.
# (d) Prove graceful degradation: a cache outage must not be an outage.
aws elasticache reboot-cache-cluster --cache-cluster-id course-cache \
  --cache-node-ids-to-reboot 0001
# Keep requesting throughout. Expect: higher latency, ZERO errors.
# If you see errors, your cache client is a hard dependency. Fix that.

Cleanup:

aws elasticache delete-cache-cluster --cache-cluster-id course-cache
aws dynamodb delete-table --table-name orders
aws dynamodb delete-table --table-name idempotency

Production relevance: cache-aside with jitter, coalescing, negative caching and graceful degradation is the complete pattern. Most production caches implement the first half and discover the second half during an incident.

Production Considerations

ConcernPractice
Access patterns firstWrite the query list before the table. A DynamoDB table designed from entities will need a rewrite.
On-demand to startSwitch to provisioned + autoscaling once you have a month of data. Break-even around 15–20% utilization.
Watch GSI throttles separatelyA throttled GSI throttles base-table writes. This is how a table goes down.
ReturnConsumedCapacity in logsFree, exact cost telemetry per operation. Log it and you can attribute spend to endpoints.
PITR on anything that matters35 days of per-second recovery, for a small storage charge
TTL for ephemeral dataIdempotency keys, sessions, tokens. Free deletion, no WCU.
Never scan in a request pathIf you must scan, do it against an S3 export with Athena
Alarm on ThrottledRequests and SystemErrorsLeading indicators of key-design problems
Contributor InsightsShows the most-accessed keys — the direct way to find a hot partition
Cache is never a hard dependencyFalling back to the source must be the default behaviour, with a circuit breaker so the fallback does not kill the source
TTL jitter alwaysOne line. Prevents a whole class of outage.
Redis Multi-AZ with automatic failoverA single-node cache is a single point of degradation; at minimum, know that
Never KEYS in productionRedis is single-threaded. Use SCAN.

Failure Scenarios

S14 — Throttled at 40% of provisioned capacity

A table is provisioned at 20,000 WCU. It throttles at around 8,000. Adding capacity changes nothing.

QuestionAnswer
What can fail?A single partition, because the partition key is status — four values, one dominant
What happens?All PENDING_PAYMENT writes hash to one partition, capped at 1,000 WCU. The remaining 19,000 provisioned units are on partitions nothing touches.
Recovery?Not by adding capacity. This is a key-design problem.
Discriminating signalThrottledRequests high while consumed capacity is far below provisioned; CloudWatch Contributor Insights naming the hot key
Data lost / duplicated?Writes rejected; if the caller retries without idempotency, duplicates

The fixes, in order:

  1. Change the partition key to something high-cardinality (orderId). The real fix.
  2. Write sharding if a low-cardinality access pattern is genuinely required: PENDING# plus (hash % 10), then query 10 shards in parallel.
  3. Sparse GSI so the "find pending work" index contains only outstanding work and stays small.
  4. Adaptive capacity helps automatically — but it cannot exceed the per-partition ceiling, so it mitigates skew, not a single hot key.

The senior framing: "In DynamoDB, throughput is a property of your key design, not of your capacity setting. If you're throttled below your provisioned limit, stop looking at capacity and look at the keys."

S15 — A cache node restarts at 09:00 and the database falls over at 09:00:03

QuestionAnswer
What can fail?The cache node — reboot, failover, maintenance, eviction, or a synchronized TTL expiry
What happens?The cache empties. The source store has been serving 10% of read traffic and now receives 100%, instantly, with no ramp. It saturates, latency climbs, clients time out and retry, and the retries add load.
Recovery?Only once the cache refills — which requires the database to survive long enough to serve the misses
How quickly?Minutes, if the database survives. If it does not, it is a full outage.
Data lost / duplicated?No, but requests fail

The layered fix — no single mitigation is sufficient:

LayerMitigation
ExpiryTTL jitter. Keys stop expiring in lockstep.
ConcurrencyRequest coalescing. One fill per key, not 1,000.
FreshnessStale-while-revalidate. Serve the old value; refresh behind it.
TopologyReplicas. A failover does not empty the cache.
Warm-upPopulate before taking traffic; do not let a cold node into rotation.
ProtectionCircuit breaker + load shedding on the source, so overload degrades instead of collapsing (Day 12).

S16 — A "simple" nightly scan costs hundreds of dollars a month

A nightly job finds orders needing reconciliation with Scan + FilterExpression.

QuestionAnswer
What can fail?Nothing functionally. It works perfectly.
What happens?The scan reads every item in a 50 GB table to return about 200. On-demand read request units are charged for all of it. Once a night, forever, and the cost grows linearly with the table.
Recovery?A sparse GSI containing only items needing reconciliation
Discriminating signalConsumedReadCapacityUnits spiking nightly; ScannedCount vastly exceeding Count

The arithmetic:

50 GB table, 1 KB items            = about 50,000,000 items
4 items fit in one 4 KB read unit
eventually consistent, so 0.5 RCU per unit:
  50,000,000 / 4 x 0.5             = about 6,250,000 read request units per scan
nightly:  6.25M x 30 nights        = about 187M RRU
cost:     187M / 1M x $0.25        = about $47/month

Make the items 4 KB instead of 1 KB, or run it hourly instead of nightly,
and you are quickly into many hundreds of dollars - for 200 rows.

With a sparse GSI: about 200 items read per run, a fraction of a cent per month. Four to five orders of magnitude, from a key-design decision.

The general lesson, which is the whole day in one sentence: in DynamoDB, cost and latency are consequences of key design. There is no query planner to save you, and no index you can add later without having designed for it.

Security

QuestionAnswer
Who can access this?DynamoDB is IAM-only — no network endpoint of its own, no username/password. Access is entirely a policy question, and can be scoped to a table, an index, and even specific attributes or items matching a condition. Redis is inside the VPC, reachable only from app-sg.
What credentials are used?The task/execution role. No database credential exists for DynamoDB at all — which removes an entire category of leak.
Where is data encrypted?DynamoDB: at rest by default (AWS-owned key; use a CMK for control), in transit by TLS. Redis: --transit-encryption-enabled and at-rest encryption, plus AUTH/ACLs — all off by default on older cluster types, so be deliberate.
What if credentials leak?A scoped IAM policy limits damage to specific tables and operations. This is where DynamoDB's IAM-native model is genuinely better than a relational database's — you can express "may read only items whose tenantId matches this principal" as a policy condition.

Fine-grained access control, the feature worth knowing:

{
  "Effect": "Allow",
  "Action": ["dynamodb:GetItem", "dynamodb:Query"],
  "Resource": "arn:aws:dynamodb:eu-west-1:123456789012:table/orders",
  "Condition": {
    "ForAllValues:StringEquals": {
      "dynamodb:LeadingKeys": ["CUST#${aws:PrincipalTag/customerId}"],
      "dynamodb:Attributes": ["pk","sk","orderId","status","amount"]
    }
  }
}

That policy enforces multi-tenant isolation in IAM, not in application code — a caller can only ever read their own partition, and only the listed attributes. There is no equivalent in a relational database without row-level security and a per-tenant connection.

Redis-specific: enable encryption in transit and AUTH/ACLs. An unauthenticated Redis reachable from anything in the VPC is a data store with no access control at all, and it frequently holds sessions and tokens.

Performance

LeverEffect
Key designEverything. Nothing else on this list matters as much.
Eventually consistent readsHalf the cost, a few milliseconds of staleness
Projection typeINCLUDE with exactly what you need: less storage, cheaper writes, no second fetch
BatchGetItemUp to 100 items in one round trip. Remember to retry UnprocessedItems.
Parallel scanFor genuine full-table work, Segment/TotalSegments across workers
Item sizeCapacity is charged in 1 KB (write) / 4 KB (read) units, rounded up. Trimming a 4.1 KB item to 3.9 KB drops it from 2 read units to 1.
DAXMicrosecond reads for repeated keys; only worth it for read-heavy, high-repetition access
Redis pipeliningBatch commands to cut round trips
Local + remote cacheAn in-process cache with a short TTL in front of Redis fixes hot keys and removes network latency entirely
Redis single-threadedKeep commands O(1) or O(log n). No KEYS, no huge LRANGE, no unbounded Lua.

Cost

DynamoDB (us-east-1, illustrative — verify):

ItemRate
On-demand writes~$1.25 per million write request units
On-demand reads~$0.25 per million read request units
Provisioned WCU~$0.00013 per WCU-hour (about $0.095/WCU-month)
Provisioned RCU~$0.000065 per RCU-hour (about $0.047/RCU-month)
Storage~$0.25 per GB-month
PITR~$0.20 per GB-month
Streams reads~$0.02 per 100,000 read requests

The flagship order system's idempotency table (S17) — 50,000 orders/minute peak, about 260M writes/month:

On-demand:     260M / 1M x $1.25                        = about $325/month
Provisioned:   peak 833 writes/s x 1 KB = 833 WCU
               provision about 1,000 WCU with autoscaling
               1,000 x $0.095                           = about  $95/month

Saving: about $230/month, in exchange for managing autoscaling.
And with a 24-hour TTL, storage stays tiny instead of growing forever.

ElastiCache:

ItemRate
cache.t4g.micro$0.016/hr ($12/month)
cache.r7g.large$0.20/hr ($146/month)
ServerlessPer GB-hour stored + per ECPU consumed
Data transferCross-AZ applies — keep clients and cache in the same AZ where possible

Three cost lessons:

  1. A cache in front of Aurora is usually a cost saving, not just a latency one. Removing 90% of reads from a database can let you run a smaller instance.
  2. On-demand DynamoDB is a convenience premium of up to about 30x at steady load. Fine while you learn your traffic; expensive as a permanent choice.
  3. ReturnConsumedCapacity makes DynamoDB the most measurable data store you will use. Log it per operation and you can attribute cost to individual API endpoints — something almost impossible with a relational database.

Alternatives

Instead ofYou couldTrade-off
DynamoDBAurora/RDS (Day 6)Joins, ad-hoc queries, transactions, familiar tooling; one writer, size-dependent performance, operational scaling work
DynamoDBDocumentDB / MongoDBRicher queries and aggregation; you manage capacity and sharding, and it is not IAM-native
DynamoDBKeyspaces (Cassandra)Cassandra API compatibility; a smaller ecosystem on AWS
DynamoDB + RedisDynamoDB + DAXPurpose-built, write-through, no cache code to write; DynamoDB only, in-VPC, less flexible than Redis
ElastiCacheIn-process cache (Caffeine)Zero network latency, free; per-instance, not shared, and invalidation across instances is hard
ElastiCacheMemoryDBRedis API with durability — a primary store, not a cache; more expensive
Cache-asideRead replica (Day 6)A replica moves reads; a cache removes them. The cache is usually the better lever.
Redis locksDynamoDB conditional writesAtomic and durable. For correctness, prefer this.
ProvisionedOn-demandSimplicity vs cost; break-even around 15–20% utilization

Trade-offs

DecisionGainCost
DynamoDB over relationalPredictable latency at any scale, no operational scalingAccess patterns must be known up front; no joins; no ad-hoc queries
Single-table designOne round trip for related items, one capacity poolMuch harder to read; a new access pattern may need a redesign
GSI over LSIAdd/remove any time, own capacity, different PKEventually consistent only; a throttled GSI throttles base-table writes
Eventually consistent readsHalf priceA few milliseconds of staleness, occasionally visible to users
On-demandNo capacity planningUp to about 30x the cost at steady load
Conditional writesCorrectness under concurrency with no locksFailed conditions still consume capacity; every caller must handle the exception
TransactionsCross-item atomicity2x capacity; 100-item limit
A cacheLatency and cost reductionA new failure mode, stale data, invalidation complexity
Write-behind cachingFastest writesData loss on cache failure
TTL jitterPrevents stampedesSlightly less predictable expiry, which is the point

Common Mistakes

  1. Designing the table from entities instead of access patterns. The one mistake that requires a rewrite.
  2. A low-cardinality partition key — status, country, tenant tier, date. Guaranteed hot partition.
  3. Scan with FilterExpression and believing it is a WHERE clause. You pay for everything read.
  4. Expecting a GSI to be strongly consistent. It never is.
  5. Under-provisioning a GSI and taking down base-table writes.
  6. ProjectionType: ALL everywhere, doubling storage and write cost by default.
  7. Treating ConditionalCheckFailedException as an error. It is usually the answer.
  8. Ignoring UnprocessedItems from batch operations. Silent data loss.
  9. An LSI on an unbounded item collection. At 10 GB, writes to that key fail permanently and the LSI cannot be removed.
  10. No TTL on ephemeral data. An idempotency table that grows forever.
  11. A fixed cache TTL with no jitter. A synchronized expiry event waiting to happen.
  12. Updating the cache on write instead of invalidating. A race that caches the older value.
  13. The cache as a hard dependency. A cache outage becomes a service outage.
  14. KEYS on a production Redis. Single-threaded; it blocks everything.
  15. A Redis lock for correctness. Use a fencing token or a DynamoDB conditional write.

Interview Questions

2–3 YOE

  1. What is a partition key, and what is a sort key?
  2. Why does DynamoDB not support joins?
  3. What is the difference between Query and Scan?
  4. What is cache-aside?
  5. What is a GSI, and how does it differ from an LSI?
  6. What is the difference between a strongly consistent and an eventually consistent read?

4–5 YOE

  1. Your table is provisioned at 20,000 WCU and throttles at 8,000. What is happening?
  2. You need "all orders in status PENDING". Why is status a bad partition key, and what would you do instead?
  3. How do you make an operation safe when two requests might perform it simultaneously — with no locks?
  4. Explain how a cache node restart can take down a healthy database.
  5. When is a cache a better answer than a read replica?
  6. What does FilterExpression actually cost you?

Senior-Level Questions

6–7 YOE

  1. Design the DynamoDB model for a multi-tenant SaaS where tenants vary in size by four orders of magnitude. How do you stop the largest tenant from throttling the smallest?
  2. Design the caching strategy for a product catalogue of 5 million items with heavy read skew, and describe what happens on a cache failover.
  3. You need "all orders created today" on a table taking 50,000 writes/minute. Design it, and say what your design costs at read time.
  4. A team wants TransactWriteItems for every write "to be safe". Explain what it actually buys and what it costs.
  5. Design an idempotency mechanism for a payment API using DynamoDB. What do you store, keyed on what, for how long, and what is the residual failure window?

8–10 YOE

  1. A team proposes migrating a 400 GB relational table to DynamoDB "for scale". Make the case both ways, then decide — and say what evidence would change your mind.
  2. Where exactly is the consistency boundary in a "write the item, then query a GSI" flow, and what does a user observe inside it? How would you design around it?
  3. Your DynamoDB bill tripled with no traffic change. Describe your investigation and the three most likely causes.
  4. Argue that adding a cache makes a system less reliable. Then argue the opposite. Which do you believe, and under what conditions?
  5. Design the data layer for a system with both high-volume key-value access and a genuine need for ad-hoc analytical queries. How many data stores, and how do they stay consistent?

Day-End Revision

The five sentences

  1. DynamoDB hashes the partition key to pick a partition and orders items within it by sort key — which is why single-item and range reads cost the same at any table size.
  2. Design from access patterns, never entities; a low-cardinality partition key is a hot partition, and throughput is a property of key design rather than of your capacity setting.
  3. GSIs are always eventually consistent, have their own capacity, and a throttled GSI throttles base-table writes.
  4. Conditional writes give correctness under concurrency without locks, and ConditionalCheckFailedException is usually an answer rather than an error.
  5. Cache-aside with TTL jitter, request coalescing and graceful degradation is the complete pattern; a cache must never be a hard dependency.

The diagram to redraw from memory: the hot-partition comparison — pk = status versus pk = orderId.

The numbers

Max item size400 KB
Per-partition limits10 GB, 3,000 RCU, 1,000 WCU
1 RCU1 strong or 2 eventual reads of 4 KB/s
1 WCU1 write of 1 KB/s
Transactional ops2x capacity
BatchGetItem / BatchWriteItem100 / 25 items
TransactWriteItems100 items, 4 MB
LSI item collection limit10 GB
Streams retention24 hours
PITR window35 days
On-demand vs provisioned break-evenaround 15–20% utilization
TTL deletiontypically within 48 h, free

Today's trap: "DynamoDB is always faster." It is predictable at any scale, which is different. A well-indexed relational database beats it on complex access patterns, and DynamoDB punishes access patterns you did not design for.

Tomorrow needs: conditional writes (the mechanism behind idempotency), the idempotency table with TTL that you created today, eventual consistency as a general phenomenon, and the observation that a retry of unknown outcome needs a dedupe key.

Mini Assignment

Time: 50–60 minutes. Capstone contribution: the DynamoDB idempotency table and the Redis catalogue cache.

  1. In day-07/, write the access-pattern list first, in NOTES.md, before any code. Then design the keys.
  2. Implement the orders table with gsi1 (customer history) and a sparse gsi3 (pending work queue). Seed 500 orders.
  3. Implement markPaid with a conditional write and prove that 10 concurrent calls produce exactly one state change.
  4. Demonstrate the hot partition. Create a second table with pk = status, provision it at 100 WCU, drive writes at 300/s, and capture ThrottledRequests. Then fix it with write sharding and capture the difference.
  5. Do the scan-vs-query cost comparison with --return-consumed-capacity TOTAL. Report ScannedCount, Count and capacity for both. Compute the monthly cost of running the scan hourly.
  6. Add the Redis cache-aside layer with jitter, coalescing and negative caching. Cause a stampede, measure it, fix it, re-measure.
  7. Prove graceful degradation: reboot the cache node under load; requests must slow, not fail.
  8. In NOTES.md:
    • Your access-pattern table and the key design that satisfies each in one Query or GetItem.
    • The hot-partition before/after numbers, with an explanation of why more capacity did not help.
    • The scan-vs-query capacity comparison and the extrapolated monthly cost.
    • Stampede source-store read count before and after jitter + coalescing.
    • For the idempotency table: on-demand vs provisioned cost at 260M writes/month, and which you would choose.
    • One access pattern your design does not support well, and what you would do if it became a requirement.
  9. Delete the cache cluster and the tables. Commit.

Success criterion: your hot-partition demo shows throttling below provisioned capacity, and your stampede fix reduces the source-store spike by at least an order of magnitude.

AWS Documentation


Previous: Day 6 — Relational Data in Production · Next: Day 8 — Queues, Delivery Semantics and Idempotency (Phase 4)