One table to rule the reef
// written in 2023: tools, versions and prices may have changed since.
Part of my Krabber series, a Twitter clone in Go. The full source is on GitHub.
Intro
Krabber stores everything in one DynamoDB table. Accounts, molts, replies, likes, follows, the home feed, notifications, sessions, and the rate-limit counters all live together in a table literally named krabber-prod. No Postgres, no Redis, no separate store for sessions. Just one table.
People hear “single-table design” and picture some kind of magic trick. It isn’t. It’s really just a few boring rules applied consistently, and in this post I want to walk through the ones that matter so you could do the same on your next project. This is the data model half of the story; the cost post is its companion.
I. The table has two keys, and that’s it
Every item has a partition key PK and a sort key SK. That is the whole schema. DynamoDB doesn’t care what you put in them, so the keys are the design: you encode the entity type and its identity into strings, each with a prefix.
PK SK what it is
───────────────────────── ────────── ─────────────────────────
C#alice@example.com C# a crab (account)
M#<molt-id> M# a molt
C#alice@example.com F#<crab> alice follows someone
RL#signup-ip#<network> RL#<epoch> a rate-limit counter
U#alice U# a username reservation
So a crab is just keyed by a normalized email:
func crabPK(email string) string { return "C#" + normalizeEmail(email) }
func crabSK() string { return "C#" }
Those prefixes do two jobs. First, they keep entity types from colliding in the same keyspace. And second, they make begins_with queries cheap, so “every follow for this crab” is one query on PK = C#alice..., SK begins_with "F#". Simple, but it’s the whole trick.
II. Access patterns first, then indexes
In a relational database you model the data and figure out the queries later. In DynamoDB you do the opposite. You write down every read your app performs, and then you design keys and indexes so each one is a single query. Krabber has a handful of global secondary indexes, and each one earns its place against a specific access pattern:
- GSI2 looks a crab up by ID instead of by email.
CreateCrabwritesGSI2PK = GSI2SK = <crab-id>on the account item, projectedALL, because once you’ve found a crab by ID you usually want the whole thing. - GSI8 is the purge queue. When an account is deleted it goes on this index, so a background job can clean up its molts, follows, and likes without ever scanning the table.
- A few more handle feeds and lookups, projected
KEYS_ONLYwhere I only need to find the keys andALLwhere I need the item right there. Projection type is a cost lever:KEYS_ONLYis cheap to write,ALLsaves you a follow-up read. You pick per index.
Now, the rule I follow here is to add an index only when a real access pattern needs it, and to project the least I can get away with. Every index is a second copy of part of your data written on every write, so they aren’t free. Make them earn their keep.
III. Uniqueness without a UNIQUE constraint
DynamoDB has no UNIQUE, so you build it yourself with a conditional write and a marker item. When a crab claims a username, I write a U#<name> item guarded by a condition:
// write the marker only if nobody already holds it
ConditionExpression: aws.String("attribute_not_exists(PK)")
If the name is taken, the condition fails and the whole write rolls back. Same trick for “email already in use.” The marker item is the constraint. And when you change your username, the old marker is deleted and the new one claimed in the same transaction, so two krabs can never hold the same handle. It took me a minute to trust this instead of reaching for a database that hands you UNIQUE, but it works and it’s cheap.
IV. Sessions and rate limits live here too
This is the part where people get nervous, and also where it pays off the most. Sessions are just items. Rate-limit counters are just items:
func rateLimitPK(action, key string) string { return "RL#" + action + "#" + key }
func rateLimitSK(start time.Time) string { return fmt.Sprintf("RL#%d", start.Unix()) }
So “twenty signups per IP per five minutes” is a counter item keyed by the action, the network, and the time window. No Redis, no second thing to run, back up, and pay for. And because these items carry an expires_at attribute and the table has TTL turned on, DynamoDB deletes expired sessions and stale counters for me, for free. The rate-limit window from last week’s bot attack is already gone, and I never ran a cleanup job.
That last point had a nice payoff, by the way. When scrapers hit the site, I could prove they never tried to sign up, because no RL#signup-ip# items existed for that window. The security control left an audit trail for free. (I tell that whole story in the scraper post.)
V. Fan-out, and surviving a crash mid-write
The home feed, the Trench, is built by fan-out: when you molt, that molt is written once into each follower’s feed. That’s a lot of writes, and a deploy or a crash can absolutely land in the middle of one.
So a molt that’s waiting for fan-out is marked in the table. The work isn’t fire-and-forget in memory, it’s a durable flag. After a restart, the next instance sees the mark and finishes the job. In other words, the table isn’t only storage here, it’s also the queue and the checkpoint, so an interrupted fan-out resumes instead of quietly dropping half of someone’s followers.
VI. The guard rails that keep it boring
Two table settings do a lot of quiet work:
- Maximum-throughput caps on the table and every index. The table is on-demand, but a cap means a runaway loop or a traffic spike gets throttled rather than billed into the stratosphere. The app paces its own bursts, like fan-out, to stay under them. I’d rather a known ceiling than an unknown bill.
- Point-in-time recovery and deletion protection, with
prevent_destroyin Terraform on top. One table holds the entire site, so the blast radius of a badterraform applyis, well, everything. Belt and suspenders are worth it here.
VII. When single-table is the wrong call
I’m not a zealot about this, so let me be fair to the other side. Single-table shines when you know your access patterns, your entities are related, and you want one thing to operate and pay for. It’s miserable when your queries are ad hoc and analytical. Something like “all molts last Tuesday containing the word reef” is a scan, and scans are exactly where this design falls apart. Krabber’s reads are all known and keyed ahead of time, so it fits. If yours aren’t, don’t force it.
Conclusion
The payoff, though, is real: one table, one bill, one thing to back up, with sessions and rate limits included for free. For a site I pay for out of my own pocket, that boring simplicity is the whole point. If you’re building something small and you can list your access patterns on one page, give single-table a try. Thanks for reading, and may your queries stay single.