Durable fan-out: background work that survives a crash
Part of my Krabber series, a Twitter clone in Go. The full source is on GitHub.
Intro
On Krabber, your home feed is called the Trench, and it’s built by fan-out. When you post a molt, that molt gets written once into the feed of every krab who follows you. For someone with a few hundred followers that’s a few hundred writes kicked off by one click, and it all happens in the background so the post feels instant.
The interesting problem isn’t the writing. It’s what happens when the process doing that writing gets interrupted halfway through, by a deploy or a crash. In this post I’ll show how Krabber makes that background work durable without adding a queue service, because the answer is a nice reuse of the one database it already has.
I. In-process, on purpose
Krabber runs on a single long-running Beanstalk box (I explained why, not Lambda), and the fan-out runs right there in the same Go process, as goroutines. There’s a worker that does Trench fan-out and another that refreshes the public timeline on a ticker. No separate queue service, no message broker.
That’s cheap and simple, but it raises an obvious question. If the work lives in memory on one box, what happens when that box goes away? And it does go away, routinely, every time I deploy, because an immutable deploy boots a fresh instance and retires the old one. A molt that was halfway through fanning out when the old box shut down can’t just vanish.
II. The table is the queue
So the work isn’t really “in memory.” A molt that still needs fan-out is marked in DynamoDB, as a durable record that the job is pending. The worker reads that mark, does the writes, and clears it when it’s finished.
The payoff is in the crash case. If the box dies mid-fan-out, the mark is still sitting in the table, because it was never only in memory. When the next instance comes up, its worker sees the pending mark and finishes the job. The molt lands in every follower’s feed, just a little later than usual. Nobody gets a half-delivered post.
In other words, the table is doing double duty. It’s the queue that holds the pending work, and it’s the checkpoint that says how far we got. I didn’t have to run RabbitMQ or SQS to get “do this reliably, and resume if interrupted,” because the database I already have for everything else can hold a to-do item just fine.
III. Why resuming is safe
Resuming only works if doing part of the job twice doesn’t corrupt anything, and here the data model makes that free. Writing a molt into a given follower’s feed is keyed by that follower and that molt, so doing it again just writes the same item to the same key. It’s idempotent. A resumed fan-out that re-covers some ground it already covered does no harm, it just overwrites identical items.
That’s the quiet reason this pattern works: at-least-once delivery plus idempotent writes equals exactly-once in effect, without any of the machinery people usually build to get there. I lean on it the same way for the other background jobs, like expiring sign-ups that were never activated.
Conclusion
The thing I like about this is that durability came from the design, not from a new service. Because the work is marked in the same table that holds everything else, a deploy or a crash can interrupt a fan-out and the next instance quietly finishes it, and because the writes are idempotent, finishing it twice is harmless. For a one-box app that redeploys by replacing the box, that’s exactly the property you need. If you’re about to add a queue service for background work, check whether your database can hold the to-do list and the checkpoint first. Thanks for reading, and may your molts always reach the trench.