Metering every request so the bill can't surprise me

// written in 2023: tools, versions and prices may have changed since.

Part of my Krabber series, a Twitter clone in Go. The full source is on GitHub.

Intro

I pay for Krabber out of my own pocket. That one fact drives every design decision I make, and it turns cost from a thing you check on the invoice at the end of the month into a thing you test in CI, the same way you test anything else.

The goal I set was simple to say and annoying to hold: under $100 a month at 10,000 daily active krabs, with a bill that has a known ceiling even under attack. At hobby traffic it runs about $11. This post is how I keep it there, and it pairs with the single-table post on the data model underneath it.

I. You can’t optimize what you can’t see

DynamoDB bills by capacity units, which is roughly how much data a request reads or writes. The trap is that a single page can quietly cost 40 units because it does the convenient thing forty times. You won’t see it in the latency, so it won’t bother you while you’re coding. You’ll see it on the invoice, a month later, when it’s already shipped.

So I put a meter on the DynamoDB client itself. Every call adds up the capacity it consumed and tags it to the request that made it. In production every logged request then carries its own read and write units, so answering “which pages cost the most” is one query instead of a guess.

And that meter found the obvious crime right away. A feed page was costing 45–47 read units because it looked up likes, remolts, and bookmarks one molt at a time, so twenty molts meant sixty little point-lookups. I rewrote it as one range query per type and it dropped to 16–18 units. I also changed polling for new molts from constant to once a minute, and only while the tab is actually open. Before those two changes, 10,000 daily active krabs would have cost about $134 a month. After, about $81. The traffic didn’t change at all, only the shape of the queries did.

II. A test that fails when a page gets greedy

Measuring once is a chore you’ll quietly skip next quarter. So I made the measurement a test. TestCostProfile seeds 25 krabs, loads every main page, poll, and action, sums the capacity each one consumes, and fails if any of them goes over its budget:

go test ./internal/web -run TestCostProfile
# --- FAIL: TestCostProfile
#     cost_test.go: GET /trench used 46 read units, budget 20

That one assertion changes the whole dynamic. A cost regression stops being something I discover on the bill, and becomes a red build sitting right next to the line that caused it, before the code ever merges. If I write a handler that does a point-lookup in a loop, the test tells me on my laptop while I still remember why. A small script turns the same numbers into the cost table on the project page, so that table is generated from the real test, not hand-waved.

III. Where the money actually goes

Here’s the thing measuring taught me that intuition never would have. When you add it up, the breakdown is surprising:

  • Writes cost the most, and likes are 55% of them. A like isn’t one write, it writes the like, its index entry, and a notification. So the cheapest-feeling thing a user can do is the most expensive thing I store.
  • Trench fan-out is another 26% of writes, because every molt is written once more per follower. Popular krabs are expensive krabs.
  • Reads come next, and feed pages are 41% of those, mostly from loading 20 molts at a time.

I would not have pointed at “likes dominate the write bill” before I measured it. And that’s exactly the kind of thing that decides whether you come in under budget, which is why guessing wasn’t going to cut it.

IV. The ceiling, for when I’m wrong or attacked

Optimizing the average case isn’t enough, you also need a roof for the bad day. So the guard rails:

  • Maximum-throughput caps on the table and every index. Requests above the cap are throttled, not billed, so a runaway loop or a traffic spike can make the site slow but it can’t make the bill explode. The app paces its own bursts, fan-out especially, to live under the caps.
  • CloudFront on the flat-rate plan, which bundles WAF, DDoS protection, and DNS and never bills overage. The edge is where a flood lands, and the edge has a fixed price.
  • On-demand everything, so idle costs nothing. At hobby traffic the bill is mostly the instance, its public IP, and its disk, the parts that exist whether or not anyone shows up.

Put together, a bad actor can cost me latency and some throttled requests. What they can’t do is hand me a surprise invoice. And honestly that was the real requirement all along, not “cheap” but “bounded.”

Conclusion

Treating cost as a feature just means giving it the same tools as everything else: a measurement, a test, a budget, and a ceiling. The meter makes it visible, the test keeps it honest, and the caps make it safe. I didn’t make Krabber cheap by being clever once, I made it cheap by never letting a greedy query reach production without a red build first. If you run something you pay for yourself, put a meter on your most expensive dependency and wire it to a test. Thanks for reading, and may your invoices be boring.

cd ../blog