Watching an $11 site: budgets, alarms, and a canary
Part of my Krabber series, a Twitter clone in Go. The full source is on GitHub.
Intro
A site I run alone needs to tell me when something’s wrong, because I’m not sitting there watching it. Krabber has a small observability setup that covers three questions: is it up, is it misbehaving, and is it about to cost more than it should? All of it notifies one email address, and the whole thing is designed to stay inside the free allowances. Here’s how it’s put together.
I. Is it about to cost more than it should?
Since I pay for Krabber myself, cost gets first-class monitoring. There’s a daily budget of $1.50 (normal is well under that) and a monthly budget of $15, each raised in step as I let more krabs in. On top of those, Cost Anomaly Detection watches for a spend pattern that’s weird even if it’s under budget, and emails me if it sees one.
And there are two alarms specifically about sneaky costs. One watches CPUSurplusCreditsCharged, which fires the moment the instance’s burstable-CPU surcharge starts (it’s capped at about eight cents an hour, but I want to know). Another is a metric filter on the app’s own “mail cap reached” log line, because if I’m suddenly hitting the daily email cap, someone’s probably abusing signup or password reset.
II. Is it misbehaving?
A set of CloudWatch alarms, eight in all and within the ten free ones, each watch a specific failure, and all notify the same krabber-alerts email:
- Beanstalk environment health degraded or severe for five minutes.
- DynamoDB throttling, which means either an attack or that I’ve grown into my caps and should raise them.
- CloudFront 5xx rate over 5 percent, meaning the origin is failing behind the edge.
- SES bounce rate over 4 percent and complaint rate over 0.08 percent, both well below the levels where AWS starts asking questions.
Each one turns a silent problem into an email while it’s still small.
III. Is it up? The canary
Alarms watch metrics, but I also want a real end-to-end check: can an actual request get through DNS, TLS, CloudFront, the WAF, and the app, and come back right? That’s a tiny Lambda canary that runs every five minutes. It fetches /healthz and /krab/login and fails if either isn’t a 200, or if the login page is missing its form. If it fails twice in three runs, I get paged.
The interesting choice here was cost. The obvious tool is a Route 53 health check, but its many global checkers would send more than a million requests a month through CloudFront, which would blow past the Free plan’s allowance and start costing money just to ask “are you up?” My canary sends about 17,000 requests a month instead. It’s a good example of the whole project’s theme: even the monitoring has to respect the budget.
IV. And when something is wrong, where did the cost go?
When I do need to dig in, every logged request carries its own cost. Each line records rru, wru, and ddb_calls, the DynamoDB read units, write units, and call count it used, and fan-out lines record the units per molt. So a single Logs Insights query tells me which pages are expensive:
stats sum(rru), sum(wru) by path
Normally I log a 5 percent sample to keep log costs down, but flipping LOG_ALL_REQUESTS=true for a while logs every request when I’m chasing something. The logs live in CloudWatch for 14 days, which is plenty for a site this size.
Conclusion
Observability for a one-person project isn’t about dashboards you’ll never look at, it’s about being told, by email, the moment something is down, misbehaving, or getting expensive. Budgets and anomaly detection guard the bill, a handful of focused alarms catch the common failures, a cheap canary proves the whole path end to end, and per-request cost logging means I can always answer “where did it go.” It all fits inside the free tiers, because even the watching has to stay cheap. Thanks for reading, and may your alarms stay silent.