Reading the scrapersphere
Part of my Krabber series, a Twitter clone in Go. The full source is on GitHub.
Intro
This one is a sequel to the outage post. While I was busy chasing that, I noticed something else in the WAF dashboard: a spike. About 3,100 requests in a window where I normally get dozens. Roughly 2,100 allowed, 992 blocked, almost all from a single country, and almost all caught by the rate-limit rule.
So bots were crawling krabber.net hard, and they seemed to know exactly where to go. That left me two questions worth answering honestly before I panicked: did they actually do any damage, and how did they know my paths? Let’s take them in order.
I. Did they create any accounts?
The scary version of this story is “bots mass-registered thousands of fake krabs.” The real version is in the data, so I went and looked instead of guessing.
Krabber keeps everything in one DynamoDB table, and a crab (an account) is just an item whose sort key is C#. So counting accounts is one scan:
aws dynamodb scan --table-name krabber-prod \
--filter-expression 'SK = :sk' \
--expression-attribute-values '{":sk":{"S":"C#"}}' \
--projection-expression 'created_at, activated'
# 2 items. Both activated, both weeks old. Zero new during the attack.
Two accounts, both mine-adjacent, both old. Nothing new during the attack. So far so good, but I wanted to know if the bots even reached signup. Here’s a nice side effect of keeping rate limits in the same table: every signup POST writes a per-IP counter (RL#signup-ip#...) before anything else happens, so I can just scan for those too:
aws dynamodb scan --table-name krabber-prod \
--filter-expression 'begins_with(PK, :p)' \
--expression-attribute-values '{":p":{"S":"RL#signup-ip#"}}'
# 0 items, they never even knocked on signup
Zero. The bots never hit the signup endpoint at all. The nginx log backed that up, it was GET /trench, GET /sea, GET /static/... on repeat. This was crawling, not account abuse. They were walking the public pages, not trying to get in.
And even if they had tried, the signup funnel is three gates deep: a per-IP rate limit, a Turnstile bot check, and email activation before an account can do anything (unactivated accounts expire). A scraper reading HTML clears none of those. So the honest answer to “did they do damage” is no.
II. How did they know my paths?
This one bugged me more, because the bots asked for real routes like /trench, /sea, /molt/view/..., not a generic wordlist. It felt like they had a map.
They did. And I handed it to them.
I ruled out the obvious leaks first. My repo is private, so it wasn’t that. My service worker only references /offline, so it wasn’t that either. Then I opened my own robots.txt and there was the map:
User-agent: *
Disallow: /krabmin
Disallow: /settings
Disallow: /trench
Disallow: /search
A scanner’s first move is to fetch robots.txt and crawl the Disallow lines, because that’s exactly where people put the interesting stuff. My logs showed hits on /trench, a disallowed path, which is the fingerprint of a bot treating robots.txt as a treasure map. The rest of the routes are just links in the page HTML, visible to anyone who loads the site. Nothing was cracked. The paths were never secret.
So the lesson landed on me, not the bots: robots.txt is advisory, and it’s public. Listing an admin path there doesn’t protect it, it advertises it. The paths that actually need protecting are already behind auth, and they don’t need to be named in a file that scanners read first.
III. What actually defends a one-box site
Krabber is a single Elastic Beanstalk instance. There’s no autoscaling to ride out a flood, so the defense has to live at the edge. Here’s what actually carries the weight, roughly in order of importance, and notice that none of it is robots.txt.
First, the origin is unreachable except through CloudFront. The instance’s security group only allows port 443 from CloudFront’s managed origin-facing prefix list, and nothing else:
# 443 inbound from CloudFront's origin-facing prefix list, nothing else
prefix_list_ids = ["pl-xxxxxxxx"] # com.amazonaws.global.cloudfront.origin-facing
A scraper that finds the box’s address and tries to connect directly just times out. Everything has to come through the edge, which means everything passes WAF and the rate limits first. This is the single most important control, and it’s about one line of Terraform.
Second, WAF, layered. Behind the distribution I run a small stack of rules: a global per-IP rate limit (the one that ate the 992 requests), a tighter limit on the auth endpoints, the AWS IP-reputation list, the anonymous-IP list (VPNs, Tor exit nodes, datacenters, which is where scrapers live), and a geo-block for countries I’ll never have krabs in. The AWS managed groups are free, so this is cheap.
Third, caching where I can. Static assets and the generated avatars are cached at the edge, so repeat crawling of those never touches the box. (HTML pages aren’t cached yet. That’s the next lever and the trickiest one, because you can only safely cache the signed-out versions.)
And fourth, the app funnel from section I: Turnstile, per-IP limits on the expensive endpoints, and email activation. Defense in depth for the one path bots actually want.
IV. What I changed afterward
The attack was absorbed, but I used it as a nudge to tighten a few things:
- Blocked the offending country at the edge with a WAF geo-match rule.
- Turned the AWS Common Rule Set from “count” (watching) to “block” (acting), after checking it wasn’t flagging my own htmx posts.
- Added the anonymous-IP managed group, since scrapers love a datacenter IP.
- Added explicit
Disallowlines inrobots.txtfor the AI and LLM crawler user-agents. It’s advisory, so the polite ones honor it and the rude ones get caught by the rules above.
The one I’m still eyeing is swapping the general rate-limit’s action from an outright block to a WAF Challenge. A real browser solves the silent challenge without noticing, while a headless scraper fails it, so you’d stop the bots without ever showing a human a wall. I haven’t flipped that one yet, but it’s next on the list.
Conclusion
Here’s the honest framing. The bots didn’t take my site down, a registrar hold did. They created nothing, broke nothing, and were soaked up by rules that cost a couple of dollars a month. The only real finding was about me: I’d published a map of my own sensitive routes and called it a security file.
So if you run a small site, go read your own robots.txt the way a scanner would, lock your origin to your CDN so nobody can skip the edge, and remember that the controls doing the real work are the boring ones. The scrapersphere is going to keep knocking either way. Thanks for reading, and may your logs stay quiet.