The day krabber.net went dark
Part of my Krabber series, a Twitter clone in Go. The full source is on GitHub.
Intro
So krabber.net went down. Not slow, not throwing an error page, just gone. My own browser couldn’t even find it.
I’ve run krabber.net on and off for years, and this was from the latest time I spun it back up. A fresh spin-up came with a fresh kind of failure, one I’d never hit on the earlier rounds, which is exactly why it took me a while to place it.
If you’ve ever shipped something the night before an outage, you know the first feeling: what did I break? I’d pushed a small mobile change the day before, so naturally I assumed it was me. It wasn’t. But it took me a few wrong turns to hear what the evidence was trying to tell me, and since “the site is just down” is a debugging exercise worth practicing, I figured I’d write down the whole walk, commands and all.
I. The stack, so the errors make sense
Krabber is deliberately boring (I wrote up the whole architecture here). One CloudFront distribution out front, one Elastic Beanstalk instance behind it, and one DynamoDB table underneath.
browser ──▶ CloudFront + WAF ──▶ Elastic Beanstalk (one box) ──▶ DynamoDB
krabber.net origin.krabber.net
CloudFront’s origin is a record named origin.krabber.net. Keep that name in mind, because it matters later.
II. Is the app even up?
I started at the bottom and worked up. First question: is the box actually serving? Beanstalk will tell you its own health without you having to log in anywhere.
aws elasticbeanstalk describe-environments --environment-name krabber-prod \
--query 'Environments[].{Status:Status,Health:Health}' --output table
# Status: Ready Health: Green
Green. So the app was up, the process was running, and the health check was passing. The problem was somewhere between the box and the world.
III. CloudFront says 502
Next I asked CloudFront directly, resolving the domain to the distribution so I could skip whatever my laptop’s DNS was doing:
curl -sS -I https://<distribution>.cloudfront.net/ -H 'Host: krabber.net'
# HTTP/2 502
# server: CloudFront
# x-cache: Error from cloudfront
A 502 from CloudFront means it couldn’t get a usable answer out of the origin. So the app was green but the edge couldn’t reach it, and because I’d recently set up HTTPS between CloudFront and the box, TLS was the first thing I suspected. That suspicion cost me a good half hour.
IV. Down the TLS rabbit hole
To check the certificate I needed to be on the box. There’s no SSH key on that instance by design, and I don’t miss it. Krabber runs on the Amazon Linux platform, which ships the SSM agent out of the box, so Session Manager just works: no key pair to store or rotate, and nothing listening on port 22. (On a lot of other images you’d be installing and wiring up the agent yourself first. More on why I picked that platform in a later post.) So I hopped straight on:
aws ssm start-session --target <instance-id>
Once I was on, I checked that nginx was alive and listening where it should be:
systemctl is-active nginx
# active
sudo ss -tlnp | grep -E ':(443|5000)'
# nginx on :443, the Go app on :5000
sudo nginx -t
# syntax is ok
Then the certificate itself. The thing CloudFront cares about is whether the origin presents a complete, trusted chain, so I had openssl walk it and print the verdict:
echo | openssl s_client -connect 127.0.0.1:443 -servername origin.krabber.net 2>/dev/null \
| grep 'Verify return code'
# Verify return code: 0 (ok)
And just to be sure the box would answer a real request, I curled it from itself:
curl -ki https://127.0.0.1:443/ -H 'Host: krabber.net'
# HTTP/1.1 200 OK
So the box was serving a clean 200 over a valid TLS chain, to itself, the whole time. The origin was fine. I’d spent half an hour on the one thing I’d touched recently, and it was innocent. Which, annoyingly, is exactly what you’d expect if you were debugging your bias instead of the system.
V. Read the actual error
Here’s the lesson I keep re-learning: read the error before you theorize about it. CloudFront’s 502 page isn’t generic. If you fetch the body instead of just the status line, it tells you why:
curl -s https://<distribution>.cloudfront.net/ -H 'Host: krabber.net'
# ...
# CloudFront wasn't able to resolve the origin domain name.
Resolve. Not connect, not handshake. CloudFront couldn’t turn origin.krabber.net into an IP address. This had been DNS the whole time, and I’d been on the box checking certificates because certificates were fresh in my mind.
VI. It was DNS
Here’s the part that threw me. When I asked my own Route 53 nameservers directly, everything resolved perfectly:
dig @<route53-nameserver> origin.krabber.net +short
# krabber-prod.<...>.elasticbeanstalk.com.
# <the right IP>
So the zone was fine and the records were all there. Yet a normal resolver, the kind CloudFront and your browser use, said the domain didn’t exist:
dig @8.8.8.8 krabber.net +short
# (nothing)
dig @8.8.8.8 krabber.net | grep status
# status: NXDOMAIN
NXDOMAIN, not SERVFAIL. SERVFAIL would have pointed me at DNSSEC. NXDOMAIN means something authoritative is saying “this name does not exist.” So I traced it from the root to see who was saying it:
dig +trace krabber.net | tail
# the .net servers hand back their own SOA, no delegation to my nameservers at all
There it was. The .net registry wasn’t handing out my delegation anymore. My zone still existed and still answered anyone who asked it directly, but the parent had stopped pointing at it, so no resolver could find it. That is a registry-level action, not anything in my account.
VII. The real cause: a 15-day timer
One whois confirmed it:
whois krabber.net | grep -i 'domain status'
# Domain Status: clientHold https://icann.org/epp#clientHold
clientHold. The registrar had told the registry to pull my domain out of DNS entirely. The “why” is almost funny. I’d registered krabber.net 15 days earlier, and ICANN requires registrars to verify the registrant’s email within 15 days. If you don’t click the link, they have to suspend the domain. Amazon’s Route 53 tracks that as a reachability status:
aws route53domains get-contact-reachability-status --domain-name krabber.net
# { "domainName": "krabber.net", "status": "PENDING" }
PENDING. The verification mail had gone to an address whose spam filter quietly held it, and on day 15 a cron job at the registrar flipped the switch. Not my deploy, not my TLS, not Beanstalk. A timer.
The fix is to verify the email. If you can’t find the original, you can have it resent:
aws route53domains resend-contact-reachability-email --domain-name krabber.net
I clicked the link, and about an hour later the registrar lifted the hold, .net put the delegation back, and the site came back on its own.
VIII. One root, three symptoms
While all of this was going on I also got an SES warning that the mail record for mail.krabber.net had “gone missing.” It hadn’t, the record was right there. But SES validates with a normal resolver too, so it saw the same NXDOMAIN as everyone else and assumed the record was gone.
Site down, CloudFront 502, SES “missing” record: three symptoms, one cause. When a single root explains everything you’re seeing, you’ve probably found it. When you’re patching three things separately, you probably haven’t.
Conclusion
A few things I’d tell myself before the next one:
- Read the error first. CloudFront told me “resolve” in plain English on the very first 502, and I SSH’d… well, SSM’d onto the box to poke at TLS anyway, because TLS was fresh in my mind. Recency isn’t evidence.
- “Green” is narrow. Beanstalk health means the app answered a health check. It says nothing about DNS, the edge, or the domain itself.
- Query the authoritative nameserver and a normal resolver, and compare. The gap between “resolves at the nameserver” and “
NXDOMAINeverywhere” is the tell for a parent-level problem, and it would have saved me the whole rabbit hole. - Check
whoisstatus early. It’s free and it’s one line, andclientHoldturns “mysteriously down” into “oh.”
The site was down for most of a day over an unclicked link in an email I never saw. The least I can do is get a blog post out of it. So if you run your own domains, go check that the registrant email is one you actually read, and put that 15-day verification somewhere you won’t miss it. Thanks for reading, and may your crabs stay afloat.