Terraform in two stacks: the one you apply by hand, and the one CI applies
Part of my Krabber series, a Twitter clone in Go. The full source is on GitHub.
Intro
All of Krabber’s AWS infrastructure is Terraform, but it’s in two separate stacks, and that split is on purpose. One stack I apply by hand, exactly once. The other is applied automatically by CI on every merge. In this post I’ll explain why it’s two instead of one, because the reason is a nice little chicken-and-egg problem.
I. The chicken and the egg
My CI pipeline authenticates to AWS with OIDC and assumes a role, and it stores Terraform state in an S3 bucket. So CI needs two things to exist before it can do anything: the IAM role it assumes, and the state bucket it reads.
But those things are themselves infrastructure. If I put them in the same Terraform that CI runs, then CI would need to create the very role it uses to run, and the bucket that holds the state it’s reading. That can’t work. Something has to create the foundation before CI can stand on it.
So the foundation is its own stack.
II. The bootstrap stack, applied once
The bootstrap stack is the foundation, and I apply it once, locally, as my own IAM user. It creates the things CI depends on and the account-wide settings you only set up once:
- the S3 bucket that holds Terraform state,
- the GitHub OIDC provider and the two CI roles (plan and deploy),
- the workload roles the app uses,
- and account hygiene like the password policy, account-wide S3 block-public-access, and the security contact.
There’s a neat moment in applying it: the bootstrap stack starts with its state in a local file, and then, once it has created the state bucket, it moves its own state into that bucket with terraform init -migrate-state. It builds the shelf and then puts itself on it.
I apply this by hand because it almost never changes, and because it creates IAM, which I deliberately do not let CI touch.
III. The prod stack, applied by CI
The prod stack is everything else: the DynamoDB table, the Beanstalk app and environment, CloudFront and the WAF, SES, the alarms. This is what changes as I build, so this is what CI applies, using the deploy role the bootstrap stack made. The deploy role can manage these services but can’t create IAM, which is exactly the boundary I want: day-to-day infrastructure changes flow through CI, but the ability to grant permissions stays in the bootstrap stack and my own hands.
A couple of details that make it pleasant. The state lives in S3 with use_lockfile = true, which is Terraform 1.10’s native S3 locking, so there’s no separate DynamoDB lock table to run just to stop two applies from colliding. And because the state can contain generated secrets, that bucket is locked down to only the identities that genuinely need it.
Conclusion
Splitting Terraform into a hand-applied bootstrap and a CI-applied prod stack isn’t extra ceremony, it’s the answer to a real ordering problem: CI can’t create the things it needs to run. Put those in a foundation you apply once, keep IAM in that foundation and out of CI’s reach, and let the pipeline own the day-to-day infrastructure on top. If your CI provisions its own cloud, draw the line between “what bootstraps the pipeline” and “what the pipeline runs,” and you’ll find it wants to be two stacks. Thanks for reading, and may your state always lock.