terraform-ops
Terraform and OpenTofu infrastructure-as-code operations - project layout, state management, module design, plan/apply safety, CI/CD pipelines, and secrets. Use for: terraform, opentofu, infrastructure as code, IaC, tfstate, terraform state, terraform module, remote backend, terr
Install
npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/terraform-ops
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
git clone https://github.com/0xDarkMatter/claude-mods.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Terraform Operations
Terraform / OpenTofu infrastructure-as-code: layout, state, modules, safety, CI/CD, secrets.
Version context (verified 2026-06): Terraform 1.15.x (BUSL-1.1 licence since 1.6) · OpenTofu 1.12.x (MPL-2.0 fork of Terraform 1.5.x). Commands below are interchangeable (terraform ↔ tofu) unless flagged. See Terraform vs OpenTofu for the decision note.
Reference Files
| File | Covers |
|---|---|
| references/state-management.md | Remote backends, locking, moved/import/removed blocks, state surgery, drift detection |
| references/module-patterns.md | Module composition, variable validation, optional/nullable, output contracts, versioning |
| references/cicd-pipelines.md | GitHub Actions plan/apply, OIDC auth, policy gates (tflint/trivy/checkov/OPA), Atlantis/HCP |
| references/security-and-secrets.md | Secrets in state, ephemeral resources, write-only arguments, SOPS/Vault, sensitive limits |
| assets/github-actions-terraform.yml | Ready-to-adapt PR-plan + OIDC-apply workflow |
| scripts/check-action-refs.sh | Staleness verifier for any workflow's uses: action refs (offline structural / live API resolve) |
The action versions pinned in
github-actions-terraform.ymlare point-in-time (verified 2026-06). Runscripts/check-action-refs.sh --livebefore adopting — a tag that was valid at write time may have been retracted or never existed (e.g.trivy-action@0.33.1vs the realv0.33.1).
Project Layout Decision Tree
How many environments / accounts?
│
├─ One environment, one team
│ └─ Single root module + tfvars. Don't over-engineer.
│
├─ Multiple environments (dev/staging/prod)
│ ├─ Need different backend/account/region per env? (usually YES for prod isolation)
│ │ └─ DIRECTORY PER ENVIRONMENT (recommended default)
│ │ environments/{dev,staging,prod}/ each a thin root calling shared modules
│ │
│ └─ Environments truly identical except a few variables, same backend account?
│ └─ Workspaces are *acceptable* — but see the workspace caveats below
│
└─ Many teams / many state files / platform engineering
└─ Directory-per-env + per-component state split (network / data / app)
Consider Terragrunt, Terraform Stacks (HCP), or OpenTofu + CI orchestration
Canonical multi-env layout
infra/
├── modules/ # Reusable child modules (no provider/backend blocks)
│ ├── network/
│ │ ├── main.tf
│ │ ├── variables.tf
│ │ ├── outputs.tf
│ │ └── versions.tf # required_providers ONLY (no provider config)
│ └── app-service/
├── environments/ # Root modules — one state file each
│ ├── dev/
│ │ ├── main.tf # module "network" { source = "../../modules/network" ... }
│ │ ├── backend.tf # remote backend, env-specific key
│ │ ├── providers.tf # provider config lives in ROOT only
│ │ ├── terraform.tfvars # committed, non-secret env values
│ │ └── versions.tf # required_version + required_providers pins
│ └── prod/
└── .tflint.hcl
Why directories usually beat workspaces
| Concern | Directories | Workspaces |
|---|---|---|
| Separate backend/account per env | Yes — each root has its own backend.tf |
No — one backend, envs differ only by state key |
| Blast radius of wrong-env apply | Low — you're physically in prod/ |
High — invisible terraform workspace select state |
| Env-specific config divergence | Natural (different main.tf if needed) | terraform.workspace conditionals creep everywhere |
| Prod IAM isolation | Per-dir CI role | Same credentials see all envs |
| Visibility in code review | Diff shows which env changed | Workspace is runtime state, not in the diff |
Workspaces fit short-lived ephemeral copies (PR preview envs) — not the dev/prod boundary. HashiCorp's own docs say workspaces are "not suitable for strong separation."
tfvars conventions
terraform.tfvars # auto-loaded — per-root committed defaults (non-secret)
*.auto.tfvars # auto-loaded — generated/local overrides
prod.tfvars # explicit only: terraform plan -var-file=prod.tfvars
TF_VAR_db_password=... # env var injection — secrets in CI, never in files
Gotcha: -var-file + directories-per-env is belt-and-braces; with workspaces it's load-bearing and one forgotten flag applies dev values to prod.
State Quick Reference
Full detail: references/state-management.md.
| Task | Command / block | Notes |
|---|---|---|
| Remote backend (AWS) | backend "s3" { bucket, key, region, use_lockfile = true } |
S3-native locking (TF ≥1.10) — DynamoDB table no longer required |
| Rename resource in code | moved { from = aws_x.a, to = aws_x.b } |
Declarative, reviewable, no CLI surgery |
| Adopt existing infra | import { to = aws_x.a, id = "i-123" } + plan -generate-config-out=gen.tf |
Config-driven import (TF ≥1.5) beats terraform import CLI |
| Forget without destroy | removed { from = aws_x.a, lifecycle { destroy = false } } |
TF ≥1.7; OpenTofu 1.12 also has lifecycle { destroy = false } on resources |
| Drift detection | terraform plan -detailed-exitcode |
Exit 0 = clean, 1 = error, 2 = drift — cron it |
| Inspect state | terraform state list / state show ADDR |
Read-only, always safe |
| Move state (last resort) | terraform state mv SRC DST |
Prefer moved blocks — see "when NOT to" below |
| Pull/push (emergency) | terraform state pull > backup.tfstate |
ALWAYS pull a backup before any surgery |
State surgery — when NOT to: if a moved/removed/import block can express it, use the block. CLI state mv/rm is immediate, unreviewed, unversioned, and a typo orphans real infrastructure. Legit uses: splitting state between roots, unwedging a failed migration. Always state pull a backup first.
Module Quick Reference
Full detail: references/module-patterns.md.
module "network" {
source = "terraform-aws-modules/vpc/aws"
version = "~> 6.0" # pin minor-float for registry modules; exact pin in prod roots
# ...
}
| Rule | Why |
|---|---|
| Composition over inheritance | Roots compose flat modules; never module-wraps-module-wraps-module |
| No provider blocks in child modules | Providers configured in root only; child declares required_providers |
validation blocks on variables |
Fail at plan with a real message, not mid-apply |
optional(type, default) in object attrs |
Callers omit fields; nullable = false rejects explicit null |
| Outputs are the contract | Output IDs/ARNs consumers need; document with description |
| Anti-pattern: thin wrappers | A module that just renames variables of another module adds a version-lag layer and zero value — call the upstream module directly |
Safety Checklist (before every apply)
□ plan output READ, not skimmed — every destroy/replace explained
□ "Plan: X to add, Y to change, Z to destroy" — does Z surprise you?
□ -/+ (replace) lines: check the "forces replacement" attribute
□ Applying the SAME saved plan that was reviewed: plan -out=tfplan → apply tfplan
□ prevent_destroy on stateful resources (db, state bucket, KMS keys)
□ Cloud-side deletion protection too (RDS deletion_protection, S3 versioning+MFA-delete)
□ No -target unless this is a declared emergency (see below)
□ for_each (stable keys), not count, for any collection that can reorder
Footguns
| Footgun | Detail | Fix |
|---|---|---|
count index shift |
Removing item 0 of a count list re-addresses every later item → destroy/recreate cascade |
for_each with stable string keys |
-target habit |
Skips dependency graph; state diverges from config; hides drift | Emergency-only (broken dependency cycle, partial outage). Follow with a full clean plan |
prevent_destroy false comfort |
Doesn't survive the block being deleted, and doesn't stop state rm + console delete |
Pair with cloud-native deletion protection |
| Dynamic blocks everywhere | dynamic for 2 static blocks is obfuscation |
Use dynamic only over genuinely variable collections |
| Unpinned providers | aws = ">= 5.0" in prod pulls a breaking major the day it ships |
~> 6.12 + commit .terraform.lock.hcl |
| Apply ≠ reviewed plan | Plan on PR, apply on merge re-plans — drift in between applies unreviewed changes | Save the plan artifact, or accept + re-review the merge plan |
resource "aws_db_instance" "main" {
deletion_protection = true # cloud-side
lifecycle {
prevent_destroy = true # terraform-side
ignore_changes = [password] # if rotated outside TF
}
}
CI/CD Quick Reference
Full detail: references/cicd-pipelines.md · template: assets/github-actions-terraform.yml.
PR opened → fmt -check → validate → tflint → trivy/checkov → plan → plan posted as PR comment
PR merged → plan (fresh) → apply, authenticated via OIDC — no long-lived cloud keys
Nightly → plan -detailed-exitcode → exit 2 ⇒ drift alert
- OIDC everywhere —
aws-actions/configure-aws-credentialswithrole-to-assume, neverAWS_ACCESS_KEY_IDsecrets. Same supply-chain doctrine as this repo's rules: short-lived tokens, no standing credentials. - Pin action SHAs in workflows (
uses: actions/checkout@<sha>), not floating tags. - Policy gates:
tflint(provider-aware lint),trivy config/checkov(misconfig scan), OPA/conftestfor org policy ("no public buckets").
| Orchestrator | Fit |
|---|---|
| Plain GitHub Actions | Default — full control, free, template in assets/ |
| Atlantis | Self-hosted PR automation, atlantis plan/apply comments, locking per dir |
| HCP Terraform / Terraform Cloud | Managed runs, Sentinel policy, state hosting; free ≤500 resources |
| Spacelift / env0 / Digger / Scalr | Commercial Atlantis-likes; Digger runs inside your Actions |
Verification — uses: ref staleness
GitHub Action versions rot: a tag gets retracted, or a workflow pins one that never existed. scripts/check-action-refs.sh lints every uses: owner/repo@ref line. It's general — pass any workflow file(s) as positionals (default: this skill's own assets/github-actions-terraform.yml).
# Structural only, no network — well-formedness of every uses: ref (CI-safe gate).
# Floating @main/@master → WARN (exit 0; use --strict to fail). Malformed → exit 4.
scripts/check-action-refs.sh --offline .github/workflows/ci.yml
# Live — resolve each ref against the GitHub API. A 404 (ref doesn't exist) → exit 10
# DRIFT; API unreachable/rate-limited → exit 7 (advisory, never fails the build, §7).
# Set GITHUB_TOKEN to dodge the unauthenticated rate limit.
GITHUB_TOKEN=$GH_PAT scripts/check-action-refs.sh --live .github/workflows/*.yml
scripts/check-action-refs.sh --json --offline | jq '.data[] | select(.status!="ok")'
--live is the check that catches the classic aquasecurity/trivy-action@0.33.1 mistake — that tag 404s; the real one is v0.33.1. Run live on a schedule (never as a blocking PR gate), offline in PR CI.
Testing Quick Reference
# tests/network.tftest.hcl — native test framework (TF ≥1.6 / OpenTofu ≥1.6)
variables { cidr = "10.0.0.0/16" }
run "valid_cidr_plan" {
command = plan # plan = fast unit-ish; apply = real integration
assert {
condition = aws_vpc.main.cidr_block == "10.0.0.0/16"
error_message = "VPC CIDR did not match input"
}
}
run "rejects_tiny_cidr" {
command = plan
variables { cidr = "10.0.0.0/30" }
expect_failures = [var.cidr] # asserts the validation block fires
}
terraform test runs every *.tftest.hcl under tests/; command = apply runs create real (then auto-destroyed) infra — use a sandbox account. Mock providers (mock_provider blocks, TF ≥1.7) fake apply without credentials. For multi-tool/Go-level orchestration (retry, real HTTP probes), Terratest is the heavyweight alternative — native terraform test covers most module CI needs first.
Secrets Quick Reference
Full detail: references/security-and-secrets.md.
| Mechanism | Version | What it does |
|---|---|---|
sensitive = true |
all | Redacts from CLI output only — value still plaintext in state |
Ephemeral resources (ephemeral "...") |
TF ≥1.10 / OpenTofu ≥1.11 | Fetch secret at run time; never persisted to state or plan |
Write-only arguments (password_wo) |
TF ≥1.11 / OpenTofu ≥1.11 | Send secret to provider; never stored in state; rotate via _wo_version |
| SOPS-encrypted tfvars | tool | Secrets encrypted at rest in git; decrypted at plan time |
| Vault / cloud secret manager | tool | Reference by ID; resource reads secret at boot, TF never sees it |
| OpenTofu state encryption | OpenTofu ≥1.7 | Client-side AES-GCM encryption of state/plan — no Terraform equivalent |
Rule zero: treat state as secret regardless. Encrypt the backend (SSE-KMS), restrict IAM on the bucket, never commit *.tfstate (gitignore it).
Terraform vs OpenTofu
| Terraform | OpenTofu | |
|---|---|---|
| Licence | BUSL-1.1 since 1.6 (no production use competing with HashiCorp; fine for normal internal use) | MPL-2.0 — genuinely open source, Linux Foundation |
| Current | 1.15.x | 1.12.x |
| Exclusive features | Stacks (HCP-tied), Terraform Cloud agents, terraform query |
State/plan encryption, provider for_each iteration, -exclude flag, early variable eval in backend/module blocks, OCI registry distribution, .tofu file extension |
| Registry | registry.terraform.io | registry.opentofu.org (mirrors most providers) |
| Compatibility | — | Forked at 1.5.x; HCL/state compatible for mainstream use, diverging feature-by-feature since |
Decision: vendors and anyone redistributing IaC tooling commercially → OpenTofu (licence risk). Teams on HCP Terraform/Sentinel → Terraform. Everyone else: either works; OpenTofu's state encryption is the single biggest technical differentiator. Migration terraform → tofu is tofu init + state-compatible up to ~1.8-era features; the gap widens each release — migrate early or commit.
Command Quick Reference
terraform init -upgrade # init / upgrade providers within constraints
terraform fmt -recursive -check # CI: fail on unformatted
terraform validate # syntax + internal consistency (no creds needed after init)
terraform plan -out=tfplan # save plan for exact-apply
terraform show -json tfplan | jq # machine-readable plan (policy tools eat this)
terraform apply tfplan # apply EXACTLY the reviewed plan
terraform plan -detailed-exitcode # 0 clean / 2 drift — for cron drift checks
terraform plan -refresh-only # show drift without proposing config changes
terraform apply -replace=aws_x.a # force recreate one resource (replaces old taint)
terraform state pull > backup.json # ALWAYS before surgery
terraform output -json # consume outputs in scripts
terraform graph | dot -Tsvg > g.svg # dependency graph
tofu init # OpenTofu: same verbs throughout
Files (claude-mods)
-
assets
-
github-actions-terraform.yml 8.7 KB
# ============================================================================= # Terraform CI/CD — PR plan + OIDC apply (GitHub Actions template) # ============================================================================= # What this gives you: # * PRs: fmt-check -> validate -> tflint -> trivy -> plan -> plan posted as # a PR comment (updated in place, not stacked) # * Merge: fresh plan + apply on main, gated by a GitHub "environment" # (add required reviewers on the environment for a manual approval) # * Auth: AWS via OIDC — NO long-lived access keys anywhere. # # Adapt before use: # 1. Set WORKING_DIR to your root module (e.g. environments/prod). # 2. Replace the two role ARNs (plan = read-only, apply = write). # 3. Pin every `uses:` to a commit SHA for your security posture — # tags are shown here for readability; SHA-pin in production: # uses: actions/checkout@08c6903cd8c0fde910a37f88322edcfb5dd907a8 # v5.0.0 # 4. OpenTofu: swap hashicorp/setup-terraform for opentofu/setup-opentofu # (input `tofu_version`) and `terraform` for `tofu` in run steps. # 5. Multi-env monorepo: turn `WORKING_DIR` into a job matrix. # ============================================================================= name: terraform on: pull_request: branches: [main] paths: ["environments/prod/**", "modules/**"] # plan only when infra changes push: branches: [main] paths: ["environments/prod/**", "modules/**"] # Least privilege by default; jobs widen only what they need. permissions: contents: read env: TF_VERSION: "1.15.5" # match required_version in versions.tf WORKING_DIR: environments/prod AWS_REGION: ap-southeast-2 TF_IN_AUTOMATION: "true" # suppresses interactive-use hints in output # One run per state file at a time. Plans queue; applies never overlap. concurrency: group: terraform-prod cancel-in-progress: false # NEVER cancel a running apply jobs: # --------------------------------------------------------------------------- # PR: static checks + plan + comment # --------------------------------------------------------------------------- plan: if: github.event_name == 'pull_request' runs-on: ubuntu-latest permissions: contents: read id-token: write # OIDC token for AWS pull-requests: write # post the plan comment defaults: run: working-directory: ${{ env.WORKING_DIR }} steps: - uses: actions/checkout@v5 - uses: hashicorp/setup-terraform@v3 with: terraform_version: ${{ env.TF_VERSION }} # --- cheap static gates first: fail fast before touching the cloud --- - name: fmt run: terraform fmt -check -recursive -diff - name: tflint uses: terraform-linters/setup-tflint@v6 - run: tflint --init && tflint --recursive working-directory: ${{ env.WORKING_DIR }} - name: trivy misconfig scan uses: aquasecurity/trivy-action@v0.33.1 with: scan-type: config scan-ref: ${{ env.WORKING_DIR }} exit-code: "1" # hard gate; set "0" while triaging a baseline severity: HIGH,CRITICAL # --- read-only cloud credentials, scoped to PR plans --- # This role's trust policy allows any branch of THIS repo, but the role # itself carries ReadOnlyAccess + state-bucket read. PR code cannot mutate. - name: Configure AWS credentials (plan role, read-only) uses: aws-actions/configure-aws-credentials@v5 with: role-to-assume: arn:aws:iam::123456789012:role/github-terraform-plan aws-region: ${{ env.AWS_REGION }} - name: init run: terraform init -input=false -lock-timeout=2m - name: validate run: terraform validate -no-color - name: plan id: plan # -lock=false: a PR plan must never block (or be blocked by) an apply run: | set -o pipefail terraform plan -input=false -no-color -lock=false -out=tfplan 2>&1 \ | tee plan.txt continue-on-error: true # we still want the comment when plan fails # Optional org-policy gate against the machine-readable plan: # - run: terraform show -json tfplan > tfplan.json # - run: conftest test --policy ../../policy tfplan.json - name: Comment plan on PR uses: actions/github-script@v8 env: PLAN_OUTCOME: ${{ steps.plan.outcome }} with: script: | const fs = require('fs'); const dir = process.env.WORKING_DIR; const marker = `### Terraform plan — \`${dir}\``; let plan = fs.readFileSync(`${dir}/plan.txt`, 'utf8'); // GitHub comment hard cap is 65,536 chars — keep the tail (the summary) if (plan.length > 60000) plan = '... (truncated, see job log)\n' + plan.slice(-60000); const status = process.env.PLAN_OUTCOME === 'success' ? '' : '\n> **PLAN FAILED** — see details below.'; const body = `${marker}${status}\n<details><summary>Show plan</summary>\n\n\`\`\`hcl\n${plan}\n\`\`\`\n</details>`; const { data: comments } = await github.rest.issues.listComments({ ...context.repo, issue_number: context.issue.number }); const prev = comments.find(c => c.body.startsWith(marker)); if (prev) { await github.rest.issues.updateComment({ ...context.repo, comment_id: prev.id, body }); } else { await github.rest.issues.createComment({ ...context.repo, issue_number: context.issue.number, body }); } - name: Fail job if plan failed if: steps.plan.outcome == 'failure' run: exit 1 # --------------------------------------------------------------------------- # Merge to main: fresh plan + apply # --------------------------------------------------------------------------- apply: if: github.event_name == 'push' && github.ref == 'refs/heads/main' runs-on: ubuntu-latest # The "production" environment is where the human gate lives: # repo Settings -> Environments -> production -> required reviewers. environment: production permissions: contents: read id-token: write defaults: run: working-directory: ${{ env.WORKING_DIR }} steps: - uses: actions/checkout@v5 - uses: hashicorp/setup-terraform@v3 with: terraform_version: ${{ env.TF_VERSION }} # Apply role: write permissions, trust policy locked to # repo:myorg/infra:environment:production # so ONLY this gated environment on main can assume it. - name: Configure AWS credentials (apply role) uses: aws-actions/configure-aws-credentials@v5 with: role-to-assume: arn:aws:iam::123456789012:role/github-terraform-apply aws-region: ${{ env.AWS_REGION }} - name: init run: terraform init -input=false -lock-timeout=5m # Fresh plan at merge time. If the world moved since PR review, this plan # differs from the reviewed one — the saved-plan apply below makes that # explicit rather than silently applying something new. - name: plan run: terraform plan -input=false -no-color -lock-timeout=5m -out=tfplan - name: apply # Applying the saved plan (not `apply -auto-approve` on config) means # we apply EXACTLY what the step above planned — no second refresh gap. run: terraform apply -input=false -no-color tfplan # --------------------------------------------------------------------------- # Nightly drift detection # --------------------------------------------------------------------------- # Move to its own file with `on: schedule` if you prefer; shown here for # completeness. Exit code 2 from -detailed-exitcode means drift. # drift: # if: github.event_name == 'schedule' # runs-on: ubuntu-latest # permissions: { contents: read, id-token: write } # steps: # - uses: actions/checkout@v5 # - uses: hashicorp/setup-terraform@v3 # with: { terraform_version: "1.15.5" } # - uses: aws-actions/configure-aws-credentials@v5 # with: # role-to-assume: arn:aws:iam::123456789012:role/github-terraform-plan # aws-region: ap-southeast-2 # - run: terraform init -input=false # working-directory: environments/prod # - name: drift check # working-directory: environments/prod # run: | # set +e # terraform plan -detailed-exitcode -lock=false -input=false -no-color # code=$? # [ "$code" -eq 2 ] && { echo "::error::Drift detected in prod"; exit 1; } # exit $code
-
-
references
-
cicd-pipelines.md 9.5 KB
# CI/CD Pipelines The contract: **every change is planned on the PR, the plan is visible to the reviewer, and apply happens from CI with short-lived credentials.** Humans never run `apply` against shared environments from laptops. Full workflow template: [../assets/github-actions-terraform.yml](../assets/github-actions-terraform.yml). ## Pipeline Shape ``` ┌─ fmt -check ─┐ PR opened ──> ├─ validate ├──> tflint ──> trivy/checkov ──> plan ──> plan as PR comment └─ (parallel) ┘ │ reviewer reads plan PR merged ──> fresh plan ──> apply (OIDC role, environment gate) Nightly ──> plan -detailed-exitcode ──> exit 2 ⇒ drift alert ``` Key decisions baked into that shape: 1. **Plan on PR, apply on merge.** The merge re-plans rather than applying the stale PR plan artifact — simpler, and the `concurrency` group serializes applies. If you need apply-exactly-what-was-reviewed, upload `tfplan` as an artifact on the PR and apply that artifact on merge; accept the trade-off that the world may have moved (the apply will fail if so, which is the safe failure). 2. **One job per root module** (matrix or separate workflows). A monorepo with `environments/{dev,prod}` plans both on PR, applies dev on merge, applies prod behind a GitHub *environment* with required reviewers. 3. **`concurrency` group per state file** so two merges can't apply concurrently (backend locking would catch it, but failing fast in CI is cleaner). ## OIDC Cloud Auth — no long-lived keys This is non-negotiable and matches the repo's supply-chain doctrine (short-lived tokens over standing credentials; a leaked workflow can't exfiltrate what doesn't exist). GitHub mints a signed JWT per job; the cloud trusts GitHub's issuer for *specific repos/branches* and returns temporary credentials. ### AWS ```yaml permissions: id-token: write # REQUIRED for OIDC contents: read steps: - uses: aws-actions/configure-aws-credentials@v5 with: role-to-assume: arn:aws:iam::123456789012:role/github-terraform-plan aws-region: ap-southeast-2 ``` Trust policy on the role — scope it tight: ```json { "Effect": "Allow", "Principal": { "Federated": "arn:aws:iam::123456789012:oidc-provider/token.actions.githubusercontent.com" }, "Action": "sts:AssumeRoleWithWebIdentity", "Condition": { "StringEquals": { "token.actions.githubusercontent.com:aud": "sts.amazonaws.com" }, "StringLike": { "token.actions.githubusercontent.com:sub": "repo:myorg/infra:ref:refs/heads/main" } } } ``` - **Two roles**: a read-only `plan` role (assumable from any branch/PR) and a write `apply` role (assumable only from `ref:refs/heads/main` or `environment:prod`). PR plans from forks then physically cannot mutate anything. - Audit the trust federation periodically — stale OIDC subjects (deleted repos, renamed branches) with live trust are exactly the entry point the 2026 supply-chain worms abused. `zizmor` catches `pull_request_target` + OIDC misconfigs statically. ### Azure / GCP equivalents ```yaml # Azure — federated credential on an app registration - uses: azure/login@v2 with: client-id: ${{ vars.AZURE_CLIENT_ID }} tenant-id: ${{ vars.AZURE_TENANT_ID }} subscription-id: ${{ vars.AZURE_SUBSCRIPTION_ID }} # GCP — workload identity federation - uses: google-github-actions/auth@v3 with: workload_identity_provider: projects/123/locations/global/workloadIdentityPools/github/providers/myorg service_account: terraform@myproj.iam.gserviceaccount.com ``` ## Plan as PR Comment The reviewer must see the plan without leaving the PR. Minimal recipe (full version in the asset): ```yaml - name: Plan id: plan run: terraform plan -no-color -input=false -out=tfplan 2>&1 | tee plan.txt - name: Comment plan on PR uses: actions/github-script@v8 with: script: | const fs = require('fs'); const plan = fs.readFileSync('plan.txt', 'utf8').slice(0, 60000); // comment size cap const body = `### Terraform plan — \`prod\`\n<details><summary>Show plan</summary>\n\n\`\`\`hcl\n${plan}\n\`\`\`\n</details>`; // find-and-update existing comment instead of stacking new ones const { data: comments } = await github.rest.issues.listComments({ ...context.repo, issue_number: context.issue.number }); const prev = comments.find(c => c.body.startsWith('### Terraform plan — `prod`')); if (prev) await github.rest.issues.updateComment({ ...context.repo, comment_id: prev.id, body }); else await github.rest.issues.createComment({ ...context.repo, issue_number: context.issue.number, body }); ``` Notes: - **Update-in-place** (find previous comment) or every push spams the PR. - Truncate: GitHub comments cap at 65,536 chars. For huge plans link to the job log and post only the resource-change summary (`terraform show -json tfplan | jq -r '.resource_changes[] | "\(.change.actions | join(",")) \(.address)"'`). - Plans can leak values — `sensitive = true` redacts in plan output, but data sources and resource attributes are not all marked. Treat the PR comment as visible to everyone with repo read. ## Policy Gates | Tool | Layer | What it catches | Invocation | |---|---|---|---| | `terraform fmt -check -recursive` | style | Unformatted code | exit ≠ 0 fails CI | | `terraform validate` | syntax | Type errors, bad references | needs `init` (use `-backend=false` for speed) | | `tflint` | lint | Provider-aware errors: invalid instance types, deprecated syntax, unused declarations | `tflint --init && tflint --recursive` | | `trivy config .` | security | Misconfig: public buckets, open SGs, unencrypted disks (absorbed tfsec's rule set) | exit codes; SARIF upload for code-scanning UI | | `checkov -d .` | security | Same space as trivy; bigger policy library, more noise — pick ONE of trivy/checkov | `--soft-fail` while triaging | | `conftest test tfplan.json` | org policy | YOUR rules in Rego/OPA: "no resources without tags", "only approved regions", "no IAM * actions" | run against `terraform show -json tfplan` | ```hcl # .tflint.hcl plugin "aws" { enabled = true version = "0.40.0" source = "github.com/terraform-linters/tflint-ruleset-aws" } rule "terraform_required_version" { enabled = true } rule "terraform_naming_convention" { enabled = true } ``` ```rego # policy/tags.rego — conftest example against plan JSON package main deny[msg] { rc := input.resource_changes[_] rc.change.actions[_] == "create" not rc.change.after.tags.Environment msg := sprintf("%s: missing required tag 'Environment'", [rc.address]) } ``` Layering guidance: fmt/validate/tflint are table stakes on every PR. trivy *or* checkov as the misconfig gate (start `--soft-fail`, ratchet to hard once the baseline is clean). conftest/OPA only when you have genuinely org-specific rules the scanners can't express — it's the highest-maintenance layer. ## Workflow Hardening Same doctrine as the rest of the repo's supply-chain rules: - **Pin actions to commit SHAs** — `uses: actions/checkout@08c6903cd8c0fde910a37f88322edcfb5dd907a8 # v5.0.0`, not `@v5`. A hijacked tag is a hijacked pipeline with cloud credentials. - `permissions:` block at workflow top, least privilege (`contents: read` default; `id-token: write` only on jobs that auth to cloud; `pull-requests: write` only on the comment job). - **Never `pull_request_target` with checkout of PR head** for terraform repos — that's RCE-with-secrets for any fork. - Plan jobs from forks: read-only role or no cloud auth at all (validate/lint only). - `step-security/harden-runner` for egress allow-listing on apply jobs if you want runtime control. - Pin the terraform version: `hashicorp/setup-terraform@<sha>` with `terraform_version: 1.15.5` (or `opentofu/setup-opentofu` with `tofu_version`), matching `required_version`. ## Drift Detection Job ```yaml on: schedule: [{ cron: "17 18 * * *" }] # nightly; odd minute avoids the top-of-hour stampede jobs: drift: permissions: { id-token: write, contents: read } steps: # ... checkout, setup, OIDC auth (read-only role), init ... - name: Detect drift run: | set +e terraform plan -detailed-exitcode -lock=false -input=false -no-color code=$? if [ "$code" -eq 2 ]; then echo "::error::Drift detected"; exit 1; fi exit $code ``` Wire the failure to Slack/issue creation. Exit 2 means *either* console drift or merged-but-unapplied config — both are findings. ## Orchestrator Alternatives | | Model | Locking | Policy | Cost | Pick when | |---|---|---|---|---|---| | **GitHub Actions (DIY)** | Workflows you own | `concurrency` groups | tflint/trivy/conftest steps | Free-ish | Default. Full control, template in assets/ | | **Atlantis** | Self-hosted server; `atlantis plan` / `atlantis apply` PR comments | Per-directory PR locks (best-in-class) | Custom workflows + conftest | Server you run | Many roots, many PRs, comment-driven culture | | **HCP Terraform** | Managed runs + state + UI | Workspace runs serialize | Sentinel / OPA | Free ≤ 500 resources, then $$ | Want managed everything; Sentinel policy; private registry | | **Spacelift / env0 / Scalr** | Commercial SaaS orchestrators | Built-in | OPA-based | $$ | Enterprise multi-IaC (Pulumi/CFN too), RBAC needs | | **Digger** | Runs inside *your* GitHub Actions | Orchestrated via PR | Pluggable | OSS core | Atlantis UX without hosting a server | OpenTofu note: Atlantis, Spacelift, env0, Digger all support `tofu` as the binary. HCP Terraform does not run OpenTofu. -
module-patterns.md 9.3 KB
# Module Patterns Modules are Terraform's unit of reuse. Most module pain comes from treating them like classes (inheritance, deep nesting, wrapping) instead of functions (flat composition, explicit inputs/outputs). ## Root vs Child Modules | | Root module | Child module | |---|---|---| | What | The directory you run `terraform` in | Anything called via `module` block | | Backend block | Yes — exactly one | **Never** | | Provider config (`provider "aws" {}`) | Yes | **Never** — declare `required_providers` only | | tfvars | Yes | No (inputs come from the caller) | | State | Owns one state file | Lives inside the caller's state | ```hcl # modules/network/versions.tf — child module declares NEEDS, not config terraform { required_version = ">= 1.10" required_providers { aws = { source = "hashicorp/aws" version = ">= 6.0, < 8.0" # modules use RANGES; roots pin tighter } } } ``` A child module with a `provider` block can't be used with `for_each`/`count`/`depends_on` and can't be removed cleanly. Pass aliased providers explicitly when needed: ```hcl module "dns" { source = "../../modules/dns" providers = { aws = aws.us_east_1 } # ACM certs for CloudFront, etc. } ``` ## Composition Over Inheritance Build small single-purpose modules; compose them in the root by wiring outputs to inputs. ```hcl # Root composes — dependencies are explicit data flow module "network" { source = "../../modules/network" cidr = "10.0.0.0/16" } module "database" { source = "../../modules/database" subnet_ids = module.network.private_subnet_ids # output -> input wiring vpc_id = module.network.vpc_id } module "app" { source = "../../modules/app-service" subnet_ids = module.network.private_subnet_ids db_connection_arn = module.database.connection_secret_arn } ``` Rules of thumb: - **Nesting depth ≤ 2.** Root → module → (occasionally) submodule. Deeper means you're rebuilding inheritance and debugging through four layers of variable plumbing. - A module should manage a *cohesive* set of resources with a clear lifecycle (a VPC and its subnets/routes — yes; "everything for the app" — no). - If a module takes 40 variables and most callers set 3, split it. - Don't create a module for a single resource unless it encodes real policy (e.g. an S3 module that enforces encryption + public-access-block on every bucket — that's policy, not wrapping). ### Anti-pattern: thin wrapper modules ```hcl # modules/our-vpc/main.tf — adds NOTHING module "vpc" { source = "terraform-aws-modules/vpc/aws" version = "~> 6.0" name = var.vpc_name # renamed from `name`. why. cidr = var.network_cidr # renamed from `cidr`. why. } ``` Costs: a second version to bump (wrapper lags upstream), a second docs surface, every upstream feature needs a wrapper variable added before anyone can use it, and `moved`-block refactors in upstream don't propagate. **Call the upstream module directly from the root.** A wrapper earns its existence only when it enforces organizational policy (mandatory tags, forced encryption, restricted instance types) — and then it should *say so* in its README. ## Variable Design ### Validation — fail at plan, not mid-apply ```hcl variable "environment" { type = string description = "Deployment environment." validation { condition = contains(["dev", "staging", "prod"], var.environment) error_message = "environment must be one of: dev, staging, prod." } } variable "cidr" { type = string validation { condition = can(cidrhost(var.cidr, 0)) && tonumber(split("/", var.cidr)[1]) <= 24 error_message = "cidr must be a valid IPv4 CIDR no smaller than /24." } } ``` Since TF 1.9, `condition` can reference *other* variables and data — cross-field validation lives on the variable, not buried in a `precondition`. ### optional() + nullable ```hcl variable "logging" { description = "Access-logging config. Omit fields for defaults." type = object({ enabled = optional(bool, true) bucket = optional(string) # null when omitted prefix = optional(string, "logs/") sample_rate = optional(number, 1.0) }) default = {} nullable = false # caller may omit the variable, but may NOT pass logging = null } ``` - `optional(type, default)` lets callers pass partial objects — the killer feature for config-object variables. Without defaults you'd force every caller to spell out every field. - `nullable = false` means an explicit `null` is rejected and the variable's own `default` is used instead. Use it on almost every variable: it converts "caller passed null, module exploded on `var.x.enabled`" into a plan-time error. - Gotcha: `optional(string)` with no default yields `null` — guard with `coalesce(...)` or `try(...)` before interpolating. ### Variable hygiene | Rule | Why | |---|---| | Every variable has `description` | It's the module's API doc (`terraform-docs` renders it) | | `sensitive = true` on secret inputs | Keeps values out of CLI output (NOT out of state — see security-and-secrets.md) | | Prefer typed objects over `map(any)` | `any` defers errors to deep inside the module | | No "pass-through everything" variables | A `extra_settings = any` variable is an API you can never change | | Defaults = safe choice, not common choice | Default to encrypted/private/protected; make callers opt *out* loudly | ## Output Contracts Outputs are the module's public API. Consumers wire them into other modules and remote-state reads — changing one is a breaking change. ```hcl output "vpc_id" { description = "ID of the created VPC." value = aws_vpc.main.id } output "private_subnet_ids" { description = "Private subnet IDs, keyed by AZ." value = { for k, s in aws_subnet.private : k => s.id } } output "db_endpoint" { description = "Writer endpoint. SENSITIVE: contains hostname only, no creds." value = aws_rds_cluster.main.endpoint } ``` - Output the **identifiers consumers need** (IDs, ARNs, endpoints, security-group IDs) — not whole resource objects (`value = aws_vpc.main` couples consumers to the provider schema and bloats state). - `sensitive = true` propagates: an output derived from a sensitive value must itself be marked sensitive or plan errors. - Treat output removal/rename like an API break: semver-major the module, or keep the old output as an alias for one cycle. ## Versioning and Pinning ```hcl # Registry module — minor-float module "vpc" { source = "terraform-aws-modules/vpc/aws" version = "~> 6.0" # >= 6.0.0, < 7.0.0 } # Git module — pin a TAG (never a branch) module "internal" { source = "git::https://github.com/myorg/tf-modules.git//network?ref=v2.3.1" } ``` | Constraint | Meaning | Use for | |---|---|---| | `~> 6.0` | ≥ 6.0, < 7.0 | Registry modules/providers in shared modules | | `~> 6.12.0` | ≥ 6.12.0, < 6.13.0 | Conservative prod roots | | `= 6.12.1` / `?ref=v2.3.1` | Exact | Prod roots wanting byte-identical builds | | `>= 6.0` (open-ended) | Anything newer | **Never in prod** — a breaking major auto-arrives | - **Commit `.terraform.lock.hcl`.** It pins exact provider versions + hashes; `terraform init -upgrade` is the deliberate act of moving within constraints. This is the same supply-chain posture as any other lockfile: a pin only protects you if it pre-dates a compromise and you don't run unconstrained upgrades in CI. - Run `terraform providers lock -platform=linux_amd64 -platform=darwin_arm64 -platform=windows_amd64` so the lockfile carries hashes for every platform your team + CI uses (OpenTofu 1.12 does this automatically at init). - Internal module registries: HCP Terraform private registry, or plain git tags + a `modules/` monorepo. Git tags are fine; the registry's value is the version-constraint syntax and docs rendering. ## Module Repo Layout ``` terraform-aws-network/ # one module per repo (registry-publishable), or modules/ monorepo ├── main.tf ├── variables.tf ├── outputs.tf ├── versions.tf ├── README.md # terraform-docs generated ├── examples/ │ └── complete/ # a runnable root that exercises the module — doubles as docs + test fixture │ └── main.tf └── tests/ └── network.tftest.hcl # native tests (see SKILL.md Testing section) ``` `terraform-docs markdown table . > README.md` keeps docs honest — wire it as a pre-commit hook or CI check. ## count vs for_each (the index-shift footgun) ```hcl # BAD: count over a list resource "aws_subnet" "private" { count = length(var.subnet_cidrs) # remove element 0 -> cidr_block = var.subnet_cidrs[count.index] # every subnet re-addresses -> destroy cascade } # GOOD: for_each over a map with stable keys resource "aws_subnet" "private" { for_each = var.subnets # { "ap-southeast-2a" = "10.0.1.0/24", ... } cidr_block = each.value availability_zone = each.key } ``` `count` is fine for "0 or 1 of this" conditionals (`count = var.enabled ? 1 : 0`) — though even there, `for_each = var.enabled ? { main = true } : {}` keeps addresses stable if it might ever become "n of this". Migrating existing `count` resources to `for_each`: write `moved` blocks for each index→key pair (see state-management.md) so nothing is destroyed. -
security-and-secrets.md 8.3 KB
# Security and Secrets The uncomfortable truth first: **Terraform state stores resource attributes in plaintext JSON** — including any password, key, token, or connection string a provider ever returned. `sensitive = true` changes what's *printed*, not what's *stored*. Every secrets strategy below is a variation on "make sure the secret never enters state, and harden state anyway." ## Threat Model | Surface | Exposure | Mitigation | |---|---|---| | State file | Plaintext attributes of every resource | Encrypted backend + IAM; OpenTofu state encryption; keep secrets out entirely (below) | | Plan files (`tfplan`, JSON) | Contain proposed values, incl. some sensitive ones | Treat plan artifacts as secrets; short retention on CI artifacts | | CLI / CI logs | Values interpolated into output | `sensitive = true`, `-no-color` log review, masked CI vars | | PR plan comments | Anyone with repo read | Summarize rather than full-dump for sensitive roots | | `.tfvars` in git | Whatever you put there | Never commit secret tfvars; SOPS-encrypt or env-var inject | | Provider credentials | Long-lived keys in CI secrets | OIDC short-lived tokens (see cicd-pipelines.md) | ## What `sensitive = true` Actually Does (and doesn't) ```hcl variable "db_password" { type = string sensitive = true # plan/apply output prints (sensitive value) } output "endpoint" { value = "${aws_db_instance.main.address}:${var.db_password}" # ERROR unless... sensitive = true # ...the output is marked too (sensitivity propagates) } ``` Does: redact from `plan`/`apply`/`output` human output; propagate taint through expressions; force derived outputs to be marked. Does **not**: encrypt anything; remove the value from state (`terraform state pull | jq` shows it plaintext); redact from `terraform output -json` (explicitly prints sensitive values); stop a provider logging it at TRACE. `ephemeral = true` on variables (TF ≥ 1.10) goes further: the value may only flow into ephemeral contexts (write-only args, provider config, locals marked ephemeral) and is never written to state or plan. ## Ephemeral Resources (TF ≥ 1.10, OpenTofu ≥ 1.11) Ephemeral resources open/fetch a value during the run and are **never persisted to state or plan**. The first-class pattern for "read a secret from a manager at apply time." ```hcl # Fetch the secret ephemerally — exists only for the duration of the run ephemeral "aws_secretsmanager_secret_version" "db" { secret_id = aws_secretsmanager_secret.db.id } # Use it via a write-only argument so it never lands in state either resource "aws_db_instance" "main" { # ... password_wo = ephemeral.aws_secretsmanager_secret_version.db.secret_string password_wo_version = 1 } ``` Available ephemeral resource types include `aws_secretsmanager_secret_version`, `aws_ssm_parameter`, `azurerm_key_vault_secret`, `google_secret_manager_secret_version`, Vault's `vault_kv_secret_v2`, and `random_password` (ephemeral variant). Ephemeral *values* can feed: write-only arguments, provider configuration, `terraform_data` triggers — anything that itself persists will error if you try. Contrast with the classic `data "aws_secretsmanager_secret_version"` — a data source's result **is stored in state**, which quietly copied your secret into the state file. Migrate those reads to `ephemeral` blocks. ## Write-Only Arguments (TF ≥ 1.11, OpenTofu ≥ 1.11) Provider-defined `*_wo` arguments accept a value, hand it to the API, and store **nothing** in state. Because nothing is stored, Terraform can't diff them — that's what the paired `*_wo_version` integer is for: ```hcl resource "aws_db_instance" "main" { password_wo = var.db_password # never in state password_wo_version = 2 # bump to push a new password } ``` - Rotate by incrementing `_wo_version` (or wiring it to a rotation timestamp/secret version). - Write-only args accept ephemeral values — the combination above (ephemeral read → write-only write) is the zero-secrets-in-state gold standard. - Caveat: not every provider/resource has `_wo` variants yet; check the provider docs. Where absent, fall back to the secret-manager-reference pattern below. ## Pattern Ladder (best to worst) ``` 1. Secret never touches Terraform at all App reads from Secrets Manager/Vault at BOOT using its IAM role; Terraform only creates the (empty or rotated-out-of-band) secret container + IAM. 2. Ephemeral read -> write-only write (TF >= 1.11) Secret transits the run in memory only. Nothing in state or plan. 3. Secret manager reference via data source (any version) data "aws_secretsmanager_secret_version" -- secret IS in state, but at least it's centrally rotated + audited. Encrypt state, restrict IAM. 4. TF_VAR_ env injection from CI secret store (any version) Keeps secrets out of git; still lands in state if assigned to a resource attribute. 5. SOPS-encrypted tfvars in git Good at-rest story for git; same state caveat as 4. 6. Plaintext in tfvars/locals committed to git Never. Rotating means rewriting git history. ``` ### Vault / cloud secret manager integration ```hcl # Terraform creates the container + access policy; VALUE is set out-of-band or by rotation lambda resource "aws_secretsmanager_secret" "db" { name = "prod/db/password" kms_key_id = aws_kms_key.secrets.arn } resource "aws_iam_role_policy" "app_reads_secret" { role = aws_iam_role.app.id policy = jsonencode({ Version = "2012-10-17" Statement = [{ Effect = "Allow", Action = "secretsmanager:GetSecretValue", Resource = aws_secretsmanager_secret.db.arn }] }) } # App fetches the secret at startup. Terraform never knows the value. ``` Vault: prefer short-TTL dynamic secrets (`vault_database_secret_backend_role`) so even a leaked credential expires in minutes. The Vault provider's classic data sources persist to state — use the ephemeral variants on TF ≥ 1.10. ### SOPS pattern ```bash # Encrypt env tfvars with a KMS key; ciphertext is committable sops --encrypt --kms arn:aws:kms:...:key/... prod.tfvars.json > prod.sops.tfvars.json ``` ```hcl # Via carlpett/sops provider data "sops_file" "secrets" { source_file = "prod.sops.tfvars.json" } locals { db_password = data.sops_file.secrets.data["db_password"] } ``` Honest accounting: decrypted values still flow into plan/state unless they terminate in write-only arguments. SOPS solves *git at-rest*, not *state at-rest*. ## Hardening State Itself Do all of this regardless of which pattern above you use: - Backend encryption: S3 SSE-KMS with a dedicated CMK; bucket policy denying un-encrypted puts; `azurerm`/`gcs` with CMK. - IAM: state bucket readable only by the CI roles + break-glass group. State read access ≈ secret read access — treat the grant accordingly. - Versioning on (recovery) + access logging on the bucket (audit). - `.gitignore`: `*.tfstate`, `*.tfstate.*`, `*.tfplan`, `.terraform/`, and crash logs (`crash.log` can embed values). - **OpenTofu state encryption** (≥ 1.7) — client-side AES-GCM over state *and* plan files, key from PBKDF2 passphrase, AWS/GCP KMS, Azure Key Vault, or external program. The strongest state story available; Terraform has no equivalent (config sample in state-management.md). Plan key rotation: `encryption` supports a `fallback` method so old state remains readable during rotation. ## Provider Credential Hygiene | Don't | Do | |---|---| | `provider "aws" { access_key = "..." }` hardcoded | Ambient auth: OIDC in CI, SSO/instance profiles locally | | Long-lived `AWS_ACCESS_KEY_ID` in CI secrets | OIDC `role-to-assume` (see cicd-pipelines.md) | | One god-role for plan and apply | Read-only plan role; write apply role gated to main/environment | | Shared human credentials for break-glass | Named identities + audited assume-role | ## Scanning and Gates - `trivy config .` / `checkov -d .` catch *misconfigurations* (public buckets, `0.0.0.0/0` ingress, unencrypted volumes) — wire into PR CI (see cicd-pipelines.md). - `gitleaks` / push-gates catch secrets *in the repo* — tfvars are a classic leak vector. - `terraform providers` + lockfile review on provider bumps: providers execute arbitrary code on your CI runner with cloud credentials. A provider is a dependency — the repo's supply-chain rules (cooldown, behavioural scan before adopting unfamiliar providers from the registry) apply in full. -
state-management.md 9.3 KB
# State Management Terraform/OpenTofu state is the mapping between config addresses and real infrastructure. It is the single most dangerous file in the project: lose it and Terraform forgets your infra; corrupt it and applies destroy the wrong things; leak it and every secret a provider ever returned is exposed. ## Remote Backends Never keep state local for anything shared or production. Remote backends give durability, locking, and team access. ### S3 (AWS) — current recommended shape ```hcl terraform { backend "s3" { bucket = "myorg-tfstate" key = "prod/network/terraform.tfstate" # one key per root module region = "ap-southeast-2" encrypt = true # SSE; pair with bucket-default SSE-KMS use_lockfile = true # S3-native locking (TF >= 1.10) } } ``` - **`use_lockfile = true` replaces the DynamoDB lock table** (Terraform ≥ 1.10, OpenTofu ≥ 1.10). It uses S3 conditional writes to create a `.tflock` object next to the state key. The old `dynamodb_table` argument still works and was the standard for a decade — you'll see it everywhere — but new setups don't need the extra table. During migration you can set both; remove `dynamodb_table` once all collaborators are ≥ 1.10. - Bucket hygiene: versioning **on** (state history = your undo), default SSE-KMS, block public access, lifecycle rule to expire old noncurrent versions after ~90 days, bucket policy restricting to the CI role + break-glass humans. - One bucket per org/account is fine; isolation comes from `key` prefixes + IAM conditions on the prefix. ### azurerm ```hcl terraform { backend "azurerm" { resource_group_name = "rg-tfstate" storage_account_name = "myorgtfstate" container_name = "tfstate" key = "prod.network.tfstate" use_azuread_auth = true # RBAC instead of storage keys } } ``` Locking is native via blob leases — nothing extra to configure. Prefer `use_azuread_auth = true` so CI uses OIDC-federated identity, not storage account keys. ### gcs ```hcl terraform { backend "gcs" { bucket = "myorg-tfstate" prefix = "prod/network" } } ``` Locking native via Cloud Storage generation preconditions. Enable object versioning on the bucket. Use workload identity federation in CI. ### HCP Terraform / Terraform Cloud ```hcl terraform { cloud { organization = "myorg" workspaces { name = "prod-network" } } } ``` State hosted, encrypted, versioned, locked by the platform; runs can execute remotely. Free tier covers ≤ 500 managed resources. Note: OpenTofu cannot use HCP Terraform as a backend (it can use the generic `remote` backend against compatible APIs). ### Backend selection | Backend | Locking | Encryption at rest | Best when | |---|---|---|---| | `s3` + `use_lockfile` | S3 conditional writes | SSE-KMS | AWS shops (default) | | `s3` + `dynamodb_table` | DynamoDB | SSE-KMS | Legacy / mixed TF < 1.10 teams | | `azurerm` | Blob lease (built-in) | Platform + CMK | Azure shops | | `gcs` | Generation precondition (built-in) | Platform + CMEK | GCP shops | | HCP Terraform | Platform | Platform | Want managed runs/policies too | | `pg` / `consul` / `kubernetes` | Yes | Varies | Niche; self-hosted constraints | ### Migrating backends ```bash # 1. Add/replace the backend block, then: terraform init -migrate-state # copies state old -> new, prompts # 2. Verify: terraform state list shows everything # 3. Delete the old state only after a clean plan ``` ## State Locking Locking prevents two concurrent applies corrupting state. It is **not** optional for teams. - `terraform apply` acquires the lock automatically; a crash can leave it stuck. - `terraform force-unlock <LOCK_ID>` — only after confirming no run is actually live (check CI). The lock ID is printed in the error. - `-lock-timeout=5m` in CI lets queued runs wait instead of failing instantly. - Locking protects against concurrent *writes*; it does not serialize plans — two PRs can both plan green and conflict at apply. Solve at the orchestration layer (Atlantis dir-locks, Actions concurrency groups — see cicd-pipelines.md). ## Declarative State Changes: moved / import / removed These blocks are the modern, code-reviewable replacements for CLI state surgery. They live in config, show up in diffs, and execute as part of a normal plan/apply. ### `moved` — refactor without destroy ```hcl # Renamed a resource moved { from = aws_instance.web to = aws_instance.frontend } # Moved a resource into a module moved { from = aws_vpc.main to = module.network.aws_vpc.main } # count -> for_each migration moved { from = aws_subnet.private[0] to = aws_subnet.private["ap-southeast-2a"] } ``` Plan shows `# aws_instance.web has moved to aws_instance.frontend` instead of destroy+create. Keep `moved` blocks around for at least one release cycle of a shared module so downstream consumers also get the move; then prune. ### `import` — adopt existing infrastructure (TF ≥ 1.5) ```hcl import { to = aws_s3_bucket.legacy id = "myorg-legacy-bucket" } ``` ```bash # Generate matching config for resources you haven't written yet: terraform plan -generate-config-out=generated.tf # Review generated.tf, clean it up (it's verbose), move into proper files, apply. ``` Why blocks beat `terraform import` CLI: the import is planned (you see exactly what will be adopted and whether config matches reality) and reviewed in the PR. The CLI command mutates state immediately with no plan. Import blocks also support `for_each` (TF ≥ 1.7) for bulk adoption. After a successful apply, delete the `import` blocks — they're one-shot. ### `removed` — forget without destroying (TF ≥ 1.7, OpenTofu ≥ 1.7) ```hcl removed { from = aws_db_instance.legacy lifecycle { destroy = false # remove from state, leave the real DB alone } } ``` Use when handing a resource to another team/state, or un-managing something Terraform should no longer own. The declarative version of `terraform state rm`. OpenTofu 1.12 additionally allows `lifecycle { destroy = false }` directly on a resource being deleted from config. ## State Surgery (CLI) — and when NOT to ```bash terraform state list # enumerate addresses (safe) terraform state show aws_vpc.main # inspect one resource (safe) terraform state pull > backup.tfstate # ALWAYS do this first terraform state mv aws_x.a aws_x.b # immediate, unreviewed rename terraform state mv aws_x.a 'module.m.aws_x.a' terraform state rm aws_x.a # forget (does NOT destroy) terraform state push fixed.tfstate # overwrite remote state (extreme) ``` **Decision rule:** if a `moved`, `removed`, or `import` block can express the change — use the block. Reasons: 1. Blocks are planned and reviewed; CLI mutations are instant and invisible to reviewers. 2. Blocks are idempotent across the team; a CLI command run by one person leaves everyone else's mental model stale. 3. A typo in `state mv` orphans a real resource: Terraform now plans to *create* a duplicate while the original drifts unmanaged. **Legitimate CLI surgery:** - Splitting one state into two roots (`state mv -state-out=...` or pull/edit/push between backends). - Recovering from a half-failed migration or a provider bug that wedged an address. - Anything on Terraform < 1.5 (no blocks available). **Protocol for any surgery:** `state pull` a timestamped backup → make the change → `terraform plan` must come back clean (or exactly the expected diff) → only then walk away. ## Drift Detection Infrastructure changes outside Terraform (console edits, autoscaling, other tooling). Detect it before it bites an apply. ```bash terraform plan -detailed-exitcode -lock=false # exit 0 -> no changes (clean) # exit 1 -> error # exit 2 -> changes pending (drift OR un-applied config) ``` - `-lock=false` so a read-only drift check never blocks a real apply. - `terraform plan -refresh-only` shows only *state vs reality* differences without proposing config-driven changes — cleaner signal for "who touched the console". - Cron it: nightly scheduled CI job, alert on exit 2 (recipe in cicd-pipelines.md). - Chronic drift on specific attributes → either codify the external process or `lifecycle { ignore_changes = [...] }` deliberately (document why). ## State File Hygiene | Rule | Why | |---|---| | `*.tfstate*` in `.gitignore` | Local state in git = secrets in git history forever | | One state per blast-radius unit | Network / data / app split — a bad apply can't take everything | | Keep states small (< ~100 resources guideline) | Plan time, lock contention, blast radius all scale with state size | | Versioned backend bucket | `state push` mistakes become a revert, not a rebuild | | Treat state as secret | Provider attributes (DB passwords, certs) sit in plaintext JSON — see security-and-secrets.md | | OpenTofu: consider state encryption | Client-side AES-GCM, keys via PBKDF2/KMS — defence even if the bucket leaks | ```hcl # OpenTofu >= 1.7 only — state + plan encryption terraform { encryption { key_provider "aws_kms" "main" { kms_key_id = "arn:aws:kms:...:key/..." key_spec = "AES_256" } method "aes_gcm" "main" { keys = key_provider.aws_kms.main } state { method = method.aes_gcm.main } plan { method = method.aes_gcm.main } } } ```
-
-
scripts
-
.gitkeep 0 B · in bundle
-
check-action-refs.sh 11.8 KB
#!/usr/bin/env bash # Staleness verifier for GitHub Actions `uses:` references in workflow YAML. # # Lints every `uses: owner/repo@ref` line in one or more workflow files. General- # purpose: not terraform-specific — point it at any GitHub Actions workflow. Two # modes per the staleness-verifier pattern (SKILL-RESOURCE-PROTOCOL §7): # --offline (default): structural-only, NO network. Asserts every `uses:` is # well-formed; floating @main/@master refs are a WARNING. # --live: resolves every owner/repo@ref against the GitHub API. # A 404 (ref does not exist) is DRIFT; rate-limit/offline # is UNAVAILABLE (advisory, never a build failure). # # Usage: check-action-refs.sh [--offline|--live] [--strict] [--json] [-q] [FILE ...] # Input: workflow YAML paths as positionals (default: the skill's own # assets/github-actions-terraform.yml). Pure grep/sed extraction — no # YAML library dependency. # Output: stdout = data only (findings list, or JSON envelope with --json) # Stderr: headers, progress, warnings, errors # Exit: 0 ok, 2 usage, 3 not-found, 4 malformed-uses, 5 missing-dep, # 7 api-unavailable (live), 10 drift (live: a ref does not resolve) # # Examples: # check-action-refs.sh --offline # check-action-refs.sh --offline .github/workflows/ci.yml # GITHUB_TOKEN=ghp_xxx check-action-refs.sh --live ci.yml deploy.yml # check-action-refs.sh --json --offline | jq '.data[] | select(.status!="ok")' set -uo pipefail EXIT_OK=0; EXIT_USAGE=2; EXIT_NOT_FOUND=3; EXIT_MALFORMED=4 EXIT_MISSING_DEP=5; EXIT_UNAVAILABLE=7; EXIT_DRIFT=10 # Terminal design system (skills/_lib/term.sh). Framing rides stderr (term_init 2); # the findings list / --json stay plain on stdout. Degrade if the lib is gone. __lib="$(cd "$(dirname "${BASH_SOURCE[0]}")/../../_lib" 2>/dev/null && pwd || true)" if [ -n "${__lib:-}" ] && [ -f "$__lib/term.sh" ]; then . "$__lib/term.sh"; term_init 2; __HAVE_TERM=1 else __HAVE_TERM=0; TERM_DOT="|"; fi SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" DEFAULT_FILE="${SCRIPT_DIR}/../assets/github-actions-terraform.yml" MODE="offline"; STRICT=0; JSON=0; QUIET=0; FILES=() while [[ $# -gt 0 ]]; do case "$1" in --offline) MODE="offline" ;; --live) MODE="live" ;; --strict) STRICT=1 ;; --json) JSON=1 ;; -q|--quiet) QUIET=1 ;; -h|--help) sed -n '2,26p' "$0" | sed 's/^# \{0,1\}//'; exit "$EXIT_OK" ;; -*) echo "ERROR: unknown flag: $1 (try --help)" >&2; exit "$EXIT_USAGE" ;; *) FILES+=("$1") ;; esac shift done [[ ${#FILES[@]} -eq 0 ]] && FILES=("$DEFAULT_FILE") command -v grep >/dev/null 2>&1 || { echo "ERROR: grep required" >&2; exit "$EXIT_MISSING_DEP"; } HAS_JQ=0; command -v jq >/dev/null 2>&1 && HAS_JQ=1 [[ "$JSON" -eq 1 && "$HAS_JQ" -eq 0 ]] && { echo '{"error":{"code":"PRECONDITION","message":"jq required for --json"}}' echo "ERROR: jq required for --json" >&2; exit "$EXIT_MISSING_DEP"; } if [[ "$MODE" == "live" ]]; then command -v curl >/dev/null 2>&1 || { echo "ERROR: curl required for --live" >&2; exit "$EXIT_MISSING_DEP"; } fi emit() { [[ "$QUIET" -eq 1 ]] && return; printf '%s\n' "$1" >&2; } # Panel framing applies to the human stderr stream when it's a TTY (or FORCE_COLOR # forces a render); piped/quiet consumers keep the legacy "=== / [TAG]" lines. PANEL=0 if [[ "$__HAVE_TERM" -eq 1 && "$QUIET" -eq 0 ]] && { [ -t 2 ] || [ -n "${FORCE_COLOR:-}" ]; }; then PANEL=1; fi __PANEL_OPEN=0 popen() { [[ "$PANEL" -eq 1 && "$__PANEL_OPEN" -eq 0 ]] || return 0 { term_panel_open terraform "action-refs ${TERM_DOT} ${MODE}"; term_panel_vert; } >&2 __PANEL_OPEN=1 } # prow <mark> <legacy-prefix> <text> — panel status row, or the legacy tagged line. prow() { if [[ "$PANEL" -eq 1 ]]; then popen; term_status_row "$1" "$3" >&2 else emit " $2 $3"; fi } # State accumulators malformed=0; drift=0; unavailable=0; warned=0 declare -a JSON_OBJS=() declare -a TEXT_ROWS=() # --- classify a `uses:` value ------------------------------------------------- # Sets globals: C_STATUS (ok|warn|malformed), C_OWNER, C_REPO, C_REF, C_KIND classify_uses() { local v=$1 C_OWNER=""; C_REPO=""; C_REF=""; C_KIND=""; C_STATUS="ok" # Local action: ./path — always valid, no network if [[ "$v" == ./* ]]; then C_KIND="local"; return; fi # Docker action: docker://image[:tag] — out of scope for ref resolution if [[ "$v" == docker://* ]]; then C_KIND="docker"; return; fi # Must contain an @ separating owner/repo[/path] from ref if [[ "$v" != *"@"* ]]; then C_STATUS="malformed"; C_KIND="action"; return; fi local path="${v%@*}" ref="${v#*@}" C_REF="$ref" # Empty ref, or empty path if [[ -z "$ref" || -z "$path" ]]; then C_STATUS="malformed"; C_KIND="action"; return; fi # path must be owner/repo[/subpath...] if [[ "$path" != */* ]]; then C_STATUS="malformed"; C_KIND="action"; return; fi C_OWNER="${path%%/*}" local rest="${path#*/}" C_REPO="${rest%%/*}" C_KIND="action" if [[ -z "$C_OWNER" || -z "$C_REPO" ]]; then C_STATUS="malformed"; return; fi # Validate ref shape: tag (vN / vN.N.N...), 40-hex SHA, or branch name. if [[ "$ref" =~ ^[0-9a-f]{40}$ ]]; then C_STATUS="ok" # pinned SHA — best practice elif [[ "$ref" == "main" || "$ref" == "master" ]]; then C_STATUS="warn" # floating default branch — advisory elif [[ "$ref" =~ ^v[0-9]+([._][0-9]+)*$ ]]; then C_STATUS="ok" # version tag elif [[ "$ref" =~ ^[A-Za-z0-9._/-]+$ ]]; then C_STATUS="ok" # other tag/branch name — structurally fine else C_STATUS="malformed" fi } # --- live resolution: does owner/repo@ref exist on GitHub? -------------------- # Echoes one of: resolved | notfound | unavailable resolve_ref() { local owner=$1 repo=$2 ref=$3 local -a auth=() [[ -n "${GITHUB_TOKEN:-}" ]] && auth=(-H "Authorization: Bearer ${GITHUB_TOKEN}") local base="https://api.github.com/repos/${owner}/${repo}" local code body url try() { local u=$1 out out=$(curl -sS -w $'\n%{http_code}' \ -H "Accept: application/vnd.github+json" \ -H "X-GitHub-Api-Version: 2022-11-28" \ "${auth[@]}" "$u" 2>/dev/null) [[ -z "$out" ]] && { echo "000"; return; } code="${out##*$'\n'}"; body="${out%$'\n'*}" echo "$code" } if [[ "$ref" =~ ^[0-9a-f]{40}$ ]]; then url="${base}/commits/${ref}" else url="${base}/git/ref/tags/${ref}" fi code=$(try "$url") # Network failure [[ "$code" == "000" ]] && { echo "unavailable"; return; } # Rate limit — advisory, never drift if [[ "$code" == "403" || "$code" == "429" ]]; then echo "unavailable"; return; fi if [[ "$code" == "200" ]]; then echo "resolved"; return; fi # Tag lookup 404'd — fall back to a commit/branch lookup before declaring drift. # A non-SHA ref can still be a branch name; the commits endpoint resolves those. if [[ "$code" == "404" && ! "$ref" =~ ^[0-9a-f]{40}$ ]]; then code=$(try "${base}/commits/${ref}") [[ "$code" == "000" ]] && { echo "unavailable"; return; } if [[ "$code" == "403" || "$code" == "429" ]]; then echo "unavailable"; return; fi [[ "$code" == "200" ]] && { echo "resolved"; return; } # 404 (no such branch) or 422 (unparseable commit-ish) — the ref does not exist [[ "$code" == "404" || "$code" == "422" ]] && { echo "notfound"; return; } echo "unavailable"; return fi # Direct SHA lookup that 404/422'd, or a tag 404 with no branch fallback path [[ "$code" == "404" || "$code" == "422" ]] && { echo "notfound"; return; } echo "unavailable" } add_json() { # file line ref status [[ "$HAS_JQ" -eq 1 ]] || return JSON_OBJS+=("$(jq -cn --arg f "$1" --argjson l "$2" --arg r "$3" --arg s "$4" \ '{file:$f, line:$l, ref:$r, status:$s}')") } if [[ "$PANEL" -eq 1 ]]; then popen; else emit "=== check-action-refs (${MODE}) ==="; fi for f in "${FILES[@]}"; do if [[ ! -f "$f" ]]; then prow bad "ERROR:" "file not found: $f" # In JSON mode still report a structured error per §5 if [[ "$JSON" -eq 1 ]]; then echo "{\"error\":{\"code\":\"NOT_FOUND\",\"message\":\"file not found: $f\"}}" fi exit "$EXIT_NOT_FOUND" fi # Extract every `uses:` value with its 1-based line number. Strip inline # comments and surrounding quotes. grep -n gives "LINE:content". while IFS= read -r entry; do [[ -z "$entry" ]] && continue lineno="${entry%%:*}" rawline="${entry#*:}" # value after `uses:` — drop leading `- ` list dash if present val="${rawline#*uses:}" val="${val%%#*}" # strip trailing comment # trim whitespace val="${val#"${val%%[![:space:]]*}"}" val="${val%"${val##*[![:space:]]}"}" # strip surrounding quotes val="${val%\"}"; val="${val#\"}" val="${val%\'}"; val="${val#\'}" [[ -z "$val" ]] && continue classify_uses "$val" case "$C_STATUS" in malformed) malformed=1 prow bad "[MALFORMED]" "${f}:${lineno} uses: ${val}" TEXT_ROWS+=("${f}:${lineno} ${val} malformed") add_json "$f" "$lineno" "$val" "malformed" ;; warn) warned=1 prow warn "[WARN floating]" "${f}:${lineno} ${val} (prefer SHA pin)" TEXT_ROWS+=("${f}:${lineno} ${val} warn") add_json "$f" "$lineno" "$val" "warn" ;; ok) if [[ "$MODE" == "live" && "$C_KIND" == "action" ]]; then res=$(resolve_ref "$C_OWNER" "$C_REPO" "$C_REF") case "$res" in resolved) prow ok "[ok]" "${f}:${lineno} ${C_OWNER}/${C_REPO}@${C_REF}" TEXT_ROWS+=("${f}:${lineno} ${val} ok") add_json "$f" "$lineno" "$val" "ok" ;; notfound) drift=1 prow bad "[DRIFT 404]" "${f}:${lineno} ${C_OWNER}/${C_REPO}@${C_REF}" TEXT_ROWS+=("${f}:${lineno} ${val} drift") add_json "$f" "$lineno" "$val" "drift" ;; unavailable) unavailable=1 prow warn "[unavailable]" "${f}:${lineno} ${C_OWNER}/${C_REPO}@${C_REF} (API unreachable/rate-limited)" TEXT_ROWS+=("${f}:${lineno} ${val} unavailable") add_json "$f" "$lineno" "$val" "unavailable" ;; esac else TEXT_ROWS+=("${f}:${lineno} ${val} ok") add_json "$f" "$lineno" "$val" "ok" fi ;; esac done < <(grep -nE '^[[:space:]]*-?[[:space:]]*uses:[[:space:]]*' "$f" 2>/dev/null) done # --- panel footer (stderr framing only) --------------------------------------- if [[ "$PANEL" -eq 1 && "$__PANEL_OPEN" -eq 1 ]]; then ph_state="healthy"; ph_text="refs well-formed" if [[ "$malformed" -eq 1 || "$drift" -eq 1 ]]; then ph_state="critical"; ph_text="findings present" elif [[ "$unavailable" -eq 1 ]]; then ph_state="warning"; ph_text="api unavailable" elif [[ "$warned" -eq 1 ]]; then ph_state="warning"; ph_text="floating refs"; fi { term_panel_vert term_panel_close "--live to resolve ${TERM_DOT} --json for data" "$(term_health "$ph_state" "$ph_text")" } >&2 fi # --- output ------------------------------------------------------------------- if [[ "$JSON" -eq 1 ]]; then printf '%s\n' "${JSON_OBJS[@]:-}" | jq -s \ --arg mode "$MODE" \ '{data: map(select(length>0)), meta: {mode:$mode, count:(map(select(length>0))|length), schema:"claude-mods.terraform-ops.action-refs/v1"}}' else # plain text: data rows to stdout (only the non-ok findings are interesting, # but emit all rows so the agent sees the full inventory) for row in "${TEXT_ROWS[@]:-}"; do [[ -n "$row" ]] && printf '%s\n' "$row" done fi # --- exit --------------------------------------------------------------------- [[ "$malformed" -eq 1 ]] && exit "$EXIT_MALFORMED" [[ "$drift" -eq 1 ]] && exit "$EXIT_DRIFT" [[ "$unavailable" -eq 1 ]] && exit "$EXIT_UNAVAILABLE" if [[ "$warned" -eq 1 && "$STRICT" -eq 1 ]]; then exit "$EXIT_MALFORMED"; fi exit "$EXIT_OK"
-
-
tests
-
run.sh 3.5 KB
#!/usr/bin/env bash # Self-test for terraform-ops — fully offline: no network, no GitHub API. # # Wraps the skill's §7 staleness verifier (scripts/check-action-refs.sh), which # lints `uses: owner/repo@ref` lines in GitHub Actions workflow YAML. Contract # (bash -n + --help), offline happy path against the shipped # assets/github-actions-terraform.yml, the --json §7 envelope (jq-guarded), and a # NEGATIVE proving the verifier rejects a malformed uses (a ref with no @ -> # exit 4 MALFORMED). --live is NEVER invoked (it resolves refs against the GitHub # API); a network blip must never fail a PR. # # Usage: bash tests/run.sh # Exit: 0 all pass, 1 one or more failures set -uo pipefail HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" SKILL="$(dirname "$HERE")" V="$SKILL/scripts/check-action-refs.sh" DEFAULT="$SKILL/assets/github-actions-terraform.yml" SB="$(mktemp -d)"; trap 'rm -rf "$SB"' EXIT PASS=0; FAIL=0 ok() { PASS=$((PASS+1)); printf ' PASS %s\n' "$1"; } no() { FAIL=$((FAIL+1)); printf ' FAIL %s\n' "$1"; } expect_exit() { [[ "$2" == "$3" ]] && ok "$1 (exit $3)" || no "$1 (want $2 got $3)"; } expect_has() { case "$3" in *"$2"*) ok "$1";; *) no "$1 (missing '$2')";; esac; } echo "=== terraform-ops self-test ===" # ── contract ────────────────────────────────────────────────────────────────── echo "-- contract --" bash -n "$V" && ok "bash -n check-action-refs.sh" || no "bash -n check-action-refs.sh" bash "$V" --help >/dev/null 2>&1; expect_exit "--help exits 0" 0 $? out="$(bash "$V" --help 2>&1)" expect_has "--help has Examples" "xamples" "$out" bash "$V" --bogus >/dev/null 2>&1; expect_exit "unknown flag -> 2" 2 $? # ── offline structural mode (§7 seam: --offline default, --live advisory) ───── echo "-- offline structural --" [[ -f "$DEFAULT" ]] && ok "shipped workflow present" || no "shipped workflow present" bash "$V" --offline >/dev/null 2>&1; expect_exit "--offline clean on shipped skill" 0 $? # --json needs jq; SKIP the envelope checks where jq is absent (a no-jq runner is # legitimate, and the verifier itself exits 5 for --json without jq). if command -v jq >/dev/null 2>&1; then out="$(bash "$V" --offline --json 2>/dev/null)" expect_has "--offline --json envelope schema" '"schema": "claude-mods.terraform-ops.action-refs/v1"' "$out" expect_has "--offline --json mode" '"mode": "offline"' "$out" else echo " SKIP --json envelope (no jq on this runner)" fi # ── negative: a uses with no @ref must be flagged malformed (exit 4) ────────── echo "-- negative --" cat > "$SB/bad.yml" <<'EOF' name: bad jobs: build: runs-on: ubuntu-latest steps: - uses: hashicorp/setup-terraform EOF bash "$V" --offline "$SB/bad.yml" >"$SB/neg.out" 2>&1 expect_exit "--offline flags missing @ref -> 4" 4 $? expect_has "finding names the malformed ref" "malformed" "$(cat "$SB/neg.out")" # ── SKILL.md sanity ─────────────────────────────────────────────────────────── echo "-- SKILL.md --" grep -q '^name: terraform-ops$' "$SKILL/SKILL.md" && ok "frontmatter name" || no "frontmatter name" grep -q 'check-action-refs.sh' "$SKILL/SKILL.md" && ok "verifier cited from SKILL.md" || no "verifier cited from SKILL.md" echo "" echo "=== $PASS passed, $FAIL failed ===" [[ "$FAIL" -eq 0 ]] || exit 1 exit 0
-
-
SKILL.md 16.2 KB
--- name: terraform-ops description: "Terraform and OpenTofu infrastructure-as-code operations - project layout, state management, module design, plan/apply safety, CI/CD pipelines, and secrets. Use for: terraform, opentofu, infrastructure as code, IaC, tfstate, terraform state, terraform module, remote backend, terraform plan, terraform apply, for_each, moved block, terraform import, drift detection, tflint, checkov, HCL." license: MIT allowed-tools: "Read Write Bash" metadata: author: claude-mods related-skills: ci-cd-ops, docker-ops, container-orchestration --- # Terraform Operations Terraform / OpenTofu infrastructure-as-code: layout, state, modules, safety, CI/CD, secrets. **Version context (verified 2026-06):** Terraform **1.15.x** (BUSL-1.1 licence since 1.6) · OpenTofu **1.12.x** (MPL-2.0 fork of Terraform 1.5.x). Commands below are interchangeable (`terraform` ↔ `tofu`) unless flagged. See [Terraform vs OpenTofu](#terraform-vs-opentofu) for the decision note. ## Reference Files | File | Covers | |------|--------| | [references/state-management.md](references/state-management.md) | Remote backends, locking, moved/import/removed blocks, state surgery, drift detection | | [references/module-patterns.md](references/module-patterns.md) | Module composition, variable validation, optional/nullable, output contracts, versioning | | [references/cicd-pipelines.md](references/cicd-pipelines.md) | GitHub Actions plan/apply, OIDC auth, policy gates (tflint/trivy/checkov/OPA), Atlantis/HCP | | [references/security-and-secrets.md](references/security-and-secrets.md) | Secrets in state, ephemeral resources, write-only arguments, SOPS/Vault, sensitive limits | | [assets/github-actions-terraform.yml](assets/github-actions-terraform.yml) | Ready-to-adapt PR-plan + OIDC-apply workflow | | [scripts/check-action-refs.sh](scripts/check-action-refs.sh) | Staleness verifier for any workflow's `uses:` action refs (offline structural / live API resolve) | > The action versions pinned in `github-actions-terraform.yml` are **point-in-time** (verified 2026-06). Run `scripts/check-action-refs.sh --live` before adopting — a tag that was valid at write time may have been retracted or never existed (e.g. `trivy-action@0.33.1` vs the real `v0.33.1`). ## Project Layout Decision Tree ``` How many environments / accounts? │ ├─ One environment, one team │ └─ Single root module + tfvars. Don't over-engineer. │ ├─ Multiple environments (dev/staging/prod) │ ├─ Need different backend/account/region per env? (usually YES for prod isolation) │ │ └─ DIRECTORY PER ENVIRONMENT (recommended default) │ │ environments/{dev,staging,prod}/ each a thin root calling shared modules │ │ │ └─ Environments truly identical except a few variables, same backend account? │ └─ Workspaces are *acceptable* — but see the workspace caveats below │ └─ Many teams / many state files / platform engineering └─ Directory-per-env + per-component state split (network / data / app) Consider Terragrunt, Terraform Stacks (HCP), or OpenTofu + CI orchestration ``` ### Canonical multi-env layout ``` infra/ ├── modules/ # Reusable child modules (no provider/backend blocks) │ ├── network/ │ │ ├── main.tf │ │ ├── variables.tf │ │ ├── outputs.tf │ │ └── versions.tf # required_providers ONLY (no provider config) │ └── app-service/ ├── environments/ # Root modules — one state file each │ ├── dev/ │ │ ├── main.tf # module "network" { source = "../../modules/network" ... } │ │ ├── backend.tf # remote backend, env-specific key │ │ ├── providers.tf # provider config lives in ROOT only │ │ ├── terraform.tfvars # committed, non-secret env values │ │ └── versions.tf # required_version + required_providers pins │ └── prod/ └── .tflint.hcl ``` ### Why directories usually beat workspaces | Concern | Directories | Workspaces | |---------|-------------|------------| | Separate backend/account per env | Yes — each root has its own `backend.tf` | No — one backend, envs differ only by state key | | Blast radius of wrong-env apply | Low — you're physically in `prod/` | High — invisible `terraform workspace select` state | | Env-specific config divergence | Natural (different main.tf if needed) | `terraform.workspace` conditionals creep everywhere | | Prod IAM isolation | Per-dir CI role | Same credentials see all envs | | Visibility in code review | Diff shows which env changed | Workspace is runtime state, not in the diff | Workspaces fit short-lived ephemeral copies (PR preview envs) — not the dev/prod boundary. HashiCorp's own docs say workspaces are "not suitable for strong separation." ### tfvars conventions ```bash terraform.tfvars # auto-loaded — per-root committed defaults (non-secret) *.auto.tfvars # auto-loaded — generated/local overrides prod.tfvars # explicit only: terraform plan -var-file=prod.tfvars TF_VAR_db_password=... # env var injection — secrets in CI, never in files ``` Gotcha: `-var-file` + directories-per-env is belt-and-braces; with workspaces it's load-bearing and one forgotten flag applies dev values to prod. ## State Quick Reference Full detail: [references/state-management.md](references/state-management.md). | Task | Command / block | Notes | |------|-----------------|-------| | Remote backend (AWS) | `backend "s3" { bucket, key, region, use_lockfile = true }` | S3-native locking (TF ≥1.10) — DynamoDB table no longer required | | Rename resource in code | `moved { from = aws_x.a, to = aws_x.b }` | Declarative, reviewable, no CLI surgery | | Adopt existing infra | `import { to = aws_x.a, id = "i-123" }` + `plan -generate-config-out=gen.tf` | Config-driven import (TF ≥1.5) beats `terraform import` CLI | | Forget without destroy | `removed { from = aws_x.a, lifecycle { destroy = false } }` | TF ≥1.7; OpenTofu 1.12 also has `lifecycle { destroy = false }` on resources | | Drift detection | `terraform plan -detailed-exitcode` | Exit 0 = clean, 1 = error, **2 = drift** — cron it | | Inspect state | `terraform state list` / `state show ADDR` | Read-only, always safe | | Move state (last resort) | `terraform state mv SRC DST` | Prefer `moved` blocks — see "when NOT to" below | | Pull/push (emergency) | `terraform state pull > backup.tfstate` | ALWAYS pull a backup before any surgery | **State surgery — when NOT to:** if a `moved`/`removed`/`import` block can express it, use the block. CLI `state mv`/`rm` is immediate, unreviewed, unversioned, and a typo orphans real infrastructure. Legit uses: splitting state between roots, unwedging a failed migration. Always `state pull` a backup first. ## Module Quick Reference Full detail: [references/module-patterns.md](references/module-patterns.md). ```hcl module "network" { source = "terraform-aws-modules/vpc/aws" version = "~> 6.0" # pin minor-float for registry modules; exact pin in prod roots # ... } ``` | Rule | Why | |------|-----| | Composition over inheritance | Roots compose flat modules; never module-wraps-module-wraps-module | | No provider blocks in child modules | Providers configured in root only; child declares `required_providers` | | `validation` blocks on variables | Fail at plan with a real message, not mid-apply | | `optional(type, default)` in object attrs | Callers omit fields; `nullable = false` rejects explicit null | | Outputs are the contract | Output IDs/ARNs consumers need; document with `description` | | **Anti-pattern: thin wrappers** | A module that just renames variables of another module adds a version-lag layer and zero value — call the upstream module directly | ## Safety Checklist (before every apply) ``` □ plan output READ, not skimmed — every destroy/replace explained □ "Plan: X to add, Y to change, Z to destroy" — does Z surprise you? □ -/+ (replace) lines: check the "forces replacement" attribute □ Applying the SAME saved plan that was reviewed: plan -out=tfplan → apply tfplan □ prevent_destroy on stateful resources (db, state bucket, KMS keys) □ Cloud-side deletion protection too (RDS deletion_protection, S3 versioning+MFA-delete) □ No -target unless this is a declared emergency (see below) □ for_each (stable keys), not count, for any collection that can reorder ``` ### Footguns | Footgun | Detail | Fix | |---------|--------|-----| | `count` index shift | Removing item 0 of a `count` list re-addresses every later item → destroy/recreate cascade | `for_each` with stable string keys | | `-target` habit | Skips dependency graph; state diverges from config; hides drift | Emergency-only (broken dependency cycle, partial outage). Follow with a full clean plan | | `prevent_destroy` false comfort | Doesn't survive the block being deleted, and doesn't stop `state rm` + console delete | Pair with cloud-native deletion protection | | Dynamic blocks everywhere | `dynamic` for 2 static blocks is obfuscation | Use `dynamic` only over genuinely variable collections | | Unpinned providers | `aws = ">= 5.0"` in prod pulls a breaking major the day it ships | `~> 6.12` + commit `.terraform.lock.hcl` | | Apply ≠ reviewed plan | Plan on PR, apply on merge re-plans — drift in between applies unreviewed changes | Save the plan artifact, or accept + re-review the merge plan | ```hcl resource "aws_db_instance" "main" { deletion_protection = true # cloud-side lifecycle { prevent_destroy = true # terraform-side ignore_changes = [password] # if rotated outside TF } } ``` ## CI/CD Quick Reference Full detail: [references/cicd-pipelines.md](references/cicd-pipelines.md) · template: [assets/github-actions-terraform.yml](assets/github-actions-terraform.yml). ``` PR opened → fmt -check → validate → tflint → trivy/checkov → plan → plan posted as PR comment PR merged → plan (fresh) → apply, authenticated via OIDC — no long-lived cloud keys Nightly → plan -detailed-exitcode → exit 2 ⇒ drift alert ``` - **OIDC everywhere** — `aws-actions/configure-aws-credentials` with `role-to-assume`, never `AWS_ACCESS_KEY_ID` secrets. Same supply-chain doctrine as this repo's rules: short-lived tokens, no standing credentials. - **Pin action SHAs** in workflows (`uses: actions/checkout@<sha>`), not floating tags. - Policy gates: `tflint` (provider-aware lint), `trivy config` / `checkov` (misconfig scan), OPA/`conftest` for org policy ("no public buckets"). | Orchestrator | Fit | |---|---| | Plain GitHub Actions | Default — full control, free, template in assets/ | | Atlantis | Self-hosted PR automation, `atlantis plan/apply` comments, locking per dir | | HCP Terraform / Terraform Cloud | Managed runs, Sentinel policy, state hosting; free ≤500 resources | | Spacelift / env0 / Digger / Scalr | Commercial Atlantis-likes; Digger runs inside your Actions | ### Verification — `uses:` ref staleness GitHub Action versions rot: a tag gets retracted, or a workflow pins one that never existed. [scripts/check-action-refs.sh](scripts/check-action-refs.sh) lints every `uses: owner/repo@ref` line. It's **general** — pass any workflow file(s) as positionals (default: this skill's own `assets/github-actions-terraform.yml`). ```bash # Structural only, no network — well-formedness of every uses: ref (CI-safe gate). # Floating @main/@master → WARN (exit 0; use --strict to fail). Malformed → exit 4. scripts/check-action-refs.sh --offline .github/workflows/ci.yml # Live — resolve each ref against the GitHub API. A 404 (ref doesn't exist) → exit 10 # DRIFT; API unreachable/rate-limited → exit 7 (advisory, never fails the build, §7). # Set GITHUB_TOKEN to dodge the unauthenticated rate limit. GITHUB_TOKEN=$GH_PAT scripts/check-action-refs.sh --live .github/workflows/*.yml scripts/check-action-refs.sh --json --offline | jq '.data[] | select(.status!="ok")' ``` `--live` is the check that catches the classic `aquasecurity/trivy-action@0.33.1` mistake — that tag 404s; the real one is `v0.33.1`. Run live on a schedule (never as a blocking PR gate), offline in PR CI. ## Testing Quick Reference ```hcl # tests/network.tftest.hcl — native test framework (TF ≥1.6 / OpenTofu ≥1.6) variables { cidr = "10.0.0.0/16" } run "valid_cidr_plan" { command = plan # plan = fast unit-ish; apply = real integration assert { condition = aws_vpc.main.cidr_block == "10.0.0.0/16" error_message = "VPC CIDR did not match input" } } run "rejects_tiny_cidr" { command = plan variables { cidr = "10.0.0.0/30" } expect_failures = [var.cidr] # asserts the validation block fires } ``` `terraform test` runs every `*.tftest.hcl` under `tests/`; `command = apply` runs create real (then auto-destroyed) infra — use a sandbox account. Mock providers (`mock_provider` blocks, TF ≥1.7) fake apply without credentials. For multi-tool/Go-level orchestration (retry, real HTTP probes), Terratest is the heavyweight alternative — native `terraform test` covers most module CI needs first. ## Secrets Quick Reference Full detail: [references/security-and-secrets.md](references/security-and-secrets.md). | Mechanism | Version | What it does | |---|---|---| | `sensitive = true` | all | Redacts from CLI output **only** — value still plaintext in state | | Ephemeral resources (`ephemeral "..."`) | TF ≥1.10 / OpenTofu ≥1.11 | Fetch secret at run time; never persisted to state or plan | | Write-only arguments (`password_wo`) | TF ≥1.11 / OpenTofu ≥1.11 | Send secret to provider; never stored in state; rotate via `_wo_version` | | SOPS-encrypted tfvars | tool | Secrets encrypted at rest in git; decrypted at plan time | | Vault / cloud secret manager | tool | Reference by ID; resource reads secret at boot, TF never sees it | | OpenTofu state encryption | OpenTofu ≥1.7 | Client-side AES-GCM encryption of state/plan — **no Terraform equivalent** | **Rule zero: treat state as secret regardless.** Encrypt the backend (SSE-KMS), restrict IAM on the bucket, never commit `*.tfstate` (gitignore it). ## Terraform vs OpenTofu | | Terraform | OpenTofu | |---|---|---| | Licence | **BUSL-1.1** since 1.6 (no production use *competing with HashiCorp*; fine for normal internal use) | **MPL-2.0** — genuinely open source, Linux Foundation | | Current | 1.15.x | 1.12.x | | Exclusive features | Stacks (HCP-tied), Terraform Cloud agents, `terraform query` | State/plan **encryption**, provider `for_each` iteration, `-exclude` flag, early variable eval in backend/module blocks, OCI registry distribution, `.tofu` file extension | | Registry | registry.terraform.io | registry.opentofu.org (mirrors most providers) | | Compatibility | — | Forked at 1.5.x; HCL/state compatible for mainstream use, diverging feature-by-feature since | **Decision:** vendors and anyone redistributing IaC tooling commercially → OpenTofu (licence risk). Teams on HCP Terraform/Sentinel → Terraform. Everyone else: either works; OpenTofu's state encryption is the single biggest technical differentiator. Migration `terraform → tofu` is `tofu init` + state-compatible up to ~1.8-era features; the gap widens each release — migrate early or commit. ## Command Quick Reference ```bash terraform init -upgrade # init / upgrade providers within constraints terraform fmt -recursive -check # CI: fail on unformatted terraform validate # syntax + internal consistency (no creds needed after init) terraform plan -out=tfplan # save plan for exact-apply terraform show -json tfplan | jq # machine-readable plan (policy tools eat this) terraform apply tfplan # apply EXACTLY the reviewed plan terraform plan -detailed-exitcode # 0 clean / 2 drift — for cron drift checks terraform plan -refresh-only # show drift without proposing config changes terraform apply -replace=aws_x.a # force recreate one resource (replaces old taint) terraform state pull > backup.json # ALWAYS before surgery terraform output -json # consume outputs in scripts terraform graph | dot -Tsvg > g.svg # dependency graph tofu init # OpenTofu: same verbs throughout ```
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.