Claude Skill

terraform-ops

Terraform and OpenTofu infrastructure-as-code operations - project layout, state management, module design, plan/apply safety, CI/CD pipelines, and secrets. Use for: terraform, opentofu, infrastructure as code, IaC, tfstate, terraform state, terraform module, remote backend, terr

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download 0xdarkmatter-claude-mods-skills_terraform-ops-3dfaf0b.zip · 33 KB
Part of 0xdarkmatter/claude-mods — 94 skills

Install

skills CLI npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/terraform-ops
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
Git git clone https://github.com/0xDarkMatter/claude-mods.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Terraform Operations

Terraform / OpenTofu infrastructure-as-code: layout, state, modules, safety, CI/CD, secrets.

Version context (verified 2026-06): Terraform 1.15.x (BUSL-1.1 licence since 1.6) · OpenTofu 1.12.x (MPL-2.0 fork of Terraform 1.5.x). Commands below are interchangeable (terraform ↔ tofu) unless flagged. See Terraform vs OpenTofu for the decision note.

Reference Files

File Covers
references/state-management.md Remote backends, locking, moved/import/removed blocks, state surgery, drift detection
references/module-patterns.md Module composition, variable validation, optional/nullable, output contracts, versioning
references/cicd-pipelines.md GitHub Actions plan/apply, OIDC auth, policy gates (tflint/trivy/checkov/OPA), Atlantis/HCP
references/security-and-secrets.md Secrets in state, ephemeral resources, write-only arguments, SOPS/Vault, sensitive limits
assets/github-actions-terraform.yml Ready-to-adapt PR-plan + OIDC-apply workflow
scripts/check-action-refs.sh Staleness verifier for any workflow's uses: action refs (offline structural / live API resolve)

The action versions pinned in github-actions-terraform.yml are point-in-time (verified 2026-06). Run scripts/check-action-refs.sh --live before adopting — a tag that was valid at write time may have been retracted or never existed (e.g. trivy-action@0.33.1 vs the real v0.33.1).

Project Layout Decision Tree

How many environments / accounts?
│
├─ One environment, one team
│  └─ Single root module + tfvars. Don't over-engineer.
│
├─ Multiple environments (dev/staging/prod)
│  ├─ Need different backend/account/region per env? (usually YES for prod isolation)
│  │  └─ DIRECTORY PER ENVIRONMENT (recommended default)
│  │     environments/{dev,staging,prod}/ each a thin root calling shared modules
│  │
│  └─ Environments truly identical except a few variables, same backend account?
│     └─ Workspaces are *acceptable* — but see the workspace caveats below
│
└─ Many teams / many state files / platform engineering
   └─ Directory-per-env + per-component state split (network / data / app)
      Consider Terragrunt, Terraform Stacks (HCP), or OpenTofu + CI orchestration

Canonical multi-env layout

infra/
├── modules/                  # Reusable child modules (no provider/backend blocks)
│   ├── network/
│   │   ├── main.tf
│   │   ├── variables.tf
│   │   ├── outputs.tf
│   │   └── versions.tf       # required_providers ONLY (no provider config)
│   └── app-service/
├── environments/             # Root modules — one state file each
│   ├── dev/
│   │   ├── main.tf           # module "network" { source = "../../modules/network" ... }
│   │   ├── backend.tf        # remote backend, env-specific key
│   │   ├── providers.tf      # provider config lives in ROOT only
│   │   ├── terraform.tfvars  # committed, non-secret env values
│   │   └── versions.tf       # required_version + required_providers pins
│   └── prod/
└── .tflint.hcl

Why directories usually beat workspaces

Concern Directories Workspaces
Separate backend/account per env Yes — each root has its own backend.tf No — one backend, envs differ only by state key
Blast radius of wrong-env apply Low — you're physically in prod/ High — invisible terraform workspace select state
Env-specific config divergence Natural (different main.tf if needed) terraform.workspace conditionals creep everywhere
Prod IAM isolation Per-dir CI role Same credentials see all envs
Visibility in code review Diff shows which env changed Workspace is runtime state, not in the diff

Workspaces fit short-lived ephemeral copies (PR preview envs) — not the dev/prod boundary. HashiCorp's own docs say workspaces are "not suitable for strong separation."

tfvars conventions

terraform.tfvars            # auto-loaded — per-root committed defaults (non-secret)
*.auto.tfvars               # auto-loaded — generated/local overrides
prod.tfvars                 # explicit only: terraform plan -var-file=prod.tfvars
TF_VAR_db_password=...      # env var injection — secrets in CI, never in files

Gotcha: -var-file + directories-per-env is belt-and-braces; with workspaces it's load-bearing and one forgotten flag applies dev values to prod.

State Quick Reference

Full detail: references/state-management.md.

Task Command / block Notes
Remote backend (AWS) backend "s3" { bucket, key, region, use_lockfile = true } S3-native locking (TF ≥1.10) — DynamoDB table no longer required
Rename resource in code moved { from = aws_x.a, to = aws_x.b } Declarative, reviewable, no CLI surgery
Adopt existing infra import { to = aws_x.a, id = "i-123" } + plan -generate-config-out=gen.tf Config-driven import (TF ≥1.5) beats terraform import CLI
Forget without destroy removed { from = aws_x.a, lifecycle { destroy = false } } TF ≥1.7; OpenTofu 1.12 also has lifecycle { destroy = false } on resources
Drift detection terraform plan -detailed-exitcode Exit 0 = clean, 1 = error, 2 = drift — cron it
Inspect state terraform state list / state show ADDR Read-only, always safe
Move state (last resort) terraform state mv SRC DST Prefer moved blocks — see "when NOT to" below
Pull/push (emergency) terraform state pull > backup.tfstate ALWAYS pull a backup before any surgery

State surgery — when NOT to: if a moved/removed/import block can express it, use the block. CLI state mv/rm is immediate, unreviewed, unversioned, and a typo orphans real infrastructure. Legit uses: splitting state between roots, unwedging a failed migration. Always state pull a backup first.

Module Quick Reference

Full detail: references/module-patterns.md.

module "network" {
  source  = "terraform-aws-modules/vpc/aws"
  version = "~> 6.0"          # pin minor-float for registry modules; exact pin in prod roots
  # ...
}
Rule Why
Composition over inheritance Roots compose flat modules; never module-wraps-module-wraps-module
No provider blocks in child modules Providers configured in root only; child declares required_providers
validation blocks on variables Fail at plan with a real message, not mid-apply
optional(type, default) in object attrs Callers omit fields; nullable = false rejects explicit null
Outputs are the contract Output IDs/ARNs consumers need; document with description
Anti-pattern: thin wrappers A module that just renames variables of another module adds a version-lag layer and zero value — call the upstream module directly

Safety Checklist (before every apply)

□ plan output READ, not skimmed — every destroy/replace explained
□ "Plan: X to add, Y to change, Z to destroy" — does Z surprise you?
□ -/+ (replace) lines: check the "forces replacement" attribute
□ Applying the SAME saved plan that was reviewed: plan -out=tfplan → apply tfplan
□ prevent_destroy on stateful resources (db, state bucket, KMS keys)
□ Cloud-side deletion protection too (RDS deletion_protection, S3 versioning+MFA-delete)
□ No -target unless this is a declared emergency (see below)
□ for_each (stable keys), not count, for any collection that can reorder

Footguns

Footgun Detail Fix
count index shift Removing item 0 of a count list re-addresses every later item → destroy/recreate cascade for_each with stable string keys
-target habit Skips dependency graph; state diverges from config; hides drift Emergency-only (broken dependency cycle, partial outage). Follow with a full clean plan
prevent_destroy false comfort Doesn't survive the block being deleted, and doesn't stop state rm + console delete Pair with cloud-native deletion protection
Dynamic blocks everywhere dynamic for 2 static blocks is obfuscation Use dynamic only over genuinely variable collections
Unpinned providers aws = ">= 5.0" in prod pulls a breaking major the day it ships ~> 6.12 + commit .terraform.lock.hcl
Apply ≠ reviewed plan Plan on PR, apply on merge re-plans — drift in between applies unreviewed changes Save the plan artifact, or accept + re-review the merge plan
resource "aws_db_instance" "main" {
  deletion_protection = true            # cloud-side
  lifecycle {
    prevent_destroy = true              # terraform-side
    ignore_changes  = [password]        # if rotated outside TF
  }
}

CI/CD Quick Reference

Full detail: references/cicd-pipelines.md · template: assets/github-actions-terraform.yml.

PR opened   → fmt -check → validate → tflint → trivy/checkov → plan → plan posted as PR comment
PR merged   → plan (fresh) → apply, authenticated via OIDC — no long-lived cloud keys
Nightly     → plan -detailed-exitcode → exit 2 ⇒ drift alert
  • OIDC everywhere — aws-actions/configure-aws-credentials with role-to-assume, never AWS_ACCESS_KEY_ID secrets. Same supply-chain doctrine as this repo's rules: short-lived tokens, no standing credentials.
  • Pin action SHAs in workflows (uses: actions/checkout@<sha>), not floating tags.
  • Policy gates: tflint (provider-aware lint), trivy config / checkov (misconfig scan), OPA/conftest for org policy ("no public buckets").
Orchestrator Fit
Plain GitHub Actions Default — full control, free, template in assets/
Atlantis Self-hosted PR automation, atlantis plan/apply comments, locking per dir
HCP Terraform / Terraform Cloud Managed runs, Sentinel policy, state hosting; free ≤500 resources
Spacelift / env0 / Digger / Scalr Commercial Atlantis-likes; Digger runs inside your Actions

Verification — uses: ref staleness

GitHub Action versions rot: a tag gets retracted, or a workflow pins one that never existed. scripts/check-action-refs.sh lints every uses: owner/repo@ref line. It's general — pass any workflow file(s) as positionals (default: this skill's own assets/github-actions-terraform.yml).

# Structural only, no network — well-formedness of every uses: ref (CI-safe gate).
# Floating @main/@master → WARN (exit 0; use --strict to fail). Malformed → exit 4.
scripts/check-action-refs.sh --offline .github/workflows/ci.yml

# Live — resolve each ref against the GitHub API. A 404 (ref doesn't exist) → exit 10
# DRIFT; API unreachable/rate-limited → exit 7 (advisory, never fails the build, §7).
# Set GITHUB_TOKEN to dodge the unauthenticated rate limit.
GITHUB_TOKEN=$GH_PAT scripts/check-action-refs.sh --live .github/workflows/*.yml

scripts/check-action-refs.sh --json --offline | jq '.data[] | select(.status!="ok")'

--live is the check that catches the classic aquasecurity/trivy-action@0.33.1 mistake — that tag 404s; the real one is v0.33.1. Run live on a schedule (never as a blocking PR gate), offline in PR CI.

Testing Quick Reference

# tests/network.tftest.hcl  — native test framework (TF ≥1.6 / OpenTofu ≥1.6)
variables { cidr = "10.0.0.0/16" }

run "valid_cidr_plan" {
  command = plan                          # plan = fast unit-ish; apply = real integration
  assert {
    condition     = aws_vpc.main.cidr_block == "10.0.0.0/16"
    error_message = "VPC CIDR did not match input"
  }
}

run "rejects_tiny_cidr" {
  command = plan
  variables { cidr = "10.0.0.0/30" }
  expect_failures = [var.cidr]            # asserts the validation block fires
}

terraform test runs every *.tftest.hcl under tests/; command = apply runs create real (then auto-destroyed) infra — use a sandbox account. Mock providers (mock_provider blocks, TF ≥1.7) fake apply without credentials. For multi-tool/Go-level orchestration (retry, real HTTP probes), Terratest is the heavyweight alternative — native terraform test covers most module CI needs first.

Secrets Quick Reference

Full detail: references/security-and-secrets.md.

Mechanism Version What it does
sensitive = true all Redacts from CLI output only — value still plaintext in state
Ephemeral resources (ephemeral "...") TF ≥1.10 / OpenTofu ≥1.11 Fetch secret at run time; never persisted to state or plan
Write-only arguments (password_wo) TF ≥1.11 / OpenTofu ≥1.11 Send secret to provider; never stored in state; rotate via _wo_version
SOPS-encrypted tfvars tool Secrets encrypted at rest in git; decrypted at plan time
Vault / cloud secret manager tool Reference by ID; resource reads secret at boot, TF never sees it
OpenTofu state encryption OpenTofu ≥1.7 Client-side AES-GCM encryption of state/plan — no Terraform equivalent

Rule zero: treat state as secret regardless. Encrypt the backend (SSE-KMS), restrict IAM on the bucket, never commit *.tfstate (gitignore it).

Terraform vs OpenTofu

Terraform OpenTofu
Licence BUSL-1.1 since 1.6 (no production use competing with HashiCorp; fine for normal internal use) MPL-2.0 — genuinely open source, Linux Foundation
Current 1.15.x 1.12.x
Exclusive features Stacks (HCP-tied), Terraform Cloud agents, terraform query State/plan encryption, provider for_each iteration, -exclude flag, early variable eval in backend/module blocks, OCI registry distribution, .tofu file extension
Registry registry.terraform.io registry.opentofu.org (mirrors most providers)
Compatibility — Forked at 1.5.x; HCL/state compatible for mainstream use, diverging feature-by-feature since

Decision: vendors and anyone redistributing IaC tooling commercially → OpenTofu (licence risk). Teams on HCP Terraform/Sentinel → Terraform. Everyone else: either works; OpenTofu's state encryption is the single biggest technical differentiator. Migration terraform → tofu is tofu init + state-compatible up to ~1.8-era features; the gap widens each release — migrate early or commit.

Command Quick Reference

terraform init -upgrade               # init / upgrade providers within constraints
terraform fmt -recursive -check       # CI: fail on unformatted
terraform validate                    # syntax + internal consistency (no creds needed after init)
terraform plan -out=tfplan            # save plan for exact-apply
terraform show -json tfplan | jq      # machine-readable plan (policy tools eat this)
terraform apply tfplan                # apply EXACTLY the reviewed plan
terraform plan -detailed-exitcode     # 0 clean / 2 drift — for cron drift checks
terraform plan -refresh-only          # show drift without proposing config changes
terraform apply -replace=aws_x.a      # force recreate one resource (replaces old taint)
terraform state pull > backup.json    # ALWAYS before surgery
terraform output -json                # consume outputs in scripts
terraform graph | dot -Tsvg > g.svg   # dependency graph
tofu init                             # OpenTofu: same verbs throughout
Files (claude-mods)
  • assets
    • github-actions-terraform.yml 8.7 KB
      # =============================================================================
      # Terraform CI/CD — PR plan + OIDC apply (GitHub Actions template)
      # =============================================================================
      # What this gives you:
      #   * PRs:   fmt-check -> validate -> tflint -> trivy -> plan -> plan posted as
      #            a PR comment (updated in place, not stacked)
      #   * Merge: fresh plan + apply on main, gated by a GitHub "environment"
      #            (add required reviewers on the environment for a manual approval)
      #   * Auth:  AWS via OIDC — NO long-lived access keys anywhere.
      #
      # Adapt before use:
      #   1. Set WORKING_DIR to your root module (e.g. environments/prod).
      #   2. Replace the two role ARNs (plan = read-only, apply = write).
      #   3. Pin every `uses:` to a commit SHA for your security posture —
      #      tags are shown here for readability; SHA-pin in production:
      #        uses: actions/checkout@08c6903cd8c0fde910a37f88322edcfb5dd907a8 # v5.0.0
      #   4. OpenTofu: swap hashicorp/setup-terraform for opentofu/setup-opentofu
      #      (input `tofu_version`) and `terraform` for `tofu` in run steps.
      #   5. Multi-env monorepo: turn `WORKING_DIR` into a job matrix.
      # =============================================================================
      
      name: terraform
      
      on:
        pull_request:
          branches: [main]
          paths: ["environments/prod/**", "modules/**"]   # plan only when infra changes
        push:
          branches: [main]
          paths: ["environments/prod/**", "modules/**"]
      
      # Least privilege by default; jobs widen only what they need.
      permissions:
        contents: read
      
      env:
        TF_VERSION: "1.15.5"            # match required_version in versions.tf
        WORKING_DIR: environments/prod
        AWS_REGION: ap-southeast-2
        TF_IN_AUTOMATION: "true"        # suppresses interactive-use hints in output
      
      # One run per state file at a time. Plans queue; applies never overlap.
      concurrency:
        group: terraform-prod
        cancel-in-progress: false       # NEVER cancel a running apply
      
      jobs:
        # ---------------------------------------------------------------------------
        # PR: static checks + plan + comment
        # ---------------------------------------------------------------------------
        plan:
          if: github.event_name == 'pull_request'
          runs-on: ubuntu-latest
          permissions:
            contents: read
            id-token: write             # OIDC token for AWS
            pull-requests: write        # post the plan comment
          defaults:
            run:
              working-directory: ${{ env.WORKING_DIR }}
          steps:
            - uses: actions/checkout@v5
      
            - uses: hashicorp/setup-terraform@v3
              with:
                terraform_version: ${{ env.TF_VERSION }}
      
            # --- cheap static gates first: fail fast before touching the cloud ---
            - name: fmt
              run: terraform fmt -check -recursive -diff
      
            - name: tflint
              uses: terraform-linters/setup-tflint@v6
            - run: tflint --init && tflint --recursive
              working-directory: ${{ env.WORKING_DIR }}
      
            - name: trivy misconfig scan
              uses: aquasecurity/trivy-action@v0.33.1
              with:
                scan-type: config
                scan-ref: ${{ env.WORKING_DIR }}
                exit-code: "1"          # hard gate; set "0" while triaging a baseline
                severity: HIGH,CRITICAL
      
            # --- read-only cloud credentials, scoped to PR plans ---
            # This role's trust policy allows any branch of THIS repo, but the role
            # itself carries ReadOnlyAccess + state-bucket read. PR code cannot mutate.
            - name: Configure AWS credentials (plan role, read-only)
              uses: aws-actions/configure-aws-credentials@v5
              with:
                role-to-assume: arn:aws:iam::123456789012:role/github-terraform-plan
                aws-region: ${{ env.AWS_REGION }}
      
            - name: init
              run: terraform init -input=false -lock-timeout=2m
      
            - name: validate
              run: terraform validate -no-color
      
            - name: plan
              id: plan
              # -lock=false: a PR plan must never block (or be blocked by) an apply
              run: |
                set -o pipefail
                terraform plan -input=false -no-color -lock=false -out=tfplan 2>&1 \
                  | tee plan.txt
              continue-on-error: true   # we still want the comment when plan fails
      
            # Optional org-policy gate against the machine-readable plan:
            # - run: terraform show -json tfplan > tfplan.json
            # - run: conftest test --policy ../../policy tfplan.json
      
            - name: Comment plan on PR
              uses: actions/github-script@v8
              env:
                PLAN_OUTCOME: ${{ steps.plan.outcome }}
              with:
                script: |
                  const fs = require('fs');
                  const dir = process.env.WORKING_DIR;
                  const marker = `### Terraform plan — \`${dir}\``;
                  let plan = fs.readFileSync(`${dir}/plan.txt`, 'utf8');
                  // GitHub comment hard cap is 65,536 chars — keep the tail (the summary)
                  if (plan.length > 60000) plan = '... (truncated, see job log)\n' + plan.slice(-60000);
                  const status = process.env.PLAN_OUTCOME === 'success' ? '' : '\n> **PLAN FAILED** — see details below.';
                  const body = `${marker}${status}\n<details><summary>Show plan</summary>\n\n\`\`\`hcl\n${plan}\n\`\`\`\n</details>`;
                  const { data: comments } = await github.rest.issues.listComments({
                    ...context.repo, issue_number: context.issue.number });
                  const prev = comments.find(c => c.body.startsWith(marker));
                  if (prev) {
                    await github.rest.issues.updateComment({ ...context.repo, comment_id: prev.id, body });
                  } else {
                    await github.rest.issues.createComment({ ...context.repo, issue_number: context.issue.number, body });
                  }
      
            - name: Fail job if plan failed
              if: steps.plan.outcome == 'failure'
              run: exit 1
      
        # ---------------------------------------------------------------------------
        # Merge to main: fresh plan + apply
        # ---------------------------------------------------------------------------
        apply:
          if: github.event_name == 'push' && github.ref == 'refs/heads/main'
          runs-on: ubuntu-latest
          # The "production" environment is where the human gate lives:
          # repo Settings -> Environments -> production -> required reviewers.
          environment: production
          permissions:
            contents: read
            id-token: write
          defaults:
            run:
              working-directory: ${{ env.WORKING_DIR }}
          steps:
            - uses: actions/checkout@v5
      
            - uses: hashicorp/setup-terraform@v3
              with:
                terraform_version: ${{ env.TF_VERSION }}
      
            # Apply role: write permissions, trust policy locked to
            #   repo:myorg/infra:environment:production
            # so ONLY this gated environment on main can assume it.
            - name: Configure AWS credentials (apply role)
              uses: aws-actions/configure-aws-credentials@v5
              with:
                role-to-assume: arn:aws:iam::123456789012:role/github-terraform-apply
                aws-region: ${{ env.AWS_REGION }}
      
            - name: init
              run: terraform init -input=false -lock-timeout=5m
      
            # Fresh plan at merge time. If the world moved since PR review, this plan
            # differs from the reviewed one — the saved-plan apply below makes that
            # explicit rather than silently applying something new.
            - name: plan
              run: terraform plan -input=false -no-color -lock-timeout=5m -out=tfplan
      
            - name: apply
              # Applying the saved plan (not `apply -auto-approve` on config) means
              # we apply EXACTLY what the step above planned — no second refresh gap.
              run: terraform apply -input=false -no-color tfplan
      
        # ---------------------------------------------------------------------------
        # Nightly drift detection
        # ---------------------------------------------------------------------------
        # Move to its own file with `on: schedule` if you prefer; shown here for
        # completeness. Exit code 2 from -detailed-exitcode means drift.
        # drift:
        #   if: github.event_name == 'schedule'
        #   runs-on: ubuntu-latest
        #   permissions: { contents: read, id-token: write }
        #   steps:
        #     - uses: actions/checkout@v5
        #     - uses: hashicorp/setup-terraform@v3
        #       with: { terraform_version: "1.15.5" }
        #     - uses: aws-actions/configure-aws-credentials@v5
        #       with:
        #         role-to-assume: arn:aws:iam::123456789012:role/github-terraform-plan
        #         aws-region: ap-southeast-2
        #     - run: terraform init -input=false
        #       working-directory: environments/prod
        #     - name: drift check
        #       working-directory: environments/prod
        #       run: |
        #         set +e
        #         terraform plan -detailed-exitcode -lock=false -input=false -no-color
        #         code=$?
        #         [ "$code" -eq 2 ] && { echo "::error::Drift detected in prod"; exit 1; }
        #         exit $code
      
  • references
    • cicd-pipelines.md 9.5 KB
      # CI/CD Pipelines
      
      The contract: **every change is planned on the PR, the plan is visible to the reviewer, and apply happens from CI with short-lived credentials.** Humans never run `apply` against shared environments from laptops.
      
      Full workflow template: [../assets/github-actions-terraform.yml](../assets/github-actions-terraform.yml).
      
      ## Pipeline Shape
      
      ```
                    ┌─ fmt -check ─┐
      PR opened ──> ├─ validate    ├──> tflint ──> trivy/checkov ──> plan ──> plan as PR comment
                    └─ (parallel)  ┘                                              │
                                                                         reviewer reads plan
      PR merged ──> fresh plan ──> apply (OIDC role, environment gate)
      Nightly   ──> plan -detailed-exitcode ──> exit 2 ⇒ drift alert
      ```
      
      Key decisions baked into that shape:
      
      1. **Plan on PR, apply on merge.** The merge re-plans rather than applying the stale PR plan artifact — simpler, and the `concurrency` group serializes applies. If you need apply-exactly-what-was-reviewed, upload `tfplan` as an artifact on the PR and apply that artifact on merge; accept the trade-off that the world may have moved (the apply will fail if so, which is the safe failure).
      2. **One job per root module** (matrix or separate workflows). A monorepo with `environments/{dev,prod}` plans both on PR, applies dev on merge, applies prod behind a GitHub *environment* with required reviewers.
      3. **`concurrency` group per state file** so two merges can't apply concurrently (backend locking would catch it, but failing fast in CI is cleaner).
      
      ## OIDC Cloud Auth — no long-lived keys
      
      This is non-negotiable and matches the repo's supply-chain doctrine (short-lived tokens over standing credentials; a leaked workflow can't exfiltrate what doesn't exist). GitHub mints a signed JWT per job; the cloud trusts GitHub's issuer for *specific repos/branches* and returns temporary credentials.
      
      ### AWS
      
      ```yaml
      permissions:
        id-token: write      # REQUIRED for OIDC
        contents: read
      
      steps:
        - uses: aws-actions/configure-aws-credentials@v5
          with:
            role-to-assume: arn:aws:iam::123456789012:role/github-terraform-plan
            aws-region: ap-southeast-2
      ```
      
      Trust policy on the role — scope it tight:
      
      ```json
      {
        "Effect": "Allow",
        "Principal": { "Federated": "arn:aws:iam::123456789012:oidc-provider/token.actions.githubusercontent.com" },
        "Action": "sts:AssumeRoleWithWebIdentity",
        "Condition": {
          "StringEquals": { "token.actions.githubusercontent.com:aud": "sts.amazonaws.com" },
          "StringLike":   { "token.actions.githubusercontent.com:sub": "repo:myorg/infra:ref:refs/heads/main" }
        }
      }
      ```
      
      - **Two roles**: a read-only `plan` role (assumable from any branch/PR) and a write `apply` role (assumable only from `ref:refs/heads/main` or `environment:prod`). PR plans from forks then physically cannot mutate anything.
      - Audit the trust federation periodically — stale OIDC subjects (deleted repos, renamed branches) with live trust are exactly the entry point the 2026 supply-chain worms abused. `zizmor` catches `pull_request_target` + OIDC misconfigs statically.
      
      ### Azure / GCP equivalents
      
      ```yaml
      # Azure — federated credential on an app registration
      - uses: azure/login@v2
        with:
          client-id: ${{ vars.AZURE_CLIENT_ID }}
          tenant-id: ${{ vars.AZURE_TENANT_ID }}
          subscription-id: ${{ vars.AZURE_SUBSCRIPTION_ID }}
      
      # GCP — workload identity federation
      - uses: google-github-actions/auth@v3
        with:
          workload_identity_provider: projects/123/locations/global/workloadIdentityPools/github/providers/myorg
          service_account: terraform@myproj.iam.gserviceaccount.com
      ```
      
      ## Plan as PR Comment
      
      The reviewer must see the plan without leaving the PR. Minimal recipe (full version in the asset):
      
      ```yaml
      - name: Plan
        id: plan
        run: terraform plan -no-color -input=false -out=tfplan 2>&1 | tee plan.txt
      
      - name: Comment plan on PR
        uses: actions/github-script@v8
        with:
          script: |
            const fs = require('fs');
            const plan = fs.readFileSync('plan.txt', 'utf8').slice(0, 60000); // comment size cap
            const body = `### Terraform plan — \`prod\`\n<details><summary>Show plan</summary>\n\n\`\`\`hcl\n${plan}\n\`\`\`\n</details>`;
            // find-and-update existing comment instead of stacking new ones
            const { data: comments } = await github.rest.issues.listComments({ ...context.repo, issue_number: context.issue.number });
            const prev = comments.find(c => c.body.startsWith('### Terraform plan — `prod`'));
            if (prev) await github.rest.issues.updateComment({ ...context.repo, comment_id: prev.id, body });
            else await github.rest.issues.createComment({ ...context.repo, issue_number: context.issue.number, body });
      ```
      
      Notes:
      
      - **Update-in-place** (find previous comment) or every push spams the PR.
      - Truncate: GitHub comments cap at 65,536 chars. For huge plans link to the job log and post only the resource-change summary (`terraform show -json tfplan | jq -r '.resource_changes[] | "\(.change.actions | join(",")) \(.address)"'`).
      - Plans can leak values — `sensitive = true` redacts in plan output, but data sources and resource attributes are not all marked. Treat the PR comment as visible to everyone with repo read.
      
      ## Policy Gates
      
      | Tool | Layer | What it catches | Invocation |
      |---|---|---|---|
      | `terraform fmt -check -recursive` | style | Unformatted code | exit ≠ 0 fails CI |
      | `terraform validate` | syntax | Type errors, bad references | needs `init` (use `-backend=false` for speed) |
      | `tflint` | lint | Provider-aware errors: invalid instance types, deprecated syntax, unused declarations | `tflint --init && tflint --recursive` |
      | `trivy config .` | security | Misconfig: public buckets, open SGs, unencrypted disks (absorbed tfsec's rule set) | exit codes; SARIF upload for code-scanning UI |
      | `checkov -d .` | security | Same space as trivy; bigger policy library, more noise — pick ONE of trivy/checkov | `--soft-fail` while triaging |
      | `conftest test tfplan.json` | org policy | YOUR rules in Rego/OPA: "no resources without tags", "only approved regions", "no IAM * actions" | run against `terraform show -json tfplan` |
      
      ```hcl
      # .tflint.hcl
      plugin "aws" {
        enabled = true
        version = "0.40.0"
        source  = "github.com/terraform-linters/tflint-ruleset-aws"
      }
      rule "terraform_required_version" { enabled = true }
      rule "terraform_naming_convention" { enabled = true }
      ```
      
      ```rego
      # policy/tags.rego — conftest example against plan JSON
      package main
      deny[msg] {
        rc := input.resource_changes[_]
        rc.change.actions[_] == "create"
        not rc.change.after.tags.Environment
        msg := sprintf("%s: missing required tag 'Environment'", [rc.address])
      }
      ```
      
      Layering guidance: fmt/validate/tflint are table stakes on every PR. trivy *or* checkov as the misconfig gate (start `--soft-fail`, ratchet to hard once the baseline is clean). conftest/OPA only when you have genuinely org-specific rules the scanners can't express — it's the highest-maintenance layer.
      
      ## Workflow Hardening
      
      Same doctrine as the rest of the repo's supply-chain rules:
      
      - **Pin actions to commit SHAs** — `uses: actions/checkout@08c6903cd8c0fde910a37f88322edcfb5dd907a8 # v5.0.0`, not `@v5`. A hijacked tag is a hijacked pipeline with cloud credentials.
      - `permissions:` block at workflow top, least privilege (`contents: read` default; `id-token: write` only on jobs that auth to cloud; `pull-requests: write` only on the comment job).
      - **Never `pull_request_target` with checkout of PR head** for terraform repos — that's RCE-with-secrets for any fork.
      - Plan jobs from forks: read-only role or no cloud auth at all (validate/lint only).
      - `step-security/harden-runner` for egress allow-listing on apply jobs if you want runtime control.
      - Pin the terraform version: `hashicorp/setup-terraform@<sha>` with `terraform_version: 1.15.5` (or `opentofu/setup-opentofu` with `tofu_version`), matching `required_version`.
      
      ## Drift Detection Job
      
      ```yaml
      on:
        schedule: [{ cron: "17 18 * * *" }]   # nightly; odd minute avoids the top-of-hour stampede
      jobs:
        drift:
          permissions: { id-token: write, contents: read }
          steps:
            # ... checkout, setup, OIDC auth (read-only role), init ...
            - name: Detect drift
              run: |
                set +e
                terraform plan -detailed-exitcode -lock=false -input=false -no-color
                code=$?
                if [ "$code" -eq 2 ]; then echo "::error::Drift detected"; exit 1; fi
                exit $code
      ```
      
      Wire the failure to Slack/issue creation. Exit 2 means *either* console drift or merged-but-unapplied config — both are findings.
      
      ## Orchestrator Alternatives
      
      | | Model | Locking | Policy | Cost | Pick when |
      |---|---|---|---|---|---|
      | **GitHub Actions (DIY)** | Workflows you own | `concurrency` groups | tflint/trivy/conftest steps | Free-ish | Default. Full control, template in assets/ |
      | **Atlantis** | Self-hosted server; `atlantis plan` / `atlantis apply` PR comments | Per-directory PR locks (best-in-class) | Custom workflows + conftest | Server you run | Many roots, many PRs, comment-driven culture |
      | **HCP Terraform** | Managed runs + state + UI | Workspace runs serialize | Sentinel / OPA | Free ≤ 500 resources, then $$ | Want managed everything; Sentinel policy; private registry |
      | **Spacelift / env0 / Scalr** | Commercial SaaS orchestrators | Built-in | OPA-based | $$ | Enterprise multi-IaC (Pulumi/CFN too), RBAC needs |
      | **Digger** | Runs inside *your* GitHub Actions | Orchestrated via PR | Pluggable | OSS core | Atlantis UX without hosting a server |
      
      OpenTofu note: Atlantis, Spacelift, env0, Digger all support `tofu` as the binary. HCP Terraform does not run OpenTofu.
      
    • module-patterns.md 9.3 KB
      # Module Patterns
      
      Modules are Terraform's unit of reuse. Most module pain comes from treating them like classes (inheritance, deep nesting, wrapping) instead of functions (flat composition, explicit inputs/outputs).
      
      ## Root vs Child Modules
      
      | | Root module | Child module |
      |---|---|---|
      | What | The directory you run `terraform` in | Anything called via `module` block |
      | Backend block | Yes — exactly one | **Never** |
      | Provider config (`provider "aws" {}`) | Yes | **Never** — declare `required_providers` only |
      | tfvars | Yes | No (inputs come from the caller) |
      | State | Owns one state file | Lives inside the caller's state |
      
      ```hcl
      # modules/network/versions.tf — child module declares NEEDS, not config
      terraform {
        required_version = ">= 1.10"
        required_providers {
          aws = {
            source  = "hashicorp/aws"
            version = ">= 6.0, < 8.0"     # modules use RANGES; roots pin tighter
          }
        }
      }
      ```
      
      A child module with a `provider` block can't be used with `for_each`/`count`/`depends_on` and can't be removed cleanly. Pass aliased providers explicitly when needed:
      
      ```hcl
      module "dns" {
        source    = "../../modules/dns"
        providers = { aws = aws.us_east_1 }   # ACM certs for CloudFront, etc.
      }
      ```
      
      ## Composition Over Inheritance
      
      Build small single-purpose modules; compose them in the root by wiring outputs to inputs.
      
      ```hcl
      # Root composes — dependencies are explicit data flow
      module "network" {
        source = "../../modules/network"
        cidr   = "10.0.0.0/16"
      }
      
      module "database" {
        source     = "../../modules/database"
        subnet_ids = module.network.private_subnet_ids   # output -> input wiring
        vpc_id     = module.network.vpc_id
      }
      
      module "app" {
        source            = "../../modules/app-service"
        subnet_ids        = module.network.private_subnet_ids
        db_connection_arn = module.database.connection_secret_arn
      }
      ```
      
      Rules of thumb:
      
      - **Nesting depth ≤ 2.** Root → module → (occasionally) submodule. Deeper means you're rebuilding inheritance and debugging through four layers of variable plumbing.
      - A module should manage a *cohesive* set of resources with a clear lifecycle (a VPC and its subnets/routes — yes; "everything for the app" — no).
      - If a module takes 40 variables and most callers set 3, split it.
      - Don't create a module for a single resource unless it encodes real policy (e.g. an S3 module that enforces encryption + public-access-block on every bucket — that's policy, not wrapping).
      
      ### Anti-pattern: thin wrapper modules
      
      ```hcl
      # modules/our-vpc/main.tf — adds NOTHING
      module "vpc" {
        source  = "terraform-aws-modules/vpc/aws"
        version = "~> 6.0"
        name    = var.vpc_name        # renamed from `name`. why.
        cidr    = var.network_cidr    # renamed from `cidr`. why.
      }
      ```
      
      Costs: a second version to bump (wrapper lags upstream), a second docs surface, every upstream feature needs a wrapper variable added before anyone can use it, and `moved`-block refactors in upstream don't propagate. **Call the upstream module directly from the root.** A wrapper earns its existence only when it enforces organizational policy (mandatory tags, forced encryption, restricted instance types) — and then it should *say so* in its README.
      
      ## Variable Design
      
      ### Validation — fail at plan, not mid-apply
      
      ```hcl
      variable "environment" {
        type        = string
        description = "Deployment environment."
        validation {
          condition     = contains(["dev", "staging", "prod"], var.environment)
          error_message = "environment must be one of: dev, staging, prod."
        }
      }
      
      variable "cidr" {
        type = string
        validation {
          condition     = can(cidrhost(var.cidr, 0)) && tonumber(split("/", var.cidr)[1]) <= 24
          error_message = "cidr must be a valid IPv4 CIDR no smaller than /24."
        }
      }
      ```
      
      Since TF 1.9, `condition` can reference *other* variables and data — cross-field validation lives on the variable, not buried in a `precondition`.
      
      ### optional() + nullable
      
      ```hcl
      variable "logging" {
        description = "Access-logging config. Omit fields for defaults."
        type = object({
          enabled       = optional(bool, true)
          bucket        = optional(string)          # null when omitted
          prefix        = optional(string, "logs/")
          sample_rate   = optional(number, 1.0)
        })
        default  = {}
        nullable = false      # caller may omit the variable, but may NOT pass logging = null
      }
      ```
      
      - `optional(type, default)` lets callers pass partial objects — the killer feature for config-object variables. Without defaults you'd force every caller to spell out every field.
      - `nullable = false` means an explicit `null` is rejected and the variable's own `default` is used instead. Use it on almost every variable: it converts "caller passed null, module exploded on `var.x.enabled`" into a plan-time error.
      - Gotcha: `optional(string)` with no default yields `null` — guard with `coalesce(...)` or `try(...)` before interpolating.
      
      ### Variable hygiene
      
      | Rule | Why |
      |---|---|
      | Every variable has `description` | It's the module's API doc (`terraform-docs` renders it) |
      | `sensitive = true` on secret inputs | Keeps values out of CLI output (NOT out of state — see security-and-secrets.md) |
      | Prefer typed objects over `map(any)` | `any` defers errors to deep inside the module |
      | No "pass-through everything" variables | A `extra_settings = any` variable is an API you can never change |
      | Defaults = safe choice, not common choice | Default to encrypted/private/protected; make callers opt *out* loudly |
      
      ## Output Contracts
      
      Outputs are the module's public API. Consumers wire them into other modules and remote-state reads — changing one is a breaking change.
      
      ```hcl
      output "vpc_id" {
        description = "ID of the created VPC."
        value       = aws_vpc.main.id
      }
      
      output "private_subnet_ids" {
        description = "Private subnet IDs, keyed by AZ."
        value       = { for k, s in aws_subnet.private : k => s.id }
      }
      
      output "db_endpoint" {
        description = "Writer endpoint. SENSITIVE: contains hostname only, no creds."
        value       = aws_rds_cluster.main.endpoint
      }
      ```
      
      - Output the **identifiers consumers need** (IDs, ARNs, endpoints, security-group IDs) — not whole resource objects (`value = aws_vpc.main` couples consumers to the provider schema and bloats state).
      - `sensitive = true` propagates: an output derived from a sensitive value must itself be marked sensitive or plan errors.
      - Treat output removal/rename like an API break: semver-major the module, or keep the old output as an alias for one cycle.
      
      ## Versioning and Pinning
      
      ```hcl
      # Registry module — minor-float
      module "vpc" {
        source  = "terraform-aws-modules/vpc/aws"
        version = "~> 6.0"        # >= 6.0.0, < 7.0.0
      }
      
      # Git module — pin a TAG (never a branch)
      module "internal" {
        source = "git::https://github.com/myorg/tf-modules.git//network?ref=v2.3.1"
      }
      ```
      
      | Constraint | Meaning | Use for |
      |---|---|---|
      | `~> 6.0` | ≥ 6.0, < 7.0 | Registry modules/providers in shared modules |
      | `~> 6.12.0` | ≥ 6.12.0, < 6.13.0 | Conservative prod roots |
      | `= 6.12.1` / `?ref=v2.3.1` | Exact | Prod roots wanting byte-identical builds |
      | `>= 6.0` (open-ended) | Anything newer | **Never in prod** — a breaking major auto-arrives |
      
      - **Commit `.terraform.lock.hcl`.** It pins exact provider versions + hashes; `terraform init -upgrade` is the deliberate act of moving within constraints. This is the same supply-chain posture as any other lockfile: a pin only protects you if it pre-dates a compromise and you don't run unconstrained upgrades in CI.
      - Run `terraform providers lock -platform=linux_amd64 -platform=darwin_arm64 -platform=windows_amd64` so the lockfile carries hashes for every platform your team + CI uses (OpenTofu 1.12 does this automatically at init).
      - Internal module registries: HCP Terraform private registry, or plain git tags + a `modules/` monorepo. Git tags are fine; the registry's value is the version-constraint syntax and docs rendering.
      
      ## Module Repo Layout
      
      ```
      terraform-aws-network/            # one module per repo (registry-publishable), or modules/ monorepo
      ├── main.tf
      ├── variables.tf
      ├── outputs.tf
      ├── versions.tf
      ├── README.md                     # terraform-docs generated
      ├── examples/
      │   └── complete/                 # a runnable root that exercises the module — doubles as docs + test fixture
      │       └── main.tf
      └── tests/
          └── network.tftest.hcl        # native tests (see SKILL.md Testing section)
      ```
      
      `terraform-docs markdown table . > README.md` keeps docs honest — wire it as a pre-commit hook or CI check.
      
      ## count vs for_each (the index-shift footgun)
      
      ```hcl
      # BAD: count over a list
      resource "aws_subnet" "private" {
        count      = length(var.subnet_cidrs)        # remove element 0 ->
        cidr_block = var.subnet_cidrs[count.index]   # every subnet re-addresses -> destroy cascade
      }
      
      # GOOD: for_each over a map with stable keys
      resource "aws_subnet" "private" {
        for_each   = var.subnets                      # { "ap-southeast-2a" = "10.0.1.0/24", ... }
        cidr_block = each.value
        availability_zone = each.key
      }
      ```
      
      `count` is fine for "0 or 1 of this" conditionals (`count = var.enabled ? 1 : 0`) — though even there, `for_each = var.enabled ? { main = true } : {}` keeps addresses stable if it might ever become "n of this". Migrating existing `count` resources to `for_each`: write `moved` blocks for each index→key pair (see state-management.md) so nothing is destroyed.
      
    • security-and-secrets.md 8.3 KB
      # Security and Secrets
      
      The uncomfortable truth first: **Terraform state stores resource attributes in plaintext JSON** — including any password, key, token, or connection string a provider ever returned. `sensitive = true` changes what's *printed*, not what's *stored*. Every secrets strategy below is a variation on "make sure the secret never enters state, and harden state anyway."
      
      ## Threat Model
      
      | Surface | Exposure | Mitigation |
      |---|---|---|
      | State file | Plaintext attributes of every resource | Encrypted backend + IAM; OpenTofu state encryption; keep secrets out entirely (below) |
      | Plan files (`tfplan`, JSON) | Contain proposed values, incl. some sensitive ones | Treat plan artifacts as secrets; short retention on CI artifacts |
      | CLI / CI logs | Values interpolated into output | `sensitive = true`, `-no-color` log review, masked CI vars |
      | PR plan comments | Anyone with repo read | Summarize rather than full-dump for sensitive roots |
      | `.tfvars` in git | Whatever you put there | Never commit secret tfvars; SOPS-encrypt or env-var inject |
      | Provider credentials | Long-lived keys in CI secrets | OIDC short-lived tokens (see cicd-pipelines.md) |
      
      ## What `sensitive = true` Actually Does (and doesn't)
      
      ```hcl
      variable "db_password" {
        type      = string
        sensitive = true        # plan/apply output prints (sensitive value)
      }
      
      output "endpoint" {
        value     = "${aws_db_instance.main.address}:${var.db_password}"  # ERROR unless...
        sensitive = true        # ...the output is marked too (sensitivity propagates)
      }
      ```
      
      Does: redact from `plan`/`apply`/`output` human output; propagate taint through expressions; force derived outputs to be marked.
      Does **not**: encrypt anything; remove the value from state (`terraform state pull | jq` shows it plaintext); redact from `terraform output -json` (explicitly prints sensitive values); stop a provider logging it at TRACE.
      
      `ephemeral = true` on variables (TF ≥ 1.10) goes further: the value may only flow into ephemeral contexts (write-only args, provider config, locals marked ephemeral) and is never written to state or plan.
      
      ## Ephemeral Resources (TF ≥ 1.10, OpenTofu ≥ 1.11)
      
      Ephemeral resources open/fetch a value during the run and are **never persisted to state or plan**. The first-class pattern for "read a secret from a manager at apply time."
      
      ```hcl
      # Fetch the secret ephemerally — exists only for the duration of the run
      ephemeral "aws_secretsmanager_secret_version" "db" {
        secret_id = aws_secretsmanager_secret.db.id
      }
      
      # Use it via a write-only argument so it never lands in state either
      resource "aws_db_instance" "main" {
        # ...
        password_wo         = ephemeral.aws_secretsmanager_secret_version.db.secret_string
        password_wo_version = 1
      }
      ```
      
      Available ephemeral resource types include `aws_secretsmanager_secret_version`, `aws_ssm_parameter`, `azurerm_key_vault_secret`, `google_secret_manager_secret_version`, Vault's `vault_kv_secret_v2`, and `random_password` (ephemeral variant). Ephemeral *values* can feed: write-only arguments, provider configuration, `terraform_data` triggers — anything that itself persists will error if you try.
      
      Contrast with the classic `data "aws_secretsmanager_secret_version"` — a data source's result **is stored in state**, which quietly copied your secret into the state file. Migrate those reads to `ephemeral` blocks.
      
      ## Write-Only Arguments (TF ≥ 1.11, OpenTofu ≥ 1.11)
      
      Provider-defined `*_wo` arguments accept a value, hand it to the API, and store **nothing** in state. Because nothing is stored, Terraform can't diff them — that's what the paired `*_wo_version` integer is for:
      
      ```hcl
      resource "aws_db_instance" "main" {
        password_wo         = var.db_password          # never in state
        password_wo_version = 2                        # bump to push a new password
      }
      ```
      
      - Rotate by incrementing `_wo_version` (or wiring it to a rotation timestamp/secret version).
      - Write-only args accept ephemeral values — the combination above (ephemeral read → write-only write) is the zero-secrets-in-state gold standard.
      - Caveat: not every provider/resource has `_wo` variants yet; check the provider docs. Where absent, fall back to the secret-manager-reference pattern below.
      
      ## Pattern Ladder (best to worst)
      
      ```
      1. Secret never touches Terraform at all
         App reads from Secrets Manager/Vault at BOOT using its IAM role;
         Terraform only creates the (empty or rotated-out-of-band) secret container + IAM.
      
      2. Ephemeral read -> write-only write          (TF >= 1.11)
         Secret transits the run in memory only. Nothing in state or plan.
      
      3. Secret manager reference via data source    (any version)
         data "aws_secretsmanager_secret_version" -- secret IS in state,
         but at least it's centrally rotated + audited. Encrypt state, restrict IAM.
      
      4. TF_VAR_ env injection from CI secret store  (any version)
         Keeps secrets out of git; still lands in state if assigned to a resource attribute.
      
      5. SOPS-encrypted tfvars in git
         Good at-rest story for git; same state caveat as 4.
      
      6. Plaintext in tfvars/locals committed to git
         Never. Rotating means rewriting git history.
      ```
      
      ### Vault / cloud secret manager integration
      
      ```hcl
      # Terraform creates the container + access policy; VALUE is set out-of-band or by rotation lambda
      resource "aws_secretsmanager_secret" "db" {
        name       = "prod/db/password"
        kms_key_id = aws_kms_key.secrets.arn
      }
      
      resource "aws_iam_role_policy" "app_reads_secret" {
        role   = aws_iam_role.app.id
        policy = jsonencode({
          Version = "2012-10-17"
          Statement = [{ Effect = "Allow", Action = "secretsmanager:GetSecretValue",
                         Resource = aws_secretsmanager_secret.db.arn }]
        })
      }
      # App fetches the secret at startup. Terraform never knows the value.
      ```
      
      Vault: prefer short-TTL dynamic secrets (`vault_database_secret_backend_role`) so even a leaked credential expires in minutes. The Vault provider's classic data sources persist to state — use the ephemeral variants on TF ≥ 1.10.
      
      ### SOPS pattern
      
      ```bash
      # Encrypt env tfvars with a KMS key; ciphertext is committable
      sops --encrypt --kms arn:aws:kms:...:key/... prod.tfvars.json > prod.sops.tfvars.json
      ```
      
      ```hcl
      # Via carlpett/sops provider
      data "sops_file" "secrets" { source_file = "prod.sops.tfvars.json" }
      locals { db_password = data.sops_file.secrets.data["db_password"] }
      ```
      
      Honest accounting: decrypted values still flow into plan/state unless they terminate in write-only arguments. SOPS solves *git at-rest*, not *state at-rest*.
      
      ## Hardening State Itself
      
      Do all of this regardless of which pattern above you use:
      
      - Backend encryption: S3 SSE-KMS with a dedicated CMK; bucket policy denying un-encrypted puts; `azurerm`/`gcs` with CMK.
      - IAM: state bucket readable only by the CI roles + break-glass group. State read access ≈ secret read access — treat the grant accordingly.
      - Versioning on (recovery) + access logging on the bucket (audit).
      - `.gitignore`: `*.tfstate`, `*.tfstate.*`, `*.tfplan`, `.terraform/`, and crash logs (`crash.log` can embed values).
      - **OpenTofu state encryption** (≥ 1.7) — client-side AES-GCM over state *and* plan files, key from PBKDF2 passphrase, AWS/GCP KMS, Azure Key Vault, or external program. The strongest state story available; Terraform has no equivalent (config sample in state-management.md). Plan key rotation: `encryption` supports a `fallback` method so old state remains readable during rotation.
      
      ## Provider Credential Hygiene
      
      | Don't | Do |
      |---|---|
      | `provider "aws" { access_key = "..." }` hardcoded | Ambient auth: OIDC in CI, SSO/instance profiles locally |
      | Long-lived `AWS_ACCESS_KEY_ID` in CI secrets | OIDC `role-to-assume` (see cicd-pipelines.md) |
      | One god-role for plan and apply | Read-only plan role; write apply role gated to main/environment |
      | Shared human credentials for break-glass | Named identities + audited assume-role |
      
      ## Scanning and Gates
      
      - `trivy config .` / `checkov -d .` catch *misconfigurations* (public buckets, `0.0.0.0/0` ingress, unencrypted volumes) — wire into PR CI (see cicd-pipelines.md).
      - `gitleaks` / push-gates catch secrets *in the repo* — tfvars are a classic leak vector.
      - `terraform providers` + lockfile review on provider bumps: providers execute arbitrary code on your CI runner with cloud credentials. A provider is a dependency — the repo's supply-chain rules (cooldown, behavioural scan before adopting unfamiliar providers from the registry) apply in full.
      
    • state-management.md 9.3 KB
      # State Management
      
      Terraform/OpenTofu state is the mapping between config addresses and real infrastructure. It is the single most dangerous file in the project: lose it and Terraform forgets your infra; corrupt it and applies destroy the wrong things; leak it and every secret a provider ever returned is exposed.
      
      ## Remote Backends
      
      Never keep state local for anything shared or production. Remote backends give durability, locking, and team access.
      
      ### S3 (AWS) — current recommended shape
      
      ```hcl
      terraform {
        backend "s3" {
          bucket       = "myorg-tfstate"
          key          = "prod/network/terraform.tfstate"   # one key per root module
          region       = "ap-southeast-2"
          encrypt      = true                # SSE; pair with bucket-default SSE-KMS
          use_lockfile = true               # S3-native locking (TF >= 1.10)
        }
      }
      ```
      
      - **`use_lockfile = true` replaces the DynamoDB lock table** (Terraform ≥ 1.10, OpenTofu ≥ 1.10). It uses S3 conditional writes to create a `.tflock` object next to the state key. The old `dynamodb_table` argument still works and was the standard for a decade — you'll see it everywhere — but new setups don't need the extra table. During migration you can set both; remove `dynamodb_table` once all collaborators are ≥ 1.10.
      - Bucket hygiene: versioning **on** (state history = your undo), default SSE-KMS, block public access, lifecycle rule to expire old noncurrent versions after ~90 days, bucket policy restricting to the CI role + break-glass humans.
      - One bucket per org/account is fine; isolation comes from `key` prefixes + IAM conditions on the prefix.
      
      ### azurerm
      
      ```hcl
      terraform {
        backend "azurerm" {
          resource_group_name  = "rg-tfstate"
          storage_account_name = "myorgtfstate"
          container_name       = "tfstate"
          key                  = "prod.network.tfstate"
          use_azuread_auth     = true        # RBAC instead of storage keys
        }
      }
      ```
      
      Locking is native via blob leases — nothing extra to configure. Prefer `use_azuread_auth = true` so CI uses OIDC-federated identity, not storage account keys.
      
      ### gcs
      
      ```hcl
      terraform {
        backend "gcs" {
          bucket = "myorg-tfstate"
          prefix = "prod/network"
        }
      }
      ```
      
      Locking native via Cloud Storage generation preconditions. Enable object versioning on the bucket. Use workload identity federation in CI.
      
      ### HCP Terraform / Terraform Cloud
      
      ```hcl
      terraform {
        cloud {
          organization = "myorg"
          workspaces { name = "prod-network" }
        }
      }
      ```
      
      State hosted, encrypted, versioned, locked by the platform; runs can execute remotely. Free tier covers ≤ 500 managed resources. Note: OpenTofu cannot use HCP Terraform as a backend (it can use the generic `remote` backend against compatible APIs).
      
      ### Backend selection
      
      | Backend | Locking | Encryption at rest | Best when |
      |---|---|---|---|
      | `s3` + `use_lockfile` | S3 conditional writes | SSE-KMS | AWS shops (default) |
      | `s3` + `dynamodb_table` | DynamoDB | SSE-KMS | Legacy / mixed TF < 1.10 teams |
      | `azurerm` | Blob lease (built-in) | Platform + CMK | Azure shops |
      | `gcs` | Generation precondition (built-in) | Platform + CMEK | GCP shops |
      | HCP Terraform | Platform | Platform | Want managed runs/policies too |
      | `pg` / `consul` / `kubernetes` | Yes | Varies | Niche; self-hosted constraints |
      
      ### Migrating backends
      
      ```bash
      # 1. Add/replace the backend block, then:
      terraform init -migrate-state          # copies state old -> new, prompts
      # 2. Verify: terraform state list shows everything
      # 3. Delete the old state only after a clean plan
      ```
      
      ## State Locking
      
      Locking prevents two concurrent applies corrupting state. It is **not** optional for teams.
      
      - `terraform apply` acquires the lock automatically; a crash can leave it stuck.
      - `terraform force-unlock <LOCK_ID>` — only after confirming no run is actually live (check CI). The lock ID is printed in the error.
      - `-lock-timeout=5m` in CI lets queued runs wait instead of failing instantly.
      - Locking protects against concurrent *writes*; it does not serialize plans — two PRs can both plan green and conflict at apply. Solve at the orchestration layer (Atlantis dir-locks, Actions concurrency groups — see cicd-pipelines.md).
      
      ## Declarative State Changes: moved / import / removed
      
      These blocks are the modern, code-reviewable replacements for CLI state surgery. They live in config, show up in diffs, and execute as part of a normal plan/apply.
      
      ### `moved` — refactor without destroy
      
      ```hcl
      # Renamed a resource
      moved {
        from = aws_instance.web
        to   = aws_instance.frontend
      }
      
      # Moved a resource into a module
      moved {
        from = aws_vpc.main
        to   = module.network.aws_vpc.main
      }
      
      # count -> for_each migration
      moved {
        from = aws_subnet.private[0]
        to   = aws_subnet.private["ap-southeast-2a"]
      }
      ```
      
      Plan shows `# aws_instance.web has moved to aws_instance.frontend` instead of destroy+create. Keep `moved` blocks around for at least one release cycle of a shared module so downstream consumers also get the move; then prune.
      
      ### `import` — adopt existing infrastructure (TF ≥ 1.5)
      
      ```hcl
      import {
        to = aws_s3_bucket.legacy
        id = "myorg-legacy-bucket"
      }
      ```
      
      ```bash
      # Generate matching config for resources you haven't written yet:
      terraform plan -generate-config-out=generated.tf
      # Review generated.tf, clean it up (it's verbose), move into proper files, apply.
      ```
      
      Why blocks beat `terraform import` CLI: the import is planned (you see exactly what will be adopted and whether config matches reality) and reviewed in the PR. The CLI command mutates state immediately with no plan. Import blocks also support `for_each` (TF ≥ 1.7) for bulk adoption.
      
      After a successful apply, delete the `import` blocks — they're one-shot.
      
      ### `removed` — forget without destroying (TF ≥ 1.7, OpenTofu ≥ 1.7)
      
      ```hcl
      removed {
        from = aws_db_instance.legacy
        lifecycle {
          destroy = false      # remove from state, leave the real DB alone
        }
      }
      ```
      
      Use when handing a resource to another team/state, or un-managing something Terraform should no longer own. The declarative version of `terraform state rm`. OpenTofu 1.12 additionally allows `lifecycle { destroy = false }` directly on a resource being deleted from config.
      
      ## State Surgery (CLI) — and when NOT to
      
      ```bash
      terraform state list                      # enumerate addresses (safe)
      terraform state show aws_vpc.main         # inspect one resource (safe)
      terraform state pull > backup.tfstate     # ALWAYS do this first
      terraform state mv aws_x.a aws_x.b        # immediate, unreviewed rename
      terraform state mv aws_x.a 'module.m.aws_x.a'
      terraform state rm aws_x.a                # forget (does NOT destroy)
      terraform state push fixed.tfstate        # overwrite remote state (extreme)
      ```
      
      **Decision rule:** if a `moved`, `removed`, or `import` block can express the change — use the block. Reasons:
      
      1. Blocks are planned and reviewed; CLI mutations are instant and invisible to reviewers.
      2. Blocks are idempotent across the team; a CLI command run by one person leaves everyone else's mental model stale.
      3. A typo in `state mv` orphans a real resource: Terraform now plans to *create* a duplicate while the original drifts unmanaged.
      
      **Legitimate CLI surgery:**
      
      - Splitting one state into two roots (`state mv -state-out=...` or pull/edit/push between backends).
      - Recovering from a half-failed migration or a provider bug that wedged an address.
      - Anything on Terraform < 1.5 (no blocks available).
      
      **Protocol for any surgery:** `state pull` a timestamped backup → make the change → `terraform plan` must come back clean (or exactly the expected diff) → only then walk away.
      
      ## Drift Detection
      
      Infrastructure changes outside Terraform (console edits, autoscaling, other tooling). Detect it before it bites an apply.
      
      ```bash
      terraform plan -detailed-exitcode -lock=false
      # exit 0 -> no changes (clean)
      # exit 1 -> error
      # exit 2 -> changes pending (drift OR un-applied config)
      ```
      
      - `-lock=false` so a read-only drift check never blocks a real apply.
      - `terraform plan -refresh-only` shows only *state vs reality* differences without proposing config-driven changes — cleaner signal for "who touched the console".
      - Cron it: nightly scheduled CI job, alert on exit 2 (recipe in cicd-pipelines.md).
      - Chronic drift on specific attributes → either codify the external process or `lifecycle { ignore_changes = [...] }` deliberately (document why).
      
      ## State File Hygiene
      
      | Rule | Why |
      |---|---|
      | `*.tfstate*` in `.gitignore` | Local state in git = secrets in git history forever |
      | One state per blast-radius unit | Network / data / app split — a bad apply can't take everything |
      | Keep states small (< ~100 resources guideline) | Plan time, lock contention, blast radius all scale with state size |
      | Versioned backend bucket | `state push` mistakes become a revert, not a rebuild |
      | Treat state as secret | Provider attributes (DB passwords, certs) sit in plaintext JSON — see security-and-secrets.md |
      | OpenTofu: consider state encryption | Client-side AES-GCM, keys via PBKDF2/KMS — defence even if the bucket leaks |
      
      ```hcl
      # OpenTofu >= 1.7 only — state + plan encryption
      terraform {
        encryption {
          key_provider "aws_kms" "main" {
            kms_key_id = "arn:aws:kms:...:key/..."
            key_spec   = "AES_256"
          }
          method "aes_gcm" "main" {
            keys = key_provider.aws_kms.main
          }
          state { method = method.aes_gcm.main }
          plan  { method = method.aes_gcm.main }
        }
      }
      ```
      
  • scripts
    • .gitkeep 0 B · in bundle
    • check-action-refs.sh 11.8 KB
      #!/usr/bin/env bash
      # Staleness verifier for GitHub Actions `uses:` references in workflow YAML.
      #
      # Lints every `uses: owner/repo@ref` line in one or more workflow files. General-
      # purpose: not terraform-specific — point it at any GitHub Actions workflow. Two
      # modes per the staleness-verifier pattern (SKILL-RESOURCE-PROTOCOL §7):
      #   --offline (default): structural-only, NO network. Asserts every `uses:` is
      #                        well-formed; floating @main/@master refs are a WARNING.
      #   --live:              resolves every owner/repo@ref against the GitHub API.
      #                        A 404 (ref does not exist) is DRIFT; rate-limit/offline
      #                        is UNAVAILABLE (advisory, never a build failure).
      #
      # Usage:   check-action-refs.sh [--offline|--live] [--strict] [--json] [-q] [FILE ...]
      # Input:   workflow YAML paths as positionals (default: the skill's own
      #          assets/github-actions-terraform.yml). Pure grep/sed extraction — no
      #          YAML library dependency.
      # Output:  stdout = data only (findings list, or JSON envelope with --json)
      # Stderr:  headers, progress, warnings, errors
      # Exit:    0 ok, 2 usage, 3 not-found, 4 malformed-uses, 5 missing-dep,
      #          7 api-unavailable (live), 10 drift (live: a ref does not resolve)
      #
      # Examples:
      #   check-action-refs.sh --offline
      #   check-action-refs.sh --offline .github/workflows/ci.yml
      #   GITHUB_TOKEN=ghp_xxx check-action-refs.sh --live ci.yml deploy.yml
      #   check-action-refs.sh --json --offline | jq '.data[] | select(.status!="ok")'
      
      set -uo pipefail
      
      EXIT_OK=0; EXIT_USAGE=2; EXIT_NOT_FOUND=3; EXIT_MALFORMED=4
      EXIT_MISSING_DEP=5; EXIT_UNAVAILABLE=7; EXIT_DRIFT=10
      
      # Terminal design system (skills/_lib/term.sh). Framing rides stderr (term_init 2);
      # the findings list / --json stay plain on stdout. Degrade if the lib is gone.
      __lib="$(cd "$(dirname "${BASH_SOURCE[0]}")/../../_lib" 2>/dev/null && pwd || true)"
      if [ -n "${__lib:-}" ] && [ -f "$__lib/term.sh" ]; then . "$__lib/term.sh"; term_init 2; __HAVE_TERM=1
      else __HAVE_TERM=0; TERM_DOT="|"; fi
      
      SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
      DEFAULT_FILE="${SCRIPT_DIR}/../assets/github-actions-terraform.yml"
      
      MODE="offline"; STRICT=0; JSON=0; QUIET=0; FILES=()
      while [[ $# -gt 0 ]]; do
        case "$1" in
          --offline)   MODE="offline" ;;
          --live)      MODE="live" ;;
          --strict)    STRICT=1 ;;
          --json)      JSON=1 ;;
          -q|--quiet)  QUIET=1 ;;
          -h|--help)   sed -n '2,26p' "$0" | sed 's/^# \{0,1\}//'; exit "$EXIT_OK" ;;
          -*)  echo "ERROR: unknown flag: $1 (try --help)" >&2; exit "$EXIT_USAGE" ;;
          *)   FILES+=("$1") ;;
        esac
        shift
      done
      
      [[ ${#FILES[@]} -eq 0 ]] && FILES=("$DEFAULT_FILE")
      
      command -v grep >/dev/null 2>&1 || { echo "ERROR: grep required" >&2; exit "$EXIT_MISSING_DEP"; }
      HAS_JQ=0; command -v jq >/dev/null 2>&1 && HAS_JQ=1
      [[ "$JSON" -eq 1 && "$HAS_JQ" -eq 0 ]] && {
        echo '{"error":{"code":"PRECONDITION","message":"jq required for --json"}}'
        echo "ERROR: jq required for --json" >&2; exit "$EXIT_MISSING_DEP"; }
      if [[ "$MODE" == "live" ]]; then
        command -v curl >/dev/null 2>&1 || { echo "ERROR: curl required for --live" >&2; exit "$EXIT_MISSING_DEP"; }
      fi
      
      emit() { [[ "$QUIET" -eq 1 ]] && return; printf '%s\n' "$1" >&2; }
      
      # Panel framing applies to the human stderr stream when it's a TTY (or FORCE_COLOR
      # forces a render); piped/quiet consumers keep the legacy "=== / [TAG]" lines.
      PANEL=0
      if [[ "$__HAVE_TERM" -eq 1 && "$QUIET" -eq 0 ]] && { [ -t 2 ] || [ -n "${FORCE_COLOR:-}" ]; }; then PANEL=1; fi
      __PANEL_OPEN=0
      popen() {
        [[ "$PANEL" -eq 1 && "$__PANEL_OPEN" -eq 0 ]] || return 0
        { term_panel_open terraform "action-refs ${TERM_DOT} ${MODE}"; term_panel_vert; } >&2
        __PANEL_OPEN=1
      }
      # prow <mark> <legacy-prefix> <text> — panel status row, or the legacy tagged line.
      prow() {
        if [[ "$PANEL" -eq 1 ]]; then popen; term_status_row "$1" "$3" >&2
        else emit "  $2 $3"; fi
      }
      
      # State accumulators
      malformed=0; drift=0; unavailable=0; warned=0
      declare -a JSON_OBJS=()
      declare -a TEXT_ROWS=()
      
      # --- classify a `uses:` value -------------------------------------------------
      # Sets globals: C_STATUS (ok|warn|malformed), C_OWNER, C_REPO, C_REF, C_KIND
      classify_uses() {
        local v=$1
        C_OWNER=""; C_REPO=""; C_REF=""; C_KIND=""; C_STATUS="ok"
        # Local action: ./path  — always valid, no network
        if [[ "$v" == ./* ]]; then C_KIND="local"; return; fi
        # Docker action: docker://image[:tag]  — out of scope for ref resolution
        if [[ "$v" == docker://* ]]; then C_KIND="docker"; return; fi
        # Must contain an @ separating owner/repo[/path] from ref
        if [[ "$v" != *"@"* ]]; then C_STATUS="malformed"; C_KIND="action"; return; fi
        local path="${v%@*}" ref="${v#*@}"
        C_REF="$ref"
        # Empty ref, or empty path
        if [[ -z "$ref" || -z "$path" ]]; then C_STATUS="malformed"; C_KIND="action"; return; fi
        # path must be owner/repo[/subpath...]
        if [[ "$path" != */* ]]; then C_STATUS="malformed"; C_KIND="action"; return; fi
        C_OWNER="${path%%/*}"
        local rest="${path#*/}"
        C_REPO="${rest%%/*}"
        C_KIND="action"
        if [[ -z "$C_OWNER" || -z "$C_REPO" ]]; then C_STATUS="malformed"; return; fi
        # Validate ref shape: tag (vN / vN.N.N...), 40-hex SHA, or branch name.
        if [[ "$ref" =~ ^[0-9a-f]{40}$ ]]; then
          C_STATUS="ok"          # pinned SHA — best practice
        elif [[ "$ref" == "main" || "$ref" == "master" ]]; then
          C_STATUS="warn"        # floating default branch — advisory
        elif [[ "$ref" =~ ^v[0-9]+([._][0-9]+)*$ ]]; then
          C_STATUS="ok"          # version tag
        elif [[ "$ref" =~ ^[A-Za-z0-9._/-]+$ ]]; then
          C_STATUS="ok"          # other tag/branch name — structurally fine
        else
          C_STATUS="malformed"
        fi
      }
      
      # --- live resolution: does owner/repo@ref exist on GitHub? --------------------
      # Echoes one of: resolved | notfound | unavailable
      resolve_ref() {
        local owner=$1 repo=$2 ref=$3
        local -a auth=()
        [[ -n "${GITHUB_TOKEN:-}" ]] && auth=(-H "Authorization: Bearer ${GITHUB_TOKEN}")
        local base="https://api.github.com/repos/${owner}/${repo}"
        local code body url
      
        try() {
          local u=$1 out
          out=$(curl -sS -w $'\n%{http_code}' \
            -H "Accept: application/vnd.github+json" \
            -H "X-GitHub-Api-Version: 2022-11-28" \
            "${auth[@]}" "$u" 2>/dev/null)
          [[ -z "$out" ]] && { echo "000"; return; }
          code="${out##*$'\n'}"; body="${out%$'\n'*}"
          echo "$code"
        }
      
        if [[ "$ref" =~ ^[0-9a-f]{40}$ ]]; then
          url="${base}/commits/${ref}"
        else
          url="${base}/git/ref/tags/${ref}"
        fi
        code=$(try "$url")
        # Network failure
        [[ "$code" == "000" ]] && { echo "unavailable"; return; }
        # Rate limit — advisory, never drift
        if [[ "$code" == "403" || "$code" == "429" ]]; then echo "unavailable"; return; fi
        if [[ "$code" == "200" ]]; then echo "resolved"; return; fi
        # Tag lookup 404'd — fall back to a commit/branch lookup before declaring drift.
        # A non-SHA ref can still be a branch name; the commits endpoint resolves those.
        if [[ "$code" == "404" && ! "$ref" =~ ^[0-9a-f]{40}$ ]]; then
          code=$(try "${base}/commits/${ref}")
          [[ "$code" == "000" ]] && { echo "unavailable"; return; }
          if [[ "$code" == "403" || "$code" == "429" ]]; then echo "unavailable"; return; fi
          [[ "$code" == "200" ]] && { echo "resolved"; return; }
          # 404 (no such branch) or 422 (unparseable commit-ish) — the ref does not exist
          [[ "$code" == "404" || "$code" == "422" ]] && { echo "notfound"; return; }
          echo "unavailable"; return
        fi
        # Direct SHA lookup that 404/422'd, or a tag 404 with no branch fallback path
        [[ "$code" == "404" || "$code" == "422" ]] && { echo "notfound"; return; }
        echo "unavailable"
      }
      
      add_json() {  # file line ref status
        [[ "$HAS_JQ" -eq 1 ]] || return
        JSON_OBJS+=("$(jq -cn --arg f "$1" --argjson l "$2" --arg r "$3" --arg s "$4" \
          '{file:$f, line:$l, ref:$r, status:$s}')")
      }
      
      if [[ "$PANEL" -eq 1 ]]; then popen; else emit "=== check-action-refs (${MODE}) ==="; fi
      
      for f in "${FILES[@]}"; do
        if [[ ! -f "$f" ]]; then
          prow bad "ERROR:" "file not found: $f"
          # In JSON mode still report a structured error per §5
          if [[ "$JSON" -eq 1 ]]; then
            echo "{\"error\":{\"code\":\"NOT_FOUND\",\"message\":\"file not found: $f\"}}"
          fi
          exit "$EXIT_NOT_FOUND"
        fi
      
        # Extract every `uses:` value with its 1-based line number. Strip inline
        # comments and surrounding quotes. grep -n gives "LINE:content".
        while IFS= read -r entry; do
          [[ -z "$entry" ]] && continue
          lineno="${entry%%:*}"
          rawline="${entry#*:}"
          # value after `uses:` — drop leading `- ` list dash if present
          val="${rawline#*uses:}"
          val="${val%%#*}"                       # strip trailing comment
          # trim whitespace
          val="${val#"${val%%[![:space:]]*}"}"
          val="${val%"${val##*[![:space:]]}"}"
          # strip surrounding quotes
          val="${val%\"}"; val="${val#\"}"
          val="${val%\'}"; val="${val#\'}"
          [[ -z "$val" ]] && continue
      
          classify_uses "$val"
      
          case "$C_STATUS" in
            malformed)
              malformed=1
              prow bad "[MALFORMED]" "${f}:${lineno}  uses: ${val}"
              TEXT_ROWS+=("${f}:${lineno}	${val}	malformed")
              add_json "$f" "$lineno" "$val" "malformed"
              ;;
            warn)
              warned=1
              prow warn "[WARN floating]" "${f}:${lineno}  ${val}  (prefer SHA pin)"
              TEXT_ROWS+=("${f}:${lineno}	${val}	warn")
              add_json "$f" "$lineno" "$val" "warn"
              ;;
            ok)
              if [[ "$MODE" == "live" && "$C_KIND" == "action" ]]; then
                res=$(resolve_ref "$C_OWNER" "$C_REPO" "$C_REF")
                case "$res" in
                  resolved)
                    prow ok "[ok]" "${f}:${lineno}  ${C_OWNER}/${C_REPO}@${C_REF}"
                    TEXT_ROWS+=("${f}:${lineno}	${val}	ok")
                    add_json "$f" "$lineno" "$val" "ok" ;;
                  notfound)
                    drift=1
                    prow bad "[DRIFT 404]" "${f}:${lineno}  ${C_OWNER}/${C_REPO}@${C_REF}"
                    TEXT_ROWS+=("${f}:${lineno}	${val}	drift")
                    add_json "$f" "$lineno" "$val" "drift" ;;
                  unavailable)
                    unavailable=1
                    prow warn "[unavailable]" "${f}:${lineno}  ${C_OWNER}/${C_REPO}@${C_REF} (API unreachable/rate-limited)"
                    TEXT_ROWS+=("${f}:${lineno}	${val}	unavailable")
                    add_json "$f" "$lineno" "$val" "unavailable" ;;
                esac
              else
                TEXT_ROWS+=("${f}:${lineno}	${val}	ok")
                add_json "$f" "$lineno" "$val" "ok"
              fi
              ;;
          esac
        done < <(grep -nE '^[[:space:]]*-?[[:space:]]*uses:[[:space:]]*' "$f" 2>/dev/null)
      done
      
      # --- panel footer (stderr framing only) ---------------------------------------
      if [[ "$PANEL" -eq 1 && "$__PANEL_OPEN" -eq 1 ]]; then
        ph_state="healthy"; ph_text="refs well-formed"
        if [[ "$malformed" -eq 1 || "$drift" -eq 1 ]]; then ph_state="critical"; ph_text="findings present"
        elif [[ "$unavailable" -eq 1 ]]; then ph_state="warning"; ph_text="api unavailable"
        elif [[ "$warned" -eq 1 ]]; then ph_state="warning"; ph_text="floating refs"; fi
        { term_panel_vert
          term_panel_close "--live to resolve ${TERM_DOT} --json for data" "$(term_health "$ph_state" "$ph_text")"
        } >&2
      fi
      
      # --- output -------------------------------------------------------------------
      if [[ "$JSON" -eq 1 ]]; then
        printf '%s\n' "${JSON_OBJS[@]:-}" | jq -s \
          --arg mode "$MODE" \
          '{data: map(select(length>0)),
            meta: {mode:$mode,
                   count:(map(select(length>0))|length),
                   schema:"claude-mods.terraform-ops.action-refs/v1"}}'
      else
        # plain text: data rows to stdout (only the non-ok findings are interesting,
        # but emit all rows so the agent sees the full inventory)
        for row in "${TEXT_ROWS[@]:-}"; do
          [[ -n "$row" ]] && printf '%s\n' "$row"
        done
      fi
      
      # --- exit ---------------------------------------------------------------------
      [[ "$malformed" -eq 1 ]] && exit "$EXIT_MALFORMED"
      [[ "$drift" -eq 1 ]] && exit "$EXIT_DRIFT"
      [[ "$unavailable" -eq 1 ]] && exit "$EXIT_UNAVAILABLE"
      if [[ "$warned" -eq 1 && "$STRICT" -eq 1 ]]; then exit "$EXIT_MALFORMED"; fi
      exit "$EXIT_OK"
      
  • tests
    • run.sh 3.5 KB
      #!/usr/bin/env bash
      # Self-test for terraform-ops — fully offline: no network, no GitHub API.
      #
      # Wraps the skill's §7 staleness verifier (scripts/check-action-refs.sh), which
      # lints `uses: owner/repo@ref` lines in GitHub Actions workflow YAML. Contract
      # (bash -n + --help), offline happy path against the shipped
      # assets/github-actions-terraform.yml, the --json §7 envelope (jq-guarded), and a
      # NEGATIVE proving the verifier rejects a malformed uses (a ref with no @ ->
      # exit 4 MALFORMED). --live is NEVER invoked (it resolves refs against the GitHub
      # API); a network blip must never fail a PR.
      #
      # Usage:   bash tests/run.sh
      # Exit:    0 all pass, 1 one or more failures
      
      set -uo pipefail
      
      HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
      SKILL="$(dirname "$HERE")"
      V="$SKILL/scripts/check-action-refs.sh"
      DEFAULT="$SKILL/assets/github-actions-terraform.yml"
      
      SB="$(mktemp -d)"; trap 'rm -rf "$SB"' EXIT
      PASS=0; FAIL=0
      ok() { PASS=$((PASS+1)); printf '  PASS  %s\n' "$1"; }
      no() { FAIL=$((FAIL+1)); printf '  FAIL  %s\n' "$1"; }
      expect_exit() { [[ "$2" == "$3" ]] && ok "$1 (exit $3)" || no "$1 (want $2 got $3)"; }
      expect_has()  { case "$3" in *"$2"*) ok "$1";; *) no "$1 (missing '$2')";; esac; }
      
      echo "=== terraform-ops self-test ==="
      
      # ── contract ──────────────────────────────────────────────────────────────────
      echo "-- contract --"
      bash -n "$V" && ok "bash -n check-action-refs.sh" || no "bash -n check-action-refs.sh"
      bash "$V" --help >/dev/null 2>&1; expect_exit "--help exits 0" 0 $?
      out="$(bash "$V" --help 2>&1)"
      expect_has "--help has Examples" "xamples" "$out"
      bash "$V" --bogus >/dev/null 2>&1; expect_exit "unknown flag -> 2" 2 $?
      
      # ── offline structural mode (§7 seam: --offline default, --live advisory) ─────
      echo "-- offline structural --"
      [[ -f "$DEFAULT" ]] && ok "shipped workflow present" || no "shipped workflow present"
      bash "$V" --offline >/dev/null 2>&1; expect_exit "--offline clean on shipped skill" 0 $?
      # --json needs jq; SKIP the envelope checks where jq is absent (a no-jq runner is
      # legitimate, and the verifier itself exits 5 for --json without jq).
      if command -v jq >/dev/null 2>&1; then
        out="$(bash "$V" --offline --json 2>/dev/null)"
        expect_has "--offline --json envelope schema" '"schema": "claude-mods.terraform-ops.action-refs/v1"' "$out"
        expect_has "--offline --json mode" '"mode": "offline"' "$out"
      else
        echo "  SKIP  --json envelope (no jq on this runner)"
      fi
      
      # ── negative: a uses with no @ref must be flagged malformed (exit 4) ──────────
      echo "-- negative --"
      cat > "$SB/bad.yml" <<'EOF'
      name: bad
      jobs:
        build:
          runs-on: ubuntu-latest
          steps:
            - uses: hashicorp/setup-terraform
      EOF
      bash "$V" --offline "$SB/bad.yml" >"$SB/neg.out" 2>&1
      expect_exit "--offline flags missing @ref -> 4" 4 $?
      expect_has "finding names the malformed ref" "malformed" "$(cat "$SB/neg.out")"
      
      # ── SKILL.md sanity ───────────────────────────────────────────────────────────
      echo "-- SKILL.md --"
      grep -q '^name: terraform-ops$' "$SKILL/SKILL.md" && ok "frontmatter name" || no "frontmatter name"
      grep -q 'check-action-refs.sh' "$SKILL/SKILL.md" && ok "verifier cited from SKILL.md" || no "verifier cited from SKILL.md"
      
      echo ""
      echo "=== $PASS passed, $FAIL failed ==="
      [[ "$FAIL" -eq 0 ]] || exit 1
      exit 0
      
  • SKILL.md 16.2 KB
    ---
    name: terraform-ops
    description: "Terraform and OpenTofu infrastructure-as-code operations - project layout, state management, module design, plan/apply safety, CI/CD pipelines, and secrets. Use for: terraform, opentofu, infrastructure as code, IaC, tfstate, terraform state, terraform module, remote backend, terraform plan, terraform apply, for_each, moved block, terraform import, drift detection, tflint, checkov, HCL."
    license: MIT
    allowed-tools: "Read Write Bash"
    metadata:
      author: claude-mods
      related-skills: ci-cd-ops, docker-ops, container-orchestration
    ---
    
    # Terraform Operations
    
    Terraform / OpenTofu infrastructure-as-code: layout, state, modules, safety, CI/CD, secrets.
    
    **Version context (verified 2026-06):** Terraform **1.15.x** (BUSL-1.1 licence since 1.6) · OpenTofu **1.12.x** (MPL-2.0 fork of Terraform 1.5.x). Commands below are interchangeable (`terraform` ↔ `tofu`) unless flagged. See [Terraform vs OpenTofu](#terraform-vs-opentofu) for the decision note.
    
    ## Reference Files
    
    | File | Covers |
    |------|--------|
    | [references/state-management.md](references/state-management.md) | Remote backends, locking, moved/import/removed blocks, state surgery, drift detection |
    | [references/module-patterns.md](references/module-patterns.md) | Module composition, variable validation, optional/nullable, output contracts, versioning |
    | [references/cicd-pipelines.md](references/cicd-pipelines.md) | GitHub Actions plan/apply, OIDC auth, policy gates (tflint/trivy/checkov/OPA), Atlantis/HCP |
    | [references/security-and-secrets.md](references/security-and-secrets.md) | Secrets in state, ephemeral resources, write-only arguments, SOPS/Vault, sensitive limits |
    | [assets/github-actions-terraform.yml](assets/github-actions-terraform.yml) | Ready-to-adapt PR-plan + OIDC-apply workflow |
    | [scripts/check-action-refs.sh](scripts/check-action-refs.sh) | Staleness verifier for any workflow's `uses:` action refs (offline structural / live API resolve) |
    
    > The action versions pinned in `github-actions-terraform.yml` are **point-in-time** (verified 2026-06). Run `scripts/check-action-refs.sh --live` before adopting — a tag that was valid at write time may have been retracted or never existed (e.g. `trivy-action@0.33.1` vs the real `v0.33.1`).
    
    ## Project Layout Decision Tree
    
    ```
    How many environments / accounts?
    │
    ├─ One environment, one team
    │  └─ Single root module + tfvars. Don't over-engineer.
    │
    ├─ Multiple environments (dev/staging/prod)
    │  ├─ Need different backend/account/region per env? (usually YES for prod isolation)
    │  │  └─ DIRECTORY PER ENVIRONMENT (recommended default)
    │  │     environments/{dev,staging,prod}/ each a thin root calling shared modules
    │  │
    │  └─ Environments truly identical except a few variables, same backend account?
    │     └─ Workspaces are *acceptable* — but see the workspace caveats below
    │
    └─ Many teams / many state files / platform engineering
       └─ Directory-per-env + per-component state split (network / data / app)
          Consider Terragrunt, Terraform Stacks (HCP), or OpenTofu + CI orchestration
    ```
    
    ### Canonical multi-env layout
    
    ```
    infra/
    ├── modules/                  # Reusable child modules (no provider/backend blocks)
    │   ├── network/
    │   │   ├── main.tf
    │   │   ├── variables.tf
    │   │   ├── outputs.tf
    │   │   └── versions.tf       # required_providers ONLY (no provider config)
    │   └── app-service/
    ├── environments/             # Root modules — one state file each
    │   ├── dev/
    │   │   ├── main.tf           # module "network" { source = "../../modules/network" ... }
    │   │   ├── backend.tf        # remote backend, env-specific key
    │   │   ├── providers.tf      # provider config lives in ROOT only
    │   │   ├── terraform.tfvars  # committed, non-secret env values
    │   │   └── versions.tf       # required_version + required_providers pins
    │   └── prod/
    └── .tflint.hcl
    ```
    
    ### Why directories usually beat workspaces
    
    | Concern | Directories | Workspaces |
    |---------|-------------|------------|
    | Separate backend/account per env | Yes — each root has its own `backend.tf` | No — one backend, envs differ only by state key |
    | Blast radius of wrong-env apply | Low — you're physically in `prod/` | High — invisible `terraform workspace select` state |
    | Env-specific config divergence | Natural (different main.tf if needed) | `terraform.workspace` conditionals creep everywhere |
    | Prod IAM isolation | Per-dir CI role | Same credentials see all envs |
    | Visibility in code review | Diff shows which env changed | Workspace is runtime state, not in the diff |
    
    Workspaces fit short-lived ephemeral copies (PR preview envs) — not the dev/prod boundary. HashiCorp's own docs say workspaces are "not suitable for strong separation."
    
    ### tfvars conventions
    
    ```bash
    terraform.tfvars            # auto-loaded — per-root committed defaults (non-secret)
    *.auto.tfvars               # auto-loaded — generated/local overrides
    prod.tfvars                 # explicit only: terraform plan -var-file=prod.tfvars
    TF_VAR_db_password=...      # env var injection — secrets in CI, never in files
    ```
    
    Gotcha: `-var-file` + directories-per-env is belt-and-braces; with workspaces it's load-bearing and one forgotten flag applies dev values to prod.
    
    ## State Quick Reference
    
    Full detail: [references/state-management.md](references/state-management.md).
    
    | Task | Command / block | Notes |
    |------|-----------------|-------|
    | Remote backend (AWS) | `backend "s3" { bucket, key, region, use_lockfile = true }` | S3-native locking (TF ≥1.10) — DynamoDB table no longer required |
    | Rename resource in code | `moved { from = aws_x.a, to = aws_x.b }` | Declarative, reviewable, no CLI surgery |
    | Adopt existing infra | `import { to = aws_x.a, id = "i-123" }` + `plan -generate-config-out=gen.tf` | Config-driven import (TF ≥1.5) beats `terraform import` CLI |
    | Forget without destroy | `removed { from = aws_x.a, lifecycle { destroy = false } }` | TF ≥1.7; OpenTofu 1.12 also has `lifecycle { destroy = false }` on resources |
    | Drift detection | `terraform plan -detailed-exitcode` | Exit 0 = clean, 1 = error, **2 = drift** — cron it |
    | Inspect state | `terraform state list` / `state show ADDR` | Read-only, always safe |
    | Move state (last resort) | `terraform state mv SRC DST` | Prefer `moved` blocks — see "when NOT to" below |
    | Pull/push (emergency) | `terraform state pull > backup.tfstate` | ALWAYS pull a backup before any surgery |
    
    **State surgery — when NOT to:** if a `moved`/`removed`/`import` block can express it, use the block. CLI `state mv`/`rm` is immediate, unreviewed, unversioned, and a typo orphans real infrastructure. Legit uses: splitting state between roots, unwedging a failed migration. Always `state pull` a backup first.
    
    ## Module Quick Reference
    
    Full detail: [references/module-patterns.md](references/module-patterns.md).
    
    ```hcl
    module "network" {
      source  = "terraform-aws-modules/vpc/aws"
      version = "~> 6.0"          # pin minor-float for registry modules; exact pin in prod roots
      # ...
    }
    ```
    
    | Rule | Why |
    |------|-----|
    | Composition over inheritance | Roots compose flat modules; never module-wraps-module-wraps-module |
    | No provider blocks in child modules | Providers configured in root only; child declares `required_providers` |
    | `validation` blocks on variables | Fail at plan with a real message, not mid-apply |
    | `optional(type, default)` in object attrs | Callers omit fields; `nullable = false` rejects explicit null |
    | Outputs are the contract | Output IDs/ARNs consumers need; document with `description` |
    | **Anti-pattern: thin wrappers** | A module that just renames variables of another module adds a version-lag layer and zero value — call the upstream module directly |
    
    ## Safety Checklist (before every apply)
    
    ```
    □ plan output READ, not skimmed — every destroy/replace explained
    □ "Plan: X to add, Y to change, Z to destroy" — does Z surprise you?
    □ -/+ (replace) lines: check the "forces replacement" attribute
    □ Applying the SAME saved plan that was reviewed: plan -out=tfplan → apply tfplan
    □ prevent_destroy on stateful resources (db, state bucket, KMS keys)
    □ Cloud-side deletion protection too (RDS deletion_protection, S3 versioning+MFA-delete)
    □ No -target unless this is a declared emergency (see below)
    □ for_each (stable keys), not count, for any collection that can reorder
    ```
    
    ### Footguns
    
    | Footgun | Detail | Fix |
    |---------|--------|-----|
    | `count` index shift | Removing item 0 of a `count` list re-addresses every later item → destroy/recreate cascade | `for_each` with stable string keys |
    | `-target` habit | Skips dependency graph; state diverges from config; hides drift | Emergency-only (broken dependency cycle, partial outage). Follow with a full clean plan |
    | `prevent_destroy` false comfort | Doesn't survive the block being deleted, and doesn't stop `state rm` + console delete | Pair with cloud-native deletion protection |
    | Dynamic blocks everywhere | `dynamic` for 2 static blocks is obfuscation | Use `dynamic` only over genuinely variable collections |
    | Unpinned providers | `aws = ">= 5.0"` in prod pulls a breaking major the day it ships | `~> 6.12` + commit `.terraform.lock.hcl` |
    | Apply ≠ reviewed plan | Plan on PR, apply on merge re-plans — drift in between applies unreviewed changes | Save the plan artifact, or accept + re-review the merge plan |
    
    ```hcl
    resource "aws_db_instance" "main" {
      deletion_protection = true            # cloud-side
      lifecycle {
        prevent_destroy = true              # terraform-side
        ignore_changes  = [password]        # if rotated outside TF
      }
    }
    ```
    
    ## CI/CD Quick Reference
    
    Full detail: [references/cicd-pipelines.md](references/cicd-pipelines.md) · template: [assets/github-actions-terraform.yml](assets/github-actions-terraform.yml).
    
    ```
    PR opened   → fmt -check → validate → tflint → trivy/checkov → plan → plan posted as PR comment
    PR merged   → plan (fresh) → apply, authenticated via OIDC — no long-lived cloud keys
    Nightly     → plan -detailed-exitcode → exit 2 ⇒ drift alert
    ```
    
    - **OIDC everywhere** — `aws-actions/configure-aws-credentials` with `role-to-assume`, never `AWS_ACCESS_KEY_ID` secrets. Same supply-chain doctrine as this repo's rules: short-lived tokens, no standing credentials.
    - **Pin action SHAs** in workflows (`uses: actions/checkout@<sha>`), not floating tags.
    - Policy gates: `tflint` (provider-aware lint), `trivy config` / `checkov` (misconfig scan), OPA/`conftest` for org policy ("no public buckets").
    
    | Orchestrator | Fit |
    |---|---|
    | Plain GitHub Actions | Default — full control, free, template in assets/ |
    | Atlantis | Self-hosted PR automation, `atlantis plan/apply` comments, locking per dir |
    | HCP Terraform / Terraform Cloud | Managed runs, Sentinel policy, state hosting; free ≤500 resources |
    | Spacelift / env0 / Digger / Scalr | Commercial Atlantis-likes; Digger runs inside your Actions |
    
    ### Verification — `uses:` ref staleness
    
    GitHub Action versions rot: a tag gets retracted, or a workflow pins one that never existed. [scripts/check-action-refs.sh](scripts/check-action-refs.sh) lints every `uses: owner/repo@ref` line. It's **general** — pass any workflow file(s) as positionals (default: this skill's own `assets/github-actions-terraform.yml`).
    
    ```bash
    # Structural only, no network — well-formedness of every uses: ref (CI-safe gate).
    # Floating @main/@master → WARN (exit 0; use --strict to fail). Malformed → exit 4.
    scripts/check-action-refs.sh --offline .github/workflows/ci.yml
    
    # Live — resolve each ref against the GitHub API. A 404 (ref doesn't exist) → exit 10
    # DRIFT; API unreachable/rate-limited → exit 7 (advisory, never fails the build, §7).
    # Set GITHUB_TOKEN to dodge the unauthenticated rate limit.
    GITHUB_TOKEN=$GH_PAT scripts/check-action-refs.sh --live .github/workflows/*.yml
    
    scripts/check-action-refs.sh --json --offline | jq '.data[] | select(.status!="ok")'
    ```
    
    `--live` is the check that catches the classic `aquasecurity/trivy-action@0.33.1` mistake — that tag 404s; the real one is `v0.33.1`. Run live on a schedule (never as a blocking PR gate), offline in PR CI.
    
    ## Testing Quick Reference
    
    ```hcl
    # tests/network.tftest.hcl  — native test framework (TF ≥1.6 / OpenTofu ≥1.6)
    variables { cidr = "10.0.0.0/16" }
    
    run "valid_cidr_plan" {
      command = plan                          # plan = fast unit-ish; apply = real integration
      assert {
        condition     = aws_vpc.main.cidr_block == "10.0.0.0/16"
        error_message = "VPC CIDR did not match input"
      }
    }
    
    run "rejects_tiny_cidr" {
      command = plan
      variables { cidr = "10.0.0.0/30" }
      expect_failures = [var.cidr]            # asserts the validation block fires
    }
    ```
    
    `terraform test` runs every `*.tftest.hcl` under `tests/`; `command = apply` runs create real (then auto-destroyed) infra — use a sandbox account. Mock providers (`mock_provider` blocks, TF ≥1.7) fake apply without credentials. For multi-tool/Go-level orchestration (retry, real HTTP probes), Terratest is the heavyweight alternative — native `terraform test` covers most module CI needs first.
    
    ## Secrets Quick Reference
    
    Full detail: [references/security-and-secrets.md](references/security-and-secrets.md).
    
    | Mechanism | Version | What it does |
    |---|---|---|
    | `sensitive = true` | all | Redacts from CLI output **only** — value still plaintext in state |
    | Ephemeral resources (`ephemeral "..."`) | TF ≥1.10 / OpenTofu ≥1.11 | Fetch secret at run time; never persisted to state or plan |
    | Write-only arguments (`password_wo`) | TF ≥1.11 / OpenTofu ≥1.11 | Send secret to provider; never stored in state; rotate via `_wo_version` |
    | SOPS-encrypted tfvars | tool | Secrets encrypted at rest in git; decrypted at plan time |
    | Vault / cloud secret manager | tool | Reference by ID; resource reads secret at boot, TF never sees it |
    | OpenTofu state encryption | OpenTofu ≥1.7 | Client-side AES-GCM encryption of state/plan — **no Terraform equivalent** |
    
    **Rule zero: treat state as secret regardless.** Encrypt the backend (SSE-KMS), restrict IAM on the bucket, never commit `*.tfstate` (gitignore it).
    
    ## Terraform vs OpenTofu
    
    | | Terraform | OpenTofu |
    |---|---|---|
    | Licence | **BUSL-1.1** since 1.6 (no production use *competing with HashiCorp*; fine for normal internal use) | **MPL-2.0** — genuinely open source, Linux Foundation |
    | Current | 1.15.x | 1.12.x |
    | Exclusive features | Stacks (HCP-tied), Terraform Cloud agents, `terraform query` | State/plan **encryption**, provider `for_each` iteration, `-exclude` flag, early variable eval in backend/module blocks, OCI registry distribution, `.tofu` file extension |
    | Registry | registry.terraform.io | registry.opentofu.org (mirrors most providers) |
    | Compatibility | — | Forked at 1.5.x; HCL/state compatible for mainstream use, diverging feature-by-feature since |
    
    **Decision:** vendors and anyone redistributing IaC tooling commercially → OpenTofu (licence risk). Teams on HCP Terraform/Sentinel → Terraform. Everyone else: either works; OpenTofu's state encryption is the single biggest technical differentiator. Migration `terraform → tofu` is `tofu init` + state-compatible up to ~1.8-era features; the gap widens each release — migrate early or commit.
    
    ## Command Quick Reference
    
    ```bash
    terraform init -upgrade               # init / upgrade providers within constraints
    terraform fmt -recursive -check       # CI: fail on unformatted
    terraform validate                    # syntax + internal consistency (no creds needed after init)
    terraform plan -out=tfplan            # save plan for exact-apply
    terraform show -json tfplan | jq      # machine-readable plan (policy tools eat this)
    terraform apply tfplan                # apply EXACTLY the reviewed plan
    terraform plan -detailed-exitcode     # 0 clean / 2 drift — for cron drift checks
    terraform plan -refresh-only          # show drift without proposing config changes
    terraform apply -replace=aws_x.a      # force recreate one resource (replaces old taint)
    terraform state pull > backup.json    # ALWAYS before surgery
    terraform output -json                # consume outputs in scripts
    terraform graph | dot -Tsvg > g.svg   # dependency graph
    tofu init                             # OpenTofu: same verbs throughout
    ```
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related