infra-iac-terraform
Infrastructure as Code with HashiCorp Terraform
Install
npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/infra-iac-terraform/skills/infra-iac-terraform
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
git clone https://github.com/agents-inc/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Terraform Patterns
Quick Guide: Declarative infrastructure using HCL. Pin provider versions in
required_providersand commit.terraform.lock.hcl. Use remote backends with state locking for team collaboration. Preferfor_eachovercountfor non-identical resources. Usemovedblocks for refactoring,importblocks for adopting existing infrastructure. Validate inputs withvalidationblocks and infrastructure withprecondition/postcondition. Keep modules flat, composable, and single-purpose. Runterraform fmtandterraform validatebefore every commit.
<critical_requirements>
CRITICAL: Before Using This Skill
All code must follow project conventions in CLAUDE.md
(You MUST pin provider versions with constraints in required_providers and commit .terraform.lock.hcl to version control)
(You MUST use a remote backend with state locking for any shared or production infrastructure)
(You MUST use for_each with a map/set for non-identical resources -- count causes index-shift destruction on removal)
(You MUST never store secrets in .tf files, .tfvars, or state -- use environment variables (TF_VAR_*) or your secrets manager)
(You MUST run terraform plan and review the diff before every terraform apply -- never apply blindly)
</critical_requirements>
Detailed Resources:
- examples/core.md - Resource definitions, variables, outputs, locals, data sources, provider configuration
- examples/modules.md - Module structure, composition, versioning, registry publishing
- examples/state.md - Remote backends, state locking, moved/import/removed blocks, workspaces
- examples/patterns.md - for_each, dynamic blocks, lifecycle, conditions, validations
- reference.md - Decision frameworks, CLI cheat sheet, file naming conventions
Auto-detection: Terraform, OpenTofu, HCL, .tf files, terraform init, terraform plan, terraform apply, terraform fmt, terraform validate, required_providers, terraform block, resource block, data source, module block, variable block, output block, locals, backend configuration, remote state, state locking, moved block, import block, for_each, count, dynamic block, lifecycle, precondition, postcondition, .terraform.lock.hcl, tfvars, provider configuration
When to use:
- Writing or reviewing Terraform/OpenTofu configuration files (
.tf) - Defining cloud resources, data sources, modules, variables, and outputs
- Managing state backends, locking, and multi-environment deployments
- Refactoring infrastructure with
moved,import, andremovedblocks - Structuring reusable modules for team or registry consumption
When NOT to use:
- Application code deployment logic (that belongs in CI/CD pipelines)
- Container orchestration configuration (Kubernetes manifests, Helm charts)
- One-off scripting tasks better handled by shell scripts or CLI tools
Key patterns covered:
- Provider pinning, lock files, and version constraints
- Resource definitions with meta-arguments (
for_each,count,depends_on,lifecycle) - Variable validation, locals for derived values, output descriptions
- Remote backend configuration with state locking
- Module structure (flat composition, single-purpose modules)
- Refactoring with
moved,import, andremovedblocks - Custom conditions (
precondition,postcondition,checkblocks) - Dynamic blocks for repeated nested configuration
- Environment management (directory-based vs workspaces)
<red_flags>
RED FLAGS
High Priority:
- Missing
.terraform.lock.hclin version control -- different team members get different provider versions, causing plan drift and mysterious failures - Local state for shared infrastructure -- no locking means concurrent applies corrupt state; no remote backup means state loss is catastrophic
- Secrets in
.tfor.tfvarsfiles -- committed to version control, visible in state file; useTF_VAR_*environment variables or your secrets manager - Using
countfor non-identical resources -- removing an item shifts indices, destroying and recreating unrelated resources terraform applywithout reviewing the plan -- auto-approve in production is how you delete databases- Unpinned provider versions --
version = ">= 5.0"without an upper bound allows major version upgrades that break everything
Medium Priority:
depends_onwhen an expression reference suffices --depends_oncauses overly conservative plans; let Terraform infer dependencies from expressions- Deeply nested module trees -- modules calling modules calling modules are hard to debug; keep the tree flat and compose at the root
ignore_changes = all-- Terraform will never update the resource again, even for intentional changes; be specific about which attributes to ignore- No
descriptionon variables and outputs -- undocumented inputs/outputs make modules unusable for anyone but the author - Hardcoded values instead of variables -- makes modules non-reusable; parameterize anything that changes between environments
Gotchas & Edge Cases:
- Backend configuration cannot use variables, locals, or expressions -- values must be literal strings or passed via
-backend-configduringterraform init prevent_destroydoes not prevent destruction if you remove the resource block entirely -- it only preventsterraform destroyon the resource while the block existsfor_eachkeys must be known at plan time -- they cannot reference resource attributes that are computed during applymovedblocks are processed once duringterraform plan/apply-- remove them after the migration is applied to keep configuration cleanterraform statesubcommands (mv, rm, pull, push) bypass safety checks -- usemoved/removedblocks instead for auditable, reviewable refactoringdatasources are read during planning by default -- if they depend on resources being created in the same apply, usedepends_onto defer the readsensitive = trueon variables prevents the value from appearing in plan output but does NOT encrypt it in state -- state encryption is a separate concernterraform fmtonly formats.tffiles in the current directory -- useterraform fmt -recursiveto format all subdirectoriestoset()deduplicates -- if your list has duplicates,for_each = toset(var.list)silently drops them.tfvarsfiles are auto-loaded only if namedterraform.tfvarsor*.auto.tfvars-- other filenames require explicit-var-fileflag
</red_flags>
<critical_reminders>
CRITICAL REMINDERS
All code must follow project conventions in CLAUDE.md
(You MUST pin provider versions with constraints in required_providers and commit .terraform.lock.hcl to version control)
(You MUST use a remote backend with state locking for any shared or production infrastructure)
(You MUST use for_each with a map/set for non-identical resources -- count causes index-shift destruction on removal)
(You MUST never store secrets in .tf files, .tfvars, or state -- use environment variables (TF_VAR_*) or your secrets manager)
(You MUST run terraform plan and review the diff before every terraform apply -- never apply blindly)
Failure to follow these rules will cause state corruption, accidental resource destruction, secret exposure, and non-reproducible infrastructure.
</critical_reminders>
Files (skills)
-
examples
-
core.md 8.4 KB
# Terraform - Core Examples > Resource definitions, variables, outputs, locals, data sources, and provider configuration. See [SKILL.md](../SKILL.md) for pattern summaries and [reference.md](../reference.md) for decision frameworks. **Additional Examples:** - [modules.md](modules.md) - Module structure, composition, versioning - [state.md](state.md) - Remote backends, state locking, moved/import/removed blocks - [patterns.md](patterns.md) - for_each, dynamic blocks, lifecycle, conditions --- ## Pattern 1: Provider Configuration ### Basic Provider with Default Tags ```hcl # providers.tf provider "aws" { region = var.aws_region default_tags { tags = { Project = var.project_name Environment = var.environment ManagedBy = "terraform" } } } ``` **Why good:** `default_tags` ensures every AWS resource gets baseline tags for cost allocation and compliance without repeating tags on every resource block. ### Provider Aliases for Multi-Region ```hcl # providers.tf provider "aws" { region = "us-east-1" alias = "us_east" } provider "aws" { region = "eu-west-1" alias = "eu_west" } # main.tf -- reference alias explicitly resource "aws_s3_bucket" "us_logs" { provider = aws.us_east bucket = "${var.project_name}-logs-us" } resource "aws_s3_bucket" "eu_logs" { provider = aws.eu_west bucket = "${var.project_name}-logs-eu" } ``` **Why good:** Aliases enable multi-region deployments from a single configuration. Resources without an explicit `provider` use the default (non-aliased) provider. --- ## Pattern 2: Resource Argument Ordering Follow the official style guide ordering: meta-arguments first, resource arguments, nested blocks, lifecycle last. ### Good Example ```hcl resource "aws_instance" "app" { for_each = var.app_instances # 1. Meta-arguments first ami = data.aws_ami.ubuntu.id # 2. Resource arguments instance_type = each.value.instance_type subnet_id = each.value.subnet_id tags = merge( { Name = "app-${each.key}" }, var.extra_tags, ) root_block_device { # 3. Nested blocks volume_size = each.value.volume_size_gb encrypted = true } lifecycle { # 4. Lifecycle last create_before_destroy = true } } ``` **Why good:** Consistent ordering makes resources scannable; meta-arguments at top tell you how many instances exist; lifecycle at bottom is where you look for special behavior. ### Bad Example ```hcl resource "aws_instance" "app" { lifecycle { # BAD: lifecycle buried at top create_before_destroy = true } tags = { Name = "app" } # BAD: tags before core config ami = "ami-12345678" # BAD: hardcoded AMI for_each = var.instances # BAD: meta-argument after other args instance_type = "t3.micro" # BAD: hardcoded instance type } ``` **Why bad:** Inconsistent ordering forces readers to scan the entire block to understand meta-arguments and lifecycle; hardcoded values prevent reuse across environments. --- ## Pattern 3: Variables with Types and Validation ### Simple Variables ```hcl # variables.tf variable "environment" { type = string description = "Deployment environment" validation { condition = contains(["dev", "staging", "production"], var.environment) error_message = "Environment must be dev, staging, or production." } } variable "instance_count" { type = number description = "Number of application instances" default = 2 validation { condition = var.instance_count >= 1 && var.instance_count <= 10 error_message = "Instance count must be between 1 and 10." } } ``` ### Complex Variable Types ```hcl variable "ingress_rules" { type = list(object({ from_port = number to_port = number protocol = string cidr_blocks = list(string) description = string })) description = "List of ingress rules for the security group" default = [] } variable "app_instances" { type = map(object({ instance_type = string subnet_id = string volume_size_gb = number })) description = "Map of application instance configurations keyed by name" } ``` **Why good:** Structured types enforce shape at plan time. Map keys become stable `for_each` keys. Descriptions make modules self-documenting. ### Sensitive Variables ```hcl variable "database_password" { type = string description = "Database master password" sensitive = true # Redacted from plan/apply output validation { condition = length(var.database_password) >= 16 error_message = "Database password must be at least 16 characters." } } ``` **Gotcha:** `sensitive = true` only redacts from CLI output. The value is still stored in plain text in state. For true secret protection, use state encryption (OpenTofu) or a secrets manager data source. --- ## Pattern 4: Outputs ```hcl # outputs.tf output "vpc_id" { value = aws_vpc.main.id description = "ID of the VPC" } output "private_subnet_ids" { value = [for s in aws_subnet.private : s.id] description = "List of private subnet IDs" } output "database_endpoint" { value = aws_db_instance.main.endpoint description = "Database connection endpoint" sensitive = true # Contains host:port, may be sensitive } ``` **Key rules:** Every output needs a `description`. Use `sensitive = true` for outputs containing connection strings, IPs, or credentials. Outputs are the module's public API -- treat them as a contract. --- ## Pattern 5: Locals for Derived Values Use locals to name complex expressions and avoid repetition. Keep them deterministic -- they should not change between runs. ### Good Example ```hcl locals { # Derived from variables -- deterministic name_prefix = "${var.project_name}-${var.environment}" # Common tags merged from variable + computed values common_tags = merge( var.extra_tags, { Project = var.project_name Environment = var.environment ManagedBy = "terraform" }, ) # Conditional logic named for clarity is_production = var.environment == "production" # Pre-computed map for for_each subnet_config = { for idx, cidr in var.private_subnet_cidrs : "private-${idx}" => { cidr = cidr availability_zone = var.availability_zones[idx % length(var.availability_zones)] } } } ``` **Why good:** Named locals make resource blocks readable. `is_production` is clearer than repeating `var.environment == "production"` in multiple resources. Pre-computed maps with `for` expressions keep `for_each` arguments clean. ### Bad Example ```hcl locals { # BAD: Chained locals that are hard to follow step1 = [for x in var.items : x if x.enabled] step2 = [for x in local.step1 : merge(x, { processed = true })] step3 = { for x in local.step2 : x.name => x } # BAD: Local that should be a variable (caller-controlled) instance_type = "t3.micro" } ``` **Why bad:** Deep local chains obscure what `step3` actually contains. If a value should be set by the caller, it belongs in a variable, not a local. --- ## Pattern 6: Data Sources Data sources read existing infrastructure. They are evaluated during planning by default. ### Good Example ```hcl # Fetch latest Ubuntu AMI data "aws_ami" "ubuntu" { most_recent = true owners = ["099720109477"] # Canonical filter { name = "name" values = ["ubuntu/images/hvm-ssd-gp3/ubuntu-noble-24.04-amd64-server-*"] } filter { name = "virtualization-type" values = ["hvm"] } } # Read existing VPC by tag data "aws_vpc" "main" { filter { name = "tag:Name" values = ["${var.project_name}-vpc"] } } # Reference in resources resource "aws_instance" "web" { ami = data.aws_ami.ubuntu.id subnet_id = data.aws_vpc.main.id # ... } ``` **Why good:** Data sources reference existing infrastructure without managing it. AMI lookups always get the latest patched image. VPC lookup by tag is more resilient than hardcoded IDs. ### Gotcha: Data Source Timing ```hcl # BAD: Data source depends on resource created in same apply data "aws_instance" "web" { instance_id = aws_instance.web.id # Not yet created during plan! } # GOOD: Use depends_on to defer the read data "aws_instance" "web" { depends_on = [aws_instance.web] instance_id = aws_instance.web.id } ``` **Why:** Data sources run during planning. If they reference resources that don't exist yet, the plan fails. `depends_on` defers the read until after the dependency is created. -
modules.md 6.8 KB
# Terraform - Module Examples > Module structure, composition, versioning, and registry patterns. See [SKILL.md](../SKILL.md) for core concepts and [reference.md](../reference.md) for the module decision framework. **Additional Examples:** - [core.md](core.md) - Resource definitions, variables, outputs, locals, data sources - [state.md](state.md) - Remote backends, state locking, moved/import/removed blocks - [patterns.md](patterns.md) - for_each, dynamic blocks, lifecycle, conditions --- ## Pattern 1: Standard Module Structure Every module follows the same file layout. `README.md` presence makes a nested module public-facing. ``` modules/ vpc/ main.tf # Resource definitions variables.tf # Input variables with type, description, validation outputs.tf # Output values with description README.md # Usage documentation (required for public modules) locals.tf # Local values (optional, if complex) data.tf # Data sources (optional, if needed) ``` ### Module: variables.tf ```hcl # modules/vpc/variables.tf variable "cidr_block" { type = string description = "CIDR block for the VPC" validation { condition = can(cidrhost(var.cidr_block, 0)) error_message = "Must be a valid CIDR block (e.g., 10.0.0.0/16)." } } variable "environment" { type = string description = "Environment name used for resource naming and tagging" } variable "private_subnet_cidrs" { type = list(string) description = "CIDR blocks for private subnets" default = [] } variable "public_subnet_cidrs" { type = list(string) description = "CIDR blocks for public subnets" default = [] } variable "enable_nat_gateway" { type = bool description = "Whether to create a NAT gateway for private subnet internet access" default = false } ``` ### Module: main.tf ```hcl # modules/vpc/main.tf resource "aws_vpc" "this" { cidr_block = var.cidr_block enable_dns_hostnames = true enable_dns_support = true tags = { Name = "${var.environment}-vpc" } } resource "aws_subnet" "private" { for_each = { for idx, cidr in var.private_subnet_cidrs : "private-${idx}" => { cidr = cidr, az_index = idx } } vpc_id = aws_vpc.this.id cidr_block = each.value.cidr availability_zone = data.aws_availability_zones.available.names[ each.value.az_index % length(data.aws_availability_zones.available.names) ] tags = { Name = "${var.environment}-${each.key}" Tier = "private" } } data "aws_availability_zones" "available" { state = "available" } ``` ### Module: outputs.tf ```hcl # modules/vpc/outputs.tf output "vpc_id" { value = aws_vpc.this.id description = "ID of the created VPC" } output "private_subnet_ids" { value = [for s in aws_subnet.private : s.id] description = "List of private subnet IDs" } ``` **Why good:** Consistent file layout makes any module instantly navigable. Every variable has type + description + validation. Every output has a description. The module does one thing (VPC networking). --- ## Pattern 2: Root Module Composition Root modules compose child modules. Keep the tree flat -- root calls modules directly, modules do not call other modules. ### Good Example ```hcl # environments/production/main.tf module "vpc" { source = "../../modules/vpc" cidr_block = "10.0.0.0/16" environment = "production" private_subnet_cidrs = ["10.0.1.0/24", "10.0.2.0/24", "10.0.3.0/24"] public_subnet_cidrs = ["10.0.101.0/24", "10.0.102.0/24"] enable_nat_gateway = true } module "compute" { source = "../../modules/compute" environment = "production" vpc_id = module.vpc.vpc_id subnet_ids = module.vpc.private_subnet_ids instance_type = "t3.large" instance_count = 3 } module "database" { source = "../../modules/database" environment = "production" vpc_id = module.vpc.vpc_id subnet_ids = module.vpc.private_subnet_ids } ``` **Why good:** Flat composition. Root module is the orchestrator -- it wires outputs from one module to inputs of another. Each module is independently reusable. ### Bad Example ```hcl # BAD: Module calling another module (nested) # modules/app/main.tf module "vpc" { source = "../vpc" # Module-to-module dependency cidr_block = var.cidr_block } module "compute" { source = "../compute" vpc_id = module.vpc.vpc_id # Tight coupling subnet_ids = module.vpc.private_subnet_ids } ``` **Why bad:** The `app` module now depends on both `vpc` and `compute` modules -- it can't be used without them. Debugging requires tracing through multiple module layers. The root module loses visibility into what's being created. --- ## Pattern 3: Module Versioning and Sources ### Registry Module (Versioned) ```hcl module "vpc" { source = "terraform-aws-modules/vpc/aws" version = "~> 5.0" # Pin to major version name = "${var.environment}-vpc" cidr = "10.0.0.0/16" azs = ["us-east-1a", "us-east-1b", "us-east-1c"] private_subnets = ["10.0.1.0/24", "10.0.2.0/24", "10.0.3.0/24"] public_subnets = ["10.0.101.0/24", "10.0.102.0/24", "10.0.103.0/24"] enable_nat_gateway = true } ``` ### Git Source (Tag-Pinned) ```hcl module "internal_module" { source = "git::https://github.com/org/terraform-modules.git//modules/vpc?ref=v2.1.0" cidr_block = "10.0.0.0/16" environment = var.environment } ``` ### Local Source ```hcl module "vpc" { source = "./modules/vpc" cidr_block = "10.0.0.0/16" environment = var.environment } ``` **Key rules:** - Registry modules: always pin with `version = "~> X.Y"` (pessimistic constraint) - Git sources: always pin with `?ref=vX.Y.Z` (exact tag) - Local modules: no version pinning needed (changes apply immediately) - Never use `ref=main` or `ref=HEAD` for git sources in production -- pins to a moving target --- ## Pattern 4: Module with Optional Features Use `count` or `for_each` conditional patterns to make module features optional. ```hcl # modules/vpc/main.tf # NAT Gateway -- only when enabled resource "aws_nat_gateway" "this" { count = var.enable_nat_gateway ? 1 : 0 allocation_id = aws_eip.nat[0].id subnet_id = values(aws_subnet.public)[0].id tags = { Name = "${var.environment}-nat" } } resource "aws_eip" "nat" { count = var.enable_nat_gateway ? 1 : 0 domain = "vpc" } # Route table for private subnets -- conditional NAT route resource "aws_route" "private_nat" { count = var.enable_nat_gateway ? 1 : 0 route_table_id = aws_route_table.private.id destination_cidr_block = "0.0.0.0/0" nat_gateway_id = aws_nat_gateway.this[0].id } ``` **Why good:** `count = var.enable_flag ? 1 : 0` is the standard pattern for conditional resource creation. Dev environments skip the NAT gateway (cost savings), production enables it. -
patterns.md 9.8 KB
# Terraform - Advanced Patterns > for_each, dynamic blocks, lifecycle meta-arguments, custom conditions, and expressions. See [SKILL.md](../SKILL.md) for pattern summaries and [reference.md](../reference.md) for decision frameworks. **Additional Examples:** - [core.md](core.md) - Resource definitions, variables, outputs, locals, data sources - [modules.md](modules.md) - Module structure, composition, versioning - [state.md](state.md) - Remote backends, state locking, moved/import/removed blocks --- ## Pattern 1: for_each with Maps ### Good Example ```hcl variable "buckets" { type = map(object({ versioning = bool lifecycle_days = number })) description = "Map of S3 buckets to create, keyed by bucket purpose" } resource "aws_s3_bucket" "this" { for_each = var.buckets bucket = "${var.project_name}-${each.key}" tags = { Name = each.key Purpose = each.key } } resource "aws_s3_bucket_versioning" "this" { for_each = { for k, v in var.buckets : k => v if v.versioning } bucket = aws_s3_bucket.this[each.key].id versioning_configuration { status = "Enabled" } } ``` **Why good:** Map keys are stable identifiers. Removing a bucket from the map only destroys that specific bucket. The versioning resource uses a filtered `for` expression to only create for buckets with `versioning = true`. ### Calling with map variable ```hcl # terraform.tfvars buckets = { logs = { versioning = true lifecycle_days = 90 } artifacts = { versioning = false lifecycle_days = 30 } } ``` --- ## Pattern 2: for_each with toset When you have a simple list of unique values (no complex object needed), convert to a set. ```hcl variable "team_members" { type = list(string) description = "List of IAM user names to create" default = ["alice", "bob", "carol"] } resource "aws_iam_user" "team" { for_each = toset(var.team_members) name = each.value tags = { ManagedBy = "terraform" } } ``` **Gotcha:** `toset()` deduplicates. If your list has `["alice", "alice", "bob"]`, only two users are created. This is silent -- no warning. --- ## Pattern 3: Conditional Resource Creation Use `count = var.flag ? 1 : 0` for resources that should optionally exist. ```hcl variable "enable_monitoring" { type = bool description = "Whether to create monitoring resources" default = false } resource "aws_cloudwatch_metric_alarm" "high_cpu" { count = var.enable_monitoring ? 1 : 0 alarm_name = "${var.project_name}-high-cpu" comparison_operator = "GreaterThanThreshold" evaluation_periods = 2 metric_name = "CPUUtilization" namespace = "AWS/EC2" period = 300 statistic = "Average" threshold = 80 alarm_description = "CPU utilization exceeds 80% for 10 minutes" } # Reference conditional resource carefully output "alarm_arn" { value = var.enable_monitoring ? aws_cloudwatch_metric_alarm.high_cpu[0].arn : null description = "ARN of the CPU alarm, null if monitoring disabled" } ``` **Why good:** Boolean variable makes the feature toggle explicit. Output uses ternary to handle the case where the resource doesn't exist. `count = condition ? 1 : 0` is the standard Terraform pattern for conditional resources. --- ## Pattern 4: Dynamic Blocks Generate repeated nested blocks from a variable-length collection. ### Good Example ```hcl variable "ingress_rules" { type = list(object({ from_port = number to_port = number protocol = string cidr_blocks = list(string) description = string })) description = "Ingress rules for the security group" } resource "aws_security_group" "web" { name = "${var.project_name}-web-sg" description = "Security group for web servers" vpc_id = var.vpc_id dynamic "ingress" { for_each = var.ingress_rules content { from_port = ingress.value.from_port to_port = ingress.value.to_port protocol = ingress.value.protocol cidr_blocks = ingress.value.cidr_blocks description = ingress.value.description } } egress { from_port = 0 to_port = 0 protocol = "-1" cidr_blocks = ["0.0.0.0/0"] description = "Allow all outbound traffic" } } ``` **Why good:** Module callers can define as many ingress rules as needed. The egress block is static and written literally (not dynamic) because it's the same for every caller. ### When NOT to Use Dynamic Blocks ```hcl # BAD: Dynamic block for a fixed set of 2 rules -- just write them out dynamic "ingress" { for_each = [ { port = 80, desc = "HTTP" }, { port = 443, desc = "HTTPS" }, ] content { from_port = ingress.value.port to_port = ingress.value.port protocol = "tcp" cidr_blocks = ["0.0.0.0/0"] description = ingress.value.desc } } # GOOD: Write the two blocks literally -- clearer and easier to read ingress { from_port = 80 to_port = 80 protocol = "tcp" cidr_blocks = ["0.0.0.0/0"] description = "HTTP" } ingress { from_port = 443 to_port = 443 protocol = "tcp" cidr_blocks = ["0.0.0.0/0"] description = "HTTPS" } ``` **Why:** HashiCorp's own docs recommend writing nested blocks literally where possible. Dynamic blocks add indirection -- only use them in modules where the caller needs to control the set. --- ## Pattern 5: Lifecycle Meta-Arguments ### prevent_destroy for Critical Resources ```hcl resource "aws_db_instance" "main" { identifier = "${var.project_name}-db" engine = "postgres" engine_version = "16.4" instance_class = var.db_instance_class # ... lifecycle { prevent_destroy = true } } resource "aws_s3_bucket" "terraform_state" { bucket = "myorg-terraform-state" lifecycle { prevent_destroy = true } } ``` **Gotcha:** `prevent_destroy` only blocks `terraform destroy` while the resource block exists. If you remove the entire resource block from your `.tf` files, Terraform will destroy the resource regardless. ### create_before_destroy for Zero-Downtime ```hcl resource "aws_launch_template" "app" { name_prefix = "${var.project_name}-" image_id = data.aws_ami.app.id instance_type = var.instance_type lifecycle { create_before_destroy = true } } ``` **When to use:** Resources behind load balancers, ASG launch templates, TLS certificates, DNS records -- anything where destroying before creating causes downtime. ### ignore_changes for Externally Managed Attributes ```hcl resource "aws_autoscaling_group" "app" { name = "${var.project_name}-asg" desired_capacity = var.initial_capacity min_size = var.min_capacity max_size = var.max_capacity launch_template { id = aws_launch_template.app.id } lifecycle { ignore_changes = [desired_capacity] # Auto-scaling changes this } } ``` **Why:** Without `ignore_changes`, every `terraform apply` resets `desired_capacity` to the Terraform-configured value, undoing auto-scaling decisions. ### replace_triggered_by ```hcl resource "aws_instance" "app" { ami = var.ami_id instance_type = var.instance_type lifecycle { replace_triggered_by = [ aws_launch_template.app.latest_version, # Force replacement on template change ] } } ``` **When to use:** Force resource replacement when a dependency changes that Terraform's normal dependency graph does not detect. --- ## Pattern 6: Preconditions, Postconditions, and Checks ### Precondition (Validate Before Creation) ```hcl data "aws_ami" "app" { most_recent = true owners = ["self"] filter { name = "name" values = ["${var.project_name}-*"] } } resource "aws_instance" "app" { ami = data.aws_ami.app.id instance_type = var.instance_type lifecycle { precondition { condition = data.aws_ami.app.architecture == "x86_64" error_message = "AMI must be x86_64 architecture for this instance type." } } } ``` ### Postcondition (Validate After Creation) ```hcl resource "aws_db_instance" "main" { identifier = "${var.project_name}-db" engine = "postgres" instance_class = var.db_instance_class # ... lifecycle { postcondition { condition = self.status == "available" error_message = "Database instance did not reach 'available' status." } } } ``` ### Check Block (Warnings, Non-Blocking) ```hcl check "health_check" { data "http" "app_health" { url = "https://${aws_lb.app.dns_name}/health" } assert { condition = data.http.app_health.status_code == 200 error_message = "Application health check failed after deployment." } } ``` **Key difference:** - `precondition` -- blocks plan if condition fails (validate assumptions) - `postcondition` -- blocks apply if condition fails (validate guarantees) - `check` -- produces a warning but does NOT block plan or apply (informational) --- ## Pattern 7: For Expressions Transform collections inline. Use for filtering, mapping, and restructuring data. ```hcl # Filter a map -- only production instances locals { prod_instances = { for name, config in var.instances : name => config if config.environment == "production" } } # Map transformation -- extract specific field locals { instance_arns = [for inst in aws_instance.app : inst.arn] } # Restructure -- list of objects to map keyed by name locals { user_map = { for user in var.users : user.name => user } } # Conditional values with ternary in for locals { instance_types = { for name, config in var.instances : name => ( config.environment == "production" ? "t3.large" : "t3.micro" ) } } ``` **Key syntax:** `{ for k, v in map : new_key => new_value if condition }` produces a map. `[ for v in list : expression ]` produces a list. Add `...` after the value to group by key: `{ for k, v in map : group_key => v... }`. -
state.md 8.1 KB
# Terraform - State Management Examples > Remote backends, state locking, moved/import/removed blocks, and workspace patterns. See [SKILL.md](../SKILL.md) for core concepts and [reference.md](../reference.md) for state organization decision frameworks. **Additional Examples:** - [core.md](core.md) - Resource definitions, variables, outputs, locals, data sources - [modules.md](modules.md) - Module structure, composition, versioning - [patterns.md](patterns.md) - for_each, dynamic blocks, lifecycle, conditions --- ## Pattern 1: Remote Backend with State Locking ### S3 Backend (AWS) ```hcl # backend.tf terraform { backend "s3" { bucket = "myorg-terraform-state" key = "production/network/terraform.tfstate" region = "us-east-1" encrypt = true dynamodb_table = "terraform-state-locks" } } ``` **Why good:** S3 provides durable storage with versioning for rollback. DynamoDB table provides state locking -- prevents two people from running `terraform apply` simultaneously and corrupting state. `encrypt = true` encrypts state at rest. ### GCS Backend (Google Cloud) ```hcl terraform { backend "gcs" { bucket = "myorg-terraform-state" prefix = "production/network" } } ``` **Note:** GCS has built-in state locking -- no separate lock table needed. ### Azure Backend ```hcl terraform { backend "azurerm" { resource_group_name = "terraform-state-rg" storage_account_name = "myorgterraformstate" container_name = "tfstate" key = "production/network/terraform.tfstate" } } ``` --- ## Pattern 2: Partial Backend Configuration Backend blocks cannot use variables. Use partial configuration to keep environment-specific values out of code. ### Backend with Placeholders ```hcl # backend.tf -- shared across environments terraform { backend "s3" { # key is the only value that differs per environment # bucket, region, dynamodb_table passed via -backend-config } } ``` ### Config Files per Environment ```hcl # config/production.hcl bucket = "myorg-terraform-state" key = "production/terraform.tfstate" region = "us-east-1" encrypt = true dynamodb_table = "terraform-state-locks" ``` ```hcl # config/staging.hcl bucket = "myorg-terraform-state" key = "staging/terraform.tfstate" region = "us-east-1" encrypt = true dynamodb_table = "terraform-state-locks" ``` ### Usage ```bash # Initialize with environment-specific backend config terraform init -backend-config=config/production.hcl terraform init -backend-config=config/staging.hcl ``` **Why good:** Backend configuration is separated from Terraform code. Same `.tf` files work for all environments. No secrets in version-controlled files. --- ## Pattern 3: State Organization by Layer Split state by infrastructure layer to minimize blast radius and speed up plans. ``` infrastructure/ network/ # VPC, subnets, route tables, NAT gateways backend.tf # key = "prod/network/terraform.tfstate" main.tf outputs.tf # Exports VPC ID, subnet IDs for other layers compute/ # Instances, ASGs, load balancers backend.tf # key = "prod/compute/terraform.tfstate" main.tf data.tf # Reads network outputs via terraform_remote_state data/ # RDS, ElastiCache, S3 buckets backend.tf # key = "prod/data/terraform.tfstate" main.tf ``` ### Cross-Layer References with terraform_remote_state ```hcl # compute/data.tf -- reads network layer's outputs data "terraform_remote_state" "network" { backend = "s3" config = { bucket = "myorg-terraform-state" key = "production/network/terraform.tfstate" region = "us-east-1" } } # compute/main.tf -- uses network outputs resource "aws_instance" "app" { subnet_id = data.terraform_remote_state.network.outputs.private_subnet_ids[0] # ... } ``` **Why good:** Changing a compute instance doesn't risk breaking the network layer. Each layer has a faster plan (fewer resources). Teams can work on different layers independently. **Gotcha:** `terraform_remote_state` creates a hard coupling to the backend configuration of the source layer. If the source layer's state key changes, all consumers must be updated. --- ## Pattern 4: Moved Blocks for Refactoring Use `moved` blocks to rename resources or extract them into modules without destroying infrastructure. ### Rename a Resource ```hcl # Before: resource was named "web_server" # After: renamed to "app_server" moved { from = aws_instance.web_server to = aws_instance.app_server } resource "aws_instance" "app_server" { # ... same configuration ... } ``` ### Move Resource into a Module ```hcl # Before: resource was at root level # After: moved into module "compute" moved { from = aws_instance.app to = module.compute.aws_instance.app } module "compute" { source = "./modules/compute" # ... } ``` ### Migrate from count to for_each ```hcl # Before: count-based # resource "aws_iam_user" "team" { # count = 3 # name = var.team_members[count.index] # } # After: for_each-based moved { from = aws_iam_user.team[0] to = aws_iam_user.team["alice"] } moved { from = aws_iam_user.team[1] to = aws_iam_user.team["bob"] } moved { from = aws_iam_user.team[2] to = aws_iam_user.team["carol"] } resource "aws_iam_user" "team" { for_each = toset(["alice", "bob", "carol"]) name = each.value } ``` **Key rules:** - Always run `terraform plan` after adding moved blocks -- verify "will be moved" messages - Remove moved blocks after the migration is applied (they are one-time operations) - Moved blocks are processed during plan/apply, not retroactively --- ## Pattern 5: Import Blocks for Adopting Existing Resources Bring pre-existing infrastructure under Terraform management. ```hcl # Import an existing S3 bucket import { to = aws_s3_bucket.existing_logs id = "my-company-logs-bucket" } resource "aws_s3_bucket" "existing_logs" { bucket = "my-company-logs-bucket" tags = { Name = "logs" ManagedBy = "terraform" } } ``` **Workflow:** 1. Add the `import` block and matching `resource` block 2. Run `terraform plan` -- Terraform shows what it will import and any config drift 3. Adjust the resource block until the plan shows no changes after import 4. Run `terraform apply` to execute the import 5. Remove the `import` block (one-time operation) **Gotcha:** You must write the resource block to match the existing resource's configuration, or Terraform will try to modify it on the next apply. --- ## Pattern 6: Removed Blocks Remove a resource from Terraform management without destroying the actual infrastructure. ```hcl # Stop managing this resource -- it was moved to another tool removed { from = aws_instance.legacy_app lifecycle { destroy = false # Keep the resource, just forget about it } } ``` **When to use:** Migrating resources to another Terraform configuration, handing off management to another tool, or cleaning up state without destroying infrastructure. --- ## Pattern 7: Workspaces for Ephemeral Environments Workspaces share the same `.tf` files but maintain separate state files. Best for environments that are structurally identical. ```hcl # Use workspace name for environment-specific values locals { environment = terraform.workspace instance_type = { dev = "t3.micro" staging = "t3.small" production = "t3.large" } } resource "aws_instance" "app" { instance_type = local.instance_type[local.environment] tags = { Environment = local.environment } } ``` ```bash # Create and switch workspaces terraform workspace new staging terraform workspace select staging terraform plan -var-file=staging.tfvars terraform apply -var-file=staging.tfvars ``` **When to use:** Environments that differ only in variable values (instance size, count, domain). **When not to use:** Environments with different resources, providers, or Terraform versions -- use directory-based separation instead. **Risk:** It's easy to forget which workspace is active. Always verify with `terraform workspace show` before applying.
-
-
reference.md 8.7 KB
# Terraform Quick Reference Decision frameworks, CLI cheat sheet, file naming conventions, and version constraint syntax. --- ## Decision Frameworks ### Environment Management: Directories vs Workspaces ``` Are environments structurally different (different resources, providers, or versions)? ├─ YES → Directory-based separation (environments/prod/, environments/staging/) │ - Each environment is self-contained and explicit │ - Different Terraform/provider versions per environment │ - Easier CI/CD isolation and access control └─ NO → Are environments identical except for variable values? ├─ YES → Workspaces with .tfvars per environment │ - terraform workspace select prod + terraform apply -var-file=prod.tfvars │ - Single codebase, multiple state files │ - Risk: easy to apply to wrong workspace └─ NO → Hybrid: directories for permanent envs, workspaces for ephemeral ``` ### State Organization: Monolith vs Layered ``` How many resources in a single state file? ├─ < 50 → Single state file is fine ├─ 50-200 → Consider splitting by layer (network, compute, data) └─ > 200 → Must split by layer - network/ (VPC, subnets, route tables) - compute/ (instances, ASGs, load balancers) - data/ (databases, caches, queues) - monitoring/ (alarms, dashboards) Benefits: smaller blast radius, faster plans, independent deploys ``` ### Module: Local vs Registry ``` Is this module used in multiple repositories? ├─ YES → Publish to registry (private or public) │ - Semantic versioning for controlled upgrades │ - source = "registry.example.com/org/module/provider" └─ NO → Is it used in multiple root modules within this repo? ├─ YES → Local module in modules/ directory │ - source = "./modules/vpc" └─ NO → Inline resources (no module needed) ``` ### for_each vs count vs neither ``` Creating multiple instances of a resource? ├─ NO → No meta-argument needed ├─ YES → Are instances identical except for count? │ ├─ YES → count is acceptable (e.g., count = var.enable ? 1 : 0) │ └─ NO → Do instances have unique identifiers? │ ├─ YES → for_each with map (keyed by identifier) │ └─ NO → for_each with toset() (keyed by value) ``` ### When to use lifecycle meta-arguments ``` Is this a critical resource (database, DNS zone, state bucket)? ├─ YES → prevent_destroy = true └─ NO → Does replacement cause downtime? ├─ YES → create_before_destroy = true └─ NO → Are external processes modifying attributes? ├─ YES → ignore_changes = [specific_attributes] └─ NO → No lifecycle block needed ``` --- ## File Naming Conventions (Official Style Guide) | File | Purpose | | -------------- | ----------------------------------------------------------- | | `terraform.tf` | `terraform` block: `required_version`, `required_providers` | | `backend.tf` | Backend configuration | | `providers.tf` | Provider blocks and configuration | | `main.tf` | Resource and data source definitions | | `variables.tf` | Input variable declarations (alphabetical) | | `outputs.tf` | Output declarations (alphabetical) | | `locals.tf` | Local value definitions | | `data.tf` | Data source blocks (if too many for `main.tf`) | | `versions.tf` | Alternative name for `terraform.tf` (common) | **For larger codebases:** Split `main.tf` by logical group: `network.tf`, `compute.tf`, `storage.tf`, `iam.tf`. --- ## HCL Style Rules - **Indentation:** 2 spaces per nesting level - **Alignment:** Align `=` signs for consecutive single-line arguments at the same level - **Naming:** `snake_case` for resources, variables, outputs, locals, modules - **Comments:** Use `#` (not `//` or `/* */`) - **Blank lines:** One between top-level blocks, one between arguments and nested blocks - **Meta-argument order:** `count`/`for_each` first, resource args next, nested blocks after, `lifecycle`/`depends_on` last - **Format:** Run `terraform fmt -recursive` before every commit - **Validate:** Run `terraform validate` to catch syntax and type errors --- ## Version Constraint Syntax | Constraint | Meaning | Example | | --------------- | ----------------------------------------- | -------------------- | | `= 1.0.0` | Exact version | Only 1.0.0 | | `>= 1.0.0` | Minimum version (no upper bound -- risky) | 1.0.0 and above | | `~> 1.0` | Pessimistic: allows 1.x, blocks 2.0 | 1.0 through 1.99 | | `~> 1.0.0` | Pessimistic: allows 1.0.x, blocks 1.1.0 | 1.0.0 through 1.0.99 | | `>= 1.0, < 2.0` | Explicit range | 1.0 through 1.99 | **Best practice:** Use `~> MAJOR.MINOR` for providers (allows patch updates, blocks breaking changes). Use `>= MAJOR.MINOR.0, < NEXT_MAJOR.0.0` for explicit ranges. --- ## CLI Cheat Sheet ### Core Workflow ```bash terraform init # Download providers, initialize backend terraform init -upgrade # Upgrade providers within constraints terraform init -backend-config=prod.hcl # Partial backend config terraform validate # Check syntax and types (no state access) terraform fmt -recursive # Format all .tf files terraform plan # Preview changes (always review before apply) terraform plan -out=tfplan # Save plan for exact apply terraform apply tfplan # Apply saved plan (no re-planning) terraform apply # Plan + apply interactively terraform destroy # Destroy all managed resources (caution!) ``` ### State Management ```bash terraform state list # List all resources in state terraform state show aws_instance.web # Show resource details terraform state pull # Download state (read-only inspection) # Prefer moved/import/removed blocks over CLI state commands for auditable changes ``` ### Import (Legacy CLI -- prefer import blocks) ```bash terraform import aws_s3_bucket.logs my-existing-bucket # Better: use import block in .tf file (reviewable, repeatable) ``` ### Workspace Commands ```bash terraform workspace list # List workspaces terraform workspace new dev # Create workspace terraform workspace select prod # Switch workspace terraform workspace show # Show current workspace ``` --- ## .gitignore for Terraform Projects ```gitignore # Local .terraform directories **/.terraform/* # .tfstate files (state should be remote, never committed) *.tfstate *.tfstate.* # Crash log files crash.log crash.*.log # Plan files (may contain secrets) *.tfplan out.plan # Override files (local-only overrides) override.tf override.tf.json *_override.tf *_override.tf.json # CLI configuration (user-specific) .terraformrc terraform.rc # DO commit .terraform.lock.hcl (provider version pinning) # !.terraform.lock.hcl -- ensure this is NOT gitignored ``` --- ## Common Terraform Functions | Function | Purpose | Example | | ---------------------------- | ----------------------------------- | ---------------------------------------------------- | | `lookup(map, key, default)` | Safe map lookup with fallback | `lookup(var.amis, var.region, "ami-default")` | | `try(expr, fallback)` | First expression that doesn't error | `try(var.config.name, "default")` | | `coalesce(vals...)` | First non-null, non-empty value | `coalesce(var.custom_name, local.generated_name)` | | `merge(maps...)` | Merge maps (last wins on conflict) | `merge(local.default_tags, var.extra_tags)` | | `flatten(list_of_lists)` | Flatten nested lists one level | `flatten([var.public_subnets, var.private_subnets])` | | `toset(list)` | Convert list to set (deduplicates) | `for_each = toset(var.team_members)` | | `cidrsubnet(prefix, new, n)` | Calculate subnet CIDR | `cidrsubnet("10.0.0.0/16", 8, 0)` = `10.0.0.0/24` | | `templatefile(path, vars)` | Render template file | `templatefile("user-data.sh.tpl", { env = "prod" })` | | `jsonencode(value)` | Convert to JSON string | `jsonencode(local.policy_document)` | | `format(fmt, vals...)` | Printf-style formatting | `format("web-%s-%02d", var.env, count.index)` | -
SKILL.md 17.2 KB
--- name: infra-iac-terraform description: Infrastructure as Code with HashiCorp Terraform --- # Terraform Patterns > **Quick Guide:** Declarative infrastructure using HCL. Pin provider versions in `required_providers` and commit `.terraform.lock.hcl`. Use remote backends with state locking for team collaboration. Prefer `for_each` over `count` for non-identical resources. Use `moved` blocks for refactoring, `import` blocks for adopting existing infrastructure. Validate inputs with `validation` blocks and infrastructure with `precondition`/`postcondition`. Keep modules flat, composable, and single-purpose. Run `terraform fmt` and `terraform validate` before every commit. --- <critical_requirements> ## CRITICAL: Before Using This Skill > **All code must follow project conventions in CLAUDE.md** **(You MUST pin provider versions with constraints in `required_providers` and commit `.terraform.lock.hcl` to version control)** **(You MUST use a remote backend with state locking for any shared or production infrastructure)** **(You MUST use `for_each` with a map/set for non-identical resources -- `count` causes index-shift destruction on removal)** **(You MUST never store secrets in `.tf` files, `.tfvars`, or state -- use environment variables (`TF_VAR_*`) or your secrets manager)** **(You MUST run `terraform plan` and review the diff before every `terraform apply` -- never apply blindly)** </critical_requirements> --- **Detailed Resources:** - [examples/core.md](examples/core.md) - Resource definitions, variables, outputs, locals, data sources, provider configuration - [examples/modules.md](examples/modules.md) - Module structure, composition, versioning, registry publishing - [examples/state.md](examples/state.md) - Remote backends, state locking, moved/import/removed blocks, workspaces - [examples/patterns.md](examples/patterns.md) - for_each, dynamic blocks, lifecycle, conditions, validations - [reference.md](reference.md) - Decision frameworks, CLI cheat sheet, file naming conventions --- **Auto-detection:** Terraform, OpenTofu, HCL, .tf files, terraform init, terraform plan, terraform apply, terraform fmt, terraform validate, required_providers, terraform block, resource block, data source, module block, variable block, output block, locals, backend configuration, remote state, state locking, moved block, import block, for_each, count, dynamic block, lifecycle, precondition, postcondition, .terraform.lock.hcl, tfvars, provider configuration **When to use:** - Writing or reviewing Terraform/OpenTofu configuration files (`.tf`) - Defining cloud resources, data sources, modules, variables, and outputs - Managing state backends, locking, and multi-environment deployments - Refactoring infrastructure with `moved`, `import`, and `removed` blocks - Structuring reusable modules for team or registry consumption **When NOT to use:** - Application code deployment logic (that belongs in CI/CD pipelines) - Container orchestration configuration (Kubernetes manifests, Helm charts) - One-off scripting tasks better handled by shell scripts or CLI tools **Key patterns covered:** - Provider pinning, lock files, and version constraints - Resource definitions with meta-arguments (`for_each`, `count`, `depends_on`, `lifecycle`) - Variable validation, locals for derived values, output descriptions - Remote backend configuration with state locking - Module structure (flat composition, single-purpose modules) - Refactoring with `moved`, `import`, and `removed` blocks - Custom conditions (`precondition`, `postcondition`, `check` blocks) - Dynamic blocks for repeated nested configuration - Environment management (directory-based vs workspaces) --- <philosophy> ## Philosophy Terraform is a declarative infrastructure-as-code tool. You describe the desired end-state; Terraform determines the steps to reach it. The HCL configuration language is designed to be human-readable and machine-parseable. **Core principles:** - **Declarative, not imperative** -- describe what you want, not how to get there - **State is the source of truth** -- Terraform tracks what it manages via state; protect it accordingly - **Modules are the unit of reuse** -- keep them flat, composable, and single-purpose - **Pin everything** -- provider versions, Terraform version, module versions; reproducibility is non-negotiable - **Plan before apply** -- always review the diff; never apply blindly in production **OpenTofu compatibility:** OpenTofu is an open-source fork (MPL 2.0) that is syntax-compatible with Terraform 1.5.x. The patterns in this skill apply to both tools. OpenTofu uses `.tofu` file extensions for OpenTofu-only features and adds native state encryption. </philosophy> --- <patterns> ## Core Patterns ### Pattern 1: Provider and Version Pinning Pin Terraform version and all provider versions. Commit `.terraform.lock.hcl` to version control. ```hcl # terraform.tf terraform { required_version = ">= 1.9.0, < 2.0.0" required_providers { aws = { source = "hashicorp/aws" version = "~> 5.0" # Allows 5.x, blocks 6.0 } } } ``` **Why this matters:** Without version constraints, `terraform init` on different machines downloads different provider versions, causing inconsistent plans and mysterious drift. The lock file pins exact versions and cryptographic hashes. **Version constraint syntax:** `= 1.0.0` (exact), `>= 1.0.0` (minimum), `~> 1.0` (allows 1.x, blocks 2.0), `>= 1.0, < 2.0` (range). See [examples/core.md](examples/core.md) for full provider configuration with aliases and default tags. --- ### Pattern 2: Resource Definitions and Meta-Arguments Resources follow a standard argument ordering: meta-arguments first, resource arguments next, nested blocks after, lifecycle last. ```hcl resource "aws_instance" "web" { count = var.instance_count # Meta-argument first ami = data.aws_ami.ubuntu.id instance_type = var.instance_type tags = { Name = "web-${count.index}" } lifecycle { # Lifecycle block last create_before_destroy = true } } ``` **Key meta-arguments:** `count`, `for_each`, `depends_on`, `provider`, `lifecycle`. Place meta-arguments at the top, separated from resource arguments by a blank line. See [examples/core.md](examples/core.md) for argument ordering and naming conventions. --- ### Pattern 3: for_each over count Use `for_each` with a map or set for non-identical resources. `count` uses numeric indices -- removing an item from the middle shifts all subsequent indices, causing unnecessary destruction and recreation. ```hcl # for_each with a map -- stable keys, safe removal resource "aws_iam_user" "team" { for_each = toset(var.team_members) # ["alice", "bob", "carol"] name = each.value } # Removing "bob" only destroys bob's user -- alice and carol are untouched ``` ```hcl # count -- index-based, dangerous on removal resource "aws_iam_user" "team" { count = length(var.team_members) name = var.team_members[count.index] } # Removing "bob" (index 1) shifts carol from index 2 to 1 -- carol gets destroyed and recreated ``` **When count is acceptable:** Identical resources where the only difference is the count (e.g., `count = var.enable_feature ? 1 : 0` for conditional creation). See [examples/patterns.md](examples/patterns.md) for for_each with maps, sets, and conditional patterns. --- ### Pattern 4: Variables with Validation Every variable needs `type`, `description`, and validation where constraints exist. Use named locals for derived values. ```hcl variable "environment" { type = string description = "Deployment environment (dev, staging, production)" validation { condition = contains(["dev", "staging", "production"], var.environment) error_message = "Environment must be dev, staging, or production." } } variable "instance_type" { type = string description = "EC2 instance type for the application server" default = "t3.micro" } ``` **Key rules:** Always set `type` and `description`. Use `validation` blocks to catch invalid input at plan time. Never use `default` for secrets -- force the caller to provide them. See [examples/core.md](examples/core.md) for complex variable types, sensitive variables, and output definitions. --- ### Pattern 5: Remote Backend with State Locking Never use local state for shared infrastructure. Remote backends provide locking, versioning, and team collaboration. ```hcl # backend.tf terraform { backend "s3" { bucket = "my-terraform-state" key = "prod/network/terraform.tfstate" region = "us-east-1" encrypt = true dynamodb_table = "terraform-locks" # State locking } } ``` **Critical:** Backend configuration cannot use variables or locals -- values must be literal or passed via `-backend-config` flags during `terraform init`. Use partial configuration for dynamic values. **State file hierarchy:** Organize state keys by environment and layer (e.g., `prod/network/`, `prod/compute/`, `staging/network/`) to minimize blast radius. See [examples/state.md](examples/state.md) for backend configuration, partial config, and state organization. --- ### Pattern 6: Module Structure and Composition Modules are the unit of reuse. Keep modules flat, single-purpose, and composable. ``` modules/ vpc/ main.tf # Resources variables.tf # Inputs outputs.tf # Outputs README.md # Usage docs (makes it public-facing) compute/ main.tf variables.tf outputs.tf ``` ```hcl # Root module calling child modules module "vpc" { source = "./modules/vpc" cidr_block = "10.0.0.0/16" environment = var.environment } module "compute" { source = "./modules/compute" subnet_ids = module.vpc.private_subnet_ids environment = var.environment } ``` **Key principle:** Keep the module tree flat. Deeply nested modules (module calling module calling module) are hard to debug and reuse. Prefer composition at the root level. See [examples/modules.md](examples/modules.md) for module versioning, registry sources, and internal vs published modules. --- ### Pattern 7: Refactoring with moved, import, and removed Blocks Refactor infrastructure without destroying resources. ```hcl # Rename a resource -- state updated, infrastructure untouched moved { from = aws_instance.web_server to = aws_instance.app_server } # Adopt existing infrastructure into Terraform management import { to = aws_s3_bucket.logs id = "my-existing-bucket-name" } # Remove from Terraform management without destroying the resource removed { from = aws_instance.legacy lifecycle { destroy = false } } ``` **Always run `terraform plan` after adding these blocks** to verify Terraform interprets the refactoring correctly. Look for "move" and "import" messages in the plan output. See [examples/state.md](examples/state.md) for module refactoring with moved blocks and bulk imports. --- ### Pattern 8: Lifecycle Meta-Arguments Control how Terraform manages resource lifecycle. ```hcl resource "aws_db_instance" "main" { # ... configuration ... lifecycle { prevent_destroy = true # Block accidental deletion create_before_destroy = true # Zero-downtime replacement ignore_changes = [tags] # External process manages tags } } ``` **When to use each:** - `prevent_destroy` -- databases, DNS zones, state buckets (critical resources) - `create_before_destroy` -- load balancers, instances behind ASGs (zero-downtime) - `ignore_changes` -- auto-scaling managed attributes, externally tagged resources - `replace_triggered_by` -- force replacement when a dependency changes that Terraform does not detect See [examples/patterns.md](examples/patterns.md) for lifecycle combinations and replace_triggered_by. --- ### Pattern 9: Custom Conditions (Preconditions, Postconditions, Checks) Validate assumptions before provisioning and guarantees after. ```hcl resource "aws_instance" "web" { ami = data.aws_ami.ubuntu.id instance_type = var.instance_type lifecycle { precondition { condition = data.aws_ami.ubuntu.architecture == "x86_64" error_message = "AMI must be x86_64 architecture." } postcondition { condition = self.public_ip != "" error_message = "Instance must have a public IP assigned." } } } ``` **Precondition vs postcondition vs check:** - `precondition` -- validates assumptions before creation (blocks plan) - `postcondition` -- validates guarantees after creation (blocks apply) - `check` block -- validates infrastructure state without blocking operations (warnings only) See [examples/patterns.md](examples/patterns.md) for check blocks and variable validation patterns. --- ### Pattern 10: Dynamic Blocks Generate repeated nested blocks from collections. Use sparingly -- overuse hurts readability. ```hcl resource "aws_security_group" "web" { name = "web-sg" dynamic "ingress" { for_each = var.ingress_rules content { from_port = ingress.value.from_port to_port = ingress.value.to_port protocol = ingress.value.protocol cidr_blocks = ingress.value.cidr_blocks } } } ``` **When to use:** Reusable modules where the number of nested blocks varies per caller. **When not to use:** Write nested blocks literally when the set is small and fixed. Dynamic blocks cannot generate meta-argument blocks (`lifecycle`, `provisioner`). See [examples/patterns.md](examples/patterns.md) for dynamic block patterns with conditionals. </patterns> --- <red_flags> ## RED FLAGS **High Priority:** - **Missing `.terraform.lock.hcl` in version control** -- different team members get different provider versions, causing plan drift and mysterious failures - **Local state for shared infrastructure** -- no locking means concurrent applies corrupt state; no remote backup means state loss is catastrophic - **Secrets in `.tf` or `.tfvars` files** -- committed to version control, visible in state file; use `TF_VAR_*` environment variables or your secrets manager - **Using `count` for non-identical resources** -- removing an item shifts indices, destroying and recreating unrelated resources - **`terraform apply` without reviewing the plan** -- auto-approve in production is how you delete databases - **Unpinned provider versions** -- `version = ">= 5.0"` without an upper bound allows major version upgrades that break everything **Medium Priority:** - **`depends_on` when an expression reference suffices** -- `depends_on` causes overly conservative plans; let Terraform infer dependencies from expressions - **Deeply nested module trees** -- modules calling modules calling modules are hard to debug; keep the tree flat and compose at the root - **`ignore_changes = all`** -- Terraform will never update the resource again, even for intentional changes; be specific about which attributes to ignore - **No `description` on variables and outputs** -- undocumented inputs/outputs make modules unusable for anyone but the author - **Hardcoded values instead of variables** -- makes modules non-reusable; parameterize anything that changes between environments **Gotchas & Edge Cases:** - Backend configuration cannot use variables, locals, or expressions -- values must be literal strings or passed via `-backend-config` during `terraform init` - `prevent_destroy` does not prevent destruction if you remove the resource block entirely -- it only prevents `terraform destroy` on the resource while the block exists - `for_each` keys must be known at plan time -- they cannot reference resource attributes that are computed during apply - `moved` blocks are processed once during `terraform plan`/`apply` -- remove them after the migration is applied to keep configuration clean - `terraform state` subcommands (mv, rm, pull, push) bypass safety checks -- use `moved`/`removed` blocks instead for auditable, reviewable refactoring - `data` sources are read during planning by default -- if they depend on resources being created in the same apply, use `depends_on` to defer the read - `sensitive = true` on variables prevents the value from appearing in plan output but does NOT encrypt it in state -- state encryption is a separate concern - `terraform fmt` only formats `.tf` files in the current directory -- use `terraform fmt -recursive` to format all subdirectories - `toset()` deduplicates -- if your list has duplicates, `for_each = toset(var.list)` silently drops them - `.tfvars` files are auto-loaded only if named `terraform.tfvars` or `*.auto.tfvars` -- other filenames require explicit `-var-file` flag </red_flags> --- <critical_reminders> ## CRITICAL REMINDERS > **All code must follow project conventions in CLAUDE.md** **(You MUST pin provider versions with constraints in `required_providers` and commit `.terraform.lock.hcl` to version control)** **(You MUST use a remote backend with state locking for any shared or production infrastructure)** **(You MUST use `for_each` with a map/set for non-identical resources -- `count` causes index-shift destruction on removal)** **(You MUST never store secrets in `.tf` files, `.tfvars`, or state -- use environment variables (`TF_VAR_*`) or your secrets manager)** **(You MUST run `terraform plan` and review the diff before every `terraform apply` -- never apply blindly)** **Failure to follow these rules will cause state corruption, accidental resource destruction, secret exposure, and non-reproducible infrastructure.** </critical_reminders>
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.