operating-infra
Author, inspect, troubleshoot, review, and apply (after explicit
Install
npx skills add https://github.com/alexei-led/cc-thingz/tree/master/src/skills/operating-infra
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install alexei-led-cc-thingz@llmmart
git clone https://github.com/alexei-led/cc-thingz.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole alexei-led/cc-thingz collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Operate Infrastructure
Work from files, plans, logs, and read-only commands; edit repo files freely, but touch live resources only under these rules:
- Before any cloud command, confirm identity (
aws sts get-caller-identity --profile <profile>;gcloud auth list,gcloud config list), passing profile, project, region, and zone explicitly instead of relying on CLI defaults. - Before any live change that's destructive, costly, or externally visible (apply, upgrade, rollout, delete, destroy, stop, resize, scale, IAM, bucket, network, DDL/DML, rollback): show identity, exact resources (ARNs or names), blast radius, irreversibility, and the plan/diff/inventory behind them, then wait for explicit confirmation.
- Every apply, upgrade, or rollout, regardless of blast radius: confirm the exact destination first (account, context, namespace, workspace, or release — name production explicitly), run the validation gates below on the same rendered artifact, show the plan or diff with create/modify/delete counts, and apply only that same reviewed artifact (the saved plan file or rendered manifest), only after explicit confirmation of that exact artifact and destination — never apply to production without it.
- After applying, verify rollout status, pod health, or Terraform outputs/state, and name the rollback path. On apply failure or a timed-out/degraded rollout: stop, report status and rollback options, and ask before any rollback.
- Without write access, return proposed changes (file, change, reason) instead of applying them.
For troubleshooting: rank likely causes, gather one safe signal at a time, propose the next step. For authoring: pick the smallest pattern keeping ownership, state boundaries, and least privilege.
In GitHub Actions this skill owns workflow structure, triggers, permissions, runners, actions, environments, secrets, caching, and concurrency — not the shell body of a run: step (writing-shell); mixed changes use both.
References
Load every reference that matches the stack:
- Terraform/OpenTofu files, modules, state, or plans → terraform.md
- Kubernetes manifests or
kustomization.yaml→ kubernetes.md Chart.yaml, Helm values, or chart templates → helm.md- GitHub workflow YAML → github-actions.md
Dockerfileor image build/release → dockerfile.md- AWS: EC2, ECS, Lambda, S3, RDS, IAM, CloudWatch → aws.md
- GCP: GCS, Compute Engine, IAM, Pub/Sub, Cloud SQL, quotas, Cloud Logging → gcp.md
- Cloud Run services, revisions, traffic, or logs → cloud-run.md
- BigQuery queries, tables, datasets, or cost → bigquery.md
- Linux services, hosts, processes, disks, or networks → linux.md
- Applying, upgrading, rolling out, or any deploy request, including a bare "deploy this" → deploying.md
Validation gates
Run the gates for changed types when the tools exist; report each skipped gate and why.
- Terraform/OpenTofu (or
tofuequivalents):fmt,init -backend=falsewhen possible,validate,plan,tflint,checkovortrivy config. - Kubernetes/Kustomize: render first, then
kubeconformagainst the target version, thenkube-linter,kubescape,conftest, orkyverno. - Helm:
helm lint,helm templatefor every relevant values file, the Kubernetes gates on the output, andhelm diffbefore an upgrade counts as safe. - Dockerfile/images:
hadolint,trivy. - GitHub Actions:
actionlint,zizmor. - Cloud CLI: inventory, cost estimate or dry-run when available, and IAM/quota checks before mutation.
Done when the relevant build/test/lint checks pass on what you changed, or you name each check that did not run and why.
Output
INFRA RESULT
Scope: <files/resources/environment>
Identity: <account/project/profile/region or not applicable>
Status: DONE | NEEDS CONFIRMATION | BLOCKED | FAILED
Evidence: <file:line, plan/log/status summary, command result>
Changes or proposal: <minimal change or next step>
Validation: <gate — pass/fail/skipped>
Next: <safe next action, confirmation request, or none>
Files (cc-thingz)
-
.agentbundler
-
targets
-
claude.json 945 B
{ "frontmatterPatch": { "agent": "engineer", "allowed-tools": [ "Read", "Bash", "Grep", "Glob", "Bash(terraform *)", "Bash(tofu *)", "Bash(kubectl *)", "Bash(kustomize *)", "Bash(helm *)", "Bash(docker *)", "Bash(actionlint *)", "Bash(zizmor *)", "Bash(tflint *)", "Bash(checkov *)", "Bash(trivy *)", "Bash(kubeconform *)", "Bash(kube-linter *)", "Bash(kubescape *)", "Bash(conftest *)", "Bash(hadolint *)", "Bash(syft *)", "Bash(grype *)", "Bash(cosign *)", "Bash(gcloud *)", "Bash(gsutil *)", "Bash(bq *)", "Bash(aws *)", "Bash(jq *)", "Bash(yq *)", "Bash(systemctl *)", "Bash(journalctl *)", "AskUserQuestion" ], "argument-hint": "[task | --dry-run | --apply <environment>]", "context": "fork", "user-invocable": true } }
-
-
-
references
-
aws.md 1.2 KB
# AWS Before suggesting a recursive S3 delete, check versioning, lifecycle, replication, and object count. ## Service checks - EC2: instance state, system status, security groups, subnet, IAM role, user data, attached volumes, and recent CloudWatch metrics. - ECS: cluster, service events, desired/running count, task definition, image digest, target group health, and task logs. - Lambda: runtime, timeout, memory, environment, IAM role, trigger source, recent errors, and log group. - S3: bucket policy, public access block, encryption, versioning, lifecycle, replication, and object ownership. - RDS: engine/version, storage, backups, maintenance window, parameter and subnet groups, security groups, and snapshots before risky changes. - IAM: scope policies to least privilege. Wildcard actions/resources and cross-account trust are review findings. ## Troubleshooting - Auth errors: SSO session, profile, region, permission boundary, SCPs, and resource policy. - Throttling: find the service quota/API, add pagination and backoff, and avoid unbounded list calls. - Network failures: VPC, subnet, route table, security group, NACL, DNS, and load balancer target health. - Missing logs: service role permissions, log group/stream name, region, and retention. -
bigquery.md 1.4 KB
# BigQuery ## Cost safety - Dry-run non-trivial queries first, and convert bytes scanned to cost when the user asks or the query may be large. - Set maximum bytes billed on exploratory queries. - Select only needed columns; `LIMIT` does not reduce scan cost. - Filter on the partition column of time-partitioned tables unless the user explains why a full scan is needed. Use clustered filters when they match the table design. ## Workflow - Confirm project, dataset, table, partition field, and date range, then read schema and partitioning before writing the query. - For samples, limit rows and skip expensive computed fields. - For exports and destination tables, confirm target, format, write disposition or overwrite behavior, access controls, and cost first. - DDL, DML, and table or dataset deletion are destructive work. ## Cost helper `scripts/bq-cost-check.py` runs a dry-run only; it never prompts or executes the query: ```bash uv run python scripts/bq-cost-check.py "SELECT ..." --location <location> --max-bytes <bytes> --json ``` - Default output reports exact processed bytes and GiB. - `--price-per-tib <USD>` must come from the tariff for the project's location and billing model; there is no built-in rate. `--max-usd` requires it and gives an estimate, not a billing guarantee. - Exit 1: threshold exceeded or dry-run failed. Exit 2: invalid arguments. - The preflight does not limit the later query's billing; set maximum bytes billed on that query too. -
cloud-run.md 1.1 KB
# Cloud Run ## Service model - Track service, revision, traffic split, image digest, region, service account, ingress, authentication, concurrency, CPU/memory, timeout, and min/max instances. - Use immutable image digests for production diagnosis and rollback clarity. - Keep environment variables non-secret; reference Secret Manager for secrets. - Use least-privilege runtime service accounts. Confirm ingress and invoker IAM before changing public/private access. - Deploy, traffic migration, and rollback are deployment work. ## Troubleshooting - Start with service status, latest ready revision, traffic target, recent revision errors, and logs. - Startup failures: container port, command/entrypoint, env/secret references, image architecture, and startup probe. - Request failures: authentication, ingress, URL, load balancer/serverless NEG config, timeout, concurrency, and application logs. - Cold starts or latency: min instances, CPU allocation, concurrency, image size, startup work, VPC connector, and downstream latency. - Egress failures: VPC connector, route, firewall, DNS, private service access, and service account permissions. -
deploying.md 4.5 KB
# Applying and Deploying Read this before any `apply`, `upgrade`, `rollout`, or production change to an already-validated Terraform, Helm, Kustomize, or Kubernetes artifact. ## Hard rules - Never invent deploy paths, release names, workspaces, namespaces, accounts, or environments. If one is unclear, ask. - Authorization binds to the reviewed artifact and destination (SKILL.md's apply/upgrade/rollout rule): account, context, namespace, workspace, chart/version, release, and values or var files. Production authorization names the exact environment. Ambiguous, partial, or mismatched authorization means stop and ask again. - Blocked validation, missing plan/diff evidence, or an unshown destructive change: stop before confirmation. - Never run a command with an unresolved placeholder. Keep secrets out of evidence and hashes. - Push no images and trigger no CI workflows here. - Default to validating only (dry run); apply only once the user explicitly asks to apply. - In a background or non-interactive run, never apply: validate and report only — a confirmation from earlier in the conversation doesn't authorize an apply in that run. - Asked to describe the workflow rather than run it: give the ordered steps below. Unknown details (destination, release, values) become the questions to ask, not a blocked report. ## Workflow 1. Detect the infra type and target details from repo files. 2. Run the apply evidence below for the detected type, plus the Validation gates this skill lists, on the same rendered artifact. Record each unavailable tool as a skipped check with its reason. 3. Show the pre-flight report. Stop here for a dry run. 4. Record the artifact and input hashes, then get confirmation of the exact artifact and destination unless already authorized. 5. Immediately before apply, verify the artifact and input hashes still match. Changed inputs need revalidation and renewed authorization. 6. Apply with one of the allowed commands below. 7. Verify the changed resources: rollout status, pod health, or Terraform outputs/state. ```markdown ## Pre-flight: READY | BLOCKED - Environment / type: <env> / <terraform|helm|kustomize|kubernetes> - Destination: <account/context/namespace/workspace/release> - Evidence: `<command>` — <summary> - Resources: create <n>, modify <n>, delete/replace <n> - Risks: <destructive changes, CRDs/hooks, missing evidence, or none> - Rollback: <helm rollback <release> <rev> | kubectl rollout undo | revert the commit, re-plan, and confirm again> ``` ## Apply evidence by type Validate and apply against the same explicit destination and frozen inputs. - Kubernetes: `kubectl --context <context> --namespace <namespace> diff -f <reviewed-rendered-file>`, then `apply --dry-run=server` (`--dry-run=client` only without cluster access, and say so). Confirm the namespace and any referenced Secrets/ConfigMaps exist. List deletions and immutable-field changes (selectors, PVCs) before confirmation. - Kustomize: `kustomize build <overlay> > <reviewed-rendered-file>`, then the Kubernetes evidence on that file. Apply that same file; do not rebuild the overlay between review and apply. The overlay path matches the target environment. - Helm: freeze the chart package, dependencies, and every values file/flag. `helm lint`, `helm template` for the target values, then `helm diff upgrade` (or `kubectl diff` on the rendered output without `helm-diff`). The values file matches the target environment. Call out CRDs and hooks before confirmation. Record `helm history` as the rollback target. - Terraform/OpenTofu: `fmt -check`, `init -backend=false` when safe, `validate`, then `plan -out=tfplan` and `show -no-color tfplan`. Confirm workspace, backend, and var files match the target; shared environments need a locked remote backend. List every destroy/replace before confirmation. ## Allowed apply commands - `terraform apply tfplan` - `helm upgrade --install <release> <pinned-local-chart> --kube-context <context> --namespace <namespace> --values <reviewed-values-file>` - `kubectl --context <context> --namespace <namespace> apply -f <reviewed-rendered-file>` Write deployment logs only where the repo already has that convention. ## Output After an actual apply run, start with one header, then its fields: ```text DRY RUN COMPLETE: Status; Environment; Types; Validation; Plan/Diff; Blockers; Skipped AWAITING CONFIRMATION: Environment; Type; Command; Destructive changes; Confirmation needed DEPLOYMENT COMPLETE: Environment; Type; Status; Applied; Verification; Rollback option DEPLOYMENT BLOCKED | DEPLOYMENT FAILED: Environment; Type; Reason; Evidence; Next step ``` -
dockerfile.md 987 B
# Dockerfile and Images ## Build - Use multi-stage builds for compiled languages and dependency-heavy runtimes. - Copy dependency manifests before application code to keep the cache useful. - Use minimal runtime images (distroless, slim, scratch) or a pinned base that fits debugging needs. - Pin base images by version or digest; deployable images never use `latest`. - Use `COPY` plus an explicit fetch-and-verify step instead of `ADD` for remote URLs. - Run `syft`, `grype`, and `cosign` when SBOM, vulnerability, or provenance evidence matters. ## Runtime - Run as non-root, with a read-only root filesystem when the app allows it. - Keep secrets, tokens, SSH keys, and cloud credentials out of build args and layers. - Keep health checks free of secrets and external side effects; match exposed ports to the app contract. ## `.dockerignore` Exclude VCS data, local caches, secrets, test artifacts, and build output. Keep files the build reads, such as lockfiles and metadata. -
gcp.md 1.1 KB
# GCP Use `--quiet` on destructive commands only after the user confirmed the exact resources. ## Service checks - Compute Engine: instance state, zone, machine type, boot disk, service account, tags, firewall rules, metadata, serial output, and recent logs. - GCS: IAM, public access prevention, uniform bucket-level access, versioning, lifecycle, retention, soft delete, and object count. - IAM: least privilege at the service-account level. Broad project roles and user-managed keys are review findings. - Pub/Sub: subscription backlog, dead-letter and retry policy, push endpoint health, and ack deadlines. - Cloud SQL: backups, maintenance window, flags, network exposure, IAM/database users, and connection errors before risky changes. ## Troubleshooting - Auth errors: ADC vs user credentials, service account impersonation, IAM role, and org policy. - Quota errors: report quota name, region, current limit, requested amount, and service. - Region/zone errors: verify the resource location before editing config. - Missing logs: service account permission, log filter, project, region, and retention. -
github-actions.md 1.3 KB
# GitHub Actions ## Workflow rules - Separate CI, release, deploy, and security-scan workflows when their triggers or permissions differ. - Set workflow and job permissions explicitly; default to `contents: read`. - Pin every external action, including `actions/*`, by full commit SHA with a version comment (`uses: actions/checkout@<sha> # v6`). Only local `./` actions and local reusable workflows are exempt; pin remote reusable workflows by SHA like actions. - Use OIDC for cloud auth instead of long-lived cloud keys in secrets. - Add concurrency to deployments and workflows that must not overlap. - Cache only dependency and build caches that are safe to restore across branches. - Use reusable workflows only when the contract is stable and the caller controls inputs clearly. ## Blockers - Broad permissions, unpinned actions, write tokens on `pull_request_target` or fork PRs, untrusted input interpolated into `run:`, and secret exposure. - Cloud deploy jobs without environment protection, approval gates, or an explicit project/account/region. ## Other checks - Matrix jobs: artifact names are unique, and later jobs consume the intended artifact set. - Release pipelines: include provenance, SBOM, vulnerability scan, and image signing when supply chain matters. - Run `checkov` on this workflow when it's scanned alongside other IaC changes. -
helm.md 982 B
# Helm ## When to use Helm Weigh environments, values variation, release packaging, and who owns the chart. - Helm: the app ships as a chart, has many optional components, or needs release history, rollback metadata, and chart versioning. - Kustomize: small environment deltas without templating. ## Chart rules - Keep naming and labels in helpers. - Quote `appVersion`; bump chart `version` on every chart change. - Default the image tag to `appVersion` or an explicit immutable tag. - Keep values small and flat; avoid nesting that mirrors Kubernetes YAML. - Gate optional subcharts with `condition` and `enabled` flags. - Use one chart with per-environment values files, not a chart per environment. - Reference existing secrets or an external secret manager; keep secrets out of values files. - Surface immutable-field, selector, and PVC changes and resource deletions before any upgrade. - Add chart-testing for reusable charts and helm-unittest for complex conditionals. -
kubernetes.md 1.7 KB
# Kubernetes ## Tool choice - Raw manifests: small stable resources with little environment variation. - Kustomize: overlays, environment deltas, and patching without templating. - Helm: packaged apps, third-party charts, or heavy templating. - Terraform: cluster and cloud resources whose lifecycle sits outside the Kubernetes API. ## Workload defaults - Pin image tags; never `latest`. - Set requests and sane limits; avoid limits that cause predictable throttling or OOMs. - Use readiness probes for routing; add liveness probes only when a restart is a real recovery path. - Run as non-root, disable privilege escalation, drop capabilities, and use a read-only root filesystem where possible. - Add network policies when namespace isolation matters. - Use PodDisruptionBudgets and topology spread for production availability when replicas allow it. - Keep labels stable: `app.kubernetes.io/name`, `instance`, `component`, `part-of`, and `managed-by`. - Check Service and Ingress selectors against pod labels. - Selector, PVC, and other immutable-field changes may need a migration instead of a rollout. ## Secrets and config - Keep real secret values out of manifests. Use External Secrets Operator, CSI secret drivers, SOPS, Sealed Secrets, or a cloud secret manager. - Separate config from secrets. Mount high-risk secrets as files rather than environment variables. ## Troubleshooting - Check, in order: events, rollout status, pod status, container logs, image pull errors, probes, resource pressure, and service endpoints. - Networking: Service selectors, EndpointSlices, NetworkPolicies, DNS, ingress/controller logs, and cloud load balancer state. - Scheduling: node selectors, taints/tolerations, requests, affinity, topology spread, and quota. -
linux.md 1001 B
# Linux Hosts and Services Capture current state before recovery steps that could lose data. Reboot, kill, delete, chmod/chown, package upgrades, and firewall changes need a clear target and blast radius first. ## Service troubleshooting - Check service state, recent journal entries, config paths, environment files, ports, dependencies, and restart history. - Failed starts: unit files, permissions, missing files, bind addresses, config validation, and exit codes. - Log-heavy failures: narrow by time window and service name. ## Resources - CPU: load, top processes, thread count, and recent deploys or cron jobs. - Memory: RSS, OOM killer messages, cgroups, swap, and leak patterns. - Disk: filesystem fullness, inode exhaustion, mount state, deleted-but-open files, and log growth. - Network: listener ports, DNS, routes, firewall, TLS certs, proxy config, and packet loss. Use modern tools when installed (`btop`, `duf`, `dust`, `procs`, `mtr`, `jq`/`yq`) and fall back to standard ones. -
terraform.md 1.7 KB
# Terraform and OpenTofu Use Terraform for cloud resource lifecycle, shared infrastructure, policy-controlled state, and repeatable environments, not for one-off fixes whose source of truth is elsewhere. ## Module boundaries - Foundation modules hold slow-changing shared primitives: networks, shared IAM, org policy, base logging, and shared secrets plumbing. - App/environment modules hold app-owned resources: service accounts, bindings, runtime config, queues, buckets, databases, and deploy-time wiring. - Pass explicit inputs and outputs (IDs, self-links, names, emails, regions, subnet names); never read sibling state implicitly. - Align state boundaries with ownership and rollout risk, so frequent app deploys never touch shared foundations. - Use small root modules per environment or bounded service area instead of one global state file. - Keep provider and version constraints explicit. ## Safety - A plan is evidence, not permission to apply. Surface every create/update/delete count and stop on an unexpected destroy or replace. - Run Conftest on the plan JSON when policy depends on planned values. - Use targeted apply only for narrowly explained recovery work. - Keep remote state encrypted and locked. Never commit state, plan files with secrets, or provider credentials. - Mark sensitive outputs. Keep secrets out of variable files unless they are encrypted and intentionally managed. ## Troubleshooting - Lock errors: identify the lock owner before forcing an unlock. - Drift: compare state, plan, and live resource ownership before importing or changing code. - Provider auth failures: verify identity, project/account, region, and required IAM before editing modules. - Quota failures: report requested resource, region, quota name, and current limit.
-
-
scripts
-
bq-cost-check.py 5 KB
#!/usr/bin/env python3 """Estimate BigQuery query cost before running. Usage: bq-cost-check.py "SELECT * FROM table" """ from __future__ import annotations import argparse import json import math import re import subprocess import sys BQ_TIMEOUT_SECS = 120 # Matches bq's dry-run prose: "...will process 1,234 bytes of data." Anchored # on "process ... bytes" so a version banner or job id elsewhere in the output # can't be mistaken for the byte count. BYTES_PROSE_RE = re.compile(r"process\s+([\d,]+)\s+bytes", re.IGNORECASE) def estimate_bytes(query: str, location: str | None = None) -> int: try: result = subprocess.run( [ "bq", *([f"--location={location}"] if location else []), "query", "--dry_run", "--use_legacy_sql=false", "--format=json", query, ], capture_output=True, text=True, timeout=BQ_TIMEOUT_SECS, ) except FileNotFoundError: raise SystemExit("Error: bq CLI is not installed") from None except subprocess.TimeoutExpired: raise SystemExit( f"Error: bq dry-run timed out after {BQ_TIMEOUT_SECS}s" ) from None if result.returncode != 0: sys.stderr.write(result.stdout + result.stderr) raise SystemExit("Error: bq dry-run failed") # `bq query --dry_run` prints job metadata to stderr; --format=json gives # an empty list on stdout. The byte count surfaces in stderr text: # "Query successfully validated. Assuming the tables are not modified, # running this query will process N bytes of data." # Newer bq versions also support --format=prettyjson with totalBytesProcessed # in stderr-rendered JSON. We parse both. def find(obj: object) -> int | None: if isinstance(obj, dict): if "totalBytesProcessed" in obj: value = obj["totalBytesProcessed"] if isinstance(value, bool) or not re.fullmatch(r"[0-9]+", str(value)): raise SystemExit("Error: invalid totalBytesProcessed in bq output") return int(value) for value in obj.values(): hit = find(value) if hit is not None: return hit elif isinstance(obj, list): for value in obj: hit = find(value) if hit is not None: return hit return None for blob in (result.stdout, result.stderr): try: data = json.loads(blob) except json.JSONDecodeError: match = BYTES_PROSE_RE.search(blob) if match: return int(match.group(1).replace(",", "")) else: value = find(data) if value is not None: return value raise SystemExit( "Error: could not parse bq output:\n" + result.stdout + result.stderr ) def main(argv: list[str] | None = None) -> int: parser = argparse.ArgumentParser( description="Dry-run a query; report bytes and optional cost." ) parser.add_argument("query") parser.add_argument("--location", help="BigQuery job location") parser.add_argument("--price-per-tib", type=float, help="Explicit USD per TiB rate") parser.add_argument("--max-bytes", type=int, help="Exit 1 above this byte count") parser.add_argument( "--max-usd", type=float, help="Exit 1 above this estimated USD cost" ) parser.add_argument("--json", action="store_true") args = parser.parse_args(argv) for name in ("price_per_tib", "max_bytes", "max_usd"): value = getattr(args, name) if value is not None and (value < 0 or not math.isfinite(value)): parser.error(f"--{name.replace('_', '-')} must be finite and nonnegative") if args.max_usd is not None and args.price_per_tib is None: parser.error("--max-usd requires --price-per-tib") n = estimate_bytes(args.query, args.location) cost = n / 1024**4 * args.price_per_tib if args.price_per_tib is not None else None exceeded = (args.max_bytes is not None and n > args.max_bytes) or ( args.max_usd is not None and cost is not None and cost > args.max_usd ) if args.json: print( json.dumps( { "bytes_processed": n, "location": args.location, "price_per_tib_usd": args.price_per_tib, "estimated_cost_usd": cost, "threshold_exceeded": exceeded, } ) ) else: print(f"Query will scan: {n} bytes ({n / 1024**3:.2f} GiB)") if cost is not None: print(f"Estimated cost: ${cost:.4f} at ${args.price_per_tib:g}/TiB") print("Estimate excludes free tiers, discounts, and capacity pricing.") if exceeded: print("Threshold exceeded", file=sys.stderr) return 1 if exceeded else 0 if __name__ == "__main__": sys.exit(main())
-
-
SKILL.md 5 KB
--- description: Author, inspect, troubleshoot, review, and apply (after explicit confirmation) infrastructure across IaC, Kubernetes, cloud resources, containers, CI/CD, and Linux hosts. Use when changing Terraform/OpenTofu, Kubernetes, Helm, Kustomize, Dockerfiles, GitHub Actions workflow/job/permissions semantics, AWS, GCP, Cloud Run, BigQuery, IAM, logs, instances, or service health, or when the user says "deploy", "deploy to staging", "terraform apply", "helm upgrade", "kubectl apply", "rollout", "deploy check", "validate deployment", or "validate infrastructure". NOT for shell scripts, generic command pipelines, or only the shell body inside `run:` steps (see writing-shell). name: operating-infra --- # Operate Infrastructure Work from files, plans, logs, and read-only commands; edit repo files freely, but touch live resources only under these rules: - Before any cloud command, confirm identity (`aws sts get-caller-identity --profile <profile>`; `gcloud auth list`, `gcloud config list`), passing profile, project, region, and zone explicitly instead of relying on CLI defaults. - Before any live change that's destructive, costly, or externally visible (apply, upgrade, rollout, delete, destroy, stop, resize, scale, IAM, bucket, network, DDL/DML, rollback): show identity, exact resources (ARNs or names), blast radius, irreversibility, and the plan/diff/inventory behind them, then wait for explicit confirmation. - Every apply, upgrade, or rollout, regardless of blast radius: confirm the exact destination first (account, context, namespace, workspace, or release — name production explicitly), run the validation gates below on the same rendered artifact, show the plan or diff with create/modify/delete counts, and apply only that same reviewed artifact (the saved plan file or rendered manifest), only after explicit confirmation of that exact artifact and destination — never apply to production without it. - After applying, verify rollout status, pod health, or Terraform outputs/state, and name the rollback path. On apply failure or a timed-out/degraded rollout: stop, report status and rollback options, and ask before any rollback. - Without write access, return proposed changes (file, change, reason) instead of applying them. For troubleshooting: rank likely causes, gather one safe signal at a time, propose the next step. For authoring: pick the smallest pattern keeping ownership, state boundaries, and least privilege. In GitHub Actions this skill owns workflow structure, triggers, permissions, runners, actions, environments, secrets, caching, and concurrency — not the shell body of a `run:` step (writing-shell); mixed changes use both. ## References Load every reference that matches the stack: - Terraform/OpenTofu files, modules, state, or plans → [terraform.md](references/terraform.md) - Kubernetes manifests or `kustomization.yaml` → [kubernetes.md](references/kubernetes.md) - `Chart.yaml`, Helm values, or chart templates → [helm.md](references/helm.md) - GitHub workflow YAML → [github-actions.md](references/github-actions.md) - `Dockerfile` or image build/release → [dockerfile.md](references/dockerfile.md) - AWS: EC2, ECS, Lambda, S3, RDS, IAM, CloudWatch → [aws.md](references/aws.md) - GCP: GCS, Compute Engine, IAM, Pub/Sub, Cloud SQL, quotas, Cloud Logging → [gcp.md](references/gcp.md) - Cloud Run services, revisions, traffic, or logs → [cloud-run.md](references/cloud-run.md) - BigQuery queries, tables, datasets, or cost → [bigquery.md](references/bigquery.md) - Linux services, hosts, processes, disks, or networks → [linux.md](references/linux.md) - Applying, upgrading, rolling out, or any deploy request, including a bare "deploy this" → [deploying.md](references/deploying.md) ## Validation gates Run the gates for changed types when the tools exist; report each skipped gate and why. - Terraform/OpenTofu (or `tofu` equivalents): `fmt`, `init -backend=false` when possible, `validate`, `plan`, `tflint`, `checkov` or `trivy config`. - Kubernetes/Kustomize: render first, then `kubeconform` against the target version, then `kube-linter`, `kubescape`, `conftest`, or `kyverno`. - Helm: `helm lint`, `helm template` for every relevant values file, the Kubernetes gates on the output, and `helm diff` before an upgrade counts as safe. - Dockerfile/images: `hadolint`, `trivy`. - GitHub Actions: `actionlint`, `zizmor`. - Cloud CLI: inventory, cost estimate or dry-run when available, and IAM/quota checks before mutation. Done when the relevant build/test/lint checks pass on what you changed, or you name each check that did not run and why. ## Output ```text INFRA RESULT Scope: <files/resources/environment> Identity: <account/project/profile/region or not applicable> Status: DONE | NEEDS CONFIRMATION | BLOCKED | FAILED Evidence: <file:line, plan/log/status summary, command result> Changes or proposal: <minimal change or next step> Validation: <gate — pass/fail/skipped> Next: <safe next action, confirmation request, or none> ```
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.