Claude Skill

sota-secrets-management

State-of-the-art secrets management for building and auditing software. Use whenever a task involves creating, storing, injecting, rotating, or scanning for credentials — or reviewing code/infrastructure for secret leaks and misuse. BUILD mode: implementing secrets handling; AUDI

LLM Mart · 0 points · 8 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download martinholovsky-SOTA-skills-skills_sota-secrets-management-c26df6b.zip · 53 KB
Part of martinholovsky/sota-skills — 38 skills

Install

skills CLI npx skills add https://github.com/martinholovsky/SOTA-skills/tree/main/skills/sota-secrets-management
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install martinholovsky-sota-skills@llmmart
Git git clone https://github.com/martinholovsky/SOTA-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole martinholovsky/sota-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

SOTA Secrets Management

Purpose

Eliminate static secrets where possible; where not possible, make every secret short-lived, narrowly scoped, runtime-injected, auditable, and rotatable without downtime. This skill covers the full lifecycle (generation → distribution → storage → use → rotation → revocation → expiry), storage backends, application handling patterns, leak detection, incident remediation, and per-credential-type rules. It serves two workflows: BUILD (write correct secrets handling into new or existing code) and AUDIT (sweep a repo for secret issues and report findings).

The hierarchy of preference, always:

  1. No secret at all — workload identity / OIDC federation / cloud IAM roles.
  2. Short-lived, auto-issued secret — Vault dynamic creds, STS tokens, SPIRE SVIDs.
  3. Long-lived secret in a managed backend — secret manager + rotation + audit log.
  4. Encrypted secret in the repo — SOPS+age / sealed-secrets, GitOps only.
  5. Plaintext secret anywhere — never acceptable.

When you write or review code, push the design as far up this hierarchy as the platform allows, and document why if you stop below level 2.

BUILD mode

Use when implementing anything that consumes or manages a credential.

  1. Classify the credential. Type (DB cred, API key, signing key, TLS key, …) determines the rules — read rules/05-credential-types.md for the matching section before writing code.
  2. Try to eliminate it. Cloud-to-cloud or CI-to-cloud calls should use workload identity (OIDC, IAM roles, SPIFFE) — see rules/01-lifecycle-and-workload-identity.md. Only fall back to a stored secret when no federation path exists (e.g., third-party SaaS API key).
  3. Pick the storage backend per environment using the decision table in rules/02-storage-backends.md. Never invent a custom encrypted store.
  4. Wire injection at runtime — file mount or env var populated by the platform, fetched via SDK with caching/TTL, never baked into images or code. Patterns and good/bad pairs in rules/03-application-patterns.md.
  5. Design rotation before shipping. Every secret needs: an owner, a rotation procedure that works with zero downtime (dual-secret / kid overlap), an expiry or rotation interval, and a revocation path. If you cannot answer "how do we rotate this at 3am during an incident," the design is not done.
  6. Add guardrails: pre-commit scanning config, CI secret-scan gate, .gitignore entries for .env* and key files, log-scrubbing for the new secret's shape (rules/04-detection-and-remediation.md).
  7. Self-review against the Audit checklist at the end of every rules file you used.

AUDIT mode

Use when asked to find secret leaks/misuse in an existing repo.

Sweep procedure

  1. Tooling pass (if available): run gitleaks git --redact . and/or trufflehog filesystem . (and git log history scan when the repo has history). Treat tool output as candidates, not verdicts — verify each hit.
  2. Manual grep pass for what tools miss. Sweep at minimum:
    • High-entropy strings and known prefixes: AKIA, ASIA, ghp_, gho_, ghs_, ghu_, ghr_, github_pat_, xoxb-, xoxp-, sk-, sk_live_, rk_live_, AIza, ya29., glpat-, npm_, dop_v1_, shpat_, eyJhbGciOi (inline JWTs), -----BEGIN ([A-Z0-9]+ )*PRIVATE KEY( BLOCK)?----- (ERE; rules/04 §6 has the tested command).
    • Assignment patterns: (password|passwd|pwd|secret|token|api[_-]?key|auth)\s*[:=]\s*['"][^'"]{6,}.
    • Connection strings with embedded creds: ://[^/:@\s]+:[^@\s]+@ (postgres, mysql, mongodb, amqp, redis URLs).
    • Files: .env* tracked in git, *.pem, *.p12, *.pfx, *.key, *.jks, *.keystore, id_rsa*, credentials.json, serviceaccount*.json, kubeconfig, .npmrc/.pypirc with tokens, terraform.tfstate (state files contain plaintext secrets).
  3. History pass: git log -p / gitleaks git --log-opts for secrets removed from HEAD but live in history. A secret deleted in a later commit is still leaked — severity is unchanged.
  4. Handling pass (misuse, not just leaks): secrets in log statements, error messages, exception payloads, crash/telemetry dumps, URLs/query strings, CLI args (visible in ps), Dockerfile ENV/ARG, docker-compose environment: literals, Kubernetes manifests with stringData/base64 secrets committed, CI YAML with inline values, debug endpoints dumping config, world-readable key files, missing rotation/expiry on long-lived tokens, overly broad token scopes.
  5. Verify and triage each finding: is the value real (test-shaped? placeholder? entropy?), is it currently valid, what blast radius. Never call a credential live by invoking it against production without explicit permission; judge from context.

Severity conventions

Severity Definition Examples
Critical Valid (or must-assume-valid) secret exposed to anyone with repo/log access Live cloud key in code or git history; DB password in a public image; private signing key committed
High Secret exposed in a narrower channel, or handling that will leak under normal operation Secret logged at info level; cred in CLI args; .env with real values tracked in private repo; token in URL; tfstate with secrets in VCS
Medium Weak lifecycle or weak protection of an otherwise contained secret No rotation for years; long-lived token where OIDC is available; overly broad scope; secret in env var where platform supports file mounts; world-readable key file; weak generation (low entropy)
Low Hygiene gaps with no current exposure Missing pre-commit/CI scanning; .env.example containing realistic-looking values; no .gitignore for key files; missing audit logging on secret access

Confirmed-fake placeholders (changeme, xxx, <YOUR_KEY>, obvious test fixtures) are not findings, but note them as Low if they are realistic enough to mask real leaks in scans.

Finding format

Report every finding as:

[SEVERITY] path/to/file.py:123 — RULE-ID short title
  Evidence: the offending line, with the secret value REDACTED (show prefix + length only)
  Why: one sentence of impact
  Fix: concrete remediation (rotate first, then remove; target pattern to adopt)
  Effort: trivial | small | medium | large

Order findings Critical → Low. End the audit with: counts per severity, whether git history is affected (if so, remediation must include rotation + history purge per rules/04-detection-and-remediation.md), and the top 3 systemic fixes. Never reproduce a discovered secret in full in your report — redact to first 4 chars + length.

Rules index

File Read this when...
rules/01-lifecycle-and-workload-identity.md Generating secrets (entropy/length), setting rotation/expiry/revocation policy, replacing static secrets with OIDC federation, SPIFFE/SPIRE, cloud IAM roles, GitHub Actions OIDC
rules/02-storage-backends.md Choosing where a secret lives: Vault/OpenBao, AWS/GCP/Azure secret managers, SOPS+age, sealed-secrets/external-secrets in Kubernetes, env vars vs file mounts (no cross-deployment shared mounts), in-memory handling and zeroization, consumer-key end-to-end encryption (§8)
rules/03-application-patterns.md Writing app code that consumes secrets: config layering, runtime injection, caching/TTL, keeping secrets out of code/VCS/logs/errors/URLs/argv/crash dumps, per-env separation, least-privilege scoping, access audit logging; never-persist-raw (§2.1) — text whose producer you do not control is stored as a digest plus a structural skeleton, because shape redaction is an enumeration
rules/04-detection-and-remediation.md Setting up gitleaks/trufflehog, pre-commit hooks, CI gates; responding to a leak (rotate-first, purge history, assume compromised); honeytokens; running an AUDIT sweep; agent session transcripts as a credential store (§7) — inventorying and scoping the scan of ~/.claude/projects/**, and why a tool that sweeps it is a secret-processing tool
rules/05-credential-types.md Handling a specific credential class: DB creds, API keys, signing keys, TLS private keys, SSH keys, JWT secrets and kid rotation, data keys vs KMS envelope encryption, .env discipline

Top-10 non-negotiables

  1. A leaked secret is rotated first, scrubbed second. History rewriting without rotation is theater — assume every secret that ever touched git, logs, or a ticket is compromised.
  2. No secrets in source code or VCS, ever — including "temporarily," tests against real services, example files with live values, and committed .env files.
  3. Prefer no secret to a managed secret: if OIDC federation / IAM roles / SPIFFE can replace a static credential (CI→cloud, service→cloud, pod→service), use it. Static keys for cloud access from CI are a defect, not a choice.
  4. Every secret has an expiry or rotation interval and a documented zero-downtime rotation procedure (dual-secret overlap, JWT kid, DB dual users). Unrotatable = misdesigned.
  5. Generate secrets with a CSPRNG, ≥256 bits of entropy for opaque tokens (≥32 random bytes before encoding); never derive from timestamps, UUIDv4-as-secret, or human-chosen strings.
  6. Secrets never appear in: logs, error messages, exception payloads, crash dumps, URLs or query strings, CLI arguments, ps output, Dockerfile layers, image env, shell history, or telemetry. Scrub at the logger and wrap in redacting types.
  7. Inject at runtime, never at build time. Images, artifacts, and bundles are secret-free; the platform (orchestrator, secret-manager SDK, CSI driver) supplies values when the process starts — file mounts preferred over env vars where supported.
  8. Least privilege and per-environment separation: one credential per consumer per environment, scoped to the minimum actions/resources; dev/staging/prod never share secrets and prod values are unreadable from non-prod.
  9. Application data encryption uses KMS envelope encryption (encrypt data with a data key, wrap the data key with a KMS key); never hardcode or hand-manage raw encryption keys.
  10. Scanning is mandatory and layered: pre-commit (gitleaks/trufflehog) on every developer machine, a blocking CI gate, and periodic full-history scans. Detection without a leak runbook is incomplete — keep the rotate→revoke→purge→monitor runbook current.
Files (sota-skills)
  • rules
    • 01-lifecycle-and-workload-identity.md 21.1 KB
      # 01 — Secret Lifecycle & Workload Identity
      
      Scope: generation, distribution, storage policy, rotation, revocation, expiry; replacing static
      secrets with workload identity (OIDC federation, SPIFFE/SPIRE, cloud IAM roles, GitHub Actions
      OIDC). Read this before creating any new credential.
      
      ## 1. Generation
      
      **Use a CSPRNG. Always.** `secrets` (Python), `crypto.randomBytes` (Node), `crypto/rand` (Go),
      `SecureRandom` (Java/Ruby), `openssl rand`. Never `random`, `Math.random()`, `rand()`,
      timestamps, PIDs, or hashes of any of those — they are predictable and recoverable.
      
      **Entropy floor: 256 bits (32 random bytes) for opaque tokens** (API keys, session secrets,
      webhook signing secrets, HMAC keys). 128 bits is the absolute minimum for short-lived,
      rate-limited tokens; default to 256 because the cost is zero. Encode with base64url or hex —
      length on the wire is irrelevant; entropy before encoding is what counts.
      
      **UUIDv4 is not a secret.** It has 122 bits of randomness but many generators are not
      cryptographically seeded, and UUIDs leak into logs/URLs by convention. Use a real random token.
      
      ```python
      # BAD — predictable, low entropy, leaks intent into the value
      api_key = hashlib.md5(f"{user_id}{time.time()}".encode()).hexdigest()
      reset_token = str(uuid.uuid4())
      
      # GOOD — 256-bit CSPRNG token, with an identifiable prefix for scanners
      api_key = "myapp_sk_" + secrets.token_urlsafe(32)
      reset_token = secrets.token_urlsafe(32)
      ```
      
      **Prefix your own tokens** (`myapp_sk_`, `myapp_pat_`): prefixes enable secret scanners
      (gitleaks, GitHub secret scanning partner program) to detect leaks of *your* credentials, and
      make audit greps trivial. The prefix carries zero entropy cost.
      
      **Asymmetric keys:** Ed25519 by default for signing (SSH, artifact signing; JWT algorithm choice: rules/05 §6);
      ECDSA P-256 where Ed25519 unsupported; RSA only for legacy interop and then ≥3072 bits.
      Generate keys *where they will live* (HSM, KMS, TPM, target host) so the private key never
      transits — `aws kms create-key`, `ssh-keygen` on the client, CSR-based TLS issuance. A private
      key that was ever in a chat, email, or ticket is compromised.
      
      **Human passwords are not machine secrets.** Machine-to-machine credentials are never
      human-memorable strings. If a human must create a shared secret (rare), generate it with a
      password manager at ≥24 random characters.
      
      ## 2. Distribution
      
      **Secrets move through exactly one channel: the secret store.** Producer writes to
      Vault/cloud secret manager; consumer reads from it with its own identity. Never distribute via
      Slack, email, tickets, wikis, READMEs, or "I'll paste it in the PR comment." Any secret that
      traversed such a channel is leaked — rotate it.
      
      **Bootstrap (the "secret zero" problem):** the credential that lets a workload reach the secret
      store must itself not be a static secret. Solve it with platform identity:
      
      - Cloud VMs/containers: instance metadata → IAM role (no credential at all).
      - Kubernetes: projected service account tokens → Vault Kubernetes auth / cloud workload identity.
      - Bare metal / multi-cloud: SPIFFE/SPIRE node attestation (TPM, cloud metadata, join token).
      - CI: OIDC token from the CI provider (see §5).
      
      If a design document contains the phrase "we'll put the bootstrap token in an env var," send it
      back.
      
      **Human access** to production secrets goes through the secret store's UI/CLI with SSO + MFA,
      is break-glass only, and is audit-logged. Day-to-day operation should never require a human to
      *see* a secret value.
      
      **When a password must reach a person** (a new account, a vendor portal, a device PIN), never
      send it in the same message or channel as the username. Prefer a one-time, short-lived
      set-password link so no reusable secret travels at all; failing that, deliver the value over a
      mutually authenticated channel or a second, separate channel (username by email, password via
      a password-manager share or an authenticated messaging app) and force a change on first use
      (generation and first-use rules: sota-code-security rules/02 §5). OWASP: Secrets Management
      cheat sheet.
      
      ## 3. Rotation, revocation, expiry
      
      **Every secret has, at creation time:** an owner (team), a maximum lifetime or rotation
      interval, a documented zero-downtime rotation procedure, and a revocation path. Record these in
      the secret's metadata/tags. A secret missing any of the four is an audit finding (Medium).
      Those four are the floor; the full record also carries the secret **type** and **purpose**,
      its **consumers** and who may read it, **dependencies that break on rotation**, an
      **incident contact**, the **exposure impact** and **data classification** that set its
      handling, and lifecycle timestamps **with the actor** for creation, first use, each rotation
      and deletion (most stores record the timestamps in their audit log; the metadata points to
      them rather than duplicating them). OWASP: Secrets Management cheat sheet.
      
      **Rotation intervals (defaults, tighten for higher-value targets):**
      
      | Credential | Max lifetime / rotation |
      |---|---|
      | Cloud STS / OIDC-issued tokens | 15 min – 1 h (automatic) |
      | Vault dynamic DB creds | TTL ≤ 24 h (automatic) |
      | Service-to-SaaS API keys | 90 days |
      | Webhook/HMAC signing secrets | 180 days, dual-secret overlap |
      | JWT signing keys | 90 days via `kid` rotation (see rules/05) |
      | TLS leaf certs | ≤ 90 days (ACME automation) |
      | Long-lived static keys (last resort) | 90 days, with a ticket explaining why they exist |
      
      **Event triggers rotate now, whatever the calendar says.** Two apply to every class: suspected
      compromise, and **a person who could read the value leaving the organisation or changing role**.
      Revoking their account does not revoke what they may have copied — a shared database password,
      a team API key, an encryption key they once exported. So the offboarding checklist names every
      shared credential the person could reach and rotates each (for an encryption key: new key
      version, then re-wrap/re-encrypt, rules/05 §7). The cheaper fix is structural: per-identity or
      dynamic credentials (§4–6) leave nothing shared to rotate — disabling the identity is enough.
      OWASP: Cryptographic Storage cheat sheet; Database Security cheat sheet.
      
      **Zero-downtime rotation = overlap, not swap.** The universal pattern:
      
      1. Issue new secret (version N+1) alongside old (N) — both valid. Stage it as *pending*,
         not current.
      2. **Prove N+1 works before promoting it**: set it on the target system, then authenticate
         with it. Only a passing test moves the "current" label; a failing one leaves N current and
         alerts. AWS Secrets Manager's rotation functions encode this as createSecret → setSecret →
         testSecret → finishSecret, with the new value held under `AWSPENDING` until finishSecret
         moves `AWSCURRENT`; a custom rotation function whose test step is empty skips the only
         check that the new credential is usable.
      3. Roll out consumers to N+1 (redeploy or let TTL-based cache refresh pick it up).
      4. Verify no traffic uses N (audit logs, metrics) — this proves N is unused, not that N+1
         works; that was step 2.
      5. Revoke N.
      
      Verifiers (webhook receivers, JWT validators) must accept *both* during the window; issuers
      switch to N+1 immediately. If your system can only hold one value at a time, fix that before
      setting a rotation schedule — otherwise rotation means an outage and will therefore never happen.
      
      **Revocation must be possible in minutes, not days.** Test it: for each credential class, know
      the exact command/console action that kills it (`aws iam update-access-key --status Inactive`,
      Vault lease revoke, CRL/OCSP or short-lived certs, API key delete endpoint). If revocation
      requires "contact the vendor," shorten the lifetime to compensate.
      
      **Expiry beats rotation.** A credential that dies on its own at T+1h needs no rotation calendar,
      no cleanup job, and limits blast radius automatically. This is the core argument for §4–5.
      
      ```yaml
      # BAD — static key created once, lives forever, rotation is a wiki page nobody reads
      aws_access_key_id: AKIA****************   # created 2022, owner unknown
      
      # GOOD — secret metadata makes lifecycle enforceable
      metadata:
        owner: payments-team
        type: vendor-api-key
        purpose: card capture in checkout
        consumers: [checkout-api, refunds-worker]
        readers: [role/checkout-api-prod, role/refunds-prod]
        breaks_on_rotation: [refunds-worker webhook verifier]
        incident_contact: payments-oncall
        classification: restricted          # drives handling and exposure impact
        exposure_impact: live charges, card-network fines
        rotation_interval: 90d
        rotation_runbook: runbooks/rotate-stripe-key.md
        expires: 2026-09-01
        # created/rotated/deleted timestamps + actor: from the store's audit log, not hand-kept
      ```
      
      ## 4. Short-lived beats long-lived — the workload identity ladder
      
      A static secret is a liability with a half-life; an identity-issued credential is a claim that
      expires. Whenever both ends of a connection can speak to a common trust authority, **eliminate
      the static secret entirely**:
      
      | Connection | Replace static secret with |
      |---|---|
      | App on AWS → AWS service | IAM role (instance profile / IRSA / EKS Pod Identity / ECS task role) |
      | App on GCP → GCP service | Attached service account (metadata server), Workload Identity on GKE |
      | App on Azure → Azure service | Managed Identity (system- or user-assigned) |
      | CI job → cloud account | CI OIDC federation (GitHub Actions, GitLab, CircleCI, Buildkite → AWS/GCP/Azure) |
      | CI job → package registry | Trusted publishing: per-job CI OIDC exchanged for a short-lived publish grant (npm, PyPI, RubyGems, crates.io) — no stored publish token |
      | Cross-cloud (GCP app → AWS) | Workload identity federation: exchange GCP-issued OIDC token for AWS STS creds |
      | Pod → Vault/OpenBao | Kubernetes auth (projected SA token) or JWT/OIDC auth |
      | Service → service (mTLS) | SPIFFE/SPIRE-issued X.509 SVIDs, auto-rotated |
      | App → database | Vault dynamic creds, or cloud IAM DB auth (RDS IAM, Cloud SQL IAM, Azure AD auth) |
      
      **Detection rule (AUDIT):** a long-lived cloud access key (`AKIA…`, GCP service account JSON
      key file, Azure client secret) used *from inside that same cloud or from a major CI provider*
      is a Medium finding minimum — the federation path exists and isn't used. The same key in VCS is
      Critical.
      
      **SPIFFE/SPIRE** for platform-agnostic identity: SPIRE attests nodes (cloud metadata, TPM,
      join token) and workloads (k8s SA, unix attestor, binary hash), then issues short-lived X.509
      or JWT SVIDs with SPIFFE IDs (`spiffe://trust-domain/ns/prod/sa/payments`). Use it when you
      need mTLS service identity across heterogeneous infrastructure without a cloud IAM common
      denominator. SVIDs rotate automatically (default ~1h) — workloads consume them via the Workload
      API socket, never from disk-managed certs.
      
      ## 5. GitHub Actions OIDC (and CI federation generally)
      
      **Never store cloud keys in CI secret variables when the provider supports OIDC.** GitHub
      Actions, GitLab CI, CircleCI, Bitbucket, and Buildkite all issue per-job OIDC tokens; AWS, GCP,
      and Azure all accept them. The same applies to package publishing: npm, PyPI, RubyGems, and
      crates.io accept per-job CI OIDC ("trusted publishing") — a stored long-lived publish token in
      CI where the registry supports trusted publishing is a Medium finding (see rules/05 §8).
      
      ```yaml
      # BAD — long-lived key stored in repo/org secrets; any workflow (or exfiltrating
      # dependency in any workflow) can use it forever
      - uses: aws-actions/configure-aws-credentials@v6
        with:
          aws-access-key-id: ${{ secrets.AWS_ACCESS_KEY_ID }}
          aws-secret-access-key: ${{ secrets.AWS_SECRET_ACCESS_KEY }}
      
      # GOOD — per-job 1h credential, no stored secret, scoped by trust policy
      permissions:
        id-token: write
        contents: read
      steps:
        - uses: aws-actions/configure-aws-credentials@v6  # v6+: needs node24-capable runners
          with:
            role-to-assume: arn:aws:iam::123456789012:role/deploy-myapp
            aws-region: eu-central-1
      ```
      
      **Lock the trust policy down** — this is where OIDC deployments go wrong:
      
      - Condition on `aud` (`sts.amazonaws.com`) **and** `sub`. Pin `sub` to repo + ref or
        environment: `repo:my-org/my-app:ref:refs/heads/main` or
        `repo:my-org/my-app:environment:prod`. A trust policy matching `repo:my-org/*:*` lets any
        repo in the org assume the deploy role — High finding.
      - **Immutable subject claims:** repos created, renamed or transferred after 2026-07-15 get a
        `sub` carrying numeric IDs (`repo:my-org@123456/my-app@456789:ref:refs/heads/main`); older
        repos keep the name-only form unless the owner opts in. A condition in the old form stops
        matching after the switch (the role is denied, not widened), and a name-only condition trusts
        whoever holds that name later. Pin the `@id` form, or `repository_id`/`repository_owner_id`
        where the cloud can condition on them (source: docs.github.com OIDC reference).
      - One role per repo/purpose, least-privilege policy (deploy role ≠ admin).
      - Use GitHub *environments* with required reviewers for prod-deploy roles so the OIDC `sub`
        claim can't be minted from an unreviewed branch.
      - **Keep the human behind the job attributable.** A CI call to the store or cloud should be
        traceable to the person who triggered or defined the run, not only to the deploy role.
        GitHub's OIDC token carries `actor`, `actor_id` and `run_id` claims; map them into what the
        target records (a GCP WIF attribute mapping may use any token claim). On AWS,
        `configure-aws-credentials` names every session `GitHubActions` by default and applies no
        session tags on the OIDC path, so set `role-session-name` to include the run id — it is
        workflow-chosen, so join it to the CI provider's own run log for the actor rather than
        trusting it alone.
      - **Re-verify the federation config on a schedule** (quarterly, and after any org/repo rename
        or transfer): read every trust policy / attribute condition back and confirm `sub` pinning
        and the role-per-repo mapping still hold — they drift silently. OWASP: Secrets Management
        cheat sheet.
      
      The same pattern applies to GCP Workload Identity Federation (attribute conditions on
      `assertion.repository`) and Azure federated credentials (subject identifier pinning):
      
      ```hcl
      # GCP WIF — GOOD: provider restricted to one repo, mapped to a dedicated SA
      resource "google_iam_workload_identity_pool_provider" "github" {
        attribute_condition = "assertion.repository == 'my-org/my-app'"
        attribute_mapping   = { "google.subject" = "assertion.sub" }
        oidc { issuer_uri = "https://token.actions.githubusercontent.com" }
      }
      # BAD: no attribute_condition -> any GitHub repo on earth can attempt the exchange,
      # gated only by IAM bindings you must get perfect everywhere
      ```
      
      **Residual CI secrets** (the SaaS API keys OIDC can't replace): store in the CI provider's
      secret store scoped to the narrowest level (environment > repo > org), mark masked/protected,
      and remember CI secrets are exposed to *every step* of jobs that receive them — including
      third-party actions/orbs. Pin third-party actions by commit SHA, not tag; a retagged action is
      a secret-exfiltration vector with your token in hand.
      
      ## 6. Dynamic secrets in practice
      
      The pattern that operationalizes "short-lived beats long-lived" for things OIDC doesn't cover:
      
      ```hcl
      # Vault database engine — each app instance gets its own 1h DB user
      resource "vault_database_secret_backend_role" "payments" {
        backend             = "database"
        name                = "payments-prod"
        db_name             = "payments-pg"
        creation_statements = [
          "CREATE ROLE \"{{name}}\" WITH LOGIN PASSWORD '{{password}}' VALID UNTIL '{{expiration}}';",
          "GRANT SELECT, INSERT, UPDATE ON ALL TABLES IN SCHEMA app TO \"{{name}}\";"
        ]
        default_ttl = 3600
        max_ttl     = 14400
      }
      ```
      
      Operational rules for dynamic/leased credentials:
      
      - **Renew before expiry, re-fetch on revocation.** Use the agent/SDK renewal loop (Vault Agent,
        AWS SDK credential providers do this natively); hand-rolled consumers must renew at ~2/3 of
        TTL and treat renewal failure as "fetch a new lease," not "crash."
      - **Size `max_ttl` to your deploy cadence** — if leases outlive deployments, every deploy
        naturally rotates creds and `max_ttl` is a backstop, not a disruption.
      - **Revocation drills:** revoke a live lease in staging quarterly and confirm the app recovers
        via re-fetch (rules/03 §4 retry-on-auth-failure). A dynamic-secrets setup that's never been
        revoked in anger usually breaks the first real time.
      - **Watch lease floods:** a crash-looping pod minting a new DB user per restart can exhaust
        connection/user limits — alert on lease-creation rate per role.
      
      ## 7. Storage policy summary
      
      (Backends in detail: rules/02.) Lifecycle rules that apply regardless of backend:
      
      - **One source of truth per secret.** Copies in two stores will diverge and one will rot
        unrotated. Replicate via the store's own mechanism (e.g., external-secrets sync), never by
        hand.
      - **Version every secret** and keep N-1 readable during rotation windows only.
      - **Tag with owner + rotation metadata** (§3) so expiring-secret reports are automatable.
      - **Audit log every read** in production; alert on reads from unexpected principals. The
        minimum field set and per-stage detections are rules/03 §8.
      - **Maintain a secret inventory.** You cannot rotate what you don't know exists. The secret
        store's listing *is* the inventory only if everything lives there — which is the point.
        Quarterly: export all secrets with age + owner + last-accessed; flag anything unowned,
        unread in 90 days (delete candidates — unused secrets are pure liability), or older than its
        rotation interval (overdue).
      - **Decommissioning:** when a service dies, its secrets are revoked the same week — not left
        "in case we roll back." Tie secret deletion into service-retirement checklists.
      
      ```text
      # Quarterly inventory review — the three queries that matter
      1. Secrets with no owner tag            -> assign or delete
      2. Secrets not read in 90d (audit logs) -> delete (after a deprecation notice)
      3. Secrets older than rotation_interval -> rotate now, fix the automation that missed them
      ```
      
      ## Audit checklist
      
      - [ ] All token/key generation uses a CSPRNG with ≥256-bit entropy; no UUIDs, timestamps, or
            hashes-of-predictables used as secrets.
      - [ ] Internally minted tokens carry a scannable prefix.
      - [ ] Asymmetric keys are Ed25519/P-256 (RSA ≥3072 only for interop) and generated where they
            live; no private key ever moved through chat/email/tickets.
      - [ ] No secret distributed outside the secret store; bootstrap uses platform identity, not a
            static "secret zero."
      - [ ] Every secret has owner, rotation interval, zero-downtime rotation runbook, and a tested
            revocation path; intervals meet the table in §3.
      - [ ] Rotation uses overlap (dual-secret / versioned), not in-place swap.
      - [ ] **Rotation tests the pending credential before promoting it (§3)** — High for a rotation
            function that promotes without a test step. Handlers with a finish step but no test step:
            `grep -rlE 'finish_?[Ss]ecret' . | xargs -r grep -LE 'test_?[Ss]ecret'`
            (an empty test body passes this probe — read the hits and the non-hits' test step).
      - [ ] Secret metadata carries type, purpose, consumers, readers, rotation dependencies,
            incident contact, classification/exposure impact, with lifecycle timestamps and actors
            traceable in the store's audit log (§3) — Low per missing field class.
      - [ ] **Initial human passwords never travel with the username (§2)** — Medium for a
            notification template that renders a password value:
            `grep -rniE '\{\{[[:space:]]*[a-z_.]*(password|passwd|pwd)[a-z_]*[[:space:]]*\}\}' templates/ | grep -viE '(url|link)[[:space:]]*\}\}'`
      - [ ] **CI access to the store/cloud is attributable to the triggering person (§5)** and the
            federation config is re-read on a schedule — Low for AWS OIDC jobs left on the constant
            default session name: `grep -rlE 'configure-aws-credentials' .github/workflows | xargs -r grep -L 'role-session-name'`
      - [ ] Rotation runbooks and the offboarding checklist trigger rotation of every shared secret a
            leaver or role-changer could read, not only on a schedule or confirmed compromise (§3) —
            Medium. Runbooks that never mention it:
            `grep -rLiE 'offboard|leaver|leaves|departure' runbooks/`
      - [ ] No long-lived cloud keys where IAM roles / managed identity / workload identity
            federation are available (in-cloud workloads, CI jobs, cross-cloud calls).
      - [ ] CI→cloud auth uses OIDC with `aud` + pinned `sub` trust conditions; one least-privilege
            role per repo/purpose; prod roles gated by reviewed environments.
      - [ ] Package publishing from CI uses trusted publishing (registry OIDC) where the registry
            supports it; no long-lived publish tokens stored in CI.
      - [ ] Database access uses dynamic creds or IAM DB auth where the engine supports it; leased
            credentials renew at ~2/3 TTL and recover from revocation; revocation drills performed.
      - [ ] Third-party CI actions/orbs pinned by SHA in secret-bearing jobs; CI secrets scoped to
            environment level, masked, and absent from fork-triggered workflows.
      - [ ] Secret reads are audit-logged in production with alerting on anomalous principals.
      - [ ] A secret inventory exists (store listing + owner/age/last-read), reviewed quarterly;
            unused and orphaned secrets deleted; retired services' secrets revoked promptly.
      
    • 02-storage-backends.md 19.1 KB
      # 02 — Storage Backends
      
      Scope: choosing and operating the place a secret lives — Vault/OpenBao, cloud secret managers,
      SOPS+age for GitOps, Kubernetes secret delivery (sealed-secrets, external-secrets, CSI), env
      vars vs file mounts, and in-memory handling. Read this when deciding *where* a secret goes.
      
      ## 1. Decision table
      
      | Situation | Backend |
      |---|---|
      | Single-cloud app (AWS/GCP/Azure) | That cloud's secret manager + IAM; lowest operational cost, native audit |
      | Multi-cloud / on-prem / dynamic creds / PKI / transit encryption needed | Vault or OpenBao |
      | Kubernetes consuming secrets from any of the above | External Secrets Operator or Secrets Store CSI driver pulling from the backend |
      | GitOps repo must carry secret material (no reachable store at deploy time) | SOPS+age (preferred) or sealed-secrets |
      | Local development | Per-dev `.env` (untracked) with fake/dev-scoped values, or `vault`/cloud CLI with dev profile; never shared prod values |
      | Raw encryption keys for app data | Never stored directly — KMS envelope encryption (rules/05 §7) |
      
      Rules that override the table:
      
      - **Never build a custom secret store** (encrypted DB column with a key in config, "our own
        crypto service," secrets in Redis/Consul KV without encryption + ACLs). Custom stores fail at
        audit logging, rotation, and access control — the parts that matter.
      - **Terraform/Pulumi state contains plaintext secrets** regardless of backend. State must live
        in an encrypted, access-controlled backend (S3+SSE+restricted IAM, Terraform Cloud) and never
        in VCS. Prefer ephemeral resources / `write-only` arguments (Terraform ≥1.11) so secret values
        never enter state at all.
      
      ## 2. Vault / OpenBao
      
      OpenBao is the open-source fork (post Vault-BSL), now a Linux Foundation project (verify
      the latest release at openbao.org); operationally interchangeable below. Choose
      Vault/OpenBao when you need any of: **dynamic secrets** (DB creds, cloud creds minted on
      demand), **PKI** (internal CA with short-lived certs), **transit** (encryption-as-a-service so
      apps never hold keys), or a single store spanning clouds/on-prem.
      
      Non-negotiables:
      
      - **Auth via platform identity, not tokens:** Kubernetes auth (projected SA tokens with
        audience), AWS/GCP/Azure auth methods, OIDC/JWT for CI, AppRole only as a last resort and then
        with response-wrapped SecretIDs and short TTLs. A long-lived Vault token in an env var
        defeats the point — High finding.
      - **Policies are deny-by-default, path-scoped, one per consumer.** A policy granting
        `path "secret/*" { capabilities = ["read"] }` is a Medium finding.
      - **Prefer dynamic engines over KV.** A KV entry holding a static DB password is the fallback;
        the `database/` engine issuing 1h creds per pod is the target (rules/05 §1).
      - **Short token/lease TTLs** (≤1h default, renewable) so a stolen token expires; audit devices
        enabled and shipped off-box; auto-unseal via cloud KMS, recovery keys Shamir-split among ≥3
        officers and never stored together.
      - **Treat the vault itself as a tier-0 asset** (same bar as the IdP, identity-access rules/02):
        run it **HA** (Raft ≥3 nodes / clustered) so one node loss isn't an outage; take **encrypted
        Raft/storage snapshots on a schedule, ship them off-box, and DR-restore-drill** them — an
        unrecoverable secrets store is a total outage *and* the place every dynamic credential dies;
        define a **break-glass** path (sealed root token / extra recovery-key quorum, stored offline)
        for when auth backends or the network are down, and **log + alert on every use** of it.
      
      ```hcl
      # BAD — god policy bound to a token pasted into CI variables
      path "secret/*" { capabilities = ["read", "list"] }
      
      # GOOD — per-service, per-env policy bound to Kubernetes auth role
      path "secret/data/prod/payments/*" { capabilities = ["read"] }
      # role binding: namespace=prod, serviceaccount=payments, token_ttl=1h
      ```
      
      ## 3. Cloud secret managers
      
      **AWS Secrets Manager** (rotation Lambdas, RDS-integrated rotation, cross-account access via
      resource policies; use SSM Parameter Store `SecureString` for cheap non-rotating config-grade
      secrets). **GCP Secret Manager** (versions + IAM conditions, CMEK optional). **Azure Key Vault**
      (secrets/keys/certs in one service; use RBAC mode, not legacy access policies).
      
      Rules:
      
      - **IAM is the perimeter.** Grant `GetSecretValue`-equivalent per secret ARN/name per principal.
        `secretsmanager:*` on `*`, or project-wide `secretmanager.secretAccessor`, is a Medium finding
        (High in prod). Separate read (apps) from write/rotate (rotation function, admins).
      - **Reference by name + stage/version label** (`AWSCURRENT`/`latest`) so rotation needs no
        deploy; pin exact versions only for break-glass rollback.
      - **Use native rotation** where it exists (Secrets Manager rotation functions for RDS/Redshift/
        DocumentDB); otherwise schedule rotation via your own function and the dual-secret pattern.
      - **The rotation function is a privileged deputy — lock its role.** Its execution role is
        assumable only by that function (not shared with other functions or humans), its policy
        names the exact secret ARNs it rotates (AWS's examples also pin KMS use with a
        `kms:EncryptionContext:SecretARN` condition), and only the secret manager may invoke it.
        The staged create → set → **test** → promote sequence it must run is rules/01 §3.
      - **Enable and route audit logs** (CloudTrail data events, GCP Data Access logs, Key Vault
        diagnostics) — reads must be attributable. Alert on `GetSecretValue` from unexpected
        principals or regions.
      - **Replicate via the service** (multi-region replication) rather than copying values into a
        second region's store by hand.
      
      ## 4. SOPS + age (GitOps)
      
      Use when the deployment model is "git is the source of truth" and a runtime secret store isn't
      reachable at render time (Flux/Argo without ESO, edge clusters, bootstrap configs).
      
      - **Encrypt with age recipients or cloud KMS keys; commit only the encrypted file.** With KMS
        recipients you get IAM-controlled decryption + audit logs — prefer KMS recipients for prod,
        age keys for personal/dev.
      - **`creation_rules` in `.sops.yaml` per environment**, different keys per env, so a dev key
        cannot decrypt prod files. `encrypted_regex: '^(data|stringData)$'` for k8s manifests keeps
        diffs reviewable while ciphering values.
      - The **age private key is the new crown jewel**: it lives in the secret manager or KMS, never
        in the repo, never shared in chat. Losing control of it = every encrypted file in history is
        plaintext; rotate recipients with `sops updatekeys` and re-encrypt, then rotate the *secrets
        themselves* (history still holds old ciphertexts decryptable by the leaked key).
      - Flux has native SOPS support; Argo CD needs a plugin — confirm decryption happens in the
        controller, not in a CI step that writes plaintext manifests to an artifact (High finding).
      
      ```yaml
      # .sops.yaml — GOOD: per-env keys, value-only encryption
      creation_rules:
        - path_regex: k8s/prod/.*\.yaml
          encrypted_regex: ^(data|stringData)$
          kms: arn:aws:kms:eu-central-1:123456789012:key/prod-sops
        - path_regex: k8s/dev/.*\.yaml
          encrypted_regex: ^(data|stringData)$
          age: age1devteamkeyq...
      ```
      
      ## 5. Kubernetes secret delivery
      
      Kubernetes `Secret` objects are base64-encoded, **not encrypted**, in etcd by default. Baseline
      hardening regardless of delivery mechanism: enable etcd encryption at rest with a KMS provider,
      RBAC-restrict `get`/`list`/`watch` on secrets (no wildcard cluster roles), and never commit a
      `Secret` manifest with real `data:`/`stringData:` to git (Critical if real values).
      
      Delivery options, in order of preference:
      
      1. **External Secrets Operator (ESO):** syncs from Vault/AWS/GCP/Azure into k8s Secrets via
         `ExternalSecret` CRs. Source of truth stays in the real store; rotation propagates via
         `refreshInterval`. The committed CR contains only *references* — safe for GitOps.
      2. **Secrets Store CSI driver:** mounts secrets as files directly from the backend, optionally
         without creating a k8s Secret at all (strongest: nothing in etcd). Use when apps can read
         files and you want zero etcd footprint.
      3. **Vault Agent / Vault Secrets Operator:** sidecar/operator renders secrets to a shared
         memory volume with template support and lease renewal — best for dynamic creds.
      4. **sealed-secrets:** asymmetric-encrypted `SealedSecret` CRs committed to git; controller
         decrypts in-cluster. Simple, but the controller key becomes critical state (back it up,
         rotate it), secrets are cluster/namespace-bound, and rotation means re-sealing commits.
         Prefer ESO when a real backend exists; prefer SOPS when you want backend-free GitOps with
         easier key handling.
      
      ```yaml
      # ExternalSecret — GOOD: git holds only a reference; rotation flows automatically
      apiVersion: external-secrets.io/v1
      kind: ExternalSecret
      metadata: { name: payments-db, namespace: prod }
      spec:
        refreshInterval: 5m
        secretStoreRef: { name: aws-prod, kind: ClusterSecretStore }
        target: { name: payments-db, creationPolicy: Owner }
        data:
          - secretKey: password
            remoteRef: { key: prod/payments/db, property: password }
      ```
      
      Comparison when choosing among the GitOps-capable options:
      
      | | ESO | CSI driver | sealed-secrets | SOPS |
      |---|---|---|---|---|
      | Source of truth | External store | External store | Git (encrypted) | Git (encrypted) |
      | Lands in etcd | Yes (as Secret) | Optional/no | Yes | Yes (after decrypt) |
      | Rotation effort | Automatic (refresh) | Automatic (remount) | Re-seal + commit | Re-encrypt + commit |
      | Needs reachable backend at runtime | Yes | Yes | No | No (decrypt at apply) |
      | Critical key to protect | Backend creds (workload identity) | Same | Controller keypair | age/KMS key |
      
      Consumption: **mount as files (`volumeMounts` from Secret/CSI), not env vars**, when the app
      permits — see §6. Set `defaultMode: 0400` on secret volumes. Remember mounted Secrets update
      in-place (~1m propagation) but `subPath` mounts do **not** update — avoid `subPath` for
      secrets you intend to rotate without restarts.
      
      ## 6. Env vars vs file mounts
      
      **Default to file or in-memory delivery; an env var is the exception you justify.** Some
      guidance forbids env vars for secrets outright; the position here is narrower but the default
      is the same — a file mount (tmpfs) or a fetch into process memory from the store (rules/03 §4).
      Know why, and apply the table:
      
      | Property | Env var | File mount (tmpfs) |
      |---|---|---|
      | Visible to child processes | Yes — entire environment is inherited | No, unless fd/path passed |
      | Leaks via crash dumps / `/proc/<pid>/environ` / debug endpoints (`phpinfo`, Spring `/env`) | Commonly | Rarely |
      | Appears in `docker inspect`, ECS/k8s pod spec describes | Yes | No (only the mount path) |
      | Rotatable without restart | No (env is fixed at exec) | Yes — re-read or inotify-watch the file |
      | Captured by error trackers that snapshot env (Sentry default off, but common in homegrown handlers) | Yes | No |
      
      Rules:
      
      - Env vars are acceptable for: 12-factor apps on platforms where mounts are impractical
        (most PaaS), short-lived processes, dev. Mitigate: don't pass env to children
        (`env -i`, explicit allowlists), scrub env in error handlers, never log full env.
      - Env vars are **not** acceptable when: the platform offers file mounts (Kubernetes, systemd
        `LoadCredential=`; ECS injects secrets natively only as env vars, so there a file needs a
        sidecar writing to a task-scoped volume), the secret must rotate without restarts, or the
        process spawns untrusted/third-party children.
      - **Never** put secrets in: Dockerfile `ENV`/`ARG` (baked into image layers/metadata — Critical
        in pushed images; use BuildKit `--mount=type=secret` for build-time needs),
        docker-compose literals committed to git, systemd unit `Environment=` lines (world-readable
        via `systemctl show`; use `LoadCredential=`), or process command lines (rules/03 §3).
      - File-mounted secrets: tmpfs-backed (k8s Secret volumes already are), mode `0400`, owned by the
        app user, path conventional (`/run/secrets/<name>` — Docker/compose secrets land there too).
      - **Mount a secret only into the workload that consumes it.** Never put secrets on a volume
        or directory several deployments share — a node `hostPath` such as `/etc/app-secrets`, a
        shared PVC or NFS export, a compose volume mounted by every service. Each consumer then
        reads every other consumer's credentials, and a compromise of the weakest pod is a
        compromise of all of them (a Secret volume or CSI mount is per-pod by construction).
        OWASP: Proactive Controls 2024 C2; Secrets Management cheat sheet.
      
      ```dockerfile
      # BAD — secret baked into a layer forever; `docker history` shows it
      ARG NPM_TOKEN
      RUN echo "//registry.npmjs.org/:_authToken=${NPM_TOKEN}" > ~/.npmrc && npm ci
      
      # GOOD — BuildKit secret mount: exists only during this RUN, in no layer
      RUN --mount=type=secret,id=npmrc,target=/root/.npmrc npm ci
      # build: docker build --secret id=npmrc,src=$HOME/.npmrc .
      ```
      
      ```ini
      # systemd — BAD: world-readable via `systemctl show -p Environment myapp`
      [Service]
      Environment=DB_PASSWORD=s3cr3t
      # GOOD: LoadCredential mounts root-owned source to a private, service-only path
      [Service]
      LoadCredential=db_password:/etc/credstore/myapp.db_password
      # app reads $CREDENTIALS_DIRECTORY/db_password; SetCredentialEncrypted= for TPM-sealed values
      ```
      
      **Docker Compose / Swarm:** use top-level `secrets:` (file- or external-backed) mounted at
      `/run/secrets/`, never `environment:` literals; for local dev, `env_file: .env` with an
      untracked `.env` is acceptable (rules/05 §8).
      
      ```yaml
      # BAD — secret in pod spec env, visible in `kubectl describe`, frozen until restart
      env:
        - name: DB_PASSWORD
          value: "s3cr3t..."          # literal: Critical if committed
      # GOOD
      volumeMounts:
        - { name: db-creds, mountPath: /run/secrets/db, readOnly: true }
      volumes:
        - name: db-creds
          secret: { secretName: payments-db, defaultMode: 0400 }
      ```
      
      ## 7. In-memory handling
      
      Threats at app level: crash dumps and core files, swap, heap-dumping debug endpoints, error
      trackers serializing objects, language-runtime introspection. Proportionate measures:
      
      - **Hold secrets in the narrowest scope for the shortest time.** Fetch → use → drop the
        reference. Don't stash secrets on long-lived global config objects that every error handler
        serializes.
      - **Wrap in a redacting type** whose `__repr__`/`toString`/`Debug`/JSON-serialization yields
        `[REDACTED]` (Pydantic `SecretStr`, Rust `secrecy::SecretBox`, your own 20-line wrapper).
        This converts whole classes of log/trace leaks into non-events. (Examples: rules/03 §2.)
      - **Zeroization** (overwriting buffers after use) is best-effort in GC languages — strings are
        immutable and copied; use `byte[]`/`bytearray` and wipe where the runtime allows (JVM
        `Arrays.fill`, .NET `CryptographicOperations.ZeroMemory`). In Rust/C use `zeroize`/
        `explicit_bzero`. Worth doing for long-lived master keys; don't pretend it's airtight.
      - **Swap/core**: at app level, prefer platform controls — encrypted swap or none (standard on
        k8s nodes), `ulimit -c 0` / `RLIMIT_CORE=0` and `prctl(PR_SET_DUMPABLE, 0)` for processes
        holding master keys; `mlock` only for small, genuinely critical buffers (libsodium
        `sodium_mlock`). Don't `mlock` the whole heap.
      - **Disable heap-dump/debug endpoints in prod** (JMX heap dump, `/debug/pprof/heap` exposed
        publicly, py-spy on prod boxes by default). These are remote secret extraction tools.
      
      ## 8. Consumer-key (end-to-end) encryption of secrets
      
      At-rest and envelope encryption (rules/05 §7) protect a secret from someone who steals the
      store's *disk*; they do not protect it from the store itself, its operators, or anything that
      can call it with a valid identity. One step further: **encrypt the secret to the consuming
      workload's own public key** (age recipients, HPKE — RFC 9180 — or a KMS key only that
      workload's identity may decrypt with), so the store, the sync controller and every hop in
      between handle only ciphertext and plaintext exists solely inside the consumer.
      
      - **Worth it when** the store or relay sits outside your trust boundary (a third-party SaaS
        store, a vendor-operated CI, a multi-party pipeline), when regulation requires that the
        platform operator cannot read a value, or for a small set of crown-jewel keys.
      - **The cost is key distribution:** every consumer needs a private key that is itself
        protected (TPM/TEE-sealed, or a KMS key bound to the workload identity — otherwise you
        have moved secret zero, not removed it); adding or rotating a consumer means re-encrypting
        for it; and the store can no longer mint dynamic secrets, rotate for you, or record which
        plaintext was used. For most in-cluster services workload identity plus a per-consumer
        store policy (§2–3) is the better trade.
      - SOPS files with per-environment recipients (§4) are the familiar instance — the pattern
        holds only if the recipient key lives with the consumer rather than in the same store.
        OWASP: Secrets Management cheat sheet.
      
      ## Audit checklist
      
      - [ ] Backend matches the decision table; no custom/homegrown secret stores; no secrets in
            Consul/Redis/plain DB columns.
      - [ ] Terraform/Pulumi state not in VCS; state backend encrypted and access-controlled;
            write-only/ephemeral used for new secret-bearing resources.
      - [ ] Vault/OpenBao: platform-identity auth (no static tokens), deny-by-default path-scoped
            policies, short TTLs, audit devices on, KMS auto-unseal, dynamic engines preferred to KV.
      - [ ] Cloud secret managers: per-secret per-principal IAM (no wildcard access), native or
            scheduled rotation, audit/data-access logs enabled and alerted.
      - [ ] SOPS: per-env keys in `.sops.yaml`, encrypted_regex for k8s, age/KMS private keys never in
            repo; decryption happens in controller, not CI artifacts.
      - [ ] Kubernetes: etcd encryption at rest, secrets RBAC tightened, no real `Secret` manifests in
            git; ESO/CSI/Vault-agent used; secret volumes `0400`, read-only, and no `subPath` mounts
            on rotating secrets.
      - [ ] No secrets in Dockerfile ENV/ARG, image layers, compose files, systemd `Environment=`, or
            pod-spec literals; build-time secrets via BuildKit secret mounts.
      - [ ] File mounts preferred over env vars where the platform supports them; env-var usage has
            child-process and error-handler scrubbing mitigations. Env-delivered secrets on a
            platform that offers mounts (review each hit against §6; Medium when it spawns
            third-party children or must rotate): `grep -rn 'secretKeyRef' --include='*.yaml' --include='*.yml' .`
      - [ ] **No secret on a volume shared across deployments (§6)** — High for a node `hostPath` or
            shared PVC holding credentials for several workloads:
            `grep -rnE '^[[:space:]]*path:[[:space:]]*"?/[^[:space:]]*(secret|credential|credstore)' --include='*.yaml' --include='*.yml' .`
      - [ ] Rotation functions run under a role only they can assume, scoped to the secrets they
            rotate, invocable only by the secret manager (§3).
      - [ ] Where the store or relay is outside the trust boundary, secrets are encrypted to the
            consumer's key and that private key is itself sealed or identity-bound (§8).
      - [ ] Secrets wrapped in redacting types; debug/heap-dump endpoints disabled in prod; core
            dumps disabled for key-holding processes; long-lived keys zeroized where the runtime
            permits.
      
    • 03-application-patterns.md 21 KB
      # 03 — Application Patterns
      
      Scope: how application code obtains, holds, and uses secrets — config layering, runtime
      injection, caching/TTL, the no-leak surfaces (code, VCS, logs, errors, URLs, argv, dumps),
      per-environment separation, least-privilege scoping, and audit logging. Read this when writing
      or reviewing any code path that touches a credential.
      
      ## 1. Config layering
      
      Separate **config** (non-secret: hostnames, flags, pool sizes — committable) from **secrets**
      (never committable). One loader, explicit precedence, fail-fast:
      
      ```
      defaults (in code) < config file (in repo) < secret backend / mounted files < env vars (overrides)
      ```
      
      Rules:
      
      - **Secrets enter only through the secret layer** (mounted file, secret-manager fetch, env var
        injected by the platform). The committed config file may contain the *reference*
        (`db_password_secret: projects/p/secrets/db-pw`) — never the value.
      - **Validate at startup, fail fast and loud — but redacted.** Missing/blank secret → exit
        non-zero with the secret's *name*, never its partial value. Don't limp along to a 3am
        connection error.
      - **No "default secrets."** A fallback like `os.getenv("JWT_SECRET", "dev-secret")` ships the
        dev secret to prod the day the env var is mistyped — High finding. Defaults are for
        non-secrets only.
      - **Support `*_FILE` indirection** (e.g., `DB_PASSWORD_FILE=/run/secrets/db`) so the same image
        runs on env-var PaaS and file-mount platforms.
      
      ```python
      # BAD — silent fallback, secret in code, mixed into committable config
      JWT_SECRET = os.getenv("JWT_SECRET", "super-secret-dev-key")
      
      # GOOD — required, file-or-env, redacting wrapper, fail-fast
      def required_secret(name: str) -> SecretStr:
          if path := os.getenv(f"{name}_FILE"):
              return SecretStr(Path(path).read_text().strip())
          if val := os.getenv(name):
              return SecretStr(val)
          raise SystemExit(f"FATAL: secret {name} not provided")  # name only, never value
      JWT_SECRET = required_secret("JWT_SECRET")
      ```
      
      ## 2. Never in code, VCS, logs, errors, or dumps
      
      **Code/VCS:** no literal secret anywhere in the repo — source, tests, fixtures, comments,
      example configs, notebooks, lockfiles (`.npmrc` lines in lockfile diffs), or git history.
      "It's a private repo" changes nothing: contractors, CI runners, laptop theft, repo-visibility
      flips, and AI coding tools all read private repos. Tests use generated-at-runtime fakes or
      clearly impossible placeholders (`test-key-000…`); integration tests pull real (dev-scoped)
      creds from the same secret layer as the app.
      
      **Logs:** the most common real-world leak. Enforce in layers:
      
      1. **Redacting types** (rules/02 §7): `SecretStr`, `secrecy`, custom wrappers — accidental
         `log.info(f"cfg={config}")` prints `[REDACTED]`.
      2. **Logger-level scrubbing:** a processor/filter that masks known keys (`password`, `token`,
         `secret`, `authorization`, `cookie`, `set-cookie`, `x-api-key`) and known value shapes
         (your token prefixes, `Bearer [A-Za-z0-9_\-\.]+`, `AKIA\w{16}`).
      3. **Never log:** full request/response headers, full connection strings/URLs (strip userinfo:
         `postgres://user:****@host/db`), decoded JWT payloads with embedded secrets, full env, full
         config objects.
      4. **Never persist raw** — the tier the three above cannot reach, for text whose producer
         you do not control (§2.1).
      
      
      ```python
      # BAD — leaks the whole DSN (with password) on every connection failure
      logger.error("db connect failed: %s", dsn)
      # GOOD
      logger.error("db connect failed host=%s db=%s user=%s", host, dbname, user)
      ```
      
      ```python
      # GOOD — structlog/logging processor as a backstop for everything the above misses
      SECRET_KEYS = {"password", "passwd", "secret", "token", "api_key", "authorization",
                     "cookie", "set-cookie", "x-api-key", "private_key", "client_secret"}
      SECRET_SHAPES = re.compile(r"(myapp_(sk|pat)_\S+|AKIA\w{16}|Bearer\s+[\w\-.~+/]+=*|eyJ[\w-]{10,}\.[\w-]+\.[\w-]+)")
      def scrub(_, __, event):
          for k in list(event):
              if k.lower() in SECRET_KEYS:
                  event[k] = "[REDACTED]"
              elif isinstance(event[k], str):
                  event[k] = SECRET_SHAPES.sub("[REDACTED]", event[k])
          return event
      ```
      
      The processor is the *backstop*, not the plan — redacting types and disciplined log statements
      come first; the processor catches the third-party library that logs a request object.
      
      ### 2.1 Text you did not produce: persist a digest and a skeleton, never the datum
      
      Tiers 1–3 assume you own the producer. You cannot wrap someone else's bytes in a
      `SecretStr`, so when the thing being stored is **arbitrary text from a source you do not
      control** — a swept corpus, shell history, a crash dump, a scraped page, an agent
      transcript (`rules/04` §7), a diff from an untrusted branch — the guidance silently
      degrades to tier 2 alone. **Tier 2 is an enumeration, and enumerations lose.**
      
      Field-reported, with the run that killed it: a shape-redaction fix reusing an existing
      `sk-…`/`AKIA…`/`ghp_…` matcher was **defeated by its own regression test on the first
      run**, by a GitLab PAT (`glpat-…`) and a Slack webhook URL. Five weeks later the same
      corpus acquired four more live keys, one of them a **bare 64-character hex string with no
      prefix at all** — a value no prefix rule can ever match, and one that a generic
      "long random-looking string" rule cannot separate from a hash, a UUID or a minified bundle.
      
      So do not persist the datum. Persist what diagnosis actually needs:
      
      ```python
      _MAX_LINE = 160
      KEEP = set(" \t|&;()<>{}[]$\"'`\\=/-.:@#*?!~+,")   # this grammar's punctuation, nothing else
      
      def safe_line(line: str) -> dict[str, str]:
          """The only representation of a swept line this tool ever writes down."""
          return {
              "sha256": hashlib.sha256(line.encode("utf-8", "replace")).hexdigest()[:16],
              "length": str(len(line)),
              "shape": "".join(c if c in KEEP else "x" for c in line[:_MAX_LINE]),
          }
      ```
      
      The digest correlates the same line across two runs and lets a human grep their *own*
      source for it; the length bounds complexity; the shape preserves the structure a diagnosis
      needs while every run of payload characters collapses to `x`. **It is leak-proof by
      construction rather than by list** — it never has to know what a secret looks like, so no
      unenumerated format defeats it. That is the same reasoning the library applies to
      structural controls elsewhere, moved one step earlier: to the decision to persist at all.
      
      Scope it honestly, in the rule and in review:
      
      - This is for data kept to **diagnose or correlate**, never for data you must read back. If
        you need the value, you need a secret store, not this.
      - **The shape channel is a real, if narrow, disclosure.** For shell it is punctuation; for
        another grammar pick the analogous skeleton and *say what it keeps*. Do not describe it
        as zero-leak.
      - Truncation is a second, independent bound — set it deliberately, not by accident.
      - **stdout is a persistence surface too.** A console that prints what the file redacts is
        the same leak with a worse retention policy.
      
      **Error messages & exceptions:** exceptions traverse trust boundaries — API error bodies, error
      trackers, support tickets. Never embed a credential in an exception message
      (`raise AuthError(f"key {api_key} rejected")` — High). Configure Sentry/Rollbar/etc. with
      `send_default_pii=False`, server-side and client-side data scrubbers for your secret key names
      and token prefixes; review breadcrumbs (HTTP breadcrumbs can capture auth headers).
      
      **Crash dumps / telemetry:** error handlers that serialize "the request" or "the config" for
      diagnostics will capture `Authorization` headers and secret fields. Allowlist what diagnostics
      capture; never snapshot full env or full headers. Disable verbose framework debug pages in prod
      (Django DEBUG, Flask debugger, Rails verbose errors, Spring Boot `/actuator/env` &
      `/actuator/heapdump` — all dump config/env to the caller; exposed actuator is High).
      
      ## 3. Never in URLs or CLI arguments
      
      **URLs/query strings** (`?api_key=…`, `?token=…`, creds in URL userinfo): persisted by access
      logs on every hop (LB, CDN, proxy, app), browser history, `Referer` headers, and link sharing.
      Use the `Authorization` header. Exception: short-lived (≤1h) presigned/signed URLs designed for
      it — and even those don't belong in long-term logs. A long-lived API key as a query param is
      High.
      
      **CLI args:** `ps`-visible to every user on the host, captured in shell history, audit logs
      (auditd execve), and CI logs.
      
      ```bash
      # BAD — password visible in `ps aux`, shell history, CI log
      mysql -u app -pS3cr3t! prod
      curl -H "Authorization: Bearer $TOKEN" https://api...   # OK-ish: env expansion not in ps,
                                                              # but lands in shell history if literal
      # GOOD — read from file/env/stdin
      mysql --defaults-extra-file=/run/secrets/my.cnf prod
      curl -H @/run/secrets/auth_header https://api...
      PGPASSWORD unset; use ~/.pgpass (0600) or peer/IAM auth
      ```
      
      When *writing* CLIs: accept secrets via env var, `--token-file`, or stdin prompt — never as a
      positional/flag value; if you must accept a flag, document the `ps` exposure and prefer
      `--token-file`. When *invoking* subprocesses from code, pass secrets via env (explicitly
      constructed, not inherited wholesale) or files — never interpolated into the command string
      (also an injection risk).
      
      ## 4. Runtime injection, caching, and TTLs
      
      **Inject at runtime, not build time.** Images/artifacts are environment-agnostic and
      secret-free; the platform supplies values at start (rules/02 §5–6). Grep CI for
      `docker build --build-arg.*SECRET` — build args persist in image history (Critical in pushed
      images).
      
      **Fetching from a secret manager — cache deliberately:**
      
      - **Cache in memory with a TTL** (5–15 min typical): per-request fetches add latency, cost, and
        a hard runtime dependency on the store; no caching means an outage of the secret manager is an
        outage of you.
      - **Serve stale on refresh failure** (with alerting) rather than crashing a running app; only
        startup requires a successful fetch.
      - **React to rotation:** TTL expiry naturally picks up new versions; for file mounts, re-read on
        auth failure or watch the file (k8s updates mounted Secrets in-place within ~1m). The
        **retry-on-auth-failure pattern** makes rotation seamless: on 401/auth-failed, force-refresh
        the cached secret once and retry before erroring.
      - Official caching helpers exist — AWS Secrets Manager caching libraries, Vault Agent template
        cache — prefer them to hand-rolled caches.
      - **Calling a local secrets endpoint** (a serverless secrets extension, an agent sidecar on
        `localhost`) is still a network call to a credential service — write the client like one:
        - **Authenticate every request** with the runtime's own token and **fail fast if it is
          absent**. The AWS Parameters and Secrets Lambda Extension (port 2773 by default) requires
          the `X-Aws-Parameters-Secrets-Token` header set to the function's session token, which
          AWS's samples resolve from the SDK credential chain at call time.
        - **URL-encode the secret identifier.** AWS's own samples interpolate `secretId` raw into
          the query string, yet secret names may contain `+`, `=` and `/` (CreateSecret's documented
          set) — characters query parsers commonly reinterpret (`+` as a space) — and an id built
          from any input can smuggle `&versionStage=…` into the request.
        - **Set a timeout on your request.** The extension's own upstream call has no timeout by
          default (`SECRETS_MANAGER_TIMEOUT_MILLIS` = 0), so an unbounded client call can hang the
          invocation until the platform kills it.
        - **Check the status before reading the body.** The samples parse `SecretString` whatever
          came back; an error body then surfaces as a confusing parse failure — or, worse, an error
          string used as a credential.
        OWASP: Serverless FaaS Security cheat sheet.
      
      ```python
      # GOOD — TTL cache + forced refresh on auth failure
      class SecretCache:
          def __init__(self, fetch, ttl=300):
              self._fetch, self._ttl, self._val, self._at = fetch, ttl, None, 0.0
          def get(self, force=False):
              if force or self._val is None or time.monotonic() - self._at > self._ttl:
                  try:
                      self._val, self._at = self._fetch(), time.monotonic()
                  except Exception:
                      if self._val is None: raise          # startup: fail fast
                      log.warning("secret refresh failed; serving cached")  # runtime: stale + alert
              return self._val
      
      def call_api():
          r = client.call(token=cache.get())
          if r.status == 401:
              r = client.call(token=cache.get(force=True))  # rotation happened mid-TTL
          return r
      ```
      
      ## 5. Client-side code has no secrets
      
      Anything shipped to a browser, mobile app, or desktop client is **public**: JS bundles, source
      maps, APKs/IPAs (trivially decompiled), Electron asar archives. Framework env prefixes that
      inline values into bundles are the standard footgun:
      
      ```bash
      # BAD — *_PUBLIC_ prefixes inline the value into the shipped bundle
      NEXT_PUBLIC_STRIPE_SECRET_KEY=sk_live_...     # Critical: secret key in browser JS
      VITE_OPENAI_API_KEY=sk-...                    # Critical: anyone can read the bundle
      # GOOD — secret stays server-side; client calls your backend, which holds the key
      NEXT_PUBLIC_STRIPE_PUBLISHABLE_KEY=pk_live_... # publishable keys are designed to be public
      ```
      
      - Rule of thumb: if removing it from the client breaks security, it cannot live in the client.
        Proxy third-party APIs through your backend (which also lets you rate-limit and attribute).
      - Mobile: API keys in `strings.xml`/`Info.plist`/compiled constants are extracted in minutes;
        obfuscation is not protection. Use backend proxying, per-user short-lived tokens issued after
        auth, and attestation (Play Integrity / App Attest) where abuse matters.
      - Keys that are *designed* public (Stripe publishable, Firebase config, Maps browser keys) are
        fine in clients but must be restriction-locked at the vendor (referrer/bundle-id/API
        restrictions) — an unrestricted "public" key is a Medium finding.
      - AUDIT: grep built artifacts, not just source — `grep -rE 'sk_live_|AKIA' dist/ build/` —
        and source maps, which may sit anywhere (a bare `*.map` glob aborts under zsh when none match):
        `find . -name '*.map' -not -path '*/node_modules/*' -exec grep -HE 'sk_live_|AKIA' {} +`;
        also check which source maps are published to prod.
      
      ## 6. Per-environment separation
      
      - **One secret per consumer per environment.** dev/staging/prod never share a value; a leak in
        staging must not touch prod. Shared values are a High finding for prod, Medium otherwise.
      - **Namespacing enforces it:** separate cloud accounts/projects per env (best), or per-env
        paths/prefixes (`secret/prod/...` vs `secret/dev/...`) with IAM/policies that make prod
        unreadable from non-prod principals — including developers' day-to-day identities.
      - **Non-prod gets non-prod scopes:** Stripe test keys (`sk_test_`), sandbox tenants, dev-scoped
        DB users. If a vendor offers no sandbox, treat the non-prod key as prod-sensitive.
      - **CI:** prod-deploy credentials only in protected contexts (GitHub environments with
        reviewers, protected branches); PR builds from forks get *no* secrets
        (`pull_request_target` exfiltration is a classic — never check out and execute fork code in a
        secret-bearing context).
      
      ## 7. Least-privilege scoping of tokens
      
      Every credential answers: who uses it, for which actions, on which resources, until when?
      
      - **Scope down at issuance:** GitHub fine-grained PATs (specific repos + permissions) over
        classic PATs; cloud IAM policies listing actions+resources, not `*`; DB users with table-level
        grants, not superuser; API keys with vendor-side restrictions (Stripe restricted keys, Google
        API key referrer/IP/API restrictions).
      - **One credential per consumer.** A key shared by three services cannot be rotated, scoped, or
        attributed independently; its blast radius is the union. Sharing is a Medium finding.
      - **Read paths get read-only creds.** The reporting service does not reuse the app's read-write
        DB user.
      - AUDIT: flag `*:*` IAM attached to app roles (High in prod), org-wide classic PATs in CI
        (High), DB superuser in an app connection string (High), unrestricted vendor keys (Medium).
      
      ## 8. Audit logging of secret access
      
      - **Backend side:** enable the store's access logs (CloudTrail data events for Secrets Manager,
        GCP Data Access logs, Key Vault diagnostics, Vault audit devices) and ship them off-system.
        Alert on: reads by unexpected principals, first-time principal/secret pairs, reads from new
        geographies, spikes, and reads of honeytokens (rules/04 §5).
      - **App side:** log secret *usage events* — "rotated DB cred picked up," "token refresh failed,"
        "auth failure forced refresh" — with secret *names*, never values. These logs are how you
        verify a rotation completed (rules/01 §3 step 4).
      - **Attribution requires per-consumer creds** (§7) — a shared key makes every audit log entry
        ambiguous.
      - **Minimum record per event**, from the store or around it: who asked (principal, and the
        human behind a CI run — rules/01 §5), for which target system and role, whether it was
        approved or rejected, each use, expiry, value updates, authentication and authorization
        errors, and every administrative action on the store itself (policy, auth-method, audit-
        device, key changes). A store whose admin actions are not logged can switch its own audit
        off unseen.
      - **One detection per lifecycle stage:** creation outside the provisioning pipeline, a
        rotation that did not happen on schedule (or happened outside it), revocation, and
        deletion — each an alert or a report, not only a log line.
      - **Use of a revoked or expired secret is a signal, not noise:** log it and alert — it means
        a consumer missed a rotation or someone is replaying a stolen value.
      - **Alert when a dynamic credential is used from somewhere its consumer is not** (a source
        network or egress address outside the workload's own) — a leased credential lifted from a
        pod is otherwise indistinguishable from the pod. OWASP: Secrets Management cheat sheet.
      
      ## Audit checklist
      
      - [ ] Single config loader with explicit precedence; secrets only via secret layer; references
            not values in committed config; startup fails fast (redacted) on missing secrets; no
            default/fallback secret values in code.
      - [ ] No literal secrets in source, tests, fixtures, comments, examples, or git history
            (history scan done — rules/04).
      - [ ] Redacting wrapper types in use; logger-level scrubbing for secret key names and token
            shapes; connection strings/URLs logged without userinfo; no full-header, full-env, or
            full-config logging.
      - [ ] **Text whose producer you do not control is never persisted raw (§2.1)** — a swept
            corpus, shell history, crash dump, scraped page or agent transcript is stored as a
            digest plus a structural skeleton, with the shape channel's disclosure stated and
            stdout covered by the same rule; shape/prefix redaction alone is not accepted as the
            control, because it is an enumeration
      - [ ] Exceptions and error-tracker payloads carry no secret values; PII/data scrubbers
            configured; debug pages and Spring actuator env/heapdump endpoints disabled or
            authenticated in prod.
      - [ ] No secrets in URLs/query strings (except short-lived signed URLs); none passed as CLI
            arguments (in code, scripts, CI, and docs); subprocesses get explicit env/files, not
            inherited environments or interpolated command strings.
      - [ ] No build-time secret injection (`--build-arg`, Dockerfile ENV); runtime injection only.
      - [ ] No secrets in client-shipped code: bundles, source maps, and mobile binaries scanned;
            `NEXT_PUBLIC_`/`VITE_`/`REACT_APP_` vars contain only public values; designed-public keys
            vendor-restricted; third-party APIs proxied server-side.
      - [ ] Secret-manager fetches cached with TTL, stale-on-error with alerting, forced refresh on
            auth failure; rotation verified end-to-end without restarts.
      - [ ] Per-environment secrets fully separated (accounts/paths + IAM); prod unreadable from
            non-prod; vendor test keys in non-prod; fork PRs receive no secrets; prod deploys gated.
      - [ ] Tokens scoped least-privilege (actions, resources, expiry); one credential per consumer;
            read-only creds on read paths; no wildcard IAM/superuser DB users on app credentials.
      - [ ] Backend access logs enabled, shipped, and alerted; app logs usage events by name only;
            rotation completion verifiable from logs.
      - [ ] **Secret audit trail complete (§8)** — requester, target and role, approval/rejection,
            use, expiry, updates, authn/authz errors and store admin actions recorded; detections
            exist for create/rotate/revoke/delete, for any use of a revoked or expired secret, and
            for a dynamic credential used off its consumer's network — Medium per missing class
            (judgment: read the alert rules against this list).
      - [ ] **Local secrets-endpoint clients (§4)** send the runtime token, URL-encode the id, set
            a timeout and check status — Medium. Callers with no timeout at all:
            `grep -rlE 'localhost:2773|127\.0\.0\.1:2773|port: *2773' . | xargs -r grep -LiE 'timeout'`
      
    • 04-detection-and-remediation.md 23.3 KB
      # 04 — Detection & Remediation
      
      Scope: secret scanning (gitleaks, trufflehog), pre-commit hooks, CI gates, git history hygiene,
      the post-leak runbook (rotate-first), and honeytokens. Read this when setting up scanning, when
      running an AUDIT sweep, or the moment a leak is discovered. Pipeline-gate *layering* and
      org-wide push protection are owned by sota-devsecops rules/05 §5.2; this file owns tool
      configuration and leak **response** (rotation, history cleanup).
      
      ## 1. Scanning layers
      
      Defense in depth — each layer catches what the previous missed:
      
      | Layer | Tool/mechanism | Blocks? |
      |---|---|---|
      | Editor/local | gitleaks pre-commit hook (`gitleaks git --pre-commit --staged`) | Yes — before commit exists |
      | Push | server-side pre-receive hook / GitHub push protection | Yes — before history is shared |
      | CI | gitleaks/trufflehog job on every PR + full-history scan scheduled weekly | Yes — fail the build |
      | Platform | GitHub Advanced Security / GitLab secret detection, partner-program auto-revocation | Detect + sometimes auto-revoke |
      | Runtime | honeytokens (§5), backend access-log anomaly alerts (rules/03 §8) | Detect use |
      
      Pre-commit alone is insufficient (devs skip hooks, `--no-verify`); CI alone is too late (the
      secret already left the laptop and entered shared history). Run both.
      
      ### gitleaks
      
      ```yaml
      # .pre-commit-config.yaml
      repos:
        - repo: https://github.com/gitleaks/gitleaks
          rev: vX.Y.Z      # an exact release tag (latest stable: github.com/gitleaks/gitleaks/releases) — pre-commit rejects a wildcard like v8.x
          hooks:
            - id: gitleaks   # runs `gitleaks git --pre-commit --redact --staged --verbose`
      ```
      
      ```toml
      # .gitleaks.toml — extend defaults, add your own token prefixes (rules/01 §1)
      [extend]
      useDefault = true
      [[rules]]
      id = "myapp-api-key"
      description = "MyApp internal API key"
      regex = '''myapp_(sk|pat)_[A-Za-z0-9_\-]{32,}'''
      [[allowlists]]                                  # gitleaks ≥ 8.25.0; the older single [allowlist] table is superseded
      paths = ['''testdata/fake_keys\.json''']     # narrow, path-based; never allowlist by rule id
      ```
      
      CI (gitleaks ≥ 8.19 subcommands): `gitleaks git --redact .` on PRs — diff-aware via
      `--log-opts="origin/main.."` for speed — plus a scheduled full-history scan without log-opts
      (needs a `fetch-depth: 0` checkout) so old commits are rechecked as rules improve. The legacy
      `detect`/`protect` spellings still run but are deprecated aliases.
      
      ### trufflehog
      
      Complements gitleaks: ~800 detectors **with verification** — it calls the credential's own API
      to check liveness, collapsing false positives.
      
      ```bash
      trufflehog git file://. --results=verified --fail       # CI gate: verified-live secrets only
      trufflehog filesystem /path --results=verified,unknown  # audit sweep: include unverifiable
      trufflehog docker --image myorg/app:latest              # images: layers, env, files
      ```
      
      Use `--results=verified` for blocking gates (near-zero false positives; it supersedes the
      older `--only-verified`, now a hidden flag); use the broader mode for audits — an
      unverifiable secret is still a finding, just triaged manually. Also point trufflehog at
      non-git surfaces: S3 buckets, container images, CI logs exports — secrets leak there too.
      
      ### Scanner hygiene
      
      - **Allowlists are path- and fingerprint-scoped, reviewed in PRs.** A blanket
        `allowlist regex = '''.*test.*'''` silently exempts `tests/prod_credentials.py` — Medium.
      - **`--redact` everywhere**: scanner output goes to CI logs; unredacted output re-leaks the
        secret into a new surface.
      - **Inline `# gitleaks:allow` comments require justification** in the same line/commit; audit
        them — they are where real leaks hide (`grep -rn "gitleaks:allow"`).
      - Scanners miss: secrets in *binary* files, novel formats without rules, encrypted blobs with
        weak keys, and anything entropy-shaped below thresholds. The manual grep pass in SKILL.md
        AUDIT mode exists for this reason.
      - **Own the detector set, class by class.** Keep a list of every secret class the org holds
        and a rule (default or custom) for each: long-lived and hard-to-rotate tokens, connection
        strings with userinfo, 2FA/TOTP seeds (`otpauth://` URIs, base32 seed fields), session
        tokens and cookies, private and SSH keys, cloud keys, and whole secret-bearing config
        files (kubeconfig, `.npmrc`, `credentials.json`). A class with no rule is a class the gate
        cannot see; review the list when a new vendor or token format arrives.
      - **One standard fake value per type, org-wide**, used in every test, fixture and doc, so an
        allowlist can name those exact values (or their fingerprints) instead of a path or a regex
        that also swallows real keys. Vendors' documented example values often work — gitleaks
        8.30.1's default config, checked here, skipped the AWS documentation key
        `AKIAIOSFODNN7EXAMPLE` while flagging a random `AKIA…` beside it — but confirm each against
        your scanner rather than assuming. OWASP: Secrets Management cheat sheet.
      - **A working-tree scan is not limited to tracked files, and does not read `.gitignore`.**
        `gitleaks dir` (and directory scanners generally) treat the path as a plain directory, so an
        untracked artifact — a log, an editor backup, a crash dump, a tool's scratch output — is in
        scope for the gate while being **invisible to `git status`**. Verified 2026-09-15 on gitleaks
        8.30.1 with both arms: the same planted credential in a *visible* untracked file and in a
        **gitignored** one both reported `leaks found: 1`, with `git status --porcelain` empty
        throughout. **When a secret gate fails while the history pass is clean, list the findings by
        file before you read any diff** — the answer is often a file that was never yours, and the
        first instinct, suspecting the commit, burns the most time. Field-reported: 22 hits, all
        `generic-api-key` false positives on `key=value` shapes in one daemon log dump, left in the
        tree by tooling.
      
      ### Adopting scanning on a legacy repo (baseline workflow)
      
      Turning on a blocking scanner against a repo with years of history floods the team and gets the
      gate disabled within a week. Instead:
      
      1. Full-history scan once; triage every hit (real-and-live / real-but-rotated / false positive).
      2. **Rotate all real-and-live findings now** (§3) — the baseline is an incident list, not an
         ignore list.
      3. Fingerprint the remainder into a baseline (`gitleaks` `--baseline-path` over a prior
         report) so the gate only fails on *new* findings; commit the baseline and review changes to
         it like code. trufflehog has no baseline file: scan only new commits (`--since-commit <base>`)
         and mark accepted lines with an inline `trufflehog:ignore` comment — `--exclude-detectors`
         disables a whole detector class, which is a snooze, not a baseline.
      4. Burn the baseline down on a schedule; it should shrink monotonically. A growing baseline
         file means the gate is being used as a snooze button — Medium finding.
      
      ## 2. CI gates and repo hygiene
      
      - **Blocking, not advisory.** A secret-scan job that's `allow_failure: true` is decoration.
      - **Scan the diff on PRs, the full history on schedule**, and **scan built images** before push
        (`trufflehog docker`) — build args and COPY'd `.env` files surface here.
      - **`.gitignore` preloads** in every repo template: `.env`, `.env.*`, `!.env.example`, `*.pem`,
        `*.key`, `*.p12`, `*.pfx`, `*.jks`, `id_rsa*`, `*.kubeconfig`, `credentials.json`,
        `terraform.tfstate*`, `.netrc`. gitignore is a guardrail, not a control — files added with
        `git add -f` still need the scanner to catch them.
      - **GitHub push protection** (Settings → *Security and quality* → Advanced Security; at
        repository level it needs GitHub Secret Protection) on for all repos/orgs; it blocks pushes,
        web-UI commits, uploads and REST API writes containing supported secret patterns server-side,
        including from devs without hooks. **By default anyone with write access can bypass it** by
        giving a reason — restrict that with delegated bypass — and it covers only the newest token
        format of a provider's patterns, never passwords. Coverage expands continuously (GitHub
        changelog 2026-03-10: 28 new detectors, 39 more push-protected by default; validity checks;
        owner/expiry metadata for a few token types) — treat alerts marked *active* by validity
        checks as §3 incidents, not backlog.
      - **Fork PR safety:** secret-bearing workflows never run on `pull_request` from forks; audit
        any `pull_request_target` usage that checks out PR code (classic exfil vector — High).
      - **Install-time harvesting worms:** the Shai-Hulud npm worms (Sept/Nov 2025, hundreds of
        packages) ran a bundled TruffleHog on developer machines and CI at install time, exfiltrated
        hits to newly created public GitHub repos, self-propagated via stolen npm/GitHub tokens, and
        registered victims as self-hosted Actions runners. Defenses: `--ignore-scripts` (or isolated,
        token-free install steps) in CI, no long-lived registry/VCS/cloud tokens on laptops or runners
        (rules/01 §4 OIDC instead), and treat sudden public-repo creation or self-hosted-runner
        registration under an org identity as a §3 leak trigger.
      
      ## 3. Post-leak runbook: rotate first
      
      A secret that touched a commit, log, ticket, chat, or paste is **compromised** — period.
      Scrapers index public commits in well under a minute; "we force-pushed quickly" is not
      mitigation; private repos only shrink, not eliminate, the audience. Execute in this order:
      
      1. **Rotate/revoke immediately — before any cleanup.** Issue the replacement, deploy consumers,
         revoke the leaked value (rules/01 §3). If revocation breaks things, that's an availability
         bug you fix *after* killing the credential — a live leaked key is worse than downtime.
         For cloud keys also check for attacker persistence: new keys/users/roles created by the
         leaked identity, modified trust policies.
      2. **Assess blast radius:** backend access logs (CloudTrail etc.) for use of the leaked
         credential since the leak timestamp — not since discovery. Unexplained use → escalate to
         incident response; this is now a breach investigation, not a hygiene task.
      3. **Purge from history (§4)** — only after rotation. Purging an unrotated secret just
         advertises where it was.
      4. **Close the hole:** which layer (§1) should have caught it? Add the missing rule/hook/gate.
         Add the leaked secret's *shape* to scanner config so recurrence is caught.
      5. **Record it:** timeline, blast radius, fix. Leak frequency per quarter is the KPI for your
         scanning posture.
      
      Severity-of-response calibration:
      
      | Leaked | Response tempo |
      |---|---|
      | Cloud root/admin key, signing key, KMS-adjacent creds | Drop everything; rotate within the hour; full IR engagement |
      | Prod service credential (DB, API key with write scope) | Same day; check access logs before and after rotation |
      | Read-only / non-prod / sandbox credential | Within 24h; still rotate — non-prod creds pivot into prod via reused patterns and shared infra |
      | Honeytoken | No rotation needed; treat as breach signal for the planted surface (§5) |
      
      **Key or CA compromise needs a written plan before it happens** — the rotate-first steps
      above assume one credential; a signing key or a CA key invalidates everything it vouched for.
      The plan names: who to contact (internal owners, the CA, relying parties, customers where
      contracts require it); **how to re-key** (new key pair, reissue every certificate or token
      under it — rules/05 §3–4, sota-network-security rules/06 for PKI); a **key/certificate
      inventory linked to where each is deployed**, so reissue is a query, not an archaeology dig;
      **revocation that relying parties actually enforce** — CRL checking switched on, or lifetimes
      short enough to be the revocation; **monitoring that re-keying completed** (no endpoint still
      presenting the old key, no token still verifying under it); and **what the compromised key
      already signed or encrypted** — artifacts to re-sign or distrust, data to re-encrypt, signatures
      made after the compromise time to reject. OWASP: Key Management cheat sheet.
      
      **Leaks outside git** follow the same runbook with a different purge step: secrets pasted into
      CI logs (purge/expire the log retention), chat (delete + rotate; assume exported), issue
      trackers, error trackers (scrub events via API), AI tool transcripts (§7 — the developer's own
      agent-session files, which are a standing credential store rather than only a post-leak
      location), and pastebins (report for takedown). The rotate-first rule is identical — purging is best-effort everywhere; rotation is
      the only reliable mitigation.
      
      ## 4. Git history hygiene
      
      Removing a secret from HEAD does nothing — `git log -p`, every clone, and every fork still hold
      it. To actually purge:
      
      ```bash
      # Preferred: git-filter-repo (BFG is the older alternative)
      pip install git-filter-repo
      # replacements.txt:  LEAKED_VALUE==>REMOVED
      git filter-repo --replace-text replacements.txt          # rewrites all history
      # or drop whole files everywhere:
      git filter-repo --invert-paths --path config/secrets.yml
      git push --force --all && git push --force --tags
      ```
      
      Then, in order of how often it's forgotten:
      
      - **Every collaborator re-clones.** Old clones still contain the secret and will reintroduce
        the old history on a careless push (protect branches against non-fast-forward from stale
        clones).
      - **Forks keep the old commits** — you cannot rewrite someone else's fork. On GitHub,
        dangling/forked commits remain fetchable by SHA even after rewrite; contact support to
        garbage-collect cached views, and **treat the value as permanently public regardless** —
        which is why rotation came first.
      - **PRs, issues, CI logs, artifacts, package registries** may quote the secret — search and
        scrub those surfaces too (`gh api` search, CI log retention purge, yank published packages
        that embed it).
      - Re-run the full-history scan to confirm zero hits before closing the incident.
      
      ```text
      # BAD incident response (common): delete the line, normal commit, move on
      #   -> secret still in history, still valid, now flagged as interesting
      # GOOD: rotate -> verify no abuse -> filter-repo -> force-push -> re-clone fleet -> rescan
      ```
      
      ### Posture metrics
      
      Track quarterly; these tell you whether the program works:
      
      - **Time-to-detect** (commit → alert) and **time-to-rotate** (alert → leaked cred dead) per
        incident; the second number is the one that matters and should be hours, not days.
      - **New verified leaks per quarter** (trend down), **baseline size** (trend down),
        **% repos with pre-commit + CI gate + push protection** (trend to 100%).
      - **Honeytoken alert drill freshness** — last test-fire date per planted surface.
      
      ## 5. Honeytokens
      
      Plant credentials that have **no legitimate use** and alarm on *any* use — they detect breaches
      of the surfaces scanners can't watch (stolen laptops, leaked backups, insider snooping, supply
      chain).
      
      - **Sources:** Canarytokens (free: AWS keys, fake DB creds, files), Thinkst Canary, GitGuardian
        honeytoken, or roll your own — a real but permissionless AWS key whose *use* (CloudTrail
        `GetCallerIdentity` from an unknown IP) triggers an alert.
      - **Where to plant:** private repos (a fake `.env` in an internal repo detects repo compromise),
        CI variable groups, wikis/Notion, S3 backup buckets, developer-laptop `~/.aws/credentials`
        (extra profile), container images, secret-manager entries no app reads (access-log alert).
      - **Make them indistinguishable** from real credentials in naming and placement; document them
        in a register *outside* the planted surfaces so responders can tell drill from breach.
      - **Alert = assume breach of that surface.** Triage where the token was planted, not the token
        itself (it has no privileges).
      - AUDIT note: before reporting a "live AWS key" finding, consider it may be a honeytoken —
        and never *verify* candidate keys by using them against the provider without explicit
        permission; usage may trip someone's alarm or constitute unauthorized access. Judge liveness
        from context (key shape, age, references), or hand to the owner to check.
      
      ## 6. Audit sweep quick reference
      
      Condensed from SKILL.md AUDIT mode — the grep set when tools aren't available:
      
      ```bash
      # Known prefixes & key blocks
      grep -rInE '(AKIA|ASIA)[A-Z0-9]{16}|gh[pousr]_[A-Za-z0-9]{36}|github_pat_|xox[bpars]-|sk_live_|rk_live_|(^|[^A-Za-z0-9_-])sk-(proj-|svcacct-|admin-|ant-(api|admin)[0-9]{2}-)?[A-Za-z0-9_-]{20,}|AIza[A-Za-z0-9_\-]{35}|glpat-|npm_[A-Za-z0-9]{36}|dop_v1_|shpat_' .
      grep -rIlE -- '-----BEGIN ([A-Z0-9]+ )*PRIVATE KEY( BLOCK)?-----' .   # RSA/EC/OPENSSH/ENCRYPTED/PKCS#8/PGP
      # Assignments & connection strings
      grep -rInE '(password|passwd|pwd|secret|token|api[_-]?key)\s*[:=]\s*["'"'"'][^"'"'"']{6,}' --include='*.py' --include='*.js' --include='*.ts' --include='*.go' --include='*.rb' --include='*.java' --include='*.yml' --include='*.yaml' --include='*.json' --include='*.tf' --include='*.sh' --include='*.env' --include='*.cfg' --include='*.ini' --include='*.properties' .
      grep -rInE '[a-z+]+://[^/:@[:space:]]+:[^@[:space:]]+@' .
      # Dangerous tracked files
      git ls-files | grep -E '\.env($|\.)|\.pem$|\.key$|\.p12$|\.pfx$|\.jks$|id_rsa|credentials\.json|terraform\.tfstate|kubeconfig|\.npmrc$|\.netrc$'
      # History (HEAD-clean but leaked)
      git log --all -p --unified=0 | grep -E '^(\+).*(AKIA|ghp_|PRIVATE KEY|sk_live_)' | head -50
      # Scanner suppressions hiding bodies
      grep -rn 'gitleaks:allow\|nosec\|trufflehog:ignore' .
      ```
      
      Triage every hit per SKILL.md severity table; redact values in the report (prefix + length).
      
      ## 7. Agent session transcripts are a credential store
      
      Coding-agent harnesses persist the full model context to disk — Claude Code writes
      `~/.claude/projects/<project>/<session>.jsonl`, and every comparable tool has an
      equivalent. **Assume that file holds every secret the agent was ever shown**, including
      files the harness loaded on its own initiative rather than because the repo asked it to.
      
      Field-reported 2026-09-07 and verified clause by clause: a harness loaded a repo's `.env`
      into context under a header reading *"project instructions, checked into the codebase"*.
      `git ls-files --error-unmatch .env` errored (untracked); `git check-ignore -v .env` named
      the `.gitignore` line that excludes it; no `@`-import of it existed in the agent file; the
      mode was `600`. **Every clause of that label was false**, and four live API keys were then
      sitting verbatim, one line each, in the transcript. A harness's account of *why* it read a
      file is a claim, not evidence — check it the way you would any other
      (`sota-skill-security` rules/02 §2).
      
      - **Inventory the path.** It belongs on the same list as `.env`, shell history, IDE
        settings and CI variables. It is missing from most checklists because it is newer than
        they are, not because it is safe.
      - **Scope the scan before you run it.** A scanner pointed at that directory returns every
        secret the agent has seen across *every* project, not only this one — that is a triage
        budget, not a finding count. Run it with `--redact` (§1) so the scan output is not a
        third copy of the value.
      - **Any tool that reads it is a secret-processing tool.** Sweeping transcripts as an input
        corpus — a linter, an analytics script, a differential oracle — puts that tool under
        `rules/03` §2.1: persist a digest and a skeleton, never the line.
      - **A tool whose input path lies outside the repo has an attack surface that changes with
        no commit to the repo.** Nothing in the diff tells you the corpus gained four live keys
        overnight. Re-establish what the location holds before each run; it is not a fixed
        property of the tool's design.
      - **Calibrate before escalating.** Owner-only file → owner-only file on one machine is an
        **expanded local surface, not a disclosure**, and not on its own a reason to rotate. The
        same incident arc that produced this finding had already ordered one unnecessary rotation
        by reading a scanner's *rule names* instead of its *values*. Read the values; then §3 if
        they are real.
      
      This section is the after-the-fact inventory; the preventive half — keep secret files out of
      the tree the assistant reads, and out of its editor and terminal context — is rules/05 §8.
      
      Audit: a repo tool that reads an agent-transcript directory and writes any part of a line
      verbatim = **High**; the same tool emitting digests and skeletons = no finding. A secret
      inventory that does not name the transcript path = **Medium**.
      
      ## Audit checklist
      
      - [ ] gitleaks (or equivalent) pre-commit hook in `.pre-commit-config.yaml` and documented in
            dev setup; custom rules cover internal token prefixes.
      - [ ] Blocking CI secret-scan on every PR; scheduled full-history scan; built images scanned;
            scanner output redacted.
      - [ ] **A red working-tree scan against a clean history pass was triaged by file, not by
            diff.** Directory scanners do not honour `.gitignore`, so an untracked artifact
            `git status` never shows is still in scope — confirm whether the finding is even in a
            file the repo owns before treating it as a leak, and before treating it as noise.
      - [ ] **The agent-session transcript directory is on the secret inventory (§7)** — scanned
            with scope decided in advance and `--redact` on, and every repo tool that reads it
            treated as a secret-processing tool (`rules/03` §2.1). Severity calibrated on the
            values, not on a scanner's rule names: owner-only to owner-only on one machine is an
            expanded surface, not a disclosure.
      - [ ] GitHub push protection / server-side pre-receive scanning enabled org-wide.
      - [ ] Allowlists narrow (path/fingerprint), justified, and reviewed; all inline
            `gitleaks:allow` suppressions audited.
      - [ ] `.gitignore` covers `.env*`, key/cert files, tfstate, kubeconfig, `.netrc`; no such files
            tracked (`git ls-files` check).
      - [ ] Fork PRs receive no secrets; `pull_request_target` does not execute fork code.
      - [ ] CI installs run with lifecycle scripts disabled (or in token-free steps); alerts exist
            for sudden public-repo creation and self-hosted-runner registration under org identities.
      - [ ] Leak runbook exists, is current, and orders rotate → assess (logs since leak time) →
            purge → harden → record; revocation paths tested per credential class.
      - [ ] **A key/CA-compromise plan exists (§3)** — contacts, re-key method, linked key/cert
            inventory, enforced revocation, re-key completion monitoring, and triage of what the
            key signed or encrypted — High where an org runs its own CA or signing keys without one.
            Runbooks that never mention it: `grep -rLiE 'key compromise|CA compromise|re-?key' runbooks/`
      - [ ] **Detector set covers every secret class held (§1)** and tests use one standard fake
            per type — Medium. A gitleaks config with no custom rule for your own formats:
            `grep -q '^\[\[rules\]\]' .gitleaks.toml || echo 'no custom [[rules]] (or no .gitleaks.toml)'`
      - [ ] Past incidents: history actually rewritten (filter-repo + force-push + re-clone), forks/
            PR quotes/CI logs scrubbed, full-history rescan clean, and the leaked values rotated.
      - [ ] Legacy-repo adoption used a triaged baseline (live findings rotated first); baseline file
            reviewed like code and shrinking over time.
      - [ ] Non-git leak surfaces (CI logs, chat, issue/error trackers) covered by the runbook with
            retention/scrub procedures.
      - [ ] Honeytokens planted in at least repos + CI + one backup surface; register maintained
            out-of-band; alerts route to incident response and have been test-fired.
      - [ ] No audit practice involves invoking discovered credentials against providers without
            explicit owner permission.
      
    • 05-credential-types.md 21.4 KB
      # 05 — Credential Types
      
      Scope: type-specific rules for database credentials, API keys, signing keys, TLS private keys,
      SSH keys, JWT signing secrets (with `kid` rotation), encryption keys vs KMS envelope
      encryption, and `.env` file discipline. Read the matching section before implementing or
      auditing a specific credential class. General lifecycle rules (rules/01) still apply.
      
      ## 1. Database credentials
      
      Preference order:
      
      1. **IAM database auth** (RDS/Aurora IAM, Cloud SQL IAM, Azure AD for Postgres/SQL): no
         password exists; the driver presents a short-lived token. Use where the engine and driver
         support it.
      2. **Vault/OpenBao dynamic creds** (`database/` engine): per-instance users with 1–24h TTL,
         auto-revoked. Each pod gets its own user → perfect attribution, instant revocation.
      3. **Static password in a secret manager** with scheduled rotation — the floor, not the goal.
      
      Rules:
      
      - **Zero-downtime static rotation = dual users.** Two DB users (`app_a`, `app_b`) with
        identical grants; rotate the idle one's password, flip the app's secret to it, repeat next
        cycle. Single-user rotation always has a race between password change and config propagation.
        AWS Secrets Manager's "alternating users" rotation strategy implements exactly this.
      - **App users are least-privilege** (schema/table grants, no `SUPERUSER`/`GRANT OPTION`/DDL for
        runtime users; migrations run as a separate, more-privileged, more-protected user).
      - **Connection strings:** assemble from parts at runtime; never log the assembled DSN
        (rules/03 §2); never commit one with userinfo (`postgres://app:pw@…` in code/compose is
        Critical with real values).
      - Require TLS to the DB (`sslmode=verify-full`) — a rotated password helps little if creds
        cross the wire in clear inside a "trusted" network.
      
      ```sql
      -- Dual-user rotation, cycle N (app currently uses app_b):
      ALTER ROLE app_a WITH PASSWORD :'new_pw' VALID UNTIL 'infinity';
      -- update secret manager entry "db-user" -> {user: app_a, password: new_pw}
      -- wait for consumer cache TTL + verify via pg_stat_activity that app_b sessions drain
      -- cycle N+1 rotates app_b; the idle user is always the one being rotated
      ```
      
      ## 2. API keys (third-party SaaS)
      
      Usually irreplaceable by federation — manage the static secret well:
      
      - **Store in the secret manager**, runtime-injected, cached with TTL (rules/03 §4).
      - **Scope at the vendor:** Stripe restricted keys per function, Google API key
        IP/referrer/API restrictions, GitHub fine-grained PATs, read-only variants wherever offered.
        One key per consuming service per environment — never share the org-wide key across services
        (rules/03 §7).
      - **Test/live separation:** `sk_test_` in non-prod, `sk_live_` only in prod paths; CI uses test
        keys exclusively. A live-mode key in CI variables for "integration tests" is a High finding.
      - **Rotation ≤90 days** using the vendor's dual-key/overlap mechanism if present; if the vendor
        supports only one active key, schedule a brief maintenance flip and document it.
      - **Webhook signing secrets are API-key-class secrets:** verify signatures with
        constant-time comparison, support two active secrets during rotation, reject stale timestamps
        (replay).
      - AUDIT: vendor key prefixes are high-signal greps (`sk_live_`, `xoxb-`, `SG.`, `key-`,
        Twilio `SK[0-9a-fA-F]{32}` API-key SIDs; Twilio `AC[0-9a-fA-F]{32}` is an Account SID — an
        identifier, so look beside it for the auth token or key secret). A live vendor key in
        frontend bundles/mobile apps is Critical — client-shipped code is public.
      
      ## 3. Signing keys (code/artifact/webhook/general-purpose)
      
      - **Private keys live in KMS/HSM and never leave.** Sign by calling `kms:Sign` / Key Vault
        sign / PKCS#11 — the app holds a *permission*, not a key. Exportable software signing keys
        are a Medium finding when a KMS path exists; a committed one is Critical.
      - **Artifact/code signing:** prefer keyless (Sigstore cosign with OIDC identity + Rekor
        transparency log) — no long-lived key exists at all. If long-lived keys are mandated
        (Android keystores, Apple certs), store in HSM-backed services, restrict signing to a locked
        CI lane with audit-logged, reviewed triggers.
      - **Separate keys per purpose.** The webhook-HMAC key ≠ the JWT key ≠ the artifact key; reuse
        couples blast radii and blocks independent rotation.
      - **Publish verification material, version it** (public keys / certs with IDs), and rotate
        with overlap so old artifacts remain verifiable per your policy.
      
      ## 4. TLS private keys
      
      - **Automate issuance: ACME everywhere** (Let's Encrypt/ZeroSSL via cert-manager, Caddy,
        certbot, cloud LB-managed certs). ≤90-day certs make key compromise self-limiting and force
        the automation that prevents 2am expiry outages. The industry mandates this direction:
        CA/Browser Forum ballot SC-081v3 caps public TLS validity at 200 days since March 2026,
        dropping to 100 days (March 2027) and 47 days (March 2029) — manual issuance is being
        regulated out of existence. A multi-year manually-installed cert is a
        Medium finding (process smell), expired-soon without automation is operationally urgent.
      - **Generate the key where it terminates** (CSR flow); never email/chat a `.key` or `.pfx`.
        Wildcard cert keys copied to N servers multiply exposure N-fold — prefer per-host/per-SAN
        certs, or distribute via secret manager with per-host access if a wildcard is unavoidable.
        When a wildcard is justified, **confine its key to one tier**: terminate TLS for it at a
        single reverse proxy / edge layer so the private key sits on one system, and backends get
        their own per-host certs. **Issue it one level down** (`*.svc.example.org`, not
        `*.example.org`): a wildcard matches exactly one left-most label (RFC 9525 §6.3), so a
        narrow parent bounds which names a stolen key can impersonate. OWASP: Transport Layer
        Security cheat sheet.
      - **Permissions `0400`,** owner = the terminating process's user; key files outside web roots
        and build contexts (a `.pem` inside a Docker build context ends up in the image).
      - **Internal mTLS:** run a private CA (Vault PKI engine, cert-manager + internal issuer,
        SPIFFE/SPIRE) issuing ≤24h–90d certs; never hand-manage internal certs or, worse, disable
        verification (`InsecureSkipVerify: true`, `verify=False`, `NODE_TLS_REJECT_UNAUTHORIZED=0`
        in committed code is a High finding — it's adjacent to secrets because it nullifies them).
      - Compromise/leak of a key: revoke (CRL/OCSP), reissue on a new key, and rotate anything that
        transited sessions lacking forward secrecy. Committed key + cert pairs in repos are Critical
        even if expired — check the *key* for reuse in newer certs.
      
      ## 5. SSH keys
      
      - **Ed25519, per human, per device,** passphrase-protected, ideally hardware-backed
        (`sk-ssh-ed25519` FIDO2 keys, or platform agents like Secretive/TPM). Never share a private
        key between people or copy one to a second machine — generate a new key there.
      - **Better: short-lived SSH certificates** (Vault SSH CA, Teleport, Smallstep): users get
        certs valid for hours after SSO+MFA; servers trust the CA, not 400 stale `authorized_keys`
        entries. Eliminates key sprawl and offboarding gaps. At minimum, inventory and expire
        `authorized_keys` entries; remove on offboarding the same day.
      - **Machine SSH (CI → server) is a smell:** prefer pull-based deploys (GitOps, image pulls) or
        cloud-native exec (SSM Session Manager — no open port 22, IAM-audited). If unavoidable: a
        dedicated keypair per pipeline, `from=` and `command=` restrictions in `authorized_keys`,
        stored in CI secret store, rotated ≤90d.
      - **Deploy keys / repo SSH keys:** read-only unless write is proven necessary; one per
        consumer.
      - **Agent hygiene:** `ForwardAgent no` by default (a compromised host can use your agent);
        prefer `ProxyJump`. `IdentitiesOnly yes` to avoid spraying every loaded key.
      - AUDIT: `id_rsa`/`id_ed25519` files in repos or images are Critical (even passphrase-protected
        — assume crackable); `~/.ssh` COPY'd into Dockerfiles is a classic (grep Dockerfiles for
        `ssh` and `COPY.*\.ssh`); known leaked pattern: private key in a "dotfiles" repo.
      
      ## 6. JWT signing secrets and `kid` rotation
      
      - **Prefer asymmetric (ES256 as the portable default; Ed25519 where the library supports it;
        RS256 for interop) over HS256** whenever any party other
        than the issuer verifies tokens: with HS256 every verifier holds the *signing* secret and can
        mint tokens; with asymmetric, verifiers hold only public keys. HS256 is acceptable only
        issuer-verifies-own-tokens (e.g., session tokens in a monolith) — and then the secret is
        ≥256-bit CSPRNG (rules/01 §1), not a passphrase.
      - **Pin the algorithm at verification.** Accept exactly the expected `alg`; never `alg: none`;
        never let an attacker downgrade RS256→HS256 (verifier treating the public key as an HMAC
        secret — the classic confusion attack). Library config: explicit `algorithms=["ES256"]`. RFC 9864
        deprecates the polymorphic JOSE `EdDSA` identifier in favour of the fully-specified `Ed25519`;
        libraries differ (e.g. `jose` accepts `Ed25519`, PyJWT still names it `EdDSA`) — check yours.
      - **`kid` rotation (zero-downtime):**
        1. Generate key N+1; add to the published JWKS (`/.well-known/jwks.json`) alongside N.
        2. Switch issuance to N+1 (`kid` header = N+1's id).
        3. Keep N in JWKS until max token TTL elapses (all N-signed tokens expired).
        4. Remove N from JWKS; destroy the private key.
        Verifiers select the key by `kid` and cache JWKS with HTTP cache headers (≤15m TTL) plus
        refresh-on-unknown-`kid` — that last behavior is what makes emergency rotation fast.
      - **Emergency rotation** (key leaked): publish N+1, issue with N+1, remove N immediately —
        accepting that live N-signed tokens die. Forced logout beats forged admin tokens. Keep
        token TTLs short (≤15m access tokens) so even routine rotation windows are short.
      - AUDIT greps: `eyJhbGciOi` literals in code/tests (inline real JWTs — decode and check for
        prod claims), `JWT_SECRET` with a default value, `algorithms` lists containing both HS and RS
        variants, `verify=False`/`verify_signature: False` options, JWKS endpoints serving a single
        never-rotated key with no `kid`.
      
      ```python
      # BAD — alg from attacker-controlled header; shared secret in code with fallback
      jwt.decode(tok, SECRET or "devsecret", algorithms=[jwt.get_unverified_header(tok)["alg"]])
      
      # GOOD — pinned alg, key chosen by kid from cached JWKS, refresh on unknown kid
      hdr = jwt.get_unverified_header(tok)
      key = jwks.get(hdr["kid"]) or jwks.refresh_and_get(hdr["kid"])  # handles rotation
      claims = jwt.decode(tok, key, algorithms=["ES256"], audience="api://payments")
      ```
      
      ## 7. Encryption keys vs KMS envelope encryption
      
      **Never hand-manage raw data-encryption keys** (a hex key in config decrypting DB fields is a
      High finding — it combines the worst of secrets and crypto). Use envelope encryption:
      
      - **Pattern:** KMS holds the root key (non-exportable). Per object/record/tenant, generate a
        **data encryption key (DEK)** via `GenerateDataKey`; encrypt the payload locally with the
        plaintext DEK (AES-256-GCM); store the *encrypted* DEK alongside the ciphertext; zeroize the
        plaintext DEK (rules/02 §7). Decrypt = ask KMS to unwrap the stored DEK, then decrypt locally.
      - **Why:** the root key never exists outside the HSM; access is IAM-controlled and audit-logged
        per operation; "rotation" of the root key is a KMS toggle (old versions still unwrap old
        DEKs); revoking an app's decrypt permission instantly bricks its access without touching data.
      - **Key hierarchy and rotation:** root (KMS auto-rotation — AWS default 365 days, configurable
        90–2560) → optional per-tenant key → DEK per object. Re-encrypting data is only needed if a
        *DEK* is compromised; root rotation requires nothing. Per-tenant keys also give you crypto-shredding (destroy tenant key =
        tenant data unrecoverable) for deletion compliance.
      - **Use a maintained client library** (AWS Encryption SDK, Tink, cloud KMS client envelopes) —
        they handle DEK caching (bound: time *and* message count), nonce management, and
        algorithm-suite headers. Hand-rolled AES around KMS calls gets nonce reuse wrong.
      - **Encryption context / AAD:** bind ciphertexts to their identity (`tenant_id`, `record_id`)
        so ciphertext can't be swapped between rows; it also lands in KMS audit logs.
      - **Wrapped and stored keys need integrity, not only confidentiality.** A wrapped DEK or a
        key file an attacker can silently alter or swap is a key they can choose. Wrap with an
        authenticated construction: a KMS wrap (AWS KMS documents that it protects data keys with
        authenticated encryption, with the encryption context in the AAD), AES Key Wrap (RFC 3394
        — unwrap must reject the key when the integrity check value does not come back; the Python
        `cryptography` package raises `InvalidUnwrap` on a one-bit change, checked 49.0.0), or an
        AEAD such as AES-GCM with the key's identity as AAD. A DEK wrapped with bare CBC/CTR/ECB is
        a High finding: tampering goes undetected. OWASP: Key Management cheat sheet.
      - **Back up the keys whose loss is data loss.** Crypto-shredding cuts both ways: a root or
        tenant key lost by accident deletes that data as surely as a deliberate destroy. For every
        key that is the only path to product-critical data, keep a backup in a separate failure
        domain (offline/cold, or a second KMS/HSM account or region, protected to the same bar as
        the original), split custody so no single person can restore it alone, and a **restore
        drill** on a schedule — a backup that has never been restored is a hope. Never back up
        signing or authentication keys (sota-code-security rules/04 on key loss vs escrow).
        OWASP: Secrets Management cheat sheet.
      - AUDIT: literal 32/64-hex-char "ENCRYPTION_KEY" values (High; Critical if committed),
        AES-ECB or static-IV usage near such keys, `Fernet(key)` with a key from code, DEKs stored
        *unencrypted* next to data, KMS `Decrypt` permission granted on `*`.
      
      ```python
      # BAD — raw key in config, hand-rolled crypto, no rotation story
      cipher = AES.new(bytes.fromhex(os.environ["ENCRYPTION_KEY"]), AES.MODE_ECB)
      
      # GOOD — KMS envelope per record, context-bound, wrapped DEK stored with ciphertext
      dek = kms.generate_data_key(KeyId=ROOT_KEY_ARN, KeySpec="AES_256",
                                  EncryptionContext={"tenant": tenant_id})
      ct = AESGCM(dek["Plaintext"]).encrypt(nonce := os.urandom(12), payload, tenant_id.encode())
      store(record_id, ct, nonce, wrapped_dek=dek["CiphertextBlob"])
      del dek  # drop plaintext DEK immediately; zeroize where the runtime allows
      ```
      
      ## 8. `.env` file discipline
      
      `.env` files are a dev convenience, not a production secret store.
      
      - **Never committed:** `.env`, `.env.local`, `.env.production` etc. in `.gitignore` from repo
        creation; a tracked `.env` with real values is High (Critical if prod). Check history, not
        just HEAD (rules/04 §4).
      - **`.env.example` is committed** and lists every variable with empty or obviously-fake values
        (`STRIPE_KEY=sk_test_REPLACE_ME`) — it is documentation. Realistic-looking values in the
        example file mask scanner findings (Low) and get copy-pasted into real use.
      - **Local values are dev-scoped** (test keys, local DB) — never paste prod values into a
        laptop `.env`; that file is outside every audit log and backup policy. If devs need prod-like
        data access, that's a gated, logged, temporary credential from the secret store
        (`vault login` + dynamic creds, `aws-vault exec`), not a static copy.
      - **Git-ignored is not assistant-invisible.** An AI coding assistant with workspace access
        reads the working tree, not the index: a `.env` or key file that `.gitignore` keeps out of
        commits is still in its reach, and whatever it reads can be sent to the model provider and
        kept in a session transcript (rules/04 §7). Hold even dev credentials outside the tree — a
        vault/keychain, or injected at launch (`op run`, `aws-vault exec`) — and treat the tool's
        exclusion list as a filter, not a boundary: Claude Code's documented `Read(./.env)` deny
        covers its file tools and the shell reads it recognises (`cat`, `head`), but not a script
        that opens the file itself; only its OS sandbox blocks every process. While an assistant can
        see your editor or terminal, do not open `.env`/key files or paste a credential into the
        shell — the open file and terminal output are context too. Audit: an ignored secret file in
        the tree of a repo where assistants are used = **Medium** if it holds shared or prod
        credentials, **Low** for dev-only test values. OWASP: Secure Coding with AI cheat sheet.
      - **Production:** the platform injects env/files from the secret manager (rules/02 §6);
        a `.env` file on a prod host or COPY'd into an image (grep Dockerfiles for `COPY .env`,
        and check `.dockerignore` excludes `.env*`) is a High finding.
      - **Permissions `0600`**; don't `source .env` in shells with `set -x` or in CI steps with
        echoed commands; don't print env in debug scripts.
      - Loader behavior: dotenv loads only in dev (`if NODE_ENV !== 'production'`-style guards or
        dev-only dependency), so a stray file can't silently override prod config.
      
      ```js
      // BAD — dotenv unconditionally. NOT because the file beats the platform: it does not.
      // Measured on dotenv 17.4.2 — default config() leaves an already-set variable alone, and
      // only config({ override: true }) lets the file win. The real risk is the opposite shape:
      // a stray .env SUPPLIES a variable the platform forgot to set, so a value nobody reviewed
      // becomes the effective config and looks like it came from the platform.
      require("dotenv").config();
      
      // GOOD — dev-only, explicit, and never overriding real environment
      if (process.env.NODE_ENV !== "production") {
        require("dotenv").config({ override: false });
      }
      ```
      
      ```gitignore
      # Repo-template .env discipline (see rules/04 §2 for the full ignore set)
      .env
      .env.*
      !.env.example
      ```
      
      Package-manager rc files are `.env`-class: `.npmrc`/`.yarnrc.yml` with `_authToken`, `.pypirc`
      with passwords, `pip.conf` with index creds, `.netrc`. Keep tokens out of the committed rc —
      use env interpolation (`//registry.npmjs.org/:_authToken=${NPM_TOKEN}`) so the file is safe to
      commit and the token rides the secret layer; a literal token in any rc file in VCS is High.
      Registry publish tokens are no longer ordinary long-lived secrets: after the Shai-Hulud worm,
      npm permanently revoked all classic tokens (Dec 2025, creation disabled) and caps granular
      write tokens at 90 days — publish from CI via trusted publishing (per-job OIDC, no stored
      token; rules/01 §4) and treat any residual registry token as a ≤90d rotating secret.
      
      ## Audit checklist
      
      - [ ] DB access uses IAM auth or dynamic creds where supported; static passwords rotate via
            dual users; app DB users least-privilege; no DSNs with userinfo in code/logs; TLS to DB.
      - [ ] API keys: per-service per-env, vendor-side restricted, test keys in non-prod/CI, ≤90d
            rotation, none in client-shipped code; webhook signatures constant-time + dual-secret +
            replay-protected.
      - [ ] Signing keys non-exportable in KMS/HSM (or Sigstore keyless); one key per purpose;
            verification material published/versioned; signing operations audit-logged.
      - [ ] TLS: ACME automation, ≤90d certs, keys generated in place, `0400`, no key files in
            repos/images/build contexts; internal mTLS via private CA; no disabled cert verification.
      - [ ] SSH: Ed25519 per-person per-device or SSH CA certs; `authorized_keys` inventoried and
            pruned on offboarding; no private keys in repos/images; no unrestricted CI SSH keys;
            agent forwarding off.
      - [ ] JWT: asymmetric alg for multi-verifier setups, algorithm pinned, no `none`/confusion
            paths, `kid`-based JWKS rotation with refresh-on-unknown-`kid`, ≤15m access-token TTL,
            no JWT secrets with defaults, no real JWTs in fixtures.
      - [ ] Field/object encryption uses KMS envelope (wrapped DEKs, encryption context, maintained
            SDK); no raw keys in config; per-tenant keys where deletion guarantees are needed.
      - [ ] **Wrapped/stored keys are integrity-protected (§7)** — KMS wrap, RFC 3394 key wrap or an
            AEAD; a DEK wrapped with unauthenticated CBC/CTR/ECB = **High**:
            `grep -rniE '(wrap|dek|kek).*(MODE_(CBC|ECB|CTR)|AES/(CBC|ECB|CTR)|aes-(128|256)-(cbc|ecb|ctr))|(MODE_(CBC|ECB|CTR)|AES/(CBC|ECB|CTR)|aes-(128|256)-(cbc|ecb|ctr)).*(wrap|dek|kek)' .`
      - [ ] **Keys whose loss is data loss are backed up (§7)** in a separate failure domain, under
            split custody, with a dated restore drill — **High** for a production root/tenant key
            with no restorable backup (judgment: read the key register and the last drill record).
      - [ ] **Wildcard TLS keys confined (§4)** to one terminating tier and issued below the apex —
            Medium when the same wildcard key is referenced by more than one host's config:
            `grep -rniE '(ssl_certificate_key|SSLCertificateKeyFile|secretName|key_?file|private_key)[^#]*wildcard' .`
      - [ ] `.env*` untracked (HEAD and history) with `0600`; `.env.example` fake-valued; dotenv
            dev-only and non-overriding; `.dockerignore` excludes `.env*`; no prod values in local
            env files; rc files (`.npmrc`, `.pypirc`, `.netrc`) use env interpolation, never literal
            tokens; registry publishing uses trusted publishing (OIDC), not stored long-lived tokens.
      - [ ] **No secret files inside the working tree where an AI assistant can read them** (§8) —
            Medium for shared/prod values, Low for dev-only. List the ignored ones the index hides:
            `git ls-files --others --ignored --exclude-standard | grep -E '(^|/)(\.env(rc|\.[^/]+)?|[^/]*\.(pem|key|p12|pfx))$' | grep -v -E '(^|/)\.env\.(example|sample|template)$'`
      
  • SKILL.md 11.5 KB
    ---
    name: sota-secrets-management
    description: >-
      State-of-the-art secrets management for building and auditing software. Use
      whenever a task involves creating, storing, injecting, rotating, or scanning
      for credentials — or reviewing code/infrastructure for secret leaks and
      misuse. BUILD mode: implementing secrets handling; AUDIT mode: sweeping a
      repo for leaked, hardcoded, or mishandled secrets. Trigger keywords: secret,
      secrets management, credential, API key, token, password, private key,
      signing key, JWT secret, TLS key, SSH key, database password, connection
      string, .env, dotenv, environment variable, Vault, OpenBao, AWS Secrets
      Manager, GCP Secret Manager, Azure Key Vault, KMS, envelope encryption,
      SOPS, age, sealed-secrets, external-secrets, workload identity, OIDC
      federation, SPIFFE, SPIRE, IAM role, GitHub Actions OIDC, short-lived
      credential, rotation, revocation, key expiry, gitleaks, trufflehog, secret
      scanning, leaked key, hardcoded secret, committed secret, git history purge,
      honeytoken, pre-commit hook, kid rotation.
    ---
    
    # SOTA Secrets Management
    
    ## Purpose
    
    Eliminate static secrets where possible; where not possible, make every secret short-lived,
    narrowly scoped, runtime-injected, auditable, and rotatable without downtime. This skill covers
    the full lifecycle (generation → distribution → storage → use → rotation → revocation → expiry),
    storage backends, application handling patterns, leak detection, incident remediation, and
    per-credential-type rules. It serves two workflows: **BUILD** (write correct secrets handling
    into new or existing code) and **AUDIT** (sweep a repo for secret issues and report findings).
    
    The hierarchy of preference, always:
    
    1. **No secret at all** — workload identity / OIDC federation / cloud IAM roles.
    2. **Short-lived, auto-issued secret** — Vault dynamic creds, STS tokens, SPIRE SVIDs.
    3. **Long-lived secret in a managed backend** — secret manager + rotation + audit log.
    4. **Encrypted secret in the repo** — SOPS+age / sealed-secrets, GitOps only.
    5. **Plaintext secret anywhere** — never acceptable.
    
    When you write or review code, push the design as far up this hierarchy as the platform allows,
    and document why if you stop below level 2.
    
    ## BUILD mode
    
    Use when implementing anything that consumes or manages a credential.
    
    1. **Classify the credential.** Type (DB cred, API key, signing key, TLS key, …) determines the
       rules — read `rules/05-credential-types.md` for the matching section before writing code.
    2. **Try to eliminate it.** Cloud-to-cloud or CI-to-cloud calls should use workload identity
       (OIDC, IAM roles, SPIFFE) — see `rules/01-lifecycle-and-workload-identity.md`. Only fall back
       to a stored secret when no federation path exists (e.g., third-party SaaS API key).
    3. **Pick the storage backend** per environment using the decision table in
       `rules/02-storage-backends.md`. Never invent a custom encrypted store.
    4. **Wire injection at runtime** — file mount or env var populated by the platform, fetched via
       SDK with caching/TTL, never baked into images or code. Patterns and good/bad pairs in
       `rules/03-application-patterns.md`.
    5. **Design rotation before shipping.** Every secret needs: an owner, a rotation procedure that
       works with zero downtime (dual-secret / kid overlap), an expiry or rotation interval, and a
       revocation path. If you cannot answer "how do we rotate this at 3am during an incident,"
       the design is not done.
    6. **Add guardrails:** pre-commit scanning config, CI secret-scan gate, `.gitignore` entries for
       `.env*` and key files, log-scrubbing for the new secret's shape
       (`rules/04-detection-and-remediation.md`).
    7. **Self-review against the Audit checklist** at the end of every rules file you used.
    
    ## AUDIT mode
    
    Use when asked to find secret leaks/misuse in an existing repo.
    
    ### Sweep procedure
    
    1. **Tooling pass (if available):** run `gitleaks git --redact .` and/or
       `trufflehog filesystem .` (and `git log` history scan when the repo has history). Treat tool
       output as candidates, not verdicts — verify each hit.
    2. **Manual grep pass** for what tools miss. Sweep at minimum:
       - High-entropy strings and known prefixes: `AKIA`, `ASIA`, `ghp_`, `gho_`, `ghs_`, `ghu_`,
         `ghr_`, `github_pat_`, `xoxb-`, `xoxp-`, `sk-`, `sk_live_`, `rk_live_`, `AIza`, `ya29.`,
         `glpat-`, `npm_`, `dop_v1_`, `shpat_`, `eyJhbGciOi` (inline JWTs), `-----BEGIN ([A-Z0-9]+ )*PRIVATE KEY( BLOCK)?-----` (ERE; rules/04 §6 has the tested command).
       - Assignment patterns: `(password|passwd|pwd|secret|token|api[_-]?key|auth)\s*[:=]\s*['"][^'"]{6,}`.
       - Connection strings with embedded creds: `://[^/:@\s]+:[^@\s]+@` (postgres, mysql, mongodb,
         amqp, redis URLs).
       - Files: `.env*` tracked in git, `*.pem`, `*.p12`, `*.pfx`, `*.key`, `*.jks`, `*.keystore`,
         `id_rsa*`, `credentials.json`, `serviceaccount*.json`, `kubeconfig`, `.npmrc`/`.pypirc`
         with tokens, `terraform.tfstate` (state files contain plaintext secrets).
    3. **History pass:** `git log -p` / `gitleaks git --log-opts` for secrets removed from HEAD
       but live in history. A secret deleted in a later commit is **still leaked** — severity is
       unchanged.
    4. **Handling pass (misuse, not just leaks):** secrets in log statements, error messages,
       exception payloads, crash/telemetry dumps, URLs/query strings, CLI args (visible in `ps`),
       Dockerfile `ENV`/`ARG`, docker-compose `environment:` literals, Kubernetes manifests with
       stringData/base64 secrets committed, CI YAML with inline values, debug endpoints dumping
       config, world-readable key files, missing rotation/expiry on long-lived tokens, overly broad
       token scopes.
    5. **Verify and triage** each finding: is the value real (test-shaped? placeholder? entropy?),
       is it currently valid, what blast radius. Never call a credential live by invoking it against
       production without explicit permission; judge from context.
    
    ### Severity conventions
    
    | Severity | Definition | Examples |
    |---|---|---|
    | **Critical** | Valid (or must-assume-valid) secret exposed to anyone with repo/log access | Live cloud key in code or git history; DB password in a public image; private signing key committed |
    | **High** | Secret exposed in a narrower channel, or handling that will leak under normal operation | Secret logged at info level; cred in CLI args; `.env` with real values tracked in private repo; token in URL; tfstate with secrets in VCS |
    | **Medium** | Weak lifecycle or weak protection of an otherwise contained secret | No rotation for years; long-lived token where OIDC is available; overly broad scope; secret in env var where platform supports file mounts; world-readable key file; weak generation (low entropy) |
    | **Low** | Hygiene gaps with no current exposure | Missing pre-commit/CI scanning; `.env.example` containing realistic-looking values; no `.gitignore` for key files; missing audit logging on secret access |
    
    Confirmed-fake placeholders (`changeme`, `xxx`, `<YOUR_KEY>`, obvious test fixtures) are not
    findings, but note them as Low if they are realistic enough to mask real leaks in scans.
    
    ### Finding format
    
    Report every finding as:
    
    ```
    [SEVERITY] path/to/file.py:123 — RULE-ID short title
      Evidence: the offending line, with the secret value REDACTED (show prefix + length only)
      Why: one sentence of impact
      Fix: concrete remediation (rotate first, then remove; target pattern to adopt)
      Effort: trivial | small | medium | large
    ```
    
    Order findings Critical → Low. End the audit with: counts per severity, whether git history is
    affected (if so, remediation must include rotation + history purge per
    `rules/04-detection-and-remediation.md`), and the top 3 systemic fixes. **Never reproduce a
    discovered secret in full in your report** — redact to first 4 chars + length.
    
    ## Rules index
    
    | File | Read this when... |
    |---|---|
    | `rules/01-lifecycle-and-workload-identity.md` | Generating secrets (entropy/length), setting rotation/expiry/revocation policy, replacing static secrets with OIDC federation, SPIFFE/SPIRE, cloud IAM roles, GitHub Actions OIDC |
    | `rules/02-storage-backends.md` | Choosing where a secret lives: Vault/OpenBao, AWS/GCP/Azure secret managers, SOPS+age, sealed-secrets/external-secrets in Kubernetes, env vars vs file mounts (no cross-deployment shared mounts), in-memory handling and zeroization, consumer-key end-to-end encryption (§8) |
    | `rules/03-application-patterns.md` | Writing app code that consumes secrets: config layering, runtime injection, caching/TTL, keeping secrets out of code/VCS/logs/errors/URLs/argv/crash dumps, per-env separation, least-privilege scoping, access audit logging; **never-persist-raw (§2.1)** — text whose producer you do not control is stored as a digest plus a structural skeleton, because shape redaction is an enumeration |
    | `rules/04-detection-and-remediation.md` | Setting up gitleaks/trufflehog, pre-commit hooks, CI gates; responding to a leak (rotate-first, purge history, assume compromised); honeytokens; running an AUDIT sweep; **agent session transcripts as a credential store (§7)** — inventorying and scoping the scan of `~/.claude/projects/**`, and why a tool that sweeps it is a secret-processing tool |
    | `rules/05-credential-types.md` | Handling a specific credential class: DB creds, API keys, signing keys, TLS private keys, SSH keys, JWT secrets and kid rotation, data keys vs KMS envelope encryption, .env discipline |
    
    ## Top-10 non-negotiables
    
    1. **A leaked secret is rotated first, scrubbed second.** History rewriting without rotation is
       theater — assume every secret that ever touched git, logs, or a ticket is compromised.
    2. **No secrets in source code or VCS, ever** — including "temporarily," tests against real
       services, example files with live values, and committed `.env` files.
    3. **Prefer no secret to a managed secret:** if OIDC federation / IAM roles / SPIFFE can replace
       a static credential (CI→cloud, service→cloud, pod→service), use it. Static keys for cloud
       access from CI are a defect, not a choice.
    4. **Every secret has an expiry or rotation interval and a documented zero-downtime rotation
       procedure** (dual-secret overlap, JWT `kid`, DB dual users). Unrotatable = misdesigned.
    5. **Generate secrets with a CSPRNG, ≥256 bits of entropy** for opaque tokens (≥32 random bytes
       before encoding); never derive from timestamps, UUIDv4-as-secret, or human-chosen strings.
    6. **Secrets never appear in:** logs, error messages, exception payloads, crash dumps, URLs or
       query strings, CLI arguments, `ps` output, Dockerfile layers, image env, shell history, or
       telemetry. Scrub at the logger and wrap in redacting types.
    7. **Inject at runtime, never at build time.** Images, artifacts, and bundles are
       secret-free; the platform (orchestrator, secret-manager SDK, CSI driver) supplies values
       when the process starts — file mounts preferred over env vars where supported.
    8. **Least privilege and per-environment separation:** one credential per consumer per
       environment, scoped to the minimum actions/resources; dev/staging/prod never share secrets
       and prod values are unreadable from non-prod.
    9. **Application data encryption uses KMS envelope encryption** (encrypt data with a data key,
       wrap the data key with a KMS key); never hardcode or hand-manage raw encryption keys.
    10. **Scanning is mandatory and layered:** pre-commit (gitleaks/trufflehog) on every developer
        machine, a blocking CI gate, and periodic full-history scans. Detection without a leak
        runbook is incomplete — keep the rotate→revoke→purge→monitor runbook current.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related