deps-audit
Audit Hex deps for supply-chain security risk — bidi chars, compile-time exec, maintainer changes, typosquats, CVEs. Use after mix deps.update, when checking if a package upgrade is safe, or reviewing mix.lock PR diffs.
Install
npx skills add https://github.com/oliver-kriska/claude-elixir-phoenix/tree/main/plugins/elixir-phoenix/skills/deps-audit
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oliver-kriska-claude-elixir-phoenix@llmmart
git clone https://github.com/oliver-kriska/claude-elixir-phoenix.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oliver-kriska/claude-elixir-phoenix collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Hex Dependency Audit
Non-mutating supply-chain audit for Hex packages. Runs an 8-rule MVP catalogue
against changed packages, enriches with Hex API metadata, wraps existing tools
(mix hex.audit, mix_audit, OSV-Scanner), and emits a triage table.
When to Use
- After
mix deps.updateormix deps.getbrought in new versions - On PRs that touch
mix.lock(pre-merge gate) - Before manually updating a single package (
--preview <pkg>) - When investigating a dependency you don't recognize
Iron Laws
- NEVER claim a diff is clean without inspecting it. Run all 8 rules
on the unpacked NEW tarball. "Looks fine" without a tool run is a false
pass. Always write
.claude/deps-audit/last-run.json— its absence is evidence the audit didn't actually run. - NEVER install
mix_audit/osv-scanner— even if asked. Detect, warn with install instructions, skip cleanly if missing. If the user says "install it," respond with the install command (e.g.,mix deps.add mix_audit --only dev) and do not execute it. The audit skill is non-mutating;mix.exs/mix.lockare off-limits regardless of consent. - NEVER promote a finding to BLOCK without rule citation. Every finding
shows
rule_id,severity,file:line,snippet,message. No handwaving. - NEVER fetch from Hex API without rate-limiting. Cap at 5 req/sec. Cache metadata 7 days, top-500 list 1 day.
- NEVER run the audit on already-committed lock changes silently —
tell the user which mode (A/B/C) is active and which
(old, new)pairs resolved. - LLM triage only above threshold. Native rules + Semgrep + YARA
are deterministic. The
hex-deps-triageragent runs only when score10 (1 BLOCK or 3+ WARNs), and its verdicts are advisory — never auto-suppress a finding without human review.
Operating Modes
| Mode | Trigger | Old source | New source |
|---|---|---|---|
| B (default) | /phx:deps-audit |
git show HEAD:mix.lock |
working mix.lock |
| C (PR) | /phx:deps-audit --base main |
git show <ref>:mix.lock |
working mix.lock |
| A (preview) | /phx:deps-audit --preview httpoison |
locked version | Hex API latest |
See ${CLAUDE_SKILL_DIR}/references/operating-modes.md for full resolver logic.
Execution Flow
Default = full 8-rule scan with streaming progress. --quick opts out
to CVE + retirement only. See ${CLAUDE_SKILL_DIR}/references/execution-flow.md.
Step 1: Resolve the diff
Parse the mix.lock Erlang term format for both old and new sources. Emit a
list of {pkg, old_version, new_version} tuples. Surface
new-only and removed-only packages separately (a removed package is not
audited; a brand-new package gets old_version = nil and skips diff-only
rules).
See ${CLAUDE_SKILL_DIR}/references/diff-resolver.md for shell + mix run -e snippets per mode and the JSON output contract.
Step 2: Fetch tarballs (per-run tmpdir)
For each (pkg, old, new):
mix hex.package fetch <pkg> <old> --unpack -o ${AUDIT_TMPDIR}/tarballs/<pkg>/<old>/
mix hex.package fetch <pkg> <new> --unpack -o ${AUDIT_TMPDIR}/tarballs/<pkg>/<new>/
All ephemeral artifacts live under ${AUDIT_TMPDIR} (driver-owned, removed
on exit). See ${CLAUDE_SKILL_DIR}/references/audit-tmpdir.md and
${CLAUDE_SKILL_DIR}/references/tarball-fetcher.md.
Step 3: Run the 8 MVP rules on each NEW tarball
| # | Rule | Sev | Method |
|---|---|---|---|
| 1 | Bidi Unicode control chars in .ex/.exs/.erl |
BLOCK | grep |
| 2 | Code.eval_* / :erlang.apply with non-literal MFA at module scope |
BLOCK | AST (Sourceror or regex+scope) |
| 3 | System.cmd / :os.cmd / Port.open at compile time |
BLOCK | AST |
| 4 | :erlang.binary_to_term/1 on literal without :safe |
BLOCK | AST |
| 5 | New :git/:path dep in mix.exs (vs old) |
BLOCK | AST diff |
| 6 | Maintainer change between versions | BLOCK | Hex API |
| 7 | Base64 blobs >256 chars outside priv/static/, test/fixtures/, assets/ |
WARN | regex |
| 8 | Levenshtein ≤2 from top-500 + download delta >1000× | BLOCK | Hex API + fuzzy |
Full catalogue (35 rules, MVP marked) in ${CLAUDE_SKILL_DIR}/references/heuristics.md.
Bash + mix run -e implementations for all 8 MVP rules in
${CLAUDE_SKILL_DIR}/references/rules-impl.md (single-pass NEW + diff rules +
Hex API rules, with run_all_rules master loop).
Step 4: External tool wrappers (parallel)
mix hex.audit— retired-package check, always availablemix_audit— CVE check via GHSA, if installed (else warn + skip; do NOT install)osv-scanner— CVE check via OSV.dev, if installed (else warn + skip; do NOT install)
See ${CLAUDE_SKILL_DIR}/references/external-tools.md for detection, output parsing, and severity mapping per tool.
Step 5: Hex API enrichment (per package)
GET /api/packages/:name— owners, downloads, inserted_atGET /api/packages/:name/releases/:version— per-release publisher- Compute:
days_since_publish,owner_age_days,download_velocity
Cap at 5 req/sec. Per-run cache under ${AUDIT_TMPDIR}/hex-api/.
See ${CLAUDE_SKILL_DIR}/references/hex-api.md for endpoint contracts,
caching strategy, Rule 6/8 detection, and Levenshtein implementation.
Step 5.5: Apply hex_vet.exs ledger (if present)
If hex_vet.exs exists at project root, vetted-version findings are
downgraded to INFO. Unvetted versions retain their severity.
Lock-vs-ledger disagreement: lock wins. See the deps-vet skill's
hex-vet schema doc for the "Lock-vs-ledger disagreement" section.
Use /phx:deps-vet <pkg> <version> (separate skill) to add entries.
Step 5.7: Differential subtract
When run with DIFFERENTIAL=1 (default), findings that existed in the
OLD tarball are downgraded to INFO. Net-new signals reach the renderer
at full severity. See ${CLAUDE_SKILL_DIR}/references/differential.md.
Step 5.8: LLM triage (when score > threshold)
For packages where the aggregate score exceeds 10, the
hex-deps-triager sonnet agent reads finding + diff windows and
produces structured verdicts (confidence, verdict, rationale,
fp_reasons[]). A context-supervisor consolidates verdicts
across packages into triage/consolidated.md. Main skill reads only
the consolidated file.
See ${CLAUDE_SKILL_DIR}/references/llm-triage.md.
Step 6: Score & render
Per-package weighted sum: BLOCK = 10, WARN = 3, INFO = 1. Risk band: 0 clean · 1–5 low · 6–15 medium · 16+ high.
Output:
- Stdout: markdown table —
pkg | old → new | risk | findings | diff.hex.pm | maintainer-changeplus a per-package detail section for any non-clean row. - Sidecar (MANDATORY): Write
.claude/deps-audit/last-run.json. The Phase 3 gate reads this; an audit that doesn't write it is a no-op for the gate. Always emit, even on clean runs.
--json flag emits JSON to stdout instead of markdown. See
${CLAUDE_SKILL_DIR}/references/output-renderer.md for table format,
sidecar schema, exit-code rubric, and --quiet mode.
Out of scope / Phase 3 surface
- NEVER modify
mix.lock,mix.exs, or any project file (non-mutating) - NEVER auto-install missing tools (warn + skip)
- Gate
mix deps.{get,update,compile}viadeps-audit-gate.sh. See${CLAUDE_SKILL_DIR}/references/hook.md. - Prompt for
/phx:compoundafter BLOCK findings — corpus self-feeds. - Emit SARIF 2.1.0 via
--sarif <path>and gate CI via--ci.
References
${CLAUDE_SKILL_DIR}/references/heuristics.md— full 35-rule catalogue${CLAUDE_SKILL_DIR}/references/rules-impl.md— bash +mix run -efor the 8 MVP rules${CLAUDE_SKILL_DIR}/references/operating-modes.md— Mode A/B/C resolver${CLAUDE_SKILL_DIR}/references/diff-resolver.md— shell snippets, lock parser${CLAUDE_SKILL_DIR}/references/tarball-fetcher.md— fetch wrapper, parallel cap, cache prune${CLAUDE_SKILL_DIR}/references/external-tools.md—mix_audit,osv-scannerwrappers${CLAUDE_SKILL_DIR}/references/hex-api.md— endpoint contracts, rate limit, Rule 6/8${CLAUDE_SKILL_DIR}/references/output-renderer.md— markdown, JSON v1, exit codes, SARIF${CLAUDE_SKILL_DIR}/references/testing.md— smoke runner, fixture matrix${CLAUDE_SKILL_DIR}/references/differential.md/llm-triage.md— Phase 2 NDJSON subtract + triager${CLAUDE_SKILL_DIR}/references/semgrep.md/yara.md— Phase 2 precision layers (soft deps)${CLAUDE_SKILL_DIR}/references/cassettes.md/sarif.md/hook.md/ci-integration.md— Phase 3 surface${CLAUDE_SKILL_DIR}/references/trusted-publishers.md/skill-checklist.md— upstream + eval${CLAUDE_SKILL_DIR}/references/audit-tmpdir.md— Phase 5 per-run ephemeral storage contract${CLAUDE_SKILL_DIR}/references/execution-flow.md/differential-cve.md— Phase 5 default scan + CVE diff
Files (claude-elixir-phoenix)
-
priv
-
semgrep
-
elixir-supply-chain.yaml 2.4 KB
rules: - id: elixir-compile-time-http message: HTTP fetch at compile time inside __before_compile__ languages: [elixir] severity: ERROR pattern-either: - patterns: - pattern-inside: | defmacro __before_compile__($ENV) do ... end - pattern: System.cmd("curl", $ARGS) - patterns: - pattern-inside: | defmacro __before_compile__($ENV) do ... end - pattern: HTTPoison.get(...) - id: elixir-eval-string message: Code.eval_string with non-literal argument languages: [elixir] severity: ERROR pattern-either: - pattern: Code.eval_string($PAYLOAD) - pattern: Code.eval_quoted($PAYLOAD) pattern-not-either: - pattern: Code.eval_string("$LITERAL") - pattern: Code.eval_quoted("$LITERAL") - id: elixir-dynamic-system-cmd message: System.cmd with variable first argument languages: [elixir] severity: WARNING pattern: System.cmd($CMD, $ARGS) pattern-not: System.cmd("$LITERAL", $ARGS) - id: elixir-base64-to-eval message: Base.decode64 result fed directly to Code.eval_string languages: [elixir] severity: ERROR pattern-either: - pattern: Code.eval_string(Base.decode64!($X)) - patterns: - pattern: | $X = Base.decode64!(...) ... Code.eval_string($X) - id: elixir-on-load-with-side-effects message: __on_load__ callback running System.cmd / File.write languages: [elixir] severity: ERROR pattern-either: - patterns: - pattern-inside: | def __on_load__() do ... end - pattern: System.cmd(...) - patterns: - pattern-inside: | def __on_load__() do ... end - pattern: File.write(...) - id: elixir-binary-to-term-literal message: :erlang.binary_to_term called without :safe option languages: [elixir] severity: ERROR patterns: - pattern: :erlang.binary_to_term($X) - pattern-not: :erlang.binary_to_term($X, [:safe]) - id: elixir-erlang-apply-dynamic message: :erlang.apply with non-literal module or function name languages: [elixir] severity: ERROR pattern: :erlang.apply($MOD, $FUN, $ARGS) pattern-not: :erlang.apply($MOD, :$LITERAL_ATOM, $ARGS)
-
-
yara
-
hex-malware.yar 2.6 KB · in bundle
-
-
-
references
-
audit-tmpdir.md 8.6 KB
# Audit Tmpdir — Ephemeral Per-Run Storage The audit skill maintains **no persistent on-disk cache**. Every audit invocation creates a fresh tmpdir, writes all working artifacts there, and removes the tmpdir on exit (via `trap EXIT`). The only persistent files the skill writes to the project are: - `.claude/deps-audit/last-run.json` — sidecar consumed by the Phase 3 PreToolUse gate - `.claude/deps-audit/policy.exs` — Phase 3 gate policy config (manually created by the user; not generated by the audit) Everything else — tarballs, mix_audit JSON, CVE diffs, hex API metadata, findings NDJSON, SARIF output — lives in `${AUDIT_TMPDIR}` for the duration of one audit run only. ## Why ephemeral, not persistent The Phase 1-4 design used `.claude/deps-audit/cache/` with a 7-day metadata TTL and `prune_cache()` for tarballs >30 days old. The 2026-05-13 design review identified three problems: 1. **Cache invalidation is hard.** Rules change between releases, lock resolution semantics drift, and stale cache entries silently bypass updated logic. The Phase 4 `cache_signature.json` plan tried to protect against rule-version skew, but adds complexity that ephemeral storage sidesteps entirely. 2. **Audit semantics demand reproducibility.** A run that says "no vulnerabilities" must reflect today's GHSA cache and today's Hex API, not last week's. Persistent caches encourage exactly the kind of "but it passed yesterday" false comfort that supply-chain tooling must resist. 3. **The latency cost is bounded and acceptable.** A 25-package audit takes 60-90s cold (per virgil dogfood). The 4-way parallel tarball fetcher is the same in both architectures — caching only saves the second and subsequent runs of the *same* diff, which is a rare workflow in practice (people don't re-audit identical locks). ## Driver pattern Every entry-point into the audit MUST establish `AUDIT_TMPDIR` and register the EXIT trap before any other work: ```bash audit_main() { # Establish per-run tmpdir + cleanup trap. This is the FIRST thing # the driver does — before parsing args, before resolving the diff, # before any tool invocation that writes state. AUDIT_TMPDIR="$(mktemp -d -t phx-deps-audit-XXXXXX)" || { echo "ERROR: mktemp failed" >&2 return 2 } export AUDIT_TMPDIR # shellcheck disable=SC2064 trap "rm -rf '${AUDIT_TMPDIR}'" EXIT # ... resolve diff, fetch tarballs, run rules, diff CVEs, render ... # All wrappers read AUDIT_TMPDIR from env. } ``` The trap fires on: - Normal exit (rc=0 success path) - Non-zero exit (any function `return N` or `exit N` that bubbles up) - Ctrl-C / SIGTERM (bash propagates these to EXIT) - Shell errors under `set -e` (also flow through EXIT) ## Cross-tool-call handoff (CRITICAL — runtime reality) `audit_main()` above describes **one** long-lived shell. The actual runtime is different: each step of `/phx:deps-audit` is a **separate Bash tool call** with a fresh shell. Consequences: 1. **`export AUDIT_TMPDIR` does NOT survive** between tool calls. The var set in call N is gone in call N+1. 2. **Shell functions do NOT survive** either. `export -f` is bash-only (no-op under the user's zsh) and useless across calls regardless. 3. **The EXIT trap in call N does NOT fire for call N+1.** Cleanup must be an explicit final command (`rm -rf "$AUDIT_TMPDIR"`), not a trap. **Pattern — persist the path, re-read it everywhere.** First command: ```bash AUDIT_TMPDIR="$(mktemp -d -t phx-deps-audit-XXXXXX)" printf '%s' "$AUDIT_TMPDIR" > "${TMPDIR:-/tmp}/phx-audit-dir.txt" echo "AUDIT_TMPDIR=$AUDIT_TMPDIR" ``` Every subsequent command re-reads it as its first line: ```bash AUDIT_TMPDIR="$(cat "${TMPDIR:-/tmp}/phx-audit-dir.txt")" ``` Final command tears down both: ```bash AUDIT_TMPDIR="$(cat "${TMPDIR:-/tmp}/phx-audit-dir.txt")" rm -rf "$AUDIT_TMPDIR" "${TMPDIR:-/tmp}/phx-audit-dir.txt" ``` **Quoted-heredoc trap.** When a command writes a helper script via a heredoc, a *quoted* delimiter (`<<'EOF'`) does NOT expand `${AUDIT_TMPDIR}` — the script gets a literal `${AUDIT_TMPDIR}` that is empty in its own shell, so paths collapse to `/tarballs/...` and you get `mkdir: /tarballs: Read-only file system`. Either bake the resolved path in with an **unquoted** heredoc: ```bash AUDIT_TMPDIR="$(cat "${TMPDIR:-/tmp}/phx-audit-dir.txt")" cat > "${AUDIT_TMPDIR}/fetch.sh" <<EOF AUDIT_TMPDIR="${AUDIT_TMPDIR}" EOF ``` …or keep the heredoc quoted and pass the path as `$1`. Never assume the helper inherits the env. If a subagent or helper spawns its own `mktemp -d` (e.g., `_mix_audit_run_with_lock` for tmpdir lock-swap), that's a separate short-lived dir with its own `RETURN`/`EXIT` trap — nested cleanup, not a leak. The audit-level `AUDIT_TMPDIR` is the outer envelope. ## Standard subdirectories Convention (not enforced by code — every wrapper can place files wherever, but consistency aids debugging): ``` ${AUDIT_TMPDIR}/ ├── lock.old # diff-resolver: HEAD/base mix.lock ├── lock.new # diff-resolver: working mix.lock ├── diff.json # diff-resolver: {changed, added, removed} ├── tarballs/<pkg>/<version>/ # unpacked Hex tarballs ├── hex-api/ │ ├── packages/<pkg>.json │ └── top-500.json ├── mix-audit.json # mix_audit single-state output ├── cves_old.json # mix_audit on OLD lock (differential) ├── cves_new.json # mix_audit on NEW lock (differential) ├── diff_cves.jsonl # diff_cves.py output ├── findings.jsonl # native rule findings (current run) └── findings.old.jsonl # native rule findings (OLD tarballs) ``` `last-run.json` is **not** under `${AUDIT_TMPDIR}` — it's copied to `.claude/deps-audit/last-run.json` as the very last step of rendering, just before the tmpdir is torn down. ## Migration from Phase 1-4 paths | Phase 1-4 path | Phase 5+ path | |----------------|---------------| | `.claude/deps-audit/cache/lock.old` | `${AUDIT_TMPDIR}/lock.old` | | `.claude/deps-audit/cache/lock.new` | `${AUDIT_TMPDIR}/lock.new` | | `.claude/deps-audit/cache/diff.json` | `${AUDIT_TMPDIR}/diff.json` | | `.claude/deps-audit/cache/<pkg>/<version>/` | `${AUDIT_TMPDIR}/tarballs/<pkg>/<version>/` | | `.claude/deps-audit/cache/hex-api/...` | `${AUDIT_TMPDIR}/hex-api/...` | | `.claude/deps-audit/cache/mix-audit.json` | `${AUDIT_TMPDIR}/mix-audit.json` | | `.claude/deps-audit/cache/cves_old.json` | `${AUDIT_TMPDIR}/cves_old.json` | | `.claude/deps-audit/cache/cves_new.json` | `${AUDIT_TMPDIR}/cves_new.json` | | `.claude/deps-audit/cache/diff_cves.jsonl` | `${AUDIT_TMPDIR}/diff_cves.jsonl` | | `.claude/deps-audit/cache/findings.jsonl` | `${AUDIT_TMPDIR}/findings.jsonl` | | `.claude/deps-audit/last-run.json` | **unchanged — persistent** | | `.claude/deps-audit/policy.exs` | **unchanged — persistent (user-owned)** | References still containing the old `.claude/deps-audit/cache/` paths are pre-Phase-5 wording that will be migrated incrementally. ## Removed: `prune_cache()` The tarball-fetcher.md `prune_cache()` function (Phase 1) walked `.claude/deps-audit/cache/` and `rm -rf`'d entries older than 30 days. This function no longer exists — the EXIT trap obsoletes it. ## Removed: `cache_signature.json` Phase 4's planned `cache_signature.json` (a sha256 of rule files + commit SHA used to invalidate the cache on plugin upgrade) is dropped. No cache, no invalidation problem. ## Fixture isolation Smoke fixtures (`lab/deps-audit/smoke-test/{fixtures,corpus}.d/`) operate inside the runner's own `mktemp -d`, NOT inside `AUDIT_TMPDIR`. This is by design: fixtures must be runnable without the full audit driver established. The fixture conventions are documented in `testing.md`. ## One-time migration: clean stale `.claude/deps-audit/cache/` Projects audited under Phase 1-4 may have leftover tarballs at `.claude/deps-audit/cache/<pkg>/<version>/`. These are cruft under the ephemeral model — never read, never refreshed, just disk noise. The driver SHOULD remove this directory at startup, once, if present: ```bash # Run BEFORE establishing AUDIT_TMPDIR. Idempotent. migrate_stale_cache() { local stale=".claude/deps-audit/cache" if [ -d "${stale}" ]; then echo "info: removing stale Phase 1-4 cache at ${stale}" >&2 rm -rf "${stale}" fi } ``` This is a one-shot cleanup, not a recurring prune (there's no cache to prune anymore). After the first v2.12.0 run on a project, the directory stays gone. ## Hook compatibility The Phase 3 PreToolUse gate hook (`deps-audit-gate.sh`) reads `.claude/deps-audit/last-run.json` only. It doesn't touch `${AUDIT_TMPDIR}` at all, so hook semantics are unchanged. -
cassettes.md 8.3 KB
# VCR cassettes — Hex API fixtures for Rules 6 + 8 Rules 6 (maintainer change) and 8 (typosquat) hit the Hex API. To keep smoke fast and offline, we ship JSON cassettes that mock the two endpoints those rules consume. ## Iron Laws 1. **Cassettes are SHA-pinned.** Each cassette has a `_meta.sha` recording the sha256 of the response body at capture time. Rules that depend on a cassette validate the SHA before consuming it — silent corruption is worse than no cassette. 2. **NEVER auto-refresh.** Cassettes are committed artifacts. A maintainer's owner-change in real life MUST update a cassette in a real PR with an explicit reviewer — not via a CI auto-bump. 3. **Cassette mode opts in via env var.** `HEX_API_BASE` defaults to `https://hex.pm/api`; tests set `HEX_API_BASE=file://test-assets/hex-api-cassettes/`. Production audits never touch the cassettes. 4. **Empty cassette ≠ no maintainer.** When a cassette is absent, skip the rule with a logged warning, never silently pass. ## Endpoints covered | Endpoint | Cassette filename | Used by | |----------|-------------------|---------| | `GET /api/packages/:name` | `<pkg>.packages.json` | Rule 6, Rule 8 | | `GET /api/packages/:name/releases/:version` | `<pkg>.releases.<v>.json` | Rule 6 | ## Cassette layout ```text lab/deps-audit/test-assets/hex-api-cassettes/ # repo only, not shipped ├── phoenix.packages.json ├── phoenix.releases.1.7.20.json ├── phoenix.releases.1.7.21.json ├── jason.packages.json ├── jason.releases.1.4.4.json ├── phoeniix.packages.json # synthetic typosquat for Rule 8 └── _meta.json # SHA index, capture timestamps ``` ## `_meta.json` shape ```json { "captured_at": "2026-05-12T18:00:00Z", "capture_source": "https://hex.pm/api", "files": { "phoenix.packages.json": { "sha256": "abc123...", "endpoint": "/api/packages/phoenix", "captured_at": "2026-05-12T18:00:00Z" } } } ``` ## Response shape — `<pkg>.packages.json` Mirrors `hex.pm` API verbatim (only fields we consume): ```json { "name": "phoenix", "downloads": { "all": 192345678, "recent": 2345678 }, "owners": [ {"username": "chrismccord", "email": "chris@example.com"}, {"username": "team-phoenix", "email": "team@example.com"} ], "inserted_at": "2014-04-21T22:33:00Z", "updated_at": "2026-04-15T10:00:00Z", "latest_stable_version": "1.7.21" } ``` ## Response shape — `<pkg>.releases.<v>.json` ```json { "version": "1.7.21", "inserted_at": "2026-04-15T10:00:00Z", "publisher": { "username": "chrismccord", "email": "chris@example.com" }, "checksum": "0123456789abcdef...", "retired": null } ``` ## Capturing a cassette ```bash # Helper script — capture.sh pkg=$1 ver=$2 out_dir=lab/deps-audit/test-assets/hex-api-cassettes curl -fsSL "https://hex.pm/api/packages/${pkg}" \ | jq '.' > "${out_dir}/${pkg}.packages.json" if [ -n "${ver}" ]; then curl -fsSL "https://hex.pm/api/packages/${pkg}/releases/${ver}" \ | jq '.' > "${out_dir}/${pkg}.releases.${ver}.json" fi # Update _meta.json with sha + timestamp. python3 -c " import json, hashlib, sys from datetime import datetime, timezone meta_path = '${out_dir}/_meta.json' meta = json.load(open(meta_path)) if open(meta_path, 'r').readable() else {'files': {}} for fname in ['${pkg}.packages.json', '${pkg}.releases.${ver}.json']: path = '${out_dir}/' + fname try: body = open(path, 'rb').read() meta['files'][fname] = { 'sha256': hashlib.sha256(body).hexdigest(), 'captured_at': datetime.now(timezone.utc).isoformat() } except FileNotFoundError: pass meta['captured_at'] = datetime.now(timezone.utc).isoformat() json.dump(meta, open(meta_path, 'w'), indent=2) " ``` ## Consumer pattern (Rules 6 + 8) ```bash hex_api_get() { local endpoint="$1" if [[ "${HEX_API_BASE:-https://hex.pm/api}" == file://* ]]; then local base="${HEX_API_BASE#file://}" local cassette cassette=$(printf '%s' "${endpoint}" \ | sed -E 's|^/api/packages/([^/]+)$|\1.packages.json|; s|^/api/packages/([^/]+)/releases/(.+)$|\1.releases.\2.json|') cat "${base}/${cassette}" 2>/dev/null || { echo "cassette missing: ${cassette}" >&2 return 1 } else curl -fsSL "${HEX_API_BASE:-https://hex.pm/api}${endpoint}" fi } ``` ## Cassettes shipped with Phase 2 Phase 2 ships cassettes for: - The 10 synthetic malicious fixtures' supporting packages - A sample of 5 benign top-100 packages for smoke calibration - The 5 real-world calibration packages (`hex_core`, `hex`, `rebar3`, `tls_certificate_check`, plus mix.lock crossing CVE-2026-23940) Full top-100 cassettes are NOT shipped — they regenerate via the seed job (see `seed.md`). ## Validation in smoke The smoke `runner.sh` does not currently consume cassettes (Rules 6 + 8 aren't in the offline smoke surface). When Phase 2 wires them in, `runner.sh` will gain: ```bash export HEX_API_BASE="file://${HARNESS_ROOT}/../test-assets/hex-api-cassettes" ``` Per-fixture `expected.txt` then asserts on `rule:6` / `rule:8` counts. ## Lifecycle — Phase 3 monthly regen + drift detection Cassettes captured ad-hoc go stale without a refresh schedule. Phase 3 ships a monthly regeneration workflow and drift detection at audit runtime. ### Monthly regen workflow `.github/workflows/cassette-regen.yml` runs on the 5th of every month (staggered from the seed-regen run on the 1st) and on manual `workflow_dispatch`: 1. Check out repo, install jq + Python 3. 2. Iterate over the top-100 seed list (`lab/deps-audit/smoke-test/corpus.d/benign-100.txt`). 3. For each package, call `bash lab/deps-audit/capture.sh <pkg>` against live `https://hex.pm/api`. The script updates `_meta.json` SHA entries. 4. If any cassette content changed, open a PR via `peter-evans/create-pull-request@v6` with summary "Monthly cassette regen: N packages updated." ### 403 fallback Some org GitHub policies deny `GITHUB_TOKEN` PR creation ("Actions cannot create pull requests"). The workflow checks the `create-pull-request` action's exit + status, and on 403: 1. Uploads the regenerated `test-assets/hex-api-cassettes/` tree as a workflow artifact (90-day retention). 2. Writes a job summary: "cassettes regenerated as artifact; PR creation requires a repo admin to enable Settings → Actions → Allow GitHub Actions to create pull requests." 3. Exits 0 — the regen ran successfully even if the PR didn't. 4. Optional: posts to `${SLACK_WEBHOOK_URL}` if the secret exists. The artifact path is `cassettes-regen-<run-id>.zip`. Maintainers download it, run `tar xf …`, commit manually. ### Drift detection at audit runtime The audit body reads `_meta.json` before consuming any cassette: ```bash expected_sha=$(jq -r ".files[\"${cassette}\"].sha256" "${meta_path}") actual_sha=$(shasum -a 256 "${cassette_path}" | awk '{print $1}') if [ "${expected_sha}" != "${actual_sha}" ]; then echo "phx-deps-audit: cassette stale — ${cassette} (sha mismatch)" >&2 # Continue, but emit INFO so the renderer surfaces "stale cassette" in output. emit_finding rule:6 severity:info \ "Cassette ${cassette} SHA drift — consider regenerating." fi ``` Drift on a single cassette doesn't fail the audit. It surfaces in the renderer's "carried-over risks" section. If >10% of cassettes drift in one run, the renderer's tail prints a one-liner pointing at the regen workflow. ### Per-PR capture pattern (manual) When a user adds a `hex_vet.exs` entry for a package without a cassette, document the manual flow in the PR: ```bash bash lab/deps-audit/capture.sh <pkg> <ver> git add lab/deps-audit/test-assets/hex-api-cassettes/ ``` Reviewers should diff the cassette body and the `_meta.json` SHA update together — a maintainer change in the captured response should match the package's actual maintainer history. This catches the scenario where a malicious cassette is committed alongside a seemingly-benign code change. ### Why monthly, not weekly Weekly regen would catch drift sooner but doubles the PR noise. The audit body's drift detection covers the "I haven't run regen in 6 weeks" case by surfacing stale cassettes as INFO — users get a clear nudge without the maintainer overhead of weekly merges. Re-evaluate after the first 6 months of operation. -
ci-integration.md 6.3 KB
# CI integration — `--ci` flag and sample workflows Phase 3 adds non-interactive `--ci` mode to `/phx:deps-audit` for use in GitHub Actions, CircleCI, GitLab CI, and similar runners. Pairs with `--sarif <path>` (Phase 2) to feed results into the host platform's code-scanning UI. ## `--ci` semantics `--ci` makes three behavioral changes to the default audit: 1. **No interactive prompts.** `AskUserQuestion` would block forever in CI. `--ci` short-circuits all prompts to their conservative default (skip vetting, never auto-approve). 2. **Strict exit codes.** Three exit codes drive the gating decision: | Exit | Meaning | |------|---------| | 0 | Audit clean — no BLOCK findings | | 1 | One or more BLOCK findings — fail CI | | 2 | Required tool missing (e.g., `mix_audit` not installed and CI strictness demands it) — fail CI as misconfiguration, not a security finding | 3. **Machine-only output.** Stdout becomes JSON unless `--sarif <path>` is also supplied (in which case SARIF lands in the file and stdout stays JSON for log readability). `--ci` implies `--no-llm` by default (LLM triage adds wall-time and non-determinism unsuitable for CI). Override with `--ci --llm` when running an explicit pre-merge audit job that has the budget. ## GitHub Actions ```yaml name: Hex deps audit on: pull_request: paths: ['mix.lock', 'mix.exs', 'hex_vet.exs'] push: branches: [main] jobs: deps-audit: runs-on: ubuntu-latest permissions: contents: read security-events: write # required for upload-sarif steps: - uses: actions/checkout@v6 with: fetch-depth: 0 # full history so --base can diff against main - uses: erlef/setup-beam@v1 with: otp-version: '27.x' elixir-version: '1.18.x' - name: Audit Hex deps run: | mix phx.deps_audit --ci \ --base origin/main \ --sarif audit.sarif - name: Upload SARIF to Code Scanning if: always() # upload even on audit failure so reviewers see findings uses: github/codeql-action/upload-sarif@v3 with: sarif_file: audit.sarif category: phx-deps-audit ``` The `if: always()` on upload-sarif is critical: when the audit exits 1 on BLOCK findings, the job has `failed` status, and a strict `if: success()` would skip the upload — leaving reviewers with no in-UI feedback on what failed. ## CircleCI ```yaml version: 2.1 jobs: deps-audit: docker: - image: hexpm/elixir:1.18.0-erlang-27.0.0-alpine-3.18.0 steps: - checkout - run: name: Audit Hex deps command: | mix local.hex --force mix deps.get --only-prod mix phx.deps_audit --ci --base main --sarif /tmp/audit.sarif - store_artifacts: path: /tmp/audit.sarif destination: phx-deps-audit-sarif workflows: pr: jobs: [deps-audit] ``` CircleCI lacks a native SARIF surface but `store_artifacts` keeps the output downloadable. Pair with a separate job that posts a PR comment referencing the artifact URL. ## GitLab CI ```yaml deps-audit: stage: test image: hexpm/elixir:1.18.0-erlang-27.0.0-alpine-3.18.0 before_script: - mix local.hex --force - mix deps.get --only-prod script: - mix phx.deps_audit --ci --base "${CI_DEFAULT_BRANCH}" --sarif audit.sarif artifacts: when: always paths: [audit.sarif] reports: sast: audit.sarif # GitLab Ultimate SAST integration ``` GitLab Ultimate accepts SARIF as a SAST report. For lower tiers, fall back to `artifacts.paths` and manual download. ## Drone CI ```yaml kind: pipeline type: docker name: deps-audit steps: - name: audit image: hexpm/elixir:1.18.0-erlang-27.0.0-alpine-3.18.0 commands: - mix local.hex --force - mix deps.get --only-prod - mix phx.deps_audit --ci --base main --sarif audit.sarif when: paths: include: [mix.lock, mix.exs, hex_vet.exs] ``` Drone has no built-in SARIF UI; pair with a Slack notification step that links to the run. ## Mix-task pattern (Phase 3+ companion) The `--ci` flag works with `/phx:deps-audit` inside Claude Code AND with the planned `mix phx.deps_audit` Mix task shipped by the companion `phx_deps_vet` Hex package. Both surfaces share the same exit-code rubric so CI workflows are portable between teams using CC and teams using the Mix task directly. Until the companion package ships, CI users run the audit via the Claude Code CLI: ```bash claude code -m "/phx:deps-audit --ci --base origin/main --sarif audit.sarif" ``` This requires the runner has Claude Code installed and authenticated. For teams without that infrastructure, defer CI gating until the companion Mix task ships. ## `mix format` post-Write pattern for `.exs` data files Several workflows in this plugin scriptedly Write `.exs` files (`hex_vet.exs` ledger updates, `hex_vet_seed.exs` regen). The PostToolUse format hook in Claude Code rejects unformatted output — when CI scripts use `Code.format_string!` + `File.write!` to round- trip the file, the result still drifts from `mix format` if the `.formatter.exs` has unusual locals_without_parens or similar. Always run `mix format <path>` immediately after writing an `.exs` data file, before committing: ```bash mix run --no-mix-exs -e "..." # write the file via Code.format_string! mix format hex_vet.exs # belt-and-braces project-formatter alignment git add hex_vet.exs ``` This pattern recurs across `seed-regen.yml`, `cassette-regen.yml`, and any future regen workflow. Document it in those workflows' README sections. ## Flag matrix | Flag | CI use | Default | |------|--------|---------| | `--ci` | enable non-interactive mode | off | | `--sarif <path>` | emit SARIF 2.1.0 | off | | `--base <ref>` | diff against ref (Mode C) | `origin/main` in CI | | `--no-llm` | skip LLM triage | implied by `--ci` | | `--llm` | force LLM triage (override `--ci` default) | off | | `--no-differential` | disable Phase 2 NDJSON subtract | off | | `--json` | machine-readable stdout | implied by `--ci` | | `--strict` | promote all WARN to BLOCK | off | `--strict` is useful for "high-stakes deploy" pipelines: it makes any WARN finding exit-code-1. Most CI jobs leave it off and rely on `hex_vet.exs` `policy.block_on_unvetted: :strict` for gating instead. -
diff-resolver.md 5.6 KB
# Diff Resolver — Lock File Comparison Resolves `(pkg, old_version, new_version)` tuples for each mode. The plugin has no `lib/` directory, so this lives as shell + `mix run -e` snippets invoked from the skill body. Pseudocode is the spec; the snippets below are the runtime. ## Lock file format `mix.lock` is an Erlang term map: ```elixir %{ "phoenix" => {:hex, :phoenix, "1.7.14", "checksum", :mix, [...], "hexpm", "hash"}, "ecto" => {:hex, :ecto, "3.13.2", ...}, ... } ``` Position 3 (zero-indexed: 2) is the version string. The Erlang term format has been stable since Hex 0.20 (2019). ## Parsing — Elixir route (authoritative) Use `Code.eval_file/1` to get a proper map: ```bash mix run --no-deps-check --no-compile -e ' lock = Code.eval_file("mix.lock") |> elem(0) for {pkg, tup} <- lock, do: IO.puts("#{pkg}\t#{elem(tup, 2)}") ' 2>/dev/null ``` Output: tab-separated `pkg<TAB>version` lines. Reliable across all Hex versions. Requires the project to compile; for very early/broken states use the shell fallback below. > **Always redirect stderr.** Modern `mix.lock` files use quoted keys > (`"phoenix":`), so `Code.eval_file("mix.lock")` prints a > `found quoted keyword … please omit the quotes` **warning per > package** to stderr. On a 60-package lock that is tens of KB of > noise that gets persisted as an oversized tool result. The > `2>/dev/null` above is mandatory, not optional. **Never** inspect > the lock with `git diff … mix.lock | head` either — a real lock > diff is tens of KB and blows the tool-result budget; go straight to > `git show HEAD:mix.lock` + this parser. ## Parsing — Shell fallback (when Mix won't run) ```bash # Extract pkg + version pairs from raw mix.lock without mix awk -F'"' '/^ "/ {pkg=$2; getline; getline; if (match($0,/"[0-9][^"]+"/)) {print pkg"\t"substr($0,RSTART+1,RLENGTH-2)}}' mix.lock ``` Approximate — works for vanilla `:hex` entries, may misparse `:git` / `:path` deps (which is fine, they're flagged by Rule 5 anyway). ## Mode B — working vs HEAD ```bash : "${AUDIT_TMPDIR:?AUDIT_TMPDIR not set — driver must establish per-run tmpdir}" git show HEAD:mix.lock > "${AUDIT_TMPDIR}/lock.old" 2>/dev/null \ || echo '%{}' > "${AUDIT_TMPDIR}/lock.old" cp mix.lock "${AUDIT_TMPDIR}/lock.new" ``` Then run the parser on both files and diff in Elixir: ```bash mix run --no-deps-check --no-compile -e ' parse = fn path -> {map, _} = Code.eval_file(path) Map.new(map, fn {k, v} -> {k, elem(v, 2)} end) end old_map = parse.("${AUDIT_TMPDIR}/lock.old") new_map = parse.("${AUDIT_TMPDIR}/lock.new") changed = for {pkg, nv} <- new_map, ov = old_map[pkg], nv != ov, do: {pkg, ov, nv} added = for {pkg, nv} <- new_map, !Map.has_key?(old_map, pkg), do: {pkg, nil, nv} removed = for {pkg, ov} <- old_map, !Map.has_key?(new_map, pkg), do: {pkg, ov, nil} IO.puts(Jason.encode!(%{changed: changed, added: added, removed: removed})) ' > ${AUDIT_TMPDIR}/diff.json ``` If `Jason` isn't available (some early-stage projects), substitute `:erlang.term_to_binary` + Base64, or fall back to `inspect/2` and parse text. ## Mode C — `--base <ref>` Identical to Mode B but substitute the HEAD source: ```bash git show "${BASE_REF}:mix.lock" > ${AUDIT_TMPDIR}/lock.old 2>/dev/null \ || { echo "ERROR: ${BASE_REF}:mix.lock not found"; exit 2; } ``` Validate `<ref>` before use: ```bash git rev-parse --verify "${BASE_REF}^{commit}" >/dev/null 2>&1 \ || { echo "ERROR: ${BASE_REF} is not a valid git ref"; exit 2; } ``` ## Mode A — `--preview [pkg...]` Locked version = position 3 of the entry. Latest version = Hex API. ```bash # Locked side mix run --no-deps-check --no-compile -e ' {map, _} = Code.eval_file("mix.lock") Map.new(map, fn {k, v} -> {k, elem(v, 2)} end) |> Jason.encode!() |> IO.puts() ' > ${AUDIT_TMPDIR}/lock.locked.json # Latest side — query Hex API for each requested package for pkg in "$@"; do curl -fsSL \ -H "Accept: application/vnd.hex+json" \ "https://hex.pm/api/packages/${pkg}" \ | jq -r '.releases[0].version' \ > "${AUDIT_TMPDIR}/${pkg}.latest" done ``` Cap at 50 packages — warn and exit if `$#` > 50: ```bash if [ "$#" -gt 50 ]; then echo "WARN: --preview capped at 50 packages, got $#. Specify a subset." exit 2 fi ``` If no packages specified, expand to all keys from `mix.lock` (still capped). ## Output contract The resolver emits one JSON object to `${AUDIT_TMPDIR}/diff.json`: ```json { "mode": "B" | "C" | "A", "base": "HEAD" | "<ref>" | null, "changed": [["phoenix", "1.7.14", "1.7.20"], ...], "added": [["new_pkg", null, "0.1.0"], ...], "removed": [["old_pkg", "1.2.3", null], ...] } ``` Downstream steps (fetch, rule run, render) read this single file. ## Edge cases | Case | Behavior | |------|----------| | No `mix.lock` in HEAD (initial commit) | Treat all working entries as `added` | | Working `mix.lock` missing | ERROR — bail with `mix deps.get` suggestion | | `:git` entry in lock | Version is the SHA. Rule 5 flags it; diff still emits the tuple | | `:path` entry | Same as `:git` — flagged by Rule 5 | | Lockfile in conflict (`<<<<<<< HEAD`) | ERROR — refuse to audit, suggest resolving first | | Working ≡ HEAD (no changes) | Empty `changed/added/removed` arrays → renderer prints "No dep changes since HEAD" | ## Test fixtures `test/fixtures/deps-audit/lockfiles/` (created in Component 8): - `clean.lock` — no changes vs `clean.lock` (baseline) - `bump.lock` — single minor bump (`phoenix 1.7.14 → 1.7.20`) - `added.lock` — new package added - `git-dep.lock` — `:git` entry (triggers Rule 5) - `conflict.lock` — merge conflict markers (resolver should refuse) -
differential-cve.md 5.6 KB
# Differential CVE Pass Phase 5 capability. The audit runs external CVE scanners against both the OLD and NEW `mix.lock` and diffs the result, so the user learns **what each dep update actually patched** — not just "no current vulnerabilities." ## Why `mix_audit` only sees the current lock. If a 25-package update closes four CVEs disclosed in the last two weeks (real virgil example, 2026-05-12), the user just sees "No vulnerabilities found" — and never learns they were exposed for up to 12 days. The differential pass turns that into a security changelog. ## Architecture ``` mix_audit_diff (references/external-tools.md) ├─► run mix_audit against OLD mix.lock → cves_old.json └─► run mix_audit against NEW mix.lock → cves_new.json │ ▼ scripts/diff_cves.py │ ┌─────────┼──────────┐ ▼ ▼ ▼ patched introduced still_exposed (info) (block) (block) │ ▼ output-renderer.md headline: "🚨 N updates patched real CVEs" ``` ## Lock-swap mechanics Three strategies considered. We use **tmpdir copy** (option b). | Option | Verdict | |--------|---------| | (a) `MIX_LOCKFILE` env var | Mix 1.18 has no such variable — rejected | | **(b) tmpdir copy** | Chosen — simplest, no git state mutation | | (c) `git worktree add` | Mutates `.git/`; branch state to clean up; rejected | Implementation in `_mix_audit_run_with_lock` (external-tools.md): ```bash tmpdir="$(mktemp -d -t deps-audit-XXXXXX)" cp mix.exs config/* "${tmpdir}/" cp "${OLD_LOCK_FILE}" "${tmpdir}/mix.lock" ( cd "${tmpdir}" && mix deps.audit --format json ) rm -rf "${tmpdir}" ``` Crucially, the real `mix.lock` is **never touched**. Iron Law #2 (non-mutating) holds. `mix deps.audit` reads the lock and queries its locally-cached GHSA advisory DB. It does **not** require `mix deps.get` to run — the tmpdir copy works offline. ## CVE finding schema Each finding emitted by `diff_cves.py` (NDJSON, one per line): | Field | Type | Notes | |-------|------|-------| | `category` | `patched`/`introduced`/`still_exposed` | Set membership | | `rule_id` | `"ext:mix-audit:diff"` | Distinct from `ext:mix-audit` | | `severity` | `info`/`warn`/`block` | Per category × raw_severity | | `ghsa_id` | `GHSA-xxxx-xxxx-xxxx` | Primary key | | `cve_id` | `CVE-2026-12345` | Display only — may be absent | | `package` | `decimal` | Primary key | | `old_version` | `2.3.0` or `null` | Null when `introduced` | | `new_version` | `3.1.0` or `null` | Null when `patched` | | `severity_label` | `critical`/`high`/`moderate`/`low` | mix_audit raw | | `title` | `"DoS via unbounded exponent"` | From advisory | | `disclosed_at` | `"2026-05-07"` | ISO date or null | | `exposure_days` | `5` | `today - disclosed_at` | | `message` | `"decimal 2.3.0 → 3.1.0: CVE-... (high) — DoS"` | Pre-rendered | ## Severity per category | Category | Raw severity | Final | |----------|--------------|-------| | `patched` | (any) | `info` | | `introduced` | critical/high | `block` | | `introduced` | moderate | `warn` | | `still_exposed` | critical/high | `block` | | `still_exposed` | moderate | `warn` | `patched` is intentionally `info` — the update **is** the fix. The renderer (`output-renderer.md`) lifts these to the headline section despite their low severity, so the user sees the security changelog before any other report content. ## Set key `(ghsa_id, package)`. Version is deliberately NOT part of the key: a CVE that affected OLD `decimal 2.3.0` and ALSO affects NEW `decimal 2.4.0` is "still exposed" — the bump didn't address the CVE. When the GHSA advisory has no ID (rare), we fall back to the CVE ID. If neither is present, the entry is skipped (can't deduplicate). ## Caller flow Invoked by the audit driver after `mix_audit_run` (Step 4 of SKILL.md Execution Flow): ```bash # All artifacts live under ${AUDIT_TMPDIR} (per-run tmpdir established # by the driver — see references/audit-tmpdir.md). The driver's # `trap "rm -rf ${AUDIT_TMPDIR}" EXIT` removes everything on exit. # 1. Resolve OLD lock (Mode B: HEAD; Mode C: --base ref). git show "${BASE_REF:-HEAD}:mix.lock" > "${AUDIT_TMPDIR}/lock.old" # 2. Run mix_audit against both states (emits cves_old.json + cves_new.json). OLD_LOCK_FILE="${AUDIT_TMPDIR}/lock.old" \ NEW_LOCK_FILE=mix.lock \ mix_audit_diff # 3. Compute diff. python3 scripts/diff_cves.py \ --old "${AUDIT_TMPDIR}/cves_old.json" \ --new "${AUDIT_TMPDIR}/cves_new.json" \ --out "${AUDIT_TMPDIR}/diff_cves.jsonl" \ --summary # 4. Renderer reads diff_cves.jsonl, prepends the headline section # when any patched/introduced/still_exposed findings exist. ``` ## Disabling `--no-differential` skips the diff pass and runs only single-state `mix_audit_run` against the NEW lock (Phase 4 behavior). The CLI flag maps to setting `DIFFERENTIAL=0` in the driver. ## Failure modes | Failure | Behavior | |---------|----------| | OLD lock not in git (new project) | Skip diff; treat all NEW CVEs as `introduced` | | `mix_audit` not installed | Skip diff entirely (warn) | | Tmpdir copy fails (no disk) | Surface error; DO NOT fall back to mutating real lock | | GHSA cache stale | Emit freshness WARN; continue (results may miss new disclosures) | | OLD lock parse error | Skip diff; emit WARN | ## Related - `external-tools.md` — `mix_audit_run` wrapper + `mix_audit_diff` driver - `output-renderer.md` — security-changelog headline section - `differential.md` — Phase 2 native-rule differential (file-scoped, distinct from this CVE pass) - `../scripts/diff_cves.py` — the diff implementation -
differential.md 6.7 KB
# Differential mode — NDJSON set-subtract Phase 2 reduces the false-positive baseline by running each rule pass on the **OLD tarball** as well as the **NEW tarball**, then subtracting the findings that already existed in the dependency. Only the truly new signals reach the renderer at full severity. This isn't a function refactor — Phase 1's rules emit NDJSON to `findings.jsonl`. Differential mode runs them twice and diffs the output stream. ## Iron Laws 1. **Set-subtract on findings, not source.** Diffing files would require AST normalization and is fragile. Diffing on finding keys is rule-aware and stable across reformatting. 2. **Polymorphic keys, not a single tuple.** Three keying shapes covering all 8 rules — picking one universal key would have to degrade to "rule_id + everything," which is no key at all. 3. **Carried-over signals are INFO, not silenced.** A finding that existed before AND still exists IS still real — it's just not the news. Downgrade severity, emit with `differential: carried`. 4. **Cache invalidates on rule change.** The cache key MUST include a rules-checksum (`sha256` of `rules-impl.md` mtime + commit SHA). Otherwise a rule-semantics change leaves stale findings around. ## Architecture ``` Phase 1 engine: bash rules → findings.jsonl on NEW Phase 2 engine: bash rules → findings.jsonl on NEW bash rules → findings.old.jsonl on OLD ← new pass ↓ scripts/diff_findings.py ↓ new_signals.jsonl (in NEW, not OLD) info_signals.jsonl (in both — downgraded) dropped_signals.jsonl (in OLD, not NEW) ``` ## Polymorphic keying Single-tuple keys are fragile because Phase 1 emits **two finding shapes** (file-scoped and package-scoped) and Rule 5 is its own thing. The differ keys per rule_id: | Rule kind | Rules | Key shape | |-----------|-------|-----------| | File-scoped | 1, 2, 3, 4, 7 | `(rule_id, file, fn_name, sha256(snippet)[:12])` | | Mix-deps diff | 5 | `(rule_id, dep_name, kind)` where kind ∈ {git, path} | | Package-scoped | 6, 8 | `(rule_id, pkg)` | ### `fn_name` extraction — AST walk, with a regex fallback For file-scoped rules, the snippet alone is insufficient because two identical snippets in different functions are different findings. The extraction walks the source upward from the finding's `line` to the nearest enclosing `def` / `defp` / `defmacro` / `defmacrop`: - **Preferred** — rule emitters call `Code.string_to_quoted/2` and attach `fn_name` to the JSON object before writing it. Phase 2's emit helper extends Phase 1's: ```bash fn_name_from_ast() { # $1 = file, $2 = line mix run --no-mix-exs -e " {:ok, ast} = File.read!(\"$1\") |> Code.string_to_quoted() # Walk AST collecting {fn_name, start_line, end_line}; return # the innermost named def enclosing line ${2}, or 'module_scope'. IO.write(MyDifferential.fn_name_for_line(ast, ${2})) " } ``` - **Fallback** — `diff_findings.py` does a regex walk upward when `fn_name` is missing. Less accurate (won't see `defmacro` inside a `quote do`), but sufficient for first-pass stability. ### Why include `sha256(snippet)` in file-scoped keys? Functions get reformatted between releases — line numbers move, but the snippet text shifts less. Hashing 12 chars of the snippet gives us a stable identity even when the file is reformatted, while still distinguishing two `:erlang.binary_to_term(blob)` calls in different spots. ## Master-runner extension `run_all_rules` in `rules-impl.md` accepts an optional `OLD_DIR` environment variable. When set, every per-tarball rule (1, 2, 3, 4, 7) runs **twice** — once on `${new_dir}` writing `findings.jsonl`, once on `${old_dir}` writing `findings.old.jsonl`. Rule 5 already takes both dirs natively. Rules 6 and 8 (Hex API) are package-scoped and do not benefit from differential mode (the package itself is the unit). Once both NDJSON streams exist, the skill body invokes: ```bash python3 "${CLAUDE_SKILL_DIR}/scripts/diff_findings.py" \ --new "${AUDIT_TMPDIR}/findings.jsonl" \ --old "${AUDIT_TMPDIR}/findings.old.jsonl" \ --new-out "${AUDIT_TMPDIR}/new_signals.jsonl" \ --info-out "${AUDIT_TMPDIR}/info_signals.jsonl" \ --dropped-out "${AUDIT_TMPDIR}/dropped_signals.jsonl" ``` The renderer (see `output-renderer.md`) consumes `new_signals.jsonl` for the primary report. `info_signals.jsonl` is appended to a collapsible "carried-over risks" section. `dropped_signals.jsonl` informs the changelog-style "X risks resolved since previous version" line — useful as positive signal in PR review. ## Added-package mode A net-new dependency has no OLD tarball. The differ accepts `--added-package-mode emit-all` (default) and writes every finding to `new_signals.jsonl`. This matches Phase 1 behavior — a brand-new package gets the full audit. The alternative (`--added-package-mode skip`) is intentionally available but discouraged: skipping is the correct call only when the caller already knows the package was vetted at this version via `hex_vet.exs` (see `hex-vet.md`), and even then the ledger applies a downgrade rather than full skip. ## No persistent cache As of v2.12.0, the audit maintains no on-disk findings cache. Every run writes findings.jsonl / findings.old.jsonl / *_signals.jsonl into `${AUDIT_TMPDIR}` and the driver's EXIT trap removes everything when the audit completes. See `audit-tmpdir.md` for the full storage contract. This obsoletes earlier plans for `<rules-checksum>`-keyed cache directories and the `cache_signature.json` plugin-version-skew guard — without a persistent cache, neither problem exists. Every audit reflects the current rule set, current scorer, current GHSA advisories, and the current Hex API state. The latency cost is bounded: ~60-90s for a 25-package diff (per virgil dogfood), parallelized 4-way at the tarball fetcher. Re-auditing an unchanged lock is a rare workflow, so caching that case is low-value. ## When the differ is NOT used - `--no-differential` — debugging mode; emits Phase 1 findings.jsonl unchanged. - Single-version audit (no OLD ref, no `mix.lock` change). Differ falls back to added-package emit-all. ## Performance The differ is O(n + m) where n = NEW findings, m = OLD findings. With 8 rules × ~5 findings per package per rule × 20 packages in a typical `mix.lock` PR, a deps PR is ~800 findings. Differ runs in <200 ms. Cost dominator is the **rule pass on OLD** — running rules 1-4 + 7 twice per package roughly doubles the audit's wall time. Planned mitigation (Phase 3, not implemented): keep OLD findings cached and re-run only when the OLD tarball mtime changes. -
execution-flow.md 5.7 KB
# Execution Flow (Phase 5) Authoritative pacing for `/phx:deps-audit`. The SKILL.md "Execution Flow" section is a thin pointer; the runtime contract lives here. ## Default = full scan The default invocation runs **all 8 MVP rules** plus all available external tools (`mix hex.audit`, `mix_audit`, `osv-scanner`) plus Hex API enrichment plus the differential CVE pass. **There is no interactive choice prompt** — heuristics run by default. A "run heuristics?" prompt is a silent-failure footgun on large diffs: the user picks "no" to save time, the skill reports "no findings," and a bidi-char trojan lands in the lock with no warning. ``` /phx:deps-audit ← runs the full pipeline /phx:deps-audit --base main /phx:deps-audit --preview httpoison ``` ## `--quick`: CVE + retirement only The opt-out is `--quick`. When set, the audit skips the 8 native rules, Hex API enrichment, and the differential CVE pass — keeping only: - `mix hex.audit` (retired-package check) - `mix_audit` (GHSA CVE check, current lock only) Use case: very large diffs (>50 packages), CI gates where wall-time matters more than novel-attack detection, or pre-PR triage where the maintainer just wants to confirm no known-CVE updates are merging dirty. Target latency on a 25-package update: **<10s**. ``` /phx:deps-audit --quick ``` Flag synonyms considered: `--cve-only`, `--no-heuristics`, `--fast`. **`--quick`** was chosen for brevity and cultural alignment with `mix test --quick`. ## Streaming progress The full scan emits one line per package per phase to stdout: ``` [ 1/25] cowboy 2.13.0 — fetching tarball... [ 2/25] cowlib 2.15.0 — fetching tarball... ... [ 1/25] cowboy 2.13.0 — running rule 1 (bidi)... [ 1/25] cowboy 2.13.0 — running rule 2 (eval)... ... [ 1/25] cowboy 2.13.0 — Hex API: owners, downloads [12/25] decimal 3.1.0 — running rule 8 (typosquat)... [25/25] phoenix 1.8.7 — done (0 findings) ``` Streaming progress is the default: facing a silent terminal for 60-90s, users assume the skill has hung and Ctrl-C before results come back. Format spec: - `[N/M]` — package counter (current/total), zero-padded to align - `pkg ver` — package name and NEW version - ` — ` em-dash separator - Phase verb in present continuous (`fetching`, `running`, `enriching`) - Phase noun (optional but encouraged) Render to stdout, not stderr, so the user sees it interactively but machine consumers (`--json`, `--sarif`) can suppress with `--quiet`. ## Suppressing progress | Flag | Effect | |------|--------| | `--quiet` | No streaming output; final result only | | `--json` | Implies `--quiet`; JSON to stdout | | `--sarif PATH` | Implies `--quiet`; SARIF to PATH | | `--ci` | Implies `--quiet`; exits non-zero on BLOCK | ## Phase ordering Default pipeline in order. Each step is independent and parallelism-safe unless noted. 1. **Resolve diff** — `(pkg, old, new)` tuples from mix.lock pair 2. **Fetch tarballs** — 4-way parallel; both OLD and NEW versions 3. **Run 8 native rules** on NEW tarballs (per-package, parallel up to 4) 4. **External CVE pass** — `mix_audit_diff` (OLD + NEW), then `diff_cves.py` 5. **External retirement** — `mix hex.audit` (current lock) 6. **Hex API enrichment** — owner, downloads, rule 6/8 (rate-limited 5/s) 7. **`hex_vet.exs` ledger** — downgrade vetted versions to INFO 8. **Differential native subtract** — Phase 2 `diff_findings.py` 9. **LLM triage** (if score >10 per package) — `hex-deps-triager` agent 10. **Render** — markdown table + sidecar JSON, with security-changelog headline `--quick` skips steps 1, 3, 5b, 6, 7, 8, 9 — keeping only steps 1-lite, 2-lite (NEW only), 5a (`mix hex.audit`), 4-lite (`mix_audit` on NEW only, no diff), and 10-lite (no headline, just table). ## When to break from these defaults | Situation | Override | |-----------|----------| | CI gate, time-sensitive | `--quick --ci` | | Pre-merge PR review | (default — full scan) | | One-off package preview | `--preview pkg [pkg...]` (skips diff resolution) | | Already-vetted batch | Run normally; `hex_vet.exs` handles downgrades | | Air-gapped CI (no Hex API) | `--no-hex-api` (Phase 4 flag) | ## Wall-time budgets | Diff size | Default scan | `--quick` | |-----------|--------------|-----------| | 1-5 packages | 5-15s | <3s | | 10-25 packages | 30-90s | <10s | | 50-100 packages | 2-4 min | 15-30s | | 200+ packages | 5-10 min | <60s | Measured on a real 25-package update (~75s default) with the 4-way parallel fetcher. ## `--trace` flag (Iron Law #1 auditability) Without a trace there is a verification gap: when the audit reports "8 rules clean", there's no on-disk evidence the rules actually ran. The tmpdir is gone, no Bash trace is captured by ccrider, and a fast model could in principle synthesize a plausible verdict from the lock-diff text alone — exactly the Iron Law #1 false pass. `--trace` writes a verifiable audit log to `.claude/deps-audit/last-run.trace.log` documenting every shell command the audit ran (one per line, prefixed with timestamp + duration): ``` 2026-05-13T11:01:50.123Z [3.4s] mix hex.package fetch decimal 2.3.0 --unpack -o /tmp/phx-deps-audit-XXX/tarballs/decimal/2.3.0 2026-05-13T11:01:53.501Z [3.1s] mix hex.package fetch decimal 2.4.1 --unpack -o /tmp/phx-deps-audit-XXX/tarballs/decimal/2.4.1 2026-05-13T11:01:56.612Z [0.4s] grep -rP '[\x{202a}-\x{202e}\x{2066}-\x{2069}]' /tmp/phx-deps-audit-XXX/tarballs/decimal/2.4.1 ... ``` The trace file is the single piece of evidence that Iron Law #1 was upheld for a given run. It's not under `${AUDIT_TMPDIR}` — it survives audit completion so a reviewer can verify after the fact. The model SHOULD enable `--trace` by default when the gate hook is configured (`policy.exs` exists), and otherwise leave it opt-in. -
external-tools.md 15.2 KB
# External Tool Wrappers Three external CVE scanners layered on top of the 8 MVP rules. All optional except `mix hex.audit` (ships with mix). **Never install — even if asked.** Detect, warn with install instructions, skip cleanly. ## If the user asks "install mix_audit and re-run" **Refuse the install. Run the audit with what's available.** The audit skill is non-mutating. `mix.exs`, `mix.lock`, and the project build state must stay untouched. If the user explicitly requests an install, respond with the exact command for them to run, then continue the audit with `mix_audit` skipped: ``` I can't install mix_audit — the audit skill is non-mutating and won't modify mix.exs/mix.lock. Run this yourself, then re-invoke me: mix deps.add mix_audit --only dev mix deps.get Meanwhile, I'll continue with mix_audit skipped. CVE coverage via GHSA will be missing for this run; mix hex.audit (retirement check) and the 8 heuristic rules still cover novel-attack detection. ``` This is consent-resistant by design — a skill that mutates the project "because the user asked" is indistinguishable, from a security review perspective, from one that mutates on its own. See SKILL.md Iron Law #2. ## 1. `mix hex.audit` — retired-package check Ships with Mix. Zero-config. Detects packages explicitly retired by their maintainers via Hex.pm (security, deprecated, invalid, renamed, etc.). ```bash hex_audit() { mix hex.audit 2>&1 | tee "${AUDIT_TMPDIR}/hex-audit.txt" } ``` Output format (line per retirement): ``` phoenix_html 2.14.3 Reason: invalid Message: Upgrade to 4.x for new HTML escaping API ``` Parse with awk: ```bash awk ' /^[a-z_][a-z0-9_]* [0-9]/ { pkg=$1; ver=$2 } /Reason:/ { reason=$2 } /Message:/ { msg=substr($0, 11); print pkg"|"ver"|"reason"|"msg } ' "${AUDIT_TMPDIR}/hex-audit.txt" ``` Severity mapping: `security` → BLOCK · `invalid` / `deprecated` / `renamed` → WARN. **FP rate:** ~0%. Always integrate. ## 2. `mix_audit` — CVE check via GitHub Advisory Database Hex package (`{:mix_audit, "~> 2.1", only: [:dev, :test], runtime: false}`). Checks GHSA `pkg:hex` advisories. ```bash mix_audit_run() { if ! mix help deps.audit >/dev/null 2>&1; then cat >&2 <<'EOF' WARN: mix_audit not installed — skipping CVE check via GHSA. To enable: mix archive.install hex mix_audit # or add to mix.exs: # {:mix_audit, "~> 2.1", only: [:dev, :test], runtime: false} EOF _mix_audit_warn_cve_corpus_overlap >&2 || true return 0 fi local lock_file="${LOCK_FILE:-mix.lock}" local out_path="${MIX_AUDIT_OUT:-${AUDIT_TMPDIR:?AUDIT_TMPDIR not set}/mix-audit.json}" # GHSA freshness check (Phase 5). If the cache directory is >N hours # old, emit a WARN — recent disclosures may be missing. Threshold is # 24h by default; override via GHSA_MAX_AGE_HOURS. _mix_audit_check_ghsa_freshness >&2 || true # First `mix deps.audit` of a session triggers a full dependency # compile (`Compiling N files (.ex)` for every dep). This is # EXPECTED — `mix_audit` has `runtime: false` but `mix` still # ensures the dep tree is built. It is slow (tens of seconds) and # noisy, NOT a failure. Always send JSON to a file and only `tail` # stdout for the verdict; never echo the raw compile log. if [ "${lock_file}" = "mix.lock" ]; then # Fast path: scan the project as-is. mix deps.audit --format json 2>/dev/null > "${out_path}" else # Differential path: scan against a non-default lock. We MUST NOT # mutate the project's mix.lock (Iron Law #2). Strategy: copy the # project to a tmpdir, swap in the requested lock, run mix # deps.audit, discard. _mix_audit_run_with_lock "${lock_file}" "${out_path}" fi } # Phase 5 — differential CVE pass. # # Run mix_audit against both OLD and NEW mix.lock states to detect # CVEs patched by the update (the actionable narrative — "you were # exposed for N days"). Use tmpdir copy strategy because: # - MIX_LOCKFILE env var is not Mix-supported as of 1.18. # - git worktree mutates .git/ and creates branch state. # - tmpdir copy is the simplest non-mutating option. # # Inputs (env): # OLD_LOCK_FILE — path to OLD mix.lock (e.g., from `git show HEAD:mix.lock`) # NEW_LOCK_FILE — path to NEW mix.lock (default: working mix.lock) # # Outputs (under per-run tmpdir — see references/audit-tmpdir.md): # ${AUDIT_TMPDIR}/cves_old.json # ${AUDIT_TMPDIR}/cves_new.json mix_audit_diff() { local old_lock="${OLD_LOCK_FILE:?OLD_LOCK_FILE required for differential pass}" local new_lock="${NEW_LOCK_FILE:-mix.lock}" : "${AUDIT_TMPDIR:?AUDIT_TMPDIR not set — driver must establish per-run tmpdir first}" LOCK_FILE="${old_lock}" MIX_AUDIT_OUT="${AUDIT_TMPDIR}/cves_old.json" \ mix_audit_run LOCK_FILE="${new_lock}" MIX_AUDIT_OUT="${AUDIT_TMPDIR}/cves_new.json" \ mix_audit_run } # Internal: run mix deps.audit against an arbitrary mix.lock without # mutating the project. Copies project to tmpdir, swaps lock, runs. # # Defense-in-depth for Iron Law #2: unsets MIX_* env vars that could # otherwise redirect Mix back to the real project (MIX_DEPS_PATH, # MIX_LOCKFILE, MIX_PROJECT, etc.). MIX_HOME is preserved so the Hex # cache is reused. _mix_audit_run_with_lock() { local lock_src="$1" out_path="$2" # Fail fast if the requested lock doesn't exist. Silent fall-through # would produce a false-green ("no vulnerabilities") report — the # worst possible failure mode for a security tool. if [ ! -f "${lock_src}" ]; then echo "ERROR: lock file not found: ${lock_src}" >&2 return 2 fi local tmpdir tmpdir="$(mktemp -d -t deps-audit-XXXXXX)" || { echo "ERROR: mktemp failed" >&2 return 2 } # Trap cleanup at function scope — caller may have its own EXIT trap. # shellcheck disable=SC2064 trap "rm -rf '${tmpdir}'" RETURN # Reflink/hardlink-friendly copy. Exclude _build and deps to keep # the copy cheap; mix will resolve deps from Hex cache anyway. cp mix.exs "${tmpdir}/" 2>/dev/null || true [ -f mix.exs.lock ] && cp mix.exs.lock "${tmpdir}/" 2>/dev/null || true [ -d config ] && cp -R config "${tmpdir}/" 2>/dev/null || true # cp -L: dereference symlinks (don't trust a symlinked lock to point # where the caller thinks). Bubble cp failure to the caller. if ! cp -L "${lock_src}" "${tmpdir}/mix.lock"; then echo "ERROR: cannot copy ${lock_src} to tmpdir" >&2 return 2 fi ( cd "${tmpdir}" || exit 1 # Defense-in-depth: scrub MIX_* vars that could redirect Mix to # the real project. Keep MIX_HOME (Hex cache reuse) and PATH. unset MIX_DEPS_PATH MIX_BUILD_PATH MIX_LOCKFILE MIX_PROJECT MIX_ENV # mix deps.audit doesn't require deps to be fetched — it reads the # lock and queries the local GHSA cache. Avoid `mix deps.get` to # keep the run offline-fast and non-mutating. mix deps.audit --format json 2>/dev/null > "${out_path}" ) } # Internal: when mix_audit is skipped because not installed, check if # the diff contains packages with recent EEF CNA CVEs and emit a loud # alert. The 2026-05-13 enaia-main dogfood exposed this gap: a diff # touched `decimal 2.3 → 2.4`, `phoenix 1.8.5 → 1.8.7`, `postgrex 0.22.0 # → 0.22.2` — all three patch real EEF CNA CVEs (32686/32689/32687) — # and the audit emitted the same generic "install mix_audit" hint as # for any other skip, with no signal that this specific diff would have # triggered three CVE matches. # # The corpus list is hand-curated from cna.erlef.org and refreshed in # tandem with smoke fixtures (corpus.d/). Conservative pattern: only # packages with at least one published CVE in the last 12 months. _mix_audit_warn_cve_corpus_overlap() { local diff_json="${AUDIT_TMPDIR:-}/diff.json" [ -f "${diff_json}" ] || return 0 # Subset of EEF CNA-tracked Hex packages with recent CVEs. Keep # narrow — false positives here erode trust faster than misses. local corpus="decimal phoenix postgrex bandit cowlib plug ecto" local hits=() for pkg in ${corpus}; do if jq -e --arg p "${pkg}" '.changed[]? | select(.package == $p)' \ "${diff_json}" >/dev/null 2>&1; then hits+=("${pkg}") fi done [ "${#hits[@]}" -eq 0 ] && return 0 cat <<EOF 🚨 mix_audit is not installed AND this diff touches packages with known recent CVEs on the EEF CNA list: ${hits[*]} The audit cannot confirm whether this update patches or introduces any of those CVEs. Install mix_audit and re-run for coverage: mix archive.install hex mix_audit Canonical Elixir CVE list: https://cna.erlef.org/cves/ EOF } # Internal: warn if the GHSA cache is staler than GHSA_MAX_AGE_HOURS. # The mix_audit Hex package ships its advisory database as part of the # Hex archive; it's refreshed when the user runs `mix deps.audit` # directly, but the cache itself can lag behind cna.erlef.org. _mix_audit_check_ghsa_freshness() { local max_age="${GHSA_MAX_AGE_HOURS:-24}" local cache_dir="" # Locate the GHSA advisory cache. mix_audit stores it under # _build/<env>/lib/mix_audit/priv/advisories/ when installed as a dep, # or under ~/.mix/archives/ when installed as an archive. Try both. for candidate in \ "_build/dev/lib/mix_audit/priv/advisories" \ "_build/test/lib/mix_audit/priv/advisories" \ "$HOME/.mix/archives/mix_audit"*; do if [ -d "${candidate}" ]; then cache_dir="${candidate}" break fi done [ -z "${cache_dir}" ] && return 0 # Can't locate cache — skip warn. local mtime now age_hours # macOS stat -f %m, GNU stat -c %Y. Try both. mtime="$(stat -f %m "${cache_dir}" 2>/dev/null || stat -c %Y "${cache_dir}" 2>/dev/null)" [ -z "${mtime}" ] && return 0 now="$(date +%s)" age_hours=$(( (now - mtime) / 3600 )) if [ "${age_hours}" -gt "${max_age}" ]; then cat <<EOF WARN: GHSA advisory cache is ${age_hours} hours old (>${max_age}h threshold). Recent disclosures (last 7 days) may be missing. Refresh with: mix deps.update mix_audit # if installed as dep mix archive.install hex mix_audit --force # if installed as archive Canonical Elixir CVE list: https://cna.erlef.org/cves/ EOF return 1 fi return 0 } ``` Output is JSON: ```json { "pass": false, "vulnerabilities": [ { "advisory": {"id": "GHSA-xxxx-xxxx-xxxx", "cve": "CVE-2026-12345", "title": "...", "description": "...", "severity": "high", "patched_versions": "~> 1.2.3"}, "dependency": {"package": "...", "version": "..."} } ] } ``` Severity mapping: `critical` / `high` → BLOCK · `moderate` → WARN · `low` → INFO. **FP rate:** ~0% (advisory DB is curated). Coverage gap: GHSA has fewer Hex entries than npm/RustSec, so absence is not proof of safety. ## 2a. GHSA cache freshness (Phase 5) `mix_audit` ships its GHSA advisory database as part of the Hex package or archive. The advisory DB is **not** refreshed automatically on each `mix deps.audit` invocation — it updates only when the user explicitly runs `mix deps.update mix_audit` or `mix archive.install hex mix_audit --force`. This matters because the EEF CNA disclosure cadence has accelerated: 14 Hex package CVEs were published in April-May 2026 alone (e.g., `postgrex 0.22.0 → 0.22.1` patched a SQL-injection CVE disclosed **the same day** the 2026-05-12 virgil dogfood ran). A 1-week-old advisory cache will miss these. **Behavior:** `mix_audit_run` checks the mtime of the GHSA advisory directory before running. If older than `GHSA_MAX_AGE_HOURS` (default 24), it emits a WARN to stderr with refresh instructions and a link to the canonical EEF CNA list (<https://cna.erlef.org/cves/>). **How to refresh manually:** ```bash # If installed as a dep in mix.exs: mix deps.update mix_audit # If installed as a Mix archive: mix archive.install hex mix_audit --force # Verify freshness: ls -la _build/dev/lib/mix_audit/priv/advisories/ | head ``` **Override the threshold:** ```bash GHSA_MAX_AGE_HOURS=72 /phx:deps-audit # tolerate 3-day-old cache GHSA_MAX_AGE_HOURS=1 /phx:deps-audit # paranoid mode ``` **Future (Phase 6+):** auto-refresh via a PreToolUse hook that watches for `mix deps.audit` invocations and triggers `mix deps.update mix_audit` if the cache is >24h old. Deferred until we can prove the refresh itself is non-flaky (Hex registry can rate-limit or 503). **Why not a `mix deps.update`?** Mutating `mix.lock` violates Iron Law #2. The wrapper warns and continues — the user must refresh explicitly. This is consent-resistant by the same logic as the mix_audit-install refusal above. ## 3. `osv-scanner` — CVE check via OSV.dev Standalone Go binary (`go install github.com/google/osv-scanner@latest`). v2.3.5+ supports Elixir/Hex. ```bash osv_scan() { if ! command -v osv-scanner >/dev/null 2>&1; then cat >&2 <<'EOF' WARN: osv-scanner not installed — skipping CVE check via OSV.dev. To enable: go install github.com/google/osv-scanner@latest # or: brew install osv-scanner EOF return 0 fi osv-scanner \ --lockfile mix.lock \ --format json \ > "${AUDIT_TMPDIR}/osv-scan.json" 2>/dev/null || true } ``` Output is JSON with `results[].packages[].vulnerabilities[]`. Each vulnerability has `id` (OSV ID), `aliases` (CVE list), `severity` (array of CVSS strings). Severity mapping: parse highest CVSS score from `severity[].score`: - ≥ 9.0 → BLOCK (critical) - ≥ 7.0 → BLOCK (high) - ≥ 4.0 → WARN (medium) - < 4.0 → INFO (low) **FP rate:** ~0%. **Why integrate both `mix_audit` and `osv-scanner`?** GHSA and OSV.dev have non-overlapping coverage. Running both catches more real-world CVEs. ## Aggregation into findings format Each external-tool finding maps to the same shape as MVP-rule findings, with `rule_id = "ext:<tool>"`: ```elixir %{ rule_id: "ext:hex-audit" | "ext:mix-audit" | "ext:osv-scanner", severity: :block | :warn | :info, file: nil, # CVEs are package-level, not file-level line: nil, snippet: "<advisory-id>", message: "<title or description>" } ``` Attach to the per-package finding list before scoring. ## Parallelism All three tools run independently. Spawn in background: ```bash hex_audit & pid_hex=$! mix_audit_run & pid_mix=$! osv_scan & pid_osv=$! wait $pid_hex $pid_mix $pid_osv ``` `mix hex.audit` and `mix deps.audit` may contend on the `mix` lock; if so, serialize the two mix-based scanners and only parallelize `osv-scanner`. ## Exit-code handling | Tool | 0 | Non-zero | |------|---|----------| | `mix hex.audit` | No retirements | Retirements found (informational, NOT fatal) | | `mix deps.audit` | No CVEs | CVEs found | | `osv-scanner` | No CVEs | CVEs found OR scan error | Treat non-zero as "findings to parse", not "skill failure". The skill itself returns 0 unless the *audit infrastructure* fails (missing `mix`, bad network, corrupt cache). ## Why not Snyk / Phylum / Endor? | Tool | Why skipped | |------|-------------| | Snyk CLI | Paid for org use, signal duplicates `mix_audit` for free | | Phylum | Thin Hex support (per 2026 research) | | Endor Labs | No reliable BEAM reachability — not credible | | Semgrep SC | Paid tier; OSS Semgrep covered separately in Phase 2 | | Socket.dev | No Hex support; we **reimplement** their signal model | ## Future: SARIF output (Phase 2) `osv-scanner --format sarif` and `mix deps.audit --format sarif` (proposed) would let us emit a single SARIF file for GitHub Code Scanning. Deferred to Phase 2 alongside Semgrep ruleset. -
heuristics.md 8.3 KB
# Hex Supply-Chain Heuristics — Full Catalogue Full 35-rule catalogue. **Rules marked ✅ MVP are implemented in Phase 1.** Remaining rules are deferred to Phase 2. Severity scale: - **BLOCK** — high-confidence malicious indicator. Refuse to "pass" the audit. - **WARN** — suspicious but plausibly legitimate. Surface for human review. - **INFO** — context only, never alone. Scoring weights: BLOCK = 10 · WARN = 3 · INFO = 1. ## Category 1 — Compile-Time Code Execution (7 rules) | # | Rule | Sev | Method | FP | MVP | |---|------|-----|--------|----|-----| | 1.1 | Top-level expressions outside `def`/`defp`/`defmacro` body | WARN | AST module-body walk | ~5% | — | | 1.2 | `Code.eval_string` / `Code.eval_quoted` with non-literal arg | BLOCK | AST | ~1% | ✅ Rule 2 | | 1.3 | `:erlang.apply(Mod, Fun, Args)` with non-literal MFA at module scope | BLOCK | AST | ~1% | ✅ Rule 2 | | 1.4 | `System.cmd` / `:os.cmd` / `Port.open` at compile time | BLOCK | AST scope check | ~3% | ✅ Rule 3 | | 1.5 | `@on_load` / `@on_definition` callbacks | WARN | Attribute scan | ~10% | — | | 1.6 | Mix alias wrapping `deps.get` / `deps.update` | WARN | Parse `mix.exs` aliases | ~5% | — | | 1.7 | Excessive macro density (>20% of file is `defmacro`) | INFO | AST node counting | ~15% | — | ## Category 2 — Network Egress During Compile (5 rules) | # | Rule | Sev | Method | FP | MVP | |---|------|-----|--------|----|-----| | 2.1 | `:httpc.request` / `HTTPoison.*` / `Req.*` / `Tesla.*` at module load | BLOCK | AST scope check | ~2% | — (covered partially by 1.4 if shelling out) | | 2.2 | `:gen_tcp.connect` / `:ssl.connect` outside function bodies | BLOCK | AST | ~1% | — | | 2.3 | NIF socket calls (`:inet.*`) at compile time | BLOCK | AST | ~1% | — | | 2.4 | Webhook-style URLs (Discord/Slack/IPFS) in source | WARN | Regex on `.ex`/`.exs` | ~10% | — | | 2.5 | DNS exfil patterns (long subdomains, base64 in host) | WARN | Regex | ~12% | — | ## Category 3 — Obfuscation / Hidden Payloads (6 rules) | # | Rule | Sev | Method | FP | MVP | |---|------|-----|--------|----|-----| | 3.1 | Base64 string literals >256 chars outside `priv/static/`, `test/fixtures/`, `assets/` | WARN | Regex with directory exclude | ~8% | ✅ Rule 7 | | 3.2 | Hex binaries >128 bytes (`<<0x..., ...>>` literal) | WARN | AST literal scan | ~6% | — | | 3.3 | `:erlang.binary_to_term/1` on literal without `:safe` opt | BLOCK | AST | ~2% | ✅ Rule 4 | | 3.4 | Unicode homoglyph identifiers (Cyrillic/Greek lookalikes in `def` names) | WARN | Char-class regex | ~3% | — | | 3.5 | Bidi control characters in source (CVE-2021-42574 Trojan Source) | BLOCK | Grep `[--]` | ~0% | ✅ Rule 1 | | 3.6 | String concat to hide reserved words (`"Sy" <> "stem" <> ".cmd"`) | WARN | AST concat-of-string-literals scan | ~5% | — | ## Category 4 — NIF / Port Driver Red Flags (4 rules) | # | Rule | Sev | Method | FP | MVP | |---|------|-----|--------|----|-----| | 4.1 | Native source presence (`c_src/`, `native/`, `.c`, `.rs`, `.zig`) | INFO | `find` | ~30% | — | | 4.2 | Build scripts invoking `curl` / `wget` / `sh -c` | BLOCK | Grep in `Makefile`, `*.sh`, `build.rs` | ~3% | — | | 4.3 | Precompiled binaries in `priv/` (any non-text file > 10KB) | WARN | `file -i` MIME check | ~15% | — | | 4.4 | `rustler` / `zigler` / `:erlang.load_nif` newly added | WARN | AST diff old vs new `mix.exs` | ~10% | — | ## Category 5 — Typosquatting & Impersonation (5 rules) | # | Rule | Sev | Method | FP | MVP | |---|------|-----|--------|----|-----| | 5.1 | Levenshtein ≤ 2 from top-500 Hex packages + download delta >1000× | BLOCK | Hex API + fuzzy distance | ~1% | ✅ Rule 8 | | 5.2 | Homoglyph package names (`phoeniх` with Cyrillic х) | BLOCK | Unicode confusable check | ~0.5% | — | | 5.3 | Identical package description but different author | WARN | Hex API description diff | ~5% | — | | 5.4 | New author + sudden traction (>1k DLs in week 1) | WARN | Hex API + insertedat math | ~8% | — | | 5.5 | Cross-ecosystem name reuse (npm/PyPI package with same name + different author) | INFO | npm/PyPI registry lookup | ~20% | — | ## Category 6 — Maintainer / Publishing Signals (5 rules) | # | Rule | Sev | Method | FP | MVP | |---|------|-----|--------|----|-----| | 6.1 | Maintainer change between versions | BLOCK | Hex API `owners` diff | ~2% | ✅ Rule 6 | | 6.2 | Single-maintainer package with >500 transitive dependents | INFO | Hex API + reverse deps | ~0% | — | | 6.3 | Anomalous version bump (major skip, non-semver) | WARN | Version string parsing | ~5% | — | | 6.4 | `days_since_publish < 7` ("let it cook" rule) | WARN | Hex API `inserted_at` | ~30% | — | | 6.5 | Yank-then-republish at same version | BLOCK | Hex API `retirements` history | ~0.5% | — | ## Category 7 — Diff-Based Checks (6 rules) | # | Rule | Sev | Method | FP | MVP | |---|------|-----|--------|----|-----| | 7.1 | New top-level expressions vs old version | WARN | AST diff of module bodies | ~7% | — | | 7.2 | New `System.cmd` / network calls vs old | BLOCK | AST diff | ~4% | — | | 7.3 | New files in `priv/` (esp. binary) | WARN | `diff` of file lists | ~10% | — | | 7.4 | New `:git` / `:path` dep in `mix.exs` | BLOCK | AST diff of `deps/0` | ~5% | ✅ Rule 5 | | 7.5 | License change between versions | WARN | Read `LICENSE` / `mix.exs` `:licenses` | ~5% | — | | 7.6 | `.hex` / `metadata.config` tampering (checksum mismatch) | BLOCK | `mix hex.audit` already covers retirement; this extends | ~1% | — | ## Category 8 — Mix-Specific (4 rules) | # | Rule | Sev | Method | FP | MVP | |---|------|-----|--------|----|-----| | 8.1 | `:git` / `:path` dep in `mix.exs` (any, not just new) | INFO | Parse `mix.exs` | ~40% (legit local dev) | — | | 8.2 | Custom compilers (`:compilers` modified) | WARN | Parse `project/0` | ~10% | — | | 8.3 | Aliases wrapping `deps.get` / `deps.update` | WARN | Parse aliases | ~5% | — | | 8.4 | `override: true` on transitive deps | INFO | Parse `deps/0` | ~20% | — | ## MVP Summary — 8 Rules The 8 MVP rules (Phase 1) target the highest-confidence indicators with aggregate FP <5%: | MVP # | From | Rule | |-------|------|------| | 1 | 3.5 | Bidi Unicode control chars | | 2 | 1.2 + 1.3 | `Code.eval_*` / `:erlang.apply` non-literal at module scope | | 3 | 1.4 | `System.cmd` / `:os.cmd` / `Port.open` at compile time | | 4 | 3.3 | `:erlang.binary_to_term/1` literal without `:safe` | | 5 | 7.4 | New `:git`/`:path` dep vs old `mix.exs` | | 6 | 6.1 | Maintainer change between versions | | 7 | 3.1 | Base64 >256 chars outside allowlisted dirs | | 8 | 5.1 | Typosquat: Levenshtein ≤ 2 + download delta >1000× | Covers Trojan Source, compile-time RCE, BEAM deser, dep confusion, account takeover, and typosquatting. Each finding shape: ```elixir %{ rule_id: 1..8, severity: :block | :warn | :info, file: "lib/foo.ex" | nil, line: integer | nil, snippet: "...", message: "..." } ``` ## Prior Art - **Trojan Source / CVE-2021-42574** — bidi control chars in source. Rule 1. - **diff-CodeQL (Froh et al., SCORED '23)** — 41 CodeQL queries on npm diffs, 1.4% FP. Rules 7.1–7.4 port the diff-of-findings pattern. - **Socket.dev signal model** — multi-signal scoring. Our weighted sum is a simplified version. - **`cargo-vet` audit ledger** — Phase 2 `hex_vet.exs` ledger derives from it. - **`npq` pre-flight pattern** — score before download. Mode A (`--preview`) replicates this. - **CVE-2026-21619** — `hex_core` unsafe `binary_to_term` (RCE). Motivates Rule 4. ## Tuning Notes - Rule 7 (base64) dominates noise. If FP rate exceeds 10% in real use, raise threshold to 512 chars or add more exclude directories. - Rule 8 (typosquat) requires daily refresh of top-500 list. Cache in `${AUDIT_TMPDIR}/hex-api/top-500.json` with 24h TTL. - Rule 6 (maintainer change) needs cross-referencing release timestamps with owner-list timestamps. Hex API `inserted_at` per release helps. ## Out of Scope (Phase 2) - Differential rules 7.1–7.6 (full old vs new diff per file) — currently only 7.4 (`mix.exs` deps) is in MVP because it's the cheapest diff. - All NIF rules (Category 4) — requires `file -i` MIME detection, deferred. - Semgrep ruleset (Category 1–2 expanded) — adds binary dependency. - YARA byte-pattern scan — adds binary dependency. - LLM triage on high-score findings — Phase 2 once FP rate proven. -
hex-api.md 7.5 KB
# Hex API Enrichment Adds metadata (owners, downloads, publish dates) to each audited package. Drives Rule 6 (maintainer change) and Rule 8 (typosquat download delta). ## Endpoints | Call | Endpoint | Why | |------|----------|-----| | Package metadata | `GET https://hex.pm/api/packages/:name` | Owners, downloads, inserted_at | | Per-release metadata | `GET https://hex.pm/api/packages/:name/releases/:version` | Per-release publisher, inserted_at | | Top-500 by downloads | `GET https://hex.pm/api/packages?sort=downloads&page=1..7` | Typosquat denominator (Rule 8) | All require `Accept: application/vnd.hex+json` header. ## Computed signals per package ```elixir %{ pkg: "phoenix", owners: ["chrismccord", "josevalim", "..."], owner_age_days: 3650, # min(inserted_at across owners) downloads_all: 50_000_000, downloads_recent: 800_000, # weekly inserted_at: "2014-04-17T...", days_since_publish_latest: 14, download_velocity: 800_000 / 7, # downloads/day, recent release_publisher: "josevalim" # at the version being audited } ``` ## Rate limit **5 req/sec.** Hex API doesn't publish official limits but the Hex.pm team has stated this is the polite ceiling. **Use `python3` + `urllib`, not curl.** A `curl "...${var}"` wrapper is fragile to whitespace/newline contamination in the interpolated path (dogfood: `curl: (3) Malformed input to a URL function`). `python3` is already a hard dependency of the audit and is robust here. Canonical: ```bash hex_api_get() { # usage: hex_api_get <api-path> e.g. packages/phoenix AUDIT_TMPDIR="$(cat "${TMPDIR:-/tmp}/phx-audit-dir.txt")" python3 - "$1" "$AUDIT_TMPDIR" <<'PY' import sys, os, json, time, urllib.request, pathlib path, tmp = sys.argv[1].strip(), sys.argv[2] cache = pathlib.Path(tmp, "hex-api", path.replace("/", "_") + ".json") ttl = 7 * 86400 if cache.is_file() and time.time() - cache.stat().st_mtime < ttl: print(cache.read_text()); sys.exit(0) cache.parent.mkdir(parents=True, exist_ok=True) req = urllib.request.Request( "https://hex.pm/api/" + path, headers={"Accept": "application/vnd.hex+json", "User-Agent": "phx-deps-audit/0.1 (+claude-elixir-phoenix)"}) with urllib.request.urlopen(req, timeout=20) as r: body = r.read().decode() cache.write_text(body) time.sleep(0.2) # 5 req/sec ceiling print(body) PY } ``` The `sleep 0.2` blocks the calling process; keep Hex calls serial (parallelism 1) or the effective rate becomes `parallelism / 0.2`. > **curl fallback** (only if `python3` is unavailable, which the audit > otherwise assumes): `curl -fsSL -H "Accept: application/vnd.hex+json" > "https://hex.pm/api/${path}"` — but trim the path first > (`path="${path//[$'\n\r\t ']/}"`) to avoid the malformed-URL failure. ## Cache TTL | Resource | TTL | Justification | |----------|-----|---------------| | Package metadata | 7 days | Owners change rarely; we want fresh-ish data | | Per-release metadata | 30 days | Immutable once published | | Top-500 list | 24 hours | Daily refresh is industry standard | Override with `--no-cache` for debugging. ## Top-500 list — typosquat denominator ```bash fetch_top_500() { AUDIT_TMPDIR="$(cat "${TMPDIR:-/tmp}/phx-audit-dir.txt")" python3 - "$AUDIT_TMPDIR" <<'PY' import sys, time, json, urllib.request, pathlib cache = pathlib.Path(sys.argv[1], "hex-api", "top-500.json") if cache.is_file() and time.time() - cache.stat().st_mtime < 24 * 3600: sys.exit(0) cache.parent.mkdir(parents=True, exist_ok=True) out = [] for page in range(1, 8): req = urllib.request.Request( f"https://hex.pm/api/packages?sort=downloads&page={page}", headers={"Accept": "application/vnd.hex+json"}) with urllib.request.urlopen(req, timeout=20) as r: out += json.loads(r.read().decode()) time.sleep(0.5) cache.write_text(json.dumps(out)) PY } ``` > curl fallback: same `for page in 1..7; curl … | jq -s 'add'` loop as > before — use only if `python3` is missing. 7 pages × 100 packages/page = 700 entries; we use the first 500 (the tail has thin signal for typosquat denominator anyway). Total fetch: ~3 seconds on cold cache, free on warm. ## Computing signals ```bash package_signals() { local pkg="$1" local data data=$(hex_api_get "packages/${pkg}") jq -r --arg today "$(date -u +%Y-%m-%dT%H:%M:%SZ)" ' { pkg: .name, owners: [.owners[]?.username], downloads_all: .downloads.all, downloads_recent: .downloads.recent, inserted_at: .inserted_at, latest_version: .releases[0].version, latest_version_inserted_at: .releases[0].inserted_at } | tojson ' <<<"${data}" } ``` `owner_age_days` requires a second call per owner (to `/users/:username`) — deferred to Phase 2 to keep rate-limit pressure low. For Phase 1, we treat "owner list" as a static set and only diff between versions (Rule 6). ## Rule 6 — Maintainer change detector ```bash maintainer_change() { local pkg="$1" old_ver="$2" new_ver="$3" local old_pub new_pub old_pub=$(hex_api_get "packages/${pkg}/releases/${old_ver}" \ | jq -r '.publisher.username // empty') new_pub=$(hex_api_get "packages/${pkg}/releases/${new_ver}" \ | jq -r '.publisher.username // empty') if [ -n "${old_pub}" ] && [ -n "${new_pub}" ] && [ "${old_pub}" != "${new_pub}" ]; then echo "BLOCK|6|maintainer changed: ${old_pub} → ${new_pub}" fi } ``` Note: `publisher` is the user who *published the release*. `owners` is the package-level list. The two CAN diverge (publisher is a delegate of an owner). For Phase 1 we flag *publisher* change at release boundary. ## Rule 8 — Typosquat detector ```bash typosquat_check() { local pkg="$1" # Pre-fetched top-500 list fetch_top_500 jq -r --arg pkg "${pkg}" ' .[] | select(.name != $pkg) | [.name, .downloads.all] | @tsv ' ${AUDIT_TMPDIR}/hex-api/top-500.json \ | while IFS=$'\t' read -r candidate dl_count; do local dist dist=$(levenshtein "${pkg}" "${candidate}") if [ "${dist}" -le 2 ]; then local target_dl target_dl=$(hex_api_get "packages/${pkg}" | jq -r '.downloads.all // 0') if [ "${dl_count}" -gt $((target_dl * 1000)) ]; then echo "BLOCK|8|typosquat candidate: '${pkg}' (${target_dl} DLs) vs '${candidate}' (${dl_count} DLs, distance ${dist})" fi fi done } ``` Levenshtein helper (inline awk impl or `string_distance` Hex pkg if available): ```bash levenshtein() { awk -v a="$1" -v b="$2" 'BEGIN { la = length(a); lb = length(b) if (la == 0) { print lb; exit } if (lb == 0) { print la; exit } for (i = 0; i <= la; i++) d[i,0] = i for (j = 0; j <= lb; j++) d[0,j] = j for (i = 1; i <= la; i++) for (j = 1; j <= lb; j++) { c = (substr(a,i,1) == substr(b,j,1)) ? 0 : 1 v = d[i-1,j] + 1 h = d[i,j-1] + 1 diag = d[i-1,j-1] + c d[i,j] = (v < h ? (v < diag ? v : diag) : (h < diag ? h : diag)) } print d[la,lb] }' } ``` ## Failure modes | Failure | Behavior | |---------|----------| | Rate-limit response (429) | Sleep 5s, retry once. Second 429 → skip API enrichment for this pkg with WARN | | 404 on package | Mark `unknown` (likely renamed or removed); skip API rules | | Network timeout | Use stale cache if available; else skip with WARN | | Malformed JSON | Log + skip (Hex API is stable; this means corruption) | | Cache disk full | Bypass cache, fall back to direct call | ## Why no `mix hex.search` / `mix hex.info`? Both are interactive shells under the hood — slow to parse, not designed for machine output. Direct HTTP is 5-10× faster. -
hook.md 5.4 KB
# PreToolUse `deps-audit-gate.sh` — tiered fast-path Phase 3 wires the audit into the hot loop: every `mix deps.get`, `mix deps.update`, and `mix deps.compile` runs through a tiered hook before mix actually executes. The hook must stay invisible on the common path (sub-second), block on real risk, and never lock users out of their own machine. ## Why tiered A non-tiered hook would either: - **Run the full pipeline** (30-90s) on every `mix deps.get` — unusable during feature work where `deps.get` runs hourly. - **Cache the full audit** — but cache-fills are still 30-90s, and the first `mix deps.get` after a `mix.lock` change is the highest-anxiety moment. Three tiers split the workload so each `mix deps.*` invocation pays only for what changed: | Tier | Budget | Work | Exit | |------|--------|------|------| | 0 | <200ms | lock-SHA cache lookup | 0 silent | | 1 | <2s | bidi grep + `:git`/`:path` diff | 0 with hint, or 2 to block | | 2 | unbounded | full Phase 2 pipeline (only `:full` mode) | 0 or 2 | ## Tier 0: cache hit Read `.claude/deps-audit/last-run.json`. If the lock-SHA matches the cached run AND `audit_passed: true` AND the policy mode is unchanged, exit 0 immediately. Cost: one `shasum` + one `jq` per dep operation. The policy-mode equality check is necessary because a user who flipped from `false` to `:strict` between runs must not get a stale "passed" from the warn-only era. ## Tier 1: deterministic fast rules Only two Phase 1 rules run inline: - **Rule 1 — bidi/RLO chars in `mix.lock`.** One `perl` scan over a small file. Either present or not. - **Rule 5 — new `:git` or `:path` deps in `mix.exs`.** Git-diffs against `${PHX_DEPS_AUDIT_BASE:-origin/main}` to isolate ADDED non-Hex deps. Re-locks of existing `:git` deps are ignored. Both rules are **zero false positive** by design and produce stable NDJSON findings compatible with the Phase 2 differ output. If both rules return clean, the gate prints a one-line stderr hint ("`Run /phx:deps-audit for full pipeline`") and exits 0. Rules 2, 3, 4, 7, 8 are intentionally deferred to Tier 2 — they require unpacking tarballs and lose the <2s budget on the very first new package. ## Tier 2: full pipeline (opt-in only) Tier 2 invocation is NOT chained from the hook. When `block_on_unvetted` is `:full`, the hook still only runs Tiers 0+1, blocks on Tier 1 findings, and points the user at `/phx:deps-audit` for the full Tier 2 pipeline. Reason: hook-budget exhaustion via the Bash 600s timeout is worse UX than an explicit "run the audit" message. The `/phx:deps-audit` skill body owns Tier 2 — when invoked manually, it can take 30-90s, run subagents, and update `last-run.json` so the next hook invocation gets a Tier 0 hit. ## Policy enforcement The hook reads `policy.block_on_unvetted` from `hex_vet.exs` (see `deps-vet/references/hex-vet.md` for the tri-mode schema) and gates findings accordingly: | Mode | On Tier 1 findings | |------|--------------------| | `false` | Print summary, exit 0 (warn-only) | | `:new_only` (default) | Block ONLY on rule 1 (bidi) and rule 5 (new dep). Both ARE "new" signals by definition. | | `:strict` | Block on any Tier 1 finding | | `:full` | Same as `:strict` for hook scope; full pipeline runs via skill body | Override: `PHX_SKIP_DEPS_AUDIT=1 mix deps.get` bypasses the gate entirely. The escape hatch is critical for emergency unblocks and CI environments where the audit runs separately. ## Latency budget rationale - **Tier 0 cache hit**: dominant case during feature work where `mix.lock` doesn't change between calls. Target: <200ms p95 so the hook is invisible to the developer. - **Tier 1**: target <2s p95. Bidi grep on `mix.lock` is ~50ms. Rule 5's `git show` + `comm` against a 100-line `mix.exs` is ~200ms. Slowest path is git-fetch when `PHX_DEPS_AUDIT_BASE` is stale; users should pre-fetch in pre-commit or CI. - **Tier 2**: explicitly unbounded; only reached on user opt-in via `/phx:deps-audit`. ## Failure modes - **No `mix.lock`** → exit 0 (no deps to audit yet; first `deps.get`). - **No `hex_vet.exs`** → mode defaults to `false` (warn-only). Hook still runs Tiers 0+1, prints findings as warnings, never blocks. The user gets visibility without the friction. - **Corrupt `last-run.json`** → Tier 0 returns false, Tier 1 runs. Worst case: 2s of work, then a fresh `last-run.json` overwrites. - **`PHX_DEPS_AUDIT_BASE` unreachable** → Rule 5 treats whole `mix.exs` as new (false-positive favored over silent-pass). Set the env to a local ref (`HEAD~1`) for offline work. ## Hook registration In `hooks/hooks.json`, the gate runs under `PreToolUse → Bash` alongside `block-dangerous-ops.sh`. The `if: "Bash(*mix deps.*)"` condition keeps the script silent on non-deps commands — the script itself also re-checks the command, defense-in-depth. ```json { "type": "command", "if": "Bash(*mix deps.*)", "command": "${CLAUDE_PLUGIN_ROOT}/hooks/scripts/deps-audit-gate.sh", "timeout": 5, "statusMessage": "Tiered deps audit..." } ``` The 5-second timeout caps Tier 1 hard. If the hook itself runs over, the exit code propagates as "hook failed" — the user sees the timeout, not a stuck `mix` invocation. ## CI usage In CI, set `PHX_SKIP_DEPS_AUDIT=1` for `mix deps.get` steps and run `mix phx.deps_audit --ci` as a separate job — CI wants determinism and full Tier 2 output, not the fast-path hook. See `ci-integration.md` for sample workflows. -
llm-triage.md 5.8 KB
# LLM triage — two-tier with context-supervisor Phase 2 adds optional LLM-assisted triage on top of the deterministic rule layer. The deterministic rules carry the security load; the LLM compresses evidence and surfaces likely-FPs vs likely-TPs to the human reviewer. **No finding is auto-suppressed without human review.** ## Iron Laws 1. **TRIAGE IS ADVISORY.** LLM verdicts adjust the *renderer's ordering and visual emphasis*, never the underlying severity in the NDJSON stream. Audit findings remain the source of truth. 2. **THRESHOLD-GATED.** Triage runs only when a package's weighted score exceeds 10 (BLOCK=10, WARN=3, INFO=1). Below threshold, the deterministic output is the final word. 3. **TWO-TIER MANDATORY.** Per-package triagers write JSON files; a `context-supervisor` consolidates. Main skill reads ONLY the consolidated file. Reading per-package outputs directly in the main context blows the budget on 5+ packages. 4. **NO INVENTED FINDINGS.** Each verdict maps 1:1 to an input finding. Post-call validator drops verdicts whose `rule_id + file + line` triple isn't in the input set. ## Architecture ``` deps-audit body │ ▼ score > 10 for package P? │ ├─► YES → spawn hex-deps-triager (sonnet) for P │ ↓ writes triage/<pkg>-<ts>.json │ ▼ (wait for all triagers) context-supervisor │ ▼ reads triage/<pkg>-*.json files │ ▼ writes triage/consolidated.md │ deps-audit body reads consolidated.md ``` ## Two-tier rationale A 20-package audit run produces ~3-5KB of findings.jsonl. Each triager fetches 3-line diff windows × N findings — easily 50KB context per package. Running 10 triagers in parallel and then synthesizing in the main context exhausts the budget on synthesize. Splitting the work: per-package triager has its own context, writes a small structured verdict (1-3KB). The supervisor's input is N × 3KB ≈ 30KB even at 10 packages — well within the supervisor's window. Main skill reads a single consolidated.md (<5KB). ## Triager spawn pattern The skill body builds one input JSON per package, then spawns triagers in parallel via the Agent tool: ```bash # Pseudocode — the skill's actual body uses Task() blocks for pkg in $(jq -r 'keys[]' high_score.json); do build_triager_input "${pkg}" > "/tmp/triage-in-${pkg}.json" spawn_task --subagent_type hex-deps-triager \ --prompt "Triage findings for ${pkg}. Input: /tmp/triage-in-${pkg}.json Output: .claude/deps-audit/triage/${pkg}-$(date +%s).json" done wait ``` After all triagers complete, the skill spawns `context-supervisor`: ```bash spawn_task --subagent_type context-supervisor \ --prompt "Consolidate triage verdicts. Inputs: .claude/deps-audit/triage/*.json Output: .claude/deps-audit/triage/consolidated.md Group by verdict (likely_malicious, needs_human, likely_benign). Surface model+confidence per verdict. Sum across packages." ``` ## Input JSON contract See `hex-deps-triager.md` agent. Each input contains: - `package`, `version`, `previous_version` - `tarball_dir`, `previous_tarball_dir` (absolute paths) - `findings[]` — list with `rule_id`, `file`, `line`, `snippet`, `message`, `diff_window` (3 lines context) - `hex_metadata` — downloads, owners, release_publisher, inserted_at Build via `mix run --no-mix-exs -e` over the existing findings.jsonl plus the hex-api cache. Keep diff_window ≤200 chars per finding to control input size. ## Output JSON contract ```json { "package": "<pkg>", "version": "<version>", "model": "sonnet", "summary": "<one-line per-package summary>", "verdicts": [ { "rule_id": <int>, "file": "<file>", "line": <int>, "confidence": <0.0..1.0>, "verdict": "likely_benign | needs_human | likely_malicious", "rationale": "<one paragraph>", "fp_reasons": ["<reason1>", "<reason2>"] } ] } ``` Validation pass after each triager: 1. Verify the file parses as JSON. 2. Verify every `verdicts[].{rule_id, file, line}` triple appears in the input `findings[]`. Mismatches are dropped with a warning. 3. Verify `verdict` and `confidence` ranges. Out-of-range verdicts → `needs_human` with `confidence: 0.0`. ## Consolidated output The `context-supervisor` produces a markdown file shaped like: ```markdown # Triage Consolidated — N packages, M findings ## Summary - Likely malicious: 1 (1 finding) - Needs human review: 3 (5 findings) - Likely benign: 7 (12 findings) ## Likely malicious (BLOCK by default) ### phoenix_extras 0.2.0 → 0.3.0 (score 22, sonnet conf 0.92) - rule 3 (compile-time exec) at lib/init.ex:14 — System.cmd to unknown URL inside __before_compile__. New maintainer + 1mo old. ## Needs human review ... ## Likely benign (consider INFO downgrade) ... ``` The deps-audit renderer reads this and slots the sections above the per-package findings tables. ## Token budget | Component | Budget | Actual (10 pkg audit) | |-----------|--------|----------------------| | Per-triager input | ≤15KB | ~5-10KB | | Per-triager output | ≤3KB | ~1-2KB | | Consolidator input | 10 × 3KB | ~25KB | | Consolidator output | ≤5KB | ~3KB | | Main skill triage read | ≤5KB | ~3KB | Compared to the naive approach (main reads 10 × 10KB = 100KB), the two-tier flow saves ~95% of the post-triage context burn. ## When NOT to invoke LLM triage - Package score ≤ 10 → deterministic output suffices. - Mode A (`--preview`) on a single package not yet locked → user is exploring; the audit table is enough. - `--no-llm` flag passed by the user. - No `ANTHROPIC_API_KEY` available in the environment — fall back cleanly to deterministic output with a warning. -
operating-modes.md 5.1 KB
# Operating Modes — A / B / C The audit engine is **mode-agnostic**: ``` audit(pkg, old_version, new_version) → findings[] ``` Modes only change how `(old, new)` pairs get resolved. The engine never inspects how the diff came to be. ## Mode B — Default (working vs HEAD) ``` /phx:deps-audit ``` Compares the working-tree `mix.lock` against `git show HEAD:mix.lock`. | | | |---|---| | **Old source** | `git show HEAD:mix.lock` | | **New source** | working `mix.lock` | | **Use case** | Post-`mix deps.update` safety net. Pre-commit gate. | | **Cost** | Cheapest mode — no remote calls for diff resolution | | **Why default** | Catches Igniter / auto-update footguns retroactively, before commit | If HEAD has no `mix.lock` (initial commit, brand new dep), every locked package is treated as a NEW package (`old_version = nil`). Diff-only rules (Rule 5, Rule 6) are skipped for new packages; static rules (1, 2, 3, 4, 7) still run. ## Mode C — PR / branch comparison ``` /phx:deps-audit --base main /phx:deps-audit --base origin/main /phx:deps-audit --base abc1234 ``` Compares the working-tree `mix.lock` against `git show <ref>:mix.lock`. | | | |---|---| | **Old source** | `git show <ref>:mix.lock` | | **New source** | working `mix.lock` | | **Use case** | CI on PRs that touch `mix.lock`. Pre-PR self-review. | | **Cost** | Same as Mode B | | **CI hint** | `/phx:deps-audit --base origin/main --json` for machine-readable output | `<ref>` is any valid Git revision: branch name, tag, commit SHA, `HEAD~N`. ## Mode A — Preview (locked vs Hex latest) ``` /phx:deps-audit --preview # all locked deps vs latest /phx:deps-audit --preview httpoison # one package /phx:deps-audit --preview httpoison req # multiple ``` Compares the locked version against the latest version on Hex.pm. | | | |---|---| | **Old source** | Version in working `mix.lock` | | **New source** | Latest version from Hex API (`GET /api/packages/:name`) | | **Use case** | "If I run `mix deps.update X`, what lands?" Pre-update analysis. | | **Cost** | Adds 1 Hex API call per package (cached 1h) | | **Limitation** | Does not resolve transitive deps. Only the named packages. | If no packages are specified, all packages from `mix.lock` are previewed. Cap at 50 packages to avoid Hex API spam — show a warning and stop if exceeded. ## Resolver pseudocode ```elixir def resolve_pairs(mode) do case mode do :working_vs_head -> old = parse_lock(git_show("HEAD:mix.lock")) new = parse_lock(File.read!("mix.lock")) diff_pairs(old, new) {:base, ref} -> old = parse_lock(git_show("#{ref}:mix.lock")) new = parse_lock(File.read!("mix.lock")) diff_pairs(old, new) {:preview, packages} -> locked = parse_lock(File.read!("mix.lock")) pkgs = if packages == [], do: Map.keys(locked), else: packages Enum.map(pkgs, fn pkg -> latest = hex_api_latest(pkg) {pkg, locked[pkg], latest} end) end end defp diff_pairs(old_map, new_map) do changed = for {pkg, new_v} <- new_map, old_v = old_map[pkg], new_v != old_v do {pkg, old_v, new_v} end added = for {pkg, new_v} <- new_map, !Map.has_key?(old_map, pkg) do {pkg, nil, new_v} end removed = for {pkg, old_v} <- old_map, !Map.has_key?(new_map, pkg) do {pkg, old_v, nil} end %{changed: changed, added: added, removed: removed} end ``` (Pseudocode — actual implementation lives inline in the skill body as shell or `mix run -e` snippets. Plugin has no `lib/` directory.) ## `mix.lock` format Erlang term format: ```elixir %{ "phoenix" => {:hex, :phoenix, "1.7.14", "checksum", :mix, [...], "hexpm", "hash"}, ... } ``` Position 3 is the version string. Use `Code.eval_file("mix.lock")` or a simple regex to extract; the format has been stable for years. ## Why Mode B is the default 1. **Zero new infra** — `git show` is one shell call 2. **Retroactive Igniter check** — catches auto-update footguns before commit 3. **Natural pre-commit hook integration** — Phase 3 PreToolUse hook just invokes Mode B on the working-tree diff 4. **Mode A needs more work** — Hex API resolver for latest versions plus transitive resolution. Strictly more code paths than Mode B. ## What modes deliberately do NOT do - **No transitive resolution in Mode A** — we audit only the packages named or already in `mix.lock`. Use `mix deps.tree` for full transitive. - **No fetch of removed packages** — a `removed: [...]` package is reported in the table as "removed", but no rule runs on it. There's no NEW to audit. - **No automatic mode promotion** — the user picks. Default = B. ## Mode selection examples | Situation | Command | |-----------|---------| | Just ran `mix deps.update`, want pre-commit check | `/phx:deps-audit` | | Reviewing a PR that bumped `mix.lock` | `/phx:deps-audit --base origin/main` | | Considering whether to update `httpoison` | `/phx:deps-audit --preview httpoison` | | Curious which deps could update | `/phx:deps-audit --preview` (capped at 50) | | Just merged main, want fresh audit on whole tree | `git fetch && /phx:deps-audit --base origin/main~1` | -
output-renderer.md 9.8 KB
# Output Renderer Two outputs from every audit: 1. **Stdout** — markdown table + per-package detail (terminal-first) 2. **Sidecar** — `.claude/deps-audit/last-run.json` (machine-readable) ## Scoring weights & risk bands ``` BLOCK = 10 points WARN = 3 points INFO = 1 point Per-package risk = sum of findings' points Risk band: 0 → clean 1–5 → low 6–15 → medium 16+ → high ``` Risk emoji used in markdown for skim-readability: | Band | Emoji | Meaning | |------|-------|---------| | clean | ✅ | No findings | | low | 🟢 | INFO / minor WARN | | medium | 🟡 | Multiple WARNs or 1 BLOCK | | high | 🔴 | Multiple BLOCKs | (Emoji is the *only* place this skill uses Unicode glyphs in output; the rest of the renderer is ASCII to keep diff/grep-friendly.) ## Security-changelog headline (Phase 5) When `diff_cves.py` emits any `patched`, `introduced`, or `still_exposed` findings, the renderer **prepends a headline section** before the package table — the actionable summary of what the update changed. ### Render order ``` 1. (BLOCK headline) introduced + still_exposed CVEs ← if any 2. (INFO headline) patched CVEs (the security changelog) 3. existing markdown table 4. per-package detail sections 5. editorial framing (major bumps, etc.) ``` `patched` is `info` severity but lifts to the headline regardless — it's the user-facing security story for the update. ### Patched (informational, lifted to top) ```markdown # Hex Dependency Audit — Mode B (working vs HEAD) 🚨 4 of 25 updates patched real CVEs. You were exposed: - decimal 2.3.0 → 3.1.0: CVE-2026-32686 (high) — DoS via unbounded exponent Disclosed 2026-05-07. 6 days exposed. - bandit 1.10.3 → 1.11.0: CVE-2026-39805 (high) — HTTP/1.1 request smuggling Disclosed 2026-05-01. 12 days exposed. (+1 more CVE in this bump) - phoenix 1.8.5 → 1.8.7: CVE-2026-32689 (high) — long-poll NDJSON DoS Disclosed 2026-05-05. 8 days exposed. - postgrex 0.22.0 → 0.22.1: CVE-2026-32687 (critical) — SQL injection in Notifications.listen/3 Disclosed 2026-05-12. 1 day exposed. **Recommendation:** ship these updates ASAP. ``` Format spec: - Leading `🚨 N of M` line ONLY when patched findings exist — single emoji per report, never per-finding. - One bullet per (package, GHSA) pair. Bundle multi-CVE bumps with `(+N more CVE in this bump)` to keep the list scannable. - `Disclosed YYYY-MM-DD. N days exposed.` from `exposure_days` field. - Final line: "**Recommendation: ship these updates ASAP.**" ### Introduced (regression — BLOCK at top) ```markdown 🛑 BLOCKED — 1 update INTRODUCED a CVE (regression): - examplepkg 1.0.0 → 1.0.1: CVE-2026-99999 (critical) — Compromised release introduces RCE Disclosed 2026-05-10. This update should NOT be merged. This is a regression: the OLD version did not have this CVE; the NEW version does. Investigate the release (`mix hex.package diff <pkg> <old> <new>`) before proceeding. ``` Format spec: - `🛑 BLOCKED — N update(s) INTRODUCED a CVE` - Bullet per finding, ending "This update should NOT be merged." - Followup line with the `mix hex.package diff` command stub. ### Still exposed (didn't fix it — BLOCK at top) ```markdown 🛑 BLOCKED — 1 CVE STILL EXPOSED after this update: - decimal 2.3.0 → 2.3.1: CVE-2026-32686 (high) — DoS via unbounded exponent The fix is in decimal >= 3.0.0; this bump did not address the CVE. Disclosed 2026-05-07. **Recommendation:** bump further (mix deps.update decimal to >= 3.0.0). ``` Format spec: - `🛑 BLOCKED — N CVE(s) STILL EXPOSED after this update` - Include `patched_versions` constraint from advisory ("fix is in X >= Y"). - Recommend a more-aggressive bump. ### Combining If both `introduced` and `still_exposed` exist, render both headlines (introduced first — it's the more urgent failure mode). `patched` only renders if NO blockers exist for the same packages — the user has bigger problems than the security changelog when something is blocked. ## Markdown table — top section ```markdown # Hex Dependency Audit — Mode B (working vs HEAD) Audited 4 changed · 1 added · 0 removed packages. Tools run: mix hex.audit ✓ · mix_audit ✓ · osv-scanner ✗ (not installed) | Package | Change | Risk | Findings | diff.hex.pm | |---------|--------|------|----------|-------------| | phoenix | 1.7.14 → 1.7.20 | ✅ clean | — | [view](https://diff.hex.pm/diff/phoenix/1.7.14..1.7.20) | | ecto | 3.13.2 → 3.13.4 | 🟢 low (3) | 1× WARN: base64 in priv/img | [view](https://diff.hex.pm/diff/ecto/3.13.2..3.13.4) | | req | 0.5.0 → 0.5.1 | 🔴 high (23) | 2× BLOCK · 1× WARN — maintainer changed | [view](https://diff.hex.pm/diff/req/0.5.0..0.5.1) | | **new_logger** (added) | — → 0.1.0 | 🔴 high (10) | 1× BLOCK: typosquat of `logger` (50× DLs) | [view](https://hex.pm/packages/new_logger) | ``` When the security-changelog headline above already covered a package, the table still includes it (consistent grain) — but the per-package detail section refers back to the headline rather than repeating CVE text. ## Markdown — per-package detail (only for non-clean) For every row with score > 0, emit a detail section in order: ```markdown ## req — 🔴 high (score 23) Maintainer changed: alice_dev → bob_unknown (between 0.5.0 and 0.5.1) ### Findings - **BLOCK · rule 6 · maintainer change** `release publisher`: bob_unknown (was alice_dev) GHSA: n/a · CVE: n/a - **BLOCK · rule 3 · System.cmd at compile time** `lib/req/setup.ex:14` System.cmd("curl", ["-fsSL", url]) Triggered inside `__before_compile__/1`. - **WARN · rule 7 · base64 blob >256 chars** `lib/req/templates.ex:42` "TG9yZW0gaXBzdW0gZG9sb3Igc2l0IGFtZXQs..." (412 chars) ``` Layout rules: - Header includes risk emoji, band name, and score in parens - One blank line between findings - Code blocks for snippets are indented 4 spaces, never fenced (so they render cleanly even when output is piped through grep) - File:line shown as plain `path:line` for terminal hyperlinking - diff.hex.pm link **only** appears in the top table, not per-finding ## Markdown footer ```markdown --- **Aggregate risk:** 🔴 high (1 package over threshold) Re-run after fix: `/phx:deps-audit` Inspect one package: `/phx:deps-audit --preview req` Compare against main: `/phx:deps-audit --base origin/main` Detailed findings: `.claude/deps-audit/last-run.json` ``` ## `--json` flag Replaces the markdown stdout with the same data the sidecar would receive. Useful for CI consumers. Schema: ```json { "version": 1, "generated_at": "2026-05-12T10:32:18Z", "mode": "B", "base": "HEAD", "tools": { "hex_audit": {"available": true, "ran": true}, "mix_audit": {"available": true, "ran": true}, "osv_scanner": {"available": false, "ran": false} }, "summary": { "changed": 4, "added": 1, "removed": 0, "packages_with_findings": 2, "highest_risk_band": "high", "blocks_total": 3, "warns_total": 2, "infos_total": 0 }, "packages": [ { "pkg": "req", "old_version": "0.5.0", "new_version": "0.5.1", "diff_url": "https://diff.hex.pm/diff/req/0.5.0..0.5.1", "risk_score": 23, "risk_band": "high", "maintainer_change": {"from": "alice_dev", "to": "bob_unknown"}, "findings": [ { "rule_id": 6, "severity": "block", "file": null, "line": null, "snippet": "alice_dev → bob_unknown", "message": "Maintainer changed between 0.5.0 and 0.5.1" }, { "rule_id": 3, "severity": "block", "file": "lib/req/setup.ex", "line": 14, "snippet": "System.cmd(\"curl\", [\"-fsSL\", url])", "message": "System.cmd at compile time (inside __before_compile__/1)" } ], "external_findings": [] } ] } ``` Schema versioned with `"version": 1` so Phase 3 hook can detect incompatible upgrades without parsing. ## Sidecar file Always written to `.claude/deps-audit/last-run.json` regardless of `--json` flag. Phase 3 PreToolUse hook reads this file to detect "recently audited" state — if `generated_at` is within the last 10 minutes AND the working `mix.lock` has the same SHA-256, allow `mix deps.get`/`update` without re-audit prompt. ## Quiet mode `--quiet` suppresses clean rows from the markdown table. Useful for CI/pre-commit hooks that should only chime on findings. ## Exit code rubric | Outcome | Exit code | |---------|-----------| | All packages clean | 0 | | Some WARNs, no BLOCKs | 0 | | Any BLOCK finding | 2 | | Audit infrastructure failed (missing tools, bad network) | 3 | Exit `2` is the conventional CC plugin convention for "findings present, human review needed." Exit `3` separates "you can't trust this audit" from "this audit caught something." ## Implementation entry point The renderer reads `${AUDIT_TMPDIR}/findings.json` (a flat array written by each rule + tool wrapper) and the original `diff.json` from the resolver, then emits both outputs. ```bash render() { local fmt="${1:-markdown}" local findings="${AUDIT_TMPDIR}/findings.json" local diff="${AUDIT_TMPDIR}/diff.json" case "${fmt}" in markdown) render_markdown "${diff}" "${findings}" ;; json) render_json "${diff}" "${findings}" ;; esac write_sidecar "${diff}" "${findings}" > .claude/deps-audit/last-run.json } ``` `render_markdown` and `render_json` are jq programs (kept inline in the skill body — see [implementation skeleton in heuristics.md](heuristics.md) for the per-rule shape contract findings must obey). ## Anti-pattern: emoji-only signals Some renderers use emoji as the *only* severity marker. Don't. Always include the band name in text (`high`, `medium`, etc.) so the output is greppable and accessible to terminals without emoji rendering. -
rules-impl.md 18 KB
# MVP Rule Implementations Eight detection routines. Each emits zero or more findings as NDJSON lines (one JSON object per line) to the file named by `${FINDINGS_FILE}` (defaults to `${AUDIT_TMPDIR}/findings.jsonl`). The renderer aggregates by package. ## Portability floor Every new native regex rule MUST use **perl**, not `grep -P` or `grep -E '{n,}'` for n > 255. macOS ships BSD grep, which: - Lacks `-P` (no PCRE — Unicode character classes don't work). - Caps interval quantifier `{n,}` at 255. Perl is preinstalled on every supported platform (macOS, Linux, WSL, Alpine via `apk add perl`). Treat perl as the cross-platform floor — even Linux GNU-grep environments work without change. Same applies to `comm` and `diff` for differential mode: BSD `comm` requires pre-sorted inputs and BSD `diff` lacks `--no-dereference`. Prefer Python (`scripts/diff_findings.py`) or jq for diffs. ## Common finding shape ```json {"pkg": "phoenix", "version": "1.7.20", "rule_id": 3, "severity": "block", "file": "lib/foo.ex", "line": 14, "snippet": "System.cmd(\"curl\", url)", "message": "System.cmd called at module top level"} ``` `file` and `line` are nullable (for package-level findings like Rule 6). `snippet` is truncated to 200 chars. ## Helper: emit a finding ```bash emit() { jq -n -c \ --arg pkg "$1" --arg version "$2" --argjson rule_id "$3" \ --arg severity "$4" --arg file "$5" --arg line "$6" \ --arg snippet "$7" --arg message "$8" \ '{pkg:$pkg, version:$version, rule_id:$rule_id, severity:$severity, file:($file|select(.!="")), line:(if $line=="" then null else ($line|tonumber) end), snippet:$snippet, message:$message}' \ >> "${FINDINGS_FILE:-${AUDIT_TMPDIR}/findings.jsonl}" } ``` All rules below assume `$pkg` and `$ver` are set and `$tarball_dir` points to the unpacked NEW version (e.g., `${AUDIT_TMPDIR}/tarballs/phoenix/1.7.20/`). --- ## Rule 1 — Bidi Unicode control chars (BLOCK) CVE-2021-42574 Trojan Source. Grep for the 9 directional-override control chars across all source files in the unpacked tarball. ```bash rule_1_bidi() { local pkg="$1" ver="$2" dir="$3" # PUA + bidi overrides: U+202A..U+202E, U+2066..U+2069, U+200E, U+200F, U+061C. # macOS BSD grep lacks -P (PCRE), so use perl for Unicode classes. find "${dir}" \( -name '*.ex' -o -name '*.exs' -o -name '*.erl' \) -print0 \ | xargs -0 perl -CSD -ne ' if (/[\x{202A}-\x{202E}\x{2066}-\x{2069}\x{200E}\x{200F}\x{061C}]/) { chomp; printf("%s\x1f%s\x1f%s\n", $ARGV, $., $_); } ' 2>/dev/null \ | while IFS=$'\x1f' read -r file line snippet; do emit "${pkg}" "${ver}" 1 "block" \ "${file#${dir}/}" "${line}" "${snippet:0:200}" \ "Bidi/directional Unicode control char in source (Trojan Source CVE-2021-42574)" done } ``` **FP rate:** ~0%. Legit uses of bidi chars in code source are exceedingly rare; localization strings live in `.po` files (not scanned). --- ## Rule 2 — `Code.eval_*` / `:erlang.apply` non-literal at module scope (BLOCK) Detects dynamic code evaluation called outside a function body. Uses `Code.string_to_quoted/2` to get an AST and walks it. ```bash rule_2_eval() { local pkg="$1" ver="$2" dir="$3" find "${dir}" -name '*.ex' -o -name '*.exs' | while read -r file; do mix run --no-deps-check --no-compile -e " path = System.argv() |> List.first() {:ok, ast} = path |> File.read!() |> Code.string_to_quoted(file: path, columns: true) scan = fn ast, scan -> case ast do # def/defp/defmacro body — skip subtree (function-scope eval is fine) {form, _, _} when form in [:def, :defp, :defmacro, :defmacrop] -> :ok # Top-level eval — flag {{:., _, [{:__aliases__, _, [:Code]}, op]}, meta, [arg | _]} when op in [:eval_string, :eval_quoted] -> unless is_binary(arg) and op == :eval_string and String.length(arg) < 5 do IO.puts(\"#{meta[:line]}|Code.#{op}|non-literal eval at module scope\") end # :erlang.apply with non-literal MFA {{:., _, [:erlang, :apply]}, meta, [m, _f, _a]} when not is_atom(m) -> IO.puts(\"#{meta[:line]}|:erlang.apply|dynamic MFA at module scope\") {_, _, children} when is_list(children) -> Enum.each(children, &scan.(&1, scan)) list when is_list(list) -> Enum.each(list, &scan.(&1, scan)) _ -> :ok end end scan.(ast, scan) " -- "${file}" 2>/dev/null \ | while IFS='|' read -r line snippet message; do emit "${pkg}" "${ver}" 2 "block" \ "${file#${dir}/}" "${line}" "${snippet}" "${message}" done done } ``` **FP rate:** ~1%. Legit eval at module scope is almost never seen; macros that build code use `quote do ... end`, not `Code.eval_string`. --- ## Rule 3 — `System.cmd` / `:os.cmd` / `Port.open` at compile time (BLOCK) Compile-time means: inside `defmacro`, `__before_compile__`, `__after_compile__`, or at the module top level. Same AST walk as Rule 2, different match patterns. ```bash rule_3_compile_exec() { local pkg="$1" ver="$2" dir="$3" find "${dir}" -name '*.ex' -o -name '*.exs' | while read -r file; do mix run --no-deps-check --no-compile -e " path = System.argv() |> List.first() {:ok, ast} = path |> File.read!() |> Code.string_to_quoted(file: path, columns: true) scan = fn ast, in_compile, scan -> case ast do # entering a function body — out of compile scope {form, _, _} = node when form in [:def, :defp] -> :ok # entering a compile-time callback {form, _, children} when form in [:defmacro, :defmacrop] and is_list(children) -> Enum.each(children, &scan.(&1, true, scan)) {:__before_compile__, _, _} -> Enum.each(elem(ast, 2) || [], &scan.(&1, true, scan)) {:__after_compile__, _, _} -> Enum.each(elem(ast, 2) || [], &scan.(&1, true, scan)) # System.cmd / :os.cmd / Port.open at compile scope {{:., _, [{:__aliases__, _, [:System]}, :cmd]}, m, _} when in_compile -> IO.puts(\"#{m[:line]}|System.cmd|System.cmd at compile time\") {{:., _, [:os, :cmd]}, m, _} when in_compile -> IO.puts(\"#{m[:line]}|:os.cmd|:os.cmd at compile time\") {{:., _, [{:__aliases__, _, [:Port]}, :open]}, m, _} when in_compile -> IO.puts(\"#{m[:line]}|Port.open|Port.open at compile time\") {_, _, children} when is_list(children) -> Enum.each(children, &scan.(&1, in_compile, scan)) list when is_list(list) -> Enum.each(list, &scan.(&1, in_compile, scan)) _ -> :ok end end # Top-level expressions in a module body run at compile time scan.(ast, true, scan) " -- \"${file}\" 2>/dev/null \ | while IFS='|' read -r line snippet message; do emit \"${pkg}\" \"${ver}\" 3 \"block\" \ \"${file#${dir}/}\" \"${line}\" \"${snippet}\" \"${message}\" done done } ``` **FP rate:** ~3%. Some legit build-tool packages (`make`-style wrappers) shell out at compile time. Manual review needed when flagged. --- ## Rule 4 — `:erlang.binary_to_term/1` on literal without `:safe` (BLOCK) Unsafe deserialization (CVE-2026-21619 in hex_core itself). Detect calls to `:erlang.binary_to_term/1` without `[:safe]` in the second arg. ```bash rule_4_binary_to_term() { local pkg="$1" ver="$2" dir="$3" find "${dir}" -name '*.ex' -o -name '*.exs' | while read -r file; do mix run --no-deps-check --no-compile -e " path = System.argv() |> List.first() {:ok, ast} = path |> File.read!() |> Code.string_to_quoted(file: path, columns: true) scan = fn ast, scan -> case ast do # arity-1: no opts at all {{:., _, [:erlang, :binary_to_term]}, m, [_]} -> IO.puts(\"#{m[:line]}|:erlang.binary_to_term/1|missing :safe option\") # arity-2: opts present but :safe not in list {{:., _, [:erlang, :binary_to_term]}, m, [_, opts]} when is_list(opts) -> unless :safe in opts do IO.puts(\"#{m[:line]}|:erlang.binary_to_term/2|:safe not in opts\") end {_, _, children} when is_list(children) -> Enum.each(children, &scan.(&1, scan)) list when is_list(list) -> Enum.each(list, &scan.(&1, scan)) _ -> :ok end end scan.(ast, scan) " -- \"${file}\" 2>/dev/null \ | while IFS='|' read -r line snippet message; do emit \"${pkg}\" \"${ver}\" 4 \"block\" \ \"${file#${dir}/}\" \"${line}\" \"${snippet}\" \"${message}\" done done } ``` **FP rate:** ~2%. Internal serialization formats sometimes pass trusted binaries — but those should still use `:safe`. The fix is one keyword. --- ## Rule 5 — New `:git` / `:path` dep in `mix.exs` (BLOCK) Diff `deps/0` between OLD and NEW versions. Flag any new entry that uses `:git:` or `:path:` (not `:hex` — hex is the trusted default). ```bash rule_5_new_git_path() { local pkg="$1" ver="$2" new_dir="$3" old_dir="$4" [ -z "${old_dir}" ] && return 0 # no old version → skip diff rule extract_deps() { mix run --no-deps-check --no-compile -e " path = System.argv() |> List.first() {:ok, ast} = File.read!(path) |> Code.string_to_quoted() # locate deps/0 in module body deps = Macro.prewalk(ast, [], fn {:def, _, [{:deps, _, _} | _]} = node, acc -> {node, [node | acc]} other, acc -> {other, acc} end) |> elem(1) |> List.first() case deps do nil -> :ok {:def, _, [_head, [do: {:__block__, _, _} | _]]} -> :ok # unusual {:def, _, [_head, [do: list]]} when is_list(list) -> for entry <- list do case entry do {dep, _, opts} when is_atom(dep) and is_list(opts) -> cond do Keyword.has_key?(opts, :git) -> IO.puts(\"#{dep}|git|#{inspect(opts[:git])}\") Keyword.has_key?(opts, :path) -> IO.puts(\"#{dep}|path|#{inspect(opts[:path])}\") true -> :ok end _ -> :ok end end _ -> :ok end " -- "${1}" 2>/dev/null } local old_deps new_deps old_deps=$(extract_deps "${old_dir}/mix.exs" 2>/dev/null || true) new_deps=$(extract_deps "${new_dir}/mix.exs" 2>/dev/null || true) # New entries = in new_deps but not in old_deps comm -13 <(echo "${old_deps}" | sort -u) <(echo "${new_deps}" | sort -u) \ | while IFS='|' read -r dep kind src; do [ -z "${dep}" ] && continue emit "${pkg}" "${ver}" 5 "block" \ "mix.exs" "" "{:${dep}, ${kind}: ${src}}" \ "New ${kind} dep added vs old version — bypasses Hex's checksumming" done } ``` **FP rate:** ~5%. Some libraries legit use `:git` for transitive forks (e.g., a pinned `phoenix_html` fork). Manual review acceptable. --- ## Rule 6 — Maintainer change between versions (BLOCK) Implemented in `references/hex-api.md` as `maintainer_change()`. Reproduced here for completeness: ```bash rule_6_maintainer_change() { local pkg="$1" old_ver="$2" new_ver="$3" ver="$4" [ -z "${old_ver}" ] && return 0 local out out=$(maintainer_change "${pkg}" "${old_ver}" "${new_ver}") if [ -n "${out}" ]; then local message=$(echo "${out}" | cut -d'|' -f3) emit "${pkg}" "${ver}" 6 "block" \ "" "" "${message}" "${message}" fi } ``` Depends on Hex API `/api/packages/:name/releases/:version` returning a `publisher.username` field. **FP rate:** ~2%. --- ## Rule 7 — Base64 blob > 256 chars outside allowlisted dirs (WARN) Regex for long base64-ish strings, excluding `priv/static/`, `test/fixtures/`, `assets/`, `node_modules/` (if vendored). ```bash rule_7_base64() { local pkg="$1" ver="$2" dir="$3" # macOS BSD grep caps repetition count at 255; use perl for the >=256 match. find "${dir}" \( -name '*.ex' -o -name '*.exs' \) \ -not -path '*/priv/*' -not -path '*/test/*' \ -not -path '*/assets/*' -not -path '*/node_modules/*' -print0 \ | xargs -0 perl -ne ' if (/"[A-Za-z0-9+\/]{256,}={0,2}"/) { chomp; my $s = substr($_, 0, 200); printf("%s\x1f%s\x1f%s\n", $ARGV, $., $s); } ' 2>/dev/null \ | while IFS=$'\x1f' read -r file line snippet; do emit "${pkg}" "${ver}" 7 "warn" \ "${file#${dir}/}" "${line}" "${snippet}" \ "Base64-like string literal >256 chars outside priv/static, test/fixtures, assets/" done } ``` **Portability note:** Rules 1 and 7 use perl instead of `grep -P` / `grep -E '{256,}'` because (a) macOS BSD grep lacks `-P`, and (b) BSD grep caps `{n,}` at 255. Linux GNU grep handles both natively, but perl works identically on both — net win. **FP rate:** ~8%. License headers, embedded SVG/PNG, fixture data. Bump threshold to 512 if too noisy in real use. The directory exclude list catches the most common legit cases. --- ## Rule 8 — Typosquat (Levenshtein ≤ 2 + 1000× download delta) (BLOCK) Implemented in `references/hex-api.md` as `typosquat_check()`. Reproduced: ```bash rule_8_typosquat() { local pkg="$1" ver="$2" typosquat_check "${pkg}" \ | while IFS='|' read -r sev rule message; do [ -z "${sev}" ] && continue emit "${pkg}" "${ver}" 8 "block" \ "" "" "${message}" "${message}" done } ``` Depends on the top-500 cache (fetched daily) and per-package download counts. **FP rate:** ~1%. --- ## Master runner ```bash run_all_rules() { # Phase 2 differential mode: emit NEW findings to findings.jsonl AND # OLD findings to findings.old.jsonl in the same loop. When DIFFERENTIAL=0, # behave like Phase 1 (no OLD pass). : > ${AUDIT_TMPDIR}/findings.jsonl : > ${AUDIT_TMPDIR}/findings.old.jsonl local differential="${DIFFERENTIAL:-1}" jq -c '.changed[], .added[]' ${AUDIT_TMPDIR}/diff.json \ | while IFS= read -r row; do pkg=$(echo "${row}" | jq -r '.[0]') old=$(echo "${row}" | jq -r '.[1]') new=$(echo "${row}" | jq -r '.[2]') [ "${new}" = "null" ] && continue new_dir="${AUDIT_TMPDIR}/tarballs/${pkg}/${new}" old_dir="" [ "${old}" != "null" ] && old_dir="${AUDIT_TMPDIR}/tarballs/${pkg}/${old}" # --- NEW pass: emit to findings.jsonl --- FINDINGS_FILE="${FINDINGS_FILE:-${AUDIT_TMPDIR}/findings.jsonl}" \ rule_1_bidi "${pkg}" "${new}" "${new_dir}" & FINDINGS_FILE="${FINDINGS_FILE:-${AUDIT_TMPDIR}/findings.jsonl}" \ rule_2_eval "${pkg}" "${new}" "${new_dir}" & FINDINGS_FILE="${FINDINGS_FILE:-${AUDIT_TMPDIR}/findings.jsonl}" \ rule_3_compile_exec "${pkg}" "${new}" "${new_dir}" & FINDINGS_FILE="${FINDINGS_FILE:-${AUDIT_TMPDIR}/findings.jsonl}" \ rule_4_binary_to_term "${pkg}" "${new}" "${new_dir}" & FINDINGS_FILE="${FINDINGS_FILE:-${AUDIT_TMPDIR}/findings.jsonl}" \ rule_7_base64 "${pkg}" "${new}" "${new_dir}" & wait # --- OLD pass (differential mode): emit to findings.old.jsonl --- if [ "${differential}" = "1" ] && [ -n "${old_dir}" ] && [ -d "${old_dir}" ]; then FINDINGS_FILE=${AUDIT_TMPDIR}/findings.old.jsonl \ rule_1_bidi "${pkg}" "${old}" "${old_dir}" & FINDINGS_FILE=${AUDIT_TMPDIR}/findings.old.jsonl \ rule_2_eval "${pkg}" "${old}" "${old_dir}" & FINDINGS_FILE=${AUDIT_TMPDIR}/findings.old.jsonl \ rule_3_compile_exec "${pkg}" "${old}" "${old_dir}" & FINDINGS_FILE=${AUDIT_TMPDIR}/findings.old.jsonl \ rule_4_binary_to_term "${pkg}" "${old}" "${old_dir}" & FINDINGS_FILE=${AUDIT_TMPDIR}/findings.old.jsonl \ rule_7_base64 "${pkg}" "${old}" "${old_dir}" & wait fi # Diff rules (intrinsically diff-aware — single pass). rule_5_new_git_path "${pkg}" "${new}" "${new_dir}" "${old_dir}" # Hex API rules — package-scoped, not differential. rule_6_maintainer_change "${pkg}" "${old}" "${new}" "${new}" rule_8_typosquat "${pkg}" "${new}" done # Set-subtract NEW vs OLD into new_signals / info_signals / dropped_signals. if [ "${differential}" = "1" ]; then python3 "${CLAUDE_SKILL_DIR}/scripts/diff_findings.py" \ --new ${AUDIT_TMPDIR}/findings.jsonl \ --old ${AUDIT_TMPDIR}/findings.old.jsonl \ --new-out ${AUDIT_TMPDIR}/new_signals.jsonl \ --info-out ${AUDIT_TMPDIR}/info_signals.jsonl \ --dropped-out ${AUDIT_TMPDIR}/dropped_signals.jsonl fi } ``` `FINDINGS_FILE` defaults to `findings.jsonl` for backward compat. The `emit()` helper at the top of this file must be updated to redirect to `${FINDINGS_FILE}` instead of the hard-coded path — that's a one-line change inside `emit()`. ## Anti-FP notes - **Rule 7 (base64)** is the dominant noise source. Always include `priv/` in the exclude list. Bumping threshold to 512 chars drops most legit embedded SVG. - **Rules 2 + 3** rely on `Code.string_to_quoted/2` — files with syntax errors silently skip. This is acceptable (the package wouldn't compile anyway). Log skipped files in `--verbose` mode for debugging. - **Rule 5** trips on `:path` deps used for monorepo umbrella apps. The intent of the rule is "dep added that bypasses Hex" — manual approval is reasonable when the user knows the path dep. ## Phase 2 additions - **Differential mode** — `DIFFERENTIAL=1` (default) emits findings on OLD as well as NEW, then NDJSON set-subtracts via `scripts/diff_findings.py`. See `differential.md`. - **Optional precision layers** — `semgrep --config priv/semgrep/` and `yara -r priv/yara/` run alongside native rules when available; both are soft deps. See `semgrep.md` and `yara.md`. - **LLM triage** — high-score packages get verdicts via the `hex-deps-triager` sonnet agent with a `context-supervisor` consolidating per-package output. See `llm-triage.md`. ## Out of scope (Phase 3+) - Sourceror dependency for richer AST patterns (currently using built-in `Code.string_to_quoted/2`). - Sobelow on unpacked tarballs (would cover Rules 2–4 with stronger taint analysis). - Multi-agent orchestrator (5 specialist auditors) — Phase 2 native + Semgrep + YARA cover MVP. - PreToolUse hook on `mix deps.get` — needs Phase 2 ledger to be solid. -
sarif.md 6.1 KB
# SARIF output — `--sarif <path>` Phase 2 adds SARIF 2.1.0 emission so audit findings can flow into the same UIs developers already use for SAST: VS Code's `sarif-viewer`, GitHub's Code Scanning alerts, and JetBrains' Qodana surfaces. ## Iron Laws 1. **SARIF is additive, NOT replacement.** Markdown table on stdout stays the default. `--sarif <path>` writes SARIF alongside; it does not silence the other outputs. 2. **No hard tabs in any emitted YAML/JSON example.** SARIF JSON uses 2-space indent. Examples in this doc and docs/ use 2-space indent too — never tabs (markdown lint MD010). 3. **Schema-validate every run.** Smoke must include a SARIF round-trip: `--sarif /tmp/out.sarif` → validate against schemastore SARIF 2.1.0 schema → re-load. Catches mapping bugs that ad-hoc users won't catch. 4. **Stable `ruleId`.** Use `phx-deps-audit/rule-<N>` (e.g. `phx-deps-audit/rule-3`). Don't change the prefix across versions — GitHub Code Scanning de-duplicates by `ruleId` over PR history. ## SARIF 2.1.0 contract A SARIF log is a JSON document with one or more `runs`. We emit a single run per audit: ```json { "$schema": "https://json.schemastore.org/sarif-2.1.0.json", "version": "2.1.0", "runs": [ { "tool": { "driver": { "name": "phx-deps-audit", "version": "2.10.0", "informationUri": "https://github.com/oliver-kriska/claude-elixir-phoenix", "rules": [ { "id": "phx-deps-audit/rule-1", "name": "BidiUnicodeControlChar", "shortDescription": {"text": "Bidi Unicode control char in source"}, "fullDescription": {"text": "Detects directional-override Unicode control characters in source files (Trojan Source CVE-2021-42574)."}, "helpUri": "https://github.com/oliver-kriska/claude-elixir-phoenix/blob/main/plugins/elixir-phoenix/skills/deps-audit/references/heuristics.md#rule-1", "defaultConfiguration": {"level": "error"} } ] } }, "results": [ { "ruleId": "phx-deps-audit/rule-3", "level": "error", "message": { "text": "System.cmd at compile time in lib/init.ex:14 — package phoenix_extras 0.2.0" }, "locations": [ { "physicalLocation": { "artifactLocation": {"uri": "lib/init.ex"}, "region": {"startLine": 14, "snippet": {"text": "System.cmd(\"curl\", [\"-fsSL\", \"https://attacker.example\"])"}} }, "logicalLocations": [ {"name": "phoenix_extras", "kind": "package"}, {"name": "0.2.0", "kind": "version"} ] } ], "properties": { "package": "phoenix_extras", "version": "0.2.0", "previous_version": "0.1.0", "differential": "new" } } ] } ] } ``` ### Severity mapping | Phase 1 severity | SARIF `level` | SARIF `kind` | |------------------|---------------|--------------| | block | error | fail | | warn | warning | fail | | info | note | informational | SARIF distinguishes `level` (severity) from `kind` (true/false positive). We emit `kind: "fail"` for block/warn and `kind: "informational"` for INFO — matching how Code Scanning UIs group results. ## Mapping logic ```python # scripts/findings_to_sarif.py — outline import json import sys from pathlib import Path LEVELS = {"block": "error", "warn": "warning", "info": "note"} def finding_to_result(f, package, version, previous_version): loc = { "physicalLocation": { "artifactLocation": {"uri": f.get("file") or "mix.exs"}, "region": {"startLine": f.get("line") or 1} } } if f.get("snippet"): loc["physicalLocation"]["region"]["snippet"] = {"text": f["snippet"]} loc["logicalLocations"] = [ {"name": package, "kind": "package"}, {"name": version, "kind": "version"}, ] return { "ruleId": f"phx-deps-audit/rule-{f['rule_id']}", "level": LEVELS.get(f.get("severity", "warn"), "warning"), "message": {"text": f.get("message", "")}, "locations": [loc], "properties": { "package": package, "version": version, "previous_version": previous_version, "differential": f.get("differential", "new"), }, } ``` The skill body invokes this script over `new_signals.jsonl` (or `findings.jsonl` when `--no-differential`). ## Schema validation Smoke target (post-Phase 2): ```bash python3 -m jsonschema -i out.sarif \ https://json.schemastore.org/sarif-2.1.0.json ``` `pip install jsonschema` is a build-time dep, NOT a runtime dep — production audits don't validate (overhead). Validation runs in CI only. ## VS Code sarif-viewer ```text 1. Install: `code --install-extension MS-SarifVSCode.sarif-viewer` 2. Run audit: `/phx:deps-audit --sarif .claude/audit.sarif` 3. In VS Code, open `.claude/audit.sarif` — the SARIF panel auto-opens and lets you jump to file:line for each result. ``` ## GitHub upload-sarif action For projects that want CI gating, add a workflow step: ```yaml - name: Run deps-audit run: | claude code -m "/phx:deps-audit --sarif audit.sarif" - name: Upload SARIF to Code Scanning uses: github/codeql-action/upload-sarif@v3 with: sarif_file: audit.sarif category: phx-deps-audit ``` Note the 2-space indent (NOT tabs) — MD010 lint and the YAML parser both reject hard tabs. The `category` field separates phx-deps-audit results from other Code Scanning sources in the GitHub UI. ## Limitations - **No fix suggestions.** SARIF supports `fixes[]`, but deps-audit doesn't propose patches (it's a deps-level tool, not a code rewriter). Field omitted. - **No code flows.** Could compute `codeFlows[]` for Rule 3 (`__before_compile__` → `System.cmd`) but adds parser complexity for little reviewer-side value. Deferred. - **Single tool.** Each audit emits one SARIF run with one driver. Multi-tool SARIF (Semgrep + YARA + native rules in one file) is a Phase 3 enhancement when those layers are stable. -
semgrep.md 6.6 KB
# Semgrep ruleset — optional precision layer Native rules (`rules-impl.md`) carry the high-severity detection load. Semgrep is an **optional precision layer** that stacks on top — findings from both sources merge into the same NDJSON stream and flow through the differ + renderer. ## Iron Laws 1. **SOFT DEPENDENCY.** If `semgrep` is not installed, the audit skips the layer with a one-line install hint. Never auto-install, never block the audit run on its absence. 2. **DEFENSE IN DEPTH, NOT REPLACEMENT.** Semgrep findings stack on native findings — they don't shadow them. A pattern caught by both layers shows up as two findings (with different `rule_id` namespaces), and the differ + ledger de-dup naturally. 3. **NO `--diff-depth`.** Semgrep's `--diff-depth N` flag is for git-diff scanning. We run on full unpacked tarballs and diff findings ourselves via `diff_findings.py`. Don't conflate. 4. **NORMALIZE TO PHASE 1 SHAPE.** Semgrep JSON output is parsed and coerced into the existing finding JSON shape — same field names, same severity vocabulary. The renderer doesn't know which layer a finding came from. ## Starter ruleset — `priv/semgrep/elixir-supply-chain.yaml` ```yaml # Phase 2 starter ruleset. Sevens rules covering the patterns most # robust to AST detection vs. ad-hoc regex. Soft dep — install # instructions in the README. rules: - id: elixir-compile-time-http message: HTTP fetch at compile time inside __before_compile__ languages: [elixir] severity: ERROR pattern-either: - patterns: - pattern-inside: | defmacro __before_compile__($ENV) do ... end - pattern: System.cmd("curl", $ARGS) - patterns: - pattern-inside: | defmacro __before_compile__($ENV) do ... end - pattern: HTTPoison.get(...) - id: elixir-eval-string message: Code.eval_string with non-literal argument languages: [elixir] severity: ERROR pattern-either: - pattern: Code.eval_string($PAYLOAD) - pattern: Code.eval_quoted($PAYLOAD) pattern-not: Code.eval_string("$LITERAL") pattern-not: Code.eval_quoted("$LITERAL") - id: elixir-dynamic-system-cmd message: System.cmd with variable first argument languages: [elixir] severity: WARNING pattern: System.cmd($CMD, $ARGS) pattern-not: System.cmd("$LITERAL", $ARGS) - id: elixir-base64-to-eval message: Base.decode64 result fed directly to Code.eval_string languages: [elixir] severity: ERROR pattern-either: - pattern: Code.eval_string(Base.decode64!($X)) - patterns: - pattern: | $X = Base.decode64!(...) ... Code.eval_string($X) - id: elixir-on-load-with-side-effects message: __on_load__ callback running System.cmd / File.write languages: [elixir] severity: ERROR pattern-either: - patterns: - pattern-inside: | def __on_load__() do ... end - pattern: System.cmd(...) - patterns: - pattern-inside: | def __on_load__() do ... end - pattern: File.write(...) - id: elixir-binary-to-term-literal message: :erlang.binary_to_term called without :safe option languages: [elixir] severity: ERROR patterns: - pattern: :erlang.binary_to_term($X) - pattern-not: :erlang.binary_to_term($X, [:safe]) - id: elixir-erlang-apply-dynamic message: :erlang.apply with non-literal module or function name languages: [elixir] severity: ERROR pattern: :erlang.apply($MOD, $FUN, $ARGS) pattern-not: :erlang.apply($MOD, :$LITERAL_ATOM, $ARGS) ``` ## Subprocess invocation ```bash run_semgrep() { local tarball_dir="$1" command -v semgrep >/dev/null 2>&1 || { echo "semgrep: not installed (skipping). Install via 'brew install semgrep'." >&2 return 0 } semgrep \ --config "${CLAUDE_SKILL_DIR}/priv/semgrep/elixir-supply-chain.yaml" \ --lang elixir \ --json \ --quiet \ --error \ --metrics off \ "${tarball_dir}" 2>/dev/null \ | jq -c '.results[]?' \ | while IFS= read -r r; do # Normalize Semgrep JSON to Phase 1 finding shape. jq -n -c \ --arg pkg "${PKG}" --arg version "${VER}" \ --arg rule_id "$(echo "${r}" | jq -r '.check_id')" \ --arg severity "$(echo "${r}" | jq -r '.extra.severity | ascii_downcase')" \ --arg file "$(echo "${r}" | jq -r '.path')" \ --argjson line "$(echo "${r}" | jq -r '.start.line')" \ --arg snippet "$(echo "${r}" | jq -r '.extra.lines | .[0:200]')" \ --arg message "$(echo "${r}" | jq -r '.extra.message')" \ '{pkg:$pkg, version:$version, rule_id:("semgrep/" + $rule_id), severity: (if $severity == "error" then "block" elif $severity == "warning" then "warn" else "info" end), file:$file, line:$line, snippet:$snippet, message:$message}' \ >> "${FINDINGS_FILE:-${AUDIT_TMPDIR}/findings.jsonl}" done } ``` Note the `rule_id` namespacing: native rules use integers 1-8; Semgrep rules use string prefixes (`semgrep/elixir-eval-string`). The differ's polymorphic key extraction handles string rule_ids via the `unknown` fallback path — a conservative high-entropy key keeps Semgrep findings stable across re-runs. ## Severity mapping | Semgrep severity | Phase 1 severity | |------------------|------------------| | ERROR | block | | WARNING | warn | | INFO | info | ## Performance Semgrep typically runs in ~5 seconds per package for the 7-rule starter set. On a 20-package audit run, that's 100s additional — considered acceptable for the precision boost. Parallelism via the `run_all_rules` master loop (rules are already backgrounded). ## When NOT to enable Semgrep - CI environments where install adds >2 minutes — pin a Docker image in CI rather than installing fresh each run. - Codebases with non-standard Elixir dialects (Phoenix LiveView HEEx templates inside `.ex` strings) — Semgrep's Elixir parser may stumble. Native rules don't care about embedded HEEx. ## Future ruleset growth The starter set covers the highest-precision wins. Additions in priority order: 1. Module-attribute persistence of decoded payloads (cross-line data-flow). 2. `:os.cmd` and `Port.open` in compile-time contexts. 3. ETF (Erlang term format) literal patterns in source. 4. `Cachex.put_or_create` with dynamic module names. Each new rule MUST land alongside a synthetic fixture in `lab/deps-audit/smoke-test/fixtures.d/` plus an entry in this file's table. -
skill-checklist.md 3.9 KB
# Skill Checklist — Phase 1 lessons baked in Pre-flight for any new skill or agent added to `claude-elixir-phoenix`. Captures the concrete gotchas that cost Phase 1 implementation ~15 minutes each. Apply this list **before** running `make eval`. ## Frontmatter - `name:` — kebab-case, namespaced when surfaced as a command (`phx:deps-audit`). Plain `deps-audit` works for internal skills. - `description:` — **≤200 characters** and **≥3 keywords** from the scorer's Elixir/Phoenix domain list (see `lab/eval/matchers.py:198`). Vague descriptions like "Audit deps" score zero on triggering. Effective pattern: `"<action> <domain> for <risk surface> — <specifics>. Use <when>."` Example: `"Audit Hex dep updates for supply-chain security risk — bidi chars, compile-time exec, maintainer changes. Use after mix deps.update."` - `argument-hint:` — quote the value if it contains `[...]`. Unquoted square brackets get parsed as YAML flow sequences. Wrong: `argument-hint: [--base <ref>] [--json]` Right: `argument-hint: "[--base <ref>] [--json]"` ## Headings - The Iron Laws heading must be literally `## Iron Laws` — no em-dash, no parenthetical. The eval scorer's `section_exists` check does a literal string match. Variants like `## Iron Laws — Never Violate` fail completeness. ## Body - Keep SKILL.md under ~150 lines. Command skills get ~185. - Reference paths use `${CLAUDE_SKILL_DIR}/references/<file>.md`. Bare `references/<file>.md` paths break in subagents because the cwd is the project, not the plugin install. - Iron Laws section: numbered list with concrete prohibitions and a one-sentence rationale per law. Aim for 4-6 laws. ## Markdown lint - Lists need a blank line above and below (MD032). - No hard tabs in code fences. Use 2-space indent. Even one tab in a Makefile snippet trips MD010 and fails CI. - Code fences must declare a language (` ```bash ` not just ` ``` `). ## Eval scorer - `make eval` runs `git diff` to find changed files. **Untracked files are invisible.** `git add` the new skill/agent before running, or `make eval-all` for the full pass. - Trigger-accuracy is cached per skill description hash. After tuning the description, the cache invalidates on its own. - Use `make eval-fix` to see exact failures and get auto-fix suggestions. - **Agents use a different scorer than skills.** `make eval` routes correctly. Manual scoring during development must pick the right one: - Skills → `python3 -m lab.eval.scorer <path>` - Agents → `python3 -m lab.eval.agent_scorer <path>` - Running `scorer.py` against an agent over-restricts (skill-shaped thresholds applied to agent content) and fails on dimensions the agent scorer skips. If a manual run reports completeness-0 on an obviously-complete agent, you used the wrong scorer. ## Agent-specific - No `permissionMode` on any agent. Claude Code ignores it on plugin agents, and the plugin directory's security scan fails `bypassPermissions`. - Read-only agents (no Write): set `disallowedTools: Edit, NotebookEdit` (NOT Write — agents need Write to save their own findings file). Set `omitClaudeMd: true`. - `effort:` must match model: `low` for haiku, `medium` for sonnet and opus. Mismatch fails the consistency check. ## Description keyword reference The fixed Elixir/Phoenix domain list is in `lab/eval/matchers.py`. Effective single-word triggers include: `audit`, `security`, `review`, `hex`, `mix`, `liveview`, `ecto`, `oban`, `phoenix`, `elixir`, `migration`, `changeset`, `genserver`, `supervisor`, `compile`, `test`, `debug`, `performance`. Aim for ≥3 of these in the description. ## Quick verification flow ```bash git add plugins/elixir-phoenix/skills/<new-skill>/ make eval-fix # see failures, get fix suggestions # Apply fixes, then: make eval # confirm pass ``` If only one skill or agent changed, `make eval` runs in <30 s. The full `make eval-all` over 51 skills + 26 agents takes ~5 min. -
tarball-fetcher.md 7 KB
# Tarball Fetcher — `mix hex.package fetch` (ephemeral, per-run) Fetches and unpacks Hex tarballs for both old and new versions of each changed package. All artifacts live in `${AUDIT_TMPDIR}/tarballs/` and are removed by the driver's EXIT trap (see `audit-tmpdir.md`). ## Tmpdir layout ``` ${AUDIT_TMPDIR}/ ├── lock.old # diff-resolver: HEAD/base mix.lock ├── lock.new # diff-resolver: working mix.lock ├── diff.json # diff-resolver: {changed, added, removed} ├── hex-api/ │ ├── packages/<pkg>.json │ └── top-500.json ├── tarballs/ │ ├── phoenix/ │ │ ├── 1.7.14/ # unpacked tarball — old │ │ └── 1.7.20/ # unpacked tarball — new │ └── <pkg>/<version>/ ├── mix-audit.json └── cves_{old,new}.json ``` `.claude/deps-audit/last-run.json` is the only **persistent** output — written at the end of rendering, just before the tmpdir is torn down. ## Single-version fetch ```bash fetch_version() { local pkg="$1" version="$2" : "${AUDIT_TMPDIR:?AUDIT_TMPDIR not set}" local dest="${AUDIT_TMPDIR}/tarballs/${pkg}/${version}" # Skip if already fetched this run (e.g., transitive dep listed twice # in the diff). The audit is ephemeral, so this is the only kind of # cache hit that exists. if [ -f "${dest}/hex_metadata.config" ]; then return 0 fi mkdir -p "$(dirname "${dest}")" mix hex.package fetch "${pkg}" "${version}" --unpack -o "${dest}" 2>&1 \ | grep -v "Fetching\|Unpacked\|^$" || true if [ ! -f "${dest}/hex_metadata.config" ]; then echo "ERROR: fetch failed for ${pkg} ${version}" >&2 return 2 fi } ``` `mix hex.package fetch` exit code is 0 on success and 1 on network/checksum failure. Always verify `hex_metadata.config` exists before reporting success — `mix` is occasionally non-zero on success and zero on transient failure. ## Bulk fetch from `diff.json` Reads the resolver's output and fetches every `(pkg, old?, new?)` pair: ```bash fetch_from_diff() { jq -c ' (.changed[] | [.[0], .[1], .[2]]), (.added[] | [.[0], "_skip_old_", .[2]]), (.removed[] | [.[0], .[1], "_skip_new_"]) ' "${AUDIT_TMPDIR}/diff.json" \ | while IFS= read -r row; do pkg=$(echo "$row" | jq -r '.[0]') old=$(echo "$row" | jq -r '.[1]') new=$(echo "$row" | jq -r '.[2]') [ "$old" != "_skip_old_" ] && [ "$old" != "null" ] && fetch_version "$pkg" "$old" [ "$new" != "_skip_new_" ] && [ "$new" != "null" ] && fetch_version "$pkg" "$new" done } ``` `added` packages have no old to fetch. `removed` packages have no new to fetch. Both forms produce a single tarball; rules that need both versions (Rule 5 dep diff, Rule 6 maintainer diff) skip these entries. ## Parallelism **4-way parallel fetch is the default.** Do **not** use `xargs … bash -c 'fetch_version …'` — that needs `export -f fetch_version`, which is bash-only (no-op under zsh) and does not survive across separate Bash tool calls anyway (see `audit-tmpdir.md` "Cross-tool-call handoff"). Materialize a self-contained script with the tmpdir path **baked in** (unquoted heredoc), then fan it out with `xargs`: ```bash AUDIT_TMPDIR="$(cat "${TMPDIR:-/tmp}/phx-audit-dir.txt")" cat > "${AUDIT_TMPDIR}/fetch.sh" <<EOF #!/bin/bash AUDIT_TMPDIR="${AUDIT_TMPDIR}" pkg="\$1"; ver="\$2" out="\${AUDIT_TMPDIR}/tarballs/\${pkg}/\${ver}" mkdir -p "\$out" mix hex.package fetch "\$pkg" "\$ver" --unpack -o "\$out" >/dev/null 2>&1 \\ && echo "OK \$pkg \$ver" || echo "ERR \$pkg \$ver" EOF chmod +x "${AUDIT_TMPDIR}/fetch.sh" jq -r '(.changed[]|"\(.[0]) \(.[2])"),(.added[]|"\(.[0]) \(.[2])")' \ "${AUDIT_TMPDIR}/diff.json" \ | xargs -P 4 -n 2 "${AUDIT_TMPDIR}/fetch.sh" ``` Cap at 4 parallel fetches — `hex.pm` is fine with this and avoids rate-limit headers (`X-Ratelimit-Remaining`). Higher concurrency (`-P 8` or `-P 16`) trips Hex CDN throttling without meaningful speedup. ## Latency budget Measured on the 2026-05-12 virgil dogfood (25-package update, residential ISP). The ephemeral-tmpdir architecture means **every run pays the cold-fetch cost** — there is no on-disk cache to reuse between audit invocations. | Diff size | Wall time | |-----------|-----------| | 1-5 packages | 15-25s | | 10-25 packages | **60-90s** | | 50-100 packages | 3-5 min | | 200+ packages | 8-15 min | Earlier docs claimed 5-10 min was realistic for 25 packages and assumed a persistent cache could amortize the cost. That was wrong on both counts: the 4-way parallel fetcher is fast, and re-auditing identical locks is a rare workflow (people don't re-audit a lock they haven't touched). **Implications for UX:** - Default scan on a 25-package update completes in ~75s total. The user sees streaming progress (see `execution-flow.md`) so the wait feels active, not stalled. - `--quick` mode skips the tarball fetch entirely (it doesn't need unpacked sources — `mix_audit` and `mix hex.audit` read the lock). Target latency: <10s for 25 packages. - The `${AUDIT_TMPDIR}` is wiped on driver exit. There is no "second run is faster" — by design. ## No prune step Pre-Phase-5 versions of this fetcher included a `prune_cache()` walking `.claude/deps-audit/cache/` and removing entries older than 30 days. **That function is gone.** The driver's `trap "rm -rf ${AUDIT_TMPDIR}" EXIT` makes pruning irrelevant — every audit starts fresh and ends clean. If you're upgrading from a pre-2.12 release and `.claude/deps-audit/cache/` exists in your project, it's safe to delete: `rm -rf .claude/deps-audit/cache/`. The new code never writes there. Keep `.claude/deps-audit/last-run.json` and `.claude/deps-audit/policy.exs` — those are the persistent hook contract. ## `.gitignore` rule The `${AUDIT_TMPDIR}` lives under `${TMPDIR}` (system temp), not in the project tree — no .gitignore entry needed for the working artifacts. Keep `.claude/deps-audit/last-run.json` tracked? **Default: ignore.** It's a snapshot that becomes stale; the next audit regenerates it. Phase 3 PreToolUse hook reads it to detect "recently audited" but the file is expected to be local. `.claude/deps-audit/policy.exs` (Phase 3 gate config) is user-owned — track or ignore per your team's policy. ## Failure modes | Failure | Behavior | |---------|----------| | Network timeout to `hex.pm` | Print warning, retry once, then BLOCK with exit 2 | | Package not on `hex.pm` (e.g., `:git` dep) | Skip with note, audit proceeds | | Disk full | Bail immediately — `mix hex.package fetch` will fail loudly | | Concurrent audit on same project | Each run has its own `mktemp -d`; no contention | | `${TMPDIR}` unwritable | Driver `mktemp -d` fails at startup, returns 2 | ## Hex API alternative (transport-only) For Mode A `--preview` we already hit `GET /api/packages/:name` for the latest version. The tarball is also at: ``` https://repo.hex.pm/tarballs/<pkg>-<version>.tar ``` But that needs manual `tar -xf` + checksum verification. Using `mix hex.package fetch` is one line and handles signature checking. Stick with mix. -
testing.md 6 KB
# Testing — Fixtures & Smoke Test The plugin has no Elixir ExUnit harness of its own (it distributes skills, agents, hooks — no `lib/` or `test/`). Tests for `/phx:deps-audit` rules live as a bash smoke-test runner that materializes synthetic fixtures and asserts each rule fires. ## Why heredoc fixtures, not committed `.ex` files The plugin's PostToolUse hook (`format-elixir.sh`) treats every committed `.ex` and `.exs` file as project source and runs `mix format --check-formatted` against it. Our fixtures are intentionally malformed: - Rule 1 fixture contains raw U+202E bidi bytes (Trojan Source). - Rule 2 fixture calls `Code.eval_string(@payload)` at module top level — syntactically valid but semantically hostile. - Rule 5 fixture pair carries an added `:git` dep that should be flagged. Committing these would (a) flag the plugin's own format hook on every write and (b) imply the plugin authors endorse the code as exemplary Elixir. Both bad. Storing the fixture content as `setup.sh` heredocs under `smoke-test/fixtures.d/<name>/` keeps them version-controlled without polluting the plugin's lint surface. ## Harness layout (Phase 2) The harness lives in the plugin repo at `lab/deps-audit/`, outside the installed skill, so plugin installs and generated targets never ship it. ``` lab/deps-audit/ ├── smoke-test/ │ ├── runner.sh # driver — loads every fixtures.d/<name>/ │ ├── lib/detectors.sh # shared rule detectors (perl/grep/awk) │ ├── fixtures.d/<name>/ # one dir per fixture │ │ ├── setup.sh # heredoc'd fixture content, writes into $FIXTURE_DIR │ │ └── expected.txt # rule:N op:>= count:1 assertions │ └── corpus.d/ # real-package lists (benign-100, EEF CNA CVEs) ├── test-assets/hex-api-cassettes/ # recorded hex.pm responses (Rules 6 + 8) └── capture.sh # cassette capture helper ``` Real tarballs for `corpus.d/` lists are fetched by the skill's runtime loader, `scripts/fetch_tarball.sh`. `runner.sh` discovers fixtures automatically — drop a new directory in `fixtures.d/` and it runs next pass. ## Running the smoke test ```bash bash lab/deps-audit/smoke-test/runner.sh ``` Expected output (~1 second): ``` Running 7 fixture(s) under .../fixtures.d: ok 00_clean rule:1 == 0 (got 0) ok 00_clean rule:2 == 0 (got 0) ok 00_clean rule:3 == 0 (got 0) ok 00_clean rule:4 == 0 (got 0) ok 00_clean rule:7 == 0 (got 0) ok 01_bidi rule:1 >= 1 (got 1) ok 02_eval rule:2 >= 1 (got 1) ok 03_compile_exec rule:3 >= 1 (got 1) ok 04_binary_to_term rule:4 >= 1 (got 1) ok 05_git_dep rule:5 >= 1 (got 1) ok 07_base64 rule:7 >= 1 (got 1) smoke: 7 pass, 0 fail ``` Exit `0` = all pass. Exit `1` with `fail` lines otherwise. ## Coverage matrix | Rule | Fixture | Detector under test | Asserts | |------|---------|---------------------|---------| | Rule 1 (bidi) | `01_bidi/` — file with raw U+202E byte | perl `[\x{202A}-\x{202E}\x{2066}-\x{2069}\x{200E}\x{200F}\x{061C}]` | ≥1 finding | | Rule 2 (eval) | `02_eval/` — top-level `Code.eval_string(@payload)` | grep `^[[:space:]]*Code\.eval_(string\|quoted)\(` | ≥1 finding | | Rule 3 (compile exec) | `03_compile_exec/` — `System.cmd` inside `__before_compile__` | awk scope tracker | ≥1 finding | | Rule 4 (binary_to_term) | `04_binary_to_term/` — `:erlang.binary_to_term(blob)` | grep `:erlang\.binary_to_term\([^,]+\)\s*$` | ≥1 finding | | Rule 5 (new :git dep) | `05_git_dep/{old,new}` — new `git:` keyword | grep diff of `git:` count | ≥1 new dep | | Rule 7 (base64) | `07_base64/` — 308-char base64 literal | perl `"[A-Za-z0-9+/]{256,}={0,2}"` | ≥1 finding | | All | `00_clean/` — benign module | All 5 single-tarball rules | 0 findings each | **Rules 6 and 8 are not smoke-tested** — both require live Hex API calls and would make the smoke flaky/slow. They have unit-test stubs in `references/hex-api.md` (synthetic JSON fixtures for the parser, no network). Phase 2 will add VCR-style HTTP cassettes for full coverage. ## Detectors in the smoke test vs full `rules-impl.md` The smoke test uses **lightweight detector approximations** (single-line grep/perl/awk patterns) for speed. The full detectors in `rules-impl.md` use AST walks (`Code.string_to_quoted/2`) and richer scope tracking — they're more accurate but require a working Mix install. The smoke test's job is to catch regressions in the *fixture shape* (e.g., did we accidentally strip the bidi byte during a refactor?), not to validate the AST detectors themselves. When AST-detector behaviour is in question, run the deps-audit skill against a real `mix.lock` change and compare output to the smoke detectors. Discrepancies that favour the AST detector are usually correct; discrepancies that favour the smoke detector are usually bugs in the AST detector worth fixing. ## When to update fixtures - A new rule lands → add a new `fixtures.d/<NN>_<name>/` directory with `setup.sh` + `expected.txt`. Runner picks it up automatically. - An existing detector changes its output shape → adjust the `expected.txt` assertion, not the fixture, unless the fixture itself no longer represents the hostile pattern. - A detector emits unexpected findings on `00_clean/` → that's a false positive regression. Investigate before relaxing the assertion. ## Phase 2 test plan - VCR cassettes for Rules 6 and 8 (Hex API). - Real malicious-package replay fixtures (synthetic but modelled on axios / event-stream patterns). - FP audit on the top-50 most-installed Hex packages (manual review of clean output — listed in plan success criteria). - Property-based testing for the rule combiner using StreamData (would need an Elixir test harness for the plugin, deferred). ## CI integration For now the smoke test is run manually. To wire into the eval pipeline, add a `smoke` target in `Makefile` whose recipe runs `runner.sh` from the harness root. Deferred until corpus fetch from `scripts/fetch_tarball.sh` is reliable enough for CI gating (it depends on hex.pm reachability). -
trusted-publishers.md 3.2 KB
# Hex.pm trusted-publishers — upstream tracking and plugin stance Hex.pm has an open issue tracking server-side provenance attestation for package releases — analogous to npm's "provenance" or PyPI's trusted-publishers feature. - Upstream issue: <https://github.com/hexpm/hexpm/issues/1193> - EEF Ægis roadmap: <https://security.erlef.org/aegis/roadmap/hex-vulnerability-handling.html> When this lands, the plugin can defer registry-level provenance checks to Hex.pm itself rather than reproducing them client-side. ## Plugin stance Phase 1+2+3 work assumes **no registry-side attestation**. Every audit runs locally against tarball contents because that's the only signal available today. As soon as Hex.pm publishes a trusted-publishers API, the plugin adopts a hybrid model: 1. **Prefer registry signal.** A package with a verified trusted publisher (e.g., release built and signed by a GitHub Actions workflow on `main`) gets a positive trust signal — comparable to `:safe_to_run` in the `hex_vet.exs` ledger. 2. **Keep tarball rules as the floor.** Tarball-level audit stays primary for packages without trusted-publisher attestation, and for defense-in-depth even on attested packages — registry-side verification doesn't catch a malicious build pipeline. 3. **Surface attestation gap as a finding.** Once a meaningful share of the ecosystem adopts trusted publishers, packages WITHOUT attestation become the outliers worth flagging. ## Placeholder rule When the API ships, register a new rule slot: ```text Rule 9 — Missing trusted-publisher attestation (INFO → WARN over time) Method: Hex API `/api/packages/:name/releases/:version` reads the `provenance` field (or equivalent). Severity: INFO until adoption > 20% of top-500; WARN after. False positives: dropped before adoption threshold; the absence isn't useful signal until most packages do attest. ``` The placeholder lives here, not in `heuristics.md`, until the upstream API contract stabilizes. Adding it prematurely would lock the plugin into a guessed schema. ## Decision log The Phase 2 corpus review (2026-05-12) noted: **zero verified maintainer-account compromises** in the Hex ecosystem. Trusted publishers reduce the attack surface for the class of compromise the plugin protects against — when registry-side attestation exists, the client-side maintainer-change rule (rule 6) becomes a sanity check rather than a primary defense. This is good news for everyone except the plugin's value proposition. We track upstream because: (a) if attestation ships before an incident, Phase 3's hook value-prop narrows; (b) when it ships, the plugin should adopt — not compete — to stay aligned with where the ecosystem is heading. ## Action items (gated on upstream) - [ ] Watch hexpm#1193 for API design merge - [ ] When API merges, implement Rule 9 placeholder against staging - [ ] When >20% of top-500 attest, promote Rule 9 to WARN default - [ ] When >80% of top-500 attest, reconsider whether tarball rules are still worth the wall-time cost on attested packages No action is required from plugin users today. This doc exists so contributors and downstream-aware users know the plugin's roadmap intersects upstream registry work. -
yara.md 4.8 KB
# YARA byte-pattern layer — optional defense in depth YARA scans unpacked tarballs for byte-pattern signatures that are cheaper to detect as byte-scans than as AST walks: large base64 blobs, embedded BEAM bytecode, gzip/zip magic in source files, cross-ecosystem attack signatures translated from npm. ## Iron Laws 1. **SOFT DEPENDENCY.** If `yara` is absent, skip the layer with an install hint to stderr. Never block the audit on its absence. 2. **DEFENSE IN DEPTH.** YARA findings stack on native + Semgrep findings. Cross-ecosystem signatures (event-stream, flatmap-stream, XZ-style magic) live here precisely because AST detectors don't see byte patterns inside string literals. 3. **`yara` IS THE ONLY SUPPORTED BINARY.** Not `yara-x` (Rust port, different rule semantics). The plugin's starter rules target YARA 4.x. Test compatibility before swapping engines. 4. **NORMALIZE TO PHASE 1 SHAPE.** YARA findings parse into NDJSON with `rule_id` namespaced as `yara/<rule_name>`. Severity comes from the rule's `meta.severity` field. ## Starter rules — `priv/yara/hex-malware.yar` The shipped starter file has 6 rules: | Rule | Severity | Purpose | |------|----------|---------| | `large_base64_blob` | warn | Base64 ≥256 chars — faster than perl regex | | `beam_magic_in_source` | block | BEAM bytecode magic `FOR1` inside source | | `gzip_magic_in_source` | warn | Gzip magic — embedded payload | | `zip_magic_in_source` | warn | Zip magic — embedded payload | | `event_stream_flatmap_signature` | warn | npm event-stream attack literal | | `curl_attacker_pattern` | warn | curl + http:// (non-TLS) URL pattern | All rules emit through the metadata's `rule_id` and `severity` fields so the parser stays generic. ## Subprocess invocation ```bash run_yara() { local tarball_dir="$1" command -v yara >/dev/null 2>&1 || { echo "yara: not installed (skipping). Install via 'brew install yara'." >&2 return 0 } # -r recursive, -s show matched strings, -m show metadata. yara -r -s -m \ "${CLAUDE_SKILL_DIR}/priv/yara/hex-malware.yar" \ "${tarball_dir}" 2>/dev/null \ | parse_yara_output_to_ndjson } parse_yara_output_to_ndjson() { # YARA output format (line-oriented): # <rule_name> [meta1="val",meta2="val"] <file_path> # 0x<offset>:$<string_id>: <string_content> # # Group rule + per-match lines; emit one NDJSON per match. local current_rule="" current_file="" current_meta="{}" while IFS= read -r line; do if [[ "${line}" =~ ^([a-z_][a-z0-9_]*)\ \[(.+)\]\ (.+)$ ]]; then current_rule="${BASH_REMATCH[1]}" current_meta="${BASH_REMATCH[2]}" current_file="${BASH_REMATCH[3]}" elif [[ "${line}" =~ ^0x[0-9a-f]+: ]]; then # severity comes from metadata 'severity="..."'. local severity=warn [[ "${current_meta}" =~ severity=\"([^\"]+)\" ]] && severity="${BASH_REMATCH[1]}" jq -n -c \ --arg pkg "${PKG}" --arg version "${VER}" \ --arg rule_id "yara/${current_rule}" \ --arg severity "${severity}" \ --arg file "${current_file}" \ --arg snippet "${line:0:200}" \ --arg message "YARA: ${current_rule}" \ '{pkg:$pkg, version:$version, rule_id:$rule_id, severity:$severity, file:$file, line:null, snippet:$snippet, message:$message}' \ >> "${FINDINGS_FILE:-${AUDIT_TMPDIR}/findings.jsonl}" fi done } ``` YARA emits `file:line` only for hex offsets, not source lines — matches get `line: null`. The renderer handles null line numbers (it already does for Rules 6 + 8). ## Severity from metadata YARA's compiled rules don't carry severity natively. We use a `meta.severity` string ("block" / "warn" / "info") parsed by the NDJSON normalizer. Rules without a severity meta default to "warn". ## Performance YARA is fast — typical run is <500ms per package for the starter set. Negligible vs. the 5s Semgrep adds. Both layers can run in parallel with the native rule loop via shell backgrounding. ## When NOT to enable YARA - Air-gapped CI without `yara` installed and no network to fetch it — let the soft-dep skip take care of it. - Codebases with large legitimate binary blobs (e.g., embedded graphics in `priv/static/`). YARA's magic-byte rules trip on those — the existing path-exclude logic (skip `priv/`, `assets/`, `test/fixtures/`) should cover it, but verify on first run. ## Rule growth roadmap 1. **XZ-style backdoor markers** — pull from public IoC feeds when YARA-format rules become available. 2. **Macro-bytecode in string literals** — Erlang `.beam` magic variants beyond the `FOR1` header. 3. **Reflection-loader strings** — `:code.load_binary`, `Module.create` with non-literal args (also caught by AST rules, but cheaper here for first-pass triage). Each new rule MUST have a synthetic fixture in `lab/deps-audit/smoke-test/fixtures.d/` + an entry in the table above.
-
-
scripts
-
diff_cves.py 8.9 KB
#!/usr/bin/env python3 """ diff_cves.py — CVE set difference for /phx:deps-audit differential pass. Reads cves_old.json and cves_new.json (produced by mix_audit_diff in references/external-tools.md) and categorizes CVEs into three sets: - patched: in OLD, not in NEW (the security changelog) - introduced: in NEW, not in OLD (regression — block) - still_exposed: in both (didn't fix it — block) Input shape (mix_audit JSON): { "pass": false, "vulnerabilities": [ { "advisory": { "id": "GHSA-xxxx-xxxx-xxxx", "cve": "CVE-2026-12345", "title": "...", "description": "...", "severity": "high", "patched_versions": "~> 1.2.3", "disclosure_date": "2026-05-07" // optional }, "dependency": {"package": "decimal", "version": "2.3.0"} } ] } Keying: (ghsa_id, package). Version is intentionally NOT part of the key — a CVE that affected OLD 2.3.0 and ALSO affects NEW 2.4.0 is "still exposed", regardless of version drift. Output (NDJSON, one finding per line): { "category": "patched" | "introduced" | "still_exposed", "rule_id": "ext:mix-audit:diff", "severity": "block" | "warn" | "info", "ghsa_id": "GHSA-xxxx-xxxx-xxxx", "cve_id": "CVE-2026-12345", "package": "decimal", "old_version": "2.3.0", // null when category == "introduced" "new_version": "3.1.0", // null when category == "patched" "severity_label": "high", "title": "Decimal DoS via unbounded exponent", "disclosed_at": "2026-05-07", "exposure_days": 5, "message": "decimal 2.3.0 → 3.1.0: CVE-2026-32686 (high) ..." } Severity mapping per category: - patched → info (informational — the update is the fix) - introduced → block (the update regressed security) - still_exposed → block (update didn't address the CVE) The renderer (references/output-renderer.md) lifts patched findings to the headline section despite their low severity. Usage: python3 diff_cves.py \\ --old cves_old.json \\ --new cves_new.json \\ --out diff_cves.jsonl """ from __future__ import annotations import argparse import json import sys from datetime import date, datetime from pathlib import Path from typing import Any, Iterable def _normalize_severity(raw: str | None) -> str: if not raw: return "moderate" s = raw.lower().strip() if s in {"critical"}: return "critical" if s in {"high"}: return "high" if s in {"medium", "moderate"}: return "moderate" if s in {"low"}: return "low" return s def _patched_severity(category: str, raw_sev: str) -> str: """Map mix_audit severity → deps-audit severity per category. patched → info (the update IS the fix; informational) introduced → block (regression — block the update) still_exposed → block (didn't fix it — block until further update) """ if category == "patched": return "info" if raw_sev in {"critical", "high"}: return "block" if raw_sev == "moderate": return "warn" return "info" def _key(vuln: dict[str, Any]) -> tuple[str, str]: advisory = vuln.get("advisory", {}) or {} dependency = vuln.get("dependency", {}) or {} ghsa_id = advisory.get("id") or advisory.get("cve") or "" package = dependency.get("package") or "" return (ghsa_id, package) def _exposure_days(disclosed_at: str | None) -> int | None: if not disclosed_at: return None try: d = datetime.strptime(disclosed_at, "%Y-%m-%d").date() except ValueError: return None delta = date.today() - d return max(delta.days, 0) def _load_vulns(path: Path | None) -> list[dict[str, Any]]: if path is None or not path.exists(): return [] try: data = json.loads(path.read_text(encoding="utf-8")) except json.JSONDecodeError as e: print(f"diff_cves: invalid JSON in {path}: {e}", file=sys.stderr) return [] if isinstance(data, list): return data # already a list of vulns return data.get("vulnerabilities", []) or [] def _build_finding( category: str, vuln_old: dict[str, Any] | None, vuln_new: dict[str, Any] | None, ) -> dict[str, Any]: # Prefer NEW for introduced/still_exposed (current state), OLD for # patched (what got fixed). primary = vuln_new if vuln_new and category != "patched" else ( vuln_old or vuln_new or {} ) advisory = primary.get("advisory", {}) or {} old_dep = (vuln_old or {}).get("dependency", {}) or {} new_dep = (vuln_new or {}).get("dependency", {}) or {} package = old_dep.get("package") or new_dep.get("package") or "" old_version = old_dep.get("version") new_version = new_dep.get("version") disclosed_at = advisory.get("disclosure_date") or advisory.get("disclosed_at") ghsa_id = advisory.get("id") or "" cve_id = advisory.get("cve") or "" raw_sev = _normalize_severity(advisory.get("severity")) severity = _patched_severity(category, raw_sev) title = advisory.get("title") or "" # Compose a human-readable message tailored per category. if category == "patched": message = ( f"{package} {old_version} → {new_version}: " f"{cve_id or ghsa_id} ({raw_sev}) — {title}" ) elif category == "introduced": message = ( f"REGRESSION: {package} {new_version} introduces " f"{cve_id or ghsa_id} ({raw_sev}) — {title}" ) else: # still_exposed delta = ( f"{old_version} → {new_version}" if old_version and new_version and old_version != new_version else (new_version or old_version or "") ) message = ( f"STILL EXPOSED: {package} {delta} remains vulnerable to " f"{cve_id or ghsa_id} ({raw_sev}) — {title}" ) finding = { "category": category, "rule_id": "ext:mix-audit:diff", "severity": severity, "ghsa_id": ghsa_id, "cve_id": cve_id, "package": package, "old_version": old_version, "new_version": new_version, "severity_label": raw_sev, "title": title, "disclosed_at": disclosed_at, "exposure_days": _exposure_days(disclosed_at), "message": message, } return finding def diff_cves( old: list[dict[str, Any]], new: list[dict[str, Any]] ) -> list[dict[str, Any]]: old_by_key = {_key(v): v for v in old} new_by_key = {_key(v): v for v in new} findings: list[dict[str, Any]] = [] for key, vuln in old_by_key.items(): if not key[0]: # skip vulns with no GHSA/CVE id continue if key in new_by_key: findings.append(_build_finding("still_exposed", vuln, new_by_key[key])) else: findings.append(_build_finding("patched", vuln, new_by_key.get(key))) for key, vuln in new_by_key.items(): if not key[0]: continue if key not in old_by_key: findings.append(_build_finding("introduced", None, vuln)) # Stable sort: patched first (headline), then introduced/still_exposed # (blockers), each group sorted by package then ghsa_id. category_order = {"patched": 0, "introduced": 1, "still_exposed": 2} findings.sort( key=lambda f: (category_order.get(f["category"], 9), f["package"], f["ghsa_id"]) ) return findings def write_ndjson(path: Path, rows: Iterable[dict[str, Any]]) -> int: path.parent.mkdir(parents=True, exist_ok=True) n = 0 with path.open("w", encoding="utf-8") as fh: for r in rows: fh.write(json.dumps(r, separators=(",", ":")) + "\n") n += 1 return n def main(argv: list[str]) -> int: p = argparse.ArgumentParser(description="CVE set difference for deps-audit") p.add_argument("--old", required=True, help="cves_old.json from mix_audit") p.add_argument("--new", required=True, help="cves_new.json from mix_audit") p.add_argument("--out", default="diff_cves.jsonl", help="NDJSON output path") p.add_argument( "--summary", action="store_true", help="print a one-line summary to stderr (count per category)", ) args = p.parse_args(argv) old = _load_vulns(Path(args.old)) new = _load_vulns(Path(args.new)) findings = diff_cves(old, new) n = write_ndjson(Path(args.out), findings) if args.summary: counts = {"patched": 0, "introduced": 0, "still_exposed": 0} for f in findings: counts[f["category"]] = counts.get(f["category"], 0) + 1 print( f"diff_cves: {n} findings — patched={counts['patched']} " f"introduced={counts['introduced']} " f"still_exposed={counts['still_exposed']}", file=sys.stderr, ) return 0 if __name__ == "__main__": sys.exit(main(sys.argv[1:])) -
diff_findings.py 7.8 KB
#!/usr/bin/env python3 """ diff_findings.py — NDJSON set-subtract for /phx:deps-audit differential mode. Reads two findings.jsonl streams (NEW and OLD) and emits three NDJSON streams: signals that are new in this version, signals shared across both versions (downgraded to INFO), and signals dropped since OLD. Keying is polymorphic per rule: - Rules 1, 2, 3, 4, 7 (file-scoped): (rule_id, file, fn_name, sha256(snippet)[:12]) - Rule 5 (mix.exs dep diff): (rule_id, dep_name, kind) # kind ∈ {git, path} - Rules 6, 8 (package-scoped): (rule_id, pkg) Added-package mode: when --old is omitted or empty, every NEW finding is emitted as a new signal (no subtraction). This matches Phase 1 behavior and is the documented decision for net-new dependencies. Cache invalidation: the caller is responsible for namespacing the cache by a rules-checksum (sha256 of references/rules-impl.md mtime + commit SHA). This script is content-pure and does not maintain its own cache. Usage: python3 diff_findings.py --new findings.jsonl --old findings.old.jsonl \\ --new-out new_signals.jsonl --info-out info_signals.jsonl \\ --dropped-out dropped_signals.jsonl """ from __future__ import annotations import argparse import hashlib import json import re import sys from pathlib import Path from typing import Any, Iterable FILE_SCOPED_RULES = {1, 2, 3, 4, 7} MIX_EXS_RULE = 5 PACKAGE_SCOPED_RULES = {6, 8} def sha12(s: str) -> str: return hashlib.sha256((s or "").encode("utf-8")).hexdigest()[:12] # Match `def`, `defp`, `defmacro`, `defmacrop` with their name. The AST # walk in rule emitters is preferred (see `fn_name_for_line` notes in # differential.md), but this regex walk is the portable fallback when no # Elixir tooling is present. _FN_DEF = re.compile( r"^\s*(?:def|defp|defmacro|defmacrop)\s+([a-z_][A-Za-z0-9_?!]*)" ) def fn_name_for_line(source_lines: list[str], line_no: int) -> str: """Walk upward from line_no to find enclosing named function. Falls back to 'module_scope' for top-level code (no enclosing def) or 'anonymous' for code inside an anonymous fn. The differ favors stability over precision — even an approximate function name is a better stability anchor than line number alone. """ if not source_lines or line_no < 1: return "module_scope" for i in range(min(line_no, len(source_lines)) - 1, -1, -1): m = _FN_DEF.match(source_lines[i]) if m: return m.group(1) return "module_scope" def finding_key(finding: dict[str, Any]) -> tuple: """Polymorphic key for set-subtraction. Phase 1 emitters do not write fn_name. When unset, we derive it cheaply from snippet text only — full AST walks happen in the rule layer, not here. """ rule_id = finding.get("rule_id") if rule_id in FILE_SCOPED_RULES: snippet = finding.get("snippet", "") or "" fn_name = finding.get("fn_name") or _approx_fn_from_snippet(snippet) return ( "file", rule_id, finding.get("file", ""), fn_name, sha12(snippet), ) if rule_id == MIX_EXS_RULE: return ( "mix", rule_id, finding.get("dep_name", "") or _dep_name_from_message(finding), finding.get("kind", "git"), ) if rule_id in PACKAGE_SCOPED_RULES: return ("pkg", rule_id, finding.get("pkg", "")) # Unknown rule_id → fall back to a conservative high-entropy key so # diff never silently drops signals. return ("unknown", rule_id, json.dumps(finding, sort_keys=True)) def _approx_fn_from_snippet(snippet: str) -> str: m = _FN_DEF.match(snippet) return m.group(1) if m else "module_scope" def _dep_name_from_message(finding: dict[str, Any]) -> str: # Rule 5 message format: 'new :git dep "phoenix_extras"' msg = finding.get("message", "") or "" m = re.search(r'"([a-z_][a-z0-9_]*)"', msg) return m.group(1) if m else "" def load_ndjson(path: Path | None) -> list[dict[str, Any]]: if path is None or not path.exists(): return [] out: list[dict[str, Any]] = [] with path.open("r", encoding="utf-8") as fh: for ln, raw in enumerate(fh, 1): raw = raw.strip() if not raw: continue try: out.append(json.loads(raw)) except json.JSONDecodeError as e: print( f"diff_findings: skip {path}:{ln} (invalid JSON: {e})", file=sys.stderr, ) return out def write_ndjson(path: Path, rows: Iterable[dict[str, Any]]) -> int: path.parent.mkdir(parents=True, exist_ok=True) n = 0 with path.open("w", encoding="utf-8") as fh: for r in rows: fh.write(json.dumps(r, separators=(",", ":")) + "\n") n += 1 return n def diff( new: list[dict[str, Any]], old: list[dict[str, Any]] ) -> tuple[list[dict[str, Any]], list[dict[str, Any]], list[dict[str, Any]]]: new_by_key = {finding_key(f): f for f in new} old_keys = {finding_key(f) for f in old} new_signals: list[dict[str, Any]] = [] info_signals: list[dict[str, Any]] = [] for key, f in new_by_key.items(): if key in old_keys: downgraded = dict(f) downgraded["severity"] = "info" downgraded["differential"] = "carried" info_signals.append(downgraded) else: promoted = dict(f) promoted["differential"] = "new" new_signals.append(promoted) new_keys = set(new_by_key.keys()) dropped = [] for f in old: if finding_key(f) not in new_keys: d = dict(f) d["differential"] = "dropped" dropped.append(d) return new_signals, info_signals, dropped def main(argv: list[str]) -> int: p = argparse.ArgumentParser(description="NDJSON differential for deps-audit findings") p.add_argument("--new", required=True, help="findings.jsonl on NEW version") p.add_argument("--old", required=False, help="findings.old.jsonl on OLD version") p.add_argument("--new-out", default="new_signals.jsonl") p.add_argument("--info-out", default="info_signals.jsonl") p.add_argument("--dropped-out", default="dropped_signals.jsonl") p.add_argument( "--added-package-mode", choices=("emit-all", "skip"), default="emit-all", help="how to handle an absent --old (default: emit-all = Phase 1 behavior)", ) args = p.parse_args(argv) new_path = Path(args.new) old_path = Path(args.old) if args.old else None new = load_ndjson(new_path) if old_path is None or not old_path.exists() or not load_ndjson(old_path): if args.added_package_mode == "skip": print( f"diff_findings: no OLD findings; --added-package-mode=skip → 0 signals", file=sys.stderr, ) write_ndjson(Path(args.new_out), []) write_ndjson(Path(args.info_out), []) write_ndjson(Path(args.dropped_out), []) return 0 all_new = [dict(f, differential="new") for f in new] write_ndjson(Path(args.new_out), all_new) write_ndjson(Path(args.info_out), []) write_ndjson(Path(args.dropped_out), []) print( f"diff_findings: no OLD; emitted {len(all_new)} signals as NEW", file=sys.stderr, ) return 0 old = load_ndjson(old_path) new_signals, info_signals, dropped = diff(new, old) n_new = write_ndjson(Path(args.new_out), new_signals) n_info = write_ndjson(Path(args.info_out), info_signals) n_dropped = write_ndjson(Path(args.dropped_out), dropped) print( f"diff_findings: {n_new} new, {n_info} carried/info, {n_dropped} dropped", file=sys.stderr, ) return 0 if __name__ == "__main__": sys.exit(main(sys.argv[1:])) -
fetch_tarball.sh 3.1 KB
#!/usr/bin/env bash # scripts/fetch_tarball.sh — fetch real Hex tarballs into the local cache. # Used by /phx:deps-vet (single-vet) and by calibration runs against the # benign corpus; the offline smoke fixtures never call it. # # Usage: # bash scripts/fetch_tarball.sh phoenix 1.7.21 # bash scripts/fetch_tarball.sh --batch batch.txt # "<pkg> <version>" per line # bash scripts/fetch_tarball.sh --prune # drop tarballs >30 days old # # Cache layout: # ${AUDIT_TMPDIR}/corpus/<pkg>/<version>/ # ├── <pkg>-<version>.tar # raw Hex tarball # └── contents/ # extracted source # # Soft dependency: requires a Mix project context for `mix hex.package fetch`. # If invoked outside a Mix project, falls back to direct repo.hex.pm download. set -u CACHE_ROOT="${HEX_AUDIT_CACHE:-${HOME}/.cache/phx-deps-audit/corpus}" mkdir -p "${CACHE_ROOT}" fail() { echo "fetch: $*" >&2; exit 1; } info() { echo "fetch: $*" >&2; } fetch_one() { local pkg="$1" ver="$2" local dest="${CACHE_ROOT}/${pkg}/${ver}" if [ -d "${dest}/contents" ]; then info "cached: ${pkg} ${ver}" return 0 fi mkdir -p "${dest}/contents" if command -v mix >/dev/null 2>&1 && [ -f mix.exs ]; then # In a Mix project — use the canonical tool. mix hex.package fetch "${pkg}" "${ver}" --output "${dest}/${pkg}-${ver}.tar" \ --unpack >/dev/null 2>&1 || fail "mix hex.package fetch failed for ${pkg} ${ver}" # mix --unpack writes to a folder named ${pkg}-${ver}/; move into contents/ if [ -d "${dest}/${pkg}-${ver}" ]; then mv "${dest}/${pkg}-${ver}"/* "${dest}/contents/" 2>/dev/null || true rmdir "${dest}/${pkg}-${ver}" 2>/dev/null || true fi else # Fallback: download tarball directly via curl. No checksum verification # at this level — production audits validate via hex_metadata.config. local url="https://repo.hex.pm/tarballs/${pkg}-${ver}.tar" info "no mix project; falling back to direct download ${url}" curl -fsSL "${url}" -o "${dest}/${pkg}-${ver}.tar" \ || fail "curl failed for ${pkg} ${ver}" (cd "${dest}/contents" && tar -xf "../${pkg}-${ver}.tar") \ || fail "tar extraction failed for ${pkg} ${ver}" # Hex inner archive is contents.tar.gz inside the outer tar if [ -f "${dest}/contents/contents.tar.gz" ]; then tar -xzf "${dest}/contents/contents.tar.gz" -C "${dest}/contents/" \ || fail "inner contents.tar.gz extraction failed" fi fi info "fetched: ${pkg} ${ver}" } prune() { # Drop tarballs older than 30 days (cache TTL). find "${CACHE_ROOT}" -type d -mtime +30 -exec rm -rf {} + 2>/dev/null || true info "pruned entries older than 30 days" } main() { case "${1:-}" in --prune) prune ;; --batch) [ -f "$2" ] || fail "batch file not found: $2" while IFS=' ' read -r pkg ver; do [ -z "${pkg}" ] && continue [[ "${pkg}" =~ ^# ]] && continue fetch_one "${pkg}" "${ver}" done < "$2" ;; '') fail "usage: fetch_tarball.sh <pkg> <version> | --batch <file> | --prune" ;; *) [ -n "${2:-}" ] || fail "version required" fetch_one "$1" "$2" ;; esac } main "$@" -
findings_to_sarif.py 5.5 KB
#!/usr/bin/env python3 """ findings_to_sarif.py — emit SARIF 2.1.0 from phx-deps-audit NDJSON. Reads one NDJSON file (Phase 2 `new_signals.jsonl` by default; supply `findings.jsonl` when running --no-differential), writes a SARIF log to the path given as the second argument. Each input line is a JSON object with at minimum: rule_id, severity, optionally: file, line, snippet, message, package, version, previous_version, differential. Usage: python3 findings_to_sarif.py <findings.jsonl> <out.sarif> [--plugin-version 3.0.0] Output: SARIF 2.1.0 JSON conforming to https://json.schemastore.org/sarif-2.1.0.json """ import argparse import json import sys from pathlib import Path LEVELS = {"block": "error", "warn": "warning", "info": "note"} KINDS = {"block": "fail", "warn": "fail", "info": "informational"} RULE_DESCRIPTIONS = { 1: ("BidiUnicodeControlChar", "Bidi Unicode control char in source", "Detects directional-override Unicode control characters in source files (Trojan Source CVE-2021-42574)."), 2: ("DynamicEvalAtModuleScope", "Code.eval_* or :erlang.apply with non-literal MFA", "Detects evaluator calls with runtime-determined target at module scope."), 3: ("CompileTimeShellExec", "System.cmd / :os.cmd / Port.open at compile time", "Detects shell-exec calls in compile-time macros."), 4: ("UnsafeBinaryToTerm", ":erlang.binary_to_term/1 on literal without :safe", "Detects unsafe deserialization without the :safe option."), 5: ("NewNonHexDep", "New :git or :path dep in mix.exs", "Detects dependencies that bypass the Hex registry."), 6: ("MaintainerChange", "Maintainer change between versions", "Detects Hex package ownership change between the audited and previous version."), 7: ("LargeBase64Blob", "Base64 blob >256 chars outside priv/static/, test/fixtures/, assets/", "Detects suspiciously large base64 strings in source."), 8: ("TyposquatCandidate", "Typosquat candidate (Levenshtein + download delta)", "Detects packages with names ≤2 edits from top-500 and >1000x download delta."), } def rule_definition(rule_id: int) -> dict: name, short, full = RULE_DESCRIPTIONS.get( rule_id, (f"Rule{rule_id}", f"Rule {rule_id}", f"phx-deps-audit rule {rule_id}"), ) return { "id": f"phx-deps-audit/rule-{rule_id}", "name": name, "shortDescription": {"text": short}, "fullDescription": {"text": full}, "helpUri": ( "https://github.com/oliver-kriska/claude-elixir-phoenix/blob/main/" f"plugins/elixir-phoenix/skills/deps-audit/references/heuristics.md#rule-{rule_id}" ), "defaultConfiguration": {"level": "error"}, } def finding_to_result(f: dict) -> dict: rule_id = f["rule_id"] severity = f.get("severity", "warn") package = f.get("package", "") version = f.get("version", "") region = {"startLine": int(f.get("line") or 1)} if f.get("snippet"): region["snippet"] = {"text": f["snippet"]} loc = { "physicalLocation": { "artifactLocation": {"uri": f.get("file") or "mix.exs"}, "region": region, } } if package: loc["logicalLocations"] = [ {"name": package, "kind": "package"}, {"name": version, "kind": "version"}, ] return { "ruleId": f"phx-deps-audit/rule-{rule_id}", "level": LEVELS.get(severity, "warning"), "kind": KINDS.get(severity, "fail"), "message": {"text": f.get("message") or RULE_DESCRIPTIONS.get(rule_id, ("", "", ""))[1]}, "locations": [loc], "properties": { "package": package, "version": version, "previous_version": f.get("previous_version", ""), "differential": f.get("differential", "new"), }, } def load_ndjson(path: Path) -> list[dict]: if not path.exists(): return [] out = [] with path.open() as fh: for raw in fh: raw = raw.strip() if not raw: continue out.append(json.loads(raw)) return out def build_sarif(findings: list[dict], plugin_version: str) -> dict: rule_ids = sorted({f["rule_id"] for f in findings}) return { "$schema": "https://json.schemastore.org/sarif-2.1.0.json", "version": "2.1.0", "runs": [ { "tool": { "driver": { "name": "phx-deps-audit", "version": plugin_version, "informationUri": "https://github.com/oliver-kriska/claude-elixir-phoenix", "rules": [rule_definition(rid) for rid in rule_ids], } }, "results": [finding_to_result(f) for f in findings], } ], } def main(argv: list[str]) -> int: p = argparse.ArgumentParser(description="Convert phx-deps-audit NDJSON to SARIF 2.1.0") p.add_argument("findings", type=Path, help="NDJSON findings file") p.add_argument("out", type=Path, help="SARIF output path") p.add_argument("--plugin-version", default="3.0.0", help="Plugin version for tool.driver.version") args = p.parse_args(argv) findings = load_ndjson(args.findings) sarif = build_sarif(findings, args.plugin_version) args.out.write_text(json.dumps(sarif, indent=2) + "\n") print(f"sarif: {len(findings)} findings → {args.out}", file=sys.stderr) return 0 if __name__ == "__main__": sys.exit(main(sys.argv[1:]))
-
-
SKILL.md 9.2 KB
--- name: deps-audit description: Audit Hex deps for supply-chain security risk — bidi chars, compile-time exec, maintainer changes, typosquats, CVEs. Use after mix deps.update, when checking if a package upgrade is safe, or reviewing mix.lock PR diffs. effort: medium argument-hint: "[--base <ref> | --preview [pkg...]] [--quick] [--json] [--sarif <path>] [--ci] [--strict] [--no-differential] [--no-llm | --llm] [--trace]" --- # Hex Dependency Audit Non-mutating supply-chain audit for Hex packages. Runs an 8-rule MVP catalogue against changed packages, enriches with Hex API metadata, wraps existing tools (`mix hex.audit`, `mix_audit`, OSV-Scanner), and emits a triage table. ## When to Use - After `mix deps.update` or `mix deps.get` brought in new versions - On PRs that touch `mix.lock` (pre-merge gate) - Before manually updating a single package (`--preview <pkg>`) - When investigating a dependency you don't recognize ## Iron Laws 1. **NEVER claim a diff is clean without inspecting it.** Run all 8 rules on the unpacked NEW tarball. "Looks fine" without a tool run is a false pass. **Always write `.claude/deps-audit/last-run.json`** — its absence is evidence the audit didn't actually run. 2. **NEVER install `mix_audit` / `osv-scanner` — even if asked.** Detect, warn with install instructions, skip cleanly if missing. If the user says "install it," respond with the install command (e.g., `mix deps.add mix_audit --only dev`) and **do not execute it**. The audit skill is non-mutating; `mix.exs` / `mix.lock` are off-limits regardless of consent. 3. **NEVER promote a finding to BLOCK without rule citation.** Every finding shows `rule_id`, `severity`, `file:line`, `snippet`, `message`. No handwaving. 4. **NEVER fetch from Hex API without rate-limiting.** Cap at 5 req/sec. Cache metadata 7 days, top-500 list 1 day. 5. **NEVER run the audit on already-committed lock changes silently** — tell the user which mode (A/B/C) is active and which `(old, new)` pairs resolved. 6. **LLM triage only above threshold.** Native rules + Semgrep + YARA are deterministic. The `hex-deps-triager` agent runs only when score > 10 (1 BLOCK or 3+ WARNs), and its verdicts are advisory — never auto-suppress a finding without human review. ## Operating Modes | Mode | Trigger | Old source | New source | |------|---------|-----------|-----------| | **B** (default) | `/phx:deps-audit` | `git show HEAD:mix.lock` | working `mix.lock` | | **C** (PR) | `/phx:deps-audit --base main` | `git show <ref>:mix.lock` | working `mix.lock` | | **A** (preview) | `/phx:deps-audit --preview httpoison` | locked version | Hex API latest | See `${CLAUDE_SKILL_DIR}/references/operating-modes.md` for full resolver logic. ## Execution Flow Default = full 8-rule scan with streaming progress. `--quick` opts out to CVE + retirement only. See `${CLAUDE_SKILL_DIR}/references/execution-flow.md`. ### Step 1: Resolve the diff Parse the `mix.lock` Erlang term format for both old and new sources. Emit a list of `{pkg, old_version, new_version}` tuples. Surface new-only and removed-only packages separately (a removed package is not audited; a brand-new package gets `old_version = nil` and skips diff-only rules). See `${CLAUDE_SKILL_DIR}/references/diff-resolver.md` for shell + `mix run -e` snippets per mode and the JSON output contract. ### Step 2: Fetch tarballs (per-run tmpdir) For each `(pkg, old, new)`: ``` mix hex.package fetch <pkg> <old> --unpack -o ${AUDIT_TMPDIR}/tarballs/<pkg>/<old>/ mix hex.package fetch <pkg> <new> --unpack -o ${AUDIT_TMPDIR}/tarballs/<pkg>/<new>/ ``` All ephemeral artifacts live under `${AUDIT_TMPDIR}` (driver-owned, removed on exit). See `${CLAUDE_SKILL_DIR}/references/audit-tmpdir.md` and `${CLAUDE_SKILL_DIR}/references/tarball-fetcher.md`. ### Step 3: Run the 8 MVP rules on each NEW tarball | # | Rule | Sev | Method | |---|------|-----|--------| | 1 | Bidi Unicode control chars in `.ex`/`.exs`/`.erl` | BLOCK | grep | | 2 | `Code.eval_*` / `:erlang.apply` with non-literal MFA at module scope | BLOCK | AST (Sourceror or regex+scope) | | 3 | `System.cmd` / `:os.cmd` / `Port.open` at compile time | BLOCK | AST | | 4 | `:erlang.binary_to_term/1` on literal without `:safe` | BLOCK | AST | | 5 | New `:git`/`:path` dep in `mix.exs` (vs old) | BLOCK | AST diff | | 6 | Maintainer change between versions | BLOCK | Hex API | | 7 | Base64 blobs >256 chars outside `priv/static/`, `test/fixtures/`, `assets/` | WARN | regex | | 8 | Levenshtein ≤2 from top-500 + download delta >1000× | BLOCK | Hex API + fuzzy | Full catalogue (35 rules, MVP marked) in `${CLAUDE_SKILL_DIR}/references/heuristics.md`. Bash + `mix run -e` implementations for all 8 MVP rules in `${CLAUDE_SKILL_DIR}/references/rules-impl.md` (single-pass NEW + diff rules + Hex API rules, with `run_all_rules` master loop). ### Step 4: External tool wrappers (parallel) - `mix hex.audit` — retired-package check, always available - `mix_audit` — CVE check via GHSA, if installed (else warn + skip; do NOT install) - `osv-scanner` — CVE check via OSV.dev, if installed (else warn + skip; do NOT install) See `${CLAUDE_SKILL_DIR}/references/external-tools.md` for detection, output parsing, and severity mapping per tool. ### Step 5: Hex API enrichment (per package) - `GET /api/packages/:name` — owners, downloads, inserted_at - `GET /api/packages/:name/releases/:version` — per-release publisher - Compute: `days_since_publish`, `owner_age_days`, `download_velocity` Cap at 5 req/sec. Per-run cache under `${AUDIT_TMPDIR}/hex-api/`. See `${CLAUDE_SKILL_DIR}/references/hex-api.md` for endpoint contracts, caching strategy, Rule 6/8 detection, and Levenshtein implementation. ### Step 5.5: Apply `hex_vet.exs` ledger (if present) If `hex_vet.exs` exists at project root, vetted-version findings are **downgraded to INFO**. Unvetted versions retain their severity. Lock-vs-ledger disagreement: lock wins. See the deps-vet skill's hex-vet schema doc for the "Lock-vs-ledger disagreement" section. Use `/phx:deps-vet <pkg> <version>` (separate skill) to add entries. ### Step 5.7: Differential subtract When run with `DIFFERENTIAL=1` (default), findings that existed in the OLD tarball are downgraded to INFO. Net-new signals reach the renderer at full severity. See `${CLAUDE_SKILL_DIR}/references/differential.md`. ### Step 5.8: LLM triage (when score > threshold) For packages where the aggregate score exceeds 10, the `hex-deps-triager` sonnet agent reads finding + diff windows and produces structured verdicts (`confidence`, `verdict`, `rationale`, `fp_reasons[]`). A `context-supervisor` consolidates verdicts across packages into `triage/consolidated.md`. Main skill reads only the consolidated file. See `${CLAUDE_SKILL_DIR}/references/llm-triage.md`. ### Step 6: Score & render Per-package weighted sum: BLOCK = 10, WARN = 3, INFO = 1. Risk band: 0 clean · 1–5 low · 6–15 medium · 16+ high. Output: 1. **Stdout:** markdown table — `pkg | old → new | risk | findings | diff.hex.pm | maintainer-change` plus a per-package detail section for any non-clean row. 2. **Sidecar (MANDATORY):** Write `.claude/deps-audit/last-run.json`. The Phase 3 gate reads this; an audit that doesn't write it is a no-op for the gate. Always emit, even on clean runs. `--json` flag emits JSON to stdout instead of markdown. See `${CLAUDE_SKILL_DIR}/references/output-renderer.md` for table format, sidecar schema, exit-code rubric, and `--quiet` mode. ## Out of scope / Phase 3 surface - **NEVER modify** `mix.lock`, `mix.exs`, or any project file (non-mutating) - **NEVER auto-install** missing tools (warn + skip) - **Gate** `mix deps.{get,update,compile}` via `deps-audit-gate.sh`. See `${CLAUDE_SKILL_DIR}/references/hook.md`. - **Prompt** for `/phx:compound` after BLOCK findings — corpus self-feeds. - **Emit** SARIF 2.1.0 via `--sarif <path>` and gate CI via `--ci`. ## References - `${CLAUDE_SKILL_DIR}/references/heuristics.md` — full 35-rule catalogue - `${CLAUDE_SKILL_DIR}/references/rules-impl.md` — bash + `mix run -e` for the 8 MVP rules - `${CLAUDE_SKILL_DIR}/references/operating-modes.md` — Mode A/B/C resolver - `${CLAUDE_SKILL_DIR}/references/diff-resolver.md` — shell snippets, lock parser - `${CLAUDE_SKILL_DIR}/references/tarball-fetcher.md` — fetch wrapper, parallel cap, cache prune - `${CLAUDE_SKILL_DIR}/references/external-tools.md` — `mix_audit`, `osv-scanner` wrappers - `${CLAUDE_SKILL_DIR}/references/hex-api.md` — endpoint contracts, rate limit, Rule 6/8 - `${CLAUDE_SKILL_DIR}/references/output-renderer.md` — markdown, JSON v1, exit codes, SARIF - `${CLAUDE_SKILL_DIR}/references/testing.md` — smoke runner, fixture matrix - `${CLAUDE_SKILL_DIR}/references/differential.md` / `llm-triage.md` — Phase 2 NDJSON subtract + triager - `${CLAUDE_SKILL_DIR}/references/semgrep.md` / `yara.md` — Phase 2 precision layers (soft deps) - `${CLAUDE_SKILL_DIR}/references/cassettes.md` / `sarif.md` / `hook.md` / `ci-integration.md` — Phase 3 surface - `${CLAUDE_SKILL_DIR}/references/trusted-publishers.md` / `skill-checklist.md` — upstream + eval - `${CLAUDE_SKILL_DIR}/references/audit-tmpdir.md` — Phase 5 per-run ephemeral storage contract - `${CLAUDE_SKILL_DIR}/references/execution-flow.md` / `differential-cve.md` — Phase 5 default scan + CVE diff
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.