Claude Skill

swmm-experiment-audit

Consolidate Agentic SWMM run artifacts into auditable provenance, comparison records, and local Obsidian audit notes. Use after any SWMM build/run/QA attempt, successful or failed, when an agent or CLI workflow needs a traceable record of inputs, commands, artifacts, metrics, QA

LLM Mart · 0 points · 6 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download zhonghao1995-agentic-swmm-workflow-skills_swmm-experiment-audit-2d743b9.zip · 29 KB
Part of zhonghao1995/agentic-swmm-workflow — 18 skills

Install

skills CLI npx skills add https://github.com/Zhonghao1995/agentic-swmm-workflow/tree/main/skills/swmm-experiment-audit
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install zhonghao1995-agentic-swmm-workflow@llmmart
Git git clone https://github.com/Zhonghao1995/agentic-swmm-workflow.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole zhonghao1995/agentic-swmm-workflow collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

SWMM Experiment Audit

Part of Agentic SWMM — install the project first for the executable toolchain (aiswmm CLI, SWMM solver, MCP servers).

What this skill provides

  • A standard audit layer for Agentic SWMM runs.
  • Consolidation of dispersed manifest.json, QA JSON, logs, metrics, and artifact paths.
  • Machine-readable outputs for reproducibility and review.
  • Obsidian-compatible Markdown notes for human research records.
  • Default local Obsidian export into a clean English audit vault.
  • Automatic update of the Obsidian Experiment Audit Index.
  • Optional run-to-run comparison for baseline/scenario or before/after parser validation.

This skill records what happened. It does not run SWMM, build models, invent missing artifacts, or replace module-level validation.

When to use this skill

Use this skill after any of these events:

  • swmm-end-to-end completes successfully.
  • swmm-end-to-end stops or fails after producing partial artifacts.
  • A user wants an Obsidian-ready experiment note for an existing run directory.
  • A user wants to compare two run directories.
  • A run needs evidence for reproducibility, metric provenance, QA status, or paper claims.

Do not use this skill as a substitute for swmm-runner, swmm-builder, or calibration tools. Run the model first, then audit the run directory.

Output contract

For every audited run, write these files into the run's 09_audit/ directory unless explicit output paths are provided:

  • experiment_provenance.json — machine-readable provenance (immutable).
  • experiment_note.md — human-readable Obsidian digest.
  • model_diagnostics.json — deterministic SWMM screening checks.
  • comparison.json — run-to-run comparison (only when --compare-to is given).

experiment_provenance.json is the machine-readable source for:

  • run identity
  • repo state
  • tool versions
  • command trace
  • input hash records
  • artifact index
  • metrics with source artifacts and source tables
  • QA checks
  • detected warnings and limitations

comparison.json records differences against another run when --compare-to is provided. If no comparison target is provided, it still records that no comparison was requested.

experiment_note.md is Obsidian-compatible Markdown with YAML frontmatter. It should stay readable in GitHub as plain Markdown.

The direct script (audit_run.py) writes a copy of the audit note to the Obsidian vault by default; pass --no-obsidian to suppress it. The canonical CLI (aiswmm audit) does not export to Obsidian by default; pass --obsidian to enable it.

When Obsidian export is active, the note is written into:

~/Documents/Agentic-SWMM-Obsidian-Vault/20_Audit_Layer/Experiment_Audits

and the index is updated at:

~/Documents/Agentic-SWMM-Obsidian-Vault/20_Audit_Layer/Experiment Audit Index.md

CLI

The canonical CLI is aiswmm audit. It wraps the underlying script with backup of prior audit files, MOC regeneration (runs/INDEX.md), and the M2 audit → memory auto-trigger. The direct script path is also supported and is the primary option for MCP / Mode-0 calls.

Canonical CLI (aiswmm audit)

Obsidian export is opt-in with aiswmm audit — pass --obsidian to copy the note to the default vault. Without --obsidian the note is written only into 09_audit/ inside the run directory.

aiswmm audit --run-dir runs/acceptance/latest

With comparison:

aiswmm audit --run-dir runs/acceptance/codex-check-peakfix \
  --compare-to runs/acceptance/codex-check

With explicit metadata and Obsidian export:

aiswmm audit --run-dir runs/real-todcreek-minimal \
  --case-name "Tod Creek minimal" \
  --workflow-mode "minimal real-data fallback" \
  --objective "Verify real-data SWMM execution and preserve provenance." \
  --obsidian

Skip the M2 memory auto-trigger (useful for acceptance / benchmark runs):

aiswmm audit --run-dir runs/acceptance/latest --no-memory

Run-to-run comparison with the standalone verb:

aiswmm compare --run-a runs/baseline --run-b runs/scenario

Direct script path (python3 scripts/audit_run.py)

Use when calling from MCP or when the full aiswmm install is unavailable. Obsidian export is on by default in the script; disable with --no-obsidian.

Initialize a first-user Obsidian vault:

python3 skills/swmm-experiment-audit/scripts/init_obsidian_vault.py
python3 skills/swmm-experiment-audit/scripts/audit_run.py \
  --run-dir runs/acceptance/latest \
  --no-obsidian

With comparison:

python3 skills/swmm-experiment-audit/scripts/audit_run.py \
  --run-dir runs/acceptance/codex-check-peakfix \
  --compare-to runs/acceptance/codex-check \
  --no-obsidian

With explicit metadata including --case-id (records the case slug in experiment_provenance.json for cross-session memory recall):

python3 skills/swmm-experiment-audit/scripts/audit_run.py \
  --run-dir runs/real-todcreek-minimal \
  --case-id tod-creek \
  --case-name "Tod Creek minimal" \
  --workflow-mode "minimal real-data fallback" \
  --objective "Verify real-data SWMM execution and preserve provenance." \
  --no-obsidian

With an Obsidian vault folder:

python3 skills/swmm-experiment-audit/scripts/audit_run.py \
  --run-dir runs/acceptance/latest \
  --obsidian-dir "/path/to/Obsidian/Agentic SWMM/04_Experiments"

Audit rules

  • Always preserve relative paths when the artifact is inside the repository.
  • Include absolute paths in JSON only when useful for local traceability.
  • Record SHA256 for existing file artifacts when feasible.
  • Record artifact role, producer, and downstream use.
  • Preserve command return codes, stdout paths, stderr paths, and timings when available.
  • Treat failed or partial runs as auditable. Missing artifacts should be recorded as missing, not invented.
  • Keep metrics tied to source artifacts and source tables.
  • For SWMM peak flow, prefer Node Inflow Summary / Maximum Total Inflow.
  • Use Outfall Loading Summary / Max Flow only as fallback for outfalls.
  • Do not extract peak flow from Node Depth Summary; that table reports depth and HGL, not flow.

Relationship to swmm-end-to-end

swmm-end-to-end is the executor and orchestrator.

swmm-experiment-audit is the recorder and auditor.

The agent should run this audit skill after every build/run/QA attempt, even when the workflow stops early or fails. The audit output should reference whatever artifacts exist in the run directory and clearly mark missing or incomplete evidence.

Obsidian support

The generated experiment_note.md is designed for Obsidian:

  • YAML frontmatter
  • stable headings
  • tables for QA, metrics, and artifact index
  • relative paths for vault portability
  • no chat transcript or conversational content

The note is also valid GitHub Markdown, so it can be committed as an example or exported as supplementary evidence if desired.

The default local vault is:

~/Documents/Agentic-SWMM-Obsidian-Vault

It is organized for first-time Obsidian use:

00_Home/
10_Memory_Layer/
20_Audit_Layer/
30_Evidence_Layer/
40_Skill_Evolution/
90_Templates/

The direct script always writes canonical outputs into the run directory and copies the note to the Obsidian vault unless --no-obsidian is given. The canonical CLI (aiswmm audit) writes canonical outputs only; add --obsidian to also copy to the vault.

In both paths, use --obsidian-dir and --obsidian-index to target a vault location other than the default ~/Documents/Agentic-SWMM-Obsidian-Vault.

Files (agentic-swmm-workflow)
  • scripts
    • audit_run.py 90.9 KB
      #!/usr/bin/env python3
      from __future__ import annotations
      
      import argparse
      import hashlib
      import json
      import re
      import subprocess
      import platform
      import sys
      from datetime import datetime, timezone
      from pathlib import Path
      from typing import Any
      
      
      REPO_ROOT = Path(__file__).resolve().parents[3]
      DEFAULT_OBSIDIAN_VAULT = Path.home() / "Documents" / "Agentic-SWMM-Obsidian-Vault"
      DEFAULT_OBSIDIAN_AUDIT_DIR = DEFAULT_OBSIDIAN_VAULT / "20_Audit_Layer" / "Experiment_Audits"
      DEFAULT_OBSIDIAN_AUDIT_INDEX = DEFAULT_OBSIDIAN_VAULT / "20_Audit_Layer" / "Experiment Audit Index.md"
      
      
      def now_utc() -> str:
          return datetime.now(timezone.utc).isoformat(timespec="seconds")
      
      
      def read_json(path: Path) -> dict[str, Any]:
          if not path.exists():
              return {}
          try:
              parsed = json.loads(path.read_text(encoding="utf-8"))
          except (OSError, json.JSONDecodeError):
              return {}
          return parsed if isinstance(parsed, dict) else {}
      
      
      def write_json(path: Path, obj: Any) -> None:
          path.parent.mkdir(parents=True, exist_ok=True)
          path.write_text(json.dumps(obj, indent=2, sort_keys=True), encoding="utf-8")
      
      
      def write_text(path: Path, text: str) -> None:
          path.parent.mkdir(parents=True, exist_ok=True)
          path.write_text(text, encoding="utf-8")
      
      
      def obsidian_wikilink(note_path: Path) -> str:
          return f"[[{note_path.stem}]]"
      
      
      def update_obsidian_audit_index(
          index_path: Path,
          provenance: dict[str, Any],
          out_provenance: Path,
          out_comparison: Path,
          obsidian_note: Path,
      ) -> None:
          index_path.parent.mkdir(parents=True, exist_ok=True)
          if index_path.exists():
              text = index_path.read_text(encoding="utf-8")
          else:
              text = "\n".join(
                  [
                      "---",
                      "project: Agentic SWMM",
                      "type: experiment-audit-index",
                      "status: active",
                      "---",
                      "",
                      "# Experiment Audit Index",
                      "",
                      "## Audited Runs",
                      "",
                      "| Run ID | Status | Audit Note | Provenance JSON | Comparison JSON | Last Updated |",
                      "|---|---|---|---|---|---|",
                      "",
                  ]
              )
      
          marker = "|---|---|---|---|---|---|"
          if marker not in text:
              text = text.rstrip() + "\n\n## Audited Runs\n\n| Run ID | Status | Audit Note | Provenance JSON | Comparison JSON | Last Updated |\n|---|---|---|---|---|---|\n"
      
          run_id = str(provenance.get("run_id") or obsidian_note.stem)
          status = str(provenance.get("status") or "unknown")
          generated_at = str(provenance.get("generated_at_utc") or now_utc())
          row = (
              f"| `{run_id}` | {status} | {obsidian_wikilink(obsidian_note)} | "
              f"`{out_provenance}` | `{out_comparison}` | {generated_at} |"
          )
      
          lines = text.splitlines()
          filtered: list[str] = []
          for line in lines:
              if line.startswith(f"| `{run_id}` |"):
                  continue
              filtered.append(line)
      
          insert_at = None
          for i, line in enumerate(filtered):
              if line.strip() == marker:
                  insert_at = i + 1
                  break
      
          if insert_at is None:
              filtered.extend(["", "## Audited Runs", "", "| Run ID | Status | Audit Note | Provenance JSON | Comparison JSON | Last Updated |", marker])
              insert_at = len(filtered)
      
          filtered.insert(insert_at, row)
          write_text(index_path, "\n".join(filtered).rstrip() + "\n")
      
      
      def sha256_file(path: Path) -> str | None:
          if not path.exists() or not path.is_file():
              return None
          digest = hashlib.sha256()
          with path.open("rb") as f:
              for chunk in iter(lambda: f.read(1024 * 1024), b""):
                  digest.update(chunk)
          return digest.hexdigest()
      
      
      def run_git(repo_root: Path, *args: str) -> str | None:
          proc = subprocess.run(["git", *args], cwd=repo_root, capture_output=True, text=True)
          if proc.returncode != 0:
              return None
          return proc.stdout.strip()
      
      
      def get_swmm_version(repo_root: Path) -> str | None:
          try:
              proc = subprocess.run(["swmm5", "--version"], cwd=repo_root, capture_output=True, text=True)
          except FileNotFoundError:
              return None
          if proc.returncode != 0:
              return None
          return proc.stdout.strip() or None
      
      
      def relpath(path: Path, repo_root: Path) -> str:
          try:
              return str(path.resolve().relative_to(repo_root.resolve()))
          except ValueError:
              return str(path)
      
      
      def resolve_recorded_path(value: str | None, repo_root: Path) -> Path | None:
          if not value:
              return None
          p = Path(value)
          if p.is_absolute():
              return p
          return repo_root / p
      
      
      def first_existing(paths: list[Path]) -> Path | None:
          for path in paths:
              if path.exists():
                  return path
          return None
      
      
      # ADR-0004 stage numbering. ``agentic_swmm/agent/swmm_runtime/run_layout.py``
      # is the single source of truth for canonical stage names + their legacy
      # aliases (``run_layout.LEGACY_ALIASES``). This script intentionally stays
      # agentic_swmm-import-free (it runs as a standalone subprocess script), so
      # these tuples are a manually-synced mirror of that module's values —
      # canonical name first (search preference), then every legacy generation.
      # Keep in sync BY VALUE with run_layout.py if a stage is ever renumbered
      # again.
      BUILDER_STAGE_NAMES = ["05_builder", "04_builder", "builder"]
      RUNNER_STAGE_NAMES = ["06_runner", "05_runner", "runner", "01_runner"]
      QA_STAGE_NAMES = ["07_qa", "06_qa"]
      
      
      def find_stage_manifest(run_dir: Path, names: list[str]) -> Path | None:
          direct = [run_dir / name / "manifest.json" for name in names]
          found = first_existing(direct)
          if found:
              return found
          candidates: list[Path] = []
          for pattern in names:
              candidates.extend(sorted(run_dir.glob(f"**/*{pattern.strip('0123456789_')}*/manifest.json")))
          return candidates[0] if candidates else None
      
      
      def find_qa_file(run_dir: Path, filename: str) -> Path | None:
          """Locate a QA-stage file, canonical (``07_qa``) first, then ``06_qa``."""
          return first_existing([run_dir / name / filename for name in QA_STAGE_NAMES])
      
      
      def find_builder_inp(run_dir: Path) -> Path | None:
          """Locate the built INP by layout when no manifest records it.
      
          The typed tools (the Canada route among them) write
          ``05_builder/model.inp`` and no builder or top manifest. Without this
          fallback the audit called that model "missing evidence" on every such
          run, and the false flag flowed into memory_summary and failure_advice
          (live test 2026-09-03, S27). Canonical stage first, then legacy names;
          ``model.inp`` first, then a lone ``*.inp`` in the stage.
          """
          for name in BUILDER_STAGE_NAMES:
              stage = run_dir / name
              direct = stage / "model.inp"
              if direct.exists():
                  return direct
              if stage.is_dir():
                  inps = sorted(stage.glob("*.inp"))
                  if len(inps) == 1:
                      return inps[0]
          return None
      
      
      def artifact_record(
          *,
          artifact_id: str,
          role: str,
          path: Path | None,
          repo_root: Path,
          produced_by: str | None = None,
          used_for: list[str] | None = None,
          metadata: dict[str, Any] | None = None,
      ) -> dict[str, Any]:
          exists = bool(path and path.exists())
          record: dict[str, Any] = {
              "id": artifact_id,
              "role": role,
              "relative_path": relpath(path, repo_root) if path else None,
              "absolute_path": str(path.resolve()) if path and path.exists() else (str(path) if path else None),
              "exists": exists,
              "sha256": sha256_file(path) if path and exists else None,
              "produced_by": produced_by,
              "used_for": used_for or [],
          }
          if metadata:
              record["metadata"] = metadata
          return record
      
      
      def normalize_artifacts(records: list[dict[str, Any]]) -> dict[str, dict[str, Any]]:
          out: dict[str, dict[str, Any]] = {}
          for record in records:
              if record["id"] not in out:
                  out[record["id"]] = record
          return out
      
      
      _FLOW_UNITS_RE = re.compile(r"Flow Units\s*\.*\s*([A-Za-z]+)")
      
      
      def flow_units_from_rpt_text(text: str) -> str | None:
          """The report's own flow unit, or None (live finding F-52, 2026-09-02)."""
          match = _FLOW_UNITS_RE.search(text)
          return match.group(1).upper() if match else None
      
      
      def parse_node_inflow_peak(rpt_path: Path | None, node: str | None) -> dict[str, Any] | None:
          if not rpt_path or not node or not rpt_path.exists():
              return None
          text = rpt_path.read_text(errors="ignore")
          lines = text.splitlines()
          units = flow_units_from_rpt_text(text)
      
          def extract_section(title: str) -> str:
              start_idx = None
              for i, line in enumerate(lines):
                  if title.lower() in line.lower():
                      start_idx = i + 1
                      break
              if start_idx is None:
                  return ""
              block: list[str] = []
              for line in lines[start_idx:]:
                  if line.strip().startswith("*****") and block:
                      break
                  block.append(line)
              return "\n".join(block)
      
          inflow_block = extract_section("Node Inflow Summary")
          match = re.search(
              rf"^\s*{re.escape(node)}\s+\S+\s+([-+]?\d+(?:\.\d+)?)\s+([-+]?\d+(?:\.\d+)?)\s+\d+\s+(\d\d):(\d\d)",
              inflow_block,
              re.M,
          )
          if match:
              return {
                  "node": node,
                  "value": float(match.group(2)),
                  "unit": units,
                  "time_hhmm": f"{match.group(3)}:{match.group(4)}",
                  "source_section": "Node Inflow Summary",
                  "source_field": "Maximum Total Inflow",
              }
      
          outfall_block = extract_section("Outfall Loading Summary")
          fallback = re.search(
              rf"^\s*{re.escape(node)}\s+([-+]?\d+(?:\.\d+)?)\s+([-+]?\d+(?:\.\d+)?)\s+([-+]?\d+(?:\.\d+)?)\s+([-+]?\d+(?:\.\d+)?)\s*$",
              outfall_block,
              re.M,
          )
          if fallback:
              return {
                  "node": node,
                  "value": float(fallback.group(3)),
                  "unit": units,
                  "time_hhmm": None,
                  "source_section": "Outfall Loading Summary",
                  "source_field": "Max Flow",
              }
          return None
      
      
      def metric_matches_report(recorded: dict[str, Any], parsed: dict[str, Any] | None) -> bool | None:
          if parsed is None:
              return None
          value = recorded.get("value")
          if value is None:
              value = recorded.get("peak")
          if value is None:
              return None
          try:
              value_f = float(value)
          except (TypeError, ValueError):
              return None
          return abs(value_f - float(parsed["value"])) <= 1e-9
      
      
      def current_repo_state(repo_root: Path) -> dict[str, Any]:
          return {
              "root": str(repo_root),
              "git_head": run_git(repo_root, "rev-parse", "HEAD"),
              "git_branch": run_git(repo_root, "rev-parse", "--abbrev-ref", "HEAD"),
              "git_status_porcelain": run_git(repo_root, "status", "--short"),
          }
      
      
      def _derive_case_name(
          case_name_arg: str | None, top_manifest: dict[str, Any], run_dir: Path
      ) -> str:
          """Case identity used to GROUP runs in parametric memory.
      
          Precedence: an explicit ``--case-name`` > the manifest's ``case_name`` >
          the manifest's ``case_id`` > the run-dir name. The ``case_id`` fallback is
          the fix that lets runs linked via ``aiswmm run --case-id <slug>`` group
          under one case (the memory-informed read side resolves the same slug), so
          the parametric store accumulates >=2 rows per watershed instead of one
          orphan row per unique run-dir name.
          """
          manifest = top_manifest if isinstance(top_manifest, dict) else {}
          explicit = case_name_arg or manifest.get("case_name") or manifest.get("case_id")
          if explicit:
              return explicit
          # Agent runs are named runs/<date>/<session_id>_<slug>_<kind> (kind in
          # {run, chat}); extract the <slug> so they group by watershed instead of by
          # the unique timestamped dir name. Keep in sync with the same rule in
          # agentic_swmm/agent/state.py:_case_id_from_run_dir.
          m = re.match(r"^\d+_(.+)_(?:run|chat)$", run_dir.name)
          if m:
              return m.group(1)
          return run_dir.name
      
      
      def derive_status(qa: dict[str, Any], runner_manifest: dict[str, Any], files_exist: bool) -> str:
          if qa:
              fail_count = qa.get("fail_count")
              if fail_count == 0:
                  return "pass"
              if isinstance(fail_count, int) and fail_count > 0:
                  return "fail"
          if runner_manifest:
              # ``run_ok`` is the solver truth (rc==0 AND no "ERROR n:" lines
              # in the rpt). swmm5 exits 0 while writing solver errors, so
              # trusting return_code alone called solver-errored runs "pass"
              # whenever qa was absent (same honesty class as the benchmark
              # harness fix, 2026-08-08). return_code stays as the fallback
              # for legacy manifests written before run_ok existed.
              run_ok = runner_manifest.get("run_ok")
              if isinstance(run_ok, bool):
                  return "pass" if run_ok else "fail"
              return "pass" if runner_manifest.get("return_code") == 0 else "fail"
          if files_exist:
              return "pass"
          return "unknown"
      
      
      def builder_validation_ok(validation: dict[str, Any]) -> bool:
          """Read the builder manifest's validation block honestly.
      
          The block has two shapes. Builders that validate write error lists
          (``{"errors": [], "warnings": []}``): any non-empty list is a failure.
          The prepared-input workflow (``aiswmm run --inp``) writes a verdict
          (``{"status": "pass", "notes": [...]}``). Treating every truthy value
          as an error read that verdict as a failure, so every prepared-input
          run was audited "fail" and memory stamped ``qa_failed`` on healthy
          runs (live test 2026-09-03, S30). ``status`` decides when present;
          ``notes`` are informational either way.
          """
          status = validation.get("status")
          if isinstance(status, str) and status.strip():
              return status.strip().lower() in {"pass", "passed", "ok", "success"}
          return not any(bool(v) for k, v in validation.items() if k != "notes")
      
      
      def build_qa_checks(
          *,
          acceptance_report: dict[str, Any],
          builder_manifest: dict[str, Any],
          runner_manifest: dict[str, Any],
          peak_metric: dict[str, Any] | None,
          artifacts: dict[str, dict[str, Any]],
      ) -> dict[str, Any]:
          if acceptance_report.get("qa"):
              qa = dict(acceptance_report["qa"])
              qa["status"] = "pass" if qa.get("fail_count") == 0 else "fail"
              return qa
      
          checks: list[dict[str, Any]] = []
          validation = builder_manifest.get("validation")
          if isinstance(validation, dict):
              validation_ok = builder_validation_ok(validation)
              checks.append(
                  {
                      "id": "builder_input_validation",
                      "ok": validation_ok,
                      "detail": json.dumps(validation, sort_keys=True),
                  }
              )
          if runner_manifest:
              checks.append(
                  {
                      "id": "runner_return_code_zero",
                      "ok": runner_manifest.get("return_code") == 0,
                      "detail": f"return_code={runner_manifest.get('return_code')}",
                  }
              )
          rpt = artifacts.get("runner_rpt", {})
          out = artifacts.get("runner_out", {})
          if rpt or out:
              checks.append(
                  {
                      "id": "runner_outputs_exist",
                      "ok": bool(rpt.get("exists")) and bool(out.get("exists")),
                      "detail": f"rpt_exists={rpt.get('exists')} out_exists={out.get('exists')}",
                  }
              )
          if peak_metric:
              checks.append(
                  {
                      "id": "peak_metric_present",
                      "ok": peak_metric.get("value") is not None,
                      "detail": f"source_section={peak_metric.get('source_section')}",
                  }
              )
      
          failed = [c for c in checks if not c.get("ok")]
          return {
              "status": "pass" if checks and not failed else ("fail" if failed else "unknown"),
              "pass_count": len(checks) - len(failed),
              "fail_count": len(failed),
              "checks": checks,
          }
      
      
      def normalize_peak_metric(
          *,
          top_manifest: dict[str, Any],
          acceptance_report: dict[str, Any],
          runner_manifest: dict[str, Any],
          peak_json: dict[str, Any],
          minimal_manifest: dict[str, Any],
          runner_rpt_path: Path | None,
      ) -> dict[str, Any] | None:
          peak = {}
          for candidate in (
              peak_json,
              (runner_manifest.get("metrics") or {}).get("peak") or {},
              (acceptance_report.get("key_outputs") or {}).get("peak") or {},
          ):
              if candidate:
                  peak = candidate
                  break
      
          if not peak and minimal_manifest.get("qoi"):
              qoi = minimal_manifest["qoi"]
              for key, value in qoi.items():
                  if key.startswith("peak_flow_cms_at_"):
                      peak = {
                          "node": key.removeprefix("peak_flow_cms_at_"),
                          "peak": value,
                          "time_hhmm": qoi.get("time_of_peak_hhmm"),
                          "source": "manifest qoi",
                      }
                      break
      
          if not peak:
              return None
      
          node = peak.get("node")
          value = peak.get("value", peak.get("peak"))
          parsed = parse_node_inflow_peak(runner_rpt_path, node)
          source_section = peak.get("source")
          source_field = None
          if source_section == "Node Inflow Summary":
              source_field = "Maximum Total Inflow"
          elif source_section == "Outfall Loading Summary":
              source_field = "Max Flow"
      
          normalized = {
              "name": "peak_flow",
              "node": node,
              "value": value,
              "unit": (parsed or {}).get("unit") or peak.get("units") or peak.get("unit"),
              "time_hhmm": peak.get("time_hhmm"),
              "source_artifact": "runner_rpt" if runner_rpt_path else None,
              "source_section": source_section,
              "source_field": source_field,
              "source_validation": {
                  "parsed_from_report": parsed,
                  "matches_report": metric_matches_report({"value": value}, parsed),
              },
          }
      
          if parsed and normalized["source_section"] in (None, "manifest qoi"):
              normalized["source_section"] = parsed["source_section"]
              normalized["source_field"] = parsed["source_field"]
          return normalized
      
      
      def normalize_continuity(
          *,
          acceptance_report: dict[str, Any],
          runner_manifest: dict[str, Any],
          continuity_json: dict[str, Any],
      ) -> dict[str, Any] | None:
          continuity = {}
          for candidate in (
              continuity_json,
              ((runner_manifest.get("metrics") or {}).get("continuity") or {}),
          ):
              if candidate:
                  continuity = candidate
                  break
          errors = continuity.get("continuity_error_percent")
          if not errors:
              errors = (acceptance_report.get("key_outputs") or {}).get("continuity_error_percent")
          if not errors:
              return None
          return {
              "name": "continuity_error",
              "unit": "percent",
              "values": errors,
              "source_artifact": "runner_rpt",
              "source_sections": ["Runoff Quantity Continuity", "Flow Routing Continuity"],
          }
      
      
      def parse_wq_loads_from_rpt(rpt_path: Path | None) -> dict[str, Any] | None:
          """Extract water-quality load summaries from a SWMM .rpt file.
      
          Uses ``skills/swmm-water-quality/scripts/extract_wq_loads.py`` via
          importlib (same pattern as the source-decomposition hook) so the audit
          script stays self-contained.  Returns ``None`` when the rpt is absent,
          when WQ is not enabled, or when the helper script is unavailable.
          """
          if not rpt_path or not rpt_path.exists():
              return None
          helper = REPO_ROOT / "skills" / "swmm-water-quality" / "scripts" / "extract_wq_loads.py"
          if not helper.is_file():
              return None
          import importlib.util as _ilu
      
          spec = _ilu.spec_from_file_location("_extract_wq_loads_hook", helper)
          if spec is None or spec.loader is None:
              return None
          module = _ilu.module_from_spec(spec)
          sys.modules["_extract_wq_loads_hook"] = module
          try:
              spec.loader.exec_module(module)
              rpt_text = rpt_path.read_text(encoding="utf-8", errors="replace")
              result = module.extract_wq_loads(rpt_text)
          except Exception as exc:  # noqa: BLE001
              print(f"warning: extract_wq_loads hook failed: {exc}", file=sys.stderr)
              return None
          finally:
              sys.modules.pop("_extract_wq_loads_hook", None)
          if not isinstance(result, dict) or not result.get("wq_present"):
              return None
          return result
      
      
      def parse_inp_sections(inp_path: Path | None) -> dict[str, list[list[str]]]:
          if not inp_path or not inp_path.exists():
              return {}
          sections: dict[str, list[list[str]]] = {}
          current: str | None = None
          for raw in inp_path.read_text(encoding="utf-8", errors="ignore").splitlines():
              line = raw.strip()
              if not line or line.startswith(";"):
                  continue
              if line.startswith("[") and line.endswith("]"):
                  current = line[1:-1].upper()
                  sections.setdefault(current, [])
                  continue
              if current is None:
                  continue
              sections.setdefault(current, []).append(line.split())
          return sections
      
      
      def parse_duration_seconds(value: str | None) -> int | None:
          if not value:
              return None
          try:
              if ":" not in value:
                  return int(round(float(value)))
              parts = [int(float(p)) for p in value.split(":")]
          except ValueError:
              return None
          if len(parts) == 2:
              return parts[0] * 60 + parts[1]
          if len(parts) == 3:
              return parts[0] * 3600 + parts[1] * 60 + parts[2]
          return None
      
      
      def fnum(value: Any) -> float | None:
          try:
              return float(value)
          except (TypeError, ValueError):
              return None
      
      
      def diagnostic(
          check_id: str,
          severity: str,
          message: str,
          *,
          evidence: dict[str, Any] | None = None,
          recommendation: str | None = None,
      ) -> dict[str, Any]:
          out = {"id": check_id, "severity": severity, "message": message}
          if evidence:
              out["evidence"] = evidence
          if recommendation:
              out["recommendation"] = recommendation
          return out
      
      
      def parse_flooding_diagnostics(rpt_path: Path | None) -> list[dict[str, Any]]:
          if not rpt_path or not rpt_path.exists():
              return []
          text = rpt_path.read_text(encoding="utf-8", errors="ignore")
          lines = text.splitlines()
          start_idx = None
          for i, line in enumerate(lines):
              if "Node Flooding Summary" in line:
                  start_idx = i + 1
                  break
          if start_idx is None:
              return []
      
          rows: list[dict[str, Any]] = []
          for line in lines[start_idx:]:
              stripped = line.strip()
              if "No nodes were flooded" in stripped:
                  return []
              if not stripped:
                  if rows:
                      break
                  continue
              if stripped.startswith("*****"):
                  break
              parts = stripped.split()
              if len(parts) < 2 or parts[0].startswith("-") or parts[0].lower() in {"node", "name"}:
                  continue
              numeric = [fnum(part) for part in parts[1:]]
              numeric = [value for value in numeric if value is not None]
              if numeric and any(value > 0 for value in numeric):
                  rows.append(
                      {
                          "node": parts[0],
                          "values": numeric,
                          "source_section": "Node Flooding Summary",
                      }
                  )
          return [
              diagnostic(
                  "node_flooding_detected",
                  "warning",
                  "SWMM report contains node flooding values greater than zero.",
                  evidence={"flooded_nodes": rows},
                  recommendation="Inspect flooded nodes before treating the run as hydrologically acceptable.",
              )
          ] if rows else []
      
      
      # Canonical SWMM ``WARNING <n>:`` line (mirrors the ``ERROR <n>:`` pattern in
      # agentic_swmm.agent.honesty). Numbered form only, so the narrative word
      # "warning" in prose never false-positives.
      _RPT_WARNING_RE = re.compile(r"^\s*(WARNING\s+\d+:.*)$")
      
      
      def parse_rpt_warnings(rpt_path: Path | None) -> list[dict[str, Any]]:
          """Collect SWMM's own ``WARNING <n>:`` lines as one info-level diagnostic.
      
          These are SWMM's native advisories (time-step reductions, minimum
          slope/elevation enforced, illegal aspect ratios) — what a modeller sees
          first when opening the .rpt. Surfaced as **info** so the audit reports them
          without ever gating: a warning is not a failure.
          """
          if not rpt_path or not rpt_path.exists():
              return []
          text = rpt_path.read_text(encoding="utf-8", errors="ignore")
          warnings: list[str] = []
          for raw in text.splitlines():
              m = _RPT_WARNING_RE.match(raw)
              if m:
                  warnings.append(m.group(1).rstrip())
          if not warnings:
              return []
          return [
              diagnostic(
                  "swmm_native_warnings",
                  "info",
                  f"SWMM report contains {len(warnings)} native WARNING line(s).",
                  evidence={"warnings": warnings[:50], "count": len(warnings)},
                  recommendation=(
                      "Review SWMM's own warnings (e.g. routing time-step reductions, "
                      "minimum slope/elevation enforced), they often explain "
                      "continuity error or unexpected hydraulics."
                  ),
              )
          ]
      
      
      def build_model_diagnostics(
          *,
          inp_path: Path | None,
          rpt_path: Path | None,
          continuity_metric: dict[str, Any] | None,
          repo_root: Path,
      ) -> dict[str, Any]:
          diagnostics: list[dict[str, Any]] = []
          sections = parse_inp_sections(inp_path)
      
          nodes: set[str] = set()
          node_inverts: dict[str, float] = {}
          for section in ("JUNCTIONS", "OUTFALLS", "STORAGE"):
              for row in sections.get(section, []):
                  if row:
                      nodes.add(row[0])
                      inv = fnum(row[1]) if len(row) > 1 else None
                      if inv is not None:
                          node_inverts[row[0]] = inv
      
          subcatchments = {row[0] for row in sections.get("SUBCATCHMENTS", []) if row}
          raingages = {row[0] for row in sections.get("RAINGAGES", []) if row}
      
          for row in sections.get("SUBCATCHMENTS", []):
              if len(row) < 8:
                  continue
              name, rain_gage, outlet = row[0], row[1], row[2]
              area = fnum(row[3])
              imperv = fnum(row[4])
              width = fnum(row[5])
              if not rain_gage or rain_gage == "*" or rain_gage not in raingages:
                  diagnostics.append(
                      diagnostic(
                          "missing_rain_gage",
                          "error",
                          f"Subcatchment {name} references a missing rain gage.",
                          evidence={"subcatchment": name, "rain_gage": rain_gage},
                          recommendation="Define the rain gage or correct the subcatchment rain-gage reference.",
                      )
                  )
              if not outlet or outlet == "*" or (outlet not in nodes and outlet not in subcatchments):
                  diagnostics.append(
                      diagnostic(
                          "subcatchment_outlet_missing",
                          "error",
                          f"Subcatchment {name} references a missing outlet.",
                          evidence={"subcatchment": name, "outlet": outlet},
                          recommendation="Route the subcatchment to an existing node or subcatchment.",
                      )
                  )
              if area is not None and area <= 0:
                  diagnostics.append(
                      diagnostic("subcatchment_area_nonpositive", "error", f"Subcatchment {name} has non-positive area.", evidence={"subcatchment": name, "area": area})
                  )
              if width is not None and width <= 0:
                  diagnostics.append(
                      diagnostic("subcatchment_width_nonpositive", "error", f"Subcatchment {name} has non-positive width.", evidence={"subcatchment": name, "width": width})
                  )
              if imperv is not None and not (0 <= imperv <= 100):
                  diagnostics.append(
                      diagnostic(
                          "imperviousness_out_of_range",
                          "error",
                          f"Subcatchment {name} has imperviousness outside 0-100 percent.",
                          evidence={"subcatchment": name, "imperviousness": imperv},
                      )
                  )
      
          incoming: set[str] = set()
          for row in sections.get("CONDUITS", []):
              if len(row) < 4:
                  continue
              link, from_node, to_node = row[0], row[1], row[2]
              incoming.add(to_node)
              length = fnum(row[3])
              if from_node not in nodes or to_node not in nodes:
                  diagnostics.append(
                      diagnostic(
                          "conduit_node_missing",
                          "error",
                          f"Conduit {link} references a missing node.",
                          evidence={"link": link, "from_node": from_node, "to_node": to_node},
                      )
                  )
              if length is not None and length <= 0:
                  diagnostics.append(diagnostic("conduit_length_nonpositive", "error", f"Conduit {link} has non-positive length.", evidence={"link": link, "length": length}))
              if length and from_node in node_inverts and to_node in node_inverts:
                  slope = (node_inverts[from_node] - node_inverts[to_node]) / length
                  if slope < -0.001 or slope > 0.2:
                      diagnostics.append(
                          diagnostic(
                              "conduit_slope_suspicious",
                              "warning",
                              f"Conduit {link} has a suspicious invert-derived slope.",
                              evidence={"link": link, "from_node": from_node, "to_node": to_node, "slope": slope},
                              recommendation="Check node invert elevations, conduit direction, and conduit length.",
                          )
                      )
      
          for row in sections.get("OUTFALLS", []):
              if row and row[0] not in incoming:
                  diagnostics.append(
                      diagnostic(
                          "outfall_disconnected",
                          "warning",
                          f"Outfall {row[0]} is not referenced as a conduit downstream node.",
                          evidence={"outfall": row[0]},
                          recommendation="Confirm the outfall is intentionally disconnected or add upstream routing.",
                      )
                  )
      
          for row in sections.get("OPTIONS", []):
              if len(row) >= 2 and row[0].upper() == "ROUTING_STEP":
                  seconds = parse_duration_seconds(row[1])
                  if seconds is not None and seconds > 300:
                      diagnostics.append(
                          diagnostic(
                              "routing_step_large",
                              "warning",
                              "Routing step is larger than 5 minutes.",
                              evidence={"routing_step": row[1], "seconds": seconds},
                              recommendation="Use a smaller routing step for hydraulically sensitive or unstable models.",
                          )
                      )
      
          continuity_values = (continuity_metric or {}).get("values") or {}
          for name, value in continuity_values.items():
              value_f = fnum(value)
              if value_f is not None and abs(value_f) > 5.0:
                  diagnostics.append(
                      diagnostic(
                          "continuity_error_high",
                          "warning",
                          "Continuity error exceeds the 5 percent screening threshold.",
                          evidence={"continuity_type": name, "value_percent": value_f, "threshold_percent": 5.0},
                          recommendation="Inspect model stability, routing step, storage, and external inflow/outflow accounting.",
                      )
                  )
      
          diagnostics.extend(parse_flooding_diagnostics(rpt_path))
          diagnostics.extend(parse_rpt_warnings(rpt_path))
          errors = sum(1 for item in diagnostics if item.get("severity") == "error")
          warnings = sum(1 for item in diagnostics if item.get("severity") == "warning")
          infos = sum(1 for item in diagnostics if item.get("severity") == "info")
          return {
              "schema_version": "1.1",
              "generated_by": "swmm-experiment-audit",
              "generated_at_utc": now_utc(),
              "source_inp": relpath(inp_path, repo_root) if inp_path and inp_path.exists() else None,
              "source_rpt": relpath(rpt_path, repo_root) if rpt_path and rpt_path.exists() else None,
              # info-level diagnostics (SWMM native warnings) never gate status.
              "status": "fail" if errors else ("warning" if warnings else "pass"),
              "error_count": errors,
              "warning_count": warnings,
              "info_count": infos,
              "diagnostics": diagnostics,
          }
      
      
      def normalize_uncertainty_ensemble(top_manifest: dict[str, Any], repo_root: Path) -> dict[str, Any] | None:
          outputs = top_manifest.get("outputs") or {}
          summary_record = outputs.get("summary") if isinstance(outputs, dict) else None
          summary_path = None
          if isinstance(summary_record, dict):
              summary_path = resolve_recorded_path(summary_record.get("path"), repo_root)
          if summary_path is None:
              summary_path = resolve_recorded_path(top_manifest.get("summary"), repo_root)
          summary = read_json(summary_path) if summary_path else {}
          if not summary or "uncertainty" not in str(summary.get("mode", "")).lower():
              return None
      
          entropy_summary = summary.get("entropy_summary") or {}
          nodes = entropy_summary.get("nodes") or []
          entropy_nodes = []
          for node in nodes:
              if isinstance(node, dict):
                  entropy_nodes.append(
                      {
                          "node": node.get("node"),
                          "max_entropy": node.get("max_entropy"),
                          "mean_entropy": node.get("mean_entropy"),
                          "sample_count": node.get("sample_count"),
                          "entropy_json": relpath(resolve_recorded_path(node.get("entropy_json"), repo_root), repo_root)
                          if node.get("entropy_json")
                          else None,
                      }
                  )
          return {
              "mode": summary.get("mode"),
              "sample_count": summary.get("samples"),
              "seed": summary.get("seed"),
              "primary_node": summary.get("node"),
              "selected_plot_node": summary.get("selected_plot_node"),
              "peak_cms_envelope": summary.get("peak_cms_envelope"),
              "peak_percent_change_envelope": summary.get("peak_percent_change_envelope"),
              "node_ranking": summary.get("node_ranking", [])[:5],
              "entropy": {
                  "bins": entropy_summary.get("bins"),
                  "figure": relpath(resolve_recorded_path(entropy_summary.get("figure"), repo_root), repo_root)
                  if entropy_summary.get("figure")
                  else None,
                  "nodes": entropy_nodes,
              },
              "evidence_boundary": "Uncertainty ensemble summary; no observed-flow calibration metrics are implied.",
          }
      
      
      def collect_run(
          run_dir: Path,
          *,
          repo_root: Path,
          case_name: str | None = None,
          workflow_mode: str | None = None,
          objective: str | None = None,
      ) -> tuple[dict[str, Any], dict[str, Any] | None]:
          run_dir = run_dir.resolve()
          top_manifest_path = run_dir / "manifest.json"
          acceptance_report_path = run_dir / "acceptance_report.json"
          acceptance_report_md_path = run_dir / "acceptance_report.md"
          builder_manifest_path = find_stage_manifest(run_dir, BUILDER_STAGE_NAMES)
          runner_manifest_path = find_stage_manifest(run_dir, RUNNER_STAGE_NAMES)
          # The typed network_qa tool writes its report beside the INP it cites
          # (05_builder/network_qa.json, #474); the CLI pipeline writes it to the
          # QA stage. Look in both, QA stage first, or the audit calls a report
          # that exists "missing evidence" (live test 2026-09-03, S27 re-audit).
          network_qa_path = find_qa_file(run_dir, "network_qa.json") or first_existing(
              [run_dir / name / "network_qa.json" for name in BUILDER_STAGE_NAMES]
          )
          continuity_qa_path = find_qa_file(run_dir, "runner_continuity.json")
          peak_qa_path = find_qa_file(run_dir, "runner_peak.json")
          model_diagnostics_path = run_dir / "model_diagnostics.json"
      
          top_manifest = read_json(top_manifest_path)
          acceptance_report = read_json(acceptance_report_path)
          builder_manifest = read_json(builder_manifest_path) if builder_manifest_path else {}
          runner_manifest = read_json(runner_manifest_path) if runner_manifest_path else {}
          network_qa = read_json(network_qa_path) if network_qa_path else {}
          continuity_qa = read_json(continuity_qa_path) if continuity_qa_path else {}
          peak_qa = read_json(peak_qa_path) if peak_qa_path else {}
      
          minimal_manifest = top_manifest if "qoi" in top_manifest or "files" in top_manifest else {}
      
          runner_files = runner_manifest.get("files") or {}
          minimal_files = minimal_manifest.get("files") or {}
          top_outputs = top_manifest.get("outputs") or {}
      
          inp_path = resolve_recorded_path((top_outputs.get("built_inp") or {}).get("path") if isinstance(top_outputs.get("built_inp"), dict) else None, repo_root)
          if inp_path is None:
              inp_path = resolve_recorded_path((builder_manifest.get("outputs") or {}).get("inp"), repo_root)
          if inp_path is None:
              inp_path = resolve_recorded_path(minimal_files.get("inp"), repo_root)
          if inp_path is None:
              # The typed runner records the INP it ran (``inp`` + ``inp_sha256``)
              # even when the file lives outside the run, e.g. examples/. Without
              # this the audit called an external INP "missing evidence" and
              # summarize_memory turned that into a false ``missing_inp`` failure
              # pattern on a passing run (live test 2026-09-03, S29).
              inp_path = resolve_recorded_path(runner_manifest.get("inp"), repo_root)
          if inp_path is None:
              inp_path = find_builder_inp(run_dir)
      
          rpt_path = resolve_recorded_path((top_outputs.get("runner_rpt") or {}).get("path") if isinstance(top_outputs.get("runner_rpt"), dict) else None, repo_root)
          if rpt_path is None:
              rpt_path = resolve_recorded_path(runner_files.get("rpt"), repo_root)
          if rpt_path is None:
              rpt_path = resolve_recorded_path(minimal_files.get("rpt"), repo_root)
      
          out_path = resolve_recorded_path((top_outputs.get("runner_out") or {}).get("path") if isinstance(top_outputs.get("runner_out"), dict) else None, repo_root)
          if out_path is None:
              out_path = resolve_recorded_path(runner_files.get("out"), repo_root)
          if out_path is None:
              out_path = resolve_recorded_path(minimal_files.get("out"), repo_root)
      
          stdout_path = resolve_recorded_path(runner_files.get("stdout"), repo_root)
          stderr_path = resolve_recorded_path(runner_files.get("stderr"), repo_root)
          if stdout_path is None:
              stdout_path = run_dir / "stdout.txt"
          if stderr_path is None:
              stderr_path = run_dir / "stderr.txt"
      
          artifact_records = [
              artifact_record(
                  artifact_id="top_manifest",
                  role="Run-level provenance",
                  path=top_manifest_path,
                  repo_root=repo_root,
                  produced_by="workflow",
                  used_for=["run identity", "repo state", "command trace", "input/output hashes"],
              ),
              artifact_record(
                  artifact_id="acceptance_report_json",
                  role="QA summary",
                  path=acceptance_report_path,
                  repo_root=repo_root,
                  produced_by="acceptance runner",
                  used_for=["QA status", "key outputs", "artifact map"],
              ),
              artifact_record(
                  artifact_id="acceptance_report_md",
                  role="Human-readable QA report",
                  path=acceptance_report_md_path,
                  repo_root=repo_root,
                  produced_by="acceptance runner",
                  used_for=["quick QA inspection"],
              ),
              artifact_record(
                  artifact_id="builder_manifest",
                  role="Build provenance",
                  path=builder_manifest_path,
                  repo_root=repo_root,
                  produced_by="swmm-builder",
                  used_for=["build inputs", "validation diagnostics", "INP hash"],
              ),
              artifact_record(
                  artifact_id="model_inp",
                  role="SWMM input model",
                  path=inp_path,
                  repo_root=repo_root,
                  produced_by="swmm-builder",
                  used_for=["SWMM execution input"],
              ),
              artifact_record(
                  artifact_id="runner_manifest",
                  role="SWMM execution provenance",
                  path=runner_manifest_path,
                  repo_root=repo_root,
                  produced_by="swmm-runner",
                  used_for=["return code", "SWMM version", "runner metrics", "report/output paths"],
              ),
              artifact_record(
                  artifact_id="runner_rpt",
                  role="SWMM text report",
                  path=rpt_path,
                  repo_root=repo_root,
                  produced_by="swmm5 via swmm-runner",
                  used_for=["peak extraction", "continuity extraction", "report review"],
              ),
              artifact_record(
                  artifact_id="runner_out",
                  role="SWMM binary output",
                  path=out_path,
                  repo_root=repo_root,
                  produced_by="swmm5 via swmm-runner",
                  used_for=["hydrograph extraction", "plotting"],
              ),
              artifact_record(
                  artifact_id="runner_stdout",
                  role="SWMM stdout log",
                  path=stdout_path,
                  repo_root=repo_root,
                  produced_by="swmm-runner",
                  used_for=["execution debugging"],
              ),
              artifact_record(
                  artifact_id="runner_stderr",
                  role="SWMM stderr log",
                  path=stderr_path,
                  repo_root=repo_root,
                  produced_by="swmm-runner",
                  used_for=["execution debugging"],
              ),
              artifact_record(
                  artifact_id="network_qa",
                  role="Network QA result",
                  path=network_qa_path,
                  repo_root=repo_root,
                  produced_by="network QA stage",
                  used_for=["network validity check"],
              ),
              artifact_record(
                  artifact_id="continuity_qa",
                  role="Continuity QA result",
                  path=continuity_qa_path,
                  repo_root=repo_root,
                  produced_by="runner QA stage",
                  used_for=["continuity pass/fail", "continuity metric capture"],
              ),
              artifact_record(
                  artifact_id="peak_qa",
                  role="Peak metric QA result",
                  path=peak_qa_path,
                  repo_root=repo_root,
                  produced_by="runner QA stage",
                  used_for=["peak flow", "peak time", "metric source"],
              ),
              artifact_record(
                  artifact_id="model_diagnostics",
                  role="Deterministic SWMM model diagnostics",
                  path=model_diagnostics_path,
                  repo_root=repo_root,
                  produced_by="swmm-experiment-audit",
                  used_for=["SWMM-specific screening diagnostics", "modeling memory"],
              ),
          ]
          artifacts = normalize_artifacts(artifact_records)
      
          peak_metric = normalize_peak_metric(
              top_manifest=top_manifest,
              acceptance_report=acceptance_report,
              runner_manifest=runner_manifest,
              peak_json=peak_qa,
              minimal_manifest=minimal_manifest,
              runner_rpt_path=rpt_path,
          )
          continuity_metric = normalize_continuity(
              acceptance_report=acceptance_report,
              runner_manifest=runner_manifest,
              continuity_json=continuity_qa,
          )
          wq_loads = parse_wq_loads_from_rpt(rpt_path)
          model_diagnostics = build_model_diagnostics(
              inp_path=inp_path,
              rpt_path=rpt_path,
              continuity_metric=continuity_metric,
              repo_root=repo_root,
          )
          uncertainty_ensemble = normalize_uncertainty_ensemble(top_manifest, repo_root)
      
          qa = build_qa_checks(
              acceptance_report=acceptance_report,
              builder_manifest=builder_manifest,
              runner_manifest=runner_manifest,
              peak_metric=peak_metric,
              artifacts=artifacts,
          )
          status = derive_status(qa, runner_manifest, bool(inp_path and rpt_path and out_path))
      
          warnings: list[str] = []
          for warning in top_manifest.get("qa_warnings") or []:
              if isinstance(warning, dict):
                  kind = warning.get("kind") or "qa_warning"
                  boundary = warning.get("boundary")
                  value = warning.get("value_percent", warning.get("value"))
                  detail = f"{kind}: {boundary}" if boundary else str(kind)
                  if value is not None:
                      detail = f"{detail} value={value}"
                  warnings.append(detail)
              else:
                  warnings.append(str(warning))
          if (top_manifest.get("repo") or {}).get("git_status_porcelain"):
              warnings.append("The recorded Git working tree was not clean at run time.")
          if peak_metric and (peak_metric.get("source_validation") or {}).get("matches_report") is False:
              warnings.append("Recorded peak metric does not match the value parsed from the reported source section.")
          if status == "unknown":
              warnings.append("Run status is unknown because no complete QA or runner manifest was found.")
      
          # P0-3: collect ``memories_applied`` from whichever manifest layer recorded
          # it.  The swmm-runner script writes it into its ``manifest.json`` (the
          # ``runner_manifest`` or ``top_manifest`` for simple direct runs).  We
          # check the runner manifest first (most authoritative), then the top-level
          # manifest as a fallback for pipeline runs where the top manifest
          # aggregates the runner output.  An absent field is treated as ``[]`` so
          # pre-existing runs (without the field) remain backward-compatible.
          def _collect_memories_applied() -> list[str]:
              for source in (runner_manifest, top_manifest):
                  if not isinstance(source, dict):
                      continue
                  value = source.get("memories_applied")
                  if isinstance(value, list):
                      return [str(v) for v in value if v]
              return []
      
          memories_applied = _collect_memories_applied()
      
          provenance = {
              "schema_version": "1.1",
              "generated_by": "swmm-experiment-audit",
              "generated_at_utc": now_utc(),
              "run_id": top_manifest.get("run_id") or acceptance_report.get("run_id") or run_dir.name,
              "case_name": _derive_case_name(case_name, top_manifest, run_dir),
              "objective": objective,
              "workflow_mode": workflow_mode or top_manifest.get("pipeline") or acceptance_report.get("pipeline"),
              "status": status,
              "run_dir": {
                  "relative_path": relpath(run_dir, repo_root),
                  "absolute_path": str(run_dir),
              },
              "repo": top_manifest.get("repo") or current_repo_state(repo_root),
              "tools": top_manifest.get("tools")
              or {
                  "python_executable": sys.executable,
                  "python_version": sys.version.replace("\n", " "),
                  "swmm5_version": get_swmm_version(repo_root),
              },
              # ADR-0003: the runtime writes the environment fingerprint into
              # manifest.json; the audit copies it verbatim (this script stays
              # agentic_swmm-import-free). Legacy runs without the block get a
              # minimal audit-time capture so the key is always present.
              "environment": top_manifest.get("environment")
              or {
                  "python": sys.version.split()[0],
                  "platform": platform.platform(),
                  "swmm5_version": get_swmm_version(repo_root),
                  "captured_by": "audit-fallback",
              },
              "commands": top_manifest.get("commands") or [],
              "inputs": top_manifest.get("inputs") or {},
              "artifacts": artifacts,
              "metrics": {
                  "peak_flow": peak_metric,
                  "continuity_error": continuity_metric,
                  "swmm_return_code": runner_manifest.get("return_code"),
                  "builder_counts": builder_manifest.get("counts"),
                  "wq_loads": wq_loads,
              },
              "model_diagnostics": model_diagnostics,
              "uncertainty_ensemble": uncertainty_ensemble,
              "qa": qa,
              "warnings": warnings,
              # Additive field (P0-3): which memory entries were programmatically
              # applied to this run's inputs.  Always present — ``[]`` means no
              # memory was applied.  Never absent so Phase 1's ledger can rely on it.
              "memories_applied": memories_applied,
              "raw_sources": {
                  "top_manifest": relpath(top_manifest_path, repo_root) if top_manifest_path.exists() else None,
                  "acceptance_report": relpath(acceptance_report_path, repo_root) if acceptance_report_path.exists() else None,
                  "builder_manifest": relpath(builder_manifest_path, repo_root) if builder_manifest_path else None,
                  "runner_manifest": relpath(runner_manifest_path, repo_root) if runner_manifest_path else None,
              },
          }
          # Issue #189 + #196: the raw runner manifest dict travels back to the
          # caller as the second tuple element rather than as a smuggled key on
          # the persisted provenance. ``None`` means the runner manifest was
          # missing or unreadable; the renderer treats that as "unavailable".
          raw_runner_manifest = runner_manifest or None
          return provenance, raw_runner_manifest
      
      
      def artifact_sha(provenance: dict[str, Any], artifact_id: str) -> str | None:
          artifact = (provenance.get("artifacts") or {}).get(artifact_id) or {}
          return artifact.get("sha256")
      
      
      def peak_value(provenance: dict[str, Any]) -> Any:
          peak = ((provenance.get("metrics") or {}).get("peak_flow") or {})
          return peak.get("value")
      
      
      def peak_matches_report(provenance: dict[str, Any]) -> Any:
          peak = ((provenance.get("metrics") or {}).get("peak_flow") or {})
          return (peak.get("source_validation") or {}).get("matches_report")
      
      
      def build_comparison(current: dict[str, Any], baseline: dict[str, Any] | None) -> dict[str, Any]:
          if baseline is None:
              return {
                  "schema_version": "1.1",
                  "generated_by": "swmm-experiment-audit",
                  "generated_at_utc": now_utc(),
                  "comparison_available": False,
                  "reason": "No --compare-to run directory was provided.",
                  "current_run_id": current.get("run_id"),
              }
      
          checks: list[dict[str, Any]] = []
      
          def add_check(check_id: str, baseline_value: Any, current_value: Any, interpretation: str) -> None:
              checks.append(
                  {
                      "id": check_id,
                      "baseline": baseline_value,
                      "current": current_value,
                      "same": baseline_value == current_value,
                      "interpretation": interpretation,
                  }
              )
      
          add_check("status", baseline.get("status"), current.get("status"), "Overall audit status.")
          add_check(
              "git_head",
              (baseline.get("repo") or {}).get("git_head"),
              (current.get("repo") or {}).get("git_head"),
              "Code version captured by the run manifest.",
          )
          add_check("model_inp_sha256", artifact_sha(baseline, "model_inp"), artifact_sha(current, "model_inp"), "SWMM input model identity.")
          add_check("runner_out_sha256", artifact_sha(baseline, "runner_out"), artifact_sha(current, "runner_out"), "SWMM binary output identity.")
          add_check("runner_rpt_sha256", artifact_sha(baseline, "runner_rpt"), artifact_sha(current, "runner_rpt"), "SWMM text report identity; may differ because reports include timestamps.")
          add_check("peak_flow", peak_value(baseline), peak_value(current), "Recorded peak-flow metric.")
          add_check(
              "peak_source_section",
              (((baseline.get("metrics") or {}).get("peak_flow") or {}).get("source_section")),
              (((current.get("metrics") or {}).get("peak_flow") or {}).get("source_section")),
              "Report section used for peak-flow extraction.",
          )
          add_check(
              "peak_matches_report",
              peak_matches_report(baseline),
              peak_matches_report(current),
              "Whether the recorded peak value matches the value re-parsed from the source report section.",
          )
          add_check(
              "continuity_error",
              (((baseline.get("metrics") or {}).get("continuity_error") or {}).get("values")),
              (((current.get("metrics") or {}).get("continuity_error") or {}).get("values")),
              "Parsed continuity errors.",
          )
      
          warnings: list[str] = []
          for label, provenance in (("baseline", baseline), ("current", current)):
              peak = ((provenance.get("metrics") or {}).get("peak_flow") or {})
              validation = peak.get("source_validation") or {}
              if validation.get("matches_report") is False:
                  warnings.append(f"{label} peak-flow record does not match the value parsed from its source report section.")
      
          if artifact_sha(baseline, "model_inp") == artifact_sha(current, "model_inp") and peak_value(baseline) != peak_value(current):
              warnings.append("Peak flow changed while the SWMM input hash is unchanged; check parser version, metric source, or report records.")
      
          return {
              "schema_version": "1.1",
              "generated_by": "swmm-experiment-audit",
              "generated_at_utc": now_utc(),
              "comparison_available": True,
              "baseline_run_id": baseline.get("run_id"),
              "current_run_id": current.get("run_id"),
              "baseline_run_dir": (baseline.get("run_dir") or {}).get("relative_path"),
              "current_run_dir": (current.get("run_dir") or {}).get("relative_path"),
              "checks": checks,
              "warnings": warnings,
          }
      
      
      def short(value: Any, n: int = 12) -> str:
          if value is None:
              return "n/a"
          text = str(value)
          return text[:n] + "..." if len(text) > n else text
      
      
      def safe_note_name(value: Any) -> str:
          text = re.sub(r"[^A-Za-z0-9._ -]+", "-", str(value or "experiment")).strip(" .-")
          return text or "experiment"
      
      
      def readable_note_name(provenance: dict[str, Any]) -> str:
          name = provenance.get("case_name") or provenance.get("run_id") or "experiment"
          text = safe_note_name(name).replace("_", " ").replace("-", " ")
          text = re.sub(r"\s+", " ", text).strip()
          return text.title() if text else "Experiment"
      
      
      def md_table(headers: list[str], rows: list[list[Any]]) -> str:
          out = ["| " + " | ".join(headers) + " |", "| " + " | ".join("---" for _ in headers) + " |"]
          for row in rows:
              out.append("| " + " | ".join(str(cell) for cell in row) + " |")
          return "\n".join(out)
      
      
      def render_run_results_section(runner_manifest: dict[str, Any] | None) -> str:
          """Render the always-on ``## Run Results`` markdown block (PRD-183).
      
          The section surfaces a small handful of headline numbers from the
          runner ``manifest.json`` so a reader does not have to open the
          JSON to see peak / continuity / runoff. Numbers are rendered
          *verbatim* — whatever the runner wrote, the audit note shows.
      
          Defensive rendering:
      
          - ``runner_manifest`` is ``None`` (file missing or unreadable) ->
            a single ``unavailable`` line, no table; the rest of the note
            still renders.
          - Individual fields missing -> that row reads ``unavailable``.
          - ``metrics.internal_node_peak`` is rendered only when present;
            omitting it must not produce an ``unavailable`` row.
      
          Issue #189 — kept side-by-side with ``## Key Metrics`` (the source-
          of-truth view that re-parses the same numbers against their report
          sections). A leading tagline on each section makes the intent
          explicit; see the in-body ``## Run Results`` / ``## Key Metrics``
          taglines below for the actual wording.
          """
          header_lines = [
              "## Run Results",
              "",
              # Issue #189 tagline: stamps the intent of this section so the
              # reader knows it's the headline view, not the source-validation
              # view that follows in ## Key Metrics.
              "Headline numbers from the runner manifest, rendered verbatim.",
              "",
          ]
          if not isinstance(runner_manifest, dict):
              return "\n".join(
                  header_lines
                  + [
                      "Run Results unavailable: manifest.json could not be read.",
                      "",
                  ]
              )
      
          metrics = runner_manifest.get("metrics") or {}
          if not isinstance(metrics, dict):
              metrics = {}
          peak = metrics.get("peak") or {}
          if not isinstance(peak, dict):
              peak = {}
          continuity = metrics.get("continuity") or {}
          if not isinstance(continuity, dict):
              continuity = {}
          runoff = continuity.get("runoff_quantity") or {}
          if not isinstance(runoff, dict):
              runoff = {}
          flow_routing = continuity.get("flow_routing") or {}
          if not isinstance(flow_routing, dict):
              flow_routing = {}
          internal = metrics.get("internal_node_peak")
          if internal is not None and not isinstance(internal, dict):
              internal = {}
      
          # Status: PASS iff return_code == 0 and there are no error stanzas.
          # The runner manifest does not carry a uniform "error stanzas"
          # convention beyond return_code, so a non-zero (or missing) return
          # code is the only signal we trust here. Anything else is FAIL so
          # the reader is nudged to open the rest of the audit note.
          return_code = runner_manifest.get("return_code")
          if return_code == 0:
              status_cell = "PASS"
          elif return_code is None:
              status_cell = "unavailable"
          else:
              status_cell = "FAIL"
      
          def _peak_cell(payload: dict[str, Any]) -> str:
              node = payload.get("node")
              value = payload.get("peak")
              time = payload.get("time_hhmm")
              if value is None or node is None:
                  return "unavailable"
              time_part = f" at `{time}`" if time is not None else ""
              unit = payload.get("unit") or payload.get("units")
              unit_part = f" {unit}" if unit else " (flow units not recorded)"
              return f"`{value}`{unit_part} at node `{node}`{time_part}"
      
          def _continuity_cell(table: dict[str, Any]) -> str:
              value = table.get("Continuity Error (%)") if table else None
              if value is None:
                  return "unavailable"
              return f"`{value}` %"
      
          def _surface_runoff_cell(table: dict[str, Any]) -> str:
              entry = table.get("Surface Runoff") if table else None
              if not isinstance(entry, dict):
                  return "unavailable"
              col1 = entry.get("col1")
              col2 = entry.get("col2")
              if col1 is None or col2 is None:
                  return "unavailable"
              return f"`{col1}` hectare-m (`{col2}` mm)"
      
          rows: list[list[str]] = [
              ["Status", status_cell],
              ["Peak flow at outfall", _peak_cell(peak)],
              ["Continuity error, runoff quantity", _continuity_cell(runoff)],
              ["Continuity error, flow routing", _continuity_cell(flow_routing)],
              ["Total surface runoff", _surface_runoff_cell(runoff)],
          ]
          if isinstance(internal, dict) and internal:
              rows.append(["Internal node peak", _peak_cell(internal)])
      
          return "\n".join(
              header_lines
              + [
                  md_table(["Field", "Value"], rows),
                  "",
              ]
          )
      
      
      def render_pollutant_loads_section(wq_loads: dict[str, Any] | None) -> str:
          """Render the ``## Pollutant Loads`` markdown block for WQ-enabled runs.
      
          Returns ``""`` for non-WQ runs (zero behavioral change for hydrology-only).
          Surfaces:
          - Continuity error per pollutant (from Runoff Quality Continuity and
            Quality Routing Continuity sections).
          - Top subcatchment washoff loads table.
          - Outfall pollutant load table.
          """
          if not wq_loads or not wq_loads.get("wq_present"):
              return ""
      
          lines: list[str] = ["## Pollutant Loads", ""]
          pollutants = wq_loads.get("pollutants") or []
          lines.append(f"Water quality enabled. Pollutants: {', '.join(pollutants) or 'none'}.")
          lines.append("")
      
          # Continuity errors
          runoff_cont = wq_loads.get("runoff_quality_conti
    • init_obsidian_vault.py 7.9 KB
      #!/usr/bin/env python3
      from __future__ import annotations
      
      import argparse
      import json
      from pathlib import Path
      
      
      DEFAULT_VAULT = Path.home() / "Documents" / "Agentic-SWMM-Obsidian-Vault"
      
      
      FILES = {
          ".obsidian/app.json": json.dumps(
              {
                  "attachmentFolderPath": "30_Evidence_Layer/Attachments",
                  "promptDelete": False,
                  "alwaysUpdateLinks": True,
              },
              indent=2,
          )
          + "\n",
          ".obsidian/appearance.json": json.dumps(
              {
                  "baseFontSize": 16,
                  "translucency": False,
              },
              indent=2,
          )
          + "\n",
          "00_Home/Agentic SWMM Home.md": """---
      project: Agentic SWMM
      type: vault-home
      status: active
      ---
      
      # Agentic SWMM Home
      
      This vault is the local visualization and review layer for Agentic SWMM. It is designed for a first-time Obsidian user: start here, then follow the links into memory, audit evidence, and skill evolution.
      
      ## Start Here
      
      - [[Project Memory]]: stable project memory and evidence boundaries.
      - [[Experiment Audit Index]]: generated run-level audit notes.
      - [[Claim Register]]: claims, pending claims, and limitations.
      - [[Human Review Queue]]: decisions requiring human review.
      - [[Skill Proposal Log]]: proposed skill updates derived from audited evidence.
      
      ## Layer Model
      
      ```text
      Execution Layer
        SWMM scripts, MCP tools, runs, QA, plots
      
      Audit Layer
        experiment_provenance.json, comparison.json, experiment_note.md
      
      Memory Layer
        recurring lessons, known boundaries, claim discipline, next actions
      
      Skill Evolution Layer
        proposed skill edits, human review, benchmark verification
      ```
      
      ## Default Audit Destination
      
      ```text
      ~/Documents/Agentic-SWMM-Obsidian-Vault/20_Audit_Layer/Experiment_Audits
      ```
      """,
          "10_Memory_Layer/Project Memory.md": """---
      project: Agentic SWMM
      type: memory-layer
      status: active
      ---
      
      # Project Memory
      
      This note records stable project memory for local review. It should summarize durable context, not raw chat history.
      
      ## System Identity
      
      Agentic SWMM is a reproducible, auditable, agent-coordinated stormwater modeling workflow. Agents may coordinate steps, but deterministic scripts, SWMM outputs, manifests, QA records, and audit artifacts remain the evidence source.
      
      ## Working Assumptions
      
      - SWMM computation should remain deterministic and inspectable.
      - Agent decisions must be tied to explicit artifacts.
      - Missing drainage-network or calibration evidence must be reported as missing, not invented.
      - Obsidian is the local visualization and review layer, not the canonical machine-readable source.
      
      ## Related Notes
      
      - [[Experiment Audit Index]]
      - [[Claim Register]]
      - [[Validation Boundaries]]
      - [[Skill Proposal Log]]
      """,
          "10_Memory_Layer/Claim Register.md": """---
      project: Agentic SWMM
      type: claim-register
      status: active
      ---
      
      # Claim Register
      
      Use this note to separate evidence-backed claims from ideas, pending claims, and limitations.
      
      ## Evidence-Backed Claims
      
      | Claim | Evidence | Status |
      |---|---|---|
      
      ## Pending Claims
      
      | Claim | Missing Evidence | Next Check |
      |---|---|---|
      
      ## Limitations
      
      - A minimal real-data smoke test is not a full watershed validation.
      - An audit example is evidence-management evidence, not a new hydrologic result.
      - Skill updates should require human review and benchmark verification.
      """,
          "10_Memory_Layer/Human Review Queue.md": """---
      project: Agentic SWMM
      type: human-review-queue
      status: active
      ---
      
      # Human Review Queue
      
      This note collects decisions that should be reviewed before becoming claims, documentation, or skill changes.
      
      | Date | Item | Why Review Is Needed | Status |
      |---|---|---|---|
      """,
          "20_Audit_Layer/Experiment Audit Index.md": """---
      project: Agentic SWMM
      type: experiment-audit-index
      status: active
      ---
      
      # Experiment Audit Index
      
      This is the main local entry point for run-level audit notes. New audit notes generated by `swmm-experiment-audit` should appear under `20_Audit_Layer/Experiment_Audits` and be indexed here.
      
      ## How To Read An Audit
      
      1. Open the Obsidian audit note for the run.
      2. Check status, QA gates, metrics, warnings, and comparison results.
      3. Open the machine-readable `experiment_provenance.json` for exact commands, repo state, inputs, artifacts, and hashes.
      4. Use `comparison.json` when the run is compared against a baseline.
      5. Update [[Claim Register]] only when the evidence supports the claim.
      
      ## Audited Runs
      
      | Run ID | Status | Audit Note | Provenance JSON | Comparison JSON | Last Updated |
      |---|---|---|---|---|---|
      
      ## Related Notes
      
      - [[Project Memory]]
      - [[Claim Register]]
      - [[Validation Boundaries]]
      - [[Artifact Map]]
      """,
          "30_Evidence_Layer/Artifact Map.md": """---
      project: Agentic SWMM
      type: artifact-map
      status: active
      ---
      
      # Artifact Map
      
      This note explains where generated evidence lives.
      
      ## Repo Run Artifacts
      
      ```text
      runs/
      ```
      
      Each audited run should contain:
      
      ```text
      experiment_provenance.json
      comparison.json
      experiment_note.md
      ```
      """,
          "30_Evidence_Layer/Validation Boundaries.md": """---
      project: Agentic SWMM
      type: validation-boundaries
      status: active
      ---
      
      # Validation Boundaries
      
      This note prevents audit evidence from being overstated.
      
      ## Supported
      
      - Acceptance runs can support workflow execution and QA claims.
      - Audit outputs can support provenance and metric-source claims.
      - Real-data minimal runs can support smoke-test claims.
      
      ## Not Yet Supported
      
      - Full watershed validation unless complete network, calibration, and observed-flow evidence exist.
      - Final calibration performance claims without validation metrics.
      - Windows live validation unless tested on Windows.
      """,
          "40_Skill_Evolution/Skill Proposal Log.md": """---
      project: Agentic SWMM
      type: skill-proposal-log
      status: active
      ---
      
      # Skill Proposal Log
      
      This note records proposed skill updates that come from repeated audit findings. A proposal is not accepted until it has human review and benchmark verification.
      
      | Date | Proposed Skill Change | Evidence | Review Status | Verification |
      |---|---|---|---|---|
      """,
          "90_Templates/Experiment Audit Note Template.md": """---
      project: Agentic SWMM
      type: template
      ---
      
      # Experiment Audit - {{run_id}}
      
      ## Executive Summary
      
      ## Run Identity
      
      ## QA Gates
      
      ## Key Metrics
      
      ## Artifact Index
      
      ## Comparison
      
      ## Warnings
      
      ## Evidence Notes
      """,
      }
      
      
      DIRECTORIES = [
          ".obsidian",
          "00_Home",
          "10_Memory_Layer",
          "20_Audit_Layer/Experiment_Audits",
          "30_Evidence_Layer/Attachments",
          "40_Skill_Evolution",
          "90_Templates",
      ]
      
      
      def write_if_missing(path: Path, text: str, overwrite: bool) -> bool:
          path.parent.mkdir(parents=True, exist_ok=True)
          if path.exists() and not overwrite:
              return False
          path.write_text(text, encoding="utf-8")
          return True
      
      
      def parse_args() -> argparse.Namespace:
          parser = argparse.ArgumentParser(description="Initialize the Agentic SWMM Obsidian vault.")
          parser.add_argument(
              "--vault-dir",
              type=Path,
              default=DEFAULT_VAULT,
              help=f"Vault directory to initialize. Defaults to {DEFAULT_VAULT}.",
          )
          parser.add_argument(
              "--overwrite",
              action="store_true",
              help="Overwrite existing template files.",
          )
          return parser.parse_args()
      
      
      def main() -> None:
          args = parse_args()
          vault = args.vault_dir.expanduser().resolve()
      
          for directory in DIRECTORIES:
              (vault / directory).mkdir(parents=True, exist_ok=True)
      
          written = []
          skipped = []
          for relpath, text in FILES.items():
              path = vault / relpath
              if write_if_missing(path, text, args.overwrite):
                  written.append(str(path))
              else:
                  skipped.append(str(path))
      
          print(
              json.dumps(
                  {
                      "ok": True,
                      "vault": str(vault),
                      "written": written,
                      "skipped_existing": skipped,
                      "home_note": str(vault / "00_Home" / "Agentic SWMM Home.md"),
                      "audit_index": str(vault / "20_Audit_Layer" / "Experiment Audit Index.md"),
                  },
                  indent=2,
              )
          )
      
      
      if __name__ == "__main__":
          main()
      
  • SKILL.md 7.9 KB
    ---
    name: swmm-experiment-audit
    description: Consolidate Agentic SWMM run artifacts into auditable provenance, comparison records, and local Obsidian audit notes. Use after any SWMM build/run/QA attempt, successful or failed, when an agent or CLI workflow needs a traceable record of inputs, commands, artifacts, metrics, QA checks, run-to-run differences, and first-user-friendly Obsidian visualization.
    ---
    
    # SWMM Experiment Audit
    
    Part of [Agentic SWMM](https://github.com/Zhonghao1995/agentic-swmm-workflow) — install the project first for the executable toolchain (aiswmm CLI, SWMM solver, MCP servers).
    
    ## What this skill provides
    
    - A standard audit layer for Agentic SWMM runs.
    - Consolidation of dispersed `manifest.json`, QA JSON, logs, metrics, and artifact paths.
    - Machine-readable outputs for reproducibility and review.
    - Obsidian-compatible Markdown notes for human research records.
    - Default local Obsidian export into a clean English audit vault.
    - Automatic update of the Obsidian `Experiment Audit Index`.
    - Optional run-to-run comparison for baseline/scenario or before/after parser validation.
    
    This skill records what happened. It does not run SWMM, build models, invent missing artifacts, or replace module-level validation.
    
    ## When to use this skill
    
    Use this skill after any of these events:
    
    - `swmm-end-to-end` completes successfully.
    - `swmm-end-to-end` stops or fails after producing partial artifacts.
    - A user wants an Obsidian-ready experiment note for an existing run directory.
    - A user wants to compare two run directories.
    - A run needs evidence for reproducibility, metric provenance, QA status, or paper claims.
    
    Do not use this skill as a substitute for `swmm-runner`, `swmm-builder`, or calibration tools. Run the model first, then audit the run directory.
    
    ## Output contract
    
    For every audited run, write these files into the run's `09_audit/` directory unless explicit output paths are provided:
    
    - `experiment_provenance.json` — machine-readable provenance (immutable).
    - `experiment_note.md` — human-readable Obsidian digest.
    - `model_diagnostics.json` — deterministic SWMM screening checks.
    - `comparison.json` — run-to-run comparison (only when `--compare-to` is given).
    
    `experiment_provenance.json` is the machine-readable source for:
    
    - run identity
    - repo state
    - tool versions
    - command trace
    - input hash records
    - artifact index
    - metrics with source artifacts and source tables
    - QA checks
    - detected warnings and limitations
    
    `comparison.json` records differences against another run when `--compare-to` is provided. If no comparison target is provided, it still records that no comparison was requested.
    
    `experiment_note.md` is Obsidian-compatible Markdown with YAML frontmatter. It should stay readable in GitHub as plain Markdown.
    
    The direct script (`audit_run.py`) writes a copy of the audit note to the
    Obsidian vault by default; pass `--no-obsidian` to suppress it. The
    canonical CLI (`aiswmm audit`) does **not** export to Obsidian by default;
    pass `--obsidian` to enable it.
    
    When Obsidian export is active, the note is written into:
    
    ```text
    ~/Documents/Agentic-SWMM-Obsidian-Vault/20_Audit_Layer/Experiment_Audits
    ```
    
    and the index is updated at:
    
    ```text
    ~/Documents/Agentic-SWMM-Obsidian-Vault/20_Audit_Layer/Experiment Audit Index.md
    ```
    
    ## CLI
    
    The canonical CLI is `aiswmm audit`. It wraps the underlying script with
    backup of prior audit files, MOC regeneration (`runs/INDEX.md`), and the
    M2 audit → memory auto-trigger. The direct script path is also supported
    and is the primary option for MCP / Mode-0 calls.
    
    ### Canonical CLI (`aiswmm audit`)
    
    Obsidian export is **opt-in** with `aiswmm audit` — pass `--obsidian` to
    copy the note to the default vault. Without `--obsidian` the note is written
    only into `09_audit/` inside the run directory.
    
    ```bash
    aiswmm audit --run-dir runs/acceptance/latest
    ```
    
    With comparison:
    
    ```bash
    aiswmm audit --run-dir runs/acceptance/codex-check-peakfix \
      --compare-to runs/acceptance/codex-check
    ```
    
    With explicit metadata and Obsidian export:
    
    ```bash
    aiswmm audit --run-dir runs/real-todcreek-minimal \
      --case-name "Tod Creek minimal" \
      --workflow-mode "minimal real-data fallback" \
      --objective "Verify real-data SWMM execution and preserve provenance." \
      --obsidian
    ```
    
    Skip the M2 memory auto-trigger (useful for acceptance / benchmark runs):
    
    ```bash
    aiswmm audit --run-dir runs/acceptance/latest --no-memory
    ```
    
    Run-to-run comparison with the standalone verb:
    
    ```bash
    aiswmm compare --run-a runs/baseline --run-b runs/scenario
    ```
    
    ### Direct script path (`python3 scripts/audit_run.py`)
    
    Use when calling from MCP or when the full `aiswmm` install is unavailable.
    Obsidian export is **on by default** in the script; disable with `--no-obsidian`.
    
    Initialize a first-user Obsidian vault:
    
    ```bash
    python3 skills/swmm-experiment-audit/scripts/init_obsidian_vault.py
    ```
    
    ```bash
    python3 skills/swmm-experiment-audit/scripts/audit_run.py \
      --run-dir runs/acceptance/latest \
      --no-obsidian
    ```
    
    With comparison:
    
    ```bash
    python3 skills/swmm-experiment-audit/scripts/audit_run.py \
      --run-dir runs/acceptance/codex-check-peakfix \
      --compare-to runs/acceptance/codex-check \
      --no-obsidian
    ```
    
    With explicit metadata including `--case-id` (records the case slug in
    `experiment_provenance.json` for cross-session memory recall):
    
    ```bash
    python3 skills/swmm-experiment-audit/scripts/audit_run.py \
      --run-dir runs/real-todcreek-minimal \
      --case-id tod-creek \
      --case-name "Tod Creek minimal" \
      --workflow-mode "minimal real-data fallback" \
      --objective "Verify real-data SWMM execution and preserve provenance." \
      --no-obsidian
    ```
    
    With an Obsidian vault folder:
    
    ```bash
    python3 skills/swmm-experiment-audit/scripts/audit_run.py \
      --run-dir runs/acceptance/latest \
      --obsidian-dir "/path/to/Obsidian/Agentic SWMM/04_Experiments"
    ```
    
    ## Audit rules
    
    - Always preserve relative paths when the artifact is inside the repository.
    - Include absolute paths in JSON only when useful for local traceability.
    - Record SHA256 for existing file artifacts when feasible.
    - Record artifact role, producer, and downstream use.
    - Preserve command return codes, stdout paths, stderr paths, and timings when available.
    - Treat failed or partial runs as auditable. Missing artifacts should be recorded as missing, not invented.
    - Keep metrics tied to source artifacts and source tables.
    - For SWMM peak flow, prefer `Node Inflow Summary` / `Maximum Total Inflow`.
    - Use `Outfall Loading Summary` / `Max Flow` only as fallback for outfalls.
    - Do not extract peak flow from `Node Depth Summary`; that table reports depth and HGL, not flow.
    
    ## Relationship to `swmm-end-to-end`
    
    `swmm-end-to-end` is the executor and orchestrator.
    
    `swmm-experiment-audit` is the recorder and auditor.
    
    The agent should run this audit skill after every build/run/QA attempt, even when the workflow stops early or fails. The audit output should reference whatever artifacts exist in the run directory and clearly mark missing or incomplete evidence.
    
    ## Obsidian support
    
    The generated `experiment_note.md` is designed for Obsidian:
    
    - YAML frontmatter
    - stable headings
    - tables for QA, metrics, and artifact index
    - relative paths for vault portability
    - no chat transcript or conversational content
    
    The note is also valid GitHub Markdown, so it can be committed as an example or exported as supplementary evidence if desired.
    
    The default local vault is:
    
    ```text
    ~/Documents/Agentic-SWMM-Obsidian-Vault
    ```
    
    It is organized for first-time Obsidian use:
    
    ```text
    00_Home/
    10_Memory_Layer/
    20_Audit_Layer/
    30_Evidence_Layer/
    40_Skill_Evolution/
    90_Templates/
    ```
    
    The direct script always writes canonical outputs into the run directory and
    copies the note to the Obsidian vault unless `--no-obsidian` is given. The
    canonical CLI (`aiswmm audit`) writes canonical outputs only; add `--obsidian`
    to also copy to the vault.
    
    In both paths, use `--obsidian-dir` and `--obsidian-index` to target a vault
    location other than the default `~/Documents/Agentic-SWMM-Obsidian-Vault`.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related