fresh-eyes-sweep
Imported from paulrberg/agent-skills/skills/fresh-eyes-sweep.
Install
npx skills add https://github.com/PaulRBerg/agent-skills/tree/main/skills/fresh-eyes-sweep
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install paulrberg-agent-skills@llmmart
git clone https://github.com/PaulRBerg/agent-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole paulrberg/agent-skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Fresh Eyes Sweep
If these instructions are already present in the conversation from a slash or dollar invocation, follow them directly; do not invoke this skill again through a skill tool.
Inspect the requested Git scope for evidenced mistakes, fix every safe issue, and account for every mapped file. A verified no-op requires that coverage and a full pass that finds nothing new. Leave sound work unchanged; edits are not required to demonstrate a successful sweep.
--max-runtime DURATION is optional: it is a positive integer followed by m or h, such as 45m or 3h. Reject an
invalid duration, unknown option, or ambiguous positional input. When a deadline is supplied, calculate it before
auditing and reserve the final 15% for aggregate validation and reporting, clamped to 5–30 minutes and never exceeding
the total runtime. At that window, settle in-flight slices and do not start new fixes; report an incomplete sweep with
its ledger rather than overrunning the deadline.
Overnight Autonomy
At invocation, read the environment's local time. If it is strictly after 22:00 or strictly before 08:00, treat the entire run as autonomous even if it later crosses a boundary. Exactly 22:00 and 08:00 are outside this window.
During an autonomous overnight run:
- Do not ask the user any questions or pause for clarification, selection, or approval. This does not broaden the skill's authority: leave destructive, disclosure, purchase, public-contract, and other approval-dependent actions undone.
- Use the smallest safe reversible interpretation and continue all independent work. Put every ambiguity, blocked issue, and approval-dependent choice on an overnight backlog instead of interrupting the run.
- Present the backlog at the end with each item's evidence, safe disposition, impact, and decision needed. Phrase the entries as findings, not questions. Omit the section when the backlog is empty.
Ledger Interface
Resolve scripts/sweep-ledger.py from this SKILL.md. Create the scratch ledger outside the repository:
uv run "<skill-dir>/scripts/sweep-ledger.py" init \
--root <repo> --ledger <scratch.json> [<path>...]
With no paths, init maps the whole repository. With paths, it maps exactly those Git scopes. It records every tracked
and non-ignored untracked file plus each path's pre-existing worktree status. The helper does not classify generated,
vendored, binary, safe, important, or defective files.
Record an agent decision atomically only after inspecting or otherwise accounting for the path:
uv run "<skill-dir>/scripts/sweep-ledger.py" mark \
--ledger <scratch.json> --status <pending|inspected|fixed|reported|excluded> \
--path <path> [--path <path>...] [--reason <text>]
excluded requires an agent-written reason. Unknown paths or invalid batches fail without a partial update. Concurrent
mark calls serialize on a sidecar <scratch.json>.lock, so parallel subagents may mark their own paths.
uv run "<skill-dir>/scripts/sweep-ledger.py" pending --ledger <scratch.json> [--limit <n>]
uv run "<skill-dir>/scripts/sweep-ledger.py" summary --ledger <scratch.json>
pending returns the next unaccounted paths in stable order. summary returns exact status counts, pre-existing edit
count, completeness, percentage inputs, and a ten-cell bar. Use those facts directly; never estimate progress or
reimplement ledger arithmetic.
Setup
- Require Git and read applicable repository instructions. Record the worktree root, starting commit, starting status, resolved scope, and any deadline and validation window.
- Initialize the ledger for the requested scope. The agent may additionally inspect shared configuration and
instructions needed to understand that scope; do not silently widen the ledger. If
initmaps more than roughly 2,000 files and the user gave no[paths], partition the mapped ledger into bounded, system-aware directory or subsystem slices and continue without asking solely because of file count. Keep the complete requested scope in the ledger and preserve cross-slice invariants through the system map and aggregate validation. When a supplied deadline cannot cover every slice, stop at its validation window and report the resumable frontier; ask only when no safe partition can preserve a material invariant and the user must choose a narrower outcome. During an autonomous overnight run, record that choice in the overnight backlog and complete everything that remains independently safe. - Classify generated, vendored, minified, binary, and bulk-data artifacts. Validate them through their generator, schema, or invariants when line-by-line review is inappropriate, then mark them with the agent's reason.
- Build a compact system map: executable entry points, workspace or package dependency directions, public interfaces, generators and derived artifacts, external and persisted-data seams, and the owner of each material invariant. Trace the highest-risk workflows end to end before choosing slices.
- Inspect recent history and diffs, especially the newest changes, to find affected callers, dependencies, tests, configuration, and docs. Rank slices and fixes by evidenced impact: correctness, data loss, security, and externally exposed personal-data or disclosure risk first; then reliability, maintainability, measured performance, and developer experience. Treat recency as one prioritization signal, never as a substitute for coverage.
- Discover build, test, lint, typecheck, format, and codegen checks.
- Establish a baseline for every safe, relevant check before the first fix. If it is red, prioritize reproducible failures before discretionary work; defer failures that need an unclear or prohibited action while continuing with independently verifiable work.
- Preserve every pre-existing edit recorded by the ledger. Do not revert, absorb, commit, or report it as a finding.
After mapping, report ### 🔎 Sweep mapped — <files> files · <slices> slices · ledger <scratch.json>. Slice count is an
agent organization choice; file count comes from the ledger.
The ledger outlives the session. A later session resumes the same sweep by pointing at the same ledger path instead of
re-running init: pending defines the frontier, and already-accounted paths are not reinspected. Carry the ledger
path into every progress update, and name it again when reporting an incomplete sweep, so the user can hand it to the
next session.
Subagents
- Delegate independent slices when it materially improves coverage or completion time; use the smallest effective team within host limits and the user's delegation preferences.
- When the host supports model selection, choose reviewer and fixer models deliberately for the task; otherwise use the host default.
- Announce the planned fan-out in one line before launching: agent count and the model of each group.
- Cap concurrent reviewers at 4 unless the user raises it, always within the host's available concurrency.
- Record each spawned task ID in the coordinator's slice plan so a later stop request resolves against real IDs. The
ledger's
reasonrecords exclusions only. - Give writing agents stable IDs, dependency waves, exact non-overlapping write scopes, repository constraints, and required completion evidence. In every slice brief, completion evidence must include every discovered strict static gate — typecheck, lint, and format/import order — applicable to the languages in the slice's write scope, scoped as narrowly as the tool permits. Assign each repository-wide gate to one validation owner after its affected slices settle; other agents report that dependency instead of duplicating the run. Assign shared manifests, lockfiles, exports, and integration files to one sequential owner.
- Reconcile every wave before starting dependents. Use a fresh-context verifier when independent scrutiny addresses a concrete risk, such as concurrency, security, or a cross-slice invariant; a routine edit alone does not require one.
- Subagents and workers never commit. The coordinating session commits settled slices serially as checkpoint commits, so only one process touches the Git index.
- A session holds one coordination claim, and each new claim replaces the last. When the repository uses a claim-based
coordinator such as
ai-coord, claim the union of every in-flight writing slice (running agents plus the coordinator's own edits) before launching a writer. Widen to a new union, never to a scope that drops a slice still being written, and release only after every in-flight slice is reconciled and committed. - On lint-staged or other hook failures during a checkpoint commit, follow
$commit's failure-recovery guidance rather than diagnosing index contention here.
Inspect and Fix
Work through coherent slices so implementation, callers, tests, configuration, and documentation stay visible together. For each slice, reason from first principles: identify the intended outcome and required behavior from the user's request and repository evidence. Treat the current implementation as something to justify, not as a requirement. Interrogate it in this order:
- What is unnecessary, overly complicated, or based on weak assumptions? Challenge those assumptions against evidence.
- What can be deleted entirely while preserving required behavior and contracts? Check consumers and invariants before concluding that a piece is unnecessary.
- After removing unnecessary pieces, what remaining logic, interfaces, or workflow can be simplified?
Prefer deleting over simplifying, simplifying over optimizing, and optimizing over automating. Apply confirmed, safe improvements within the requested scope; this ordering does not justify dropping requirements or automating needless work.
Trace important control, data, concurrency, and error paths end to end. Hunt for concrete bugs, omissions, invalid assumptions, unhandled edges, security/reliability failures, inconsistencies, duplication, dead code, stale docs, and needless complexity. Also inspect evidenced problems in performance, dependencies, data formats and extensions, configuration, observability, accessibility, agent context, naming, and directory structure. Style preferences and unverified hunches are not findings.
At applicable external and persisted-data seams, inspect validation; domain precision and units; deterministic ordering and deduplication; idempotency and repeat-run behavior; atomicity and interruption safety; retry and pagination completeness; bounded concurrency, cancellation, and resource cleanup; and secret, log, path, temporary-file, and command safety.
Confirm each issue before editing. Fix the smallest root cause when intent is clear and verification is available. Add a
regression test when it protects a meaningful failure mode absent from existing coverage; do not add tests that merely
mirror reversible prose or configuration edits. Fix verifiable in-scope residual risks rather than reporting them. Mark
reported only for real decisions: intent is ambiguous, a safe fix would change a public contract for consumers outside
the repository, or no verification is available. Give every reported finding a recommended fix and its blast radius.
Do not add speculative features, broad refactors, or cosmetic churn.
Treat source files over 1000 lines and test files over 2000 lines as discovery candidates only. Split a file only when
cohesion, coupling, change risk, or testability establishes a better seam; line count alone is not evidence. When a
confirmed structural issue requires interface or seam redesign, use $codebase-design when available. Centralize the
invariant in its owning module, apply the deletion test to pass-through modules, introduce a seam only where behavior
actually varies, and keep callers and tests on the resulting interface.
Before changing an interface, persisted format, exported name, or path, enumerate and migrate every producer, consumer, schema, fixture, generator, export or manifest, script or recipe, check, configuration reference, and document. Search for the old identifier afterward and account for every intentional remainder. Apply dependency or framework updates, data-format or extension changes, renames, and reorganizations only when the migration is atomic, compatibility is demonstrable, and repository checks can prove it. Do not retain a performance change without a recorded baseline metric and repeatable benchmark.
If an experiment fails its evidence bar, revert only that experiment's attributable edits; never use repository-wide
clean, checkout, or reset commands. After each nontrivial change wave, run $code-polish over that wave's exact changed
file union when available. Otherwise apply the same fixed-scope contract inline: simplify only where comprehension or
defect risk measurably improves, review by severity, fix evidenced defects, and rerun the narrowest proving checks.
On long runs, post updates only after coherent slices settle, using the ledger summary's exact bar and counts. The bar means path accounting, not depth of inspection.
Tests
Review tests as code under the same first-principles questions, aiming for a smaller suite that catches the same or more defects. For each test, identify the required behavior it protects, then:
- Delete tests that are stale or unhelpful: they target removed or renamed behavior kept alive only by mocks or fixtures; cannot fail (no meaningful assertion, asserting a mock's own return value, tautologies); pin incidental implementation details or call sequences no requirement depends on; restate language, framework, or dependency behavior; or are skipped or commented out with no live reason.
- Merge tests that effectively prove the same thing: identical paths differing only in inputs become one table-driven or parameterized test; a narrower test fully subsumed by another test's assertions goes; duplicates across files collapse into the owning suite.
Before deleting or merging, confirm the protected behavior is obsolete or still covered by a named retained test; use
coverage output, or a temporary targeted break of the code, when the overlap is not obvious. A merge keeps every
distinct assertion, input, and diagnosable failure message. Keep regression tests for fixed bugs unless another test
demonstrably covers the same case. Never delete or skip a failing or flaky test to get green: fix the cause or mark it
reported. Run the affected suites before and after, and record the test-count delta.
Comments
Compare every comment with the code, callers, and history it describes. Fix only clear STALE (describes behavior the
code no longer has), ORPHANED (names a missing symbol, path, flag, or concept), MISLEADING (materially suggests
different behavior), or REDUNDANT (narrates self-explanatory code without intent, constraint, or context) comments.
Rewrite when the correct claim is proven; otherwise remove. Never change executable code merely to make a comment true,
and leave useful rationale and imperfect-but-accurate wording alone.
Treat behavior-bearing comments as code: compiler and tool directives (//go:*, build constraints, cgo preambles,
go:embed, @ts-expect-error, lint suppressions, coverage pragmas), license headers, public API docs, and concurrency,
ownership, or safety contracts. Edit them only when the tooling semantics are proven and validated by the relevant
tooling; otherwise mark them reported.
Verify and Report
Run the narrowest check proving each fix, including every discovered typecheck, lint, and format/import-order gate
applicable to its changed files, then aggregate checks scoped to changed files. Reinspect affected paths and repeat
until a pass finds no new evidenced issue. Before declaring completion, revisit the first-principles questions against
the result, including the sweep's own additions; passing checks alone does not justify unnecessary complexity. During a
supplied deadline's validation window, reconcile owned edits and run the aggregate format, lint, type, test, build, and
invariant checks justified by the final changed-file union. Compare final results with the recorded baseline. Audit
coverage, fixes, and checks against tool output before claiming completion. When the sweep pushed commits and the
repository defines CI workflows, such as .github/workflows, watch the pushed head's runs before reporting
(gh run list --commit <sha>, then gh run watch <run-id>, in the background when the host supports it); fix failures
attributable to the sweep and report the CI outcome. When changed code behaves differently by platform and local checks
covered only one, name the unverified platforms as a risk.
Lead with
### ✅ Sweep ledger complete — <accounted>/<mapped> files accounted (<inspected> inspected, <excluded> excluded) only
when helper complete is true; otherwise use ### ⛔ Sweep incomplete. Summarize fixed, reported, excluded, and check
counts, plus deleted and merged tests with the test-count delta and fixed comments when non-zero. Include a compact
Check | Baseline | Final table, changed artifacts and verified fixes, and subagent results. When non-empty, also
include reverted experiments with the failed evidence, each reported finding with its evidence, recommended fix, and
blast radius, and the overnight backlog when applicable. On ### ⛔ Sweep incomplete, name the ledger path so the next
session can resume from pending. Do not dump the scratch ledger's contents, unrelated pre-existing changes, or bulk
data; include task-relevant evidence when it materially supports the report.
In an interactive (non-overnight) run with any reported findings, end with one decision question listing them: fix all
as recommended, pick specific items, or leave them reported. Treat invocation wording that already authorizes fixing
(for example "fix any/all problems" or "I will follow your judgement") as that approval up front and skip the question,
except for destructive actions, external writes, and purchases, which still need explicit confirmation regardless of
invocation wording. On approval, run a fix wave over the approved items under the same sweep rules — confirm, fix,
verify, update the ledger, and rerun $code-polish — then report the updated ledger and check results. During an
autonomous overnight run, skip the question and leave reported findings in the overnight backlog instead.
Completion requires every mapped path accounted for, every finding fixed and verified, fixed by an approved fix wave, or reported with evidence and a pending decision, and every relevant check passing or its failure attributed.
Files (agent-skills)
-
agents
-
openai.yaml 43 B
policy: allow_implicit_invocation: false
-
-
scripts
-
sweep-ledger.py 9.2 KB
#!/usr/bin/env python3 """Maintain the deterministic coverage ledger for a fresh-eyes sweep.""" from __future__ import annotations import argparse import fcntl import json import math import os import subprocess import sys import tempfile from contextlib import contextmanager from pathlib import Path from typing import Any STATUSES = ("pending", "inspected", "fixed", "reported", "excluded") class LedgerError(ValueError): pass def git(root: Path, *args: str) -> bytes: result = subprocess.run(["git", *args], cwd=root, capture_output=True) if result.returncode: raise LedgerError(result.stderr.decode(errors="replace").strip() or f"git {' '.join(args)} failed") return result.stdout def repo_root(path: Path) -> Path: return Path(git(path, "rev-parse", "--show-toplevel").decode().strip()) def nul_paths(data: bytes) -> list[str]: return [value.decode(errors="surrogateescape") for value in data.split(b"\0") if value] def preexisting_status(root: Path) -> dict[str, list[str]]: fields = nul_paths(git(root, "status", "--porcelain=v1", "-z", "--untracked-files=all")) statuses: dict[str, list[str]] = {} index = 0 while index < len(fields): entry = fields[index] if len(entry) < 4: raise LedgerError("unexpected git status record") code, path = entry[:2], entry[3:] statuses.setdefault(path, []).append(code) if "R" in code or "C" in code: index += 1 if index >= len(fields): raise LedgerError("incomplete git rename status record") statuses.setdefault(fields[index], []).append(f"{code}:source") index += 1 return statuses def normalize_scopes(root: Path, values: list[str]) -> list[str]: if not values: return ["."] scopes: list[str] = [] for value in values: candidate = Path(value) absolute = candidate.resolve() if candidate.is_absolute() else (root / candidate).resolve() try: relative = absolute.relative_to(root.resolve()).as_posix() except ValueError as exc: raise LedgerError(f"scope is outside the repository: {value}") from exc scopes.append(relative or ".") return list(dict.fromkeys(scopes)) def in_scope(path: str, scopes: list[str]) -> bool: return any(scope == "." or path == scope or path.startswith(f"{scope.rstrip('/')}/") for scope in scopes) def atomic_write(path: Path, payload: dict[str, Any]) -> None: encoded = (json.dumps(payload, indent=2, ensure_ascii=False) + "\n").encode() descriptor, temporary_name = tempfile.mkstemp(prefix=f".{path.name}.", dir=path.parent) temporary = Path(temporary_name) try: with os.fdopen(descriptor, "wb") as handle: handle.write(encoded) handle.flush() os.fsync(handle.fileno()) if path.exists(): os.chmod(temporary, path.stat().st_mode) os.replace(temporary, path) finally: if temporary.exists(): temporary.unlink() @contextmanager def exclusive(path: Path): """Serialize read-modify-write cycles from concurrent subagents on a sidecar lock.""" lock = path.with_name(f"{path.name}.lock") with lock.open("a") as handle: fcntl.flock(handle, fcntl.LOCK_EX) try: yield finally: fcntl.flock(handle, fcntl.LOCK_UN) def load(path: Path) -> dict[str, Any]: try: payload = json.loads(path.read_text(encoding="utf-8")) except (OSError, json.JSONDecodeError) as exc: raise LedgerError(f"cannot read ledger: {exc}") from exc if payload.get("schemaVersion") != 1 or not isinstance(payload.get("files"), list): raise LedgerError("unsupported ledger schema") paths = set() for item in payload["files"]: if not isinstance(item, dict) or item.get("status") not in STATUSES or not isinstance(item.get("path"), str): raise LedgerError("ledger contains an invalid file record") if item["path"] in paths: raise LedgerError(f"ledger contains a duplicate path: {item['path']}") paths.add(item["path"]) if item["status"] == "excluded" and not item.get("reason"): raise LedgerError(f"excluded path has no reason: {item['path']}") return payload def counts(payload: dict[str, Any]) -> dict[str, int]: result = {status: 0 for status in STATUSES} for item in payload["files"]: result[item["status"]] += 1 result["mapped"] = len(payload["files"]) result["accounted"] = result["mapped"] - result["pending"] return result def summary(payload: dict[str, Any]) -> dict[str, Any]: summary_counts = counts(payload) mapped = summary_counts["mapped"] filled = None if mapped == 0 else min(10, math.floor(10 * summary_counts["accounted"] / mapped + 0.5)) return { "schemaVersion": 1, "revision": payload["revision"], "counts": summary_counts, "complete": summary_counts["pending"] == 0, "progress": None if filled is None else {"filled": filled, "empty": 10 - filled, "bar": "█" * filled + "░" * (10 - filled)}, "preexistingChangedFiles": sum(bool(item.get("preexistingStatus")) for item in payload["files"]), } def init_command(args: argparse.Namespace) -> dict[str, Any]: if args.ledger.exists(): raise LedgerError(f"ledger already exists: {args.ledger}") root = repo_root(args.root.resolve()) scopes = normalize_scopes(root, args.scopes + args.scope) tracked = set(nul_paths(git(root, "ls-files", "-z"))) untracked = set(nul_paths(git(root, "ls-files", "--others", "--exclude-standard", "-z"))) status = preexisting_status(root) paths = sorted(path for path in tracked | untracked if in_scope(path, scopes)) payload = { "schemaVersion": 1, "revision": 0, "repoRoot": str(root), "scopes": scopes, "files": [ { "path": path, "tracked": path in tracked, "preexistingStatus": status.get(path, []), "status": "pending", "reason": None, } for path in paths ], } args.ledger.parent.mkdir(parents=True, exist_ok=True) atomic_write(args.ledger, payload) return {"schemaVersion": 1, "ledger": str(args.ledger), "scopes": scopes, **summary(payload)} def mark_command(args: argparse.Namespace) -> dict[str, Any]: paths = list(dict.fromkeys(args.path)) if not paths: raise LedgerError("mark requires at least one --path") if args.status == "excluded" and not args.reason: raise LedgerError("excluded status requires --reason") with exclusive(args.ledger): payload = load(args.ledger) by_path = {item["path"]: item for item in payload["files"]} missing = [path for path in paths if path not in by_path] if missing: raise LedgerError(f"paths are not in the ledger: {', '.join(missing)}") for path in paths: by_path[path]["status"] = args.status by_path[path]["reason"] = args.reason if args.status == "excluded" else None payload["revision"] += 1 atomic_write(args.ledger, payload) return {"schemaVersion": 1, "updated": paths, "status": args.status, **summary(payload)} def pending_command(args: argparse.Namespace) -> dict[str, Any]: payload = load(args.ledger) paths = [item["path"] for item in payload["files"] if item["status"] == "pending"] if args.limit is not None: paths = paths[: args.limit] return {"schemaVersion": 1, "revision": payload["revision"], "paths": paths, "remaining": counts(payload)["pending"]} def build_parser() -> argparse.ArgumentParser: parser = argparse.ArgumentParser(description=__doc__) subparsers = parser.add_subparsers(dest="command", required=True) init = subparsers.add_parser("init") init.add_argument("scopes", nargs="*") init.add_argument("--scope", action="append", default=[]) init.add_argument("--root", type=Path, default=Path.cwd()) init.add_argument("--ledger", type=Path, required=True) init.set_defaults(handler=init_command) mark = subparsers.add_parser("mark") mark.add_argument("--ledger", type=Path, required=True) mark.add_argument("--status", choices=STATUSES, required=True) mark.add_argument("--path", action="append", default=[]) mark.add_argument("--reason") mark.set_defaults(handler=mark_command) pending = subparsers.add_parser("pending") pending.add_argument("--ledger", type=Path, required=True) pending.add_argument("--limit", type=int) pending.set_defaults(handler=pending_command) show = subparsers.add_parser("summary") show.add_argument("--ledger", type=Path, required=True) show.set_defaults(handler=lambda args: summary(load(args.ledger))) return parser def main() -> int: args = build_parser().parse_args() try: if getattr(args, "limit", 1) is not None and getattr(args, "limit", 1) < 1: raise LedgerError("--limit must be positive") result = args.handler(args) except LedgerError as exc: print(f"ERROR: {exc}", file=sys.stderr) return 64 print(json.dumps(result, indent=2, ensure_ascii=False)) return 0 if __name__ == "__main__": raise SystemExit(main())
-
-
SKILL.md 18.9 KB
--- argument-hint: "[paths] [--max-runtime DURATION]" compatibility: Requires Git and local command and edit access. disable-model-invocation: true name: fresh-eyes-sweep skill-dependencies: - codebase-design - code-polish - commit description: Audit an entire repository with fresh eyes for correctness errors, bugs, omissions, duplication, inconsistencies, stale or duplicate tests, stale comments, and other evidenced mistakes; fix every safe issue and verify the result. --- # Fresh Eyes Sweep If these instructions are already present in the conversation from a slash or dollar invocation, follow them directly; do not invoke this skill again through a skill tool. Inspect the requested Git scope for evidenced mistakes, fix every safe issue, and account for every mapped file. A verified no-op requires that coverage and a full pass that finds nothing new. Leave sound work unchanged; edits are not required to demonstrate a successful sweep. `--max-runtime DURATION` is optional: it is a positive integer followed by `m` or `h`, such as `45m` or `3h`. Reject an invalid duration, unknown option, or ambiguous positional input. When a deadline is supplied, calculate it before auditing and reserve the final 15% for aggregate validation and reporting, clamped to 5–30 minutes and never exceeding the total runtime. At that window, settle in-flight slices and do not start new fixes; report an incomplete sweep with its ledger rather than overrunning the deadline. ## Overnight Autonomy At invocation, read the environment's local time. If it is strictly after 22:00 or strictly before 08:00, treat the entire run as autonomous even if it later crosses a boundary. Exactly 22:00 and 08:00 are outside this window. During an autonomous overnight run: - Do not ask the user any questions or pause for clarification, selection, or approval. This does not broaden the skill's authority: leave destructive, disclosure, purchase, public-contract, and other approval-dependent actions undone. - Use the smallest safe reversible interpretation and continue all independent work. Put every ambiguity, blocked issue, and approval-dependent choice on an overnight backlog instead of interrupting the run. - Present the backlog at the end with each item's evidence, safe disposition, impact, and decision needed. Phrase the entries as findings, not questions. Omit the section when the backlog is empty. ## Ledger Interface Resolve `scripts/sweep-ledger.py` from this `SKILL.md`. Create the scratch ledger outside the repository: ```sh uv run "<skill-dir>/scripts/sweep-ledger.py" init \ --root <repo> --ledger <scratch.json> [<path>...] ``` With no paths, `init` maps the whole repository. With paths, it maps exactly those Git scopes. It records every tracked and non-ignored untracked file plus each path's pre-existing worktree status. The helper does not classify generated, vendored, binary, safe, important, or defective files. Record an agent decision atomically only after inspecting or otherwise accounting for the path: ```sh uv run "<skill-dir>/scripts/sweep-ledger.py" mark \ --ledger <scratch.json> --status <pending|inspected|fixed|reported|excluded> \ --path <path> [--path <path>...] [--reason <text>] ``` `excluded` requires an agent-written reason. Unknown paths or invalid batches fail without a partial update. Concurrent `mark` calls serialize on a sidecar `<scratch.json>.lock`, so parallel subagents may mark their own paths. ```sh uv run "<skill-dir>/scripts/sweep-ledger.py" pending --ledger <scratch.json> [--limit <n>] uv run "<skill-dir>/scripts/sweep-ledger.py" summary --ledger <scratch.json> ``` `pending` returns the next unaccounted paths in stable order. `summary` returns exact status counts, pre-existing edit count, completeness, percentage inputs, and a ten-cell bar. Use those facts directly; never estimate progress or reimplement ledger arithmetic. ## Setup 1. Require Git and read applicable repository instructions. Record the worktree root, starting commit, starting status, resolved scope, and any deadline and validation window. 2. Initialize the ledger for the requested scope. The agent may additionally inspect shared configuration and instructions needed to understand that scope; do not silently widen the ledger. If `init` maps more than roughly 2,000 files and the user gave no `[paths]`, partition the mapped ledger into bounded, system-aware directory or subsystem slices and continue without asking solely because of file count. Keep the complete requested scope in the ledger and preserve cross-slice invariants through the system map and aggregate validation. When a supplied deadline cannot cover every slice, stop at its validation window and report the resumable frontier; ask only when no safe partition can preserve a material invariant and the user must choose a narrower outcome. During an autonomous overnight run, record that choice in the overnight backlog and complete everything that remains independently safe. 3. Classify generated, vendored, minified, binary, and bulk-data artifacts. Validate them through their generator, schema, or invariants when line-by-line review is inappropriate, then mark them with the agent's reason. 4. Build a compact system map: executable entry points, workspace or package dependency directions, public interfaces, generators and derived artifacts, external and persisted-data seams, and the owner of each material invariant. Trace the highest-risk workflows end to end before choosing slices. 5. Inspect recent history and diffs, especially the newest changes, to find affected callers, dependencies, tests, configuration, and docs. Rank slices and fixes by evidenced impact: correctness, data loss, security, and externally exposed personal-data or disclosure risk first; then reliability, maintainability, measured performance, and developer experience. Treat recency as one prioritization signal, never as a substitute for coverage. 6. Discover build, test, lint, typecheck, format, and codegen checks. 7. Establish a baseline for every safe, relevant check before the first fix. If it is red, prioritize reproducible failures before discretionary work; defer failures that need an unclear or prohibited action while continuing with independently verifiable work. 8. Preserve every pre-existing edit recorded by the ledger. Do not revert, absorb, commit, or report it as a finding. After mapping, report `### 🔎 Sweep mapped — <files> files · <slices> slices · ledger <scratch.json>`. Slice count is an agent organization choice; file count comes from the ledger. The ledger outlives the session. A later session resumes the same sweep by pointing at the same ledger path instead of re-running `init`: `pending` defines the frontier, and already-accounted paths are not reinspected. Carry the ledger path into every progress update, and name it again when reporting an incomplete sweep, so the user can hand it to the next session. ## Subagents - Delegate independent slices when it materially improves coverage or completion time; use the smallest effective team within host limits and the user's delegation preferences. - When the host supports model selection, choose reviewer and fixer models deliberately for the task; otherwise use the host default. - Announce the planned fan-out in one line before launching: agent count and the model of each group. - Cap concurrent reviewers at 4 unless the user raises it, always within the host's available concurrency. - Record each spawned task ID in the coordinator's slice plan so a later stop request resolves against real IDs. The ledger's `reason` records exclusions only. - Give writing agents stable IDs, dependency waves, exact non-overlapping write scopes, repository constraints, and required completion evidence. In every slice brief, completion evidence must include every discovered strict static gate — typecheck, lint, and format/import order — applicable to the languages in the slice's write scope, scoped as narrowly as the tool permits. Assign each repository-wide gate to one validation owner after its affected slices settle; other agents report that dependency instead of duplicating the run. Assign shared manifests, lockfiles, exports, and integration files to one sequential owner. - Reconcile every wave before starting dependents. Use a fresh-context verifier when independent scrutiny addresses a concrete risk, such as concurrency, security, or a cross-slice invariant; a routine edit alone does not require one. - Subagents and workers never commit. The coordinating session commits settled slices serially as checkpoint commits, so only one process touches the Git index. - A session holds one coordination claim, and each new claim replaces the last. When the repository uses a claim-based coordinator such as `ai-coord`, claim the union of every in-flight writing slice (running agents plus the coordinator's own edits) before launching a writer. Widen to a new union, never to a scope that drops a slice still being written, and release only after every in-flight slice is reconciled and committed. - On lint-staged or other hook failures during a checkpoint commit, follow `$commit`'s failure-recovery guidance rather than diagnosing index contention here. ## Inspect and Fix Work through coherent slices so implementation, callers, tests, configuration, and documentation stay visible together. For each slice, reason from first principles: identify the intended outcome and required behavior from the user's request and repository evidence. Treat the current implementation as something to justify, not as a requirement. Interrogate it in this order: 1. What is unnecessary, overly complicated, or based on weak assumptions? Challenge those assumptions against evidence. 2. What can be deleted entirely while preserving required behavior and contracts? Check consumers and invariants before concluding that a piece is unnecessary. 3. After removing unnecessary pieces, what remaining logic, interfaces, or workflow can be simplified? Prefer deleting over simplifying, simplifying over optimizing, and optimizing over automating. Apply confirmed, safe improvements within the requested scope; this ordering does not justify dropping requirements or automating needless work. Trace important control, data, concurrency, and error paths end to end. Hunt for concrete bugs, omissions, invalid assumptions, unhandled edges, security/reliability failures, inconsistencies, duplication, dead code, stale docs, and needless complexity. Also inspect evidenced problems in performance, dependencies, data formats and extensions, configuration, observability, accessibility, agent context, naming, and directory structure. Style preferences and unverified hunches are not findings. At applicable external and persisted-data seams, inspect validation; domain precision and units; deterministic ordering and deduplication; idempotency and repeat-run behavior; atomicity and interruption safety; retry and pagination completeness; bounded concurrency, cancellation, and resource cleanup; and secret, log, path, temporary-file, and command safety. Confirm each issue before editing. Fix the smallest root cause when intent is clear and verification is available. Add a regression test when it protects a meaningful failure mode absent from existing coverage; do not add tests that merely mirror reversible prose or configuration edits. Fix verifiable in-scope residual risks rather than reporting them. Mark `reported` only for real decisions: intent is ambiguous, a safe fix would change a public contract for consumers outside the repository, or no verification is available. Give every `reported` finding a recommended fix and its blast radius. Do not add speculative features, broad refactors, or cosmetic churn. Treat source files over 1000 lines and test files over 2000 lines as discovery candidates only. Split a file only when cohesion, coupling, change risk, or testability establishes a better seam; line count alone is not evidence. When a confirmed structural issue requires interface or seam redesign, use `$codebase-design` when available. Centralize the invariant in its owning module, apply the deletion test to pass-through modules, introduce a seam only where behavior actually varies, and keep callers and tests on the resulting interface. Before changing an interface, persisted format, exported name, or path, enumerate and migrate every producer, consumer, schema, fixture, generator, export or manifest, script or recipe, check, configuration reference, and document. Search for the old identifier afterward and account for every intentional remainder. Apply dependency or framework updates, data-format or extension changes, renames, and reorganizations only when the migration is atomic, compatibility is demonstrable, and repository checks can prove it. Do not retain a performance change without a recorded baseline metric and repeatable benchmark. If an experiment fails its evidence bar, revert only that experiment's attributable edits; never use repository-wide clean, checkout, or reset commands. After each nontrivial change wave, run `$code-polish` over that wave's exact changed file union when available. Otherwise apply the same fixed-scope contract inline: simplify only where comprehension or defect risk measurably improves, review by severity, fix evidenced defects, and rerun the narrowest proving checks. On long runs, post updates only after coherent slices settle, using the ledger summary's exact bar and counts. The bar means path accounting, not depth of inspection. ### Tests Review tests as code under the same first-principles questions, aiming for a smaller suite that catches the same or more defects. For each test, identify the required behavior it protects, then: - Delete tests that are stale or unhelpful: they target removed or renamed behavior kept alive only by mocks or fixtures; cannot fail (no meaningful assertion, asserting a mock's own return value, tautologies); pin incidental implementation details or call sequences no requirement depends on; restate language, framework, or dependency behavior; or are skipped or commented out with no live reason. - Merge tests that effectively prove the same thing: identical paths differing only in inputs become one table-driven or parameterized test; a narrower test fully subsumed by another test's assertions goes; duplicates across files collapse into the owning suite. Before deleting or merging, confirm the protected behavior is obsolete or still covered by a named retained test; use coverage output, or a temporary targeted break of the code, when the overlap is not obvious. A merge keeps every distinct assertion, input, and diagnosable failure message. Keep regression tests for fixed bugs unless another test demonstrably covers the same case. Never delete or skip a failing or flaky test to get green: fix the cause or mark it `reported`. Run the affected suites before and after, and record the test-count delta. ### Comments Compare every comment with the code, callers, and history it describes. Fix only clear `STALE` (describes behavior the code no longer has), `ORPHANED` (names a missing symbol, path, flag, or concept), `MISLEADING` (materially suggests different behavior), or `REDUNDANT` (narrates self-explanatory code without intent, constraint, or context) comments. Rewrite when the correct claim is proven; otherwise remove. Never change executable code merely to make a comment true, and leave useful rationale and imperfect-but-accurate wording alone. Treat behavior-bearing comments as code: compiler and tool directives (`//go:*`, build constraints, cgo preambles, `go:embed`, `@ts-expect-error`, lint suppressions, coverage pragmas), license headers, public API docs, and concurrency, ownership, or safety contracts. Edit them only when the tooling semantics are proven and validated by the relevant tooling; otherwise mark them `reported`. ## Verify and Report Run the narrowest check proving each fix, including every discovered typecheck, lint, and format/import-order gate applicable to its changed files, then aggregate checks scoped to changed files. Reinspect affected paths and repeat until a pass finds no new evidenced issue. Before declaring completion, revisit the first-principles questions against the result, including the sweep's own additions; passing checks alone does not justify unnecessary complexity. During a supplied deadline's validation window, reconcile owned edits and run the aggregate format, lint, type, test, build, and invariant checks justified by the final changed-file union. Compare final results with the recorded baseline. Audit coverage, fixes, and checks against tool output before claiming completion. When the sweep pushed commits and the repository defines CI workflows, such as `.github/workflows`, watch the pushed head's runs before reporting (`gh run list --commit <sha>`, then `gh run watch <run-id>`, in the background when the host supports it); fix failures attributable to the sweep and report the CI outcome. When changed code behaves differently by platform and local checks covered only one, name the unverified platforms as a risk. Lead with `### ✅ Sweep ledger complete — <accounted>/<mapped> files accounted (<inspected> inspected, <excluded> excluded)` only when helper `complete` is true; otherwise use `### ⛔ Sweep incomplete`. Summarize fixed, reported, excluded, and check counts, plus deleted and merged tests with the test-count delta and fixed comments when non-zero. Include a compact `Check | Baseline | Final` table, changed artifacts and verified fixes, and subagent results. When non-empty, also include reverted experiments with the failed evidence, each `reported` finding with its evidence, recommended fix, and blast radius, and the overnight backlog when applicable. On `### ⛔ Sweep incomplete`, name the ledger path so the next session can resume from `pending`. Do not dump the scratch ledger's contents, unrelated pre-existing changes, or bulk data; include task-relevant evidence when it materially supports the report. In an interactive (non-overnight) run with any `reported` findings, end with one decision question listing them: fix all as recommended, pick specific items, or leave them reported. Treat invocation wording that already authorizes fixing (for example "fix any/all problems" or "I will follow your judgement") as that approval up front and skip the question, except for destructive actions, external writes, and purchases, which still need explicit confirmation regardless of invocation wording. On approval, run a fix wave over the approved items under the same sweep rules — confirm, fix, verify, update the ledger, and rerun `$code-polish` — then report the updated ledger and check results. During an autonomous overnight run, skip the question and leave `reported` findings in the overnight backlog instead. Completion requires every mapped path accounted for, every finding fixed and verified, fixed by an approved fix wave, or reported with evidence and a pending decision, and every relevant check passing or its failure attributed.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.