repo-doctor
Audit any repo against the agentic-quality doctrine — score entry docs, structure, and enforcement gates, then map each finding to its fix. Triggers on: repo doctor, repo audit, agentic quality, is this repo agent-friendly, doc drift, stale AGENTS.md, monorepo structure, nested C
Install
npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/repo-doctor
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
git clone https://github.com/0xDarkMatter/claude-mods.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Repo Doctor
Scores a repository against the agentic-quality doctrine — the cross-repo standard in rules/agentic-quality.md for code, comments, docs, and structure that a cold agent session can navigate. The rule says what good looks like; this skill measures a repo against it and maps each gap to its fix.
Read-only. The scorer never writes; remediation is always a separate, deliberate step.
Quick start
python scripts/repo-doctor.py # audit cwd, human panel
python scripts/repo-doctor.py --repo X:/path/to/repo # audit another repo
python scripts/repo-doctor.py --json | jq .data.grade # machine-readable
python scripts/repo-doctor.py --strict # CI gate: exit 10 below B
Six dimensions, 0–5 each, weighted into a letter grade:
| Dimension | Measures | Weight |
|---|---|---|
entry_docs |
AGENTS.md/CLAUDE.md present · Landmines section · length budget (~250 lines) · freshness in commits-since-touched | 2.0 |
docs_health |
README · docs/ index when >6 files · ghost links in the index | 1.5 |
comments |
contract blocks on the largest source files · section markers in files >400 lines | 2.0 |
structure |
monster files (>800 warn, >1500 crit; generated exempt) · repo-root junk | 2.0 |
enforcement |
tests · CI · single check entry point · invariant gate scripts |
1.5 |
doc_pairing |
fraction of recent feat/fix commits touching a *.md in the same commit |
1.0 |
Full rubric — what each check means, thresholds, and the fix for every finding: references/scoring-rubric.md.
Audit workflow
- Run the scorer on the target repo. On cp1252/plain terminals it degrades to ASCII automatically; nothing is written.
- Read findings top-down — they're sorted crit → warn → info. Facts (monster-file
list, pairing ratio, entry-doc age) ride in
--jsonunder.data.facts. - Verify before acting. The scorer is heuristic: a flagged 900-line file may be a justified single-writer module (then it needs the guard comment + section map + gate, not a split); a "stale" AGENTS.md may describe code that genuinely didn't change. Confirm each finding against the repo before proposing work.
- Remediate via the owning skill (below) — repo-doctor diagnoses, it does not operate. Batch fixes into small commits: entry-doc fixes first (highest leverage), then indexes, then guard comments, then splits.
- Re-run to confirm the grade moved. For fleets, loop the scorer over repo roots
with
--jsonand tabulate grades.
Remediation map — who owns each fix
| Finding | Owner |
|---|---|
| Missing/weak AGENTS.md, multi-platform doc mess | doc-scanner (generate/consolidate), template: assets/AGENTS-template.md |
| Missing docs index | Write from assets/docs-index-template.md |
| Monster file needs splitting | refactor-ops (extract-module patterns, circular-dep cautions) |
| Monster file is justified | Guard comment + section map + a scripts/check-* invariant gate (pattern in references/comment-doctrine.md) |
| Missing/weak comments | references/comment-doctrine.md — contract blocks, WHY-only, guard comments, citations |
| Decisions undocumented | adr-ops |
| Stale PLAN/roadmap | project-planner |
| Code-level debt (duplication, dead code, security) | techdebt — deliberately NOT scored here |
| New repo from scratch | scaffold + the two templates in assets/ |
Boundary: repo-doctor audits repo-level conventions; techdebt scans code-level
debt; review/code-review judge diffs. Don't blur the three.
Interpreting the two entry-doc questions
AGENTS.md vs CLAUDE.md — AGENTS.md is the single source of truth (open standard, read by all agent tooling). CLAUDE.md is legitimate only as a pointer or as Claude-specific deltas maintained in lockstep. The scorer flags apparent duplication; the decision tree and the one known-good dual-file pattern are in references/entry-docs.md.
Nested entry docs — nest only where a subsystem has its own contract (design-system package, determinism-bound engine, per-tool CLI); root file carries an ownership table linking each. Anatomy, length budgets, Landmines guidance, freshness discipline: same reference.
Monorepos
For large multi-subsystem repos the audit shifts: the root entry doc is judged as a
router (invariants + ownership table), each contracted package needs its own entry
doc and check, and cross-package invariants need mechanical gates, not prose. The
full playbook — boundaries, navigation aids, extraction signals, parallel-agent
(worktree) interplay, and the split-the-repo decision — is
references/monorepo-structure.md. Run the scorer
per-package as well as at root; a healthy root with a failing core package is the
common monorepo blind spot.
Resources
| Resource | What it owns |
|---|---|
| scripts/repo-doctor.py | The scorer: six dimensions, findings, grade; --json envelope claude-mods.repo-doctor/v1; --strict CI gate |
| references/scoring-rubric.md | Every check: what it measures, threshold, why, and the fix |
| references/comment-doctrine.md | Contract blocks, WHY-only inline, guard comments, section markers, format-at-site, citations — with good/bad examples |
| references/entry-docs.md | AGENTS.md anatomy + Landmines, AGENTS-vs-CLAUDE decision, nesting policy, freshness discipline |
| references/monorepo-structure.md | Structuring very large monorepos for agentic development |
| assets/AGENTS-template.md | Entry-doc skeleton with mandatory Landmines section |
| assets/docs-index-template.md | docs/00_INDEX.md skeleton with the two anti-rot rules baked in |
Files (claude-mods)
-
assets
-
AGENTS-template.md 1.9 KB
# Agent Instructions — <repo-name> <!-- Template: assets/AGENTS-template.md (repo-doctor skill). Budget: ~150 lines, 250 ceiling — every line is a recurring per-session token cost. Evict human setup walkthroughs to README/CONTRIBUTING; keep what agents need every session. --> <2–4 lines: what this repo is, what it produces, who consumes it. Orientation, not marketing.> ## Commands ```bash <run command> # start / serve <test command> # tests <check command> # the ONE gate: typecheck + lint + tests + invariant scripts ``` <!-- Commands must be exact and tested — agents trust these over exploration. If there's no single `check`, create one before filling this in. --> ## Landmines <!-- MANDATORY — the highest-value section. Admission test: would a competent agent plausibly trip this? Each entry: what breaks, why, the procedure. Examples of the genre: "index.html is BAKED — edit template.html, then run build_preview.py"; "growing any content pool invalidates three golden suites — regenerate with GOLDEN_UPDATE=1, procedure in docs/testing.md"; "tests OOM under default node — use NODE_OPTIONS=--max-old-space-size=8192". If you truly have none, write 'None known yet — add the first one the moment it bites.' --> 1. **<landmine>** — <what breaks, why, procedure/link> ## Structure | Path | What lives there | |---|---| | `<dir>/` | <one line> | <!-- Monorepo? This table becomes the ownership table: add Contract + Gate columns and link each own-contract package's nested AGENTS.md (repo-doctor references/monorepo-structure.md §2). --> ## Conventions - <repo-specific deltas ONLY — don't restate global rules> - <invariants: "money is integer cents everywhere", "no Math.random in engine/"> ## Pointers - Docs index: `docs/00_INDEX.md` · Decisions: `docs/adr/` · Design: `docs/<...>` -
docs-index-template.md 1.2 KB
# Docs Index <!-- Template: assets/docs-index-template.md (repo-doctor skill). Save as docs/00_INDEX.md. The two anti-rot rules are baked in below: (1) this file carries its own maintenance instruction; (2) volatile lists are DELEGATED to the filesystem, never hand-copied. One line per doc: what it is, why you'd read it. --> > **Maintenance:** adding/removing a doc in `docs/` updates this index in the same > commit. Keep entries to one line. Never add per-item lists that the filesystem > already answers — delegate them (see Decisions below). ## Canonical | Doc | What it is — why you'd read it | |---|---| | [ARCHITECTURE.md](ARCHITECTURE.md) | <system map — read before structural changes> | | [<doc>.md](<doc>.md) | <one line> | ## Decisions Architecture Decision Records live in [adr/](adr/) — **the directory is the index**: `ls docs/adr/ADR-*.md`. Newest ADR = highest number. (Managed by the `adr-ops` skill.) ## Plans & status | Doc | Liveness | |---|---| | [PLAN.md](PLAN.md) | <what's NEXT — shipped phases marked done; history lives in CHANGELOG/git> | ## Archive Superseded docs move to [archive/](archive/) with a one-line tombstone here only if they're still cited elsewhere; otherwise the move IS the record.
-
-
references
-
comment-doctrine.md 6.5 KB
# Comment Doctrine — inline documentation that survives the session The doctrine in one sentence: **comments carry what the code cannot — constraints, invariants, reasons, and formats — written at the site where a future agent would otherwise guess.** Everything else is noise the next reader pays for. Derived from a 2026-07 audit of 10 production repos: the practices below are the ones that measurably separated the agent-navigable repos (4.5+/5) from the grep-and-pray ones. Each pattern cites the real failure it prevents. --- ## 1. The contract block (top of every substantive file) 3–12 lines answering: what is this module, what invariants does it enforce or rely on, where does the reasoning live. ```python """Event-to-ORM projection dispatch. Two invariants guide the design: * Idempotency. Re-applying the same JSONL line must leave the ORM identical. * Best-effort side-effects. A payload missing a field skips its side-effect, never raises — the journal row is still written. See docs/adr/ADR-007-journal-projection.md for why events, not state sync. """ ``` ```typescript // Push orchestration: validate selection -> reserve (durable intent + locked // records) -> create ACCREC DRAFT in Xero -> finalise. Crash-safe + idempotent. // Token store: tenant's xero connector (AES-GCM at rest). See ADR-006/010. ``` The strongest form states a **law**: `// THE ONLY MODULE THAT TOUCHES D1 — enforced by scripts/check-no-raw-d1.mjs` or `// Determinism: no Math.random, no transcendentals, no Date.now — replay must be byte-identical`. A law + the gate that enforces it is what makes even a 9,000-line file safe to work in. For CLI scripts, the contract block IS the interface contract — Usage / Input / Output / Stderr / Exit codes / Examples (see docs/SKILL-RESOURCE-PROTOCOL.md for the full form). ## 2. WHY-only inline comments A comment must add information the code can't express. Test: delete the comment — if a competent reader loses nothing, it was noise. | Earns its place | Noise | |---|---| | `// order-independent so the same edited set retries to the same key` | `// build the key` | | `# 0.0 means "no prior evidence", NOT "known failure" — scheduler treats them differently` | `# set default to 0.0` | | `// curves wide of the headland; previous loop ran overland and beached a yacht in a farm dam` | `// ferry route points` | | `// budget shared across all skills: 1% of context; overflow silently drops descriptions` | `// check the budget` | State: units, ranges, encodings, ordering guarantees, failure behaviour, and the *reason* a value is what it is. Never narrate control flow. ## 3. Guard comments — protecting deliberate weirdness **The agent-specific failure mode.** An agent that can't see why the code has its strange shape will "improve" it. Real regressions this pattern prevents (all observed): - A 1,300-line single-file HTML app split into ES modules → **`file://` preview broke** (modules need HTTP). The single file was the feature. - A reused sort buffer inlined "for clarity" → per-frame allocation, **GC churn back**. - An unexplained idempotency-key format "tidied" → **retries stopped deduplicating**. The pattern — lead with `ARCHITECTURE:`/`PERF:`/`FORMAT:`, state the constraint, name what breaks: ```javascript /* ARCHITECTURE: classic script, ONE file — not modules, no bundler. file:// preview must work; ES modules fail on file:// (CORS). Splitting this file breaks "open index.html" and the zero-dep contract. */ ``` ```javascript // PERF: _sortBuf is reused across frames — do NOT inline or reallocate; // per-frame allocation regresses GC pauses on 500+ sprite scenes. ``` ```typescript // FORMAT: `${tenantId}|receivable|${clientId}|${sortedRecordIds}|${contentHash}` // recordIds sorted so retries of the same set hash identically; contentHash so // a manual-only push can't collide with a different manual set. ADR-010. ``` Anything that would fail a naive code review for a *reason* needs one of these. Cost: 2–5 lines. The alternative: a future session confidently reintroducing a solved bug. ## 4. Section markers and file maps - **>400 lines** → `// === SECTION ===` markers between logical regions. - **>800 lines** → also a top-of-file map so an agent jumps instead of scrolling: ```javascript /* Sections: STATE · MATH · RENDER · COMMANDS(+history) · INPUT · IO · BOOT All scene mutations funnel through dispatch(type, payload) — extend COMMANDS + makeInverse together when adding mutation types. */ ``` Note what that example does beyond navigation: it states the **extension contract** (COMMANDS + makeInverse move together) exactly where an agent adding a feature lands. ## 5. Formats at the construction site Wire formats, cache keys, hash inputs, file-name grammars, magic strings: document the grammar **where the string is built**, even when an ADR also covers it. An agent editing the construction line will not have the ADR in context; the comment is the last line of defense. One line of grammar + one line of why per component is enough. ## 6. Citations — making domain code modifiable by non-domain agents Embedded domain knowledge gets its source cited inline: standards (`AS 1743 guide-sign green`), papers (`Potrace 2003 §2.2 — penalty-DP optimal polygon`), external IDs (`OSM way 6261378 — the real ferry route`), RFCs, spec sections. A citation converts "magic constants an agent must not touch" into "parameters an agent can verify and adjust." ## 7. Docstrings / JSDoc Public API gets a docstring when the signature under-specifies: side effects, error behaviour, units, ordering/nullability guarantees, "who calls this and when". Private helpers with honest names need nothing. Never write ceremony docstrings that restate the name — they train readers to skip all docstrings. ## 8. What NOT to write - WHAT-narration (`// increment the counter`) - Reviewer-directed remarks (`// fixed per feedback`) — that's PR conversation, not code - Commented-out code — git remembers; delete it - Bare `TODO` — either `TODO(owner/context): action — see #issue` or nothing - Changelog-in-comments (`// 2026-07: changed X`) — git owns history ## Retrofit order (existing repo) 1. Guard comments on everything deliberately weird — highest value per line, prevents active regressions. 2. Contract blocks on the 10 largest / most-touched files. 3. Section markers + maps in >400/>800-line files. 4. Format comments at key/hash/wire construction sites. 5. Citations in domain-heavy modules. Run `scripts/repo-doctor.py` before and after — `comments` dimension should move first. -
entry-docs.md 5.9 KB
# Entry Docs — AGENTS.md anatomy, the CLAUDE.md question, nesting, freshness The entry doc is the highest-leverage file in a repo for agentic development: every session reads it, so every line is either recurring value or recurring token tax. This reference owns the anatomy and the two recurring debates (AGENTS vs CLAUDE, nested entry docs), with verdicts grounded in a 2026-07 audit of 37 repos. --- ## AGENTS.md anatomy (the template is assets/AGENTS-template.md) Order matters — put what agents need most often first: 1. **What this repo is** — 2–4 lines. Not marketing; orientation. 2. **Commands** — run / test / `check` / deploy, exact and tested. A wrong command in an entry doc costs more than no command (agents trust it over exploration). 3. **Landmines** — MANDATORY, the highest-value section. The things that break *non-obviously*: coupled golden fixtures ("growing any pool invalidates three test suites — regeneration procedure below"), generated files ("index.html is baked — edit template.html"), ordering-sensitive registries, env quirks (OOM flags, path conversions). Rule of admission: *would a competent agent plausibly trip this?* Each landmine: what breaks, why, the procedure. 4. **Structure map** — folder → what lives there, one line each. For monorepos this becomes the ownership table (see monorepo-structure.md). 5. **Conventions** — the repo-specific deltas only (style, naming, commit scope). Global conventions live in global rules; don't re-state them. 6. **Pointers** — docs index, ADR dir, design docs. Link, don't inline. **Length budget: ~150 lines target, 250 hard ceiling** (the scorer warns above 250). What gets evicted first: human setup walkthroughs (→ README/CONTRIBUTING), procedural tutorials (→ docs/), rationale essays (→ ADRs). The audit's sharpest contrast: a 269-line entry doc where every line was load-bearing (landmines, pinned commands, terminology canon) scored 5/5; a 350-line one mixing agent hazards with human setup guides scored 4/5 and cost every session the difference. ## AGENTS.md vs CLAUDE.md — the verdict Survey data (37 active repos, 2026-07): 431 AGENTS.md vs 42 CLAUDE.md, and most CLAUDE.md files were inside cloned third-party repos. AGENTS.md won; it's the open standard read by all agent tooling (Claude Code, Codex, Cursor, …). **Default: one AGENTS.md, no CLAUDE.md.** Claude Code reads AGENTS.md fine. **Acceptable dual-file pattern** (one known-good production example): AGENTS.md holds *universal invariants* (money rules, scoping, test discipline — stable, audience- agnostic); CLAUDE.md holds *Claude-specific operational deltas* (implementation-stack guardrails, crash-recovery mechanics — volatile, evolves with the work). Three conditions make it work, and all three are required: 1. **Deltas only** — CLAUDE.md never restates AGENTS.md content. 2. **Lockstep maintenance** — a change to a shared rule updates both in one commit. 3. **Both stay under budget** — two lean files, not two bloated ones. If you can't commit to all three, don't split. A CLAUDE.md that's a one-line pointer to AGENTS.md is always fine for tool compatibility. ## Nested entry docs — when subfolder AGENTS.md/CLAUDE.md earns its place Community advice says "nest CLAUDE.md everywhere." The drive-wide evidence says otherwise: every blanket-nested example found was a cloned third-party repo, and the only *home-grown* nested entry doc that worked was a subsystem with a genuinely distinct contract (a portable design-system package with its own tokens-only law, gallery, and lint gate). **Nest when the subsystem has its own contract**, meaning at least one of: - Its own invariant law (determinism engine, tokens-only design system, single-writer data layer) - Its own audience (a package consumed by other repos; a per-tool CLI in a tool farm) - Its own gate (`lint:design`, per-package `check`) **Rules for a nested entry doc:** - Deltas + local landmines ONLY — never duplicate root rules (duplication rots into contradiction, and agents can't tell which file wins). - Root AGENTS.md links every nested one in its ownership table — nesting without the router breaks discoverability. - Same length discipline, smaller budget (~60 lines). **Do not nest** to restate root conventions per-folder, to shorten paths in prompts, or because a folder is merely large. Large-but-contractless folders need a structure-map line in root, not their own doc. Note the tooling asymmetry: Claude Code auto-loads nested CLAUDE.md on demand when working in that subtree — that's the *mechanism*; the own-contract test above is the *policy* for when to use it. ## Freshness — commits, not days Measure entry-doc staleness in **commits since the doc was last touched**, never wall-clock. Audit evidence: repos with week-old mtimes hid 100+ commits of drift; the mtime looked fine because an unrelated edit touched the file. - Healthy: entry doc touched within ~15 commits (scorer threshold). - The discipline that keeps it healthy: any commit that invalidates an entry-doc claim updates the doc in the same commit — "done includes the doc touch" (rules/agentic-quality.md non-negotiable #3). - Automatable check: `git rev-list --count HEAD ^$(git rev-list -1 HEAD -- AGENTS.md)` ## README.md relationship README = first-time human (what/why, quickstart, badges, doc index). AGENTS.md = recurring agent (commands, landmines, structure). Distinct audiences, minimal overlap; each links the other. If they've converged into near-duplicates, the README is usually the one that's drifted — trim it to narrative + links. ## Generation and consolidation Missing or fragmented entry docs (CLAUDE.md + COPILOT.md + CURSOR.md …) → the `doc-scanner` skill scans, synthesizes, and consolidates into one AGENTS.md. Seed new repos from assets/AGENTS-template.md — the Landmines section ships with prompts so it can't be skipped silently. -
monorepo-structure.md 8.2 KB
# Monorepo Structure — organising very large repos for agentic development A monorepo is where agentic-quality economics bite hardest: entry-doc token cost multiplies across sessions, navigability failures multiply across subsystems, and parallel agent sessions multiply write-collision risk. This reference is the playbook, distilled from the four large multi-subsystem repos in the 2026-07 audit (a pnpm app/packages monorepo at 4.8/5, a multi-app Workers platform at 3.5/5, a Python orchestration platform at 4.4/5, and a 116-tool CLI farm at 3.0/5) — what the strong ones did that the weak ones didn't. --- ## 1. The organising principle: subsystem = contract boundary Structure the repo so every top-level unit is a **contract boundary** — a thing with its own invariants, its own tests, and a one-line answer to "what may depend on me." Not "frontend/backend", not file-kind buckets (`utils/`, `helpers/`), but: ``` apps/ deployable things (web SPA, worker API, admin) packages/ shared contracts (engine, content, design-system, sim harness) docs/ repo-level docs + index; per-package docs live IN the package scripts/ invariant gates + repo tooling (check-*.mjs pattern) migrations/ append-only, numbered ``` The test an agent applies constantly: *can I guess the path from the concept?* `packages/engine/src/rng/` and `src/xero/` pass; `src/utils2/` and `lib/misc/` fail. Every failed guess is a fan-out search an agent runs instead of one file read. ## 2. Root AGENTS.md is a router, not an encyclopedia At monorepo scale the root entry doc changes job: repo-wide invariants + an **ownership table**, nothing else. Target well under 150 lines — it's read every session regardless of which package the session touches. ```markdown ## Ownership | Subsystem | Path | Contract | Gate | |---|---|---|---| | Engine (deterministic sim) | packages/engine/ | packages/engine/AGENTS.md | engine golden tests | | Design system | packages/halcyon/ | packages/halcyon/AGENTS.md + DESIGN.md | npm run lint:design | | Data layer | src/db/ | "only module touching D1" (contract block) | scripts/check-no-raw-d1.mjs | | Xero integration | src/xero/ | ADR-006/010 | xero-saga tests | ``` Root keeps: the ownership table, cross-cutting invariants (money, auth, determinism), the `check` command, repo-wide landmines. Everything package-specific moves DOWN into that package's entry doc. The audit's weakest monorepo entry docs were encyclopedias that tried to hold every subsystem's rules at root — every session paid for all of them, and package-local changes routinely forgot to update the far-away root doc. ## 3. Nested entry docs: own-contract packages only Full policy in [entry-docs.md](entry-docs.md); the monorepo application: - Every `packages/*` with its own invariant law, audience, or gate → its own AGENTS.md (deltas + local landmines, ~60-line budget). - `apps/*` usually DON'T need one — they consume contracts, they don't define them; a structure-map line at root suffices. - A tool-farm (100+ sibling dirs of the same shape) needs a **protocol doc once** + per-tool docs only where a tool has real per-tool knowledge (auth quirks, API gotchas). Template-stamped boilerplate docs across 100 dirs are negative value — they bury the ones with real content. ## 4. Mechanical gates make big shared code safe The single strongest pattern found in the audit: **a 30-line check script converts a prose rule into a physical property of the repo.** - `check-no-raw-d1.mjs` — "only repo.ts touches the database" → a 9,560-line data layer stays *safe* (every agent knows exactly where all queries live) even while it stays unpleasant. - `lint:design` — "style only from tokens; no package→app imports" → a design-system package survives dozens of agent sessions without palette drift. - Golden/replay tests — "engine output is byte-identical across refactors" → agents refactor the engine fearlessly. Directive: **every cross-package invariant in the ownership table has a gate.** A rule without a gate is a request; agents (and humans) will eventually violate it silently. Wire all gates into ONE `check` entry point (`npm run check`, `just check`) that fans out to per-package checks — if verification isn't one command, it doesn't happen. ## 5. Navigation aids that scale past folder-guessing When the repo outgrows guessability, add explicit maps — in this order of cost: 1. **docs index** (`docs/00_INDEX.md`) once docs/ passes ~6 files. Two anti-rot rules (from the one index that stayed accurate across a large project): the index carries its own maintenance instruction, and volatile lists are delegated to the filesystem ("ADRs: see `ls docs/adr/`") instead of hand-copied. 2. **Section maps inside justified monster files** — the 9,000-line single-writer file needs a top-of-file TOC more than any other file in the repo. 3. **A function/route map** (`docs/repository-map.md`) for a monster module that can't be split yet — cheaper than the split, buys most of the navigation back. 4. **A machine-readable registry** for farms of same-shaped units (tools, connectors, generators): one `REGISTRY.json` with name/category/description/status per unit. The audit's 116-tool farm without one forced every discovery into a filesystem walk or a shell-out; one JSON read replaces both. ## 6. File-size discipline under monorepo gravity Big repos grow monster files faster (more contributors-per-file, more "just add it here"). The escalation ladder (same as rules/agentic-quality.md, applied per package): - ~400 lines → section markers - ~800 lines → split along responsibility seams (refactor-ops), OR write the guard comment justifying why not - deliberately-large files → guard comment + section map + the gate that enforces the invariant justifying them (a monster WITHOUT a gate is just debt with a story) Splitting priorities when a package's core file bloats: split by *lifecycle* first (read paths vs write paths), then by *entity*. Resist `utils.ts` as a split target — it's where guessability goes to die. ## 7. Parallel agents in one monorepo Monorepos concentrate agent traffic; collisions are structural, not behavioural. The rules (full doctrine: rules/worktree-boundaries.md, fleet-ops skill): - **One writer per checkout.** Parallel sessions get worktrees; the main checkout is landing-only. - **Partition by subsystem**, not by file list — the ownership table doubles as the lane map ("this session owns packages/engine, that one owns apps/web"). - **Land early, land often** — long-lived divergence across a shared `packages/` layer is the monorepo-specific failure; cross-package refactors belong in short, dedicated lanes that land before feature lanes rebase. - Per-package `check` gates keep landing cheap: a lane that only touched `packages/content` runs that package's suite, not 20 minutes of everything. ## 8. Extraction: when a subsystem should leave the monorepo The audited ecosystem's own rule, proven twice: **when a subsystem grows a roadmap, an asset library, or an audience of its own, it's a product — extract it** before it distorts the host repo's docs, tests, and entry-doc budget. Signals: - Its docs start explaining things no other subsystem cares about - Its issues/plans track independently of repo releases - Outside consumers appear (another repo vendors or clones it) - Its assets dwarf the host (models, sample libraries, fixtures) Extraction hygiene: the new repo gets its own AGENTS.md + README day one; the host keeps a pointer (skill/package doc → "extracted to <repo>"); no stale paths back. ## 9. Monorepo audit checklist (repo-doctor lens) Run `scripts/repo-doctor.py` at root AND per major package, then verify by hand: - [ ] Root AGENTS.md is a router (<150 lines) with an ownership table - [ ] Every own-contract package has its nested entry doc, linked from the table - [ ] Every ownership-table invariant has a named mechanical gate - [ ] One `check` command fans out to per-package checks - [ ] docs/ has an index; volatile lists delegated to filesystem - [ ] No package's core file is a gateless monster - [ ] Farms of same-shaped units have a machine-readable registry - [ ] Nothing in the repo is a stealth product overdue for extraction -
scoring-rubric.md 5.8 KB
# Scoring Rubric — what repo-doctor measures, thresholds, and the fix per finding Every check in `scripts/repo-doctor.py`, its threshold, why that threshold, and the remediation. Grade bands: A ≥ 4.5 · B ≥ 3.5 · C ≥ 2.5 · D ≥ 1.5 · F below. `--strict` exits 10 below B — that's the CI posture ("healthy or explain"). The scorer is heuristic and read-only. **Always verify a finding against the repo before proposing work** — a flagged monster file may be justified (then the fix is the guard-comment + gate route, not a split), and a "stale" entry doc may describe code that genuinely didn't change (then touch it in the next honest commit, don't churn it). --- ## entry_docs (weight 2.0) | Check | Threshold | Why | Fix | |---|---|---|---| | AGENTS.md or CLAUDE.md exists | crit if absent (score 0) | Agents enter blind; every session re-derives the repo | Generate via `doc-scanner`; seed from assets/AGENTS-template.md | | Landmines section | warn | The highest-value lines in the file; absence means non-obvious breakage is tribal knowledge | Add `## Landmines` — what breaks, why, procedure (entry-docs.md §anatomy) | | Length ≤ 250 lines | warn above | Recurring per-session token tax; bloat = human walkthroughs in the agent file | Evict setup guides to README/docs, keep deltas + landmines | | Touched within 15 commits | warn above | Commit-lag is the real staleness metric; mtime lies (audit: 100+-commit drift behind week-old mtimes) | Verify claims vs code; update in the same commit as the fix | | CLAUDE.md duplicates AGENTS.md | warn | Duplicates diverge into contradictions | Reduce to pointer or deltas-only (entry-docs.md §verdict) | ## docs_health (weight 1.5) | Check | Threshold | Why | Fix | |---|---|---|---| | README.md exists | warn | Humans (and some tools) enter here | Short narrative + quickstart + doc links | | docs/ index when > 6 md files | warn | Un-indexed doc dirs make agents hunt; the audit's index-less 16-file docs/ cost every session a search | assets/docs-index-template.md; delegate volatile lists to `ls` | | Index links resolve | warn per ghost | Dead links teach agents to distrust the index | Fix or remove; index carries its own maintenance note | | Docs missing from index | info | Coverage drift signal | Add one-line entries | No docs/ dir at all: not penalised beyond a cap (small repos legitimately keep everything in README + AGENTS.md). ## comments (weight 2.0) Sampled over the N largest source files (default 12, `--sample`), ≥100 lines each. | Check | Threshold | Why | Fix | |---|---|---|---| | Contract block | first 15 lines carry ≥3 comment/docstring lines | The cheapest orientation an agent gets; ratio drives the score | comment-doctrine.md §1 — what/invariants/refs | | Section markers in >400-line files | info per file | Jumping beats scrolling; markers are how agents Grep-navigate | `// === SECTION ===` per region; >800 also gets a top map | Not measured (deliberately): comment *density* — high density of WHAT-noise scores worse in practice than sparse WHY comments. The doctrine reference owns quality; the scorer only checks the two mechanically-checkable proxies. ## structure (weight 2.0) | Check | Threshold | Why | Fix | |---|---|---|---| | Monster files | warn ≥800, crit ≥1500 (generated files exempt via header detection) | #1 agent friction in 3 of 4 audited clusters; a 9,560-line file is grep-only territory | Split (refactor-ops), OR guard comment + section map + invariant gate (the justified-monster route) | | Justified monster | guard comment plus at least three `=== NAME ===` comment markers → info; only one signal → warn naming the missing half | The scorer honours the doctrine's own escape hatch — a justified file shouldn't nag forever; the info reminds you to verify the mechanical gate still holds | Keep the marker honest: if the gate is gone, the justification is a lie — split or restore the gate | | Repo-root junk | warn; media >1 MB or scratch-pattern names | Root is the first thing every agent lists; junk destroys signal | docs/screenshots/, dev/, scratchpad, or delete | ## enforcement (weight 1.5) | Check | Points | Why | Fix | |---|---|---|---| | Tests present | 1.5 | Without tests every agent edit is a hope | testing-ops / testgen; name adversarial tests for the adversary | | CI workflows | 1.0 | Local-only gates skip on exactly the sessions that forget | Minimal workflow running `check` | | Single `check` entry | 1.5 | If verification isn't one command, agents won't run it | `npm run check` / `just check` fanning out typecheck+lint+tests+gates | | Invariant gate scripts | 1.0 | Prose rules rot; 30-line scripts don't (the lint:db lesson) | `scripts/check-<invariant>.mjs` per ownership-table rule | ## doc_pairing (weight 1.0) Fraction of the last 60 non-merge `feat|fix|refactor|perf` commits that touch any `*.md` in the same commit. 50% pairing scores 5 (not every feature invalidates a doc — demanding 100% would reward doc-churn theatre). Below 15% → warn: docs are drifting as a matter of process, not accident. Fix is workflow, not writing: "done includes the doc touch" (rules/agentic-quality.md non-negotiable #3). --- ## Reading the output - **Findings are sorted crit → warn → info**; facts (`--json` → `.data.facts`) carry the raw numbers (monster list, pairing ratio, entry-doc age, gate inventory). - **Fix order for a C/D repo**: entry doc (biggest single lever) → docs index → guard comments on justified weirdness → `check` entry point → splits. Re-run after each batch; `entry_docs` and `docs_health` move immediately, `comments` moves with the retrofit, `doc_pairing` only moves with sustained workflow change. - **Fleet sweep**: loop `--json` over repo roots, tabulate `.data.grade` — the audit found grade variance within one developer's repos is a convention-enforcement signal, not a skill signal.
-
-
scripts
-
repo-doctor.py 24.2 KB
#!/usr/bin/env python3 """repo-doctor — score any repo against the agentic-quality doctrine. Usage: repo-doctor.py [--repo PATH] [--json] [--strict] [--top N] [--sample N] Input: a git repository (defaults to cwd); no network, no writes — read-only audit Output: TTY panel report on stdout by default; --json emits ONLY a JSON envelope {"data": {...}, "meta": {"schema": "claude-mods.repo-doctor/v1"}} on stdout with per-dimension scores (0-5), letter grade, and a findings list, each finding carrying {dim, severity, msg, path}. Severity: crit | warn | info. Stderr: progress/warnings only (never data) Exit: 0 report produced (grade >= B, or --strict not set and repo readable), 10 --strict and grade below B (CI gate), 2 usage error, 3 not a git repo Dimensions scored (0-5 each, weighted into the grade): entry_docs AGENTS.md/CLAUDE.md presence, Landmines section, length budget, freshness measured in commits-since-touched (never days) docs_health README, docs/ index when >6 files, ghost/missing index entries comments contract blocks on the largest source files; section markers in files >400 lines structure monster files (>800 warn, >1500 crit; generated/vendored exempt), repo-root junk (media, scratch artifacts) enforcement tests present, CI workflows, single `check` entry point, invariant gate scripts (the lint:db pattern) doc_pairing fraction of recent feat/fix commits that touch a *.md in the same commit — the as-you-go signature Examples: repo-doctor.py # audit the cwd, human panel repo-doctor.py --repo D:/code/other-repo # audit another repo repo-doctor.py --json | jq .data.grade # machine-readable repo-doctor.py --strict # exit 10 if grade < B (CI gate) repo-doctor.py --top 15 # show 15 findings instead of 10 Doctrine: rules/agentic-quality.md. Rubric + fixes per finding: references/scoring-rubric.md (sibling of this script's skill). """ from __future__ import annotations import argparse import json import os import re import subprocess import sys from pathlib import Path SCHEMA = "claude-mods.repo-doctor/v1" SOURCE_EXTS = { ".py", ".ts", ".tsx", ".js", ".jsx", ".mjs", ".cjs", ".go", ".rs", ".rb", ".php", ".java", ".cs", ".c", ".cc", ".cpp", ".h", ".hpp", ".sh", ".ps1", ".sql", ".vue", ".svelte", ".html", } EXCLUDE_DIRS = { ".git", "node_modules", ".venv", "venv", "dist", "build", "out", "vendor", "__pycache__", ".next", ".nuxt", "coverage", "target", ".claude", } MEDIA_EXTS = {".png", ".jpg", ".jpeg", ".gif", ".mp4", ".mov", ".webm", ".zip", ".7z", ".tar", ".gz", ".glb", ".psd"} SCRATCH_PAT = re.compile(r"^(tmp|temp|scratch|junk|old|copy|test)[-_]", re.I) GENERATED_PAT = re.compile(r"generated|do not edit|don'?t hand-edit|@generated", re.I) # Deliberateness signals for justified large files. Extend this list when the # doctrine adopts another unambiguous way to say that a file must stay whole. # Signals require INTENT wording — bare nouns like "one file"/"single file" # match innocent prose ("reads one file per call") and wrongly bless real # monsters (adversarial-review finding, 2026-07). LARGE_FILE_GUARD_SIGNALS = ( "do not split", "don't split", "never split", "must not be split", "deliberately single-file", "deliberately single file", "deliberately one file", ) LARGE_FILE_GUARD_PAT = re.compile( "|".join(re.escape(signal) for signal in LARGE_FILE_GUARD_SIGNALS), re.I) # Keep marker syntax in lockstep with the comments dim's SECTION_PAT or the # two dims contradict each other on the same body: accept any >=3-char run of # = or ═ around the name, not only the exact ===. LARGE_FILE_SECTION_PAT = re.compile( r"^\s*(?:#|//|;|<!--)\s*[=═]{3,}\s+\w[\w .-]*\s+[=═]{3,}", re.M) BOXED_SECTION_PAT = re.compile( r"(?m)^\s*#\s*=+\s*$\n^\s*#\s+\w[^\n]*$\n^\s*#\s*=+\s*$") LANDMINE_PAT = re.compile(r"landmine|gotcha|pitfall|footgun|hazard|trap|don'?t", re.I) COMMENT_LINE = re.compile(r"^\s*(#|//|/\*|\*|<!--|--|\"\"\"|''')") SECTION_PAT = re.compile(r"^\s*(#|//|/\*|<!--|--)\s*.{0,8}([=─-]{4,}|SECTION|═{4,})") FEATFIX_PAT = re.compile(r"^(feat|fix|refactor|perf)\b", re.I) MONSTER_WARN, MONSTER_CRIT = 800, 1500 ENTRY_LEAN_LINES = 250 FRESH_COMMITS = 15 DOCS_INDEX_THRESHOLD = 6 WEIGHTS = {"entry_docs": 2.0, "docs_health": 1.5, "comments": 2.0, "structure": 2.0, "enforcement": 1.5, "doc_pairing": 1.0} def eecho(msg: str) -> None: print(msg, file=sys.stderr) def git(repo: Path, *args: str) -> str: try: r = subprocess.run(["git", "-C", str(repo), *args], capture_output=True, text=True, timeout=30, encoding="utf-8", errors="replace") return r.stdout.strip() if r.returncode == 0 else "" except (OSError, subprocess.TimeoutExpired): return "" def commits_since_touch(repo: Path, rel: str) -> int | None: """How many commits landed after this file was last touched. None = untracked.""" last = git(repo, "rev-list", "-1", "HEAD", "--", rel) if not last: return None n = git(repo, "rev-list", "--count", "HEAD", f"^{last}") return int(n) if n.isdigit() else None def large_file_signals(path: Path) -> tuple[bool, bool]: """Return (guard comment, section map) for a deliberately large file.""" try: body = path.read_text(encoding="utf-8", errors="replace") except OSError: return False, False # Phrases only count inside comments/docstrings, never executable strings. comment_regions = re.findall( r"<!--.*?-->|/\*.*?\*/|'''[\s\S]*?'''|\"\"\"[\s\S]*?\"\"\"|" r"(?m:^\s*(?:#|//|;).*?$)", body, flags=re.S, ) has_guard = any(LARGE_FILE_GUARD_PAT.search(region) for region in comment_regions) marker_count = (len(LARGE_FILE_SECTION_PAT.findall(body)) + len(BOXED_SECTION_PAT.findall(body))) has_map = marker_count >= 3 return has_guard, has_map def iter_source_files(repo: Path): for root, dirs, files in os.walk(repo): dirs[:] = [d for d in dirs if d not in EXCLUDE_DIRS and not d.startswith(".")] for f in files: p = Path(root) / f if p.suffix.lower() in SOURCE_EXTS and ".min." not in f: yield p def head_lines(p: Path, n: int) -> list[str]: try: with p.open(encoding="utf-8", errors="replace") as fh: return [next(fh, "") for _ in range(n)] except OSError: return [] def count_lines(p: Path) -> int: try: with p.open("rb") as fh: return sum(1 for _ in fh) except OSError: return 0 def is_generated(p: Path) -> bool: return any(GENERATED_PAT.search(l) for l in head_lines(p, 3)) class Audit: def __init__(self, repo: Path, sample: int): self.repo = repo self.sample = sample self.findings: list[dict] = [] self.scores: dict[str, float] = {} self.facts: dict = {} def add(self, dim: str, severity: str, msg: str, path: str = "") -> None: self.findings.append({"dim": dim, "severity": severity, "msg": msg, "path": path}) # -- dimension: entry docs ------------------------------------------------ def check_entry_docs(self) -> None: score = 0.0 agents = self.repo / "AGENTS.md" claude = self.repo / "CLAUDE.md" entry = agents if agents.exists() else (claude if claude.exists() else None) self.facts["entry_doc"] = entry.name if entry else None if not entry: self.add("entry_docs", "crit", "no AGENTS.md or CLAUDE.md — agents enter blind; " "generate one (doc-scanner skill) with a Landmines section") self.scores["entry_docs"] = 0 return score += 2 text = entry.read_text(encoding="utf-8", errors="replace") lines = text.count("\n") + 1 self.facts["entry_doc_lines"] = lines if LANDMINE_PAT.search(text): score += 1 else: self.add("entry_docs", "warn", f"{entry.name} has no Landmines/gotchas section — the " "highest-value lines for agents are missing", entry.name) if lines <= ENTRY_LEAN_LINES: score += 1 else: self.add("entry_docs", "warn", f"{entry.name} is {lines} lines (budget ~{ENTRY_LEAN_LINES}) — " "agents pay this token cost every session; push walkthroughs " "into docs/ and link them", entry.name) since = commits_since_touch(self.repo, entry.name) self.facts["entry_doc_commits_since"] = since if since is not None and since <= FRESH_COMMITS: score += 1 elif since is not None: self.add("entry_docs", "warn", f"{entry.name} last touched {since} commits ago — verify it " "still describes the code, then touch it in the fixing commit", entry.name) if agents.exists() and claude.exists(): ctext = claude.read_text(encoding="utf-8", errors="replace") if len(ctext) > 400 and ctext[:400] in text: self.add("entry_docs", "warn", "CLAUDE.md appears to duplicate AGENTS.md — keep deltas " "only, or reduce CLAUDE.md to a pointer", "CLAUDE.md") else: self.add("entry_docs", "info", "both AGENTS.md and CLAUDE.md present — fine if CLAUDE.md " "is deltas-only (Gather pattern); shared rule changes must " "update both in one commit") self.scores["entry_docs"] = score # -- dimension: docs health ----------------------------------------------- def check_docs_health(self) -> None: score = 5.0 if not (self.repo / "README.md").exists(): score -= 1 self.add("docs_health", "warn", "no README.md — humans enter blind") docs = self.repo / "docs" if not docs.is_dir(): self.facts["docs_md_count"] = 0 self.scores["docs_health"] = min(score, 4.0) # small repos: fine, capped return md = sorted(p for p in docs.rglob("*.md") if not any(part in EXCLUDE_DIRS for part in p.parts)) self.facts["docs_md_count"] = len(md) index = next((docs / n for n in ("00_INDEX.md", "INDEX.md", "PLAN.md") if (docs / n).exists()), None) if len(md) > DOCS_INDEX_THRESHOLD and not index: score -= 2 self.add("docs_health", "warn", f"docs/ has {len(md)} files and no index " "(00_INDEX.md / INDEX.md / PLAN.md) — agents hunt instead of " "navigate; add one with a maintenance note", "docs/") if index: itext = index.read_text(encoding="utf-8", errors="replace") # Ghosts: only actual markdown LINK targets count — bare filename # mentions in prose/checklists ("removed DASH.md") are history, not # references. Mentions still count as "indexed" for the missing check. linked = {t.split("#")[0].split("/")[-1] for t in re.findall(r"\]\(([^)#\s]+\.md)[^)]*\)", itext) if not t.startswith("http")} mentioned = set(re.findall(r"([A-Za-z0-9_\-]+\.md)", itext)) ghosts = [t for t in linked if t != index.name and not list(docs.rglob(t)) and not (self.repo / t).exists()] missing = [p.name for p in md if p.name not in mentioned and p != index and p.parent == docs] if ghosts: score -= 1 self.add("docs_health", "warn", f"{index.name} references missing files: " f"{', '.join(ghosts[:5])}", f"docs/{index.name}") if len(missing) > 2: self.add("docs_health", "info", f"{len(missing)} docs/*.md not mentioned in {index.name}: " f"{', '.join(missing[:5])}…", f"docs/{index.name}") self.scores["docs_health"] = max(score, 0) # -- dimension: comments ---------------------------------------------------- def check_comments(self) -> None: sized = sorted(((count_lines(p), p) for p in iter_source_files(self.repo)), reverse=True) top = [(n, p) for n, p in sized[: self.sample] if n >= 100] self.facts["source_files"] = len(sized) if not top: self.scores["comments"] = 5 return contract_ok = markers_ok = markers_needed = 0 for n, p in top: rel = str(p.relative_to(self.repo)).replace("\\", "/") head = head_lines(p, 15) # A line counts as comment if it matches a comment prefix OR sits # inside an open block comment: Python triple-quoted docstrings and # PowerShell <# ... #> blocks carry prose lines with no per-line # prefix, but they ARE the contract block. in_doc = False n_comment = 0 for l in head: quotes = l.count('"""') + l.count("'''") opens = quotes % 2 == 1 or ("<#" in l and "#>" not in l) closes = "#>" in l and "<#" not in l if in_doc or COMMENT_LINE.match(l) or quotes or "<#" in l: n_comment += 1 if in_doc and (closes or quotes % 2 == 1): in_doc = False elif not in_doc and opens: in_doc = True has_contract = n_comment >= 3 if has_contract: contract_ok += 1 else: self.add("comments", "warn", f"no contract block: first 15 lines of a {n}-line file " "carry <3 comment lines — open with what/invariants/refs", rel) if n > 400 and not is_generated(p): markers_needed += 1 try: body = p.read_text(encoding="utf-8", errors="replace") except OSError: body = "" if any(SECTION_PAT.match(l) for l in body.splitlines()): markers_ok += 1 else: self.add("comments", "info", f"{n}-line file has no section markers — add " "`// === SECTION ===` so agents jump, not scroll", rel) score = 5 * (contract_ok / len(top)) if markers_needed: score = score * 0.7 + 5 * (markers_ok / markers_needed) * 0.3 self.facts["contract_block_ratio"] = round(contract_ok / len(top), 2) self.scores["comments"] = round(score, 1) # -- dimension: structure --------------------------------------------------- def check_structure(self) -> None: score = 5.0 monsters = [] for p in iter_source_files(self.repo): n = count_lines(p) if n >= MONSTER_WARN and not is_generated(p): monsters.append((n, str(p.relative_to(self.repo)).replace("\\", "/"))) monsters.sort(reverse=True) self.facts["monster_files"] = monsters[:10] penalties = [] for n, rel in monsters[:10]: has_guard, has_map = large_file_signals(self.repo / rel) if has_guard and has_map: self.add("structure", "info", f"{n}-line large file with guard comment + section map — " "verify its mechanical gate exists", rel) continue if has_guard or has_map: missing = "section map" if has_guard else "guard comment" self.add("structure", "warn", f"{n}-line large file has only half its justification — " f"missing {missing}", rel) penalties.append(0.5) continue penalties.append(1.0 if n >= MONSTER_CRIT else 0.5) if n >= MONSTER_CRIT: self.add("structure", "crit", f"{n}-line file — split by responsibility (refactor-ops) " "or justify with a guard comment + section map + a " "mechanical gate for the invariant that keeps it whole", rel) else: self.add("structure", "warn", f"{n}-line file — needs section markers now, a split " "or a written justification before it grows", rel) score -= min(3.0, sum(penalties)) junk = [] for p in self.repo.iterdir(): if p.is_file(): if (p.suffix.lower() in MEDIA_EXTS and p.stat().st_size > 1_000_000) \ or SCRATCH_PAT.match(p.name): junk.append(p.name) if junk: score -= min(2.0, 0.5 * len(junk)) self.add("structure", "warn", f"repo-root junk ({len(junk)}): {', '.join(junk[:6])} — move " "to docs/screenshots/, dev/, or delete") self.scores["structure"] = max(score, 0) # -- dimension: enforcement --------------------------------------------------- def check_enforcement(self) -> None: score = 0.0 has_tests = any((self.repo / d).is_dir() for d in ("tests", "test")) or \ bool(list(self.repo.glob("**/*.test.*"))[:1]) or \ bool(list(self.repo.glob("**/test_*.py"))[:1]) if has_tests: score += 1.5 else: self.add("enforcement", "warn", "no tests found — agents fly blind") if (self.repo / ".github" / "workflows").is_dir(): score += 1 else: self.add("enforcement", "info", "no CI workflows (.github/workflows)") check_entry = False pkg = self.repo / "package.json" if pkg.exists(): try: check_entry = "check" in json.loads( pkg.read_text(encoding="utf-8", errors="replace") ).get("scripts", {}) except (json.JSONDecodeError, OSError): pass for jf in ("justfile", "Justfile", "Makefile"): f = self.repo / jf if f.exists() and re.search(r"^check\s*:", f.read_text( encoding="utf-8", errors="replace"), re.M): check_entry = True if check_entry: score += 1.5 else: self.add("enforcement", "warn", "no single `check` entry point (npm run check / just check) — " "if it isn't one command, agents won't run it") gates = [p.name for p in (self.repo / "scripts").glob("check-*")] \ if (self.repo / "scripts").is_dir() else [] gates += [p.name for p in (self.repo / "tests").glob("*drift*")] \ if (self.repo / "tests").is_dir() else [] if gates: score += 1 self.facts["invariant_gates"] = gates[:5] else: self.add("enforcement", "info", "no invariant gate scripts (scripts/check-*.mjs pattern) — " "prose rules rot; 30-line scripts don't") self.scores["enforcement"] = min(score, 5) # -- dimension: doc pairing -------------------------------------------------- def check_doc_pairing(self) -> None: raw = git(self.repo, "log", "--no-merges", "-n", "60", "--pretty=%x01%s", "--name-only") if not raw: self.scores["doc_pairing"] = 2.5 return featfix = paired = 0 for block in raw.split("\x01"): blines = [l for l in block.strip().splitlines() if l.strip()] if not blines or not FEATFIX_PAT.match(blines[0]): continue featfix += 1 if any(l.strip().lower().endswith(".md") for l in blines[1:]): paired += 1 if featfix == 0: self.scores["doc_pairing"] = 2.5 return ratio = paired / featfix self.facts["doc_pairing_ratio"] = round(ratio, 2) self.facts["featfix_commits_sampled"] = featfix self.scores["doc_pairing"] = round(min(5, ratio * 10), 1) # 50% pairing = 5 if ratio < 0.15: self.add("doc_pairing", "warn", f"only {paired}/{featfix} recent feat/fix commits touched any " "*.md — docs are drifting; pair the doc touch with the feature") # -- roll-up ------------------------------------------------------------------- def run(self) -> dict: self.check_entry_docs() self.check_docs_health() self.check_comments() self.check_structure() self.check_enforcement() self.check_doc_pairing() total_w = sum(WEIGHTS.values()) overall = sum(self.scores[k] * WEIGHTS[k] for k in WEIGHTS) / total_w grade = ("A" if overall >= 4.5 else "B" if overall >= 3.5 else "C" if overall >= 2.5 else "D" if overall >= 1.5 else "F") sev_rank = {"crit": 0, "warn": 1, "info": 2} self.findings.sort(key=lambda f: sev_rank[f["severity"]]) return {"repo": str(self.repo), "grade": grade, "overall": round(overall, 2), "scores": self.scores, "facts": self.facts, "findings": self.findings} # ---- rendering ------------------------------------------------------------------ def render_panel(result: dict, top: int) -> None: enc = (getattr(sys.stdout, "encoding", "") or "").lower() uni = "utf" in enc or "cp65001" in enc bar_on, bar_off = ("█", "░") if uni else ("#", ".") rail = "│" if uni else "|" print(f"repo-doctor {rail} {result['repo']}") print(f"grade: {result['grade']} ({result['overall']}/5)") print() for dim, sc in result["scores"].items(): filled = round(sc) print(f" {dim:<12} {bar_on * filled}{bar_off * (5 - filled)} {sc}/5") shown = result["findings"][:top] if shown: print() for f in shown: loc = f" [{f['path']}]" if f["path"] else "" print(f" {f['severity'].upper():<5} {f['msg']}{loc}") hidden = len(result["findings"]) - len(shown) if hidden > 0: print(f" … {hidden} more (use --top {len(result['findings'])} or --json)") print() print("rubric + per-finding fixes: references/scoring-rubric.md (repo-doctor skill)") def main() -> int: ap = argparse.ArgumentParser( description="Score a repo against the agentic-quality doctrine " "(rules/agentic-quality.md). Read-only.", epilog="EXAMPLES:\n" " repo-doctor.py\n" " repo-doctor.py --repo D:/code/other-repo --json\n" " repo-doctor.py --strict # CI gate: exit 10 below grade B\n", formatter_class=argparse.RawDescriptionHelpFormatter) ap.add_argument("--repo", default=".", help="repo path (default: cwd)") ap.add_argument("--json", action="store_true", help="emit JSON envelope on stdout (no panel)") ap.add_argument("--strict", action="store_true", help="exit 10 if grade is below B") ap.add_argument("--top", type=int, default=10, help="findings to show in the panel (default 10)") ap.add_argument("--sample", type=int, default=12, help="largest source files sampled for comment checks") args = ap.parse_args() repo = Path(args.repo).resolve() if not repo.is_dir(): eecho(f"repo-doctor: not a directory: {repo}") return 2 if not (repo / ".git").exists(): eecho(f"repo-doctor: not a git repo: {repo} (init or pass --repo)") return 3 result = Audit(repo, args.sample).run() if args.json: print(json.dumps({"data": result, "meta": {"schema": SCHEMA}}, indent=2)) else: render_panel(result, args.top) if args.strict and result["grade"] not in ("A", "B"): return 10 return 0 if __name__ == "__main__": sys.exit(main())
-
-
tests
-
run.sh 8.5 KB
#!/usr/bin/env bash # repo-doctor behavioural suite — offline, self-contained. # # Builds a throwaway git repo seeded with known defects (no entry doc → then a # weak one, monster file, un-indexed docs/, root junk, no tests/check), runs # scripts/repo-doctor.py --json against it, and asserts each defect is detected # and the envelope schema holds. Then seeds the fixes and asserts the grade # improves and findings clear. Exit 0 = all pass, 1 = failure, 0 + skip message # when python3/git are unavailable. set -u HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" DOCTOR="$HERE/../scripts/repo-doctor.py" # Windows: `python3` on PATH is often the Microsoft Store shim, which prints an # install nag and exits nonzero. A candidate only counts if it actually RUNS. PY="" for cand in python3 python py; do if command -v "$cand" >/dev/null 2>&1 \ && "$cand" -c "import sys" >/dev/null 2>&1; then PY="$cand"; break fi done if [ -z "$PY" ] || ! command -v git >/dev/null 2>&1; then echo "SKIP: working python/git not available" exit 0 fi TMP="$(mktemp -d)" trap 'rm -rf "$TMP"' EXIT pass=0; fail=0 ok() { echo " PASS $1"; pass=$((pass+1)); } no() { echo " FAIL $1 — $2"; fail=$((fail+1)); } # ---- fixture: a repo with seeded defects ------------------------------------ cd "$TMP" || exit 1 git init -q -b main . git config user.email t@t.local; git config user.name t mkdir -p src docs # monster file: 1600 lines, no contract block, no section markers "$PY" -c "print('\n'.join('x = %d' % i for i in range(1600)))" > src/monster.py # 8 docs, no index for i in 1 2 3 4 5 6 7 8; do echo "# doc $i" > "docs/d$i.md"; done # root junk: scratch-pattern file + fake >1MB media echo x > tmp_probe.py "$PY" -c "open('scratch.png','wb').write(b'0'*1100000)" git add -A; git commit -qm "feat: seed" run_json() { "$PY" "$DOCTOR" --repo "$TMP" --json 2>/dev/null; } OUT="$(run_json)" get() { echo "$OUT" | "$PY" -c "import json,sys;d=json.load(sys.stdin);print($1)"; } # 1. envelope schema [ "$(get '"ok" if d["meta"]["schema"]=="claude-mods.repo-doctor/v1" else "no"')" = "ok" ] \ && ok "envelope carries repo-doctor/v1 schema" || no "schema" "$(get 'd["meta"]')" # 2. missing entry doc is a crit, entry_docs scores 0 [ "$(get 'd["data"]["scores"]["entry_docs"]')" = "0" ] \ && ok "missing AGENTS.md scores entry_docs=0" \ || no "missing entry doc" "score=$(get 'd["data"]["scores"]["entry_docs"]')" # 3. monster file detected as crit with path [ "$(get 'sum(1 for f in d["data"]["findings"] if f["severity"]=="crit" and "monster.py" in f["path"])')" = "1" ] \ && ok "1600-line file flagged crit" || no "monster detection" "not found" # 4. un-indexed docs/ flagged [ "$(get 'sum(1 for f in d["data"]["findings"] if "no index" in f["msg"])')" = "1" ] \ && ok "8-file docs/ without index flagged" || no "docs index" "not flagged" # 5. root junk flagged (both scratch pattern and big media) [ "$(get 'sum(1 for f in d["data"]["findings"] if "repo-root junk" in f["msg"])')" = "1" ] \ && ok "root junk flagged" || no "root junk" "not flagged" # 6. no check entry point flagged [ "$(get 'sum(1 for f in d["data"]["findings"] if "check" in f["msg"] and "entry point" in f["msg"])')" = "1" ] \ && ok "missing check entry point flagged" || no "check entry" "not flagged" # 7. strict mode gates: low grade → exit 10 "$PY" "$DOCTOR" --repo "$TMP" --strict >/dev/null 2>&1 [ $? -eq 10 ] && ok "--strict exits 10 below grade B" || no "--strict" "exit $?" # ---- seed the fixes, assert improvement -------------------------------------- cat > AGENTS.md <<'EOF' # Agent Instructions — fixture Tiny fixture repo for the repo-doctor suite. ## Commands just check ## Landmines 1. **monster.py is generated-style filler** — never hand-edit, regenerate via tests. ## Structure | src/ | code | EOF cat > docs/00_INDEX.md <<'EOF' # Docs Index > Maintenance: update in the same commit as any docs/ change. | [d1.md](d1.md) | doc | EOF rm tmp_probe.py scratch.png mkdir -p tests .github/workflows echo "echo ok" > tests/run.sh echo "name: ci" > .github/workflows/ci.yml printf 'check:\n\techo ok\n' > Makefile # generated-marker on the monster: exempts it from the monster penalty { echo "# @generated — filler fixture, do not edit"; cat src/monster.py; } > src/m2 \ && mv src/m2 src/monster.py git add -A; git commit -qm "fix: remediate per repo-doctor + docs: index" OUT="$(run_json)" # 8. entry_docs recovers (presence+landmines+lean; freshness=commit 0 lag) ENT="$(get 'd["data"]["scores"]["entry_docs"]')" "$PY" -c "exit(0 if float('$ENT')>=4 else 1)" \ && ok "remediated entry_docs >= 4 (got $ENT)" || no "entry recovery" "$ENT" # 9. generated marker exempts the monster [ "$(get 'sum(1 for f in d["data"]["findings"] if "monster.py" in f["path"] and f["severity"]=="crit")')" = "0" ] \ && ok "@generated header exempts monster file" || no "generated exempt" "still crit" # 10. grade improved to A/B and --strict passes GR="$(get 'd["data"]["grade"]')" "$PY" "$DOCTOR" --repo "$TMP" --strict >/dev/null 2>&1 \ && ok "remediated repo passes --strict (grade $GR)" || no "strict pass" "grade $GR exit $?" # 11-13. large-file escape hatch requires both guard and section map { echo '"""Deliberately single-file; do not split this portable script."""' printf '# === INPUT ===\n# === PROCESSING ===\n# === OUTPUT ===\n' "$PY" -c "print('\n'.join('y = %d' % i for i in range(1700)))"; } > src/justified.py { printf '# === INPUT ===\n# === PROCESSING ===\n# === OUTPUT ===\n' "$PY" -c "print('\n'.join('z = %d' % i for i in range(1700)))"; } > src/map-only.py "$PY" -c "print('\n'.join('b = %d' % i for i in range(1700)))" > src/bare.py git add -A; git commit -qm "feat: justified monster" OUT="$(run_json)" [ "$(get 'sum(1 for f in d["data"]["findings"] if "justified.py" in f["path"] and f["dim"]=="structure" and f["severity"]=="info")')" = "1" ] \ && ok "guard + map downgrades monster crit -> info" \ || no "justified monster" "severity wrong" [ "$(get 'sum(1 for f in d["data"]["findings"] if "map-only.py" in f["path"] and f["severity"]=="warn" and "guard comment" in f["msg"])')" = "1" ] \ && ok "map-only monster warns for missing guard" \ || no "map-only monster" "severity or message wrong" [ "$(get 'sum(1 for f in d["data"]["findings"] if "bare.py" in f["path"] and f["severity"]=="crit")')" = "1" ] \ && ok "bare monster remains crit" || no "bare monster" "severity wrong" # 14-16. adversarial-review regressions (GLM refuter, 2026-07) # guard-only (no map) -> WARN: discriminates the two-signal branch from the # old guard-alone justification logic, which would have fully excused it. { echo '"""Deliberately single-file; do not split this portable script."""' "$PY" -c "print('\n'.join('g = %d' % i for i in range(1700)))"; } > src/guard-only.py # innocent prose + markers -> must stay CRIT (bare "one file" is not intent) { echo '# reads one file per call and caches the result' printf '# === INPUT ===\n# === PROCESSING ===\n# === OUTPUT ===\n' "$PY" -c "print('\n'.join('p = %d' % i for i in range(1700)))"; } > src/prose-trap.py # 4-equals markers + guard -> INFO (marker-pattern parity with comments dim) { echo '"""Deliberately single-file; do not split."""' printf '# ==== INPUT ====\n# ==== PROCESSING ====\n# ==== OUTPUT ====\n' "$PY" -c "print('\n'.join('w = %d' % i for i in range(1700)))"; } > src/wide-markers.py git add -A; git commit -qm "feat: refuter regression fixtures" OUT="$(run_json)" [ "$(get 'sum(1 for f in d["data"]["findings"] if "guard-only.py" in f["path"] and f["severity"]=="warn" and "map" in f["msg"].lower())')" = "1" ] \ && ok "guard-only monster warns for missing map (pins new branch)" \ || no "guard-only monster" "severity or message wrong" # prose is not a guard -> map-only WARN branch; the defended regression is # "never INFO" (no false blessing from bare 'one file' prose). [ "$(get 'sum(1 for f in d["data"]["findings"] if "prose-trap.py" in f["path"] and f["dim"]=="structure" and f["severity"]=="warn" and "guard" in f["msg"])')" = "1" ] \ && [ "$(get 'sum(1 for f in d["data"]["findings"] if "prose-trap.py" in f["path"] and f["dim"]=="structure" and f["severity"]=="info")')" = "0" ] \ && ok "innocent 'one file' prose does not excuse a monster" \ || no "prose-trap monster" "false downgrade" [ "$(get 'sum(1 for f in d["data"]["findings"] if "wide-markers.py" in f["path"] and f["severity"]=="info")')" = "1" ] \ && ok "4-equals markers count as a section map" \ || no "wide-markers monster" "marker pattern too narrow" echo echo "repo-doctor tests: $pass passed, $fail failed" [ "$fail" -eq 0 ] && exit 0 || exit 1
-
-
SKILL.md 6.5 KB
--- name: repo-doctor description: "Audit any repo against the agentic-quality doctrine — score entry docs, structure, and enforcement gates, then map each finding to its fix. Triggers on: repo doctor, repo audit, agentic quality, is this repo agent-friendly, doc drift, stale AGENTS.md, monorepo structure, nested CLAUDE.md." license: MIT allowed-tools: "Read Bash Glob Grep Agent" metadata: author: claude-mods related-skills: "doc-scanner, adr-ops, refactor-ops, techdebt, scaffold, project-planner" --- # Repo Doctor Scores a repository against the **agentic-quality doctrine** — the cross-repo standard in [rules/agentic-quality.md](../../rules/agentic-quality.md) for code, comments, docs, and structure that a cold agent session can navigate. The rule says *what good looks like*; this skill measures a repo against it and maps each gap to its fix. Read-only. The scorer never writes; remediation is always a separate, deliberate step. ## Quick start ```bash python scripts/repo-doctor.py # audit cwd, human panel python scripts/repo-doctor.py --repo X:/path/to/repo # audit another repo python scripts/repo-doctor.py --json | jq .data.grade # machine-readable python scripts/repo-doctor.py --strict # CI gate: exit 10 below B ``` Six dimensions, 0–5 each, weighted into a letter grade: | Dimension | Measures | Weight | |---|---|---| | `entry_docs` | AGENTS.md/CLAUDE.md present · Landmines section · length budget (~250 lines) · freshness in **commits-since-touched** | 2.0 | | `docs_health` | README · docs/ index when >6 files · ghost links in the index | 1.5 | | `comments` | contract blocks on the largest source files · section markers in files >400 lines | 2.0 | | `structure` | monster files (>800 warn, >1500 crit; generated exempt) · repo-root junk | 2.0 | | `enforcement` | tests · CI · single `check` entry point · invariant gate scripts | 1.5 | | `doc_pairing` | fraction of recent feat/fix commits touching a `*.md` in the same commit | 1.0 | Full rubric — what each check means, thresholds, and the fix for every finding: [references/scoring-rubric.md](references/scoring-rubric.md). ## Audit workflow 1. **Run the scorer** on the target repo. On cp1252/plain terminals it degrades to ASCII automatically; nothing is written. 2. **Read findings top-down** — they're sorted crit → warn → info. Facts (monster-file list, pairing ratio, entry-doc age) ride in `--json` under `.data.facts`. 3. **Verify before acting.** The scorer is heuristic: a flagged 900-line file may be a justified single-writer module (then it needs the guard comment + section map + gate, not a split); a "stale" AGENTS.md may describe code that genuinely didn't change. Confirm each finding against the repo before proposing work. 4. **Remediate via the owning skill** (below) — repo-doctor diagnoses, it does not operate. Batch fixes into small commits: entry-doc fixes first (highest leverage), then indexes, then guard comments, then splits. 5. **Re-run to confirm** the grade moved. For fleets, loop the scorer over repo roots with `--json` and tabulate grades. ## Remediation map — who owns each fix | Finding | Owner | |---|---| | Missing/weak AGENTS.md, multi-platform doc mess | `doc-scanner` (generate/consolidate), template: [assets/AGENTS-template.md](assets/AGENTS-template.md) | | Missing docs index | Write from [assets/docs-index-template.md](assets/docs-index-template.md) | | Monster file needs splitting | `refactor-ops` (extract-module patterns, circular-dep cautions) | | Monster file is *justified* | Guard comment + section map + a `scripts/check-*` invariant gate (pattern in [references/comment-doctrine.md](references/comment-doctrine.md)) | | Missing/weak comments | [references/comment-doctrine.md](references/comment-doctrine.md) — contract blocks, WHY-only, guard comments, citations | | Decisions undocumented | `adr-ops` | | Stale PLAN/roadmap | `project-planner` | | Code-level debt (duplication, dead code, security) | `techdebt` — deliberately NOT scored here | | New repo from scratch | `scaffold` + the two templates in assets/ | Boundary: repo-doctor audits **repo-level conventions**; `techdebt` scans **code-level debt**; `review`/code-review judge **diffs**. Don't blur the three. ## Interpreting the two entry-doc questions **AGENTS.md vs CLAUDE.md** — AGENTS.md is the single source of truth (open standard, read by all agent tooling). CLAUDE.md is legitimate only as a pointer or as Claude-specific *deltas* maintained in lockstep. The scorer flags apparent duplication; the decision tree and the one known-good dual-file pattern are in [references/entry-docs.md](references/entry-docs.md). **Nested entry docs** — nest only where a subsystem has its own contract (design-system package, determinism-bound engine, per-tool CLI); root file carries an ownership table linking each. Anatomy, length budgets, Landmines guidance, freshness discipline: same reference. ## Monorepos For large multi-subsystem repos the audit shifts: the root entry doc is judged as a *router* (invariants + ownership table), each contracted package needs its own entry doc and `check`, and cross-package invariants need mechanical gates, not prose. The full playbook — boundaries, navigation aids, extraction signals, parallel-agent (worktree) interplay, and the split-the-repo decision — is [references/monorepo-structure.md](references/monorepo-structure.md). Run the scorer per-package as well as at root; a healthy root with a failing core package is the common monorepo blind spot. ## Resources | Resource | What it owns | |---|---| | [scripts/repo-doctor.py](scripts/repo-doctor.py) | The scorer: six dimensions, findings, grade; `--json` envelope `claude-mods.repo-doctor/v1`; `--strict` CI gate | | [references/scoring-rubric.md](references/scoring-rubric.md) | Every check: what it measures, threshold, why, and the fix | | [references/comment-doctrine.md](references/comment-doctrine.md) | Contract blocks, WHY-only inline, guard comments, section markers, format-at-site, citations — with good/bad examples | | [references/entry-docs.md](references/entry-docs.md) | AGENTS.md anatomy + Landmines, AGENTS-vs-CLAUDE decision, nesting policy, freshness discipline | | [references/monorepo-structure.md](references/monorepo-structure.md) | Structuring very large monorepos for agentic development | | [assets/AGENTS-template.md](assets/AGENTS-template.md) | Entry-doc skeleton with mandatory Landmines section | | [assets/docs-index-template.md](assets/docs-index-template.md) | `docs/00_INDEX.md` skeleton with the two anti-rot rules baked in |
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.