agent-surface-forge
Use when asked to audit or repair agent surfaces (plugins, agents, skills, CLAUDE.md/AGENTS.md, docs, prompts, commands, hooks) or improve one skill at depth. Not for agent grading: use skill-doctor.
Install
npx skills add https://github.com/OutlineDriven/outline-driven-development/tree/main/.devin/skills/agent-surface-forge
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install outlinedriven-outline-driven-development@llmmart
git clone https://github.com/OutlineDriven/outline-driven-development.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole outlinedriven/outline-driven-development collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Agent surface forge
Contract
| Field | Bound contract |
|---|---|
| Trigger | User asks to audit or repair agent surfaces, plugin configs, agent definitions, skills, CLAUDE.md/AGENTS.md, docs, prompts, commands, or hooks; or asks to iteratively improve one skill, fix skill-quality findings, or resolve scanner alerts (W001, W011, W012) on a skill. |
| Authority | Reversible local: audit mode writes only HIGH-certainty autoFix:yes findings on the named plugin/agent/skill/docs/prompt/claudemd/hooks surfaces under --apply; improve mode writes only files within the resolved skill directory; rollback is reverting any applied edit whose re-analysis introduces a new HIGH finding or restoring prior file content. No remote mutation. |
| Side effect | Local file edits to HIGH + autoFix:yes items only under --apply, or minimal contract-preserving edits inside one skill directory in improve mode; re-verifies after each edit. No suppression state, no model routing, no external analyzer binary outside scanner remediation. |
| Done | Audit mode: surface audit report produced with all HIGH findings complete in default output, MEDIUM/LOW shown only under --verbose, and under --apply each HIGH finding removed without new HIGH issues. Improve mode: a final review contains no critical or major findings and finalization edits pass scope and regression checks, or a non-success terminal status is reported without being presented as convergence; scanner remediation ends on a clean scan with information preserved and the pre-authoring checklist passed. |
Inputs
mode(optional):audit(default) orimprove. Audit runs the multi-surface sweep; improve runs the deep single-skill loop.target(optional, audit mode): path to audit; defaults to.. Taken from the first non-flag argument.--apply(optional flag, audit mode): apply only HIGH-certainty findings withautoFix: yes. Absent means report only.--verbose(optional flag): include MEDIUM and LOW findings. Default output shows HIGH plus MEDIUM/LOW summary counts.--focus=<type>(optional, audit mode): comma-separated subset ofplugin,agent,skill,docs,prompt,claudemd,hooks,cross-file. Unknown values are rejected.- Target skill path (required, improve mode): a path to a
SKILL.mdfile or a skill slug that resolves to one by searchingskills/. - Improvement scope (optional, improve mode): specific findings to fix, remaining findings from an interrupted run to resume, or quality dimensions to focus on. When the scope is scanner alerts, the scan output listing each alert as a code with its file, or the skill directory to scan to produce one. Optional
SNYK_TOKENis supplied by the operator environment, never requested or stored by this workflow.
Procedure
- Select mode.
improvewhen the user names one skill to iteratively improve, fix skill-quality findings on, or clear scanner alerts (W001, W011, W012) from; otherwiseaudit. Done when: the mode is selected.
Mode audit
- Parse intent. Extract
target,--apply,--verbose,--focus. Reject unknown focus values against the valid set above before any file access. Done when: target and flags are parsed and unknown focus values are rejected. - Discover files per analyzer family using these globs; keep concrete path lists and never pass globs to workers as the target contract:
- plugin:
plugins/*/.claude-plugin/plugin.json,.claude-plugin/plugin.json,**/.claude-plugin/plugin.json - agent:
**/agents/*.md - skill:
**/SKILL.md - claudemd:
CLAUDE.md,AGENTS.md,**/CLAUDE.md,**/AGENTS.md - docs:
docs/**,README.md,CHANGELOG.md - prompt:
commands/**,prompts/**,**/commands/*.md,**/prompts/*.md - hooks:
hooks/**,**/hooks/** - cross-file: enabled only when two or more of plugin/agent/skill/claudemd/prompt are present
Skip analyzers with no discovered files. Under
--focus, skip every non-focused family even if files exist. Done when: concrete file lists are discovered per analyzer family with non-focused families skipped.
- plugin:
- Launch parallel analyzers in one
taskbatch. Dispatch one generictaskworker for each analyzer family that has files. Give each worker the analyzer name, exact file list,--verbosestate, check catalog below, and required JSON return shape. Parallelize by analyzer family, never by file shards, because cross-surface consistency depends on seeing all relevant files for a family. Done when: one task worker is dispatched per analyzer family with files. - Check catalog per family. Every finding needs file, line or section, observed evidence, and a concrete fix; classify each by issue label: excess surface, duplication, structure, or correctness.
- plugin: manifest fields explicit and bounded;
nameanddescriptionpresent; no excess or duplicated permissions; structure valid. - agent: frontmatter parses;
nameanddescriptionpresent and bounded; no duplicated or contradictory directives; structure valid. - skill: frontmatter parses;
namematches the directory;descriptionpresent and bounded; no excess surface, duplication, or structural drift; sections internally consistent. - claudemd: directives explicit and bounded; no duplicated or contradictory rules; no stale aliases; structure valid.
- docs: referenced paths exist; no duplicated or contradictory content; structure valid.
- prompt: command and prompt definitions explicit and bounded; no duplicated commands; structure valid.
- hooks: hook scripts explicit and bounded; no duplicated handlers; structure and syntax valid.
- cross-file: links between surfaces resolve; no duplicated or contradictory definitions across families; names referenced in one surface exist in the target surface. Done when: each analyzer family has its check catalog applied.
- plugin: manifest fields explicit and bounded;
- Analyzer contract. Each worker reports only observed findings as JSON:
Use native tools directly:{ "analyzer": "plugin|agent|skill|docs|prompt|claudemd|hooks|cross-file", "findings": [ {"file":"path","line":1,"check":"missing_description","certainty":"HIGH","autoFix":false,"evidence":"...","fix":"..."} ], "summary": {"high":0,"medium":0,"low":0,"autoFixableHigh":0} }findandreadfor presence and frontmatter;searchscoped to discovered files for regex and field checks;ast-grepfor command snippets, shell patterns, and hook bodies where syntax matters; codegraph MCP for cross-file symbols, callers, and impact when indexed, falling back toast-grepplus scoped text search; repomix (pack_codebaseorpnpm dlx repomix --compress) only for large repos to build a digest for the workers. Done when: each worker returns findings in the required JSON shape. - Deduplicate. Stable key:
analyzer|file|line|check|normalized evidence. If two analyzers report the same underlying defect, keep the higher certainty; if equal, keep the one with narrower file/line evidence. Done when: findings are deduplicated by stable key. - Aggregate. Sort by certainty HIGH then MEDIUM then LOW, then analyzer, file, line. Count totals by analyzer and certainty. Done when: findings are sorted and totals are counted.
- Report. Default output: executive summary table, all HIGH findings, MEDIUM/LOW counts, and the HIGH auto-fixable list. Under
--verbose, include MEDIUM and LOW sections with issue labels. Done when: the report is emitted with the correct verbosity. - Apply guarded fixes, only when
--applyis present: filter tocertainty === HIGH && autoFix === true; group by analyzer; edit the minimal lines required; re-read changed files after each edit; never apply MEDIUM or LOW fixes automatically. Done when: all HIGH autoFixable findings are applied (under --apply) or listed as safe edits (without --apply). - Verify the fix set. Re-run only the analyzers whose files changed. A fix passes if the exact HIGH finding is gone and no new HIGH finding appears in the changed file. If a fix introduces a new HIGH issue, revert that fix and keep the finding in the report as manual. Done when: every applied fix is verified with no new HIGH findings, or reverted findings are labeled manual.
Mode improve
Run instead of steps 2-11.
- Resolve target. Locate the
SKILL.mdfile. If the path does not exist or does not contain a valid skill frontmatter block, reportinvalid-targetand stop. Done when: the target is resolved orinvalid-targetis reported. - Bound scope. The improvement scope is the single resolved skill directory. No file outside that directory may be read or written. Done when: the scope boundary is established.
- Initial review. Read the full
SKILL.mdand anyagents/openai.yamlin the skill directory. Evaluate against these quality dimensions:- Trigger clarity: does the frontmatter description and trigger predicate route precisely?
- Authority fidelity: does the body restatement match the declared authority without expansion?
- Procedure completeness: are steps numbered, executable, and free of ambiguity or missing branches?
- Semantic minimum: does every line earn its place by changing routing, authority, reads/writes, procedure, proof, failure handling, or license?
- Failure coverage: are named failure classes present with partial-result rules, rollback rules, and exact blocked-terminal output?
- Self-containment: does the body avoid pointers to other skills, AGENTS.md, system prompts, or rule files?
- Provenance: is origin, revision, license, and adaptation statement present?
Classify each finding as
critical,major, orminorand record them in the run report. Done when: findings are classified and recorded.
- Gate check. If zero critical and zero major findings exist, proceed to step 7 (finalization). If the iteration count equals max iterations, proceed to step 6 (non-converged terminal). Done when: the gate decision is made.
- Fix cycle. For each critical and major finding, in severity-then-file order: read the affected section; apply the minimal edit that resolves the finding without changing the skill's contract, trigger, authority, side-effect, or done predicate; record the diff in the run report. After all fixes for this iteration, re-run the review (step 3) with an incremented iteration counter. Done when: all critical and major findings for this iteration are fixed and the review is re-run.
- Non-converged terminal. If max iterations are exhausted with remaining critical or major findings, report
non-convergedwith the remaining finding count and severity breakdown. Do not present this as convergence. Done when: the non-converged status is reported. - Finalization. Run a scope check: confirm no edit changed the trigger predicate, authority class, side-effect target, or done predicate. Run a regression check: confirm the edited skill still parses (valid frontmatter, required sections present). If either check fails, revert the last iteration's edits and report
finalization-failed. If both pass, reportconvergedwith the final review findings. Done when: the finalization verdict is reported.
Scanner remediation replaces steps 3-7 when the improvement scope is scanner alerts (W001, W011, W012). Parse the scan output into (alert code, file) pairs. Only files named in the alert list are edited.
- Order the queue. Fix one alert at a time in order: W001, W011, then W012. Starting with the simplest alert minimizes rework when one fix surfaces another. Done when: the queue is ordered.
- Apply the restructuring rule for the alert code. Every rule preserves the original information by relocating or rephrasing it; deleting content to silence an alert is prohibited. Done when: the matching rule is applied.
- W001 (prompt injection via named MCP tool functions): replace each explicitly named tool function in body prose with a generic formulation naming the capability instead. Tool names remain acceptable in the frontmatter
allowed-toolsfield; only the body is restricted. - W011 (imperative external-content instructions): rewrite sentences that send the agent to fetch, check, or evaluate external content into passive availability statements that keep the URL and its purpose. Remove
alwaysfrom instructions involving external resources. Move tool invocations from prose checklists into code blocks. Running a tool is fine, but its remote-sourced output must not be the sole trigger for acting. - W012 (external content fetched and executed at runtime): replace
@latestwith an exact pinned version. Move install commands out of body prose into the alerted skill's frontmatter metadata install block. Pin GitHub Actions to a major version verified to exist in the action's releases. Never pipe remote content into a shell.
- W001 (prompt injection via named MCP tool functions): replace each explicitly named tool function in body prose with a generic formulation naming the capability instead. Tool names remain acceptable in the frontmatter
- Re-run the scanner after each fix. Run
SNYK_TOKEN=<token> snyk-agent-scan --skills <skill-directory>. If the binary is absent, runuvx snyk-agent-scanwithout installing it. Compare alert counts. If the count did not drop, undo that edit and choose a different restructuring; never stack unverified changes. Done when: the post-fix alert count is compared. - Queue surfaced alerts. Expect W011 fixes to surface hidden W012 alerts as URLs become prominent after restructuring. Treat each surfaced alert as a new item in the ordered queue. Done when: surfaced alerts are queued.
- Restructure likely false positives. Treat a URL in a reference-data table cell, official documentation link, frontmatter homepage link, or
alwaysoutside an external-resource sentence the same as a confirmed alert: use the passive-availability pattern. Do not override, suppress, or assume scanner error. Done when: likely false positives are restructured. - Prevent recurrence. When the scan is clean, apply the pre-authoring checklist to the edited content: no sentence with the agent acting on a URL, no
@latestin body install instructions, no MCP tool names in body prose, install commands in frontmatter, GitHub Actions versions real, tool invocations in code blocks, and noalwaysbefore external-resource instructions. Done when: the checklist passes on the edited content.
Failure and recovery
- Unknown
--focusvalue: reject before discovery; no files read or changed. - Analyzer worker returns malformed JSON or no findings array: discard that worker's result, mark the analyzer as errored in the report, and continue other analyzers. The partial result covers every family that returned valid output.
- An applied fix introduces a new HIGH finding: revert that single edit, keep the original finding in the report labeled manual, and leave other applied fixes intact.
- No files discovered for any family: report zero analyzers run; do not invent findings.
- Non-converged fix: if re-analysis cannot confirm a finding is gone after a revert, report the finding as unresolved and stop applying further fixes in that file.
invalid-target(improve mode): the target path does not exist or lacks valid skill frontmatter. Report and stop.scope-violation(improve mode): an edit would touch a file outside the resolved skill directory. Revert the edit, record the violation, and continue with remaining findings.non-converged(improve mode): max iterations exhausted with critical or major findings remaining. Report the count and severity breakdown. Do not claim convergence.finalization-failed(improve mode): the post-convergence scope or regression check fails. Revert the last iteration's edits and report the specific check failure.review-error(improve mode): the review step itself fails (unparseable skill, tool error). Record the error, skip the fix cycle, and reportreview-errorwith the error message.scanner-unavailable: the scanner binary is missing and theuvxdrop-in fails, orSNYK_TOKENis missing. Report the exact error and stop. No edit made in an unverified run may be claimed as fixed.alert-count-not-dropping: restore the pre-edit bytes and select a different restructuring. The failing edit must not survive.alert-count-rising: revert to the last state with the lowest verified count and stop the loop there.information-preservation-limit: an alert cannot clear without deleting content. Stop editing, keep all verified fixes, and return the blocking file, alert code, and constraint. Do not claim the done predicate.
Partial results (improve mode): record each iteration's findings and fixes in the run report as they happen. A mid-run interruption preserves everything emitted so far; resume by passing the remaining findings as the improvement scope.
Output
Audit mode: a markdown surface audit report with target path and flags, analyzers run, an executive summary table with HIGH/MEDIUM/LOW/autoFixableHigh counts per analyzer, every HIGH finding with file/line/check/evidence/fix, MEDIUM/LOW counts (full sections only under --verbose), and an auto-fix section listing applied edits and re-analysis results or safe edits without --apply.
Improve mode: on converged, a report listing the final review findings (all minor or informational) and the total iteration count; on non-converged, the remaining critical and major findings and the iteration count; on invalid-target, finalization-failed, or review-error, a terminal status message with the specific failure class and diagnostic detail. Scanner remediation: a per-file remediation report ordered by alert queue, listing code, restructuring, before-and-after counts, then final clean status or exact blockers, and the pre-authoring checklist result.
Files (outline-driven-development)
-
agents
-
openai.yaml 231 B
interface: display_name: "Agent Surface Forge" short_description: "Use when asked to audit or repair agent surfaces (plugins, agents, skills, CLAUDE.md/AGENTS.md, docs, prompts, commands, hooks) or improve one skill at depth."
-
-
references
-
analyzer-checks.md 13.8 KB
# Improve analyzer checks Every analyzer emits `certainty` as HIGH/MEDIUM/LOW and `autoFix` as yes/no. HIGH means the invariant violation is directly observable from file content or parsed config. MEDIUM means context-dependent but likely useful. LOW means advisory. Auto-fix policy: only `HIGH + autoFix: yes` may be applied, and only after explicit `--apply`. MEDIUM/LOW are report-only even when an individual row says the source tool once had a suggested fixer. ## Plugin analyzer Scope: `.claude-plugin/plugin.json`, package metadata near it, command definitions, MCP/tool schemas, embedded agent/command references. | Check | Certainty | autoFix | |---|---|---| | `plugin.json` missing, unreadable, or malformed JSON | HIGH | no | | Missing required plugin fields: `name`, `version`, `description` | HIGH | no | | `version` does not follow semver-like `x.y.z` format | HIGH | no | | Version mismatch between `.claude-plugin/plugin.json`, marketplace metadata, and nearby `package.json` when both are present | HIGH | yes | | Plugin declares command entries without required command fields (`name`, command path/entrypoint, or declared argument surface) | HIGH | no | | Command markdown/frontmatter missing `description` or tool-description equivalent for what the command does | HIGH | no | | Command or tool parameter object omits a `required` declaration for known required inputs | HIGH | yes | | Tool schema missing `additionalProperties: false` / strict object closure | HIGH | yes | | Tool definition missing human-readable `description` | HIGH | no | | Parameter schema missing parameter descriptions | MEDIUM | no | | Parameter schema is nested deeper than two object levels | MEDIUM | no | | Tool or command description exceeds 500 characters without adding constraints or examples | MEDIUM | no | | Plugin exposes many independent tools in one plugin; split may be clearer | LOW | no | | Tool allow-list grants broad shell/file access without command restrictions | HIGH | no | | Marketplace/plugin metadata contradicts manifest name, version, or capabilities | HIGH | no | ## Agent analyzer Scope: `**/agents/*.md` and agent-like markdown with YAML frontmatter. | Check | Certainty | autoFix | |---|---|---| | Missing YAML frontmatter block | HIGH | yes | | Frontmatter missing `name` | HIGH | no | | Frontmatter missing `description` | HIGH | no | | No `tools` field, implying unrestricted access | HIGH | no | | `tools` contains unrestricted `Bash` instead of command-scoped forms such as `Bash(git:*)` | HIGH | yes | | Missing explicit role statement (`You are`, `## Role`, identity/mission section) | HIGH | yes | | Missing output format or return contract | HIGH | no | | Missing constraints / MUST NOT / safety section | HIGH | no | | XML-like blocks are opened but not balanced/closed | HIGH | no | | Complex agent prompt has no XML or similarly parseable sections for large context blocks | MEDIUM | no | | Simple agent uses redundant step-by-step reasoning instructions | MEDIUM | no | | Complex reasoning agent lacks thinking/verification guidance | MEDIUM | no | | Vague instructions (`usually`, `maybe`, `try to`, `as needed`) drive behavior without decision criteria | MEDIUM | no | | References `CLAUDE.md` but not `AGENTS.md` where both project-memory surfaces may matter | MEDIUM | no | | Example count outside 2-5 examples for an example-driven role | LOW | no | | Prompt body exceeds roughly 2k tokens without strong structure | LOW | no | | Hardcoded `.claude/` state path instead of repo-relative or environment-aware location | HIGH | no | | Bloat: repeated rules, duplicated role statements, or long preambles not tied to behavior | LOW | no | ## Skill analyzer Scope: every `SKILL.md`. | Check | Certainty | autoFix | |---|---|---| | Missing YAML frontmatter block | HIGH | no | | Frontmatter `name` missing | HIGH | no | | Frontmatter `name` does not equal directory basename | HIGH | no | | Frontmatter `description` missing | HIGH | no | | Description lacks explicit trigger phrase ending (`Use when ...`) | MEDIUM | yes | | Skill causes side effects, writes files, runs commands, or coordinates multi-step work but lacks `disable-model-invocation: true` when it must be explicit-only | HIGH | no | | Skill lists `allowed-tools` but omits a tool required by its own workflow | MEDIUM | no | | Skill needs restricted tooling but has no allowed-tools/tool-boundary statement | MEDIUM | no | | Missing workflow section for a procedural skill | HIGH | no | | Missing validation gates for any skill that edits files, changes state, or executes commands | MEDIUM | no | | Uses broad external delegation instead of native file/search/read/AST recipes | MEDIUM | no | | Trigger phrasing is vague, purely marketing, or not action-oriented | MEDIUM | no | | Bulk tables/templates in `SKILL.md` instead of a local `references/` file | LOW | no | | Mentions non-native model routing, editor shims, or binary-cache state as required machinery | HIGH | no | | `references/` link points outside the same skill directory | HIGH | no | ## Docs analyzer Scope: `docs/**`, README, CHANGELOG, guide pages, API docs, inline code examples inside documentation. | Check | Certainty | autoFix | |---|---|---| | Internal doc link targets a missing file or missing anchor | HIGH | no | | Heading hierarchy skips levels (for example H1 → H3) | HIGH | yes | | Code block lacks language tag | HIGH | no | | Code block is syntactically invalid for its declared JSON/YAML/TOML language | HIGH | no | | Code example imports, commands, or exported symbol names that do not exist in repo search/codegraph | HIGH | no | | Version text contradicts manifest/package version in the same repo | HIGH | yes | | Filler prose adds no operational information | HIGH | no | | Verbose passage can be compressed without losing instructions | HIGH | yes | | Section exceeds about 1000 tokens and should be chunked for retrieval | MEDIUM | no | | Section mixes multiple unrelated topics | MEDIUM | no | | Section lacks local context needed for retrieval; reader must rely on preceding sections | MEDIUM | no | | Important setup or warning buried below unrelated material | MEDIUM | no | | Long block has no section headers | MEDIUM | no | | Token-reduction opportunity with no correctness impact | LOW | no | | Readability/RAG balance issue: too terse for humans or too narrative for retrieval | LOW | no | | Generic structure recommendation with no broken invariant | LOW | no | ## Prompt analyzer Scope: command prompts, reusable prompts, system-prompt drafts, prompt templates, agent task prompts. | Check | Certainty | autoFix | |---|---|---| | Vague language controls behavior without measurable criteria (`usually`, `sometimes`, `try`, `consider`) | HIGH | no | | Negative-only instruction (`don't`, `never`, `do not`) lacks the positive replacement behavior | HIGH | no | | No clear output format for a task/prompt that asks for deliverables | HIGH | yes | | Excessive aggressive emphasis likely to over-index the model | HIGH | yes | | Task description lacks concrete scope: file, scenario, constraints, or acceptance criteria | HIGH | no | | Invalid JSON inside JSON code block | HIGH | no | | Heading hierarchy skips levels | HIGH | no | | Redundant `think step by step` or equivalent generic chain-of-thought incantation | HIGH | no | | Complex prompt lacks examples/few-shot contrast | MEDIUM | no | | Examples lack good/bad contrast where pattern learning is required | MEDIUM | no | | Important instruction buried in the middle instead of top/bottom anchoring | MEDIUM | no | | Instruction lacks rationale where violating it is plausible | MEDIUM | no | | No priority order among competing instructions | MEDIUM | no | | Requests JSON output without schema or example | MEDIUM | no | | Task lacks verification criteria: tests, expected output, screenshot, or scenario | MEDIUM | no | | Investigation prompt does not point to likely sources or evidence types | MEDIUM | no | | Code block language tag likely mismatches content | MEDIUM | no | | Complex prompt lacks XML/tagged structure | LOW | no | | Example count outside 2-5 where examples exist | LOW | no | | Overly prescriptive phase list may block better reasoning | LOW | no | | Prompt exceeds roughly 2500 tokens without retrieval/chunking need | LOW | no | ## CLAUDE.md / AGENTS.md analyzer Scope: project memory and project instruction files named `CLAUDE.md`, `AGENTS.md`, or equivalent top-level agent memory. | Check | Certainty | autoFix | |---|---|---| | Required project memory file absent when user explicitly asked to improve project memory | HIGH | no | | No critical rules / priority rules section | HIGH | no | | No architecture, project structure, or where-things-live section | HIGH | no | | No commands/scripts section despite package/test/build commands existing | HIGH | no | | References a file path that does not exist | HIGH | no | | Documents a package script or command that does not exist | HIGH | no | | Hardcoded `.claude/` path for runtime state instead of repo-relative/environment-aware path | HIGH | no | | File is long enough that high-priority rules are likely ignored (>150 lines or similar) | HIGH | no | | Duplicates README content instead of carrying agent-only constraints | MEDIUM | no | | Exceeds recommended quick-reference token budget | MEDIUM | no | | Instructions are too verbose for retrieval during tool use | MEDIUM | no | | Rules lack WHY/context where misuse is likely | MEDIUM | no | | Uses Claude-specific terminology where AGENTS-compatible wording is needed | MEDIUM | no | | `CLAUDE.md` omits `AGENTS.md` compatibility note when the repo uses both surfaces | MEDIUM | no | | Includes information the agent can infer cheaply from reading manifests/code | MEDIUM | no | | Critical rules lack emphasis markers (`MUST`, `NEVER`, `CRITICAL`) | MEDIUM | no | | Too many inline examples for a memory file | LOW | no | | Deep nesting (>3 heading/list levels) | LOW | no | | Contains self-evident generic practices that should be deleted | LOW | no | ## Hooks analyzer Scope: hook config files, hook markdown definitions, shell snippets, JS/Python hook scripts, and hook references in plugin manifests. | Check | Certainty | autoFix | |---|---|---| | Hook definition file missing/unreadable | HIGH | no | | Hook markdown missing frontmatter | HIGH | no | | Hook frontmatter missing `name` | HIGH | no | | Hook frontmatter missing `description` | HIGH | no | | Dangerous command pattern: `rm -rf`, `sudo`, `chmod -R 777`, `chown -R`, destructive `git reset --hard`, force push, or kill-all style command | HIGH | no | | Shell command interpolates untrusted input without quoting | HIGH | no | | Hook writes outside repo or declared state directory | HIGH | no | | Hook runs network/deployment command without explicit user-facing gate | HIGH | no | | Hook references tool/command not declared in plugin config | HIGH | no | | Hook has side effects but lacks clear trigger event and scope | HIGH | no | | Hook swallows errors (`|| true`, empty catch, broad redirect) around validation logic | MEDIUM | no | | Hook has no timeout, bounded input, or failure mode for long-running work | MEDIUM | no | | Hook emits noisy output on every prompt/tool call | MEDIUM | no | | Hook duplicates a rule already present in project memory | MEDIUM | no | | Hook script has no comments for non-obvious policy decisions | LOW | no | ## Cross-file analyzer Scope: relationships among plugin manifests, agents, skills, prompts, commands, hooks, docs, and project memory. | Check | Certainty | autoFix | |---|---|---| | Tool consistency: prompt/agent mentions a tool not declared in frontmatter or plugin config | MEDIUM | no | | Tool consistency: declared tool is never used or justified in body | MEDIUM | no | | Workflow references a non-existent agent, command, skill, hook, or file | MEDIUM | no | | Workflow phase transitions are incomplete: later phase references missing earlier artifact | MEDIUM | no | | Agent ↔ skill mismatch: skill describes capabilities not present in any related agent/prompt body | MEDIUM | no | | Agent ↔ skill mismatch: agent performs behavior the owning skill/manifest does not describe | MEDIUM | no | | Skill allowed-tools differs from prompt/tool usage | MEDIUM | no | | Duplicate rules: same critical instruction appears across files with high token/Jaccard overlap | MEDIUM | no | | Contradiction: one file says `always/must/required`; another says `never/must not/forbidden` for overlapping token set | MEDIUM | no | | Rule drift: same named command/agent has different trigger semantics in manifest, docs, and prompt | MEDIUM | no | | Agent/prompt/command not referenced anywhere and not documented as standalone | MEDIUM | no | | Docs mention plugin/agent/skill capability missing from manifest or source config | MEDIUM | no | | Marketplace/README claims a command exists but command file is absent | HIGH | no | | Two command or agent names collide case-insensitively | HIGH | no | | Local reference points outside the repo or outside its own skill/plugin directory without explicit rationale | HIGH | no | ## Detection notes - XML imbalance: count same-named open/close tags for simple XML-like blocks; ignore Markdown autolinks and HTML void tags. Certainty HIGH only when a concrete tag is opened and not closed. - Duplicate/contradiction overlap: normalize to lowercase tokens, remove stopwords, require Jaccard overlap ≥0.65 for duplicate-rule candidates and ≥0.45 plus opposing modal verbs for contradiction candidates. Treat as MEDIUM unless names collide or a required file is absent. - Tool overexposure: unrestricted `Bash`, wildcard file/network tools, or no `tools` field on an agent are HIGH because the configured boundary is directly observable. - Bloat: LOW unless it hides required rules, contradicts frontmatter, or crosses the CLAUDE.md/AGENTS.md length gate. - Auto-fix yes rows must still be minimal: add missing delimiters/strict schema/heading increments/version sync/output-format scaffold only. Do not invent names, descriptions, policies, or tool scopes.
-
-
SKILL.md 17.5 KB
--- name: agent-surface-forge description: 'Use when asked to audit or repair agent surfaces (plugins, agents, skills, CLAUDE.md/AGENTS.md, docs, prompts, commands, hooks) or improve one skill at depth. Not for agent grading: use skill-doctor.' --- # Agent surface forge ## Contract | Field | Bound contract | |---|---| | Trigger | User asks to audit or repair agent surfaces, plugin configs, agent definitions, skills, CLAUDE.md/AGENTS.md, docs, prompts, commands, or hooks; or asks to iteratively improve one skill, fix skill-quality findings, or resolve scanner alerts (W001, W011, W012) on a skill. | | Authority | Reversible local: audit mode writes only HIGH-certainty autoFix:yes findings on the named plugin/agent/skill/docs/prompt/claudemd/hooks surfaces under --apply; improve mode writes only files within the resolved skill directory; rollback is reverting any applied edit whose re-analysis introduces a new HIGH finding or restoring prior file content. No remote mutation. | | Side effect | Local file edits to HIGH + autoFix:yes items only under --apply, or minimal contract-preserving edits inside one skill directory in improve mode; re-verifies after each edit. No suppression state, no model routing, no external analyzer binary outside scanner remediation. | | Done | Audit mode: surface audit report produced with all HIGH findings complete in default output, MEDIUM/LOW shown only under --verbose, and under --apply each HIGH finding removed without new HIGH issues. Improve mode: a final review contains no critical or major findings and finalization edits pass scope and regression checks, or a non-success terminal status is reported without being presented as convergence; scanner remediation ends on a clean scan with information preserved and the pre-authoring checklist passed. | ## Inputs - `mode` (optional): `audit` (default) or `improve`. Audit runs the multi-surface sweep; improve runs the deep single-skill loop. - `target` (optional, audit mode): path to audit; defaults to `.`. Taken from the first non-flag argument. - `--apply` (optional flag, audit mode): apply only HIGH-certainty findings with `autoFix: yes`. Absent means report only. - `--verbose` (optional flag): include MEDIUM and LOW findings. Default output shows HIGH plus MEDIUM/LOW summary counts. - `--focus=<type>` (optional, audit mode): comma-separated subset of `plugin`, `agent`, `skill`, `docs`, `prompt`, `claudemd`, `hooks`, `cross-file`. Unknown values are rejected. - Target skill path (required, improve mode): a path to a `SKILL.md` file or a skill slug that resolves to one by searching `skills/`. - Improvement scope (optional, improve mode): specific findings to fix, remaining findings from an interrupted run to resume, or quality dimensions to focus on. When the scope is scanner alerts, the scan output listing each alert as a code with its file, or the skill directory to scan to produce one. Optional `SNYK_TOKEN` is supplied by the operator environment, never requested or stored by this workflow. ## Procedure 1. Select mode. `improve` when the user names one skill to iteratively improve, fix skill-quality findings on, or clear scanner alerts (W001, W011, W012) from; otherwise `audit`. Done when: the mode is selected. ### Mode audit 2. Parse intent. Extract `target`, `--apply`, `--verbose`, `--focus`. Reject unknown focus values against the valid set above before any file access. Done when: target and flags are parsed and unknown focus values are rejected. 3. Discover files per analyzer family using these globs; keep concrete path lists and never pass globs to workers as the target contract: - plugin: `plugins/*/.claude-plugin/plugin.json`, `.claude-plugin/plugin.json`, `**/.claude-plugin/plugin.json` - agent: `**/agents/*.md` - skill: `**/SKILL.md` - claudemd: `CLAUDE.md`, `AGENTS.md`, `**/CLAUDE.md`, `**/AGENTS.md` - docs: `docs/**`, `README.md`, `CHANGELOG.md` - prompt: `commands/**`, `prompts/**`, `**/commands/*.md`, `**/prompts/*.md` - hooks: `hooks/**`, `**/hooks/**` - cross-file: enabled only when two or more of plugin/agent/skill/claudemd/prompt are present Skip analyzers with no discovered files. Under `--focus`, skip every non-focused family even if files exist. Done when: concrete file lists are discovered per analyzer family with non-focused families skipped. 4. Launch parallel analyzers in one `task` batch. Dispatch one generic `task` worker for each analyzer family that has files. Give each worker the analyzer name, exact file list, `--verbose` state, check catalog below, and required JSON return shape. Parallelize by analyzer family, never by file shards, because cross-surface consistency depends on seeing all relevant files for a family. Done when: one task worker is dispatched per analyzer family with files. 5. Check catalog per family. Every finding needs file, line or section, observed evidence, and a concrete fix; classify each by issue label: excess surface, duplication, structure, or correctness. - plugin: manifest fields explicit and bounded; `name` and `description` present; no excess or duplicated permissions; structure valid. - agent: frontmatter parses; `name` and `description` present and bounded; no duplicated or contradictory directives; structure valid. - skill: frontmatter parses; `name` matches the directory; `description` present and bounded; no excess surface, duplication, or structural drift; sections internally consistent. - claudemd: directives explicit and bounded; no duplicated or contradictory rules; no stale aliases; structure valid. - docs: referenced paths exist; no duplicated or contradictory content; structure valid. - prompt: command and prompt definitions explicit and bounded; no duplicated commands; structure valid. - hooks: hook scripts explicit and bounded; no duplicated handlers; structure and syntax valid. - cross-file: links between surfaces resolve; no duplicated or contradictory definitions across families; names referenced in one surface exist in the target surface. Done when: each analyzer family has its check catalog applied. 6. Analyzer contract. Each worker reports only observed findings as JSON: ```json { "analyzer": "plugin|agent|skill|docs|prompt|claudemd|hooks|cross-file", "findings": [ {"file":"path","line":1,"check":"missing_description","certainty":"HIGH","autoFix":false,"evidence":"...","fix":"..."} ], "summary": {"high":0,"medium":0,"low":0,"autoFixableHigh":0} } ``` Use native tools directly: `find` and `read` for presence and frontmatter; `search` scoped to discovered files for regex and field checks; `ast-grep` for command snippets, shell patterns, and hook bodies where syntax matters; codegraph MCP for cross-file symbols, callers, and impact when indexed, falling back to `ast-grep` plus scoped text search; repomix (`pack_codebase` or `pnpm dlx repomix --compress`) only for large repos to build a digest for the workers. Done when: each worker returns findings in the required JSON shape. 7. Deduplicate. Stable key: `analyzer|file|line|check|normalized evidence`. If two analyzers report the same underlying defect, keep the higher certainty; if equal, keep the one with narrower file/line evidence. Done when: findings are deduplicated by stable key. 8. Aggregate. Sort by certainty HIGH then MEDIUM then LOW, then analyzer, file, line. Count totals by analyzer and certainty. Done when: findings are sorted and totals are counted. 9. Report. Default output: executive summary table, all HIGH findings, MEDIUM/LOW counts, and the HIGH auto-fixable list. Under `--verbose`, include MEDIUM and LOW sections with issue labels. Done when: the report is emitted with the correct verbosity. 10. Apply guarded fixes, only when `--apply` is present: filter to `certainty === HIGH && autoFix === true`; group by analyzer; edit the minimal lines required; re-read changed files after each edit; never apply MEDIUM or LOW fixes automatically. Done when: all HIGH autoFixable findings are applied (under --apply) or listed as safe edits (without --apply). 11. Verify the fix set. Re-run only the analyzers whose files changed. A fix passes if the exact HIGH finding is gone and no new HIGH finding appears in the changed file. If a fix introduces a new HIGH issue, revert that fix and keep the finding in the report as manual. Done when: every applied fix is verified with no new HIGH findings, or reverted findings are labeled manual. ### Mode improve Run instead of steps 2-11. 1. **Resolve target.** Locate the `SKILL.md` file. If the path does not exist or does not contain a valid skill frontmatter block, report `invalid-target` and stop. Done when: the target is resolved or `invalid-target` is reported. 2. **Bound scope.** The improvement scope is the single resolved skill directory. No file outside that directory may be read or written. Done when: the scope boundary is established. 3. **Initial review.** Read the full `SKILL.md` and any `agents/openai.yaml` in the skill directory. Evaluate against these quality dimensions: - Trigger clarity: does the frontmatter description and trigger predicate route precisely? - Authority fidelity: does the body restatement match the declared authority without expansion? - Procedure completeness: are steps numbered, executable, and free of ambiguity or missing branches? - Semantic minimum: does every line earn its place by changing routing, authority, reads/writes, procedure, proof, failure handling, or license? - Failure coverage: are named failure classes present with partial-result rules, rollback rules, and exact blocked-terminal output? - Self-containment: does the body avoid pointers to other skills, AGENTS.md, system prompts, or rule files? - Provenance: is origin, revision, license, and adaptation statement present? Classify each finding as `critical`, `major`, or `minor` and record them in the run report. Done when: findings are classified and recorded. 4. **Gate check.** If zero critical and zero major findings exist, proceed to step 7 (finalization). If the iteration count equals max iterations, proceed to step 6 (non-converged terminal). Done when: the gate decision is made. 5. **Fix cycle.** For each critical and major finding, in severity-then-file order: read the affected section; apply the minimal edit that resolves the finding without changing the skill's contract, trigger, authority, side-effect, or done predicate; record the diff in the run report. After all fixes for this iteration, re-run the review (step 3) with an incremented iteration counter. Done when: all critical and major findings for this iteration are fixed and the review is re-run. 6. **Non-converged terminal.** If max iterations are exhausted with remaining critical or major findings, report `non-converged` with the remaining finding count and severity breakdown. Do not present this as convergence. Done when: the non-converged status is reported. 7. **Finalization.** Run a scope check: confirm no edit changed the trigger predicate, authority class, side-effect target, or done predicate. Run a regression check: confirm the edited skill still parses (valid frontmatter, required sections present). If either check fails, revert the last iteration's edits and report `finalization-failed`. If both pass, report `converged` with the final review findings. Done when: the finalization verdict is reported. Scanner remediation replaces steps 3-7 when the improvement scope is scanner alerts (W001, W011, W012). Parse the scan output into `(alert code, file)` pairs. Only files named in the alert list are edited. 1. **Order the queue.** Fix one alert at a time in order: W001, W011, then W012. Starting with the simplest alert minimizes rework when one fix surfaces another. Done when: the queue is ordered. 2. **Apply the restructuring rule for the alert code.** Every rule preserves the original information by relocating or rephrasing it; deleting content to silence an alert is prohibited. Done when: the matching rule is applied. - W001 (prompt injection via named MCP tool functions): replace each explicitly named tool function in body prose with a generic formulation naming the capability instead. Tool names remain acceptable in the frontmatter `allowed-tools` field; only the body is restricted. - W011 (imperative external-content instructions): rewrite sentences that send the agent to fetch, check, or evaluate external content into passive availability statements that keep the URL and its purpose. Remove `always` from instructions involving external resources. Move tool invocations from prose checklists into code blocks. Running a tool is fine, but its remote-sourced output must not be the sole trigger for acting. - W012 (external content fetched and executed at runtime): replace `@latest` with an exact pinned version. Move install commands out of body prose into the alerted skill's frontmatter metadata install block. Pin GitHub Actions to a major version verified to exist in the action's releases. Never pipe remote content into a shell. 3. **Re-run the scanner after each fix.** Run `SNYK_TOKEN=<token> snyk-agent-scan --skills <skill-directory>`. If the binary is absent, run `uvx snyk-agent-scan` without installing it. Compare alert counts. If the count did not drop, undo that edit and choose a different restructuring; never stack unverified changes. Done when: the post-fix alert count is compared. 4. **Queue surfaced alerts.** Expect W011 fixes to surface hidden W012 alerts as URLs become prominent after restructuring. Treat each surfaced alert as a new item in the ordered queue. Done when: surfaced alerts are queued. 5. **Restructure likely false positives.** Treat a URL in a reference-data table cell, official documentation link, frontmatter homepage link, or `always` outside an external-resource sentence the same as a confirmed alert: use the passive-availability pattern. Do not override, suppress, or assume scanner error. Done when: likely false positives are restructured. 6. **Prevent recurrence.** When the scan is clean, apply the pre-authoring checklist to the edited content: no sentence with the agent acting on a URL, no `@latest` in body install instructions, no MCP tool names in body prose, install commands in frontmatter, GitHub Actions versions real, tool invocations in code blocks, and no `always` before external-resource instructions. Done when: the checklist passes on the edited content. ## Failure and recovery - Unknown `--focus` value: reject before discovery; no files read or changed. - Analyzer worker returns malformed JSON or no findings array: discard that worker's result, mark the analyzer as errored in the report, and continue other analyzers. The partial result covers every family that returned valid output. - An applied fix introduces a new HIGH finding: revert that single edit, keep the original finding in the report labeled manual, and leave other applied fixes intact. - No files discovered for any family: report zero analyzers run; do not invent findings. - Non-converged fix: if re-analysis cannot confirm a finding is gone after a revert, report the finding as unresolved and stop applying further fixes in that file. - `invalid-target` (improve mode): the target path does not exist or lacks valid skill frontmatter. Report and stop. - `scope-violation` (improve mode): an edit would touch a file outside the resolved skill directory. Revert the edit, record the violation, and continue with remaining findings. - `non-converged` (improve mode): max iterations exhausted with critical or major findings remaining. Report the count and severity breakdown. Do not claim convergence. - `finalization-failed` (improve mode): the post-convergence scope or regression check fails. Revert the last iteration's edits and report the specific check failure. - `review-error` (improve mode): the review step itself fails (unparseable skill, tool error). Record the error, skip the fix cycle, and report `review-error` with the error message. - `scanner-unavailable`: the scanner binary is missing and the `uvx` drop-in fails, or `SNYK_TOKEN` is missing. Report the exact error and stop. No edit made in an unverified run may be claimed as fixed. - `alert-count-not-dropping`: restore the pre-edit bytes and select a different restructuring. The failing edit must not survive. - `alert-count-rising`: revert to the last state with the lowest verified count and stop the loop there. - `information-preservation-limit`: an alert cannot clear without deleting content. Stop editing, keep all verified fixes, and return the blocking file, alert code, and constraint. Do not claim the done predicate. Partial results (improve mode): record each iteration's findings and fixes in the run report as they happen. A mid-run interruption preserves everything emitted so far; resume by passing the remaining findings as the improvement scope. ## Output Audit mode: a markdown surface audit report with target path and flags, analyzers run, an executive summary table with HIGH/MEDIUM/LOW/autoFixableHigh counts per analyzer, every HIGH finding with file/line/check/evidence/fix, MEDIUM/LOW counts (full sections only under --verbose), and an auto-fix section listing applied edits and re-analysis results or safe edits without --apply. Improve mode: on `converged`, a report listing the final review findings (all minor or informational) and the total iteration count; on `non-converged`, the remaining critical and major findings and the iteration count; on `invalid-target`, `finalization-failed`, or `review-error`, a terminal status message with the specific failure class and diagnostic detail. Scanner remediation: a per-file remediation report ordered by alert queue, listing code, restructuring, before-and-after counts, then final clean status or exact blockers, and the pre-authoring checklist result.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.