Claude Skill

skill-repo

Use when creating skill repositories, standardizing or validating skill repo structure, setting up composer/release workflows, configuring split licensing (MIT + CC-BY-SA-4.0), fixing plugin.json / SKILL.md validation or version-parity errors, or releasing a skill version (versio

LLM Mart · 0 points · 1 views 2 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download netresearch-skill-repo-skill-skills_skill-repo-a1111ec.zip · 179 KB

Install

skills CLI npx skills add https://github.com/netresearch/skill-repo-skill/tree/main/skills/skill-repo
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install netresearch-skill-repo-skill@llmmart
Git git clone https://github.com/netresearch/skill-repo-skill.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole netresearch/skill-repo-skill collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Skill Repository Structure Guide

Repository Structure

{repo-name}/
├── plugin.json                  # portable manifest
├── .claude-plugin/plugin.json   # generated
├── skills/{name}/SKILL.md       # the control plane
├── README.md                    # human docs
├── LICENSE-MIT                  # code
├── LICENSE-CC-BY-SA-4.0         # content
├── composer.json                # PHP distribution
├── references/                  # detail, loaded on demand
├── scripts/                     # executables, never loaded
└── .github/workflows/
    ├── release.yml              # tag-triggered
    ├── validate.yml             # validation caller
    └── auto-merge-deps.yml      # dep auto-merge caller

Licensing (Split Model)

Path pattern License
skills/**/*.md, references/**, README.md, docs/** CC-BY-SA-4.0
scripts/**, .github/workflows/**, *.sh, *.py, *.php MIT
composer.json, plugin.json, config files MIT

SPDX: (MIT AND CC-BY-SA-4.0). Copyright: Netresearch DTT GmbH. No bare LICENSE — split files only.

SKILL.md Frontmatter

---
name: skill-name
description: "Use when <trigger conditions>"
---

Budgets (spec): name ≤64, no doubled/edge hyphen, matches its directory. description ≤1024, warn past 500 — a router, not documentation. Body ≤500 lines, warn at 300. compatibility ≤500, usually omit.

Flat discovery: references one level deep; SKILL.md names every references/*.md and every scripts/ executable; ## Contents past 100 lines. Rationale and trigger evals: skill-architecture. Audit: audit-skills.sh in the repository's top-level scripts/ (not shipped with the skill).

Manifests

Root plugin.json (Agent Plugins 1.0.0, closed field set) is the source of truth; sync-plugin-manifest.sh generates .claude-plugin/plugin.json plus Claude-only keys — agent-plugins-compat.

{
  "$schema": "https://agent-plugins.org/schemas/1.0.0/plugin.schema.json",
  "name": "skill-name",
  "version": "1.0.0",
  "license": "(MIT AND CC-BY-SA-4.0)",
  "author": {"name": "Netresearch DTT GmbH", "url": "https://www.netresearch.de"}
}

composer.json

Name must match GitHub repo. Type ai-agent-skill. No version field (from git tags). No composer.lock.

{
  "name": "netresearch/{repo-name}",
  "type": "ai-agent-skill",
  "license": "(MIT AND CC-BY-SA-4.0)",
  "require": {"netresearch/composer-agent-skill-plugin": "*"},
  "extra": {"ai-agent-skill": "skills/{name}/SKILL.md"}
}

Reusable Workflow Callers

Skill repos MUST delegate CI to skill-repo-skill reusable workflows:

# .github/workflows/validate.yml
uses: netresearch/skill-repo-skill/.github/workflows/validate.yml@main

Callers: validate.yml, release.yml (here); auto-merge-deps.yml (netresearch/.github). Auto-merge/pr-quality use pull_request_target. No inline Actions. Domain reusables: docs/ARCHITECTURE.md.

Releasing

Bump root plugin.json → sync → PR → merge → pull main → verify parity → signed tag → push → monitor Release. Tag only after bump PR merges. Multi-repo (>3) needs dry-run + approval. Never edit installed paths — release-discipline.

Installation

Marketplace, release download, Composer, npm — commands and the npm files default in installation-methods.

Validation

scripts/validate-skill.sh (layout, manifests, budgets, flat discovery). Shell portability: authoring-ci-gotchas.

Named here so they are findable: bump-version.sh, check-version-parity.sh, sync-plugin-manifest.sh, roll-changelog.py, fleet-release-github.sh, migrate-licensing.sh, validate-evals.sh — each with --help.

References (references/)

agent-plugins-compat · installation-methods · plugin-hooks · composer-setup · release-discipline · review-replies · skill-quality · repository-quality-rules · readme-template · skill-discovery-metadata · validation-checklist · marketplace-integration · materialization-contract · authoring-ci-gotchas · skill-retirement


Contributing: https://github.com/netresearch/skill-repo-skill

Files (skill-repo-skill)
  • evals
    • evals.json 20.5 KB
      [
        {
          "name": "create_new_skill_repo",
          "prompt": "Create a new skill repo from scratch called 'my-awesome-skill' with proper Netresearch structure",
          "assertions": [
            {
              "type": "content",
              "pattern": "\\.claude-plugin/plugin\\.json"
            },
            {
              "type": "content",
              "pattern": "SKILL\\.md"
            },
            {
              "type": "content",
              "pattern": "LICENSE-MIT"
            },
            {
              "type": "content",
              "pattern": "LICENSE-CC-BY-SA-4\\.0"
            },
            {
              "type": "content",
              "pattern": "composer\\.json"
            },
            {
              "type": "content",
              "pattern": "release\\.yml"
            }
          ]
        },
        {
          "name": "validate_skill_repo_structure",
          "prompt": "Our skill repo has SKILL.md, plugin.json, composer.json, README.md and a single LICENSE file. Validate the structure and tell me what is wrong.",
          "assertions": [
            {
              "type": "content_regex",
              "value": "validate-skill\\.sh",
              "description": "Points at the repo's own validator rather than eyeballing the tree"
            },
            {
              "type": "content_regex",
              "value": "LICENSE-MIT",
              "description": "Names the code half of the split licence"
            },
            {
              "type": "content_regex",
              "value": "LICENSE-CC-BY-SA-4\\.0",
              "description": "Names the content half — a single LICENSE file is the actual defect here"
            },
            {
              "type": "content_regex",
              "value": "ai-agent-skill",
              "description": "Names the composer type the validator requires"
            }
          ],
          "samples": {
            "passing": "Run scripts/validate-skill.sh. The single LICENSE file is the defect: Netresearch skill repos ship LICENSE-MIT for code and LICENSE-CC-BY-SA-4.0 for content, and the validator warns while both exist. It also checks that composer.json has type ai-agent-skill and that plugin.json and composer.json agree on the version.",
            "failing": [
              "The structure looks complete — SKILL.md, plugin.json, composer.json, README.md and LICENSE are all present, so there is nothing to fix.",
              "Add a CONTRIBUTING.md and a CHANGELOG.md; those are the usual missing pieces."
            ]
          }
        },
        {
          "name": "split_licensing_setup",
          "prompt": "What license files does a Netresearch skill repo need and what SPDX expression goes in composer.json?",
          "assertions": [
            {
              "type": "content",
              "pattern": "LICENSE-MIT"
            },
            {
              "type": "content",
              "pattern": "LICENSE-CC-BY-SA-4\\.0"
            },
            {
              "type": "content",
              "pattern": "\\(MIT AND CC-BY-SA-4\\.0\\)"
            },
            {
              "type": "content",
              "pattern": "Netresearch DTT GmbH"
            }
          ]
        },
        {
          "name": "migrate_single_license",
          "prompt": "This repo has a single LICENSE file. Migrate it to the split licensing model required by Netresearch skill repos.",
          "assertions": [
            {
              "type": "content",
              "pattern": "LICENSE-MIT"
            },
            {
              "type": "content",
              "pattern": "LICENSE-CC-BY-SA-4\\.0"
            },
            {
              "type": "content",
              "pattern": "(remove|delete|old).*LICENSE"
            }
          ]
        },
        {
          "name": "composer_json_setup",
          "prompt": "Set up composer.json for a skill repo called 'netresearch/typo3-testing-skill'. What fields are required?",
          "assertions": [
            {
              "type": "content",
              "pattern": "ai-agent-skill"
            },
            {
              "type": "content",
              "pattern": "composer-agent-skill-plugin"
            },
            {
              "type": "content",
              "pattern": "netresearch/typo3-testing-skill"
            },
            {
              "type": "content",
              "pattern": "SKILL\\.md"
            }
          ]
        },
        {
          "name": "composer_no_version_field",
          "prompt": "Should composer.json in a Netresearch skill repo carry a version field?",
          "assertions": [
            {
              "type": "content_regex",
              "value": "(?i)(no|never|omit|must not|should not)",
              "description": "Rejects the field"
            },
            {
              "type": "content_regex",
              "value": "(?i)(tag|packagist)",
              "description": "Names where the version comes from instead"
            },
            {
              "type": "content_regex",
              "value": "plugin\\.json",
              "description": "Names the manifest that DOES carry the version, which is the part that surprises people"
            },
            {
              "type": "content_regex",
              "value": "(?i)(parity|match|in sync|check-plugin-version)",
              "description": "Names the parity requirement between that manifest and the tag"
            }
          ],
          "samples": {
            "passing": "No. Packagist derives the version from the git tag, so a version field in composer.json only goes stale. The version lives in plugin.json instead, and it must match the tag — Build/Scripts/check-plugin-version.sh enforces that parity, so a release with a mismatched plugin.json fails before it ships.",
            "failing": [
              "No — Composer packages should never declare a version; it is derived from the tag. Nothing else to say.",
              "Yes, set it and bump it on every release so consumers can see the version in composer.json."
            ]
          }
        },
        {
          "name": "composer_no_lock_file",
          "prompt": "Should a Netresearch skill repo commit composer.lock?",
          "assertions": [
            {
              "type": "content_regex",
              "value": "(?i)(no|never|must not|should not)",
              "description": "Rejects committing it"
            },
            {
              "type": "content_regex",
              "value": "(?i)(librar|package|not an application|application)",
              "description": "Gives the reason: locks belong to applications, and a skill repo is a package"
            },
            {
              "type": "content_regex",
              "value": "\\.gitignore",
              "description": "Names where it is kept out, rather than leaving it as advice"
            },
            {
              "type": "content_regex",
              "value": "vendor/",
              "description": "Names the second artefact that must not be committed either"
            }
          ]
        },
        {
          "name": "plugin_json_structure",
          "prompt": "Create a valid plugin.json for a skill called 'security-audit' in .claude-plugin/",
          "assertions": [
            {
              "type": "content",
              "pattern": "\"name\""
            },
            {
              "type": "content",
              "pattern": "\"version\""
            },
            {
              "type": "content",
              "pattern": "\"skills\""
            },
            {
              "type": "content",
              "pattern": "\\./skills/"
            },
            {
              "type": "content",
              "pattern": "netresearch\\.de"
            }
          ]
        },
        {
          "name": "skillmd_frontmatter_format",
          "prompt": "Write valid SKILL.md frontmatter for a skill named 'docker-development'.",
          "assertions": [
            {
              "type": "content_regex",
              "value": "name: docker-development",
              "description": "Sets the name to the skill's directory name"
            },
            {
              "type": "content_regex",
              "value": "(?i)description:.{0,40}use when",
              "description": "Follows the convention that the description opens with the trigger"
            },
            {
              "type": "content_regex",
              "value": "(?i)(1,?536|300 char|100.{0,5}300)",
              "description": "Knows the description budget rather than writing an arbitrary blurb"
            },
            {
              "type": "content_regex",
              "value": "(?i)(license|metadata|version)",
              "description": "Includes the metadata block Netresearch skill repos require"
            }
          ]
        },
        {
          "name": "skillmd_word_limit",
          "prompt": "What is the word limit for SKILL.md, how is it counted, and where does the rest of the content go?",
          "assertions": [
            {
              "type": "content_regex",
              "value": "500",
              "description": "Names the limit"
            },
            {
              "type": "content_regex",
              "value": "references/",
              "description": "Sends the overflow to reference files"
            },
            {
              "type": "content_regex",
              "value": "(?i)(whole file|entire file|including|frontmatter)",
              "description": "Knows the count covers the whole file including frontmatter, which is what makes an added table row overflow it"
            },
            {
              "type": "content_regex",
              "value": "(?i)(validate-skill|audit-skills)",
              "description": "Names the script that enforces it, rather than describing it as a guideline"
            }
          ],
          "samples": {
            "passing": "500 words, counted by validate-skill.sh over the whole file including the frontmatter — which is why adding one references-table row can push a passing SKILL.md over. Extended content goes to references/*.md, and every reference must be reachable from SKILL.md or audit-skills.sh reports it as an orphan.",
            "failing": [
              "Keep SKILL.md under 500 words of body text and move anything longer into references/.",
              "There is no hard limit, but shorter is better — aim for a page or two."
            ]
          }
        },
        {
          "name": "reusable_workflow_caller",
          "prompt": "Create a caller workflow for validate.yml that uses the reusable workflow from skill-repo-skill.",
          "assertions": [
            {
              "type": "content_regex",
              "value": "netresearch/skill-repo-skill/\\.github/workflows/validate\\.yml@main",
              "description": "Uses the org reusable at @main"
            },
            {
              "type": "content_regex",
              "value": "(?i)(@main|not.{0,20}pin|sha)",
              "description": "Addresses the ref explicitly rather than silently SHA-pinning an org-owned reusable"
            },
            {
              "type": "content_regex",
              "value": "permissions:",
              "description": "Declares the permissions the called workflow needs, without which the run fails at startup"
            },
            {
              "type": "content_regex",
              "value": "(?i)(pull_request|push)",
              "description": "Wires the triggers"
            }
          ],
          "samples": {
            "passing": "jobs: validate: uses: netresearch/skill-repo-skill/.github/workflows/validate.yml@main with permissions: contents: read, triggered on push and pull_request. Netresearch-owned reusables stay on @main on purpose so upstream fixes propagate — the SHA-pin rule applies to third-party actions, not to these.",
            "failing": [
              "Pin it to a commit SHA like netresearch/skill-repo-skill/.github/workflows/validate.yml@a1b2c3d for supply-chain safety.",
              "Copy the jobs from validate.yml into your own workflow file so the repo is self-contained."
            ]
          }
        },
        {
          "name": "release_workflow_setup",
          "prompt": "What release workflow does a skill repo need and how is it triggered?",
          "assertions": [
            {
              "type": "content",
              "pattern": "release\\.yml"
            },
            {
              "type": "content",
              "pattern": "(tag|v\\*|push)"
            },
            {
              "type": "content",
              "pattern": "plugin\\.json"
            }
          ]
        },
        {
          "name": "release_process_steps",
          "prompt": "Walk me through releasing version 2.1.0 of a Netresearch skill repo.",
          "assertions": [
            {
              "type": "content_regex",
              "value": "plugin\\.json",
              "description": "Bumps the manifest that carries the version"
            },
            {
              "type": "content_regex",
              "value": "v2\\.1\\.0",
              "description": "Tags with the v prefix"
            },
            {
              "type": "content_regex",
              "value": "(?i)(parity|match|check-plugin-version|same version)",
              "description": "Names the parity between manifest and tag, which is what the release gate checks"
            },
            {
              "type": "content_regex",
              "value": "(?i)(-s\\b|sign-off|signed-off|DCO)",
              "description": "Keeps the sign-off requirement, which the repo's CI enforces"
            }
          ]
        },
        {
          "name": "installation_methods",
          "prompt": "What are the three ways to install a Netresearch skill?",
          "assertions": [
            {
              "type": "content",
              "pattern": "(marketplace|Marketplace)"
            },
            {
              "type": "content",
              "pattern": "(composer|Composer)"
            },
            {
              "type": "content",
              "pattern": "(release|download|Release)"
            }
          ]
        },
        {
          "name": "multi_skill_package",
          "prompt": "How do I set up composer.json and plugin.json for one repo containing two skills, jira-communication and jira-syntax?",
          "assertions": [
            {
              "type": "content_regex",
              "value": "skills/jira-communication|skills/jira-syntax",
              "description": "Uses one directory per skill under skills/"
            },
            {
              "type": "content_regex",
              "value": "(?i)\"skills\"\\s*:\\s*\\[|skills.{0,15}array",
              "description": "Lists both in the plugin manifest's skills array"
            },
            {
              "type": "content_regex",
              "value": "(?i)(one|single) (composer )?package|one repo|single package",
              "description": "Keeps it a single Composer package rather than splitting it"
            },
            {
              "type": "content_regex",
              "value": "(?i)(extra|ai-agent-skill)",
              "description": "Names the composer extra that points at the skill paths"
            }
          ]
        },
        {
          "name": "cross_platform_scripts",
          "prompt": "What portability rules must shell scripts in a skill repo follow, and why?",
          "assertions": [
            {
              "type": "content_regex",
              "value": "sed -i",
              "description": "Names the trap that actually bites"
            },
            {
              "type": "content_regex",
              "value": "(?i)(BSD|macOS)",
              "description": "Names the platform the CI matrix adds"
            },
            {
              "type": "content_regex",
              "value": "(?i)(matrix|runs-on|CI)",
              "description": "Ties the rule to the matrix rather than to general good practice"
            },
            {
              "type": "content_regex",
              "value": "(?i)(perl -i|temp file|tmp file|'' )",
              "description": "Names a portable replacement instead of only forbidding the GNU form"
            }
          ],
          "samples": {
            "passing": "The CI matrix runs on macOS as well as Linux, so every shipped script has to be BSD-portable, not just GNU-portable. The classic trap is sed -i: GNU takes sed -i 's/a/b/' file, BSD requires sed -i '' 's/a/b/' file, and each form fails on the other platform. Write to a temp file and move it, or use perl -i -pe.",
            "failing": [
              "Use POSIX sh, avoid bashisms, and quote your variables.",
              "Prefix GNU tools with g (gsed, ggrep) so macOS behaves like Linux."
            ]
          }
        },
        {
          "name": "license_path_mapping",
          "prompt": "Which license applies to scripts/*.sh files and which applies to references/*.md files in a skill repo?",
          "assertions": [
            {
              "type": "content",
              "pattern": "MIT"
            },
            {
              "type": "content",
              "pattern": "CC-BY-SA"
            }
          ]
        },
        {
          "name": "fix_validation_errors",
          "prompt": "validate-skill.sh reports: ERROR: composer.json type must be 'ai-agent-skill' and ERROR: plugin.json version does not match composer.json. Fix both.",
          "assertions": [
            {
              "type": "content_regex",
              "value": "\"type\"\\s*:\\s*\"ai-agent-skill\"|type.{0,15}ai-agent-skill",
              "description": "Sets the composer type to the required value"
            },
            {
              "type": "content_regex",
              "value": "(?i)(plugin\\.json).{0,80}(version)|version.{0,40}plugin\\.json",
              "description": "Fixes the version on the manifest side rather than guessing which file is authoritative"
            },
            {
              "type": "content_regex",
              "value": "(?i)(re-?run|again|verify).{0,30}(validate|script)|validate-skill\\.sh",
              "description": "Re-runs the validator instead of declaring it fixed"
            }
          ]
        },
        {
          "name": "readme_requirements",
          "prompt": "What sections must a Netresearch skill repo README.md contain?",
          "assertions": [
            {
              "type": "content_regex",
              "value": "## What this skill solves",
              "description": "Uses the exact level-2 heading agents grep for"
            },
            {
              "type": "content_regex",
              "value": "(?i)model delta",
              "description": "Includes the section that justifies the skill against the base model"
            },
            {
              "type": "content_regex",
              "value": "(?i)example prompts",
              "description": "Includes the example-prompts section"
            },
            {
              "type": "content_regex",
              "value": "(?i)(three|3)\\b",
              "description": "Knows the minimum number of example prompts"
            }
          ],
          "samples": {
            "passing": "Exact level-2 headings so agents can grep them: ## What this skill solves, ## Why this is a skill (model delta) — one sentence on what the model gets wrong without it plus a pointer to eval evidence — ## Use when, ## Expected outputs, ## Context requirements, and ## Example prompts with a minimum of three distinct realistic prompts.",
            "failing": [
              "Installation, Usage, Configuration, Contributing and License — the standard open-source README sections.",
              "Anything that helps a reader; there is no fixed structure, just describe the skill and how to install it."
            ]
          }
        },
        {
          "name": "auto_merge_workflow",
          "prompt": "Set up the auto-merge workflow for dependency PRs in a skill repo",
          "assertions": [
            {
              "type": "content",
              "pattern": "auto-merge-deps"
            },
            {
              "type": "content",
              "pattern": "netresearch/skill-repo-skill"
            },
            {
              "type": "content",
              "pattern": "(dependabot|renovate|pull_request)"
            }
          ]
        },
        {
          "name": "retire_superseded_skill",
          "prompt": "The 'example-legacy' skill is superseded by automation and must be decommissioned. Walk through the retirement.",
          "assertions": [
            {
              "type": "content",
              "pattern": "marketplace\\.json"
            },
            {
              "type": "content",
              "pattern": "[Dd]eprecat"
            },
            {
              "type": "content",
              "pattern": "archive"
            },
            {
              "type": "content",
              "pattern": "plugin uninstall"
            }
          ]
        },
        {
          "name": "renovate_regex_manager_verify_after_flag_change",
          "prompt": "I added --no-build to the uvx ruff invocation in .github/workflows/validate.yml. The renovate.json customManager pins that version. Anything to do?",
          "assertions": [
            {
              "type": "content",
              "pattern": "(?i)run the pattern against the real file|check the count|matchStrings"
            },
            {
              "type": "content",
              "pattern": "(?i)silent|no dependency found|nothing to update"
            }
          ]
        },
        {
          "name": "sast_job_interpreter_bounds_the_scan",
          "prompt": "Our bandit job builds its venv with `python3 -m venv` on ubuntu-latest and the scan is green. Is that enough?",
          "assertions": [
            {
              "type": "content",
              "pattern": "(?i)runner('s)? (default|image)|oldest"
            },
            {
              "type": "content",
              "pattern": "(?i)skipped|exit 0|scanned nothing|parse"
            },
            {
              "type": "content",
              "pattern": "(?i)uv venv|--seed|newest"
            }
          ]
        },
        {
          "name": "readme_static_version_badge",
          "prompt": "Our skill README starts with [![Version](https://img.shields.io/badge/version-3.0.0-blue.svg)](https://github.com/netresearch/typo3-testing-skill). Does validate-skill.sh say anything about it, and what should the badge be?",
          "assertions": [
            {
              "type": "content_regex",
              "value": "(?i)warn",
              "description": "Knows the validator reports the static badge as a warning"
            },
            {
              "type": "content_regex",
              "value": "img\\.shields\\.io/github/v/release/netresearch/",
              "description": "Names the live release badge form"
            },
            {
              "type": "content_regex",
              "value": "sort=semver",
              "description": "Keeps the semver sort so a pre-release or an old maintenance tag does not win by date"
            },
            {
              "type": "content_regex",
              "value": "(?i)(no|not|never).{0,40}(release step|bump-version|check-version-parity|updated)",
              "description": "Explains the drift: no release step updates README.md"
            }
          ],
          "samples": {
            "passing": "Yes: validate-skill.sh warns on an img.shields.io/badge/version- URL, because no release step updates README.md (bump-version.sh and check-version-parity.sh never read it), so the hand-typed 3.0.0 drifts from the latest tag. Use the live form [![Release](https://img.shields.io/github/v/release/netresearch/typo3-testing-skill?sort=semver)](https://github.com/netresearch/typo3-testing-skill/releases).",
            "failing": [
              "The badge is fine; just remember to edit the version number in the README whenever you cut a release.",
              "Badges are cosmetic and the validator ignores README content apart from the headings, so leave it as it is."
            ]
          }
        }
      ]
      
  • references
    • agent-plugins-compat.md 6.6 KB
      # Agent Plugins 1.0.0 compatibility
      
      ## Contents
      
      - Why two files
      - The portable manifest
      - Generating the Claude manifest
      - Migrating a repo
      - What validation enforces
      - Out of scope
      
      Netresearch skill repos ship **two** manifests. This page says what each one is
      for, which is authoritative, and how to migrate a repo that has only the
      Claude Code one.
      
      Spec: <https://agent-plugins.org/specification> ·
      Schema: <https://agent-plugins.org/schemas/1.0.0/plugin.schema.json> ·
      Skill format: <https://agentskills.io/specification>
      
      ## Why two files
      
      | File | Read by | Role |
      |---|---|---|
      | `./plugin.json` | every Agent Plugins client (Cursor, Copilot, …) | **source of truth** for shared metadata |
      | `./.claude-plugin/plugin.json` | Claude Code only | generated projection + Claude-only keys |
      
      Claude Code reads its manifest from `.claude-plugin/plugin.json` and nowhere
      else; Agent Plugins clients read `plugin.json` at the package root and nowhere
      else. One file cannot serve both: the portable schema is **closed**
      (`additionalProperties: false`), so `skills`, `agents`, `commands`,
      `outputStyles`, `hooks`, `mcpServers` and `metadata` are schema violations
      there. Conversely Claude Code ignores fields it does not know, which is why the
      projection into `.claude-plugin/plugin.json` is safe.
      
      `support` belongs in **neither** file. It is a composer-ism, not a Claude Code
      manifest field — the manifest reference documents `homepage` for a
      documentation URL and `metadata` as the free-form object, and nothing named
      `support`. `claude plugin validate --strict` reports `Unknown field 'support'`
      and exits 1; the claude.ai marketplace importer strips it and emits one warning
      per plugin. Put a contact URL in `homepage`, catalogue data in `metadata`, and
      issues under the `repository` host where they already live.
      
      Removing it once is not enough while a reviewer can suggest it back:
      `typo3-ddev-skill` dropped the key deliberately in `46b5fd6` ("fix: remove
      unsupported 'support' field from plugin.json", 2026-02-24) and a Copilot review
      reinstated it five weeks later in `1ae6b43`. Seven of the forty plugins in
      `netresearch/claude-code-marketplace` still carried it on 2026-09-17.
      
      Skills need no change: `skills/<name>/SKILL.md` is what both specs discover.
      
      ## The portable manifest
      
      ```json
      {
        "$schema": "https://agent-plugins.org/schemas/1.0.0/plugin.schema.json",
        "name": "skill-repo",
        "version": "1.27.0",
        "description": "Guide for structuring Netresearch skill repositories",
        "author": {
          "name": "Netresearch DTT GmbH",
          "url": "https://www.netresearch.de"
        },
        "repository": "https://github.com/netresearch/skill-repo-skill",
        "license": "(MIT AND CC-BY-SA-4.0)"
      }
      ```
      
      Allowed top-level fields — nothing else:
      
      | Field | Required | Notes |
      |---|---|---|
      | `$schema` | yes | exactly the 1.0.0 URL above |
      | `name` | yes | 1–64 chars, `a-z0-9` plus `-` and `.`, starts and ends alphanumeric, no `--` or `..`. Matches the Claude manifest name (and, for single-skill repos, the SKILL.md `name`) |
      | `version` | no | SemVer; same value as `.claude-plugin/plugin.json` and SKILL.md `metadata.version` |
      | `description` | no | short purpose statement |
      | `author` | no | object with `name`, `email`, `url` only — `"author": "Name"` is invalid |
      | `homepage`, `repository`, `license` | no | strings; license stays the SPDX expression `(MIT AND CC-BY-SA-4.0)` |
      | `keywords` | no | array of strings |
      | `extensions` | no | object keyed by **reverse-domain namespaces owned by the client vendor**. We do not invent namespaces for clients — Claude-specific data lives in `.claude-plugin/`, not here |
      
      ## Generating the Claude manifest
      
      ```bash
      bash skills/skill-repo/scripts/sync-plugin-manifest.sh            # write
      bash skills/skill-repo/scripts/sync-plugin-manifest.sh --check    # CI gate
      bash skills/skill-repo/scripts/sync-plugin-manifest.sh --repo DIR # fleet driver
      ```
      
      The script copies `name`, `version`, `description`, `author`, `homepage`,
      `repository`, `license` and `keywords` into `.claude-plugin/plugin.json`,
      preserves every Claude-only key already there, and never copies `$schema` or
      `extensions`. `--check` compares those shared values in both directions — key
      order and formatting in the Claude manifest are not enforced. Without a root
      `plugin.json` it is a no-op, so it is safe to run across repos that have not
      migrated yet.
      
      ## Migrating a repo
      
      1. Create `./plugin.json` with the fields above, copying the current values out
         of `.claude-plugin/plugin.json`. Drop `skills`/`agents` — they stay in the
         Claude manifest; delete `support` from both. Copy **every** shared field: a root manifest missing
         one (`keywords` is the one that gets dropped in practice, because minimal
         manifests from repos that never had it look like templates) fails `--check`,
         and a plain sync then *deletes* the missing key from the Claude manifest —
         that is data loss that ships, not a rendering difference (claude-coach-plugin
         v2.6.0 lost its `keywords` exactly this way, 2026-08-13).
      2. `sync-plugin-manifest.sh --check` — it passes without rewriting anything when
         the values were copied correctly. Run it without `--check` only when you want
         the canonical rendering.
      3. If the repo ships a root `SKILL.md` instead of `skills/<name>/SKILL.md`, move
         it: Agent Plugins clients do not discover a root `SKILL.md`. Claude Code
         loads either layout, so the move is safe there — but the plugin's skill name
         then comes from the directory, so keep the directory name equal to the
         frontmatter `name`.
      4. `bash skills/skill-repo/scripts/validate-skill.sh .` — 0 errors.
      
      ## What validation enforces
      
      `validate-skill.sh` checks, when `./plugin.json` exists: the `$schema` constant,
      the `name` pattern, that no field outside the closed schema is present, the
      `author`/`keywords`/`extensions` shapes, that the shared fields match
      `.claude-plugin/plugin.json`, and that at least one `skills/<name>/SKILL.md`
      exists. A **missing** `./plugin.json` is an error: the fleet finished adopting
      the manifest on 2026-08-07, so a repo without one is a new gap rather than a
      repo still waiting its turn.
      
      `check-version-parity.sh` treats the root `plugin.json` version as
      authoritative and fails when `.claude-plugin/plugin.json` disagrees.
      
      ## Out of scope
      
      Agent Plugins defines packaging only. Distribution, installation, marketplaces
      and permissions stay client-specific, so `claude-code-marketplace` entries and
      the composer/npm channels are unaffected by this standard.
      
      An `mcp.json` at the package root is the portable way to ship MCP servers. No
      Netresearch skill repo ships one today; add it only alongside a real server, and
      keep its `$schema` version equal to the one in `plugin.json`.
      
    • authoring-ci-gotchas.md 18.9 KB
      # Authoring & CI Gotchas
      
      ## Contents
      
      - Word-budget-first authoring
      - macOS / BSD portability for shell **and** test scripts
      - Lint without dirtying the worktree
      - "Skill Validation" can run more than once — wait for all of it
      - Installing the Claude Code CLI with `--ignore-scripts`
      - Generated YAML: exactly one trailing newline
      - `MD010` breaks copy-pasted Makefile snippets — exempt the fence, don't fake the tab
      - A validator over a structured value parses it; it does not pattern-match it
      - SonarCloud's `shelldre` rules fire on our own shell conventions — triage, don't comply
      - `uv pip compile --universal` drops marker-conditional deps — pin the floor
      - A Renovate regex manager fails silently in both directions
      - A SAST job's interpreter bounds what it scans, and says nothing
      - A new script is committed `100755`, not just `chmod +x` locally
      
      Process learnings from cross-session retrospectives (§§1–10 from 2026-06-27,
      §§11–13 from 2026-09-17). Companion to
      [`skill-quality.md`](skill-quality.md) (SKILL.md sizing) and the
      [`validation-checklist.md`](validation-checklist.md) (pre-completion checks).
      Each item is a habit that prevents a costly redo, not a structural rule.
      
      ## 1. Word-budget-first authoring
      
      This repo enforces a **hard 500-word cap on `SKILL.md`** (`scripts/validate-skill.sh`
      fails above 500; this repo's SKILL.md sits at ~495). The body also persists in context
      for the whole session, so every word is paid on each invocation — see
      [`skill-quality.md`](skill-quality.md) for the full cost model.
      
      **Before editing an existing `SKILL.md`, `wc -w` it first.** If it is already near
      the ceiling, decide the **architecture before you type**:
      
      - new content that is a lookup/detail → a `references/*.md` file (lazy-loaded, free
        until cited),
      - a genuinely separate capability → a new standalone skill,
      - only then, prose edits to the body itself.
      
      Do **not** open by iteratively trimming the maintainer's existing prose to free up
      room. In a real 2026-06-27 case that approach burned ~7 edit cycles shaving words off
      carefully-written copy before the obvious move — put the addition in a reference file
      and add one catalog line — was taken. Architecture decision first, word-shaving last
      (and rarely).
      
      ## 2. macOS / BSD portability for shell **and** test scripts
      
      This repo's CI matrix runs on macOS, not just Linux: `validate-agents.yml` defaults its
      `os-matrix` to `["ubuntu-latest", "macos-latest"]`. Any `*.sh` (and any test script the
      workflows execute) therefore has to be **BSD/macOS-portable**, not just GNU-portable.
      
      The classic trap a reviewer flagged: **`sed -i` is GNU-only.** BSD `sed` (macOS) requires
      an explicit empty backup-suffix argument:
      
      ```sh
      sed -i 's/a/b/' file      # GNU only — fails on macOS
      sed -i '' 's/a/b/' file   # BSD only — fails on GNU
      ```
      
      Prefer a form that needs no in-place flag at all, e.g. write to a temp file and move it,
      or use `perl -i -pe`. This complements the portability notes already in `SKILL.md`
      (`grep -E` not `-P`; `bash` shebangs not `zsh`; `[[ ]]` conditionals). If the CI matrix
      ever drops macOS these become Linux-only conveniences again — verify the matrix before
      relying on a GNU-ism.
      
      ## 3. Lint without dirtying the worktree
      
      Running the JS-based linters locally (e.g. `bunx markdownlint-cli2`) can make `bun`
      resolve and **write a lockfile**, emitting `Saved lockfile` and leaving the working tree
      dirty — exactly when you are about to commit or merge and want a clean status.
      
      Run a frozen/no-save install first so the lint step touches nothing:
      
      ```sh
      bun install --frozen-lockfile   # then run the linter
      # or invoke the linter in a no-save mode
      ```
      
      (CI itself runs markdownlint via the pinned `markdownlint-cli2-action`, so this is a
      **local-authoring** hygiene step — keep the pre-commit / pre-merge `git status` clean.)
      
      ## 4. "Skill Validation" can run more than once — wait for all of it
      
      The skill-validation gate can surface as **more than one run/check context** for a single
      PR (this repo wires it through the reusable `validate.yml`, and `lint.yml` triggers on
      both `push` to `main` and `pull_request`; reusable-workflow nesting can add further
      contexts). A lagging second instance can keep a PR `BLOCKED` for a short while **after
      the first one has already gone green**.
      
      Before concluding the gate is stuck, confirm **every** validation run/check has reported
      — `gh pr checks <n>` — rather than acting on the first green. Re-running or "fixing" a gate
      that is merely still finishing wastes a round-trip.
      
      ## 5. Installing the Claude Code CLI with `--ignore-scripts`
      
      Security scanners (SonarCloud `githubactions:S6505`) require the
      `--ignore-scripts` flag on `npm install` in workflows — but that breaks
      `@anthropic-ai/claude-code`, whose **postinstall downloads the platform-native
      binary**; the CLI then exits with "claude native binary not installed". The package documents its own
      sanctioned two-step (verified with 2.1.206):
      
      ```bash
      npm install -g --ignore-scripts @anthropic-ai/claude-code@<pinned-version>
      node "$(npm root -g)/@anthropic-ai/claude-code/install.cjs"
      claude --version   # proves the binary is in place
      ```
      
      This blocks lifecycle scripts of the whole dependency tree while running only
      the CLI's own vetted installer, explicitly.
      
      ## 6. Generated YAML: exactly one trailing newline
      
      The reusable `validate.yml` runs yamllint, whose **default config** (`extends:
      default`, `empty-lines: max-end: 0`) rejects trailing blank lines; the workflow
      writes that default only when the repo ships no `.yamllint*` of its own, so a
      repo config can override it — most skill repos don't. Batch-generated YAML
      (heredoc, `echo`, templating) routinely picks up a trailing blank line. One deploy of `auto-merge-deps.yml` across
      22 repos failed CI in every one of them on exactly this.
      
      When writing YAML programmatically, emit the content with a single trailing
      newline and verify before committing:
      
      ```bash
      printf '%s\n' "$CONTENT" > file.yml     # not: echo "$CONTENT" > file.yml
      tail -c 2 file.yml | xxd -p             # must NOT be 0a0a
      ```
      
      ## 7. `MD010` breaks copy-pasted Makefile snippets — exempt the fence, don't fake the tab
      
      A ` ```makefile ` fenced block showing a real recipe line needs a literal tab —
      `make` rejects a space there with `*** missing separator. Stop.` But
      markdownlint's `MD010` (no-hard-tabs) flags a hard tab **inside a fenced code
      block by default**, not just in prose. Substituting a single leading space to
      keep the linter quiet (observed in a skill repo's own docs, twice, across two
      separate snippets) produces a snippet that reads clean but silently fails the
      moment someone copies it into a real `Makefile`.
      
      The fix is a linter exemption, not a fake tab:
      
      ```jsonc
      // .markdownlint-cli2.jsonc — add the key to the EXISTING "config" object.
      // A separate .markdownlint.jsonc file does not merge with .markdownlint-cli2.jsonc's
      // "config" — it replaces it wholesale, silently re-enabling every other rule this
      // repo already disables there (MD013, MD033, etc.).
      {
        "config": {
          "MD010": { "ignore_code_languages": ["makefile"] }
          // ...alongside this repo's other existing "config" entries
        }
      }
      ```
      
      This keeps `MD010` enforcing real prose/other-language blocks while letting a
      ` ```makefile ` fence carry an actual, pastable tab. Verify the fix reproduces
      correctly before trusting it — write the fenced snippet to a scratch file and
      run `make -n -f <scratch-file>` against it (`-n` alone silently looks for
      `Makefile`/`makefile`/`GNUmakefile` in the current directory and ignores an
      arbitrarily named scratch file); a `make` that resolves the target confirms
      the tab survived, a lint pass alone does not.
      
      ---
      
      ## 8. A validator over a structured value parses it; it does not pattern-match it
      
      A check added to `validate-skill.sh` read `allowed-tools` with one anchored
      `grep`. It went through four review rounds, and each one found another **legal
      spelling of the same value** the pattern did not anticipate:
      
      | Round | Spelling that slipped past |
      |---|---|
      | 1 | folded scalar `>-`, and the YAML list form — the check read only the key line |
      | 2 | a blank line inside the value, which ended collection; `Bash,Read` and `"Bash"`, where the boundary was whitespace-only |
      | 3 | a YAML comment mentioning what it does *not* grant, matched as if it did |
      | 4 | flow list `[Bash]`, where `]` was not a delimiter |
      
      Widening the pattern each round buys one shape. The signal that the *approach*
      is wrong rather than incomplete is the repetition itself: the value never
      changed, only its spelling.
      
      What worked was taking the value apart instead:
      
      1. **Collect** the key line plus every continuation, ending only at a new
         top-level key — so a blank line inside a folded scalar keeps the value open.
      2. **Drop** comment lines and everything after a space-`#`, which is where a YAML
         comment starts. Anchoring on `[^)]*$` here fails on a comment that contains
         a parenthesis, which is exactly what a comment about `Bash(python3:*)` does.
      3. **Split** on whitespace, comma, bracket and quote — but only at parenthesis
         depth 0, so `Bash(git:*,make:*)` stays one entry and
         `Bash(bash ${CLAUDE_SKILL_DIR}/scripts/*)` survives its space.
      4. **Match each entry on its own**, anchored.
      
      `[Bash]` then falls out without a special case, because the bracket is a
      delimiter like any other. Two of the four rounds also broke *existing* passing
      tests when the split was naive — the suite is what caught it, so write the
      shape cases (plain, folded, literal, block list, flow list, comma-separated,
      quoted, commented) before widening anything.
      
      ## 9. SonarCloud's `shelldre` rules fire on our own shell conventions — triage, don't comply
      
      Every skill repo here is mostly shell and Python and runs SonarCloud, so a PR
      touching one script reliably reports a dozen new MAJOR code smells while the
      quality gate passes. The count is alarming and the content is not: on
      git-workflow-skill#300, sixteen new issues over a 96-line diff were
      `shelldre:S7688` (use `[[` instead of `[`), `shelldre:S7679` (assign positional
      parameters to local variables) and one `shelldre:S7682` (add an explicit
      `return` at the end of the function).
      
      The first two are house style, not drift. `pr-status.sh` uses `[` twenty-one
      times and `[[` not once, and the `check`/`check_contains` helpers are copied
      verbatim between test files. "Fixing" four new lines makes them the only ones of
      their kind in the file, which is worse than the finding.
      
      `S7682` is the one to read rather than skim, because complying with it can
      introduce the bug:
      
      ```bash
      collect_raw() {
        gh api graphql -f owner="$OWNER" … -f query='…'
      }        # no explicit return: the function's status IS gh's status
      ```
      
      The caller reads that status (`out=$(collect_raw 2>"$err"); rc=$?`) to tell a
      failed query from a successful one. `return 0` there would report every failure
      as a success — precisely the defect that PR was fixing. Adding `return $?` is a
      no-op that satisfies a linter and says nothing.
      
      So: read what the rule asks against what the code promises, resolve the
      intentional ones in the SonarCloud UI rather than contorting the code, and say
      in the PR which findings stand and why. A reviewer seeing "16 new issues" with
      no explanation has to re-derive that triage themselves.
      
      Pick the status by what is actually true of the finding — these are ordinary
      issues (`type: CODE_SMELL` from `api/issues`), not Security Hotspots, which live
      on their own endpoint with their own `Safe` / `Fixed` / `Acknowledged` review:
      
      - **Accept** — the rule read the code correctly and we are keeping it anyway.
        That is the house-style case, `S7688` and `S7679` above.
      - **False positive** — the analysis itself does not hold. `S7682` on a function
        whose exit status is its contract belongs here: the rule's premise, that a
        missing `return` is an oversight, is wrong for that function.
      
      `Safe` is not available for either; reaching for it means you are in the
      hotspot review by mistake.
      
      ## 10. `uv pip compile --universal` drops marker-conditional deps — pin the floor
      
      `--universal` is not "resolve for every Python". It resolves from the running
      interpreter's version upward, so a dependency whose environment marker excludes
      that version is omitted from the lock entirely — silently, with exit 0.
      
      The omission surfaces where the lock is *consumed*, not where it is written. Our
      hash-locked tool envs install with `--require-hashes`, and pip refuses the whole
      file rather than the one line:
      
      ```
      ERROR: In --require-hashes mode, all requirements must have their versions
      pinned with ==. These do not:
          typing_extensions<5.0,>=4.6 ... (from cyclonedx-python-lib==11.11.0)
      ```
      
      `cyclonedx-python-lib` declares `typing_extensions ; python_version < "3.13"`.
      Compiled on 3.14 the marker excludes it and the entry is never written; the
      audit job runs on the runner's default `python3` (3.12), which needs it. Compile
      and install therefore disagree, and only the install fails.
      
      Pass the floor the consumer actually runs on:
      
      ```bash
      uv pip compile --universal --generate-hashes --python-version 3.12 \
        requirements.in -o .github/requirements/pip-audit.txt
      ```
      
      Two things follow. **The compile command belongs in a comment next to the file
      it produces, with the floor in it** — ours said only
      `uv pip compile --universal --generate-hashes`, so following it on a newer
      machine reproduced the defect. And because this is shared CI, one stale lock
      fails the job in every consuming repo at once; it presents as one repo's red
      check, since the others have not run since.
      
      Verify on the consumer's interpreter, not yours:
      
      ```bash
      python3.12 -m venv /tmp/probe
      /tmp/probe/bin/pip install --dry-run --require-hashes --only-binary :all: \
        -r .github/requirements/pip-audit.txt
      ```
      
      ## 11. A Renovate regex manager fails silently in both directions
      
      A `customManagers` entry pins ad-hoc tool versions in workflow files
      (`uvx ruff@0.16.0`, `uv run --with pyyaml==6.0.3`) so Renovate opens a PR when
      they move. Both of its failure modes are invisible, because **"no dependency
      found here" and "nothing to update here" produce the same output: none.**
      
      *It stops matching.* Ours was written against `uvx <name>@<version>` with an
      optional `--from <x>` in front. Sixteen minutes later another commit hardened
      the pin to `uvx --no-build ruff@0.16.0`, and the pattern matched nothing from
      then on. The pin sat unmoved for seven weeks while ruff released eight versions,
      and nothing reported it — the repository looked up to date.
      
      *It matches the wrong token.* Widening the prefix to "any word may stand here"
      fixed that and broke the other direction: under a `datasource=pypi` comment,
      `uvx --from git+https://github.com/o/r@2843b87 tool` yields `2843b87` as the
      PyPI version of the package the comment names.
      
      Three habits follow.
      
      **Run the pattern against the real file after every change to an annotated
      command line, and check the count.** One call, and it distinguishes the two
      states the tool cannot:
      
      ```bash
      python3 - <<'PY'
      import json, re, pathlib
      cfg = json.loads(pathlib.Path("renovate.json").read_text())
      text = pathlib.Path(".github/workflows/validate.yml").read_text()
      for mgr in cfg["customManagers"]:          # iterate: a repo may carry several
          for pat in mgr["matchStrings"]:
              hits = [m.group("currentValue") for m in re.finditer(pat.replace("(?<", "(?P<"), text)]
              print(len(hits), hits)
      PY
      ```
      
      **Shape the pinned token, not the prefix.** `[A-Za-z][A-Za-z0-9._-]*@` is a
      package name; `\S+@` is also a URL. Because the repeat consumes whole
      space-separated words, a name can only begin where a word begins, so the `cli@`
      inside `git+https://…/cli@2843b87` is not a candidate and a `--from <source>` is
      excluded without enumerating the tool's flags. That matters: **Renovate runs the
      pattern through RE2, which has no lookahead**, so a flag allow-list would have to
      exclude the value-taking flags from its own generic branch to be safe.
      
      **Pin both directions in a test.** A case list that only asserts what *must*
      match cannot catch the second failure. Assert the shapes that must yield nothing
      as well, and note in the test that translating `(?<name>` to `(?P<name>` for
      Python's `re` is evidence about the pattern, not about RE2 — the proof for that
      half is the bump PR Renovate opens.
      
      ## 12. A SAST job's interpreter bounds what it scans, and says nothing
      
      `python3 -m venv` in a CI job takes the runner image's interpreter. On
      `ubuntu-latest` that is 3.12.3 today — for most repositories the **oldest**
      version their matrix tests, not the newest. Bandit parses with the `ast` of the
      interpreter it runs on, and a file it cannot parse is skipped with a warning
      while the run still exits 0. The gate reports clean instead of reporting that it
      scanned nothing.
      
      Measured with bandit 1.9.4 on one file holding a PEP 696 type-parameter default
      and a shell-injection finding under it:
      
      | Interpreter | Result |
      |---|---|
      | 3.12.14 | `Files skipped (1): syntax error while parsing AST` · exit 0 · no `B602` |
      | 3.14.7 | `B602 subprocess call with shell=True` · exit 1 |
      
      ```python
      import subprocess
      
      
      class Box[T = int]:  # PEP 696, 3.13+
          def run(self, cmd: str) -> None:
              subprocess.call(cmd, shell=True)  # B602
      ```
      
      So derive the interpreter from what the caller tests rather than inheriting it,
      and assert the venv landed on it — a venv on the wrong interpreter runs fine and
      scans less:
      
      ```bash
      BANDIT_PYTHON="$(printf '%s' "$PYTHON_VERSIONS" \
        | jq -r 'max_by(split(".") | map(gsub("[^0-9]";"") | tonumber? // 0))')"
      uv python install "$BANDIT_PYTHON"
      uv venv --seed --python "$BANDIT_PYTHON" "$RUNNER_TEMP/bandit-env"
      want="${BANDIT_PYTHON%%[!0-9.]*}"
      case "$("$RUNNER_TEMP/bandit-env/bin/python" -V)" in
        "Python $want"|"Python $want".*) ;;
        *) echo "::error::bandit venv is not on $BANDIT_PYTHON"; exit 1 ;;
      esac
      ```
      
      Two details that cost a round each. `--seed` is what puts a pip into the venv,
      which is what reads a hash-locked requirements file under `--require-hashes`
      (§10's lock installs unchanged on a newer interpreter — verify, do not assume).
      And `gsub` before `tonumber`: a caller may test a free-threaded build (`3.13t`),
      where a bare `tonumber` aborts the job. Compare on the digits and install the
      value as written — then compare the assertion on the numeric prefix too, since
      `python -V` answers `Python 3.13.x` for a `3.13t` request.
      
      ## 13. A new script is committed `100755`, not just `chmod +x` locally
      
      ruff's `EXE001` fails the build on a file that carries a shebang and is committed
      `100644`, and a local ruff run does not reproduce it — the mode in the index is
      what CI reads. `validate-skill.sh` catches it with the exact fix, but only after
      a push, so the cheap moment is when the file is created:
      
      ```bash
      chmod +x path/to/new-script.py && git update-index --chmod=+x path/to/new-script.py
      ```
      
      Worth the one line: this cost two separate CI rounds in a single session, once on
      a maintainer's PR and once on a contributor's, both on freshly added test
      scripts. Drop the shebang instead where the file is genuinely only imported — a
      module is not a script, and making it executable settles the mismatch from the
      wrong side.
      
    • composer-setup.md 5.5 KB
      # Composer Setup for Skills
      
      ## Contents
      
      - When to Add composer.json
      - composer.json Structure
      - Multi-Skill Packages
      - Publishing to Packagist
      - Skill repos that ship their own PHP code
      - Files to NOT Include
      - Validation
      - Integration with Plugin
      - Troubleshooting
      
      Guide for adding Composer distribution to Netresearch skills.
      
      ## When to Add composer.json
      
      Add `composer.json` to ALL skills **EXCEPT** those explicitly targeting non-PHP ecosystems.
      
      | Skill Type | composer.json |
      |------------|---------------|
      | PHP/TYPO3 skills | Required |
      | General skills | Required |
      | Go-specific skills | Not needed |
      | Rust-specific skills | Not needed |
      
      ## composer.json Structure
      
      ### Basic Structure
      
      ```json
      {
        "name": "netresearch/{skill-name}-skill",
        "description": "{Skill description from SKILL.md}",
        "type": "ai-agent-skill",
        "license": "(MIT AND CC-BY-SA-4.0)",
        "authors": [
          {
            "name": "Netresearch DTT GmbH",
            "email": "plugins@netresearch.de",
            "homepage": "https://www.netresearch.de/",
            "role": "Manufacturer"
          }
        ],
        "require": {
          "netresearch/composer-agent-skill-plugin": "*"
        },
        "extra": {
          "ai-agent-skill": "SKILL.md"
        }
      }
      ```
      
      ### Key Fields
      
      | Field | Value | Purpose |
      |-------|-------|---------|
      | `name` | `netresearch/{repo-name}` | Must match GitHub repo name exactly |
      | `type` | `ai-agent-skill` | Enables plugin discovery |
      | `require` | `composer-agent-skill-plugin` | Plugin dependency |
      | `extra.ai-agent-skill` | Path to SKILL.md | Skill location |
      
      ### Package Naming Convention
      
      - **Rule: composer name = GitHub repo name** (`netresearch/{repo-name}`)
      - All skill repos are named `*-skill` on GitHub
      - Examples:
        - Repo `netresearch/typo3-docs-skill` → composer name `netresearch/typo3-docs-skill`
        - Repo `netresearch/jira-skill` → composer name `netresearch/jira-skill`
      
      ## Multi-Skill Packages
      
      For packages containing multiple skills:
      
      ```json
      {
        "name": "netresearch/{name}-skill",
        "type": "ai-agent-skill",
        "extra": {
          "ai-agent-skill": [
            "skills/skill-one/SKILL.md",
            "skills/skill-two/SKILL.md"
          ]
        }
      }
      ```
      
      ### Example: jira-skill
      
      ```json
      {
        "name": "netresearch/jira-skill",
        "extra": {
          "ai-agent-skill": [
            "skills/jira-communication/SKILL.md",
            "skills/jira-syntax/SKILL.md"
          ]
        }
      }
      ```
      
      ## Publishing to Packagist
      
      ### Prerequisites
      
      1. GitHub repository is public
      2. composer.json is valid
      3. Packagist account linked to GitHub
      
      ### Steps
      
      1. Create version tag:
         ```bash
         git tag v1.0.0
         git push --tags
         ```
      
      2. Register on Packagist:
         - Go to <https://packagist.org/packages/submit>
         - Enter repository URL
         - Submit
      
      3. Enable auto-update:
         - Configure GitHub webhook for Packagist
         - Or use GitHub Actions integration
      
      ## Skill repos that ship their own PHP code
      
      A skill repo with its own `src/`, `bin/` and dependencies runs `composer install` in its own root, locally and in CI. That activates `netresearch/composer-agent-skill-plugin` on the skill repo itself, and the plugin registers the repo as a skill inside itself:
      
      - `extra.ai-agent-skill` is rewritten from the string path into an object (`{"skills": [...], "allow-skills": []}`). `validate-skill.sh` then fails with `composer.json extra.ai-agent-skill could not be parsed`.
      - A `<skills_system>` block is appended to `AGENTS.md`.
      - `composer.json.skill-trust.lock` appears next to `composer.json`.
      
      Deny the plugin in the repo's own config. The package is the skill and has nothing to install into itself. Consumers are unaffected — `allow-plugins` applies to the root package, so an installing project decides for itself:
      
      ```json
      {
        "config": {
          "allow-plugins": {
            "netresearch/composer-agent-skill-plugin": false
          }
        }
      }
      ```
      
      The field is not optional once the repo has dependencies of its own: without an entry, a non-interactive `composer install` exits 1 with `blocked by your allow-plugins config` instead of prompting.
      
      Cover the artefacts in `.gitignore`:
      
      ```gitignore
      /vendor/
      /composer.lock
      composer.json.skill-trust.lock
      ```
      
      ## Files to NOT Include
      
      **Never add these to skill repos:**
      
      - `composer.lock` - Locks are for applications, not libraries
      - `vendor/` - Dependencies installed by users
      
      ## Validation
      
      Check composer.json validity:
      
      ```bash
      # Syntax check
      composer validate
      
      # Check type
      grep '"type"' composer.json | grep "ai-agent-skill"
      
      # Check plugin requirement
      grep "composer-agent-skill-plugin" composer.json
      ```
      
      ## Integration with Plugin
      
      The `composer-agent-skill-plugin` automatically:
      
      1. Discovers installed `ai-agent-skill` packages
      2. Parses SKILL.md frontmatter
      3. Generates AGENTS.md index
      4. Provides `composer list-skills` command
      5. Provides `composer read-skill {name}` command
      
      ## Troubleshooting
      
      ### "Package not found"
      
      - Ensure package is published on Packagist
      - Check package name matches exactly
      - Run `composer clear-cache`
      
      ### "Plugin not activated"
      
      When prompted, allow the plugin:
      ```
      Do you trust "netresearch/composer-agent-skill-plugin"? [y,n,d,?]
      ```
      
      Or pre-authorize in composer.json:
      ```json
      {
        "config": {
          "allow-plugins": {
            "netresearch/composer-agent-skill-plugin": true
          }
        }
      }
      ```
      
      This is the setting for a **consuming** project. A skill repo that ships its own PHP code sets it to `false` for itself — see [Skill repos that ship their own PHP code](#skill-repos-that-ship-their-own-php-code).
      
      ### "SKILL.md not found"
      
      - Verify path in `extra.ai-agent-skill`
      - Paths must be relative from package root
      - Absolute paths are rejected for security
      
    • installation-methods.md 10.7 KB
      # Installation Methods
      
      ## Contents
      
      - Method 1: Netresearch Marketplace (Recommended)
      - Method 2: Skills Directory (no marketplace)
      - Method 3: Download Release
      - Method 4: Composer (PHP Projects)
      - Method 5: npm (Node Projects)
      - Choosing a Method
      - Directory Locations
      
      Five methods for installing Netresearch skills.
      
      ## Method 1: Netresearch Marketplace (Recommended)
      
      The marketplace aggregates all Netresearch skills in one place.
      
      ### Setup
      
      ```bash
      /plugin marketplace add netresearch/claude-code-marketplace
      ```
      
      ### Usage
      
      ```bash
      # Browse available plugins
      /plugin
      
      # Install a specific plugin
      /plugin install {plugin-name}@netresearch-claude-code-marketplace
      ```
      
      `{plugin-name}` is the `name` field of the repo's entry in the catalog's
      `marketplace.json`, which is **not** always the repo name — `jira-skill` is
      `jira-integration`, `matrix-skill` is `matrix-communication`,
      `php-ast-edit-skill` is `php-structured-edit`. Read the name from the catalog,
      never derive it from the repo slug.
      
      ### Do not point `marketplace add` at a skill repo
      
      ```bash
      # WRONG — fails with: Marketplace file not found at .../marketplace.json
      /plugin marketplace add netresearch/{repo-name}
      ```
      
      `marketplace add` requires the target repo to contain a
      `.claude-plugin/marketplace.json` **catalog**. Skill repos ship a
      `.claude-plugin/plugin.json` **plugin manifest** instead, so this fails. The
      catalog is `netresearch/claude-code-marketplace`; its entries point back at
      each skill repo as their source, so installing from it still fetches the code
      from this repo.
      
      ### Benefits
      
      - Curated collection
      - Automatic updates — the only route with `claude plugin update` and a version in `/plugin`
      - Easy discovery
      - No manual file management
      
      ## Method 2: Skills Directory (no marketplace, Claude Code 2.1.157+)
      
      Since Claude Code 2.1.157: *"Plugins in `.claude/skills` directories are now
      automatically loaded, no marketplace required."* Any folder under a skills
      directory containing `.claude-plugin/plugin.json` loads as
      `{plugin-name}@skills-dir` on the next session — discovered in place, not
      copied into the plugin cache.
      
      ### Installation
      
      ```bash
      mkdir -p ~/.claude/skills
      git clone https://github.com/netresearch/{repo-name}.git \
        ~/.claude/skills/{plugin-name}
      ```
      
      ### What loads
      
      Personal scope (`~/.claude/skills/`) loads the whole plugin — skills, agents,
      hooks, commands, `bin/`, `.mcp.json`, `.lsp.json` — with no restrictions.
      
      Project scope (`<cwd>/.claude/skills/`, checked into a repo) loads only after
      the workspace trust dialog, and restricts what runs: MCP servers need
      per-server approval, LSP servers start only after trust, and background
      monitors do not load at all. Project-scope plugins are found only in the
      session's primary working directory — they do not walk up to the repo root, so
      launch from the repo root or move there with `/cd` (2.1.246+).
      
      ### Update and removal
      
      ```bash
      git -C ~/.claude/skills/{plugin-name} pull   # update
      rm -rf ~/.claude/skills/{plugin-name}        # remove
      claude plugin disable {plugin-name}@skills-dir  # keep on disk, stop loading
      ```
      
      There is no uninstall step, because nothing was installed from a marketplace.
      `SKILL.md` edits apply immediately; changes to `hooks/`, `.mcp.json`,
      `agents/` and output styles need `/reload-plugins` or a restart.
      
      ### Trade-off vs. the marketplace
      
      No `claude plugin update`, no version in `/plugin`, no discovery — updating is
      whatever `git pull` gives you. Use it for pinning to a branch or working from
      a local checkout; prefer the marketplace otherwise.
      
      ## Method 3: Download Release
      
      Download packaged skill files from GitHub Releases.
      
      ### Steps
      
      1. Go to skill's GitHub repository
      2. Navigate to Releases page
      3. Download latest `.zip` or `.tar.gz`
      4. Extract into `~/.claude/skills/` — **not** into a `{skill-name}/` subdirectory
         of it. The archive carries its own `{skill-name}/` folder, so extracting one
         level deeper produces `~/.claude/skills/{skill-name}/{skill-name}/SKILL.md`,
         which is not loaded.
      
      ```bash
      unzip -d ~/.claude/skills {skill-name}-skill-vX.Y.Z.zip
      unzip -Z1 {skill-name}-skill-vX.Y.Z.zip | cut -d/ -f1 | sort -u   # one entry: {skill-name}
      ```
      
      ### Package Contents
      
      Release packages contain only skill-relevant files:
      - `SKILL.md`
      - `LICENSE-MIT`
      - `LICENSE-CC-BY-SA-4.0`
      - `references/`
      - `scripts/`
      - `assets/`
      - `templates/`
      
      ### Excluded from Packages
      
      - `README.md` (human documentation)
      - `.github/` (CI/CD)
      - `composer.json` (separate distribution)
      - Dev configuration files
      
      ## Method 4: Composer (PHP Projects)
      
      For PHP projects, install skills as Composer packages.
      
      ### Prerequisites
      
      1. PHP 8.2+
      2. Composer 2.1+
      3. [composer-agent-skill-plugin](https://github.com/netresearch/composer-agent-skill-plugin)
      
      ### Installation
      
      ```bash
      # Install the plugin first (once per project)
      composer require netresearch/composer-agent-skill-plugin
      
      # Install skills
      composer require netresearch/{repo-name}
      ```
      
      ### How It Works
      
      1. Plugin discovers packages with type `ai-agent-skill`
      2. Generates `AGENTS.md` index in project root
      3. Skills available via `composer read-skill {name}`
      
      ### Benefits
      
      - Version management via Composer
      - Dependency resolution
      - Project-specific skill sets
      - Easy updates with `composer update`
      
      ## Method 5: npm (Node Projects)
      
      For Node.js / TypeScript projects, install skills as npm packages discovered by `@netresearch/agent-skill-coordinator`.
      
      ### Prerequisites
      
      1. Node.js 18+ and npm 9+ (or pnpm/yarn equivalent)
      2. [`@netresearch/agent-skill-coordinator`](https://github.com/netresearch/node-agent-skill-coordinator) — peer dependency that scans `node_modules` and registers skills in `AGENTS.md`
      
      ### Installation
      
      ```bash
      npm install --save-dev \
        @netresearch/agent-skill-coordinator \
        github:netresearch/{repo-name}
      ```
      
      For pnpm, allowlist the coordinator's `postinstall` so it can write `AGENTS.md`:
      
      ```json
      {
        "pnpm": {
          "onlyBuiltDependencies": ["@netresearch/agent-skill-coordinator"]
        }
      }
      ```
      
      ### How It Works
      
      1. Coordinator's `postinstall` walks `node_modules` for packages declaring `aiAgentSkill: skills/<name>/SKILL.md`
      2. Validates frontmatter, then writes a `<skills_system>` block into the project's `AGENTS.md`
      3. Skill content is then visible to any agent reading `AGENTS.md`
      
      ### What ships in the npm tarball
      
      The `files` allowlist in `package.json` controls what npm packs. The default in `templates/package.json.template` is intentionally **minimal** — the skill payload (`skills/<name>/`), plugin metadata (`.claude-plugin/`), the canonical agent rules entry-point (`AGENTS.md`), licenses, and `README.md`:
      
      ```json
      {
        "files": [
          "skills/{skill-name}/",
          ".claude-plugin/",
          "AGENTS.md",
          "LICENSE-MIT",
          "LICENSE-CC-BY-SA-4.0",
          "README.md"
        ]
      }
      ```
      
      #### When to add a top-level data dir
      
      Add a top-level dir to `files` **only if your installed skill code reads from it at runtime** (e.g. via `$ROOT/<dir>/...` or `../<dir>/...` from a script under `skills/<name>/scripts/`). Common runtime data dirs:
      
      - `catalog/` — ship it when installer scripts read `$ROOT/catalog/*.json`.
      - `hooks/` — Claude Code's plugin loader reads `hooks/hooks.json`. Ship it if your skill ships PreToolUse/PostToolUse hooks.
      - `commands/` — slash command definitions. Ship if present.
      - `outputStyles/` — output style definitions. Ship if present.
      - `assets/` — referenced assets (images, configs). Ship if your skill content references them at install paths.
      
      #### When NOT to add a top-level dir
      
      - **Top-level `scripts/`** is typically **repo-maintenance only** (e.g. `verify-harness.sh`, `generate-dashboard.sh`). Keep it out unless your installed skill scripts read from `$ROOT/scripts/` at runtime. Runtime scripts belong under `skills/<name>/scripts/` (already covered by `skills/<name>/`).
        - Example: `context7-skill` does NOT ship top-level `scripts/` because its only file is `verify-harness.sh` (repo-maintenance).
      - `Build/` — dev-only build artifacts. Never ship.
      - `evals/`, `docs/` — repo-internal. Never ship.
      - `.github/`, lint configs (`.markdownlint*`, `.yamllint*`), `.envrc` — repo-internal. Never ship.
      
      Heuristic when looking at a top-level dir:
      
      ```text
      package-root/
        catalog/   # consumed at runtime  -> MUST be in files
        Build/     # dev-only build artifacts -> DO NOT include
        scripts/   # inspect: runtime or repo-maintenance? include only if runtime
      ```
      
      #### CI safeguard
      
      `templates/.github/workflows/npm-pack-smoke.yml.template` is a ready-to-copy GitHub Actions workflow that runs `npm pack --dry-run` on every PR and asserts:
      
      1. **No internal leakage** — fails if the tarball contains `.github/`, `evals/`, `docs/`, `Build/`, `verify-harness.sh`, lint configs, etc.
      2. **Runtime-referenced dirs are present** — greps `skills/*/scripts/` for `$ROOT/<dir>` and `../<dir>/` references; if a script reads `$ROOT/catalog` but `catalog/` is missing from the tarball, the job fails.
      
      Copy it into `.github/workflows/npm-pack-smoke.yml` in your skill repo. It catches both kinds of `files`-allowlist mistakes (over-inclusion and under-inclusion) before they ship.
      
      ### `"private": true` on `0.0.0-source`
      
      Skill repos use the placeholder version `0.0.0-source` and **must** set `"private": true` to guard against accidental `npm publish` of the placeholder. Real publishes (when/if added later) flip `private` to `false` (or remove it) at release time and set a real semver.
      
      ### Limitation: SKILL.md content only
      
      The npm path registers only `SKILL.md` content into `AGENTS.md`. **Slash commands** (defined under `commands/`) and **PreToolUse / PostToolUse hooks** (defined under `.claude-plugin/`) are loaded by Claude Code's plugin mechanism — not by the coordinator scanning `node_modules`. Repos that ship those features should document this explicitly in their README (a `> **Limitation:**` callout immediately after the npm install snippet) so consumers know they need the marketplace install for the full skill.
      
      ### Benefits
      
      - Lock-file pinning via `package-lock.json` / `pnpm-lock.yaml`
      - Renovate / Dependabot can bump skill versions like any other dep
      - No PHP / Composer required for Node-only projects
      
      ## Choosing a Method
      
      | Scenario | Recommended Method |
      |----------|-------------------|
      | General Claude Code use | Marketplace |
      | Offline/air-gapped | Release download |
      | Repo not in the catalog | Skills directory |
      | Pinning to a branch, or hacking on the skill | Skills directory |
      | PHP project | Composer |
      | Node / TypeScript project | npm |
      | CI/CD automation | Composer or npm |
      | Quick trial | Marketplace |
      
      ## Directory Locations
      
      | Method | Location |
      |--------|----------|
      | Marketplace | Managed by Claude Code |
      | Skills directory | `~/.claude/skills/{plugin-name}/` (loaded in place) |
      | Release | extract into `~/.claude/skills/`; the archive supplies `{skill-name}/` |
      | Composer | `vendor/netresearch/{repo-name}/` |
      | npm | `node_modules/@netresearch/{repo-name}/` |
      
    • marketplace-integration.md 5.8 KB
      # Marketplace Integration
      
      How skills are synced to the Netresearch Claude Code Marketplace.
      
      ## Separation from skill-repo rules
      
      - **This document** explains sync mechanics and source-vs-synced files.
      - **Discovery, SEO, catalog completeness, and orphan rules for the marketplace Hub** live only in the marketplace repository: **[`AGENTS.md` in `netresearch/claude-code-marketplace`](https://github.com/netresearch/claude-code-marketplace/blob/main/AGENTS.md)**.
      - **Single skill repository quality** (README sections, `agents/openai.yaml`, GitHub Description/Topics, discovery YAML) lives here in **skill-repo-skill** — see [`repository-quality-rules.md`](repository-quality-rules.md), [`readme-template.md`](readme-template.md), [`skill-discovery-metadata.md`](skill-discovery-metadata.md).
      
      Do **not** duplicate marketplace governance prose in skill-repo references; link instead.
      
      ## Overview
      
      The [claude-code-marketplace](https://github.com/netresearch/claude-code-marketplace) aggregates skills from individual repositories via automated sync workflows.
      
      ## Architecture
      
      ```
      ┌─────────────────┐     ┌─────────────────┐     ┌─────────────────┐
      │  skill-repo-1   │     │  skill-repo-2   │     │  skill-repo-N   │
      │  (source)       │     │  (source)       │     │  (source)       │
      └────────┬────────┘     └────────┬────────┘     └────────┬────────┘
               │                       │                       │
               └───────────────────────┼───────────────────────┘
                                       │
                                       ▼
                          ┌────────────────────────┐
                          │  claude-code-marketplace│
                          │  (aggregator)           │
                          │                         │
                          │  skills/                │
                          │  ├── skill-1/           │
                          │  ├── skill-2/           │
                          │  └── skill-N/           │
                          └────────────────────────┘
      ```
      
      ## Sync Workflow
      
      ### Automatic Sync (Scheduled)
      
      - **Frequency:** Weekly (Monday 2 AM UTC)
      - **Trigger:** GitHub Actions schedule
      - **Process:**
        1. Clone each source repository
        2. Extract semantic version from `.claude-plugin/plugin.json`
        3. Append commit date for versioning
        4. Copy skill files to `skills/` directory
        5. Update marketplace metadata
      
      ### On-Demand Sync
      
      Source repositories can trigger immediate sync via:
      
      ```yaml
      # notify-marketplace.yml in source repo
      name: Notify Marketplace
      
      on:
        release:
          types: [published]
      
      jobs:
        notify:
          runs-on: ubuntu-latest
          steps:
            - name: Trigger marketplace sync
              run: |
                gh workflow run sync-skills.yml \
                  --repo netresearch/claude-code-marketplace
              env:
                GH_TOKEN: ${{ secrets.MARKETPLACE_SYNC_TOKEN }}
      ```
      
      ## Adding a Skill to Marketplace
      
      ### Prerequisites
      
      1. Skill repository follows standard structure
      2. `.claude-plugin/plugin.json` has valid version
      3. Repository is public
      
      ### Configuration
      
      Add skill to `.sync-config.json` in marketplace repo:
      
      ```json
      {
        "skills": [
          {
            "name": "my-skill",
            "repo": "netresearch/my-skill",
            "path": "skills/my-skill"
          }
        ]
      }
      ```
      
      ### Versioning
      
      Marketplace versions combine:
      - Semantic version from `.claude-plugin/plugin.json`
      - Last commit date
      
      Example: `1.2.3-20251021`
      
      ## Sync Workflow Details
      
      ### Files Synced
      
      From source repository:
      - `SKILL.md`
      - `references/`
      - `scripts/`
      - `assets/`
      - `templates/`
      - `LICENSE-MIT`
      - `LICENSE-CC-BY-SA-4.0`
      
      ### Files NOT Synced
      
      - `README.md` (marketplace has its own)
      - `.github/`
      - `composer.json`
      - Dev configuration files
      - `claudedocs/`
      
      ### Marketplace Metadata
      
      Updated in `.claude-plugin/marketplace.json`:
      
      ```json
      {
        "skills": {
          "my-skill": {
            "version": "1.2.3-20251021",
            "description": "Skill description",
            "source": "netresearch/my-skill"
          }
        }
      }
      ```
      
      ## User Installation
      
      ### Add Marketplace
      
      ```bash
      /plugin marketplace add netresearch/claude-code-marketplace
      ```
      
      ### Browse Skills
      
      ```bash
      /plugin
      ```
      
      ### Install Skill
      
      ```bash
      /plugin install {plugin-name}@netresearch-claude-code-marketplace
      ```
      
      `{plugin-name}` is the entry's `name` in the catalog, not necessarily the repo
      name. `marketplace add` must point at the catalog repo — pointing it at an
      individual skill repo fails, because a skill repo ships
      `.claude-plugin/plugin.json`, not `.claude-plugin/marketplace.json`. See
      [`installation-methods.md`](installation-methods.md) for the marketplace-free
      skills-directory route.
      
      ## Benefits of Marketplace Distribution
      
      | Benefit | Description |
      |---------|-------------|
      | Discoverability | All skills in one place |
      | Curation | Quality-controlled collection |
      | Versioning | Automatic version tracking |
      | Updates | Sync keeps skills current |
      | Simplicity | One-command installation |
      
      ## Troubleshooting
      
      ### Skill Not Appearing
      
      1. Check `.sync-config.json` includes the skill
      2. Verify source repo is accessible
      3. Check sync workflow logs
      4. Ensure SKILL.md has valid frontmatter
      
      ### Outdated Version
      
      1. Check last sync time in marketplace README
      2. Trigger manual sync if needed
      3. Verify source repo has new commits/releases
      
      ### Sync Failures
      
      Common causes:
      - Invalid SKILL.md frontmatter
      - Missing required files
      - Network/API issues
      - Permission problems
      
      Check GitHub Actions logs in marketplace repo.
      
    • materialization-contract.md 18.6 KB
      # Materialization Contract
      
      ## Contents
      
      - Scope
      - Failure-pattern schema
      - Rule 1: Patches target source repos, never local cache
      - Rule 2: Workspace preference order
      - Rule 3: Branch convention
      - Rule 4: Commit conventions
      - Rule 5: PR template
      - Rule 6: Target area mapping
      - Rule 7: Eval format (skill-repo convention)
        - What these evals cannot measure, and why it matters
      - Rule 8: Per-private-repo confirmation
      - Rule 9: New-skill scaffolding
      - See also
      
      How external tools (notably `retro-skill`) materialize **skill improvements** or **new skills** by submitting PRs to skill repositories that follow `skill-repo-skill` conventions.
      
      ## Scope
      
      This contract covers two destinations from [retro-skill](https://github.com/netresearch/retro-skill)'s destination taxonomy:
      
      - **`skill-update`** — PR against an existing skill repo
      - **`new-skill`** — Scaffolding of a new skill repo
      
      For the full destination taxonomy and routing heuristics, see [retro-skill/references/destination-taxonomy.md](https://github.com/netresearch/retro-skill/blob/main/references/destination-taxonomy.md).
      
      ## Failure-pattern schema
      
      Every failure pattern encoded into a skill — a `skill-update`, a `new-skill`, a PR's "Came from" section, a coach `rule.md`/`antipattern.md` entry — is captured in **exactly** four fields, in this order:
      
      | Field | Content |
      |---|---|
      | **Symptom** | What was observed going wrong |
      | **Cause** | Why it happened (root cause, not the trigger) |
      | **Required behavior** | What the agent must do instead |
      | **Verification** | An eval id (`evals/evals.json` `name`) or checkpoint id (`checkpoints.yaml` key) that proves the required behavior — not free text |
      
      Verification is mandatory. It makes explicit the pipeline's existing TDD-eval-stub rule: a failure pattern without a named eval or checkpoint has no machine check.
      
      This is the single source of truth for the schema's four fields and their order. Consuming surfaces (`retro-skill`'s PR "Came from" section, `claude-coach-plugin`'s `rule.md`/`antipattern.md` templates) require the schema; they do not redefine it.
      
      ## Rule 1: Patches target source repos, never local cache
      
      Local skill cache (`~/.claude/plugins/cache/`) is overwritten on plugin update. Edits there are lost. **Always** locate the source repository before patching.
      
      Resolution order:
      
      1. `<skill-root>/.claude-plugin/plugin.json` → `repository` field
      2. `<skill-root>/composer.json` → `support.source` or `homepage`
      3. `<skill-root>/.git/config` → `remote.origin.url`
      4. Walk up parent directories from the resolved-symlink path looking for any of the above
      5. Last resort: ask the user
      
      If unresolvable: do NOT patch local cache. Either ask user or refuse.
      
      ## Rule 2: Workspace preference order
      
      For each skill-update target, select working directory in this order:
      
      1. **Existing worktree:** `~/p/<repo-name>/main/` exists AND is a clean git worktree → use it. `<repo-name>` is the full GitHub repo name (e.g. `skill-repo-skill`, not `skill-repo`).
      2. **Existing flat checkout:** `~/p/<repo-name>/` exists AND is a clean flat git checkout on `main` → use it.
      3. **Fresh clone:** Otherwise clone into `/tmp/retro-workspace/<repo-name>/`.
      
      Dirty checkouts: do NOT use. Fall back to /tmp. Always tell the user why.
      
      ## Rule 3: Branch convention
      
      ```
      feat/retro-<short-slug>
      ```
      
      `<short-slug>` is kebab-case, derived from the friction title. **Maximum 60 characters** for the slug part (full branch ≤ 71 chars including `feat/retro-` prefix; comfortable on most terminals).
      
      Examples:
      
      - `feat/retro-add-bun-trigger-keywords`
      - `feat/retro-fix-yaml-tool-choice-for-data-tools`
      
      For `new-skill`: the new repo's first branch is `main` (initial scaffolding commit).
      
      ## Rule 4: Commit conventions
      
      - **Format:** Conventional Commits (`<type>(<scope>): <summary>`)
        - Types: `feat`, `fix`, `docs`, `refactor`, `chore`
      - **No bot attribution.** Never append "Generated with Claude Code" or "Co-Authored-By: Claude"
      - **Preserve signing.** Never pass `--no-gpg-sign` or `-c commit.gpgsign=false`
      - **Preserve hooks.** Never pass `--no-verify`
      - **DCO sign-off.** All commits include `Signed-off-by: <name> <email>` trailer (use `git commit -s` or `git rebase --signoff`)
      - **Atomic.** One logical change per commit
      
      Commit message body should reference the friction:
      
      ```
      feat(triggers): include bun in skill description keywords
      
      Found via /retro session 2026-05-11: the assistant suggested npm
      in 4 turns despite the project clearly using bun. SKILL.md description
      didn't include 'bun' as a trigger.
      
      Signed-off-by: Sebastian Mendel <github@sebastianmendel.de>
      ```
      
      ## Rule 5: PR template
      
      PRs created by retro-skill should use the named template `retro.md`:
      
      ```bash
      gh pr create --template retro.md ...
      # or via URL query when opening manually:
      # https://github.com/<org>/<repo>/compare/main...<branch>?template=retro.md
      ```
      
      This invokes `.github/PULL_REQUEST_TEMPLATE/retro.md` (not the repo's default template, if any). The retro template has:
      
      - `## Summary`
      - `## Came from` (session date, finding signal ID, and the finding stated in the [failure-pattern schema](#failure-pattern-schema): Symptom → Cause → Required behavior → Verification)
      - `## Change` (concrete diff scope)
      - `## Target area` (`skills/<name>/SKILL.md` / `references/` / `scripts/` / `templates/` / `checkpoints.yaml` / `evals/evals.json`)
      - `## Learning source` (checkboxes: from /retro, reusable, scoped, eval included)
      - `## Test plan` (verification steps)
      
      ## Rule 6: Target area mapping
      
      Paths are relative to the **skill's own subdirectory** in the repo (this repo's convention: `skills/<skill-name>/...`, not repo-root).
      
      Each `skill-update` PR should touch **one primary area**; multi-area is allowed when cohesive (e.g. SKILL.md update + corresponding eval).
      
      | Target | When |
      |---|---|
      | `skills/<name>/SKILL.md` | Trigger description, workflow guidance, key principles |
      | `skills/<name>/references/*.md` | Detailed knowledge, examples, schemas |
      | `skills/<name>/scripts/*` | Mechanical operations |
      | `skills/<name>/templates/*` | Output formats |
      | `skills/<name>/checkpoints.yaml` | Quality gates — mechanical by default, LLM only for what a command cannot decide |
      | `skills/<name>/evals/evals.json` | Behavioral regression tests |
      | `tests/*` | Regression coverage for `scripts/*`, run by the `tests.yml` reusable |
      
      Multi-area PRs split unrelated concerns into separate PRs.
      
      ## Rule 7: Eval format (skill-repo convention)
      
      This repo's eval format is **a single `evals/evals.json` file** containing an array of objects. Each object has at minimum:
      
      ```json
      {
        "name": "<scenario-name>",
        "prompt": "<what to ask the agent>",
        "assertions": [
          { "type": "content", "pattern": "<regex>" },
          { "type": "tool_use", "tool": "<ToolName>", "pattern": "<regex>" }
        ]
      }
      ```
      
      Validation: `bash skills/skill-repo/scripts/validate-evals.sh`.
      
      When a `skill-update` PR changes behavior expectations, **append a new eval object to the existing `evals/evals.json` array** in the target skill's directory. Do NOT emit `evals/*.md` files — they will be rejected by the validator.
      
      Other skill repos may use different eval conventions; consult each target's `evals/` directory before submitting.
      
      ### Quality rule: every eval needs a delta-discriminating assertion
      
      Every eval needs at least one assertion a no-skill baseline answer FAILS; verify with
      `repo-root run-ab-evals.sh` (the `without` pass rate on that assertion must be <1.0). A
      keyword assertion that any generic answer would already contain (`"alpine"`, `"USER"`,
      a bare `"Plugin"`) rewards boilerplate and never tests the retro-born gotcha that is
      the skill's actual value — write the assertion against the specific fact, number, flag,
      or command shape a model can only produce by having read the skill (an internal script
      path, an exact risk multiplier, a non-obvious CLI invocation, a "do X not Y" correction
      of the model's default instinct).
      
      The runner computes this for you. It calls an assertion **discriminating** when the
      baseline fails it and the with-skill run passes it, prints
      `discriminating=<n>/<total>` on every eval line, and lists each eval that scored zero
      under a "no discriminating check" heading at the end. `ab-results.json` carries the same
      information as `discriminating_checks` and `evals_without_delta`. `--require-delta`
      turns that list into exit 1 — use it in a repo whose evals are already clean, so the
      next eval cannot arrive without evidence. Note the cost: `2 × N` completions per eval,
      so a sampled run over a 20-eval suite is 120 model calls.
      
      Zero discriminating checks does not make an eval wrong: a regression guard both arms
      pass today is doing its job. It makes it unusable as evidence *for the skill*, which is
      a different claim and the one being made whenever skill content is defended.
      
      This verification is a **LOCAL/manual step, not CI** — `run-ab-evals.sh` calls out to
      `claude -p` per eval. It is referenced only by `.github/workflows/ab-evals-schedule.yml`,
      which is disabled by team decision, so it is not part of any active CI gate.
      
      **Report the delta with the actor it was measured under**, never as a bare percentage —
      see [`skill-quality.md`](skill-quality.md) ("A delta belongs to an actor, not to a
      skill"). The runner prints the tuple and stores it in the `provenance` block.
      
      ### `samples`: the part of that quality rule CI can decide
      
      An assertion is only a non-empty string as far as structural validation used to be
      concerned, so an inverted one — demanding a word the intended answer never uses — read
      exactly like a discriminating one. Add an optional `samples` block and the validator
      decides it mechanically, with no model call:
      
      ```json
      {
        "name": "worktree_cleanup_spares_the_primary_directory",
        "prompt": "…",
        "assertions": [{ "type": "content", "pattern": "(?i)\\bswitch\\b[^\\n]{0,30}\\b(main|master)" }],
        "samples": {
          "passing": "Do not remove it — run git -C project/main switch main instead.",
          "failing": ["Remove every merged worktree, including project/main."]
        }
      }
      ```
      
      Every assertion must match `passing`, and each entry in `failing` must be rejected by at
      least one assertion. Write the `failing` samples from the answer a model actually gives
      without the skill — a straw man that shares no vocabulary with a correct answer passes
      the check while proving nothing.
      
      **Matching uses the grader's instrument, not a similar one.** `run-ab-evals.sh` grades
      with `grep -qiE` after stripping a leading `(?i)`: case-insensitive POSIX ERE, over
      `value or pattern` from *every* assertion whatever its type. The validator calls the
      same binary with the same flags, because Python's `re` disagrees with ERE in ways that
      decide real cases (`[^\n]` excludes the letter `n` in ERE, and `re` is case-sensitive),
      and a self-check measuring with a different engine certifies evals the grader would fail.
      For the same reason a pattern under `tool_use` or `content_regex` is validated too.
      
      **A `must_not` assertion is inverted in both instruments.** The validator requires that it
      does NOT match `samples.passing`; the grader counts the arm in which the pattern is
      **absent** as the passing one, and a discriminating check accordingly means the baseline
      produced the forbidden answer and the skill run did not. Until August 2026 the grader
      ignored the type and credited whichever arm *contained* the pattern, so a negative
      expectation could not be stated as an assertion at all and had to live in
      `samples.failing`. Both directions are now pinned in `tests/run-ab-evals.sh`.
      
      **On a new or tightened eval, `samples` are required, not optional** (decided in
      [retro-skill#92](https://github.com/netresearch/retro-skill/issues/92)). `eval-validate.yml`
      writes the base branch's copy of the `evals.json` out on a pull request and hands it to the
      validator as `EVALS_BASE_FILE`; every eval that is new there, or whose `assertions` value
      differs from the base copy, must carry `samples.passing`. An eval nobody touched is never
      looked at, so the fleet's existing evals are not retrofitted, and an eval carrying no
      pattern this validator can run against a sample — `expectations` only, plain-string
      assertions, or a single unparseable `*_contains` literal — is exempt, because samples no
      assertion backs are failed here. Without `EVALS_BASE_FILE` (a push build, a local run, a
      consumer on an older workflow) the requirement is off and the verdict is unchanged.
      
      Two things change for evals **without** a samples block: patterns are now checked to be
      parseable by `grep -E`, and an unparseable one fails — except under a `*_contains` type,
      where the value is usually a literal and the mismatch is the grader's, so it only warns.
      Nothing else about validation changed, and no evals.json in the fleet changes verdict.
      
      ### What these evals cannot measure, and why it matters
      
      `run-ab-evals.sh` builds the "with" arm by pasting the whole `SKILL.md` into
      `--append-system-prompt`, and runs both arms with `--tools ""`. The skill is
      force-fed. **There is no routing in this harness at all**, so no eval in any
      `evals.json` can detect the one failure a description can have: never being
      picked up.
      
      That is not hypothetical. Measured in
      [netresearch/agent-system-evals](https://github.com/netresearch/agent-system-evals),
      `typo3-conformance` — the skill for what a TYPO3 extension declares about
      supported versions — sat in the fleet across three cases about exactly that and
      was loaded in **1 of 25 trials**, because its description never named those
      artefacts. Its 30 evals all passed throughout: 29 of their prompts are phrased
      in the skill's own vocabulary (*"Check this TYPO3 extension for conformance to
      current standards"*), so they measure what it does once loaded and cannot see
      that nothing reaches it. Across the 58 installed Netresearch skills there are
      **1335 evals and no trigger evals**.
      
      Two consequences for anyone reading an A/B result:
      
      - **A green A/B says the content helps when supplied.** It says nothing about
        whether the skill is ever consulted, and the two are independent — a skill can
        win every eval and be reached by nobody.
      - **A `tool_use` assertion does not observe a tool call.** The grader matches
        `value or pattern` against the answer text with `grep -qiE`, whatever the
        type, and the arm runs with no tools. It is a text assertion wearing another
        name.
      
      **Where trigger testing belongs.** A trigger eval needs the skill *installed as
      a skill* rather than pasted, tools enabled, and the transcript inspected for an
      invocation. `agent-system-evals` already does this — its `skill_invoked`
      endpoint counts `Skill(...)` calls in the recorded trajectory, and
      `scripts/invocation-census` reports the rate per case. Write the queries in the
      format [agentskills.io documents](https://agentskills.io/skill-creation/optimizing-descriptions)
      — realistic prompts labelled `should_trigger` — keep them in
      `evals/eval_queries.json` beside `evals.json`, and take the phrasings from
      requests real users actually made rather than from the description they are
      meant to test.
      
      **Never copy them from a benchmark that measures the skill.** This advice, read
      literally, produced exactly that: the first trigger file written under it took
      its queries verbatim from `agent-system-evals`, named that repository's case
      identifiers, and recorded the routing rate each case had measured — including
      the rate for the case that was measuring the skill at the time. The benchmark's
      contamination check raised nineteen hits the first moment a fleet resolved the
      skill, and the arm had to be rebuilt and the comparison re-run.
      
      The line is between subject and sentences. A trigger eval for a conformance
      skill is about version declarations, and so is the case that measures it;
      sharing that subject is unavoidable and carries nothing. Sharing the case's
      sentences hands the agent the prompt it will be given, and a case identifier
      hands it the answer key. So write each query from the artefact it is about, in
      a maintainer's words, then check mechanically: no case identifier, no measured
      figure, and no run of words shared with any instruction in the benchmark beyond
      an ordinary noun phrase. `evals.json` is not exempt — its prompts reach the
      same fleet.
      
      ## Rule 8: Per-private-repo confirmation
      
      Before pushing to a private host (`git.netresearch.de`, `gitlab.com/<private-org>`, etc.), prompt the user. Decision is remembered per `(session, repo-url)` for the duration of the active retro-skill session; not persisted across sessions.
      
      ## Rule 9: New-skill scaffolding
      
      For `new-skill` destination, use the templates in `skills/skill-repo/templates/`:
      
      - `composer.json.template` (with `type: ai-agent-skill`, split licensing in SPDX)
      - `package.json.template` (for npm-distributable variants)
      - `LICENSE-MIT.template`, `LICENSE-CC-BY-SA-4.0.template`
      - `README.md.template`
      - `release.yml.template` (GitHub Actions release workflow)
      - `validate.yml.template` (CI caller for the reusable skill-validation workflow — without it the SKILL.md word cap, plugin.json schema, and markdown/yaml/action lints run only in local pre-commit, never in CI)
      - `pr-quality.yml.template` (PR validation)
      - `auto-merge-deps.yml.template` (Dependabot/Renovate auto-merge)
      - `pre-commit.template`
      
      Required files in the new repo:
      
      - `.claude-plugin/plugin.json` (Claude marketplace manifest)
      - `composer.json` (from template)
      - `skills/<name>/SKILL.md` (skill definition)
      - `LICENSE-MIT`, `LICENSE-CC-BY-SA-4.0` (from templates)
      - `README.md` (from template)
      - `AGENTS.md` (agent-harness convention)
      - `.gitignore`
      - One initial `skills/<name>/references/*.md` covering the friction pattern
      - One initial entry in `skills/<name>/evals/evals.json` covering the friction (TDD)
      - Optionally `skills/<name>/checkpoints.yaml` (start with structural checkpoints)
      - `.github/workflows/release.yml` (from template)
      - `.github/workflows/validate.yml` (from template — required so skill validation runs in CI, not just pre-commit)
      
      Run `bash skills/skill-repo/scripts/validate-skill.sh <new-repo-path>` to confirm the scaffold passes structural validation.
      
      Marketplace listing is a separate manual step (out of scope for this contract).
      
      ## See also
      
      - [retro-skill destination-taxonomy](https://github.com/netresearch/retro-skill/blob/main/references/destination-taxonomy.md) — Six destinations
      - [retro-skill patch-workflow](https://github.com/netresearch/retro-skill/blob/main/references/patch-workflow.md) — Workflow on the retro-skill side
      - [retro-skill eval-integration](https://github.com/netresearch/retro-skill/blob/main/references/eval-integration.md) — How retro reads evals when proposing skill-update
      - `skills/skill-repo/templates/` — Scaffolding templates
      - `skills/skill-repo/scripts/validate-evals.sh` — Eval validator
      - `skills/skill-repo/scripts/validate-skill.sh` — Structure validator
      - User memory: `feedback_preserve-commit-signing`, `feedback_merge-strategy`, `feedback_no-version-bumps-in-feature-prs`
      
    • plugin-hooks.md 2.6 KB
      # Plugin Hooks
      
      A skill repository can ship Claude Code hooks in `hooks/hooks.json`. A hook
      that registers, runs and exits 0 can still do nothing: `cli-tools-skill`
      shipped one for months that could never have fired, and no check noticed.
      Measured against the [hooks reference](https://code.claude.com/docs/en/hooks)
      and a live session (2026-09).
      
      ## Contents
      
      - Pick the event that fires
      - Read the input the event actually carries
      - Plain stdout does not reach the model
      - Verify live, not only in unit tests
      
      ## Pick the event that fires
      
      `PostToolUse` fires only after a tool call **succeeds**. A Bash command that
      exits non-zero — `command not found` is exit 127 — fires
      **`PostToolUseFailure`** instead. A hook meant to react to failures must
      register for `PostToolUseFailure`; add `PostToolUse` only for the case where a
      failing command sits inside a pipeline or list whose overall status is 0.
      
      ## Read the input the event actually carries
      
      | Event | Where the text is |
      |---|---|
      | `PostToolUseFailure` | top-level `error`: first line `Exit code N`, then stdout and stderr interleaved |
      | `PostToolUse` (Bash) | `tool_response.stdout` / `tool_response.stderr` |
      
      There is no top-level `output` or `stdout`. Match only the lines the failing
      program prints (e.g. `bash: line 1: rg: command not found`), not a phrase
      anywhere in the text, or ordinary output that quotes it triggers the hook.
      
      ## Plain stdout does not reach the model
      
      Plain stdout from these events goes nowhere the model sees. To add context
      without signalling an error, emit JSON:
      
      ```json
      {"hookSpecificOutput": {"hookEventName": "PostToolUseFailure", "additionalContext": "…"}}
      ```
      
      The other path is exit status 2: its stderr is shown to Claude as feedback on
      the failed or completed call. Use it for a genuine objection, not for advice
      on a call that behaved normally.
      
      Keep the text to what the hook knows: a shell's `command not found` means the
      name did not resolve on PATH, not that the tool is absent. Fail open — any
      exception exits 0 with no output — so a hook can never break the call it
      observes.
      
      ## Verify live, not only in unit tests
      
      Unit tests prove the script; they cannot prove Claude Code calls it, with
      this input, and shows the result. Load the checkout as a plugin and make the
      model trigger the hook:
      
      ```bash
      claude -p --plugin-dir <checkout> --allowedTools Bash \
        "Run: bash -c 'PATH=/nonexistent; rg --version'. Quote verbatim any hook context you received, or say NONE."
      ```
      
      `NONE` means the hook is dead, whatever its tests say. Repeat once after
      installing from the marketplace, since `${CLAUDE_PLUGIN_ROOT}` then points
      into the plugin cache.
      
    • readme-template.md 6 KB
      # README structure for Netresearch skill repositories
      
      Human-facing documentation for a skill repo. **Marketplace** pages may summarize this content but must not become the only place where these facts exist.
      
      **Boilerplate generator:** [`../templates/README.md.template`](../templates/README.md.template) — align templates with the headings below when updating scaffolding.
      
      ---
      
      ## First screen (above-the-fold)
      
      Without long scrolling, the reader must see answers to:
      
      1. **What problem does this skill solve?**
      2. **When should it be used?**
      3. **What does it produce?**
      4. **What inputs or project context does it need?**
      5. **How do I install it or find it in the marketplace?**
      
      **Version badge (optional):** a README that shows its version under the title uses the live release badge, linked to the releases page:
      
      ```markdown
      [![Release](https://img.shields.io/github/v/release/netresearch/{repo-name}?sort=semver)](https://github.com/netresearch/{repo-name}/releases)
      ```
      
      A static `img.shields.io/badge/version-X.Y.Z-…` URL is updated by no release step (`bump-version.sh` and `check-version-parity.sh` do not read `README.md`), so it drifts from the latest tag; `validate-skill.sh` warns on it.
      
      ---
      
      ## Required Markdown sections (English)
      
      Use **exact** level-2 headings so agents can grep them:
      
      - `## What this skill solves`
      - `## Why this is a skill (model delta)` — the claimed value category/categories from the content value rubric ([`skill-quality.md`](skill-quality.md)), **one sentence** on what the model does wrong without the skill, and a pointer to eval evidence (`evals.json` id or dashboard delta) where it exists. Distinct from `## What this skill solves`: that section names the problem; this one justifies why a skill (not the base model) solves it.
      - `## Use when`
      - `## Expected outputs`
      - `## Context requirements`
      - `## Example prompts` — **minimum three** fenced or bulleted realistic prompts (distinct scenarios).
      - `## Related skills` — slugs/URLs **or** explicit `none (justified: …)`.
      - `## Installation` — must include the Netresearch marketplace path, **both** lines (`/plugin marketplace add netresearch/claude-code-marketplace` followed by `/plugin install {plugin-name}@netresearch-claude-code-marketplace`), **or** a pointer to the org-standard install doc. The `marketplace add` target is always the catalog repo — never the skill's own repo, which has no `marketplace.json` and fails.
      - `## Contributing`
      - `## License`
      
      **Optional (skills that ship a `commands/` dir):** `## Commands` — a table or
      list enumerating **every** slash-command **and each of its modes / sub-commands
      / flags**. When you add or rename a command *or a mode*, update this list in the
      **same change** — the implementation, its reference docs, and the README/AGENTS
      menus drift apart otherwise (a mode added everywhere except the README menu is a
      common, user-visible miss). `validate-skill.sh` warns when a command's `/<name>`
      (from `commands/<name>.md`) is not mentioned in the README, but it cannot see
      mode-level drift — that part is on the author.
      
      **Optional:** `## German summary` — short paragraph if DACH/TYPO3/Oro/agency audience; technical detail may remain English. The scaffold [`../templates/README.md.template`](../templates/README.md.template) ships this as a **commented block** after `## License` — uncomment when needed.
      
      ---
      
      ## Voice: present tense, never narrate change
      
      Reference docs (README, `SKILL.md`, reference files, module headers) describe **what the skill is and how to use it, in plain present tense — as if it had always been this way**. They must **not** describe the skill by what it is *not* or how it *changed*.
      
      - **Banned framing:** "no longer", "instead of", "reframed", "previously", "used to", "X derives from Y / downstream … not its source", "it is **not** a router/wrapper/…". This is the *curse of knowledge* as **negative / apophatic documentation** — it only parses for a reader who knew a prior design.
      - **Before the first release there is no audience for change.** Nobody ran a prior public version, so any before/after framing is noise and actively misleads (readers hunt for a "router" that never existed). The same holds for unreleased, in-between edits: the diff *is* the commit; the README must not know a change happened.
      - **History lives only in `CHANGELOG` / `UPGRADING` / release notes / ADRs / the commit log** — those have a reader (someone upgrading from a version they used) who needs the contrast. The README and `SKILL.md` never do.
      - A deprecation notice ("X is deprecated, use Y") is the one legitimate non-history contrast and belongs in the changelog/upgrade doc, not the "what this is" sections.
      
      ---
      
      ## Cross-checks (machine-friendly)
      
      | Check | PASS criterion |
      | --- | --- |
      | `Why this is a skill (model delta)` | Names ≥1 rubric value category, states one model-without-skill failure, links eval evidence where it exists. |
      | `Use when` | Section exists and contains trigger phrases (ticket prefixes, stacks, commands). |
      | `Example prompts` | ≥3 prompts. |
      | `Related skills` | ≥1 link/slug **or** justified none. |
      | `Installation` | Mentions marketplace **or** documents exclusive alternate with owner approval in README. |
      | Version badge (if any) | Uses `img.shields.io/github/v/release/…?sort=semver`; no `img.shields.io/badge/version-` URL. |
      | Present-tense voice | `grep -rin -e 'no longer' -e 'reframed' -e 'previously' -e 'used to' -e 'derive' -e 'downstream' -e 'not its source' -e 'is not a router'` over README/`SKILL.md` returns nothing — change-narration belongs in `CHANGELOG`/`UPGRADING`. |
      | `Commands` (if `commands/` exists) | Every `commands/<name>.md`, and each mode/flag it documents, is enumerated in the README. |
      
      ---
      
      ## Alignment with other files
      
      - **Discovery YAML** (optional): [`skill-discovery-metadata.md`](skill-discovery-metadata.md) — keep summaries consistent with README sections.
      - **`agents/openai.yaml`**: short end-user description; must not contradict README summaries.
      - **GitHub**: see [`repository-quality-rules.md`](repository-quality-rules.md) for description/topic rules.
      
    • release-discipline.md 50.5 KB
      # Release Discipline
      
      ## Contents
      
      - Fleet sweeps: use the shipped driver, do not hand-write a new one
      - Canonical Order: Bump PR Merged → Tag Pushed
      - Pre-Release Version-Parity Check
      - Changelog Rollover in the Bump Commit
      - A reformatting hook aborts the first bump commit
      - Cache Safety: Never Edit the Installed Copy
      - Multi-Skill-Repo Release Dry-Run
      - GitLab (`git.netresearch.de`) skill repos release on tag too
      - A pushed tag is not a release — check the pattern the release job matches
      - Release notes: the generated body is the default — do not curate it away
      - Immutable-Release Caveat
      - Tag Signing (Mandatory)
      - No `--latest` Drift for Non-Default Branches
      - What a release archive contains
      - Supply-Chain Attestation
      
      Every step that caused the "30 failed plugin releases" incident, codified as rules.
      
      ## Fleet sweeps: use the shipped driver, do not hand-write a new one
      
      Everything below that a fleet sweep needs is mechanized in
      `scripts/fleet-release-github.sh` on top of the host-neutral
      `scripts/fleet-release-common.sh` engine. The driver runs `survey` →
      `manifest` (stops for human approval; versions and PR bodies are operator
      judgment the scripts never invent) → `bump` → `finish`, resumable from remote
      reality at every step. The changelog rollover below is its own tool,
      `scripts/roll-changelog.py`. Fleets on other hosts keep their driver, their
      repo list and their host specifics in their own (private) repository — that
      driver vendors the shared engine and implements the `host_*` callbacks its
      header documents; host knowledge stays with the host. Every sweep before
      these existed re-implemented this page as one-off scripts and re-earned the
      same bugs (#218 records the last one: brace-group exits, pretty-printed survey
      rows, `//` on empty strings — all now encoded as tests). Deliberately out of
      scope, by design: merging anyone else's PR, unarchiving, cloning missing
      checkouts, writing CI files (a NO-RELEASE-JOB manifest row means adopt the
      shared release CI first), and non-default-branch releases — the
      `--latest=false` rule below stays a manual concern. When a sweep needs
      something the driver lacks, extend it in a PR; a session-local driver script
      is how the next incident starts.
      
      **Run the sweep from a checkout you have just pulled, and read the FIRST bump
      commit before letting the phase continue.** `git fetch` is not `git pull`: a
      driver copy can be behind the feature you are relying on, and a copy that
      predates `FR_COMMIT_TRAILERS` ignores the variable rather than complaining about
      it — every repo reports `OK … bump PR open` and the commits carry nothing. On
      2026-09-11 that shipped four bump commits to `main` without their disclosure
      trailers before anyone opened one; auto-merge had already taken them past the
      point where an amend was possible, and rewriting `main` to fix it is worse than
      the defect. `fr_assert_trailers_landed` now fails the repo when a requested
      trailer is absent from the commit, which closes every producer of that symptom
      except the one it cannot see: a stale engine does not know the variable, so it
      cannot check for it. Hence the manual half — `git -C <skill-repo-skill> pull`,
      then `git log -1 --format=%B` on the first repo the phase touches. The same
      applies to a vendored engine: a fleet driver on another host reads *its own*
      copy, so an upstream fix reaches it only after a re-vendor.
      
      **Read that commit from the PR, not from the branch — the bump phase arms
      auto-merge, so the branch is often already merged and deleted by the time you
      look.** Both branch-shaped reads then fail, and they fail in ways that read like
      the bump never happened rather than like it succeeded:
      `gh api repos/$O/$R/commits/release/v$V` answers `422 No commit found for SHA`
      and `gh api "repos/$O/$R/commits?sha=release/v$V"` answers `404`. The PR record
      survives the branch deletion:
      
      ```bash
      N=$(fr_opened_lookup "$R")          # host-filtered, last record wins
      gh api "repos/$O/$R/pulls/$N/commits" --jq '.[].commit.message'
      ```
      
      Take the id from `fr_opened_lookup`, not from a hand-written `jq` over
      `opened.jsonl`. That file accumulates across runs and can hold both hosts, so
      filtering on `.repo` alone yields several ids — which interpolate into a
      multi-line path that either queries the wrong PR or fails outright. The helper
      filters `.repo` *and* `.host` and keeps the last match.
      
      This is the one check standing between a missing trailer and a defect that
      cannot be fixed after the merge, so it must not be the step that gets skipped
      because the documented command returned an error.
      
      **Where a policy requires agent or tool disclosure on every commit, set
      `FR_COMMIT_TRAILERS`** — newline-separated `Key: value` lines that the bump
      commit carries as git trailers, alongside the `--signoff` it already writes.
      The driver appends them verbatim and interprets none of them, so the keys stay
      the operator's decision (`Assisted-by:`, `Agent-Session:`, … ). It applies to
      both commit attempts, including the retry after a reformatting hook — the one
      place a second `git commit` could otherwise drop them. Unset, nothing changes.
      
      **Budget the API before a sweep: the survey costs ~10 calls per repo, and
      each GitHub quota pool (REST 5,000/h, GraphQL 5,000 points/h — separate
      pools) is shared across every tool, watcher and agent in the session.** A
      40-repo survey plus check-watchers plus a merge-gate poll can exhaust them
      mid-sweep — the GraphQL pool usually dies first because `gh pr merge`,
      `gh pr view --json` and pr-status.sh all draw from it, while plain REST
      still answers apart from short burst limits (2026-08-13: a sweep session hit
      exactly this between "all gates verified" and "merge"). Practice: act on a verified, unchanged head
      with ONE call instead of re-running preflight batteries — the REST merge
      takes an `sha` pin (`gh api -X PUT .../pulls/N/merge -f merge_method=merge
      -f sha=$(git rev-parse HEAD)`) whose 409 IS the freshness check — prefer one
      `--watch` over repeated status reads, and stop any watcher whose answer is
      already known.
      
      **A batch that waits on a prerequisite PR runs in its own `--workdir`.** The
      lock is per workdir and covers every phase, and a `finish` run holds it until
      each of its merge gates closes — up to `FR_MERGE_TIMEOUT` per repo. Repos held
      back for a prerequisite (an `EMPTY` changelog, a CI fix) therefore cannot be
      bumped next to a running `finish`: the second call dies with `workdir locked by
      running pid …`. Copy their plan rows into a fresh workdir and run `bump` and
      `finish` there. That is the intended shape, not a workaround — one workdir per
      batch keeps `opened.jsonl` and the logs from interleaving, which is what the
      lock exists to prevent. (2026-09-27: five repos with an empty `[Unreleased]`
      ran as a second batch while the first batch's `finish` held the lock.)
      
      ## Canonical Order: Bump PR Merged → Tag Pushed
      
      **Tag a version only after the version-bump PR is merged to the default branch.** Tagging first causes the Release workflow to run against the old code, fail CI, and produce an immutable GitHub release locked to a bad tag.
      
      ```
      WRONG: git tag -s v1.2.4 → git push → open bump PR
      RIGHT: open bump PR → merge → pull main → git tag -s v1.2.4 → git push
      ```
      
      ## Pre-Release Version-Parity Check
      
      Before pushing any tag, all version identifiers must match. This is the single check that would have prevented the 30-repo release failure.
      
      Use the shipped script `scripts/check-version-parity.sh` (in this repo under `skills/skill-repo/scripts/check-version-parity.sh`):
      
      ```bash
      # No arguments — compare plugin.json against SKILL.md metadata.version
      skills/skill-repo/scripts/check-version-parity.sh
      
      # With tag argument — also require plugin.json.version == tag (v prefix optional)
      skills/skill-repo/scripts/check-version-parity.sh v1.2.4
      
      # From a driver that iterates over many checkouts — name the repo explicitly
      skills/skill-repo/scripts/check-version-parity.sh --repo /path/to/repo v1.2.4
      ```
      
      All paths the script reads are relative to the repo root, so without `--repo`
      the root is the current directory. A fleet driver that passes the repo as a
      bare argument gets it parsed as a *tag* and the script then aborts on the
      missing `plugin.json` — after the bump has already been written to disk,
      leaving every repo in that batch dirty (observed 2026-08-06: five repos in one
      batch). Pass `--repo`, or `cd` into the checkout first.
      
      What it checks:
      
      - Root `plugin.json` (Agent Plugins manifest), when present, is the **source of truth**: bump it first, run `sync-plugin-manifest.sh`, and the check fails if `.claude-plugin/plugin.json` still disagrees. Repos that have not adopted the portable manifest are unaffected — see [`agent-plugins-compat.md`](agent-plugins-compat.md).
      - `.claude-plugin/plugin.json` has a `.version` field — exits with an error if missing.
      - `composer.json` does **not** have a `.version` field — composer versions come from the git tag via the Release workflow, so a hard-coded version drifts silently.
      - If a tag argument is provided, `plugin.json.version` equals that tag with the `v` prefix stripped.
      - Every `skills/*/SKILL.md` that declares a version in frontmatter — `metadata.version` *or* a top-level `version:` key (both forms exist in the fleet; some SKILL.md files declare none, which is fine) — matches `plugin.json.version`.
      
      If called without an argument and all parity passes, the script prints an advisory suggesting the next tag call. Run before every `git push origin vX.Y.Z`.
      
      Bump tooling must handle the same two frontmatter forms: a bump script that only rewrites the indented `metadata.version` silently leaves a top-level `version:` at the old value, and the tag pipeline's `validate:skill` then fails on exactly that mismatch (it-maintenance-skill v1.10.0 died this way; typo3-upgrade-estimator-skill nearly repeated it in the 2026-07-16 sweep).
      
      Use `scripts/bump-version.sh <version> [--apply]` rather than a per-repo helper. It writes every surface `check-version-parity.sh` validates — the root `plugin.json`, the `.claude-plugin/plugin.json` projected from it, plus *each* frontmatter `version:` line in *every* `skills/*/SKILL.md`, both forms, indentation and quoting preserved — refuses when `composer.json` carries a version, and re-runs the parity check afterwards. Dry-run by default.
      
      **The root `plugin.json` is the surface a bump is most likely to miss.** Until v1.28.0 the script wrote only the *generated* `.claude-plugin/plugin.json`, so on every repo that had adopted the portable manifest it left the source of truth at the old version, failed its own parity check, and exited 1 with the tree half-written. That is the same failure mode as the two frontmatter forms above, one level up: the surface the tooling does not know about is the one that silently keeps the old value. A fleet sweep hits it in every repo at once — 65 of them on 2026-08-08.
      
      It deliberately does not commit, tag or push. A helper that bumps and tags in one step is how the canonical order above gets skipped: the tag then lands on an unmerged branch, or on a tree where only one of the version surfaces moved. Two repos grew such a target locally (`make release` in it-account-lifecycle-skill and it-maintenance-skill) and both diverged from this page — one bumped only `plugin.json`, the other hard-coded a single skill path and produced the v1.10.0 failure above.
      
      ## Changelog Rollover in the Bump Commit
      
      If the repo maintains a `CHANGELOG.md`, the version-bump commit moves the `[Unreleased]` content under a new `## [X.Y.Z] - YYYY-MM-DD` heading and leaves a fresh empty `[Unreleased]` section. A bump commit that skips this leaves shipped content labeled `[Unreleased]` — the next release then has to relabel history after the fact (typo3-upgrade-estimator-skill shipped its entire v2.2.0 changelog block as `[Unreleased]` and it was only relabeled in v2.2.1).
      
      **An `EMPTY` manifest row is a prerequisite PR, not a bump.** The roll refuses
      an empty `[Unreleased]`, so such a repo needs its entries on the default branch
      first:
      
      1. Open one `docs(changelog)` PR per repo that fills `[Unreleased]` from what
         shipped since the last tag — the merged PR bodies *and* their diffs, since
         follow-up commits often change what a body describes.
      2. Check each entry's scope claim against the shipped reference file, not the PR
         body alone. Of five such PRs on 2026-09-27, two overstated their scope ("a
         `vcs` entry" where the reference says a `vcs` entry with a GitHub URL; "a bare
         `@login`" where the script strips inline links only), and only a bot review
         caught them.
      3. Merge it, then bump the repo in its own workdir (see "A batch that waits on a
         prerequisite PR" above).
      
      **Five heading shapes exist in the fleet**, not the two this page claimed until
      v1.29.0 — the 2026-08-08 sweep hit all of them across 63 repos:
      
      | Shape | Example | Seen in |
      |---|---|---|
      | bracketed dash | `## [1.2.3] - 2026-08-08` | most keep-a-changelog repos |
      | bare paren | `## 1.2.3 (2026-08-08)` | it-maintenance-skill, gitlab-skill |
      | bracketed em dash | `## [0.3.23] — 2026-07-02` | nr-monatliche-abrechnung |
      | linked | `## [v2.6.0](…/releases/tag/v2.6.0) — 2026-02-28` | typo3-docs-skill |
      | `[Unreleased]` only, no released heading yet | — | source-digest-skill |
      
      A roll anchored on one shape matches nothing in the others and exits 0, so the
      release ships with its content still under `Unreleased` and nobody sees an error.
      Detect the shape from the newest *released* heading and reproduce it — including
      the dash character and, for the linked form, rewriting the tag inside the URL.
      The fifth shape has no released heading to copy from; that is the first release,
      not an error, so default to the keep-a-changelog form its `[Unreleased]` bracket
      already implies rather than aborting.
      
      **Fail the roll when it changed no lines** — a no-match must not be
      indistinguishable from a successful roll. Same rule as for `metadata.version` vs
      a top-level `version:`: whichever surface the tooling does not know about is the
      one that silently keeps the old value.
      
      **A generated entry must satisfy markdownlint, because every repo lints its
      CHANGELOG in CI.** Emit a blank line after each `### Added`/`### Fixed` heading
      and around every list, or the bump PR goes red on MD022 (blanks-around-headings)
      and MD032 (blanks-around-lists) — in one repo, in every repo, all at once. This
      red-lit 16 of 63 repos mid-sweep on 2026-08-08. It is worth running
      `npx markdownlint-cli2 CHANGELOG.md` on the rolled file before committing:
      the roll is generated text, and generated text is exactly what nobody proofreads.
      
      Boundary regexes are also **fence-blind** — a scan anchored on `^##` matches a
      heading-looking line inside a fenced code block and splices the new section into
      the middle of an example. Track fences while scanning.
      
      ## A reformatting hook aborts the first bump commit
      
      A hook that rewrites a file it is checking — `pretty-format-json --autofix`,
      `black`, `ruff format`, `end-of-file-fixer` — modifies the staged tree and then
      *fails* the commit; re-staging its output and committing again is the remedy,
      and the second run passes because the file is already in the shape the hook
      wants. `fr_commit_allowlisted` therefore makes two attempts, and the log shows
      both (`COMMIT EXIT (1): 1`, `COMMIT EXIT (2): 0`). The retry re-walks the
      porcelain rather than re-adding the first pass's paths, so a hook that writes
      *outside* the version surfaces still fails the allowlist on the second pass.
      
      Before that retry existed, such a repo left the `bump` phase as
      `FAIL <repo>: commit` with a half-written worktree the resume then refused as
      dirty, and the operator had to commit by hand and re-run — `it-maintenance-skill`
      v1.15.0 and `netresearch-jira-skill` v2.11.1 both did on 2026-08-28 (#263).
      Their trigger is worth knowing because it recurs: `sync-plugin-manifest.sh`
      emits `.claude-plugin/plugin.json` in the source manifest's key order and
      `pretty-format-json` sorts the keys, so the two rewrite each other on every
      bump. Sorting the generator's output would settle that one pairing and no
      other — the same hook also escapes non-ASCII, which the fleet's descriptions
      are full of — so the retry, which is blind to *what* the hook changed, is the
      fix rather than agreeing with one formatter.
      
      ## Cache Safety: Never Edit the Installed Copy
      
      Installed skills and plugins live under `~/.claude/` (or wherever the marketplace resolves them). Editing these paths directly is always wrong — the next `/plugin update` or marketplace sync will silently overwrite your changes, taking any local fixes with it.
      
      ### Paths that are off-limits for edits
      
      - `~/.claude/skills/**`
      - `~/.claude/plugins/cache/**`
      - `~/.claude/plugins/marketplaces/**`
      - Anything inside a `.bare/` directory (git bare clone; worktree source)
      
      ### Pre-edit check
      
      Before every Write or Edit in skill-repo workflows:
      
      ```bash
      pwd_real=$(realpath .)
      case "$pwd_real" in
        */.claude/skills/*|*/.claude/plugins/*|*/.bare/*)
          echo "REFUSING to edit installed/cache path: $pwd_real"
          echo "Navigate to the source worktree first."
          exit 1
          ;;
      esac
      ```
      
      ### Recovery when edits landed in the wrong place
      
      1. Stop. Do not run `/plugin update` or `composer update` — they may wipe your edits.
      2. `diff -r ~/.claude/skills/<name>/ ~/projects/<name>-skill/main/skills/<name>/` to see what drifted.
      3. Copy the legitimate changes into the source worktree.
      4. Commit from the worktree; never from the cache.
      
      ## Multi-Skill-Repo Release Dry-Run
      
      When releasing >3 skill repos in one sweep, produce this manifest and wait for user approval before executing:
      
      ```
      Skill-Repo Release Plan (2026-04-18)
      
      | Repo                        | Current | Target  | Change type | Bump PR   | Notes                  |
      |-----------------------------|---------|---------|-------------|-----------|------------------------|
      | netresearch/git-workflow    | 1.9.0   | 1.10.0  | minor       | mine      | adds critical-rules    |
      | netresearch/github-project  | 2.10.0  | 2.11.0  | minor       | mine      | multi-repo-operations  |
      | netresearch/skill-repo      | 1.18.0  | 1.19.0  | minor       | #42 @kim  | NEEDS AUTHOR'S GO      |
      
      Preconditions (verified per repo):
        [✓] default branch CI green
        [✓] no pending version-bump PR by another author
        [✓] version-parity check passes
        [✓] working tree clean
      
      Execution order per repo:
        1. Create version-bump PR on release/vX.Y.Z branch
        2. Wait for CI green and approval
        3. Merge via merge-commit (respects atomic-commit policy)
        4. Pull main; run check-version-parity.sh vX.Y.Z
        5. Create signed tag vX.Y.Z
        6. Push tag
        7. Monitor Release workflow to green
        8. Halt all further releases if this one fails — produce rollback
      
      Reply "go" to proceed, or name repos to skip.
      ```
      
      The branch and commit names are not this file's to invent: `release/vX.Y.Z` and
      `chore(release): vX.Y.Z` are owned by `github-release-skill`
      (`commands/release.md`, steps 4 and 8). Follow that skill when the two disagree.
      
      ### A pre-existing bump PR by another author is not covered by the sweep's approval
      
      A sweep will sometimes find a repo whose version-bump PR is already open —
      opened by a colleague, days earlier, mergeable and green. Merging it is what the
      release needs, and the manifest row looks exactly like every other row, which is
      the problem: a single "go" over a 19-row table reads as approval of *your* work,
      and the author is never asked.
      
      Give the manifest a **Bump PR** column naming the author of every pre-existing
      PR, mark those rows as needing that author's go-ahead, and get it separately.
      Blanket batch approval covers only the PRs the sweep itself opens. (Observed
      2026-08-06: `ecom-orocommerce-docker-skill !5`, authored by a colleague, was
      merged inside a 19-repo batch on one blanket approval.)
      
      ### Building the manifest: fleet-survey gotchas
      
      **Prefer a remote-first survey — the GitHub API answers the whole classification without touching a checkout** (verified in the 2026-08-03 sweep: 40 repos surveyed, 7 released). Three calls per repo: `gh api repos/$O/$R/releases/latest` (tag of the latest *published release* — a 404 means the repo has never released; classify it for a first release instead of skipping), `gh api "repos/$O/$R/compare/<tag>...main"` (ahead-count, commit subjects *and* changed files in one response — enough for both the CI-only-delta filter and the bump-type decision; name the default branch explicitly, consistent with the `origin/main` guidance below), `gh api repos/$O/$R/contents/.claude-plugin/plugin.json` (prepared-vs-needs-bump). The local-checkout gotchas below then apply only to the repos that actually release.
      
      **The GitLab arm needs its own call triple — the response shapes differ.** Half
      the fleet lives on `git.netresearch.de/coding-ai`, and a GitHub-shaped `jq`
      filter returns *empty* against these rather than erroring, so the repo looks
      like it has no delta:
      
      ```bash
      P="coding-ai%2F$R"
      glab api "projects/$P/releases?per_page=1"                                  # .[0].tag_name; empty ⇒ never released
      glab api "projects/$P/repository/compare?from=$tag&to=main"                 # .commits[].title  and  .diffs[].new_path
      glab api "projects/$P/repository/files/.claude-plugin%2Fplugin.json/raw?ref=main"
      ```
      
      Note `.diffs[].new_path` where GitHub has `.files[].filename`, and
      `.commits[].title` where GitHub has `.commits[].commit.message`. GitLab's
      compare returns no `ahead_by`, so count `.commits[]` yourself. Verified in the
      2026-08-06 sweep (74 repos surveyed across both hosts, 19 released).
      
      **Compute the CI-only-delta filter, do not eyeball it.** The rule below ("CI-only
      deltas are not releases") is only usable if the survey reports it per repo. From
      the same compare response:
      
      ```bash
      jq '[.files[].filename] | map(select(test("^\\.github/|^\\.gitlab-ci\\.yml$|^renovate\\.json$") | not)) | length'
      ```
      
      A zero means the release archive would be byte-identical to the last tag — skip
      the repo and say so in the manifest. In the 2026-08-06 sweep this filter alone
      removed 13 of 32 repos that had commits since their tag.
      
      Surveying dozens of local skill-repo checkouts for "commits since last tag" hits these, verified in the 2026-07-16 sweep (16 releases):
      
      - **Three checkout layouts coexist** under the projects dir: bare + worktrees (`repo/.bare`), a plain repo at the top level (`repo/.git`, which is a pointer *file* when the top level is itself a worktree), and a plain clone one level down (`repo/main/.git`). Resolve the git dir per repo instead of assuming one shape:
      
        ```bash
        G=$([[ -d "$d/.bare" ]] && echo "$d/.bare" \
            || git -C "$d" rev-parse --absolute-git-dir 2>/dev/null \
            || git -C "$d/main" rev-parse --absolute-git-dir 2>/dev/null)
        ```
      
        `rev-parse --absolute-git-dir` handles both plain repos and worktree pointer files; it cannot discover `$d/.bare` (a worktree *parent* dir is not a repo, and upward discovery never looks into `.bare`), and it must not run first from `$d` when only `$d/main` is the repo — hence the explicit order.
      - **`git fetch --tags` fails wholesale on one stale tag** (`would clobber existing tag`), taking the branch fetch down with it. Survey with a branches-only refspec plus `git ls-remote --tags origin` and the peeled (`^{}`) SHA for the ahead-count; never resolve the tag locally.
      - **Duplicate checkouts happen** (two dirs, same `origin`). Dedupe the manifest by remote URL, not by directory name, or the same repo gets two bump PRs.
      - **A per-repo `exit` inside a `{ … } > log` block kills the whole driver.** A brace group is not a subshell, so the first failing repo silently cancels every repo after it — the batch "completes" with fewer log files than repos and nothing says so. Wrap each repo's body in `( … ) > log` instead, and compare log count against repo count before trusting the summary. (2026-08-13: a three-repo driver died on repo two; repo three was skipped, revealed only by its missing log.)
      - **Emit survey rows with `jq -nc`, not `jq -n`.** Bare `jq -n >> file` pretty-prints, so `wc -l` counts JSON lines, not rows — a 74-repo survey reported "1400 rows". Downstream `jq -c .` still parses the concatenated stream, but every row-count sanity check lies until the file is one-object-per-line.
      - **jq's `//` treats the empty string as present.** Fields captured as `""` (not null) sail straight through `.last_release // .last_tag`, so the coalesce prints nothing instead of falling back. When the producer writes empty strings, select explicitly: `if .last_release != "" then .last_release else .last_tag end`.
      - **A renamed repo answers on both names — the API redirect turns one repo into two survey rows.** `gh api repos/$O/<old-name>` follows the rename redirect and returns the new repo's data wholesale, so a fleet list that still carries the old name surveys the same repo twice, and the sweep would open two bump PRs against one repo. Dedupe survey rows by the response's `.full_name`, never by the name you queried (2026-08-13: `agents-skill` surveyed as a live repo with the exact ahead-count and version of `agent-rules-skill` — it is a redirect to it).
      - **CI-only deltas are not releases.** If every unreleased commit touches only `.github/**`, `.gitlab-ci.yml`, or `renovate.json`, the release archives would be byte-identical to the last tag — skip the repo and say so in the manifest.
      - **A bump PR with an unresolved bot review thread never auto-merges, and `finish` waits it out.** The `bump` phase arms auto-merge, so the natural reading of an open PR is "CI still running". A `copilot_code_review` (or CodeRabbit) thread on the rolled CHANGELOG blocks the merge with every check green and nothing red anywhere: `mergeStateStatus` says `BLOCKED`, the checks summary says all pass, and the merge gate in `finish` then polls until its timeout. **Query `--state all` when you sweep:** the same auto-merge arming means a PR can already be `MERGED` by the time you look, and an open-only list then returns nothing for that repo — which reads exactly like a bump that never opened a PR (2026-09-12: `agent-harness-skill` had merged and released while the sweep reported it as having no PR). Sweep the PRs the `bump` phase opened *before* running `finish`, and resolve what it finds:
      
        ```bash
        gh api graphql --paginate -f o="$O" -f r="$R" -F n="$N" -f query='
          query($o:String!,$r:String!,$n:Int!,$endCursor:String){repository(owner:$o,name:$r){
            pullRequest(number:$n){reviewThreads(first:100,after:$endCursor){
              nodes{id isResolved} pageInfo{hasNextPage endCursor}}}}}' \
          | jq -s '[.[].data.repository.pullRequest.reviewThreads.nodes[]|select(.isResolved==false)]|length'
        ```
      
        **Paginate it, and do not lower the page size.** A fixed `first: 20` returns the
        first page and says nothing about the rest, so a PR with more threads reports
        zero unresolved and sails through the gate this check exists to close — the
        false-negative twin of the implausible-rate rule, in the one place where a wrong
        zero is indistinguishable from a clean PR. `--paginate` needs the `$endCursor`
        variable and the `pageInfo` block above, and `jq -s` to slurp the pages it emits.
      
        In the 2026-09-03 sweep this hit 3 of 4 PRs in one batch — all three findings were real defects in the *generated* changelog, which is exactly the text nobody proofreads. Budget a review round for the roll's output rather than treating a bump PR as mechanical.
      - **A `FAIL <repo>: merge gate` with a CONFLICTING bump PR is usually a parallel release, not a conflict to resolve.** A sweep surveys once and acts minutes later, and another session on the same account can release one of the surveyed repos inside that window. The sweep's own bump PR then sits on a moved `main` with a conflict in `plugin.json`, every check green, `mergeable=CONFLICTING mergeState=DIRTY` — which reads as "rebase it". Rebasing is wrong whenever the parallel release is *higher*: read `origin/main`'s `plugin.json` version and the latest release first, and check whether that release's compare range already covers the commits the sweep's version targeted. When it does, the sweep's PR is redundant and a rebase would move the version backwards; close it, naming the superseding release and PR, and delete the branch. (2026-09-20: `skill-repo-skill` was released as v2.2.0 from #335 at 10:49 while a sweep held a v2.1.6 bump PR (#336) open; v2.2.0 covered both commits v2.1.6 existed for.) The engine cannot distinguish this from a real conflict today — it reports the generic merge-gate failure — so the operator makes the call.
      - **Dependency commits landing after a tag are not a second release; commits under `skills/**` are.** The same window produces both. A post-tag delta touching only `.pre-commit-config.yaml` or `renovate.json` ships nothing, because the release archive's allow-list does not carry those files. One that touches a skill's `SKILL.md`, `references/` or `scripts/` does, and a sweep that measured once will not see it: re-check the repos it released before reporting the fleet clean. (2026-09-20: of six GitHub repos with commits after their new tag, five were pre-commit bumps and one — `git-workflow-skill` — had two `references/advanced-git.md` commits, released as v1.34.1.)
      - **Archived repos are release-infeasible** (pushes rejected). A deprecated repo whose deprecation-banner commit landed *after* the last tag has never shipped its own deprecation notice — flag it for an unarchive decision instead of silently skipping.
      - **The unarchive → final release → re-archive path** (walked 2026-08-13 for claude-coach-plugin v2.6.0 and typo3-frontend-patterns-skill v1.2.1) has its own traps. A sweep that was blocked by the archive may have left a prepared, signed local bump branch — verify its signature and content, then push and reuse it rather than recreating. Expect the frozen repo to fail reusables that evolved while it slept, and fix proportionately: measure ruff with the fleet config (`select = ["E","F","W"]`, `line-length = 120`) before concluding code work exists — 439 default-rule findings were **zero** under it, so the fix was one `ruff.toml`, not a rewrite. `ruff format` also rewrites ` ```python ` blocks inside `.md` files, and a config change invalidates every earlier `--check` measurement — re-measure after the config lands, and AST-verify the formatted files before claiming the change is behavior-free. A red Template Drift paired with a Security `startup_failure` usually shares one cause: the frozen `security.yml` still passes the org-removed `GITLEAKS_LICENSE` secret — syncing `templates/skill/.github/workflows/security.yml` from `netresearch/.github` fixes both at once. Re-archive only after the Release object and its assets are verified.
      - **Cleaning up after the sweep is not an ancestry question.** The branches a sweep leaves behind get classified by whether their work is safe to drop, and `git merge-base --is-ancestor <branch> origin/main` does not answer that: a squash-merged PR puts the content on `main` under one new commit, so the branch SHAs are absent and shipped work reads as "unmerged". Use the PR/MR state, with `git cherry` as the offline fallback — `git-workflow-skill` → `references/advanced-git.md` ("Is This Branch Safe to Delete?") owns the rule and the commands. Applies to the *report* as much as the deletion: in the 2026-08-06 sweep this misclassification turned 2 genuinely unsaved commits into 18 alarming-looking branches.
      - **Read the version state from `origin/main`, never the local worktree.** Worktrees drift — some sit behind the remote, some ahead, some are dirty — so `cat repo/.claude-plugin/plugin.json` gives a misleading parity picture. Read the authoritative value with `git show origin/main:.claude-plugin/plugin.json` after fetching. `git show` takes a single literal path and does **not** glob, so enumerate the SKILL.md files first, then read each:
      
        ```bash
        for f in $(git ls-tree -r --name-only origin/main | grep -E 'skills/[^/]+/SKILL\.md$'); do
          git show "origin/main:$f"
        done
        ```
      
        Verified in the 2026-07-18 sweep (23 releases), where worktree reads disagreed with `origin/main` on ~6 repos.
      - **Survey for a `composer.json` `version` field before the sweep, not at bump
        time.** `bump-version.sh` and `check-version-parity.sh` both refuse while it is
        present, so a repo carrying one fails mid-bump with a half-written tree and has
        to be re-run after the field is removed. Four GitLab repos hit this on
        2026-08-08, and three of the four were already *behind* their own latest tag
        (`dxp-project-init` declared `0.1.0` against tag `v0.2.6`) — which is precisely
        the drift the no-version rule exists to prevent, so removing it is the fix, not
        a workaround. One extra call per repo during the survey turns a mid-sweep
        failure into a manifest row:
      
        ```bash
        glab api "projects/coding-ai%2F$R/repository/files/composer.json/raw?ref=main" | jq -e 'has("version")'
        ```
      
      - **A QUEUED GitHub check reports `conclusion: ""`, not `null` — so the obvious
        red-check filter calls it a failure.** The natural gate,
        `select(.conclusion != null and .conclusion != "SUCCESS" …)`, lets the empty
        string through and the sweep aborts on checks that were merely still starting.
        Gate on completion first and only then on the verdict:
      
        ```bash
        gh pr view "$BRANCH" --repo "$O/$R" --json mergeStateStatus,statusCheckRollup --jq '{
          p: [.statusCheckRollup[] | select((.__typename=="CheckRun" and .status!="COMPLETED")
                                          or (.__typename=="StatusContext" and .state=="PENDING")) | .name],
          f: [.statusCheckRollup[] | select(.__typename=="CheckRun" and .status=="COMPLETED"
                                          and (.conclusion|IN("SUCCESS","SKIPPED","NEUTRAL")|not)) | .name]}'
        ```
      
        Wait while `p` is non-empty; fail only on `f`. The two shapes in the rollup need
        different fields — `CheckRun` has `.status`/`.conclusion`, `StatusContext` has
        `.state` — and a filter written for one silently mis-reads the other. Note also
        that `gh pr checks --watch` only blocks *while it can run*: if its output
        redirect fails (a missing log directory), the command dies instantly and the
        very next query reads a still-pending rollup. Observed 2026-08-08.
      - **`plugin.json` ahead of the last tag ⇒ the bump already merged — tag only, no bump PR.** Classify each repo from the `origin/main` value: `plugin.json.version > last tag` means a prior bump PR already landed and the repo just needs a signed tag; `plugin.json.version == last tag` means it needs a bump PR first. Tagging a "prepared" repo at its committed version keeps parity true by construction. Don't open a bump PR for a repo that is already prepared — it would double-bump.
      - **A never-released repo can be prepared too — plan it at its committed version, do not invent one.** `FIRST-RELEASE` (no release *and* no tag) usually already carries the version someone meant to ship in `.claude-plugin/plugin.json`, so the plan row names exactly that and the bump phase reports "nothing to bump; finish will tag". The monotonicity check therefore accepts `version >= plugin.json` for this class and only refuses a version *below* it, where parity would fail at tag time — unlike `BUMP`, which still requires strictly greater, because there a released version exists to regress from. Until v1.35.0 equality was refused for every class, which left a prepared first release unrepresentable: the operator had to retype the row `TAG-ONLY` or invent a version nobody wrote (both `ecom-*` repos in the 2026-08-17 sweep).
      - **Splicing re-surveyed rows into an existing plan drops `version` and `body`.**
        After a partial re-survey (e.g. the freshness gate tripped on renovate merges),
        the regenerated `plan.skeleton.jsonl` rows carry empty `version`/`body` for
        TAG-ONLY rows the operator had already filled. Merge the fresh `head`/
        `classification` into the OLD plan rows, keeping their `version` and `body` —
        replacing rows wholesale ships a release with no notes source (2026-08-27:
        markdown-to-pdf briefly published "First release" as its v1.4.2 notes this way).
      - **`git pull origin main` merges into whatever branch the worktree has checked out.** A fleet sweep hits worktrees parked on a leftover feature branch, and `--ff-only` does not protect you — the stale branch fast-forwards onto `origin/main` "successfully" (observed 2026-08-03: a leftover `ci/*` branch silently advanced to the main tip).
      
        **Do not check out at all — tag `origin/main` directly.** Fetching, verifying
        and tagging need no working tree, which removes the hazard instead of stepping
        around it, and it is the only form that leaves a colleague's parked worktree
        untouched (in the 2026-08-06 sweep, `skill-repo-skill`'s only checkout sat on a
        leftover `release/v1.25.2` the whole time; `git switch main` there would have
        moved someone's workspace out from under them):
      
        ```bash
        git -C "$G" fetch origin --prune --quiet
        remote_tip=$(gh api "repos/$O/$R/commits/main" --jq .sha)
        [[ "$(git -C "$G" rev-parse origin/main)" == "$remote_tip" ]] || exit 1
        # parity read from the ref, never from a working tree
        git -C "$G" show "origin/main:.claude-plugin/plugin.json" | jq -r .version
        git -C "$G" tag -s "v$V" -m "v$V" "$remote_tip"
        git -C "$G" push origin "v$V"
        ```
      
        `git tag` takes any commit-ish, and it does not touch the index or the working
        tree — so this is safe even in a plain (non-bare) checkout that is mid-edit.
        Keep switch-then-pull only for the repos where you genuinely need files on disk.
      
      ## GitLab (`git.netresearch.de`) skill repos release on tag too
      
      The release-discipline above is GitHub-centric, but the `coding-ai` GitLab skill repos ship the same way: **a pushed signed tag triggers a pipeline that creates the GitLab Release** via the `claude-code-skill` CI component (`include:` in `.gitlab-ci.yml`). There is no separate release workflow file and **no manual `glab release create`** — pushing `vX.Y.Z` is sufficient; the Release object appears when the tag pipeline goes green. Confirm with `glab api "projects/coding-ai%2F<repo>/releases/v<ver>"` (and the tag pipeline via `pipelines?ref=v<ver>`). Same bump-PR-then-tag order as GitHub applies; merge the bump MR only when `detailed_merge_status == "mergeable"`.
      
      This holds only for repos whose release job is gated on the tag you actually push — see [A pushed tag is not a release](#a-pushed-tag-is-not-a-release--check-the-pattern-the-release-job-matches) below, which is where the exceptions live.
      
      ## A pushed tag is not a release — check the pattern the release job matches
      
      The sentence above ("pushing `vX.Y.Z` is sufficient") holds **only while the tag
      you push matches the pattern that repo's release job is gated on.** When it does
      not, the push succeeds, the pipeline may even go green, and no Release object is
      ever created. Nothing fails; the release is simply missing. Three variants, all
      found in the 2026-08-08 sweep:
      
      | Repo | Tag pushed | Release rule | Result |
      |---|---|---|---|
      | `dxp-project-maintenance` | `0.5.4` (bare) | component: `^v\d+\.\d+\.\d+$` | job never ran; releases hand-made for months |
      | `oro-bundle-upgrade-skill` | `v1.1.0` | no tag rule at all — CI runs only on MRs and the default branch | job does not exist; releases hand-made |
      | `tender-estimation` | `tender-estimation--v0.6.1` | `^tender-estimation--v\d+\.\d+\.\d+$` | **correct** — a deliberate non-standard convention |
      
      So the check is not "does the repo use `vX.Y.Z`" but **"does its release job's
      rule match the tag its maintainers actually cut?"** The third row is the reason:
      that prefix is deliberate, documented in the file as matching the
      `claude plugin tag` CLI, and its release job fires on it. Normalising it to
      `vX.Y.Z` would have *broken* a working release path.
      
      Before tagging a repo you have not released before, read the whole
      `.gitlab-ci.yml` / release workflow and compare the rule against the last tag:
      
      ```bash
      glab api "projects/coding-ai%2F$R/repository/tags?per_page=1" | jq -r '.[0].name'
      glab api "projects/coding-ai%2F$R/repository/files/.gitlab-ci.yml/raw?ref=main" | grep -n 'CI_COMMIT_TAG'
      ```
      
      **Read the whole file, not its head.** A release job commonly sits at the bottom,
      after the lint and validate jobs; concluding "this repo has no release job" from
      the first screenful is wrong, and in this sweep that mistake was made and acted
      on — it produced a plan to "fix" CI that was already correct.
      
      After every tag push, **verify the Release object exists rather than assuming the
      tag implied it.** If it is missing, the tag pipeline's job list says why — a rule
      that did not match shows up as the release job being absent from an otherwise
      green pipeline, not as a failure:
      
      ```bash
      glab api "projects/coding-ai%2F$R/pipelines?ref=$TAG" | jq -r '.[0].id'
      glab api "projects/coding-ai%2F$R/pipelines/<id>/jobs" | jq -r '.[]|"\(.stage)/\(.name): \(.status)"'
      ```
      
      The same trap exists on the GitHub side one level in: `node-agent-skill-coordinator`'s
      release workflow ran, but died on `Script not found "build"` because the caller
      never set `build-cmd` and the shared reusable defaults to `bun run build`. A
      caller that adopts a reusable and then does not cut a release has not tested it —
      the first tag after the adoption is the test.
      
      ## Release notes: the generated body is the default — do not curate it away
      
      Operator decision, 2026-08-27 (netresearch/skill-repo-skill#257, twice re-scoped
      and closed): GitHub's `generate_release_notes` body is the only release-notes
      mechanism with no maintenance debt — it has no requirements and just counts the
      merged PRs — so it stays the default, **including its `by @<login>` author
      credits** (they do not hurt; the no-bot-credit rule applies to *manually
      written* notes, where agent/bot credit must not be added). Do not rewrite
      generated bodies, do not strip credits, do not make releases depend on a
      CHANGELOG the repo does not maintain — in the 2026-08-27 sweep most repos had
      no CHANGELOG.md and five with one had an empty `[Unreleased]` at bump time.
      Curate a body only where it is otherwise a dead pointer: the GitLab
      `claude-code-skill` component writes a fixed "see CHANGELOG.md for details."
      description, which the GitLab fleet driver's finish replaces with the tag's
      CHANGELOG section or the plan row's body (gitlab-skill !88).
      
      A release job that dies on a transient network error (observed: cosign's
      `Post https://fulcio.sigstore.dev/…: connection reset by peer`) does not
      invalidate the tag — `gh run rerun <run-id> --failed` and re-verify the
      Release; never delete and re-push the tag for this.
      
      ## Immutable-Release Caveat
      
      Deleted GitHub releases do NOT free the tag for reuse. Once a release is published and deleted, that tag string is permanently locked as a deleted release — a new release with the same tag will fail. See `git-workflow-skill` → `references/github-releases.md`. Therefore: get it right the first time. The version-parity check above is what "right the first time" means in practice.
      
      ## Tag Signing (Mandatory)
      
      ```bash
      git tag -s vX.Y.Z -m "vX.Y.Z"         # -s: sign with GPG/SSH
      git push origin vX.Y.Z                # signed tag reaches the remote
      ```
      
      Never `git tag vX.Y.Z` (unsigned). Repos with protected tag rulesets will reject unsigned tags.
      
      ## No `--latest` Drift for Non-Default Branches
      
      When releasing from a non-default branch (e.g. a v1.x maintenance line while v2.x is default), pass `--latest=false` to avoid stealing the "Latest" badge by timestamp:
      
      ```bash
      gh release create v1.5.12 --latest=false --title "v1.5.12" --notes-file CHANGELOG-v1.5.12.md
      ```
      
      GitHub marks releases "Latest" by creation timestamp, not semver. A v1.5.12 created after v2.0.0 will become "Latest" without this flag — wrong, misleading, and often noticed only by downstream consumers.
      
      ## What a release archive contains
      
      `copy_skill()` in `.github/workflows/release.yml` packages a fixed **allow-list** per skill — `SKILL.md`, `references`, `scripts`, `assets`, `templates`, `examples`, plus `checkpoints.yaml` — and then copies the two root licence files into each skill directory:
      
      ```bash
      cp LICENSE-MIT LICENSE-CC-BY-SA-4.0 "$dst/"
      ```
      
      Two consequences worth knowing before touching files under `skills/<name>/`:
      
      - **A file outside that list never ships.** A per-skill `README.md`, a stray note, a `LICENSE` of its own — none of it reaches an archive, so removing one cannot change what consumers get. Answer "does deleting this break the release?" from the allow-list, not from intuition.
      - **Per-skill licence files are redundant by construction.** Every packaged skill already carries `LICENSE-MIT` and `LICENSE-CC-BY-SA-4.0` because the workflow puts them there. A `skills/<name>/LICENSE` adds nothing and rots silently: three such symlinks in `netresearch/matrix-skill` pointed at a root `LICENSE` that the split-licensing migration had deleted, and they stayed dangling for six months until a marketplace import named them.
      
      Adding a new packageable file type means extending the `for item in …` list in the reusable workflow — a repo cannot opt in from its own side.
      
      ## Supply-Chain Attestation
      
      Every release ships with provenance-attested archives and a Cosign-signed `SHA256SUMS.txt` that binds those archives by digest. Only the checksum file is Cosign-signed; the `.zip`/`.tar.gz` archives are integrity-protected through it — verifying the signature on `SHA256SUMS.txt` and then running `sha256sum --check` against the downloaded archive proves the archive was produced by this workflow.
      
      All of this happens in the SAME job that publishes the GitHub Release, BEFORE the assets are made public — there is no window where unsigned or unattested artifacts are downloadable. The flow, in order:
      
      1. Build `*.zip` and `*.tar.gz` archives.
      2. Generate `SHA256SUMS.txt` over them.
      3. **Cosign** keyless `sign-blob` the `SHA256SUMS.txt` → produces `SHA256SUMS.txt.sigstore.json` (Sigstore bundle format: cert + signature + Rekor inclusion proof in a single self-contained JSON; cosign v3+ default). The `.sigstore.json` extension is chosen so OSSF Scorecard's `signed-releases` probe recognises the signature — the content is identical to cosign's `--bundle` default output.
      4. **`actions/attest-build-provenance`** generates a SLSA build-provenance attestation for the archives + checksums file → published to GitHub's attestation API.
      5. **`softprops/action-gh-release`** publishes the GitHub Release with all assets attached at once.
      
      Callers must grant three permissions on the calling job:
      
      ```yaml
      # .github/workflows/release.yml in the consuming repo
      jobs:
        release:
          uses: netresearch/skill-repo-skill/.github/workflows/release.yml@main
          permissions:
            contents: write          # release upload
            id-token: write          # OIDC for sigstore (Cosign + attest-build-provenance)
            attestations: write      # GitHub native attestation API
      ```
      
      If any of those scopes is missing the job fails fast with `Resource not accessible by integration`; `contents: write` alone is not enough.
      
      ### What a release archive contains
      
      Every archive unpacks into ONE top-level folder named after the artefact:
      `<skill>/SKILL.md`, `<plugin>/.claude-plugin/`. That is not cosmetic. The OpenAI
      skills API accepts "a .zip that contains a single top-level folder", so a flat
      archive is rejected with a bare *Invalid skill* and no reason
      (netresearch/german-technical-writing-skill#2), and every repo's README tells
      readers to "extract to your agent's skills directory" — a flat archive unpacks
      `SKILL.md` and `references/` straight into `~/.claude/skills/`, on top of
      whatever else is there. `tests/release-archive-layout.sh` runs the packaging
      step out of the workflow against a fixture repository and asserts the property,
      so it cannot regress silently.
      
      ### Verify a downloaded release archive
      
      Both commands below pin verification to the **specific repository** that's expected to have produced the release. `--owner netresearch` and `https://github.com/netresearch/.*` are tempting shortcuts but match every workflow run in the org — meaning a compromised or unrelated netresearch repo could mint a valid-looking attestation against an artefact that was never released from this repo. Always pin to the named repo.
      
      ```bash
      # SLSA build provenance (GitHub-native attestation API)
      # Substitute <repo-name> with the actual skill repo, e.g. matrix-skill.
      # Archive name patterns: <skill>-skill-vX.Y.Z.zip and <plugin>-plugin-vX.Y.Z.zip.
      # --signer-workflow is required: the attestation is signed by the shared
      # reusable workflow, and --repo alone checks the signer against the consumer.
      gh attestation verify <skill-name>-skill-vX.Y.Z.zip --repo netresearch/<repo-name> \
        --signer-workflow netresearch/skill-repo-skill/.github/workflows/release.yml
      echo "rc=$?"   # success prints nothing when stdout is not a terminal
      
      # Cosign sign-blob signature on the checksums (no GitHub API needed).
      # The cert SAN reflects the SIGNER, which is the shared reusable release
      # workflow (`netresearch/skill-repo-skill`), NOT the consuming repo. Pin the
      # regex to skill-repo-skill, not the consumer, for the same reason
      # `gh attestation verify` above needs `--signer-workflow`. The org-wide form
      # `https://github.com/netresearch/.*` would accept signatures from any repo,
      # branch, or workflow in the org — too loose for supply-chain verification.
      cosign verify-blob \
        --bundle SHA256SUMS.txt.sigstore.json \
        --certificate-identity-regexp "^https://github\.com/netresearch/skill-repo-skill/\.github/workflows/release\.yml@" \
        --certificate-oidc-issuer "https://token.actions.githubusercontent.com" \
        SHA256SUMS.txt
      
      # Then verify the archive matches the (now-signed) checksum
      sha256sum --check SHA256SUMS.txt
      ```
      
      If verification fails:
      
      - `gh attestation verify` returns `Error: verifying with issuer "sigstore.dev"` when `--signer-workflow` is missing: `--repo` then also constrains the signer, and the signer is `skill-repo-skill`, not the consumer. Measured with gh 2.101.0 on github-release-skill v1.1.0 and v1.2.0; both pass with the flag, and a random file then fails with HTTP 404.
      - `gh attestation verify` returns `error: no attestations found` when `--repo` is wrong (or when the release predates this workflow).
      - `cosign verify-blob` returns `error: certificate identity does not match` when the regex is wrong, or `bundle verification failed` when `.sigstore.json` doesn't correspond to the file.
      
      ### Why one atomic job
      
      Splitting attestation into a separate `needs: release` job (the original design here) creates a race: the GitHub Release publishes BEFORE the attestation exists, so anyone downloading in that window gets unsigned, un-attested artifacts. Folding everything into the same job before the upload eliminates the window — either the whole bundle (archives + signature + provenance) ships, or nothing does.
      
      Same pattern as `netresearch/.github/.github/workflows/golib-create-release.yml` and `netresearch/typo3-ci-workflows/.github/workflows/release.yml`. No reason for skill repos to diverge.
      
      The previously-documented `with: attest: true` opt-in is gone; the input is still declared as `DEPRECATED — ignored` so any caller that still passes it doesn't error syntactically, but every release now gets provenance unconditionally. Drop the `with:` block if `attest` was its only entry (also true for `bump`).
      
    • repository-quality-rules.md 15.5 KB
      # Repository Quality Rules (skill repositories)
      
      ## Contents
      
      - Lizenz (Netresearch Split-Modell)
      - Mindestbestandteile eines Skill-Repos
      - Scripts-first (mechanisch prüfbare Regeln)
      - Pflicht für README-Oberfläche
      - SKILL.md vs. Discovery
      - allowed-tools (vorab gewährte Rechte)
      - Related Skills (Repo-Ebene)
      - Marketplace-Sync (Quelle bleibt Repo)
      - GitHub Repository SEO
      - GitHub Pages policy
      
      Prüfbare Regeln für **einzelne Skill-Repositories** (`netresearch/*-skill`).
      **Nicht** für das Marketplace-Repository — Discovery-Katalog- und SEO-Governance für den Hub liegen in **`netresearch/claude-code-marketplace`** (`AGENTS.md` dort).
      
      ---
      
      ## Lizenz (Netresearch Split-Modell)
      
      - **PASS**, wenn `LICENSE-MIT` und `LICENSE-CC-BY-SA-4.0` vorhanden sind und **keine** bare `LICENSE`-Datei existiert (siehe `validate-skill.sh` / Checkpoints).
      - **PASS für „LICENSE oder LICENSE.md“-Anforderungen außerhalb Netresearch:** Split-Lizenz gilt als erfüllte Lizenzpflicht; einzelne `LICENSE`-Datei ist hier **FAIL**, wenn sie das Split-Modell ersetzt.
      
      ---
      
      ## Mindestbestandteile eines Skill-Repos
      
      Jedes Repo **muss** die folgenden Elemente enthalten **oder** eine **explizite Begründung** in `README.md` unter z. B. `## Repository extras` (warum ein Pflichtobjekt fehlt).
      
      | Element | Prüfregel |
      | --- | --- |
      | `README.md` | Datei existiert (`validate-skill.sh`). |
      | Lizenz | `LICENSE-MIT` + `LICENSE-CC-BY-SA-4.0` (Policy). |
      | `CONTRIBUTING.md` | Datei existiert **oder** README verlinkt auf ein externes Contributing-Dokument und nennt den Ort (**ein** kanonischer Ort). |
      | `SECURITY.md` | Datei existiert **oder** README enthält Abschnitt „Security“ mit Kontakt/Ort der Policy. **Ausnahme:** klar als **private/internal-only** gekennzeichnete Repos — dann **muss** das im README stehen (`Private-only: no SECURITY.md`). |
      | `.github/pull_request_template.md` | Datei existiert **oder** Issue/PR-Richtlinie ist in `CONTRIBUTING.md` als PR-Checkliste beschrieben (mind. 5 konkrete Checkboxen). |
      | Skill-Verzeichnis mit `SKILL.md` | Pfad entspricht `.claude-plugin/plugin.json` → `skills`. |
      | `agents/openai.yaml` | Datei existiert **oder** Begründung + Alternative (z. B. „Agent Stack nicht OpenAI“) im README. |
      | `references/`, `scripts/`, `assets/` | **PASS**, wenn SKILL.md alle Referenzen erreichbar macht **oder** README erklärt bewusst schlankes Repo („no references: …“). |
      
      ---
      
      ## Scripts-first (mechanisch prüfbare Regeln)
      
      Generalisiert die Routing-Regel der retro-skill `destination-taxonomy` („mechanisch erkennbare Regel → `checkpoints.yaml`“) auf die Autorenseite: Was sich mechanisch prüfen lässt, wird nicht als Prosa ausgeliefert.
      
      - **PASS**, wenn jede mechanisch prüfbare Anforderung (Regex, Datei-Existenz, Kommando mit Exit-Code) als Eintrag in `checkpoints.yaml` **oder** als Script in `scripts/` vorliegt und der Fließtext nur mit **einer** Zeile darauf verweist.
      - **FAIL**, wenn eine mechanisch prüfbare Anforderung ausschließlich als Prosa-Anweisung existiert — der Agent muss sie dann bei jeder Anwendung neu interpretieren, und Abweichungen bleiben unentdeckt.
      - **Hinweis:** Checkpoint-Patterns unterliegen der Runner-Allowlist (`is_safe_eval_command`, automated-assessment `run-checkpoints.sh`): einzeilig, kein `bash` als Basiskommando, kein `;`/`&&`/`$()`; Repo-Scripts sind dort nicht aufrufbar — komplexe Prüfungen gehören nach `scripts/` und werden im Checkpoint als allowlist-konformes Pattern **gespiegelt, nicht aufgerufen** (Beispiel: Checkpoint SR-37 spiegelt `skills/skill-repo/scripts/check-version-parity.sh`).
      - **Hinweis:** Ein Checkpoint-Pattern wird über den Parser von `run-checkpoints.sh` geprüft, nicht über `yq -r`. Der Runner liest `checkpoints.yaml` zeilenweise mit Bash-Regex und dekodiert in Werten in doppelten Anführungszeichen nur `\"` und `\\`. Ein `\t` oder `\n` in doppelten Anführungszeichen und ein `''` in einfachen Anführungszeichen erreichen das Kommando deshalb unverändert, während `yq -r` sie auflöst. Ein Pattern, das über `yq -r` plus `bash <<<` besteht, kann im Runner deshalb bei jedem Lauf scheitern. Bevorzugt werden Skalare ohne Anführungszeichen und ohne Escape-Sequenzen, bei denen Rohzeile und dekodierter Wert gleich sind; die letzte Prüfung ist ein Lauf von `run-checkpoints.sh` gegen den echten Baum.
      - **FAIL**, wenn ein `llm_reviews`-Checkpoint einen ausführbaren Befehl am Zeilenanfang enthält. Ein Prompt, der eine Shell-Pipeline zeigt, beschreibt eine mechanische Prüfung in Prosa — sie läuft nie und kann nicht regressieren. Gehört der entscheidende Teil nach `mechanical`, bleibt im LLM-Eintrag nur die Bewertung, die das Kommando nicht leisten kann. Behält ein Eintrag bewusst beide Hälften, deklariert er das mit `# mechanical-counterpart: <ID>` im Block; `validate-skill.sh` prüft beides.
      - **FAIL**, wenn ein ausgeliefertes Script unter `scripts/` von keinem Test unter `tests/` referenziert wird. Skripte sind die ausführbare Oberfläche eines Skills, und ohne Test überlebt ein Defekt beliebig lange: eine Flottenmessung fand **276 Scripts in 27 von 33 Repos** ohne eine einzige Testdatei, darunter ein Verifier, der wegen `set -e` plus `((VAR++))` seit Jahren nach dem ersten Treffer abbrach und 11 von 12 Abschnitten übersprang. Ausgeführt werden die Tests vom Reusable `netresearch/skill-repo-skill/.github/workflows/tests.yml@main` — ein Test, den keine Pipeline startet, ist kein Gate.
      - Inhaltliche Bewertung, was überhaupt in einen Skill gehört: siehe Content value rubric in [`skill-quality.md`](skill-quality.md).
      
      ---
      
      ## Pflicht für README-Oberfläche
      
      Siehe [`readme-template.md`](readme-template.md) für die **exakten Überschriften** und [`skill-discovery-metadata.md`](skill-discovery-metadata.md) für YAML-Zusatzfelder außerhalb von `SKILL.md`.
      
      ---
      
      ## SKILL.md vs. Discovery
      
      - **`SKILL.md`**: Laufzeitverhalten, Trigger, Arbeitsablauf — siehe [`skill-quality.md`](skill-quality.md).
      - **Discovery / SEO / Marketplace-Felder**: README, `agents/openai.yaml`, optionale Metadatei(en), GitHub Description/Topics — **nicht** als zusätzliche YAML-Schlüssel im `SKILL.md`-Frontmatter für Katalogzwecke.
      
      ### Frontmatter (technische Grenze)
      
      - **Erforderlich:** `name`, `description` (mit Präfix `Use when…`).
      - **Verboten für Discovery:** eigene Schlüssel wie `slug`, `tags`, `category`, `keywords`, `seo_*` im Frontmatter.
      - **Optional** (Agent Skills / Validator): `license`, `compatibility`, `metadata`, `allowed-tools` — nur wenn technisch nötig; keine Marketing-/SEO-Felder dort. Zur Formulierung von `allowed-tools` siehe unten.
      
      ---
      
      ## allowed-tools (vorab gewährte Rechte)
      
      `allowed-tools` schaltet für den Zug, in dem der Skill aufgerufen wird, die Rückfrage ab. Das Feld beschränkt nichts: laut Doku bleibt jedes Tool aufrufbar, und für alles Nichtgenannte greifen weiterhin die normalen Berechtigungseinstellungen. Was in der Zeile steht, ist also nicht die Fähigkeit des Skills, sondern die Fläche, auf der er ohne Nachfrage arbeitet — und genau die liest ein Prüfer, der ein fremdes Repo bewertet.
      
      **Regel: Ein Skill, der eigene Skripte mitbringt, deklariert die Skripte, nicht den Interpreter.**
      
      ```yaml
      # ❌ deckt jedes bash-Kommando ab, das der Skill je absetzt
      allowed-tools: Bash(bash:*) Read Glob Grep
      
      # ✅ deckt das Skriptverzeichnis ab statt jeden bash-Aufruf
      allowed-tools: Bash(${CLAUDE_SKILL_DIR}/scripts/*) Bash(bash ${CLAUDE_SKILL_DIR}/scripts/*) Read Glob Grep
      ```
      
      Beide Formen gehören in die Zeile: ein Muster greift über das Präfix, und `./foo.sh` und `bash foo.sh` haben verschiedene. Wer Python-Skripte mitliefert, braucht `Bash(python3 ${CLAUDE_SKILL_DIR}/scripts/*)` zusätzlich — oder ruft sie direkt auf, was Ausführungsbit und Shebang voraussetzt.
      
      Das Muster begrenzt auf ein Präfix, nicht auf eine Dateimenge: `*` schließt `/` ein, ein Aufruf mit `..` im Pfad liegt also formal noch darin. Wer eine harte Grenze braucht, zählt die Skripte einzeln auf — das kostet einen Eintrag je Skript und muss beim nächsten neuen Skript nachgezogen werden. Für die meisten Skills ist das Verzeichnismuster der richtige Tausch, solange man es als das liest, was es ist.
      
      Claude Code ersetzt `${CLAUDE_SKILL_DIR}`, `${CLAUDE_PROJECT_DIR}` und in Plugin-Skills `${CLAUDE_PLUGIN_ROOT}` an zwei Stellen: im Text der `SKILL.md` und in den Bash-Regeln des Frontmatters. Deshalb funktioniert das Muster über Installationsarten und Versionsstände hinweg, ohne einen Pfad festzuschreiben. `${CLAUDE_SKILL_DIR}` ist die breitere Wahl, weil sie auch außerhalb einer Plugin-Installation gesetzt ist.
      
      Drei Fallstricke:
      
      - **Muster und Aufruf müssen zusammenpassen.** `bash x.sh` und `./x.sh` sind verschiedene Präfixe. Wenn die `SKILL.md` relative Aufrufe dokumentiert (`scripts/foo.sh PATH`), trifft eine absolute Regel sie nicht — dann kommt bei jedem Skriptaufruf eine Rückfrage, die der Nutzer wegklickt. Skill-Text und Regel gehören in denselben Commit.
      - **Werkzeuge, die der Agent selbst absetzt, bleiben separat.** Verifiziert der Skill Ergebnisse mit `git`, `jq` oder `grep`, gehören die weiter einzeln in die Zeile. Die Regel betrifft den Interpreter-Platzhalter, nicht jede Bash-Regel.
      - **Eine Shell-Variable im Skill-Text hebelt die Regel aus.** Ein Block, der mit `S=${CLAUDE_SKILL_DIR}/scripts` beginnt und dann `uv run $S/foo.py` schreibt, läuft in der Shell korrekt — die Berechtigungsprüfung sieht aber das Kommando, wie es abgesetzt wird, also das wörtliche `$S/…`. Kein einziger so geschriebener Aufruf trifft ein Muster auf den Vollpfad. Im `matrix-communication`-Skill betraf das alle 20 dokumentierten Aufrufe. Den Pfad ausschreiben; das Zeilenbudget gibt es her (dort 61 von 500 Zeilen).
      
      Nach dem Umstellen einmal im Standardmodus durchlaufen, mit unveränderten Einstellungen — nicht unter `--dangerously-skip-permissions`, und nicht mit einer `permissions.allow`-Regel, die dasselbe Kommando ohnehin freigibt. Sonst beweist keines der beiden Ergebnisse etwas: eine ausbleibende Rückfrage kann von einer anderen Freigabe kommen, und eine Rückfrage kann aus einer `ask`-Regel stammen, die `allowed-tools` unabhängig vom Muster sticht. Eine `deny`-Regel wiederum blockiert ohne jede Rückfrage, und ein zusammengesetztes Kommando braucht ohnehin für jeden Teil einen Treffer. Im Zweifel das tatsächlich abgefragte Kommando gegen die geltenden Regeln halten, bevor das Muster geändert wird.
      
      ---
      
      ## Related Skills (Repo-Ebene)
      
      - Im README oder in Discovery-YAML **als Slugs oder volle URLs** angeben.
      - **PASS**, wenn mindestens ein Eintrag **oder** die Zeile `Related skills: none (justified — …)` mit Grund vorhanden ist.
      - **FAIL**, wenn beliebige Links nur für SEO gesetzt sind (nicht fachlich nachvollziehbar).
      
      ---
      
      ## Marketplace-Sync (Quelle bleibt Repo)
      
      Bei Änderungen an Discovery-Inhalten: siehe [`marketplace-integration.md`](marketplace-integration.md). Agents **müssen** am Ende einer Änderung prüfen, ob Marketplace-Felder aktualisiert werden müssen (oder Override dokumentiert ist).
      
      ---
      
      ## GitHub Repository SEO
      
      ### Repository / skill name
      
      - **PASS:** Leads with the **rankable proper noun** of the tool/domain (e.g. `jujutsu`, `oro`, `vite`) — not a generic short alias that is crowded in search (e.g. `jj`).
      - **FAIL:** Spends name tokens on **redundant words** — `agent`, `agentic`, or `ai` in the repository name, or `skill` in the skill slug. The repository's `-skill` suffix and the marketplace already imply "agent skill"; **every** skill is for agents, so these add no discriminating signal.
      - **PASS:** The remaining token names the **distinctive function** (`-workflow`, `-conformance`, `-upgrade`, …) so the slug isn't a bare proper noun colliding with a sibling skill's `name`.
      - **PASS:** Name candidates validated against GitHub search results — `gh search repos "<candidate phrase>"` — preferring an uncontested, descriptive phrase over a crowded generic one.
      - **PASS:** Short aliases or command names (e.g. `jj`) are kept in the **description and topics** rather than spent on the slug, so command-searchers still match.
      
      - **Good:** repo `jujutsu-workflow-skill`, skill `jujutsu-workflow` (rankable noun + function; `jj` lives in description/topics).
      - **Bad:** `jj-agent-workflow-skill` (`jj` is generic/crowded; `agent` is redundant for a skill).
      
      ### Repository Description
      
      - **PASS:** String length **≤ 160** characters (count before save).
      - **PASS:** Mentions **concrete technology** (e.g. TYPO3, OroCommerce, Docker) **or** a **named task domain** (e.g. “extension PHPUnit matrix”, “Vite sitepackage build”).
      - **FAIL:** Generic phrases such as “Useful AI skill for developers”, “ultimate automation assistant”.
      - **PASS:** Understandable **without** opening `README.md`.
      
      - **Good:** `Agent skill for TYPO3 Vite setup, SCSS architecture and frontend asset integration.`
      - **Bad:** `Useful AI skill for developers.`
      
      ### GitHub Topics
      
      - **PASS:** Includes **`agent-skill`**.
      - **PASS:** At least **one** stack tag (`typo3`, `php`, `docker`, …) or domain tag (`testing`, `security`, `frontend`, …) matching the skill.
      - **FAIL:** Irrelevant trending tags just for visibility (keyword stuffing).
      - Document proposed topics in README under `## Repository extras` if maintainers cannot edit Topics immediately.
      
      ---
      
      ## GitHub Pages policy
      
      The [marketplace](https://github.com/netresearch/claude-code-marketplace) is the canonical public discovery and storytelling layer for all Netresearch Agent Skills. Repository Pages are **secondary, skill-specific documentation surfaces**.
      
      ### Default: Pages disabled
      
      Skill repositories **must not** enable GitHub Pages by default.
      
      - **PASS:** `gh api repos/netresearch/<repo>/pages` returns **HTTP 404** (Pages disabled).
      - **FAIL:** Pages is enabled without satisfying the criteria below.
      
      ### When Pages is appropriate
      
      Enable GitHub Pages only when the repository contains standalone public material that is too large, too visual, too navigational, or too strategically important to live well in `README.md`. **At least one** of the following must be true:
      
      - the documentation requires multiple pages,
      - the skill has a gallery of examples, reports, dashboards, screenshots or demos,
      - the skill publishes generated reference documentation,
      - the skill provides versioned documentation,
      - the skill is a public reference implementation,
      - the skill explains a reusable methodology or assessment model,
      - the skill has a specific SEO target that the marketplace landing cannot cover without becoming too broad.
      
      ### When Pages is NOT appropriate
      
      Do not enable Pages if the site would only duplicate:
      
      - the README,
      - the marketplace detail page (`https://github.com/netresearch/claude-code-marketplace#<slug>` or the future landing),
      - installation instructions,
      - `SKILL.md`,
      - the basic example prompts.
      
      ### Mandatory artefacts when Pages is enabled
      
      If Pages is enabled, the repository **must** include:
      
      - a short justification block in `README.md` (which criterion above is satisfied),
      - a documented canonical URL pointing at the Pages site,
      - a clear source path (default: `docs/`),
      - documented build and deployment commands (`make docs`, `npm run docs`, or equivalent — referenced from the README),
      - link-checking or equivalent validation in CI,
      - a note explaining which content belongs on Pages vs. README vs. marketplace.
      
      ### Mirroring rule
      
      Skill-specific metadata originates in the skill repository (`metadata/discovery.yaml`, README sections, `agents/openai.yaml`, GitHub settings). The marketplace consumes it. Do not duplicate the same metadata across README, Pages site and marketplace landing — pick one canonical surface per fact and link from the others.
      
    • review-replies.md 6.5 KB
      # Reviewer-Reply Boilerplate
      
      Canonical responses to recurring reviewer comments on Netresearch skill PRs (Copilot, Gemini Code Assist, peer review). Lift the fenced block verbatim or paraphrase to context. Each entry includes the criteria for whether to accept or decline.
      
      These were extracted from review patterns across 14+ skill PRs. Use them to keep responses consistent and to avoid re-litigating settled architectural decisions on every new PR.
      
      ## 1. "Set `private: true`"
      
      **Verdict:** Accept — already the template default.
      
      **Criteria:** Skill packages distributed via `github:org/repo` (not the npm registry) **must** set `"private": true` to guard against accidental `npm publish` of the placeholder `0.0.0-source` version. The current `package.json.template` bakes this in. If a reviewer flags it on a fresh PR, the package was scaffolded from an outdated template — accept the suggestion and add it.
      
      ```markdown
      Accepted. Adding `"private": true` — current scaffolding template (`skills/skill-repo/templates/package.json.template` in `netresearch/skill-repo-skill`) bakes this in to guard against accidental publish of the `0.0.0-source` placeholder. This PR was scaffolded before the template was updated.
      ```
      
      ## 2. "Pin `github:org/repo#vX.Y.Z` in install instructions"
      
      **Verdict:** Decline as primary advice. Document `#vX.Y.Z` as an opt-in for users who want reproducibility.
      
      **Criteria:** These skills are markdown content (procedural knowledge), not executable code where pinning protects against breakage. Consumers want skill-content updates by default. Pinning is an advanced opt-in.
      
      ```markdown
      Declined as primary advice. The default `github:netresearch/{repo-name}` form intentionally tracks the default branch so consumers receive skill-content updates — these skills are markdown content (procedural knowledge), not executable code where breakage matters. Pinning is an advanced use-case, not a default. Consumers can append `#vX.Y.Z` themselves for reproducibility (`github:netresearch/{repo-name}#v1.2.3`); we don't surface that in the README to keep the install path simple.
      ```
      
      (Reference: this is the response we used on [skill-repo-skill PR #82](https://github.com/netresearch/skill-repo-skill/pull/82).)
      
      ## 3. "Drop `.claude-plugin/` (or `commands/`, `outputStyles/`, `AGENTS.md`) from `files`"
      
      **Verdict:** Decline (won't-fix). This is the dual-distribution invariant.
      
      **Criteria:** Netresearch skill packages are *dual-distributed* — the same package serves both the Claude Code marketplace install path AND the npm install path. Plugin metadata, slash commands, output styles, and `AGENTS.md` are part of the skill's installable surface, not internal repo configuration. Excluding them from the npm tarball delivers a partial skill to npm consumers.
      
      ```markdown
      Declining. Netresearch skill packages are *dual-distributed* — the same tarball feeds both the Claude Code marketplace install path AND the npm install path. Plugin metadata (`.claude-plugin/plugin.json`), slash commands (`commands/`), output styles (`outputStyles/`), and the canonical `AGENTS.md` rules file are part of the skill's installable surface, not internal repo configuration.
      
      Excluding them from the npm tarball would deliver a partial skill to npm consumers — the same partial-install problem documented in [github-release-skill PR #19](https://github.com/netresearch/github-release-skill/pull/19)'s `> **Limitation:**` callout. The current `@netresearch/agent-skill-coordinator` (v0.1.x) `node_modules` scanner can't load plugin-mechanism features; those need Claude Code's plugin loader. Shipping these directories preserves the option to switch install methods without re-installing.
      ```
      
      (Reference: lifted from [skill-repo-skill PR #83](https://github.com/netresearch/skill-repo-skill/pull/83).)
      
      ## 4. "Top-level `scripts/` shouldn't ship"
      
      **Verdict:** **Accept** if `scripts/` is repo-maintenance only. **Decline** if installed code reads from `$ROOT/scripts/` at runtime.
      
      **Criteria:** This is the *opposite* call from #3. Top-level `scripts/` is **not** part of the dual-distribution surface unless the skill's runtime explicitly reads from it. To tell them apart:
      
      - **Repo-maintenance** (DO accept the suggestion, remove from `files`): scripts that only run in CI or by repo maintainers — `verify-harness.sh`, `generate-dashboard.sh`, `run-ab-evals.sh`, lint runners, release helpers. Look for invocation only in `.github/workflows/` or `Makefile` / `package.json scripts`.
      - **Runtime** (DO decline the suggestion, keep in `files`): scripts the *installed* skill executes, typically referenced from `skills/<name>/SKILL.md` or `skills/<name>/scripts/*.sh` via `$ROOT/scripts/...` or `../scripts/...`.
      
      Quick check: `grep -r '\$ROOT/scripts\|\.\./scripts' skills/`. If empty, it's repo-maintenance.
      
      ```markdown
      Accepted. `scripts/` at the repo root only contains `verify-harness.sh` (repo-maintenance, run via `.github/workflows/`). The installed skill code does not read from `$ROOT/scripts/` at runtime — runtime scripts live under `skills/<name>/scripts/` (already covered by the `skills/<name>/` entry). Removed from `files`. The npm-pack-smoke CI job will keep this honest going forward.
      ```
      
      If declining (runtime usage):
      
      ```markdown
      Declining. Top-level `scripts/` is consumed at runtime — `skills/<name>/SKILL.md` references `$ROOT/scripts/<file>.sh` for [specific feature]. Removing it from `files` would break npm consumers. The npm-pack-smoke CI job asserts this dir is present.
      ```
      
      ## 5. "`AGENTS.md` shouldn't ship"
      
      **Verdict:** Decline (won't-fix). Same dual-distribution invariant as #3.
      
      **Criteria:** `AGENTS.md` is the canonical agent rules entry point for the skill repo. npm consumers expect the same rules file marketplace consumers see — without it, agents reading the package don't get the harness contract.
      
      ```markdown
      Declining. `AGENTS.md` is the canonical agent rules entry point for the skill — npm consumers must receive the same rules file that marketplace consumers do, otherwise agents reading the installed package miss the harness contract documented there. This is part of the dual-distribution surface (see [skill-repo-skill PR #83](https://github.com/netresearch/skill-repo-skill/pull/83)).
      ```
      
      ## See Also
      
      - `installation-methods.md` — Method 4 (npm) for the full `files`-allowlist rationale.
      - `release-discipline.md` — version-parity check, multi-repo dry-run.
      - [skill-repo-skill PR #83](https://github.com/netresearch/skill-repo-skill/pull/83) — original dual-distribution decision record.
      
    • skill-architecture.md 20.4 KB
      # Skill Architecture: flat discovery, small routing metadata, details on demand
      
      ## Contents
      
      - The three surfaces
      - Budgets, and where each number comes from
      - The description is a router, not documentation
        - What a description can move, measured
      - What a reference reaches, measured
      - What gets executed, measured
      - Flat discovery: one level, always
      - What belongs in SKILL.md and what does not
      - Long references need a Contents section
      - What the validator enforces
      - Testing whether a skill triggers
      - Sources
      
      ## The three surfaces
      
      A skill is read in three stages, at three different moments. Putting a fact on the wrong one is why a capability that exists still does not get used.
      
      | Surface | When it enters context | Size | What a gap there costs |
      |---|---|---|---|
      | `name` + `description` | **always**, at startup, for every installed skill | ~100 tokens | **The only true skip.** The skill is never activated, so nothing else is ever read. |
      | `SKILL.md` body | when the skill is activated | < 5000 tokens recommended | A blind spot: the skill runs and does not know the thing exists. |
      | `references/`, `scripts/`, `assets/` | only when the agent decides to reach for one | unbounded | Nothing for a lookup table. For a rule that prevents a mistake: measured at four of six runs never opening one — see below. |
      
      > *"Metadata (~100 tokens): The `name` and `description` fields are loaded at startup for all skills. Instructions (< 5000 tokens recommended): The full `SKILL.md` body is loaded when the skill is activated. Resources (as needed): Files … are loaded only when required."*
      
      The practical consequence: **a missing capability in the description cannot be recovered later.** A missing capability in the body can at least be found by an agent that reads a reference. A missing reference costs nothing until something needs it.
      
      ## Budgets, and where each number comes from
      
      | Element | Limit | Kind | Target |
      |---|---|---|---|
      | `name` | 64 chars, `[a-z0-9-]`, no leading/trailing hyphen, no `--`, must match the directory name | **spec, hard** | 15–40 |
      | `description` | **1024 chars** | **spec, hard** | see below |
      | `compatibility` | 500 chars | **spec, hard** | one sentence, usually omit |
      | `SKILL.md` body | **< 500 lines**, < 5000 tokens | spec recommendation | 100–250 |
      | reference file | none | — | add a Contents list past 100 lines |
      
      Two traps worth naming:
      
      **Lines, not words.** The recommendation is *"Keep your main `SKILL.md` under 500 lines"*. A word budget is a different and far tighter constraint. This repository's validator counted 500 **words over the whole file** until 2026-08; a consuming repo sat at 499 of 500 and left a script out of its `SKILL.md` rather than spend fourteen words on it.
      
      **Frontmatter is not body.** Counting them together makes the description — the surface that decides whether the skill is used at all — compete with the instructions for one allowance. That trade is always the wrong way round.
      
      ## The description is a router, not documentation
      
      The description carries the entire burden of triggering. It is not a summary of the workflow.
      
      Write **what + when**, and stop:
      
      ```yaml
      description: "Use when formatting or validating text for Jira — descriptions, comments, wiki markup, or converting Markdown to Jira syntax."
      ```
      
      Not:
      
      ```yaml
      description: >
        Handles Jira content by first detecting Markdown, then converting it with
        md2jira.sh, validating the result, checking special tables, escaping
        characters and finally producing Jira wiki markup…
      ```
      
      The second wastes routing context, and it can actively harm: a description containing a condensed workflow invites the agent to follow **that** summary instead of activating the skill and reading the real instructions.
      
      On length, two pulls exist and both are real. The official guidance says *"A few sentences to a short paragraph"* and also *"Err on the side of being pushy. Explicitly list contexts where the skill applies"* — that argues for covering the scope properly. Against it: every description competes for a shared listing budget, and an over-broad one triggers when it should not.
      
      So: **1024 is a validation limit, not a target.** Past roughly 500 characters, check whether what you added is trigger information or process narration. Cut the narration; keep the contexts.
      
      ### What a description can move, measured
      
      A small model often does not open the skill at all. In two recorded evaluation
      series — same case, same fleet, one arm per description — the trials that
      *passed* made no `Skill` call in any run: nothing read the body, the references,
      or the file that had listed the required artefacts all along. Only the
      description reached the agent.
      
      What it moved, and what it did not:
      
      Each row is one arm carrying the new description against one carrying the old
      one, on the same case and fleet. Trials per arm differ by round and are stated
      in the table. "Passed" is the case's own mechanical check accepting the result;
      "opened" is an explicit `Skill` call.
      
      | the description named | trials per arm | passed, with → without | opened, with → without |
      |---|---|---|---|
      | the occasion the skill is for | 3 | wrote to the directory that renders, 3 → 0 | 0 → 0 |
      | every file that must carry the release version, sentence first | 6 | 4 → 0 (wrote the missing file 5 → 0) | 0 → 0 |
      | the same, as released, sentence second | 6 | 6 → 1 | 1 → 0 |
      | the file a newer convention replaced | 3 | 0 → 0, across three wordings | 0 → 0 |
      | a procedure: reproduce the report as a failing test before fixing it | 3 | 0 → 0 | 0 → 0 |
      
      The two release-version rows are the same idea measured twice. The first is an
      experiment branch whose description opened on the sentence naming the files;
      the second is what shipped, where that sentence sits second behind the trigger
      list. The effect survives the move, which is worth recording because a positive
      result on a branch is not a result about the release. The two are not
      distinguishable at six trials per arm, and nothing here says the released
      wording is better.
      
      So the routing rule has a second half. A description decides *whether* the skill
      is reached, and where it is not reached it is the only part of the skill that
      reaches the agent at all — the prompt, the system instructions, the tools and
      the model's own habits are still there and still decide plenty. Two consequences
      for writing one:
      
      - **A fact the agent lacks belongs in the description** — the artefacts it must
        produce, the place they belong. Not the steps: those are still narration, and
        narration still invites the agent to follow the summary instead of the skill.
      - **A description hands over a noun, not a procedure.** That last row is the
        clearest measurement of the boundary. Naming four files got the missing one
        written by agents that never opened the skill; naming the occasion — a user
        reports wrong output — did not get a test written first, and did not even get
        the skill opened, although the description had been rewritten for exactly that
        request shape. An artefact can be handed over in a sentence because it is a
        thing the agent can go and produce. A way of working cannot, because following
        it means already being inside the skill.
      - **A description cannot overturn what the model already believes.** Three
        attempts at "this file replaced that one" changed nothing. Where the skill has
        to correct a convention rather than supply a missing fact, the body is the only
        place that can do it — and the body only works when the skill is opened.
      
      Put the three together and they name what to do with a rule about *how* to work:
      it belongs in the body, near the top, and it is worth nothing until something
      opens the skill. Getting it opened is a separate problem from writing it, and
      the description is the only lever on that problem.
      
      ## What a reference reaches, measured
      
      The row above says a gap in `references/` costs "nothing — until it is needed".
      That holds for a lookup table. It does not hold for a rule that prevents a
      common mistake, and the difference is measurable.
      
      `OFR-TYPO3-UPGRADE-001`, 14 September 2026, Haiku 4.5, six trials on one fleet.
      This is a case where routing works: `skill_invoked` is 6 of 6, so every trial
      opened the skill and read the body. Counting `Read` calls against
      `references/` per trial:
      
      | trial | reference files read | outcome |
      |---|---|---|
      | 1 | 0 | passed both legs |
      | 2 | **0** | passed v14.3, lost v13.4 |
      | 3 | 3 | failed both legs |
      | 4 | 2 | passed both legs |
      | 5 | 0 | passed both legs |
      | 6 | 0 | passed both legs |
      
      Four of six opened no reference file at all. Trial 2 lost the leg that had been
      working to a rule that had been sitting in `references/upgrade-v13-to-v14.md`
      for nine days, written from an earlier occurrence of the same failure — and
      `SKILL.md` names the problem at exactly the right step and then points at the
      file: "`createMock` on one of them cannot be repaired by swapping the name — see
      `references/upgrade-v13-to-v14.md`". The agent was told a rule exists and not
      what it says.
      
      Reading is not what separates the outcomes here — trial 3 read three files and
      failed both legs, trials 5 and 6 read none and passed. Six trials say nothing
      about that either way. What they do say is the frequency: a reference is opened
      in a minority of runs even when the body points at it, so a sentence that has to
      land every time cannot live there.
      
      The rule that follows is about placement, not about length:
      
      - **A reference is for what an agent will look up once it knows it needs it** —
        a mapping, a schema, the fiftieth edge case, the two honest answers to a
        judgement call.
      - **A sentence that prevents a mistake belongs in the body**, even when the
        surrounding treatment stays in the reference. The body is read whenever the
        skill is activated; the reference is read when the agent decides to.
      
      
      ## What gets executed, measured
      
      The two measured sections above are about what reaches the agent — a
      description, a reference. This one is about what the agent then *does*, and it
      is the one the stack's purpose turns on: a skill is a shortcut only where the
      step it names gets run.
      
      These are not this repository's answer-text evals. Every count below is read
      from tool-call telemetry: the `verifier/trajectory.json` Harbor records for each
      trial in `netresearch/agent-system-evals`, where a Skill invocation is a `Skill`
      tool call and "ran it" means a `Bash` tool call whose arguments carry the step's
      own signature — the block's `types='…'` prefix, the runner's script name, the
      grep's alternation. Model `claude-haiku-4-5-20251001` under Claude Code, one
      prompt, one cold start, one session per trial. Records:
      `experiments/OFR-TYPO3-UPGRADE-001-20260918-{073044,092803,121328,152227}.json`
      (rounds twenty-four to twenty-seven in that case's `RESULTS.md`) and
      `experiments/OFR-TYPO3-EXT-001-20260828-121312.json`. p-values are Fisher's
      exact test, two-sided, from that repository's `scripts/lib/stats.py`; a cost is
      the agent's `final_metrics.total_cost_usd` for the trial in USD.
      
      One case, one model, one position. `OFR-TYPO3-UPGRADE-001` under Haiku 4.5,
      `Skill` invoked in 12 of 12 trials in every round below, and the same step of
      the same body carrying four shapes in turn, each measured on three to six
      trials against the one it replaced:
      
      | step 9 carried | trials that ran it | note |
      |---|---|---|
      | a rule, in the body | in context 3/3; the edit it names occurred 0/12 | untested on its own axis — nothing to prevent |
      | an instruction to run a script, via a variable the loader does not set | 0/3 | in context 3/3; no trial in twelve bound a variable of any kind |
      | the bare `grep` the script wraps, beside either | 2/6 | |
      | the check itself as a fenced block that runs as pasted | **5/6** | p 0.048 against the instruction; the sixth typed the grep by hand |
      | the body's other fenced block, step 10, for scale | 16/18 | p 1.000 against the block above |
      
      The line between the shapes is not position — step 10 sits below step 9 — and
      not length. It is whether the step can be executed as written. A command the
      agent already knows (`rector`, `phpstan`, `composer`) is run. A fenced block
      with nothing to look up and nothing to bind first is run at the same rate. A
      path that has to be assembled from the loader's output is not run at all:
      across twelve trials the agents used the printed absolute skill path verbatim,
      in five `ls`/`grep`/`find` calls against `references/`, and never once as a
      variable.
      
      A second case says the same thing on a different skill. `OFR-TYPO3-EXT-001`,
      six-trial round with cost declared, `typo3-conformance` in the fleet: the
      fenced grep block in that body ran in six of six equipped trials and its
      equivalent in none of six bare ones, at $0.13 against $0.32, with the same
      task outcome on both arms despite the different block-execution rates. That is
      one row and not a mechanism — the equipped arm carries eight skills and a
      workflow besides the block — but it is the shape being run, again. Measured
      since, by removing only that block from the body (record
      `experiments/OFR-TYPO3-EXT-001-20260918-170620.json`): cost did not rise, the
      arm without the block ran equivalent greps by hand from the tokens the body's
      numbered steps name in prose, and outcome held. The block is run; it is not
      where that case's saving lives.
      
      Two consequences for writing a body, and one limit.
      
      - **A step that must happen is a block that runs as pasted.** Not a sentence
        saying to do it, not a path to a script that does it. The runner in
        `automated-assessment` is reached by `/assess <skill>`, one hop; the
        sibling-path form of the same call measured 0/3 in a body and stays
        documented as that.
      - **A block that finds nothing costs a call and saves none.** On the upgrade
        case the edit the block catches arrives once in forty-two trials, and cost
        overlapped in three declared rounds. The saving the stack is for appears
        where a check fires and replaces the test run — or the ten tool calls — that
        would have found the same thing by hand. Measure it there.
      - **This is one model at one budget, and the other model measured has the
        opposite row.** Same case, same skill (`typo3-extension-upgrade`, v3.11.1
        through v3.12.5, every one of them naming `scripts/scan-deprecations.sh` in
        the body), one prompt, one cold start. `claude-opus-5` ran that script by its
        printed path in 16 of 17 valid trials, on 20 August 2026 at $20–34 a trial.
        `claude-haiku-4-5-20251001` ran it in 1 of 188, 31 August to 18 September,
        under $2 a trial. Six Opus trials of 19 August are left out as invalid: two
        steps each and no cost. The rows above are Haiku's. Whether a path
        instruction is followed is a property of the model reading it, not of the
        body, and the stack is measured on the model that does not follow it because
        that is where a shortcut has something to shorten.
      
      ## Flat discovery: one level, always
      
      > *"Keep file references one level deep from `SKILL.md`. Avoid deeply nested reference chains."*
      
      The reason is mechanical. `SKILL.md` is read in full on activation. A reference is read only if `SKILL.md` said what it contains and when to open it. A file reachable only through a second hop sits behind a door with no sign on it — nothing states what it holds or why it matters, so the agent must open the middle file speculatively and then guess again.
      
      Good:
      
      ```
      jira-communication/
      ├── SKILL.md
      ├── references/
      │   ├── jira-syntax.md
      │   ├── jira-style.md
      │   └── issue-fields.md
      └── scripts/
          ├── md2jira.sh
          └── validate-jira.sh
      ```
      
      with every one of them named in `SKILL.md`:
      
      ```markdown
      ## Resources
      
      - Jira wiki markup rules: `references/jira-syntax.md`
      - Wording and issue conventions: `references/jira-style.md`
      - Convert Markdown to Jira markup: run `scripts/md2jira.sh`
      - Validate generated markup: run `scripts/validate-jira.sh`
      ```
      
      `jira-syntax.md` may of course also say "run `scripts/md2jira.sh`". That is **redundancy for orientation**, and it is welcome. What it must not be is the *only* path by which the agent learns the script exists.
      
      Bad, and specifically bad:
      
      - `SKILL.md` → `scripts.md` → `scripts/foo.sh` — an index file that only points at scripts adds indirection with no information. Delete it and give the scripts a real `--help`.
      - `SKILL.md` → `jira-style.md` → `md2jql.sh` **as the only path** — the script is invisible unless that one reference happens to be opened.
      
      Split by topic at the **first** level rather than nesting: `topic-a.md` and `topic-b.md`, both listed, each with its own trigger.
      
      ## What belongs in SKILL.md and what does not
      
      `SKILL.md` is a **control plane**, not a handbook.
      
      In it:
      
      1. Invariants that must never be forgotten.
      2. Decisions — "if X, read/run Y".
      3. Workflow order, where order matters.
      4. Guardrails.
      5. The resource map: every reference and every executable, each with a trigger.
      6. Verification — how the agent knows it is done.
      
      Not in it: API documentation, syntax references, long examples, mappings, lookup tables, historical explanations, CLI references, schemas, the fiftieth edge case. Those go to `references/`.
      
      And anything deterministic — conversion, parsing, validation, AST manipulation, formatting, mechanical checks — belongs in `scripts/`. Scripts are **executed, never loaded**, so their body costs no context at all. A script named in `SKILL.md` costs one line and buys the only chance the agent has of knowing it exists.
      
      ## Long references need a Contents section
      
      Agents preview long files — `head`, an excerpt, a targeted search — rather than reading them whole. A contents list at the top makes the rest of the file visible anyway. Past about 100 lines, add one.
      
      For very large references (upwards of ~10k words), go further and tell the agent in `SKILL.md` what to search for, not just which file to open.
      
      ## What the validator enforces
      
      `scripts/validate-skill.sh` checks the mechanically decidable part:
      
      | Check | Level |
      |---|---|
      | `description` > 1024 chars | ERROR (spec) |
      | `description` > 500 chars | WARN |
      | `description` missing / not `Use when …` | ERROR |
      | `name` invalid, leading/trailing or doubled hyphen | ERROR (spec) |
      | `compatibility` > 500 chars | ERROR (spec) |
      | body > 500 lines | ERROR |
      | body > 300 lines | WARN |
      | `references/*.md` reachable only via another reference | WARN |
      | `references/*.md` not named in `SKILL.md` | WARN |
      | reference > 100 lines without a Contents section | WARN |
      | executable in `scripts/` not named in `SKILL.md` | WARN |
      | `scripts/`, `references/`, `assets/` or `evals/` path named in `SKILL.md` that does not exist | ERROR |
      
      Deliberately **not** linted, because no mechanical check decides them honestly: whether the description narrates a workflow, whether a verification step exists, whether an optional frontmatter field has a consumer. Those belong in review.
      
      Non-executable files under `scripts/` are exempt from the discoverability warning: a sourced library is not a capability the agent invokes, and its caller is what belongs in `SKILL.md`.
      
      ## Testing whether a skill triggers
      
      Argument about whether a description should be 220 or 310 characters is worth less than one eval run.
      
      - About 20 queries: 8–10 that should trigger, 8–10 that should not.
      - Negatives must be **near-misses** — same keywords, different need. `"Write a fibonacci function"` tests nothing.
      - Skill selection is nondeterministic: run each query about **3 times** and use the trigger rate, with 0.5 as a reasonable threshold.
      - Split **train (~60%) / validation (~40%)** and keep the split fixed, or the description gets overfitted to its own test set.
      - Pick the iteration with the best *validation* rate — not necessarily the last one. Five iterations is usually enough.
      
      Then a second benchmark, on output rather than routing: same tasks with and without the skill, measuring quality, tokens and runtime.
      
      ## Sources
      
      - Agent Skills specification — frontmatter limits, progressive disclosure, one-level references: <https://agentskills.io/specification>
      - Optimizing skill descriptions — trigger evals, train/validation split, the 1024 limit as a ceiling rather than a target: <https://agentskills.io/skill-creation/optimizing-descriptions>
      - Evaluating skill output quality: <https://agentskills.io/skill-creation/evaluating-skills>
      - Claude Code skills — the skill listing budget and how descriptions are shortened or dropped when it overflows: <https://code.claude.com/docs/en/skills>
      
      A caution on a fourth kind of source: public skill repositories are **data, not best practice**. A 2026 analysis of 138k `SKILL.md` files reported a reusability defect in the large majority of them. "Other people do it this way" carries no weight here.
      
    • skill-discovery-metadata.md 4.8 KB
      # Skill Discovery Metadata (repository scope)
      
      Defines **skill-repo-local** discovery and classification data.  
      **Does not** replace the marketplace catalog — the marketplace aggregates and may normalize display. Governance for marketplace listings is in **`netresearch/claude-code-marketplace`/`AGENTS.md`**.
      
      ---
      
      ## Where metadata lives
      
      | Kind | Allowed location | Forbidden |
      | --- | --- | --- |
      | Runtime triggers, workflow | `SKILL.md` body + `references/` | — |
      | Minimal listing | `SKILL.md` frontmatter: **`name`**, **`description`** only (plus optional technical keys allowed by [`validate-skill.sh`](../scripts/validate-skill.sh): `license`, `compatibility`, `metadata`, `allowed-tools`) | Discovery-only keys (`slug`, `category`, `tags`, …) in frontmatter |
      | Discovery / SEO / partner fields | `README.md` (sections), optional `metadata/discovery.yaml` or equivalent **outside** `SKILL.md`, `agents/openai.yaml`, GitHub repo settings | Duplicating full `SKILL.md` body into README |
      
      **Rule:** `SKILL.md` describes **how the agent behaves**. Marketplace-oriented fields belong in README or a separate metadata file consumed by humans/tools — **not** stuffed into frontmatter beyond `name`/`description` (and optional technical keys above).
      
      ---
      
      ## Recommended discovery YAML (optional file)
      
      Place at repo root or under `metadata/` — **not** inside `SKILL.md`. Example schema:
      
      ```yaml
      slug: typo3-vite
      display_name: TYPO3 Vite Frontend Pipeline
      summary:
        en: >
          Configures Vite for TYPO3 v13+ with vite-asset-collector, SCSS entrypoints, and CSP-safe asset URLs.
        de: >
          Richtet Vite für TYPO3 v13+ mit vite-asset-collector, SCSS-Entrypoints und CSP-tauglichen Asset-URLs ein.
      category: typo3-frontend
      tags:
        - vite
        - typo3
        - scss
        - frontend
      use_cases:
        - Bootstrap a Vite pipeline for a TYPO3 sitepackage.
        - Split CSS/JS entrypoints per content element.
      expected_outputs:
        - Vite config and npm scripts aligned with vite-asset-collector.
        - Documented build and deployment steps for frontend assets.
      context_requirements:
        - TYPO3 v13+ sitepackage or extension with asset collector available.
        - Node.js LTS matching project policy.
      action_level: modifies_files
      risk_level: medium
      related_skills:
        - typo3-frontend-patterns
        - dxp-frontend
      example_prompts:
        - "Add Vite 7 to our TYPO3 13 sitepackage with SCSS partials per content element."
        - "Wire vite-asset-collector entrypoints for tt_content templates."
      primary_keywords:
        - vite
        - typo3
        - sitepackage
      ```
      
      Fields:
      
      - **`slug`**: stable id; usually matches plugin/skill name.
      - **`display_name`**: public title (may differ from `name` in `SKILL.md`).
      - **`summary.en` / `summary.de`**: one short paragraph each; `de` optional unless DACH/TYPO3/Oro focus. **`summary.en` should be ≤ 300 chars** (snippet-friendly target; the marketplace enforces a hard cap of 500).
      - **`category`**: **one of** the canonical marketplace categories — `development`, `devops`, `security`, `design`, `workflow`, `productivity`. Keep this in sync with the marketplace `AGENTS.md` canonical list; do not invent ad-hoc values.
      - **`tags` / `use_cases`**: for README tables and marketplace sync.
      - **`expected_outputs` / `context_requirements`**: must mirror README sections (single source: generate README from this file or maintain parity explicitly).
      - **`related_skills`**: slugs or full GitHub URLs; see [`repository-quality-rules.md`](repository-quality-rules.md). Entries that don't yet exist in the catalog can be tagged `(planned)` or `(external)` — never invent fake links for SEO.
      - **`example_prompts`**: align with README **Example prompts** (≥3 in README per checklist).
      - **`primary_keywords`**: align with GitHub Topics + first sentence of summaries.
      
      ---
      
      ## Action level
      
      | Value | Definition |
      | --- | --- |
      | `read_only` | Reads/analyses repo or docs only; no writes. |
      | `suggests_changes` | Proposes patches/text but does not apply them. |
      | `modifies_files` | Writes or edits files in the working tree. |
      | `runs_commands` | Executes local shell/commands (build, tests, linters). |
      | `external_write` | Creates/updates data in external systems (GitHub API, Jira, Matrix, email, deployment APIs, …). |
      
      ## Risk level
      
      | Value | Definition |
      | --- | --- |
      | `low` | No durable impact or easily reversible edits. |
      | `medium` | Local file changes and/or command execution with repo impact. |
      | `high` | External writes, destructive operations, releases/deployments, security-sensitive changes. |
      
      **PASS rule:** every skill repo **should** state `action_level` and `risk_level` in discovery YAML **or** in README table „Classification“ with the same labels.
      
      ---
      
      ## Sync with marketplace
      
      When this metadata changes, open or update the corresponding marketplace entry per [`marketplace-integration.md`](marketplace-integration.md). **Do not** silently diverge.
      
    • skill-quality.md 21.4 KB
      # SKILL.md Quality Rules — Detail and Examples
      
      ## Contents
      
      - Why these rules exist
      - Content value rubric
      - Eval evidence
      - Description rules
      - Body rules
      - Reference patterns
      - Authoring checks before commit
      - Auditing
      - Sources
      
      Detailed guidance backing the summary in [`SKILL.md`](../SKILL.md) (§ SKILL.md Quality Rules).
      
      ## Why these rules exist
      
      Claude Code skills have two distinct context costs:
      
      - **Listing cost (per turn):** Only `name`, `description`, and optional `when_to_use` from each skill's frontmatter enter context on every assistant turn — whether the skill is invoked or not. Per-skill cap: 1,536 chars (`skillListingMaxDescChars`). Total listing cap: `skillListingBudgetFraction × context_window` (default `0.01`).
      - **Body cost (per invocation):** The full SKILL.md body loads when the skill is invoked, and persists for the rest of the session. Reference files (`references/*.md`) are **not** auto-loaded — the model only reads files it sees referenced.
      
      Description bytes are the always-on tax; body bytes are the on-demand tax. References are free unless the model knows they exist and decides to read them.
      
      ## Content value rubric
      
      Size rules bound how MUCH a skill costs; this rubric bounds WHAT KIND of content earns that cost. Every passage in a SKILL.md body or reference file must provide at least one of six value categories:
      
      1. **Org/project-specific knowledge** — Netresearch conventions, internal tooling, URLs, field IDs, policy decisions. (`jira` custom-field IDs; the split-licensing model.)
      2. **Version/ecosystem facts models get wrong** — post-training-cutoff changes, version-specific breaking changes, niche tool flags. (TYPO3 v14 `#108055` asset-concat removal; golangci-lint v2 config.)
      3. **Retro-born failure patterns** — symptom → cause → required behavior → verification, encoded from a real incident. (`github-project`: the auto-approve/Copilot race.)
      4. **Executable scripts/validators** — deterministic work shipped in `scripts/` or `checkpoints.yaml` instead of prose the model re-derives.
      5. **Inference suppression** — "read file X, never guess Y" rules that shut down a known guessing pattern and its retry cost. (`typo3-ddev`: never guess backend URLs — read `.ddev/config.yaml`.)
      6. **Anti-rationalization guards** — rules that stop the model from skipping steps or over-claiming. ("No 'tested' claim without pasted command output.")
      
      **Generic bloat** is content that provides none of these: restated public best practice a one-line prompt regenerates ("write tests first", "use parameterized SQL", tutorials paraphrased from public docs). The test, from Den Odell's essay (see Sources): *if you can generate the passage with a prompt, the skill doesn't need to carry it.*
      
      Two protections when applying the rubric:
      
      - Categories 3 and 5 are **first-class even when they read as prose**. Reducing wrong or unnecessary inference counts as value although the model "knows" the underlying facts — a skill that surfaces the right rule at the right moment saves the retry tokens the guess would have cost. Do not cut a failure pattern or a never-guess rule because it "looks like advice".
      - A claimed activation/recall benefit ("the model knows this but forgets") is an eval claim: back it with an A/B delta (repo-root `scripts/run-ab-evals.sh`) rather than asserting it. See "Eval evidence" below for what that delta does and does not say.
      
      ## Eval evidence
      
      Whether an eval can carry evidence at all — every eval needs at least one
      assertion a no-skill baseline fails — is specified in
      [`materialization-contract.md`](materialization-contract.md) ("Quality rule:
      every eval needs a delta-discriminating assertion"), together with the
      `samples` self-check. What follows is the other half: how to report a number
      once you have one.
      
      ### A delta belongs to an actor, not to a skill
      
      ```text
      delta = f(skill version, model, agent harness, task, tool environment)
      ```
      
      So `docker-development has +23%` is not a statement anyone can check. The
      reportable form names the tuple:
      
      ```text
      docker-development@2.4.0 · claude-code 2.0.31 · model sonnet
      · eval-set evals.json@09cc52ef (24 evals) → +23% over the no-skill baseline
      ```
      
      Every run prints that line and stores the same fields under `provenance` in
      `ab-results.json`, so a published number can be traced to the text, the actor
      and the questions that produced it. Quote them together — in a PR that
      justifies skill content, in the dashboard, in a release note.
      
      Two consequences worth stating, because they cut in opposite directions:
      
      - **A stronger actor can erase a delta.** If a later model derives the same
        content unaided, the skill stops buying behavior and only costs context.
        That is a reason to re-measure on the current actor before defending a
        passage, not a reason to keep the old number.
      - **A weaker delta is not always a weaker skill.** A harness that is bad at
        skill discovery scores the skill poorly although the skill is fine. The
        per-skill A/B eval assumes the skill is already loaded, so it cannot see
        that failure at all — that is what the system-level evals in
        [`netresearch/agent-system-evals`](https://github.com/netresearch/agent-system-evals)
        measure, by giving the agent a realistic underspecified request that names
        no skill, tool or method.
      
      ### Where the gate lives, and where it must not
      
      `run-ab-evals.sh --samples=3 --require-delta` is a **local, pre-merge**
      instrument: run it in the repo you are editing, before opening the PR that adds
      the eval. It calls the model `2 × N` times per eval, so it cannot run on every
      pull request, and it is deliberately **not** wired into the scheduled sweep
      either. The sweep's job is
      to publish numbers; a gate there would abort a skill's merge step, and a skill
      with one weak eval would quietly stop being refreshed on the dashboard while
      looking merely unchanged. The sweep therefore *reports* the evidence-free evals
      in the PR body it opens and leaves the decision to a person.
      
      What CI can decide without a model is the `samples` self-check: it proves that
      the assertions accept a correct answer and reject a wrong one, which is the
      mechanical half of "this eval discriminates". Prefer adding `samples` over
      adding a CI job that spends tokens.
      
      ### Sample each arm more than once
      
      A single completion per arm is not a verdict. Measured on this repo's own
      suite, two consecutive runs of the *same* eval text:
      
      | eval | run 1 | run 2 |
      |---|---|---|
      | `migrate_single_license` | 2 discriminating checks | 0 |
      | `release_workflow_setup` | 1 | 0 |
      
      Nothing about those evals changed between the runs. `--samples=N` therefore
      runs N completions per arm and decides every assertion by **majority over the
      samples**; `must_not` inverts each sample before the vote, and samples that
      carry no answer are dropped rather than counted as failures. The sweep uses
      3 by default, and `ab-results.json` records `samples_per_arm` so a quoted
      number always says how firm it is.
      
      `--require-delta` refuses to run below `--samples=2` — gating on a coin flip
      is the failure mode the flag exists to prevent, one level up.
      
      **Sampling reduces the noise; it does not remove it.** Two consecutive
      three-sample runs of this repo's own suite, nothing changed in between, still
      disagreed about four evals — all of them at the 0↔1 boundary, where a single
      marginal assertion decides the verdict. The aggregate held (24 vs 26
      discriminating checks, 8 vs 6 evidence-free); the membership of the list did
      not.
      
      So the two flags have different jobs:
      
      | flag | what it is for |
      |---|---|
      | `--require-delta` | a **reading** signal. Read a failure as "these evals were marginal in this run", open the two arms, decide. Do not block a merge on it. |
      | `--min-evidence-ratio=P` | the **gate**. Fails unless at least P percent of the *measured* evals carry a discriminating check — the quantity that held still across both runs. |
      
      Pick the floor from a measurement of the suite rather than from taste, and
      leave headroom for the boundary cases: this suite measured 62% and 71% on two
      consecutive runs, so a floor of 50% holds while a floor of 70% would fail
      every other time — which would be the same coin flip one level up.
      
      The cost is linear and real: `2 × N` completions per eval, so the default
      triples what a single-sample sweep costs. Use `--samples=1` for a cheap
      indicative run whose *aggregate* is still informative — a suite where half the
      evals never discriminate is telling you something either way — but do not
      quote a per-eval verdict from it.
      
      ### Read the arms before believing the delta
      
      A measured delta is only as good as the two answers behind it, and the first
      three runs of this instrument were each contaminated in a way the totals did
      not show:
      
      - **The operator's MCP servers were still attached.** `--tools ""` drops the
        built-in tools and nothing else, so answers on a workstation with Gmail and
        Drive servers configured discussed which tools existed — unequally between
        the arms. Fixed by `--strict-mcp-config` with an empty `--mcp-config`.
      - **The model still expected tools.** With them gone, an action-shaped prompt
        ("create a plugin.json") was answered with "I'll check the repo first" and a
        tool call print mode cannot serve: three output tokens, nothing to grade, and
        a *negative* delta on two evals because the with-skill arm spent its answer
        explaining the convention. Fixed by telling both arms up front that there are
        no tools.
      - **A call that never happened was graded as a wrong answer.** The CLI writes
        a session-limit notice or an API error into the same file an answer would go
        to, so an eval that was never measured scored 0 in both arms and appeared in
        the evidence-free list. Such a pair is now reported as **not measured** and
        excluded from the totals — unknown, never refuted — and `--require-delta`
        skips them while failing outright if *nothing* could be measured.
      
      All three were found by opening `scripts/ab-results/<eval>_with.txt` after a run
      that looked plausible in aggregate. When a delta is negative, or an eval reports
      zero discriminating checks, read the two files before concluding anything about
      the skill — the instrument is the more likely defect.
      
      ### Fact and trigger ownership
      
      Each fact and each trigger phrase has ONE canonical owning skill in the catalog; every other skill cross-references instead of restating. Duplicated facts drift apart at the next release, and duplicated trigger phrases compete for the model's skill selection. (Example: TYPO3 v14 breaking-change facts are owned by `typo3-conformance`; upgrade-execution phrasing by `typo3-extension-upgrade`; siblings link, they do not repeat.) This is the content-level twin of the discovery-level Mirroring rule in [`repository-quality-rules.md`](repository-quality-rules.md) — one canonical surface per fact, links from everywhere else.
      
      ## Description rules
      
      ### Caps
      
      - **Hard cap: 1,536 chars** — Claude Code truncates above this regardless of budget.
      - **Target: 100–300 chars** — fits comfortably in the listing, leaves headroom for other skills.
      - **Justify above ~500 chars** — long descriptions are reasonable when they enumerate triggers (e.g., ticket-key prefixes, file-extension lists). Keep the structure tight.
      
      ### Position matters
      
      Truncation is position-based. Put your primary trigger first. The convention is `Use when <trigger>` as the opener.
      
      ### Anti-patterns
      
      | Anti-pattern | Example |
      |---|---|
      | Marketing language | "blazingly fast", "powerful", "comprehensive" |
      | Vagueness | "a tool to help with X", "general-purpose helper" |
      | Restating the skill name | `description: "The frobnicator skill frobnicates."` |
      | Redundancy with body | duplicating the skill's intro paragraph in both fields |
      
      ### Examples
      
      **Good:**
      
      ```yaml
      description: "Use when reviewing your diff, writing a commit message, or asking what changed. Summarizes uncommitted changes and flags risky patterns."
      ```
      
      **Bad:**
      
      ```yaml
      description: "A comprehensive solution for managing your repository state with advanced semantic understanding"
      ```
      
      ## Body rules
      
      ### Size
      
      Body loads on invocation and persists for the rest of the session. The threshold goal is **per-invocation token cost** — every additional word in the body is paid every time the skill is invoked. Move what isn't strictly needed at decision time into `references/` (lazy-loaded when SKILL.md instructs).
      
      Word count translates to tokens at roughly 1.4× (English prose; code fences run higher). `audit-skills.sh` reports three tiers:
      
      - **INFO above 500 words** (~700 tokens) — modest cost; skim for split candidates.
      - **WARN above 1,000 words** (~1,400 tokens) — almost any body at this size has lookup content that could move to references.
      - **FAIL above 2,000 words** (~2,800 tokens) — body is acting as a manual; split required.
      
      These tiers are advisory across the wider skill ecosystem. **This repo enforces a stricter 500-word hard cap on its own SKILL.md** (via `skills/skill-repo/scripts/validate-skill.sh`) — that is the per-skill house rule, not a universal claim. Treat the audit tiers as triage levels for any skill repo; treat the 500-word cap as policy for the canonical skill-repo template. Several Netresearch skill repos (git-workflow, github-project, github-release, jira, …) ship the same `validate-skill.sh` hard cap, so treat any body ≥ ~490 words as hard-capped, not soft.
      
      **Adding to a body already near the cap:** measure first (`wc -w` / `scripts/audit-skills.sh`), then land the addition in a single budgeted pass — put the detail in `references/` and add only a minimal pointer (one Critical Rule line or one reference-table row) to the body, and trim an equal number of words from the fattest existing lines (verbose reference-table descriptions are the usual give) in the same edit. Iterating one or two words at a time across repeated re-validations wastes cycles. Gotcha: `wc -w` counts a spaced slash (`a / b`) as three tokens, so slash-separating during a trim can *raise* the count — use commas (`a, b`).
      
      Empirical anchor: the Netresearch skill corpus (n=52) has p95 ≈ 994 words. WARN at 1,000 catches the actual outliers in our own work. Anthropic's longer skills (e.g., `writing-skills` at 3,193 words) FAIL under this rule — that's the honest signal: even shipped skills can have content moved out. We can't fix theirs; we hold ourselves to the rule.
      
      - **Bloat signals**: body >3 KB, body >150 lines, single fenced code block >25 lines (long code examples are the primary lazy-load target).
      
      ### When to split to `references/`
      
      Move to `references/<topic>.md` when content is:
      
      - Long examples (>20 lines)
      - Configuration templates
      - API/syntax tables
      - Migration matrices
      - Detailed checklists
      - Repository directory trees (often the worst offenders)
      
      Keep in SKILL.md body:
      
      - High-level workflow
      - The "what does this skill do" summary
      - A reference catalog (see Pattern 2 below) when there are many topic-files
      - One canonical quick example
      
      ## Reference patterns
      
      Every `.md` file in `references/` must be discoverable to the model **from SKILL.md**. Three acceptable patterns:
      
      ### Pattern 1: Direct cite
      
      ```markdown
      For form validation, see [`references/forms.md`](references/forms.md).
      For multi-profile auth, see [`references/multi-profile.md`](references/multi-profile.md).
      ```
      
      Best for **≤10 reference files**. Each gets a one-line cite with the path. The model sees the path on invocation and reads on demand.
      
      ### Pattern 2: Catalog-with-convention
      
      ```markdown
      ## Reference Files (in `references/`, `.md` implied)
      
      - **Frameworks** (`*-security`): react, vue, angular, nextjs, nuxt
      - **Languages** (`*-security-features`): php, python, go, rust, javascript-typescript
      - **Cloud** (`*-security`): aws, azure, gcp
      ```
      
      Best for **10+ topic-grouped files**. The model sees:
      
      - The directory (`references/`)
      - The filename suffix convention (`-security`, `-security-features`)
      - The list of stems
      
      It can resolve any combination on demand. The `security-audit` skill uses this pattern for ~60 reference files in ~12 lines of body content.
      
      ### Pattern 3: List-and-pick
      
      ```markdown
      For language-specific guidance, list `references/` and read the file matching your stack.
      ```
      
      Use sparingly — burns a tool call (the model has to enumerate the directory). Useful when there's no naming convention but the directory is small.
      
      ### Anti-pattern: Hub-and-spoke
      
      ```markdown
      # DON'T DO THIS
      
      In SKILL.md:
      > See [references/index.md](references/index.md) for the catalog of references.
      
      In references/index.md:
      > - [forms.md](forms.md)
      > - [auth.md](auth.md)
      ```
      
      Multi-hop traversal isn't reliable. Anthropic's docs phrase the rule as "Reference supporting files from SKILL.md so Claude knows what each file contains" — direct visibility from SKILL.md is the model. The hub gets read but its leaves aren't deeply traversed. Anthropic's published skills use flat / catalog patterns, never hub-and-spoke.
      
      ### Anti-pattern: Orphan refs
      
      Files in `references/` with no path to discovery from SKILL.md (no direct cite, no catalog mention, no list-and-pick instruction). The model never reads them. Either cite them or delete them — they consume disk and signal "stale" to readers without contributing to the skill's behavior.
      
      ## Authoring checks before commit
      
      Four checks for any change to SKILL.md or a reference file. Each one targets a class of finding that reviewers otherwise catch one PR at a time.
      
      ### Spell in US English
      
      Skill prose uses US spelling: `-ize`/`-ization`, `behavior`, `color`, `authorization`, `serialized`. File names in the skill repos already follow it (security-audit-skill's `error-message-sanitization.md`), so a British spelling in the prose splits every search in two — `grep sanitisation` misses the rest of the skill, `grep sanitization` misses the new text. When unsure, count the repo (`git grep -i 'behavior' | wc -l` against `git grep -i 'behaviour' | wc -l`) and follow the majority; a repo whose norm is British stays British. Code identifiers, API names and quoted output are not prose and keep their spelling.
      
      ### Examples obey the rules of their own document
      
      When a document states a rule in prose and also shows code, every example must follow that rule. An example that contradicts its own page teaches the wrong rule, and the reader copies the example, not the sentence. Typical shapes: a page that says "never use strippable `assert()` for security checks" and then uses `assert()` in its PHP and Python samples; a helper defined as `Assertf` and called as `assertf`. Before committing, extract the rules from the prose as a checklist and read each example against it line by line. If a rule cannot be shown without breaking it, split the example or restate the rule.
      
      ### A cross-reference is a claim
      
      "See X below", "as the Y section covers", "§ Z" all assert that the target exists in this document. The pointer is easiest to get wrong exactly when the fact is familiar: it is real, but it lives in another file, another repo or your own notes. Grep the target before writing the pointer (`grep -n '<heading or phrase>' <file>`), then sweep the diff for pointers as a class before committing:
      
      ```bash
      git diff HEAD | grep -E '^\+' | grep -niE 'see |below|above|section|§'
      ```
      
      Confirm every hit. A pointer that one grep cannot confirm gets deleted — a section that stands on its own beats one that leans on a neighbor that is not there.
      
      ### Extending an enumerated set touches every surface that lists it
      
      Skills often enumerate a set of codes or concepts (pillars `G1`–`G3`, checks `R1`–`R6`, checkpoint IDs) in several files at once: the SKILL.md table, including the body cells of its rows, a lifecycle or workflow reference, an output template and its worked examples, a vocabulary reference. Adding a member to one of them leaves the others describing the old set, and a reader who follows any of those surfaces never learns the new member exists. Before opening the PR, pick one existing member as the anchor and grep the whole skill for it:
      
      ```bash
      grep -rlw 'R1' skills/<name>/
      ```
      
      Read every hit and decide whether the new member belongs there too. The same sweep applies in reverse when a member, option or helper is removed: `git grep -n '<name>'` across the repo finds the README feature lists, example configs and comments that still advertise it.
      
      ## Auditing
      
      Run `scripts/audit-skills.sh` from the repo root. It scans SKILL.md files for:
      
      - Description length violations (warn >500 chars, fail >1,536)
      - Body length tiers (info >500 words, warn >1,000, fail >2,000)
      - Long code fences (info >25 lines — primary lazy-load candidate)
      - Orphan refs (files in `references/` not reachable by any pattern)
      
      Pattern 2 detection is heuristic: a reference file counts as P2 when its stem matches a convention stated in SKILL.md (filename suffix or topic-list pattern). Files outside any P1/P2/P3 path are reported as ORPHAN.
      
      ## Sources
      
      - Anthropic Claude Code docs on skills: <https://code.claude.com/docs/en/skills>
      - Anthropic published skills: <https://github.com/anthropics/skills>
      - Settings schema: `skillListingBudgetFraction`, `skillListingMaxDescChars` (Claude Code settings.json)
      - Den Odell, "The Great Agent Skills Land Grab": <https://denodell.com/blog/the-great-agent-skills-land-grab> — the generate-it-with-a-prompt test and the eval-evidence bar for activation/recall claims
      - Xiong et al., "How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior", ACL 2026: <https://aclanthology.org/2026.acl-long.27/> — inaccurate past experience compounds into later runs, and executions that look correct can still be misleading as experience; the paper's proposal is to use later task outcomes as quality labels, which is what `/retro outcome` does
      
    • skill-retirement.md 4 KB
      # Skill Retirement
      
      How to decommission a skill that is superseded (e.g. by automation) or obsolete.
      Order matters — each step keeps the trail auditable and every action reversible.
      
      ## Workflow
      
      1. **Verify the successor.** Confirm the replacing process or tooling is live
         (ticket closed, automation running) before removing anything.
      
      2. **Remove the marketplace entry.** MR/PR against the marketplace repo:
         delete the plugin entry from `marketplace.json` **and** the README skill
         table. Run the marketplace's own validator before committing
         (`make validate` in the Netresearch internal marketplace).
      
      3. **Add a deprecation notice** at the top of the skill repo's README — name
         the successor and the tracking ticket. Push this **before** archiving:
         archived repos are read-only.
      
         ```markdown
         > **⚠️ DEPRECATED / ARCHIVED (<month year>)**
         >
         > Superseded by <successor> (<ticket>). This repository is archived.
         ```
      
      4. **Archive the repository — never delete.** Archiving is reversible and
         preserves history, tags, and MRs.
      
         ```bash
         glab repo archive https://git.netresearch.de/GROUP/PROJECT   # GitLab
         gh repo archive OWNER/REPO                                   # GitHub
         ```
      
      5. **Uninstall locally and verify propagation.**
      
         ```bash
         claude plugin uninstall <name>@<marketplace>
         claude plugin marketplace update <marketplace>   # refresh cache
         grep -c "<name>" <marketplace-cache>/.claude-plugin/marketplace.json  # expect 0
         ```
      
      6. **Document in the tracking ticket.** What was removed, MR links, archive
         status — so the ticket history is self-contained.
      
      ## Relocation: the skill moves to another repository
      
      When a skill is not retired but moves — typically into the repository whose
      scripts it wraps, so a drifting copy disappears — the plugin keeps its name
      and only the marketplace entry's `source.repo` changes. Done for `cli-tools`,
      moved from `cli-tools-skill` into `coding_agent_cli_toolset` (2026-09):
      
      1. **Port first, then move.** Fixes that exist only in the old repository
         go to the new one before anything is repointed; otherwise the move
         silently regresses them.
      2. **Ship the plugin from the new repository** (`.claude-plugin/plugin.json`
         with the same `name`, skill under `skills/<name>/`) and verify it loaded
         live before touching the marketplace — `claude -p` needs a prompt, e.g.
         `claude -p --plugin-dir <checkout> "List the skills you have loaded"`; for
         a hook, see [plugin-hooks](plugin-hooks.md).
      3. **Decide the version source.** Without `version` in `plugin.json` *and* in
         the marketplace entry, Claude Code versions the plugin by the source's
         commit SHA, so every merge to the default branch becomes available on the
         user's next plugin update (see below for what that takes). That fits
         a repository with no release flow; a fixed `version` would freeze users
         until someone bumps it. Such a repository does not belong in the release
         fleet list either — say so there, or the next refresh re-adds it.
      4. **Repoint the marketplace entry** (and the README row); every consumer
         that links into the old repository is updated in the same sweep.
      5. **Moved notice, then archive**, as in steps 3–4 above.
      
      **How existing installations switch over (measured).** Third-party
      marketplaces do not auto-update by default, and `plugin update` alone reads
      the cached catalog. Users need both:
      
      ```bash
      claude plugin marketplace update <marketplace>
      claude plugin update <name>@<marketplace>   # 1.9.2 -> <commit sha> for cli-tools
      ```
      
      Name both commands in the moved notice.
      
      ## Gotchas
      
      - **Installed plugins outlive the marketplace entry.** Removing the entry does
        not uninstall existing local installations — they keep working from the
        plugin cache. Announce the retirement (team channel) so users uninstall.
      - **README edits after archiving require unarchiving first.** Get the wording
        right before step 4.
      - **Do not delete the repo or its tags.** Released versions may still be
        referenced by lockfiles, Satis indexes, or npm/composer installs.
      
    • validation-checklist.md 12.7 KB
      # Validation checklist — skill repository changes
      
      Agents **must** walk through this list before declaring a skill-repo task complete.
      Marketplace-only checks live in **`netresearch/claude-code-marketplace`/`AGENTS.md`** — do not merge those steps here.
      
      ## Contents
      
      - README
      - SKILL.md
      - Manifests
      - Agents / OpenAI
      - Discovery metadata
      - GitHub repository SEO
      - GitHub Pages
      - Related skills
      - Marketplace sync expectations
      - Links and automation
      - Optional extended validation
      - Which validator catches what
      - claude.ai Organization settings sync
      - Third-party install scanners
      
      ## README
      
      - [ ] `README.md` contains **all** required sections listed in [`readme-template.md`](readme-template.md).
      - [ ] **Example prompts:** ≥ **3** realistic prompts in `## Example prompts`.
      - [ ] **Related skills:** declared **or** `none (justified: …)`.
      - [ ] First screen answers: problem, when, outputs, context, installation — verifiable without scrolling past ~1 screen (approx. first 40 lines).
      
      ## SKILL.md
      
      - [ ] Frontmatter includes **`name`** and **`description`**; **no** discovery-only keys (`slug`, `tags`, `category`, `keywords`, …).
      - [ ] Optional keys only if needed: `license`, `compatibility`, `metadata`, `allowed-tools` (per validator).
      - [ ] `description` starts with `Use when`.
      - [ ] Body describes triggers and use cases without duplicating full README marketing copy.
      
      ## Manifests
      
      - [ ] Root **`plugin.json`** exists and targets `https://agent-plugins.org/schemas/1.0.0/plugin.schema.json`.
      - [ ] It carries **only** the fields of the closed schema — no `skills`, `agents`, … — see [`agent-plugins-compat.md`](agent-plugins-compat.md).
      - [ ] Neither manifest carries `support`: it is not a Claude Code field either. `claude plugin validate --strict .` exits **0**.
      - [ ] `.claude-plugin/plugin.json` regenerated: `bash skills/skill-repo/scripts/sync-plugin-manifest.sh` leaves the tree clean (`--check` exits 0).
      - [ ] Version bumped in the **root** `plugin.json` (source of truth), then synced; `check-version-parity.sh` exits 0.
      - [ ] Every skill lives at `skills/<name>/SKILL.md` — a root `SKILL.md` is invisible to Agent Plugins clients.
      
      ## Agents / OpenAI
      
      - [ ] `agents/openai.yaml` exists **or** README documents exception.
      - [ ] File contains a **short, user-understandable** description of the skill (what/when).
      
      ## Discovery metadata
      
      - [ ] Optional `metadata/discovery.yaml` (or documented equivalent) matches README if used — see [`skill-discovery-metadata.md`](skill-discovery-metadata.md).
      - [ ] **`action_level`** and **`risk_level`** present in discovery YAML **or** README classification table.
      
      ## GitHub repository SEO
      
      - [ ] Repository **Description** ≤ **160** characters, names concrete tech or use case — not generic “AI assistant”.
      - [ ] **Topics** include `agent-skill` plus relevant tech/domain tags; no irrelevant stuffing.
      
      ## GitHub Pages
      
      - [ ] `gh api repos/netresearch/<repo>/pages` returns **HTTP 404** (Pages disabled — the default).
      - [ ] **If Pages is enabled:** the PR description names which `repository-quality-rules.md` Pages criterion is satisfied, and the README contains the mandatory artefacts (justification, canonical URL, source path, build/deploy commands, link-checking, content-split note).
      
      ## Related skills
      
      - [ ] Related skills list is **honest** (exists, planned, or external) — see [`repository-quality-rules.md`](repository-quality-rules.md).
      
      ## Marketplace sync expectations
      
      - [ ] README contains note or checkbox: when discovery fields change, marketplace entry must be updated **or** override documented (see [`marketplace-integration.md`](marketplace-integration.md)).
      
      ## Links and automation
      
      - [ ] All links in touched docs resolve (internal paths and GitHub URLs).
      - [ ] `bash skills/skill-repo/scripts/validate-skill.sh` exits **0** from repo root.
      - [ ] If repo has CI calling `netresearch/skill-repo-skill/.github/workflows/validate.yml`, PR checks are green.
      
      ## Optional extended validation
      
      - [ ] `scripts/audit-skills.sh` (if present in repo) reports no new orphan `references/` files.
      
      ## Which validator catches what
      
      Three checkers run over a skill repository and none of them is a superset of the others. A green run of one is not a release gate for the others.
      
      | Checker | Catches | Does **not** catch |
      |---|---|---|
      | `validate-skill.sh` | repo structure, required files, root `LICENSE-MIT` + `LICENSE-CC-BY-SA-4.0`, shared-field parity between `plugin.json` and `.claude-plugin/plugin.json` (`version` among them) | anything the Claude Code manifest schema defines; `SKILL.md` version metadata, tag parity and the `composer.json` version rule, which belong to `check-version-parity.sh` |
      | `claude plugin validate [--strict]` | unknown top-level manifest fields (`--strict` turns the warning into exit 1), malformed manifest | dangling symlinks, markup inside a `SKILL.md` description |
      | claude.ai marketplace import | unknown manifest fields, **symlinks whose target is not in the repository**, **angle brackets in a `SKILL.md` description** (reported as XML tags), **a `SKILL.md` description over 1024 characters**, **a top-level `bin/`** — see [claude.ai Organization settings sync](#claudeai-organization-settings-sync) | — |
      | Hermes `skills_guard.py` (install-time scanner) | regex matches in the skill directory that block a community install — see [Third-party install scanners](#third-party-install-scanners) | what the matched text does; a pattern matches prose as readily as code |
      
      The gap is not theoretical: on 2026-09-17 three `skills/*/LICENSE` symlinks in `netresearch/matrix-skill` had pointed at a file deleted six months earlier, and a description in `netresearch/orocommerce-skill` carried a literal `<Secret:>` placeholder. Both passed `validate-skill.sh` and `claude plugin validate --strict`; only the marketplace import named them. Run the import, or check symlinks and descriptions by hand, before assuming a repository is clean:
      
      ```bash
      git ls-files -s | awk '$1=="120000" {print $4}' | while read -r l; do [ -e "$l" ] || echo "DANGLING: $l"; done
      ```
      
      The description check has two traps, and a one-line `grep` walks into both. It must read the **frontmatter only** — an unscoped `/^description:/` also matches the SKILL.md template inside a body code fence, and this repository's own `skills/skill-repo/SKILL.md` carries `description: "Use when <trigger conditions>"` there, a false positive that reads exactly like a real defect. And it must read the **whole value**: `validate-skill.sh` accepts block scalars (`|`, `>`), so a tag on a continuation line is part of the description the import rejects while a check anchored on the `description:` line alone reports the file clean.
      
      ```bash
      python3 - skills/*/SKILL.md <<'PY'
      import re, sys
      for p in sys.argv[1:]:
          txt = open(p, encoding="utf-8").read()
          fm = txt.split("\n---", 1)[0][4:] if txt.startswith("---\n") else ""
          val, grab = [], False
          for ln in fm.split("\n"):
              if grab:
                  if ln[:1] in (" ", "\t") or not ln.strip():
                      val.append(ln); continue
                  break
              m = re.match(r"description:(.*)", ln)
              if m:
                  val.append(m.group(1)); grab = True
          for ln in val:
              if re.search(r"<[A-Za-z/]", ln):
                  print(f"{p}: {ln.strip()[:110]}")
      PY
      ```
      
      ## claude.ai Organization settings sync
      
      A plugin distributed through claude.ai **Organization settings › Plugins** is packaged by a sync that applies rules to the skill repository's own files which no local tool enforces. `claude plugin validate` 2.1.281 passes all of them: it accepts a 2033-character description as `validate .` and as `validate skills`. Measured on 2026-09-23, when the sync of an internal marketplace reported 38 warnings across 13 plugins.
      
      | Rule | Sync message | Fix |
      |---|---|---|
      | `SKILL.md` description has no `<` | `SKILL.md description cannot contain XML tags` | write a placeholder as `{NAME}`, never `<NAME>` — same meaning, same length |
      | `SKILL.md` description ≤ 1024 characters | `field 'description' in SKILL.md must be at most 1024 characters` | cut process detail into the body; keep every backticked command and quoted example phrasing |
      | No top-level `bin/` | `Plugin contains a top-level bin/ directory` | the whole plugin is skipped and stays at its last synced version; move the executables to `scripts/` and call them by full path |
      
      Two properties of the sync decide how to check for these:
      
      - **It reports only the first failure per file.** A description that is both too long and bracketed shows only the length warning; the fleet had 18 reported bracket violations and 29 real ones. Fix every instance of the reported failure class — the reported one included — then re-check.
      - **Length is the parsed YAML value**, not the raw frontmatter line. A single-quoted scalar doubles every apostrophe it contains, so the raw slice runs long: 1029 raw against 1023 parsed on one description, which decides a 1024 limit. Measure with a YAML parser.
      
      Rules on the marketplace manifest itself — HTTPS plugin source URLs, the 500-character plugin description of each marketplace entry — are marketplace checks and live with the marketplace repository. The internal GitLab fleet enforces the three rules above in the `validate:descriptions` job of `ci-components/claude-code-skill`.
      
      ## Third-party install scanners
      
      Some agents scan a skill before they install it. Hermes Agent scans community skills with `tools/skills_guard.py` in [NousResearch/hermes-agent](https://github.com/NousResearch/hermes-agent). Everything below was read at commit [`26780d5`](https://github.com/NousResearch/hermes-agent/blob/26780d55ae5023817dc20e074e07127d5234d59e/tools/skills_guard.py) (`SCANNER_VERSION = "skills-guard-v6"`); rules and severities change between versions.
      
      The module uses only the Python standard library. Run it from the repository root; `scan_skill` takes a `pathlib.Path`, and a plain string raises `AttributeError`:
      
      ```bash
      curl -fsSLo skills_guard.py https://raw.githubusercontent.com/NousResearch/hermes-agent/26780d55ae5023817dc20e074e07127d5234d59e/tools/skills_guard.py
      python3 - <<'PY'
      from pathlib import Path
      from skills_guard import scan_skill
      r = scan_skill(Path("skills/<name>"), source="community")
      print(r.verdict)
      for f in r.findings:
          if f.severity in ("critical", "high"):
              print(f.severity, f.pattern_id, f"{f.file}:{f.line}", f.match)
      PY
      ```
      
      **Verdict.** Any `critical` finding makes the verdict `dangerous`, any `high` finding makes it `caution`, otherwise it is `safe`; `medium` and `low` findings never change it. A skill installed from a repository outside the scanner's trusted list is `community`: there `caution` blocks the install unless the user passes `--force`, and `dangerous` blocks it with no override.
      
      **Matching.** Every rule is a case-insensitive regex, applied line by line to `SKILL.md` and to every file with a scanned extension (`.md`, `.json`, `.sh`, `.py`, `.yaml` and others); paths listed in a `.skillignore` are skipped. A rule matches text, not what the text does:
      
      | Rule | Severity | Matches, for example |
      |---|---|---|
      | `env_exfil_curl` | critical | `curl` followed on the same line by `$…KEY`, `$…TOKEN`, `$…SECRET` or `$…PASSWORD`, such as a documented bearer-token request |
      | `dump_all_env` | high | `printenv` or `env` followed by a pipe — also the word `env` before a Markdown table pipe, and `venv\|` in a regex alternation |
      | `ssh_dir_access` | high | `~/.ssh` or `$HOME/.ssh` anywhere, including prose |
      | `git_clone`, `unpinned_pip_install` | medium | `git clone`, `pip install` without `==` |
      | `allowed_tools_field` | low | an `allowed-tools:` frontmatter line |
      
      **Reviewing a pull request that cites a scanner verdict.**
      
      - Verify each finding against the tree: open the file and line, and decide whether the text does what the rule name says.
      - Fix a real defect on its own merits, with the tests any other fix gets.
      - Decline a change whose only effect is moving the verdict, and give the reason in the review.
      - A verdict that follows from the skill's purpose does not move by rewording: a skill whose job is writing agent configuration keeps that capability whatever its files say.
      
      Precedents: [netresearch/agent-rules-skill#94](https://github.com/netresearch/agent-rules-skill/issues/94) reported a `dangerous` verdict; checking each claim against the tree turned up four real defects, fixed with regression tests, and the proposed restructuring was declined because it would not have changed the verdict. [netresearch/concourse-ci-skill#76](https://github.com/netresearch/concourse-ci-skill/pull/76) proposes clearing `ssh_dir_access` by writing the key to `deploy_key` in the task root, then running `cd source/ansible` and passing `--private-key=../deploy_key`, which resolves to `source/deploy_key` — a file that is never written.
      
  • scripts
    • bump-version.sh 7.3 KB
      #!/usr/bin/env bash
      #
      # bump-version.sh — set the release version on every surface check-version-parity
      # validates, and nothing else.
      #
      # Usage:
      #   bump-version.sh 1.2.3                 # dry run: print the diff, write nothing
      #   bump-version.sh v1.2.3 --apply        # write the files
      #   bump-version.sh --repo DIR 1.2.3      # operate on DIR instead of cwd
      #
      # Exit codes: 0 = done (or dry run clean), 1 = refused or nothing to write.
      #
      # Behavior:
      #   * Rewrites the root plugin.json .version when the repo carries the portable
      #     Agent Plugins manifest, then projects it into .claude-plugin/plugin.json
      #     via sync-plugin-manifest.sh. Root plugin.json is the source of truth;
      #     writing only the generated manifest leaves the source behind and
      #     check-version-parity.sh then rejects the half-bumped tree.
      #   * Rewrites .claude-plugin/plugin.json .version (required, must exist)
      #   * Rewrites EVERY version: line inside the YAML frontmatter of every
      #     skills/*/SKILL.md — both the indented metadata.version and a top-level
      #     version: key, preserving indentation and quote style.
      #   * Refuses when composer.json carries a version field: the release workflow
      #     derives that from the git tag, so a bumped composer version would drift.
      #   * Verifies the result with check-version-parity.sh when it is present.
      #
      # What it deliberately does NOT do: commit, tag, or push. The release order is
      # bump PR -> merge -> pull -> signed tag -> push (references/release-discipline.md).
      # A helper that tags in the same breath is how an unsigned tag reaches the wild
      # and how a half-bumped tree gets an immutable release; both have happened.
      #
      # Why every version: line and not just the first one: SKILL.md frontmatter in
      # the fleet carries either form, and some carry both. A bump that rewrites only
      # the indented metadata.version leaves a stale top-level version: behind, which
      # the tag pipeline then rejects (it-maintenance-skill v1.10.0 died this way).
      
      set -euo pipefail
      
      # Resolved before any --repo cd: a relative invocation would otherwise not find
      # its sibling scripts once the working directory has moved.
      SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
      
      REPO_DIR=""
      VERSION=""
      APPLY=0
      
      usage() {
        echo "Usage: bump-version.sh [--repo DIR] <version> [--apply]" >&2
        exit 1
      }
      
      while [[ $# -gt 0 ]]; do
        case "$1" in
          --repo)
            REPO_DIR="${2:-}"
            [[ -n "$REPO_DIR" ]] || usage
            shift 2
            ;;
          --apply)
            APPLY=1
            shift
            ;;
          -h|--help)
            usage
            ;;
          -*)
            echo "ERROR: unknown option $1" >&2
            usage
            ;;
          *)
            [[ -z "$VERSION" ]] || usage
            VERSION="$1"
            shift
            ;;
        esac
      done
      
      [[ -n "$VERSION" ]] || usage
      [[ -z "$REPO_DIR" ]] || cd "$REPO_DIR"
      
      VERSION="${VERSION#v}"
      if [[ ! "$VERSION" =~ ^[0-9]+\.[0-9]+\.[0-9]+([-+][0-9A-Za-z.-]+)?$ ]]; then
        echo "ERROR: '$VERSION' is not a semantic version" >&2
        exit 1
      fi
      
      PLUGIN_JSON=".claude-plugin/plugin.json"
      PORTABLE_JSON="plugin.json"
      COMPOSER_JSON="composer.json"
      SYNC="$SCRIPT_DIR/sync-plugin-manifest.sh"
      PARITY="$SCRIPT_DIR/check-version-parity.sh"
      
      if [[ ! -f "$PLUGIN_JSON" ]]; then
        echo "ERROR: $PLUGIN_JSON not found — run this from the repo root or pass --repo" >&2
        exit 1
      fi
      if ! jq -e 'has("version")' "$PLUGIN_JSON" > /dev/null 2>&1; then
        echo "ERROR: $PLUGIN_JSON has no .version field" >&2
        exit 1
      fi
      
      # Mirror the parity check: composer.json must not carry a version at all, so
      # refuse rather than bump it into existence.
      if [[ -f "$COMPOSER_JSON" ]] && jq -e 'has("version")' "$COMPOSER_JSON" > /dev/null 2>&1; then
        echo "ERROR: $COMPOSER_JSON has a version field — remove it first" >&2
        echo "       Git tag is the source of truth for composer packages." >&2
        exit 1
      fi
      
      OLD_PLUGIN_VERSION=$(jq -r '.version' "$PLUGIN_JSON")
      CHANGES=0
      
      report() { # path old new
        printf '  %-46s %s -> %s\n' "$1" "$2" "$3"
      }
      
      set_json_version() { # path
        local tmp
        tmp=$(mktemp)
        # jq already terminates its output with a single newline; adding one here
        # produces a blank line at EOF that the end-of-file hook then strips back.
        jq --arg v "$VERSION" '.version = $v' "$1" > "$tmp"
        mv "$tmp" "$1"
      }
      
      # Root plugin.json is the source of truth once a repo has adopted the portable
      # Agent Plugins manifest; .claude-plugin/plugin.json is generated from it. Writing
      # only the generated manifest leaves the source at the old version, which is
      # exactly what check-version-parity.sh rejects — so the bump used to end with a
      # half-written tree and exit 1 on every repo that had adopted the manifest.
      if [[ -f "$PORTABLE_JSON" ]]; then
        if ! jq -e 'has("version")' "$PORTABLE_JSON" > /dev/null 2>&1; then
          echo "ERROR: $PORTABLE_JSON has no .version field" >&2
          exit 1
        fi
        OLD_PORTABLE_VERSION=$(jq -r '.version' "$PORTABLE_JSON")
        if [[ "$OLD_PORTABLE_VERSION" != "$VERSION" ]]; then
          report "$PORTABLE_JSON" "$OLD_PORTABLE_VERSION" "$VERSION"
          CHANGES=1
          (( APPLY )) && set_json_version "$PORTABLE_JSON"
        fi
      fi
      
      if [[ "$OLD_PLUGIN_VERSION" != "$VERSION" ]]; then
        report "$PLUGIN_JSON" "$OLD_PLUGIN_VERSION" "$VERSION"
        CHANGES=1
        if (( APPLY )); then
          # Project the source of truth rather than writing the generated manifest
          # directly, so any other shared-metadata drift is corrected in the same pass.
          if [[ -f "$PORTABLE_JSON" ]] && [[ -x "$SYNC" ]]; then
            "$SYNC" > /dev/null
          else
            set_json_version "$PLUGIN_JSON"
          fi
        fi
      fi
      
      shopt -s nullglob
      SKILL_FILES=(skills/*/SKILL.md)
      shopt -u nullglob
      
      if [[ ${#SKILL_FILES[@]} -eq 0 ]]; then
        echo "WARN: no skills/*/SKILL.md files found" >&2
      fi
      
      for skill_md in "${SKILL_FILES[@]}"; do
        # Every version: line inside the frontmatter block, both forms, indentation
        # and quote style preserved. Frontmatter only: a version: in the body stays.
        before=$(awk '
          /^---$/ { fm = !fm; next }
          fm && /^[[:space:]]*version:[[:space:]]*/ {
            line = $0
            sub(/^[[:space:]]*version:[[:space:]]*/, "", line)
            gsub(/["\047]/, "", line)
            sub(/[[:space:]]+$/, "", line)
            print line
          }
        ' "$skill_md" | sort -u | paste -sd, -)
      
        [[ -n "$before" ]] || continue
        [[ "$before" != "$VERSION" ]] || continue
      
        report "$skill_md" "$before" "$VERSION"
        CHANGES=1
      
        if (( APPLY )); then
          tmp=$(mktemp)
          awk -v ver="$VERSION" '
            /^---$/ { fm = !fm; print; next }
            fm && match($0, /^[[:space:]]*version:[[:space:]]*/) {
              indent = substr($0, 1, RLENGTH)
              rest = substr($0, RLENGTH + 1)
              # keep the original quoting: "x", '"'"'x'"'"' or bare
              q = ""
              if (rest ~ /^"/) q = "\""
              else if (rest ~ /^\047/) q = "\047"
              print indent q ver q
              next
            }
            { print }
          ' "$skill_md" > "$tmp"
          mv "$tmp" "$skill_md"
        fi
      done
      
      if (( ! CHANGES )); then
        echo "Nothing to do: every version surface is already at $VERSION"
        exit 0
      fi
      
      if (( ! APPLY )); then
        echo
        echo "Dry run — nothing written. Re-run with --apply."
        exit 0
      fi
      
      if [[ -x "$PARITY" ]]; then
        echo
        "$PARITY" "$VERSION"
      fi
      
      cat <<EOF
      
      Files written. Nothing was committed, tagged or pushed — on purpose.
      
      Next (references/release-discipline.md):
        1. commit the bump and open a PR
        2. merge it, then pull the default branch
        3. git tag -s -m "v$VERSION" "v$VERSION" && git push origin "v$VERSION"
      
      Tagging before the bump PR is merged releases the old code.
      EOF
      
    • check-version-parity.sh 4.6 KB
      #!/usr/bin/env bash
      #
      # check-version-parity.sh — verify plugin.json, composer.json, and SKILL.md
      # versions are consistent before a release.
      #
      # Usage:
      #   check-version-parity.sh               # compare plugin.json vs SKILL.md
      #   check-version-parity.sh v1.2.3        # also require plugin.json == 1.2.3
      #   check-version-parity.sh 1.2.3         # same, leading v optional
      #   check-version-parity.sh --repo DIR v1.2.3   # check DIR instead of cwd
      #
      # All paths are resolved relative to the repo root. Without --repo that root
      # is the current directory, so a fleet driver iterating over many checkouts
      # must pass --repo per repo rather than a bare path argument.
      #
      # Exit codes: 0 = parity OK, 1 = mismatch or missing version.
      #
      # Behavior:
      #   * Reads .claude-plugin/plugin.json version (required)
      #   * If SKILL.md frontmatter declares a version — metadata.version or a
      #     top-level version: key — requires it matches
      #   * composer.json MUST NOT have a version field (version is derived from
      #     git tags by the release workflow). Composer-shipped version would
      #     drift from the tag.
      #   * If a tag argument is provided, requires plugin.json version == tag
      #     with 'v' prefix stripped.
      
      set -euo pipefail
      
      REPO_DIR=""
      TAG_ARG=""
      while [[ $# -gt 0 ]]; do
        case "$1" in
          --repo)
            REPO_DIR="${2:-}"
            [[ -n "$REPO_DIR" ]] || { echo "ERROR: --repo needs a directory" >&2; exit 1; }
            shift 2
            ;;
          --repo=*)
            REPO_DIR="${1#--repo=}"
            shift
            ;;
          -h|--help)
            sed -n '3,16p' "$0" | sed 's/^# \{0,1\}//'
            exit 0
            ;;
          *)
            TAG_ARG="$1"
            shift
            ;;
        esac
      done
      
      if [[ -n "$REPO_DIR" ]]; then
        cd "$REPO_DIR" || { echo "ERROR: cannot enter $REPO_DIR" >&2; exit 1; }
      fi
      
      TAG_VERSION="${TAG_ARG#v}"  # strip leading v if present (empty stays empty)
      
      PLUGIN_JSON=".claude-plugin/plugin.json"
      PORTABLE_JSON="plugin.json"
      COMPOSER_JSON="composer.json"
      
      if [[ ! -f "$PLUGIN_JSON" ]]; then
        echo "ERROR: $PLUGIN_JSON not found in $(pwd)" >&2
        echo "       Run from the skill repo root, or pass --repo <dir>." >&2
        exit 1
      fi
      
      PLUGIN_VERSION=$(jq -r '.version // empty' "$PLUGIN_JSON")
      if [[ -z "$PLUGIN_VERSION" ]]; then
        echo "ERROR: $PLUGIN_JSON has no .version field" >&2
        exit 1
      fi
      
      # Root plugin.json (Agent Plugins 1.0.0) is the source of truth for the version
      # once a repo has adopted it; .claude-plugin/plugin.json is generated from it.
      if [[ -f "$PORTABLE_JSON" ]]; then
        PORTABLE_VERSION=$(jq -r '.version // empty' "$PORTABLE_JSON")
        if [[ -z "$PORTABLE_VERSION" ]]; then
          echo "ERROR: $PORTABLE_JSON has no .version field" >&2
          exit 1
        fi
        if [[ "$PORTABLE_VERSION" != "$PLUGIN_VERSION" ]]; then
          echo "ERROR: $PORTABLE_JSON=$PORTABLE_VERSION does not match $PLUGIN_JSON=$PLUGIN_VERSION" >&2
          echo "       Bump the root plugin.json, then run sync-plugin-manifest.sh." >&2
          exit 1
        fi
      fi
      
      # composer.json MUST NOT have a version field
      if [[ -f "$COMPOSER_JSON" ]] && jq -e 'has("version")' "$COMPOSER_JSON" > /dev/null 2>&1; then
        CJ_VERSION=$(jq -r '.version // empty' "$COMPOSER_JSON")
        echo "ERROR: $COMPOSER_JSON has a version field ($CJ_VERSION) — remove it" >&2
        echo "       Git tag is the source of truth for composer packages." >&2
        exit 1
      fi
      
      # Tag argument must match plugin.json
      if [[ -n "$TAG_VERSION" && "$PLUGIN_VERSION" != "$TAG_VERSION" ]]; then
        echo "ERROR: plugin.json=$PLUGIN_VERSION does not match tag $TAG_ARG" >&2
        exit 1
      fi
      
      # SKILL.md metadata.version (if present) must match plugin.json
      MISMATCH=0
      shopt -s nullglob
      SKILL_FILES=(skills/*/SKILL.md)
      shopt -u nullglob
      
      if [[ ${#SKILL_FILES[@]} -eq 0 ]]; then
        echo "WARN: no skills/*/SKILL.md files found" >&2
      fi
      
      for skill_md in "${SKILL_FILES[@]}"; do
        # Extract the version from YAML frontmatter: metadata.version (indented)
        # or a top-level version: key — both forms exist in the fleet.
        # Tolerate quoted or unquoted values; return empty if absent.
        SKILL_VERSION=$(awk '
          /^---$/ { fm = !fm; next }
          fm && /^[[:space:]]*version:[[:space:]]*/ {
            gsub(/^[[:space:]]*version:[[:space:]]*/, "")
            gsub(/["\047]/, "")
            gsub(/[[:space:]]+$/, "")
            print
            exit
          }
        ' "$skill_md")
      
        if [[ -n "$SKILL_VERSION" && "$SKILL_VERSION" != "$PLUGIN_VERSION" ]]; then
          echo "ERROR: $skill_md metadata.version=$SKILL_VERSION does not match plugin.json=$PLUGIN_VERSION" >&2
          MISMATCH=1
        fi
      done
      
      if (( MISMATCH )); then
        exit 1
      fi
      
      if [[ -n "$TAG_VERSION" ]]; then
        echo "OK: plugin.json and tag match at $PLUGIN_VERSION"
      else
        echo "OK: plugin.json and SKILL.md versions match at $PLUGIN_VERSION"
        echo "    (pass a tag argument like v$PLUGIN_VERSION to verify tag parity)"
      fi
      
    • fleet-release-common.sh 57.5 KB
      # shellcheck shell=bash
      # fleet-release-common.sh — shared engine for the fleet-release drivers.
      #
      # Sourced by a host driver (fleet-release-github.sh here; a private-host fleet
      # ships its own driver in its own infrastructure and vendors this engine); not
      # executable on its own. Everything host-independent lives here so a bug fixed
      # for one host is fixed for both (the 2026-08-13 sweep's phase_b scripts were
      # literally identical for 20 lines — twice).
      #
      # The driver must define these host callbacks before calling a phase:
      #   host_survey_repo REPO          emit ONE normalized survey row (compact JSON)
      #   host_remote_tip REPO BRANCH    print the remote head SHA of BRANCH
      #   host_create_pr REPO BRANCH TITLE BODY TARGET
      #                                  create the PR/MR against TARGET; print '<id> <url>'
      #   host_find_release_pr REPO BRANCH
      #                                  print '<id> <url> <state> <author>' of the
      #                                  newest PR/MR for BRANCH, or nothing
      #   host_arm_automerge REPO ID     arm auto-merge (or no-op); never fails the repo
      #   host_merge_gate REPO ID        wait until merged; 0 merged, 1 failed/timeout
      #   host_merge_commit REPO ID      print the SHA the merged PR/MR produced
      #   host_release_verify REPO TAG   wait for the Release object; 0 ok, 1 missing
      #   host_origin_suffix REPO        expected origin URL suffix, e.g. 'org/repo'
      #   host_display REPO              display identity, e.g. 'netresearch/repo'
      #
      # Normalized survey row schema (both hosts emit exactly these keys):
      #   host repo resolved default archived empty unreachable duplicate_of
      #   last_release last_tag root_plugin claude_plugin composer_has_version
      #   release_gate ci_status ahead nonci changelog_unreleased files_truncated
      #   tags_failed subjects files
      #   open_release_prs [{id,url,author,branch,title}] ci_tag_rules notes
      #
      # Every phase writes into $FR_WORKDIR:
      #   survey.jsonl   one row per repo (jq -nc: ROW COUNTS MUST EQUAL LINE COUNTS)
      #   manifest.md    the approval manifest
      #   plan.skeleton.jsonl / plan.jsonl
      #   opened.jsonl   the PRs/MRs THIS sweep opened — the only ones finish may
      #                  ever wait on or merge (a colleague's bump PR is never
      #                  covered by sweep approval; release-discipline.md)
      #   logs/<phase>/<repo>.log
      
      # ---------------------------------------------------------------------------
      # Globals (drivers may override before sourcing or via flags)
      # ---------------------------------------------------------------------------
      # shellcheck disable=SC2034  # consumed across the lib/driver file boundary
      FR_BASE_DIR="${FR_BASE_DIR:-$HOME/projects}"
      FR_POLL_SECONDS="${FR_POLL_SECONDS:-20}"
      FR_MERGE_TIMEOUT="${FR_MERGE_TIMEOUT:-1800}"
      FR_RELEASE_TIMEOUT="${FR_RELEASE_TIMEOUT:-600}"
      FR_SELF_LOGIN="${FR_SELF_LOGIN:-}"        # resolved by the driver (gh/glab)
      FR_BRANCH_PREFIX="release/v"              # owned by github-release-skill
      FR_COMMIT_PREFIX="chore(release): v"      # owned by github-release-skill
      # CI-only delta filter: commits touching only these paths make no
      # consumer-visible change, so they are not a release (release-discipline.md).
      # Note the test is "nothing a consumer can observe", not "byte-identical
      # archive": the lint configs below DO travel inside the release archive, but a
      # hook-pin bump changes nothing about the skill a consumer installs. The
      # `.github/**` entries are byte-identical as well; the others are not, and that
      # is deliberate.
      # shellcheck disable=SC2034  # consumed by the host drivers' survey callbacks
      FR_CI_ONLY_RE='^\.github/|^\.gitlab-ci\.yml$|^renovate\.json$|^\.pre-commit-config\.yaml$|^\.markdownlint(-cli2)?\.jsonc$|^\.yamllint\.yml$|^\.editorconfig$'
      # The only files a bump commit may touch.
      FR_ALLOWLIST_RE='^(plugin\.json|\.claude-plugin/plugin\.json|skills/[^/]+/SKILL\.md|CHANGELOG\.md)$'
      # Version-aware compare/sort, shared by every jq call site. Extracts the
      # TRAILING semver so custom tag conventions (tender-estimation--v0.6.1) order
      # correctly too — a parse that degrades them all to [0] makes max_by pick an
      # arbitrary old tag and misclassify the repo as diverged (seen live 2026-08-13).
      # Input with no trailing semver at all degrades to 0.0.0; the manifest flags
      # those rows as nonstandard instead of trusting the math.
      FR_JQ_VPARSE='def vparse: (capture("(?<v>[0-9]+\\.[0-9]+\\.[0-9]+([-+][0-9A-Za-z.-]+)?)$") // {v: "0.0.0"}) | .v | split("-")[0] | split("+")[0] | split(".") | map(tonumber? // 0);'
      
      FR_SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
      
      fr_tool() { # NAME — locate a skill-repo tool (bump-version.sh, roll-changelog.py,
          # …). A driver vendored into another repo has no sibling copies, so the
          # lookup falls back to the INSTALLED skill-repo skill (reading/executing
          # installed paths is fine; editing them is what cache-safety forbids).
          # FR_TOOLS_DIR pins an explicit source-of-truth checkout when set.
          local name="$1" d
          for d in "${FR_TOOLS_DIR:-}" "$FR_SCRIPT_DIR" \
              "$HOME/.claude/skills/skill-repo/scripts"; do
              [[ -n "$d" && -f "$d/$name" ]] && { echo "$d/$name"; return 0; }
          done
          d=$(find "$HOME/.claude/plugins/cache" -path '*/skills/skill-repo/scripts/'"$name" 2> /dev/null | sort -V | tail -1)
          [[ -n "$d" ]] && { echo "$d"; return 0; }
          fr_err "cannot locate $name (set FR_TOOLS_DIR to a skill-repo-skill checkout's skills/skill-repo/scripts)"
          return 1
      }
      
      fr_note() { printf '%s\n' "$*"; }
      fr_err()  { printf 'ERROR: %s\n' "$*" >&2; }
      
      fr_changelog_state() { # RAW-CONTENT — "absent" | "empty" | "has-content" |
          # "no-unreleased" | "unknown". An empty argument means the survey found no
          # CHANGELOG.md. Classification is delegated to roll-changelog.py
          # --check-unreleased so the survey warns by the SAME emptiness rule the
          # bump-time roll enforces — five repos failed mid-bump on empty
          # [Unreleased] sections in the 2026-08-27 sweep, each one visible to the
          # survey that ran hours earlier.
          local raw="$1" tool tmp state
          [[ -n "$raw" ]] || { echo "absent"; return 0; }
          tool=$(fr_tool roll-changelog.py 2> /dev/null) || { echo "unknown"; return 0; }
          tmp=$(mktemp "${TMPDIR:-/tmp}/fr-clstate.XXXXXX") || { echo "unknown"; return 0; }
          printf '%s\n' "$raw" > "$tmp"
          state=$(python3 "$tool" --check-unreleased "$tmp" 2> /dev/null) || state="unknown"
          rm -f "$tmp"
          echo "${state:-unknown}"
      }
      fr_die()  { fr_err "$*"; exit 1; }
      
      # ---------------------------------------------------------------------------
      # Preflight
      # ---------------------------------------------------------------------------
      fr_preflight() { # PHASE TOOL...
          local phase="$1"; shift
          [[ "${BASH_VERSINFO[0]:-0}" -ge 4 ]] \
              || fr_die "bash >= 4 required (this is bash ${BASH_VERSION:-unknown}); on macOS: brew install bash"
          local t
          for t in git jq "$@"; do
              command -v "$t" > /dev/null 2>&1 || fr_die "required tool missing: $t"
          done
          case "$phase" in
              bump)
                  command -v python3 > /dev/null 2>&1 \
                      || fr_die "python3 required for the changelog roll"
                  ;;
              *) : ;;
          esac
          # Never write into installed/cache paths (release-discipline.md,
          # "Cache Safety") — neither the workdir nor the checkouts base.
          local p
          for p in "$FR_WORKDIR" "$FR_BASE_DIR"; do
              case "$p" in
                  */.claude/skills*|*/.claude/plugins*|*/.bare|*/.bare/*)
                      fr_die "refusing to operate under installed/cache path: $p"
                      ;;
                  *) : ;;
              esac
          done
          mkdir -p "$FR_WORKDIR/logs/$phase"
          # Signing preflight: an agent that silently dropped the key fails every
          # repo mid-fleet; fail here instead. FR_SKIP_SIGN_CHECK=1 for GPG setups
          # the check cannot see.
          if [[ "$phase" == "bump" || "$phase" == "finish" ]] \
              && [[ "${FR_SKIP_SIGN_CHECK:-0}" != "1" ]]; then
              local fmt
              fmt=$(git config --get gpg.format 2> /dev/null || echo openpgp)
              if [[ "$fmt" == "ssh" ]] && ! ssh-add -l > /dev/null 2>&1; then
                  fr_die "gpg.format=ssh but ssh-add -l lists no key — signing every commit/tag would fail (FR_SKIP_SIGN_CHECK=1 to override)"
              fi
          fi
          fr_lock "$phase"
      }
      
      fr_lock() { # PHASE — one sweep per workdir; stale locks (dead PID) are taken over
          local lock="$FR_WORKDIR/.lock" pid
          if mkdir "$lock" 2> /dev/null; then
              echo "$$" > "$lock/pid"
          else
              pid=$(cat "$lock/pid" 2> /dev/null || echo "")
              if [[ -n "$pid" ]] && kill -0 "$pid" 2> /dev/null; then
                  fr_die "workdir locked by running pid $pid ($lock) — a second sweep on the same workdir would interleave logs and state"
              fi
              fr_note "taking over stale lock (pid ${pid:-unknown} is gone)"
              echo "$$" > "$lock/pid"
          fi
          # shellcheck disable=SC2064  # expand $lock now: it is stable and local
          trap "rm -rf '$lock'" EXIT
      }
      
      # ---------------------------------------------------------------------------
      # Small shared helpers
      # ---------------------------------------------------------------------------
      fr_vercmp() { # A B — prints -1|0|1 (numeric per component; never lexicographic:
          # bash [[ < ]] and BSD sort would call 1.10.0 older than 1.9.0)
          jq -rn --arg a "$1" --arg b "$2" \
              "$FR_JQ_VPARSE"' ($a|vparse) as $A | ($b|vparse) as $B
               | if $A > $B then "1" elif $A < $B then "-1" else "0" end'
      }
      
      fr_resolve_gitdir() { # DIR — three checkout layouts coexist (release-discipline.md)
          local d="$1"
          if [[ -d "$d/.bare" ]]; then
              echo "$d/.bare"
          elif git -C "$d" rev-parse --absolute-git-dir 2> /dev/null; then
              return 0
          elif git -C "$d/main" rev-parse --absolute-git-dir 2> /dev/null; then
              return 0
          else
              return 1
          fi
      }
      
      fr_check_origin() { # GITDIR EXPECTED — folder names lie; origin URLs don't.
          # A checkout whose origin points elsewhere would receive the wrong repo's
          # release branch. Normalize scheme/user/port away and compare host/path
          # EXACTLY — a suffix match would accept a crafted path whose tail merely
          # mimics host/org/repo.
          local gitdir="$1" want="$2" url
          url=$(git -C "$gitdir" config --get remote.origin.url 2> /dev/null || echo "")
          [[ -n "$url" ]] || { fr_err "no remote.origin.url in $gitdir"; return 1; }
          local norm="${url%.git}"
          norm="${norm#*://}"          # scheme
          norm="${norm#*@}"            # user
          if [[ "$norm" =~ ^([^/:]+):[0-9]+(/.*)$ ]]; then
              norm="${BASH_REMATCH[1]}${BASH_REMATCH[2]}"   # host:PORT/path -> host/path
          fi
          norm="${norm//:/\/}"         # scp-style host:org/repo -> host/org/repo
          if [[ "$norm" == "$want" ]]; then
              return 0
          fi
          fr_err "origin of $gitdir is '$url' (normalized $norm), expected $want — refusing to push there"
          return 1
      }
      
      fr_worktree_path() { # DIR VERSION — layout-aware worktree location
          local d="$1" v="$2"
          if [[ -d "$d/.bare" ]]; then
              echo "$d/release-$v"
          else
              echo "$d-release-$v"
          fi
      }
      
      fr_fetch_branches() { # GITDIR — branches only: one stale tag would fail the
          # whole fetch ('would clobber existing tag') and take the branches with it
          git -C "$1" fetch origin --prune --no-tags \
              "+refs/heads/*:refs/remotes/origin/*" 2>&1
      }
      
      fr_show_version() { # GITDIR REF FILE — .version of a JSON file at a ref, or empty
          git -C "$1" show "$2:$3" 2> /dev/null | jq -r '.version // empty' 2> /dev/null || true
      }
      
      fr_parity_from_ref() { # GITDIR REF VERSION — parity read from the ref, never a
          # working tree (worktrees drift; ~6 of 23 disagreed in one sweep)
          local gitdir="$1" ref="$2" want="$3" fail=0 cp root f sv
          cp=$(fr_show_version "$gitdir" "$ref" ".claude-plugin/plugin.json")
          if [[ -z "$cp" ]]; then
              fr_err "parity: $ref has no .claude-plugin/plugin.json version"
              return 1
          fi
          [[ "$cp" == "$want" ]] || { fr_err "parity: .claude-plugin/plugin.json=$cp != $want"; fail=1; }
          if git -C "$gitdir" cat-file -e "$ref:plugin.json" 2> /dev/null; then
              root=$(fr_show_version "$gitdir" "$ref" "plugin.json")
              [[ "$root" == "$want" ]] || { fr_err "parity: plugin.json=$root != $want"; fail=1; }
          fi
          # composer.json must NOT carry a version — the tag is the source of truth,
          # and TAG-ONLY repos never pass through bump-version.sh's own refusal.
          if git -C "$gitdir" show "$ref:composer.json" 2> /dev/null \
              | jq -e 'has("version")' > /dev/null 2>&1; then
              fr_err "parity: composer.json at $ref carries a version field"
              fail=1
          fi
          while IFS= read -r f; do
              [[ -n "$f" ]] || continue
              sv=$(git -C "$gitdir" show "$ref:$f" | awk '
                  # Frontmatter is the FIRST --- ... --- block only; a --- rule in
                  # the body must not reopen it (a body "version:" would then be
                  # parsed as the skill version).
                  /^---$/ { if (fm) exit; fm = 1; next }
                  fm && /^[[:space:]]*version:[[:space:]]*/ {
                      gsub(/^[[:space:]]*version:[[:space:]]*/, "")
                      gsub(/["\047]/, "")
                      gsub(/[[:space:]]+$/, "")
                      print
                      exit
                  }')
              if [[ -n "$sv" && "$sv" != "$want" ]]; then
                  fr_err "parity: $f version=$sv != $want"
                  fail=1
              fi
          done < <(git -C "$gitdir" ls-tree -r --name-only "$ref" \
              | grep -E '^skills/[^/]+/SKILL\.md$' || true)
          return "$fail"
      }
      
      fr_remote_tag_commit() { # GITDIR TAG — the COMMIT a remote tag points at, or
          # empty. Signed tags are annotated: ls-remote's plain line is the tag
          # OBJECT sha; only the peeled ^{} line is the commit. Comparing unpeeled
          # would call every resumed signed tag "wrong SHA".
          # Exit 2 when the remote could not be read at all — "cannot see the tag"
          # must never read as "tag absent", or a network blip triggers a doomed
          # duplicate tag push.
          local gitdir="$1" tag="$2" out
          out=$(git -C "$gitdir" ls-remote origin \
              "refs/tags/$tag" "refs/tags/$tag^{}" 2> /dev/null) || return 2
          if printf '%s\n' "$out" | grep -q "refs/tags/$tag\^{}$"; then
              printf '%s\n' "$out" | awk '/\^\{\}$/ { print $1 }'
          else
              printf '%s\n' "$out" | awk 'NF { print $1; exit }'
          fi
      }
      
      fr_tag_is_signed() { # GITDIR TAG — annotated AND carrying a signature block;
          # a crashed run's own tag was created with -s, anything else is suspect
          [[ "$(git -C "$1" cat-file -t "$2" 2> /dev/null)" == "tag" ]] || return 1
          git -C "$1" cat-file tag "$2" | grep -Eq -- '-----BEGIN (PGP|SSH) SIGNATURE-----'
      }
      
      fr_tag_on_tip() { # GITDIR TAG TIP — signed tag on the verified remote tip; no
          # checkout, no working tree (a parked worktree stays untouched). Idempotent
          # across the crash windows: tag already remote, tag local-only.
          local gitdir="$1" tag="$2" tip="$3" remote local_commit rc
          rc=0
          remote=$(fr_remote_tag_commit "$gitdir" "$tag") || rc=$?
          if [[ "$rc" -eq 2 ]]; then
              fr_err "cannot read remote tags for $tag (ls-remote failed) — not tagging blind"
              return 1
          fi
          if [[ -n "$remote" ]]; then
              if [[ "$remote" == "$tip" ]]; then
                  fr_note "tag $tag already on remote at the tip — skipping to verification"
                  return 0
              fi
              fr_err "tag $tag exists on remote at $remote, tip is $tip — a released tag is immutable; this needs a human"
              return 1
          fi
          if git -C "$gitdir" rev-parse -q --verify "refs/tags/$tag" > /dev/null 2>&1; then
              local_commit=$(git -C "$gitdir" rev-parse "$tag^{commit}")
              if [[ "$local_commit" == "$tip" ]] && fr_tag_is_signed "$gitdir" "$tag"; then
                  fr_note "local signed tag $tag already at the tip (crashed before push) — pushing it"
              elif [[ "$local_commit" == "$tip" ]]; then
                  fr_err "local tag $tag is at the tip but UNSIGNED — not this driver's work; inspect, 'git -C $gitdir tag -d $tag', re-run"
                  return 1
              else
                  fr_err "local tag $tag points at $local_commit, tip is $tip — inspect and 'git -C $gitdir tag -d $tag' manually"
                  return 1
              fi
          else
              git -C "$gitdir" tag -s "$tag" -m "$tag" "$tip" || { fr_err "tag -s failed"; return 1; }
          fi
          git -C "$gitdir" push origin "refs/tags/$tag"
          local ec=$?
          fr_note "TAG PUSH EXIT: $ec"
          return "$ec"
      }
      
      fr_check_log_invariant() { # LOGDIR EXPECTED — a killed batch 'completes' with
          # fewer logs than repos and nothing says so; count before trusting anything
          local logdir="$1" expected="$2" actual
          actual=$(find "$logdir" -maxdepth 1 -name '*.log' | wc -l | tr -d ' ')
          if [[ "$actual" -ne "$expected" ]]; then
              fr_err "log invariant violated: $actual logs for $expected repos in $logdir"
              return 1
          fi
          fr_note "log invariant holds: $actual/$expected"
      }
      
      fr_summarize_logs() { # LOGDIR — count explicit end-state markers; a truncated
          # log (hard kill mid-repo) is neither OK nor FAIL and must be visible
          local logdir="$1" f ok=0 failed=0 other=0
          for f in "$logdir"/*.log; do
              [[ -e "$f" ]] || continue
              if grep -q '^OK ' "$f"; then
                  ok=$((ok + 1))
              elif grep -q '^FAIL ' "$f"; then
                  failed=$((failed + 1))
                  fr_note "FAIL: $(basename "$f" .log): $(grep '^FAIL ' "$f" | tail -1)"
              else
                  other=$((other + 1))
                  fr_note "INDETERMINATE (no OK/FAIL marker — truncated?): $f"
              fi
          done
          fr_note "summary: $ok ok, $failed failed, $other indeterminate"
          [[ "$failed" -eq 0 && "$other" -eq 0 ]]
      }
      
      # ---------------------------------------------------------------------------
      # Survey
      # ---------------------------------------------------------------------------
      fr_survey() { # REPO... — remote-first; per-repo failures become rows, not aborts.
          # A re-survey of a SUBSET (the manifest's own "re-run survey --repos X"
          # advice) replaces exactly those rows and keeps the rest — truncating the
          # whole file would throw away the sweep it is trying to repair.
          local out="$FR_WORKDIR/survey.jsonl" seen="$FR_WORKDIR/.seen-resolved" repo row resolved tmp
          if [[ -s "$out" ]]; then
              tmp=$(mktemp)
              jq -c --args 'select(.repo as $r | $ARGS.positional | index($r) | not)' "$@" < "$out" > "$tmp"
              mv "$tmp" "$out"
              jq -r 'select((.resolved // "") != "") | .resolved' "$out" > "$seen"
              fr_note "keeping $(wc -l < "$out" | tr -d ' ') existing rows; re-surveying: $*"
          else
              : > "$out"
              : > "$seen"
          fi
          for repo in "$@"; do
              fr_note "surveying $(host_display "$repo") ..."
              row=$(host_survey_repo "$repo" < /dev/null) \
                  || row=$(jq -nc --arg r "$repo" --arg h "$FR_HOST" \
                      '{host:$h, repo:$r, unreachable:true}')
              # A renamed repo answers on both names (API redirect) — the SECOND row
              # for one resolved identity must not become a second bump PR.
              resolved=$(jq -r '.resolved // empty' <<< "$row")
              if [[ -n "$resolved" ]] && grep -Fxq "$resolved" "$seen"; then
                  row=$(jq -nc --arg r "$repo" --arg h "$FR_HOST" --arg d "$resolved" \
                      '{host:$h, repo:$r, duplicate_of:$d}')
              elif [[ -n "$resolved" ]]; then
                  printf '%s\n' "$resolved" >> "$seen"
              fi
              printf '%s\n' "$row" >> "$out"   # rows are compact: wc -l == repo count
          done
          fr_note ""
          fr_note "survey: $(wc -l < "$out" | tr -d ' ') rows -> $out"
          fr_note "next: $FR_DRIVER manifest --workdir $FR_WORKDIR"
      }
      
      # ---------------------------------------------------------------------------
      # Classification + manifest
      # ---------------------------------------------------------------------------
      fr_classify() { # ROW-JSON — one shared implementation for both hosts; blocks
          # outrank actions (a composer-version repo must never reach TAG-ONLY)
          jq -r --arg self "$FR_SELF_LOGIN" "$FR_JQ_VPARSE"'
              if .unreachable == true then "UNREACHABLE"
              elif (.duplicate_of // "") != "" then "DUPLICATE"
              elif .empty == true then "EMPTY"
              elif .archived == true then "BLOCKED-ARCHIVED"
              elif .composer_has_version == true then "BLOCKED-COMPOSER-VERSION"
              elif ([.open_release_prs[]? | select(.author != $self)] | length) > 0 then "BLOCKED-FOREIGN-PR"
              elif ([.open_release_prs[]? | select(.author == $self)] | length) > 0 then "OWN-PR-OPEN"
              elif .release_gate == false then "NO-RELEASE-JOB"
              # Before FIRST-RELEASE, deliberately: a failed tag query leaves
              # last_tag empty, which is indistinguishable from "never tagged" and
              # would offer an initial tag over an existing history. Placed here and
              # not in the else-branch because every branch below reads last_tag.
              elif (.tags_failed // false) == true then "SURVEY-INCOMPLETE"
              elif (.last_release // "") == "" and (.last_tag // "") == "" then "FIRST-RELEASE"
              elif (.last_tag // "") != "" and (.last_tag != (.last_release // "")) then "TAG-RELEASE-DIVERGED"
              else
                  ((.claude_plugin // "") | vparse) as $p
                  | ((.last_tag // "") | vparse) as $t
                  | if $p > $t then "TAG-ONLY"
                    elif $p < $t then "BEHIND-TAG"
                    # ahead/nonci < 0 = the compare MEASUREMENT failed, which must
                    # never read as "no delta": a transient API blip would silently
                    # drop a repo with releasable commits from the sweep.
                    elif (.ahead // 0) < 0 or (.nonci // 0) < 0 then "SURVEY-INCOMPLETE"
                    elif (.ahead // 0) == 0 then "UP-TO-DATE"
                    elif (.nonci // 0) == 0 then "SKIP-CI-ONLY"
                    else "BUMP"
                    end
              end' <<< "$1"
      }
      
      fr_manifest() {
          local survey="$FR_WORKDIR/survey.jsonl" manifest="$FR_WORKDIR/manifest.md"
          local skeleton="$FR_WORKDIR/plan.skeleton.jsonl" row cls
          [[ -s "$survey" ]] || fr_die "no survey at $survey — run: $FR_DRIVER survey first"
          [[ -n "$FR_SELF_LOGIN" ]] || fr_die "FR_SELF_LOGIN unresolved — cannot apply the foreign-PR author gate"
          : > "$skeleton"
          {
              echo "# Fleet Release Manifest ($FR_HOST, $(date +%F))"
              echo
              echo "Approval covers ONLY the PRs this sweep opens. Rows marked"
              echo "BLOCKED-FOREIGN-PR need that author's separate go-ahead."
              echo
              echo "| Repo | Class | Last tag | Release | plugin.json | Ahead | Non-CI | Changelog | CI on default | Notes |"
              echo "|---|---|---|---|---|---|---|---|---|---|"
          } > "$manifest"
          while IFS= read -u 3 -r row; do
              cls=$(fr_classify "$row")
              jq -r --arg cls "$cls" "$FR_JQ_VPARSE"'
                  def ordash: if . == null or . == "" then "—" else . end;
                  [ .repo,
                    $cls,
                    (.last_tag | ordash),
                    (.last_release | ordash),
                    (.claude_plugin | ordash),
                    ((.ahead // "") | tostring | ordash),
                    ((.nonci // "") | tostring | ordash),
                    ((.changelog_unreleased // "") | if . == "empty" then "EMPTY" else ordash end),
                    (.ci_status | ordash),
                    ([ (if (.open_release_prs // []) | length > 0
                        then "PRs: " + ([.open_release_prs[] | "\(.url) by @\(.author)"] | join("; "))
                        else empty end),
                       (if (.last_tag // "") != "" and ((.last_tag | test("^v?[0-9]+\\.[0-9]+\\.[0-9]+")) | not)
                        then "nonstandard tag convention" else empty end),
                       (if .files_truncated == true then "compare file list truncated" else empty end),
                       (if $cls == "NO-RELEASE-JOB"
                        then "no tag-triggered release job — adopt the shared release CI first; this driver does not write CI files"
                        else empty end),
                       (if $cls == "TAG-RELEASE-DIVERGED"
                        then "tag exists without a matching Release — check the release job matched the tag pattern; usually a CI fix plus a NEW version (re-plan as BUMP)"
                        else empty end),
                       (if $cls == "BLOCKED-ARCHIVED"
                        then "unarchive is a human decision; afterwards re-run survey --repos \(.repo)"
                        else empty end),
                       (if $cls == "BUMP" and (.changelog_unreleased // "") == "empty"
                        then "CHANGELOG [Unreleased] is empty — write the release entries before the bump phase, or its roll fails on exactly this"
                        else empty end),
                       (if $cls == "SURVEY-INCOMPLETE"
                        then (if (.tags_failed // false) == true
                              then "the TAG QUERY failed (this is not an untagged repo) — re-run survey --repos \(.repo)"
                              else "the compare MEASUREMENT failed (this is not a no-delta result) — re-run survey --repos \(.repo)"
                              end)
                        else empty end),
                       (if (.ci_tag_rules // "") != ""
                        then "tag rules: " + (.ci_tag_rules | gsub("\\|"; "¦"))
                        else empty end),
                       ((.notes // "") | if . == "" then empty else gsub("\\|"; "¦") end)
                     ] | join(". ") | ordash)
                  ] | "| " + join(" | ") + " |"' <<< "$row" >> "$manifest"
              case "$cls" in
                  BUMP|FIRST-RELEASE|TAG-ONLY|TAG-RELEASE-DIVERGED|OWN-PR-OPEN)
                      # TAG-ONLY carries its committed version — tagging a prepared
                      # repo at that version keeps parity true by construction.
                      # OWN-PR-OPEN carries the version its open PR branch encodes,
                      # so a crashed sweep's own PR re-enters the flow instead of
                      # being silently dropped. BUMP versions and bodies are operator
                      # judgment: the driver never invents either. `head` is the
                      # surveyed default-branch tip — finish refuses to tag a
                      # TAG-ONLY repo whose branch moved past it.
                      jq -c --arg cls "$cls" \
                          '{repo, classification: $cls, default: (.default // "main"),
                            last: (.claude_plugin // ""), last_tag: (.last_tag // ""),
                            head: (.head // ""),
                            version: (if $cls == "TAG-ONLY" then (.claude_plugin // "")
                                      elif $cls == "OWN-PR-OPEN"
                                      then ([.open_release_prs[]? | .branch
                                            | capture("^release/v(?<v>.+)$") | .v] | first // "")
                                      else "" end),
                            body: "", tag: "",
                            subjects: (.subjects // [])}' <<< "$row" >> "$skeleton"
                      ;;
              esac
          done 3< "$survey"
          {
              echo
              echo "Preconditions are per-repo MEASURED values above (CI on default,"
              echo "open PRs, composer version field) — never assumed checkmarks."
              echo
              echo "Scope: this driver releases the DEFAULT branch tip only. Maintenance-line"
              echo "releases (--latest=false handling) are out of scope — do those by hand."
          } >> "$manifest"
          cat "$manifest"
          fr_note ""
          fr_note "manifest -> $manifest"
          fr_note "plan skeleton -> $skeleton"
          fr_note "next: 1. review the manifest with the operator/user and get approval"
          fr_note "      2. cp $skeleton $FR_WORKDIR/plan.jsonl"
          fr_note "      3. fill 'version' (and 'body') for every BUMP/FIRST-RELEASE row; drop rows to skip"
          fr_note "      4. $FR_DRIVER bump --workdir $FR_WORKDIR"
      }
      
      # ---------------------------------------------------------------------------
      # Plan handling
      # ---------------------------------------------------------------------------
      fr_plan_validate() { # PLAN — refuse the whole phase before mutating anything
          local plan="$1" bad
          [[ -s "$plan" ]] || fr_die "no plan at $plan (copy plan.skeleton.jsonl and fill it)"
          jq -e . > /dev/null 2>&1 < "$plan" || fr_die "plan is not valid JSON lines: $plan"
          # Repo names become branch names, log paths and worktree paths — a slash
          # (GitLab subgroup) or space would land the log/worktree somewhere else.
          bad=$(jq -r 'select((.repo // "") | test("^[A-Za-z0-9._-]+$") | not) | .repo // "<empty>"' < "$plan")
          [[ -z "$bad" ]] || fr_die "plan rows with unusable repo names (subgroups are unsupported): $(tr '\n' ' ' <<< "$bad")"
          # One row per repo: two rows with different versions would open two bump
          # PRs against one repo.
          bad=$(jq -r '.repo' < "$plan" | sort | uniq -d)
          [[ -z "$bad" ]] || fr_die "duplicate plan rows for: $(tr '\n' ' ' <<< "$bad")"
          bad=$(jq -r 'select((.classification == "BUMP" or .classification == "FIRST-RELEASE" or .classification == "TAG-ONLY" or .classification == "OWN-PR-OPEN")
                        and ((.version // "") == "")) | .repo' < "$plan")
          [[ -z "$bad" ]] || fr_die "plan rows without a version: $(tr '\n' ' ' <<< "$bad")— fill them or drop them"
          bad=$(jq -r 'select((.version // "") != "")
                       | select((.version | test("^[0-9]+\\.[0-9]+\\.[0-9]+([-+][0-9A-Za-z.-]+)?$")) | not) | .repo' < "$plan")
          [[ -z "$bad" ]] || fr_die "plan rows with a non-semver version: $(tr '\n' ' ' <<< "$bad")"
          # Monotonicity: a planned version below or equal to the surveyed one would
          # ship a parity-consistent version REGRESSION that no later gate can see.
          # Two classes are exceptions, for opposite reasons:
          #   TAG-ONLY      the bump already merged, so the plan must name that exact
          #                 committed version and nothing else.
          #   FIRST-RELEASE there is no previous release OR tag to regress from — its
          #                 `.last` is the committed plugin.json version, not a
          #                 shipped one. Tagging it as-is is the normal case for a
          #                 repo someone already prepared, and bumping above it is
          #                 equally legal; only going BELOW the committed version is
          #                 wrong, because parity would then fail at tag time.
          #                 Rejecting equality here left a prepared first release
          #                 unrepresentable: the operator had to either retype the row
          #                 TAG-ONLY by hand or invent a version nobody wrote
          #                 (hit on both ecom-* repos in the 2026-08-17 sweep).
          bad=$(jq -r "$FR_JQ_VPARSE"'
              select((.version // "") != "" and (.last // "") != "")
              | if .classification == "TAG-ONLY"
                then select(.version != .last)
                     | "\(.repo) (TAG-ONLY must tag the committed \(.last), not \(.version))"
                elif .classification == "FIRST-RELEASE"
                then select((.version | vparse) < (.last | vparse))
                     | "\(.repo) (FIRST-RELEASE may tag or exceed the committed \(.last), not fall below it with \(.version))"
                else select((.version | vparse) <= (.last | vparse))
                     | "\(.repo) (\(.version) is not above the surveyed \(.last))"
                end' < "$plan")
          [[ -z "$bad" ]] || fr_die "plan version ordering: $(tr '\n' ';' <<< "$bad")"
      }
      
      fr_record_opened() { # REPO ID URL BRANCH — append-only; survives crashes
          jq -nc --arg h "$FR_HOST" --arg r "$1" --arg i "$2" --arg u "$3" --arg b "$4" \
              '{host: $h, repo: $r, id: $i, url: $u, branch: $b}' >> "$FR_WORKDIR/opened.jsonl"
      }
      
      fr_opened_lookup() { # REPO — id of the PR/MR this sweep opened, or empty.
          # Host-filtered: a workdir reused across both drivers must never resolve
          # the other host's PR number.
          [[ -f "$FR_WORKDIR/opened.jsonl" ]] || return 0
          jq -r --arg r "$1" --arg h "$FR_HOST" \
              'select(.repo == $r and .host == $h) | .id' "$FR_WORKDIR/opened.jsonl" | tail -1
      }
      
      # ---------------------------------------------------------------------------
      # Bump phase
      # ---------------------------------------------------------------------------
      fr_bump_one() { # ROW — runs inside the per-repo subshell; log via redirection
          local row="$1"
          local repo version body default last cls
          repo=$(jq -r '.repo' <<< "$row")
          version=$(jq -r '.version // ""' <<< "$row")
          body=$(jq -r '.body // ""' <<< "$row")
          default=$(jq -r '.default // "main"' <<< "$row")
          last=$(jq -r '.last // ""' <<< "$row")
          cls=$(jq -r '.classification // "BUMP"' <<< "$row")
          local branch="${FR_BRANCH_PREFIX}${version}"
          echo "=== $(host_display "$repo") v$version ($cls) ==="
          [[ -n "$version" ]] || { echo "FAIL $repo: plan row has no version"; return 1; }
      
          if [[ "$cls" == "TAG-ONLY" ]]; then
              echo "OK $repo v$version no bump needed (prepared repo; finish will tag)"
              return 0
          fi
          # OWN-PR-OPEN re-enters here on purpose: the remote-reality checks below
          # find the sweep's own open PR and record it instead of double-bumping.
          if [[ "$cls" != "BUMP" && "$cls" != "FIRST-RELEASE" && "$cls" != "OWN-PR-OPEN" ]]; then
              echo "FAIL $repo: classification $cls does not belong in the bump phase"
              return 1
          fi
      
          local dir="$FR_BASE_DIR/$repo" gitdir
          gitdir=$(fr_resolve_gitdir "$dir") \
              || { echo "FAIL $repo: no checkout at $dir (survey is remote-first; bump needs a local checkout — clone it deliberately, this driver will not)"; return 1; }
          fr_check_origin "$gitdir" "$(host_origin_suffix "$repo")" \
              || { echo "FAIL $repo: origin mismatch"; return 1; }
          fr_fetch_branches "$gitdir" || { echo "FAIL $repo: fetch"; return 1; }
      
          # Freshness gate: classification happened at survey time and approval is
          # human-paced. If the fleet moved since (a colleague released), writing the
          # planned version would be a parity-consistent version REGRESSION that
          # every later gate happily waves through — refuse instead.
          local cur
          cur=$(fr_show_version "$gitdir" "origin/$default" ".claude-plugin/plugin.json")
          if [[ "$cur" == "$version" ]]; then
              echo "OK $repo v$version already at $version on origin/$default (nothing to bump; finish will tag)"
              return 0
          fi
          if [[ "$cur" != "$last" ]]; then
              echo "FAIL $repo: fleet moved since survey (surveyed $last, origin/$default now $cur) — re-run survey and re-plan this repo"
              return 1
          fi
      
          # Remote reality is the resume truth, interrogated per step; the crash
          # windows (after commit, after push, after PR create) each land in one of
          # these branches instead of a dead end.
          if git -C "$gitdir" rev-parse -q --verify "refs/remotes/origin/$branch" > /dev/null 2>&1; then
              local pr prid prurl prstate prauthor
              pr=$(host_find_release_pr "$repo" "$branch" < /dev/null || true)
              if [[ -n "$pr" ]]; then
                  read -r prid prurl prstate prauthor <<< "$pr"
                  case "$prstate" in
                      merged|MERGED)
                          echo "OK $repo v$version bump already merged ($prurl) — run finish"
                          return 0
                          ;;
                      open|OPEN|opened)
                          if [[ "$prauthor" != "$FR_SELF_LOGIN" ]]; then
                              echo "FAIL $repo: open release PR $prurl by @$prauthor — a colleague's PR needs their go, not sweep approval"
                              return 1
                          fi
                          [[ -n "$(fr_opened_lookup "$repo")" ]] || fr_record_opened "$repo" "$prid" "$prurl" "$branch"
                          echo "OK $repo v$version PR already open: $prurl"
                          return 0
                          ;;
                  esac
              fi
              local bv extra
              bv=$(fr_show_version "$gitdir" "origin/$branch" ".claude-plugin/plugin.json")
              if [[ "$bv" == "$version" ]]; then
                  # Version alone does not prove the branch is a bump: an auto-merge
                  # will merge whatever else it carries unseen. Only the four version
                  # surfaces may differ from the base.
                  extra=$(git -C "$gitdir" diff --name-only "origin/$default" "origin/$branch" \
                      | grep -vE "$FR_ALLOWLIST_RE" || true)
                  if [[ -n "$extra" ]]; then
                      echo "FAIL $repo: remote branch $branch changes more than the version surfaces ($(tr '\n' ' ' <<< "$extra")) — not adopting it for auto-merge; inspect it"
                      return 1
                  fi
                  echo "branch $branch already pushed at $version (crashed before PR create) — creating the PR"
                  fr_create_and_record "$repo" "$branch" "$version" "$body" "$default" || return 1
                  echo "OK $repo v$version PR created from existing branch"
                  return 0
              fi
              echo "FAIL $repo: leftover remote branch $branch at version '${bv:-unknown}' != $version — inspect and delete it first"
              return 1
          fi
      
          local wt
          wt=$(fr_worktree_path "$dir" "$version")
          if [[ -e "$wt" ]]; then
              fr_resume_worktree "$repo" "$wt" "$branch" "$version" "$default" "$gitdir" || return 1
          else
              git -C "$gitdir" worktree add "$wt" -b "$branch" "origin/$default" \
                  || { echo "FAIL $repo: worktree add"; return 1; }
          fi
      
          if [[ "$(fr_show_version "$gitdir" "$branch" ".claude-plugin/plugin.json" 2> /dev/null)" != "$version" ]]; then
              local bump_tool roll_tool
              bump_tool=$(fr_tool bump-version.sh) || { echo "FAIL $repo: bump-version.sh not found"; return 1; }
              bash "$bump_tool" --repo "$wt" "$version" --apply \
                  || { echo "FAIL $repo: bump"; return 1; }
              if [[ -f "$wt/CHANGELOG.md" ]]; then
                  roll_tool=$(fr_tool roll-changelog.py) || { echo "FAIL $repo: roll-changelog.py not found"; return 1; }
                  python3 "$roll_tool" "$wt/CHANGELOG.md" "$version" \
                      || { echo "FAIL $repo: changelog roll"; return 1; }
                  fr_lint_changelog "$wt"
              fi
              fr_commit_allowlisted "$repo" "$wt" "$version" || return 1
          else
              echo "bump already committed on $branch (crashed before push)"
          fi
      
          # Also here, not only inside fr_commit_allowlisted: the branch above is the
          # RESUME path, where the commit was written by an earlier run and this one
          # never went through the commit function. A commit whose trailers failed the
          # assertion leaves exactly that state behind — clean worktree, commit in
          # place — so without this the next run would push the very commit the
          # previous run refused.
          fr_assert_trailers_landed "$repo" "$wt" || return 1
      
          ( cd "$wt" && git push -u origin "$branch" )
          local ec=$?
          echo "PUSH EXIT: $ec"
          [[ "$ec" -eq 0 ]] || { echo "FAIL $repo: push"; return 1; }
      
          fr_create_and_record "$repo" "$branch" "$version" "$body" "$default" || return 1
          echo "OK $repo v$version bump PR open"
      }
      
      fr_resume_worktree() { # REPO WT BRANCH VERSION DEFAULT GITDIR — reuse only what
          # is provably this sweep's own half-done work; anything else is a loud stop
          local repo="$1" wt="$2" branch="$3" version="$4" default="$5" gitdir="$6"
          local head cur
          head=$(git -C "$wt" rev-parse --abbrev-ref HEAD 2> /dev/null || echo "")
          if [[ "$head" != "$branch" ]]; then
              echo "FAIL $repo: $wt exists on branch '${head:-?}' (expected $branch) — not this sweep's worktree, inspect manually"
              return 1
          fi
          if [[ -n "$(git -C "$wt" status --porcelain)" ]]; then
              echo "FAIL $repo: $wt is dirty — a half-written bump; inspect, then 'git -C $gitdir worktree remove --force $wt'"
              return 1
          fi
          cur=$(jq -r '.version // empty' "$wt/.claude-plugin/plugin.json" 2> /dev/null || echo "")
          if [[ "$cur" == "$version" ]]; then
              echo "reusing $wt: bump already committed"
              return 0
          fi
          if [[ "$(git -C "$wt" rev-parse HEAD)" == "$(git -C "$gitdir" rev-parse "origin/$default")" ]]; then
              echo "reusing $wt: fresh worktree, bump not yet applied"
              return 0
          fi
          echo "FAIL $repo: $wt is clean but at version '$cur' on a stale base — inspect, then remove it"
          return 1
      }
      
      fr_lint_changelog() { # WT — the roll is generated text and generated text is
          # what nobody proofreads; 16 of 63 repos went red on MD022/MD032 once.
          # Advisory when no runner exists (the roll itself emits clean structure).
          # Only an ALREADY-INSTALLED binary is used — an on-demand `npx --yes`
          # would download and execute an unpinned package mid-sweep.
          local wt="$1"
          if command -v markdownlint-cli2 > /dev/null 2>&1; then
              ( cd "$wt" && markdownlint-cli2 CHANGELOG.md ) \
                  && echo "markdownlint: CHANGELOG.md clean" \
                  || echo "WARN: markdownlint flagged CHANGELOG.md — fix before merge goes red"
          else
              echo "markdownlint-cli2 not installed — structural guarantees from roll-changelog.py only"
          fi
      }
      
      fr_stage_allowlisted() { # REPO WT — stage every pending path, refusing anything
          # outside the version surfaces; -z parsing (paths with spaces exist in the
          # fleet), no add -A ever
          local repo="$1" wt="$2"
          local entry status path
          local paths=()
          while IFS= read -r -d '' entry; do
              [[ -n "$entry" ]] || continue
              status="${entry:0:2}"
              path="${entry:3}"
              case "$status" in
                  R*|C*)
                      echo "FAIL $repo: unexpected rename/copy in bump tree: $entry"
                      return 1
                      ;;
              esac
              if ! grep -qE "$FR_ALLOWLIST_RE" <<< "$path"; then
                  echo "FAIL $repo: unexpected change outside the version surfaces: $path"
                  return 1
              fi
              paths+=("$path")
          done < <(git -C "$wt" status --porcelain -z)
          [[ "${#paths[@]}" -gt 0 ]] || { echo "FAIL $repo: nothing to commit after bump"; return 1; }
          git -C "$wt" add -- "${paths[@]}"
      }
      
      fr_trailer_args() { # ARRAY_NAME — fill with --trailer args from FR_COMMIT_TRAILERS
          # Disclosure trailers (provenance, session, host) are the operator's to
          # define, not this driver's: it must not know what `Assisted-by` means, so
          # the caller supplies whole `Key: value` lines, one per line, and each
          # becomes a `git commit --trailer`. Blank lines are skipped so a
          # here-doc-assigned variable with a trailing newline behaves.
          # `--trailer` appends into the message's own trailer block, so it composes
          # with the `--signoff` above rather than competing with it, and git dedupes
          # nothing — a line already present in the message would appear twice, which
          # is why this only ever runs against the generated one-line subject.
          local -n _out="$1"
          _out=()
          [[ -n "${FR_COMMIT_TRAILERS:-}" ]] || return 0
          local line
          while IFS= read -r line; do
              [[ -n "${line//[[:space:]]/}" ]] || continue
              _out+=(--trailer "$line")
          done <<< "$FR_COMMIT_TRAILERS"
      }
      
      fr_assert_trailers_landed() { # REPO WT — every requested trailer is in the commit
          # A commit that exits 0 is not a commit that carries what was asked for.
          # The failure this exists for: a driver copy that predates FR_COMMIT_TRAILERS
          # ignores the variable entirely and commits happily, so the operator sets it,
          # reads "OK <repo> vX.Y.Z bump PR open" per repo, and finds out only if they
          # open a commit. On 2026-09-11 that shipped four bump commits to `main`
          # without their disclosure trailers before anyone looked; auto-merge had
          # taken them past the point where an amend was possible.
          #
          # A stale engine cannot warn about itself — it does not know the variable —
          # so this check cannot catch that case in the copy that has the bug. It
          # catches every other producer of the same symptom from here on: a
          # prepare-commit-msg hook that rewrites the message, a git too old for
          # --trailer, a trailer whose key git declines to append.
          local repo="$1" wt="$2" line key
          [[ -n "${FR_COMMIT_TRAILERS:-}" ]] || return 0
          local msg
          # `git interpret-trailers --parse`, not the raw %B: a hook that moves a line
          # into the message BODY leaves it findable by a plain grep while `git log
          # --grep` and every other trailer consumer no longer see it as a trailer —
          # which is the whole point of writing one. Compared with `-x` so a line is
          # matched whole; a substring match would accept a longer line that merely
          # contains the requested one.
          msg=$(git -C "$wt" log -1 --format=%B | git interpret-trailers --parse)
          while IFS= read -r line; do
              [[ -n "${line//[[:space:]]/}" ]] || continue
              key=${line%%:*}
              if ! grep -qFx -- "$line" <<< "$msg"; then
                  echo "FAIL $repo: requested trailer did not land in the commit: ${key}:"
                  echo "  set FR_COMMIT_TRAILERS but the message has no such line — is this"
                  echo "  driver copy current? git -C <skill-repo-skill> pull, then re-run."
                  return 1
              fi
          done <<< "$FR_COMMIT_TRAILERS"
          return 0
      }
      
      fr_commit_allowlisted() { # REPO WT VERSION — only the four version surfaces may
          # move. Two attempts: a reformatting hook (pre-commit's pretty-format-json
          # --autofix, black, ruff format, …) rewrites a staged file and FAILS the
          # commit, and re-staging its output is the documented remedy — without it
          # a bump dies as "FAIL <repo>: commit" and leaves a half-written worktree
          # the resume then refuses as dirty (#263: it-maintenance-skill v1.15.0 and
          # netresearch-jira-skill v2.11.1 in the 2026-08-28 sweep, where the hook
          # sorts the keys sync-plugin-manifest.sh emits unsorted). The retry re-walks
          # the porcelain instead of re-adding the first attempt's paths, so a hook
          # that writes OUTSIDE the version surfaces still fails the allowlist. A
          # deterministic rejection fails the second attempt too and reports as before.
          local repo="$1" wt="$2" version="$3"
          local attempt ec
          # Expanded below as ${trailer_args[@]+"${trailer_args[@]}"}: expanding an
          # EMPTY array as "${a[@]}" under `set -u` is an error before bash 4.4, and
          # fr_preflight admits bash >= 4 — so the plain form would break every bump
          # on exactly the older shells that guard exists for, and only for operators
          # who did NOT set FR_COMMIT_TRAILERS.
          local -a trailer_args=()
          fr_trailer_args trailer_args
          for attempt in 1 2; do
              echo "--- porcelain (attempt $attempt):"
              git -C "$wt" status --porcelain
              fr_stage_allowlisted "$repo" "$wt" || return 1
              ( cd "$wt" && git commit -S --signoff ${trailer_args[@]+"${trailer_args[@]}"} -m "${FR_COMMIT_PREFIX}${version}" )
              ec=$?
              echo "COMMIT EXIT ($attempt): $ec"   # bare exit code — a pipe here once hid a hook abort
              if [[ "$ec" -eq 0 ]]; then
                  fr_assert_trailers_landed "$repo" "$wt" || return 1
                  return 0
              fi
              # An `if`, not `cond && echo`: the latter is the loop body's last
              # command and returns 1 on the second pass, which would make the
              # function's own status depend on how the caller suppresses `set -e`.
              if [[ "$attempt" -eq 1 ]]; then
                  echo "commit failed — retrying once, in case a hook rewrote a staged file"
              fi
          done
          echo "FAIL $repo: commit"
          return 1
      }
      
      fr_create_and_record() { # REPO BRANCH VERSION BODY TARGET
          local repo="$1" branch="$2" version="$3" body="$4" target="$5" out prid prurl
          out=$(host_create_pr "$repo" "$branch" "${FR_COMMIT_PREFIX}${version}" "$body" "$target" < /dev/null) \
              || { echo "FAIL $repo: PR/MR create: $out"; return 1; }
          read -r prid prurl <<< "$out"
          fr_record_opened "$repo" "$prid" "$prurl" "$branch"
          echo "PR: $prurl"
          # Arm failure is not repo failure — finish merges plainly when green.
          host_arm_automerge "$repo" "$prid" < /dev/null || true
      }
      
      fr_bump() { # PLAN
          local plan="$1" logdir="$FR_WORKDIR/logs/bump" row repo n=0 rc=0
          fr_plan_validate "$plan"
          # Logs are per-RUN evidence: stale ones from a previous run would fail the
          # count invariant and resurrect old FAIL lines into this run's summary.
          find "$logdir" -maxdepth 1 -name '*.log' -delete 2> /dev/null || true
          while IFS= read -u 3 -r row || [[ -n "$row" ]]; do
              [[ -n "$row" ]] || continue
              repo=$(jq -r '.repo' <<< "$row")
              n=$((n + 1))
              fr_note "bump: $repo ..."
              # Subshell, never a brace group: a per-repo exit inside { } > log kills
              # the whole driver and the batch "completes" with missing logs.
              ( fr_bump_one "$row" ) > "$logdir/$repo.log" 2>&1 || rc=1
              tail -1 "$logdir/$repo.log" 2> /dev/null || rc=1
          done 3< "$plan"
          fr_check_log_invariant "$logdir" "$n" || rc=1
          fr_summarize_logs "$logdir" || rc=1
          fr_note ""
          fr_note "opened PRs recorded in $FR_WORKDIR/opened.jsonl"
          fr_note "next: $FR_DRIVER finish --workdir $FR_WORKDIR"
          return "$rc"
      }
      
      # ---------------------------------------------------------------------------
      # Finish phase
      # ---------------------------------------------------------------------------
      fr_finish_one() { # ROW — exit 0 ok, 1 pre-tag failure (loop continues),
          # 2 post-tag failure (systemic: the release machinery itself broke — the
          # loop HALTS, because "halt all further releases if this one fails" is a
          # numbered step of the spec's execution order)
          local row="$1"
          local repo version default cls tag_override
          repo=$(jq -r '.repo' <<< "$row")
          version=$(jq -r '.version // ""' <<< "$row")
          default=$(jq -r '.default // "main"' <<< "$row")
          cls=$(jq -r '.classification // "BUMP"' <<< "$row")
          tag_override=$(jq -r '.tag // ""' <<< "$row")
          local tag="${tag_override:-v$version}"
          local branch="${FR_BRANCH_PREFIX}${version}"
          local head
          head=$(jq -r '.head // ""' <<< "$row")
          echo "=== $(host_display "$repo") $tag ($cls) ==="
          [[ -n "$version" ]] || { echo "FAIL $repo: plan row has no version"; return 1; }
          case "$cls" in
              BUMP|FIRST-RELEASE|TAG-ONLY|OWN-PR-OPEN) ;;
              *)
                  echo "FAIL $repo: classification $cls does not belong in the finish phase — resolve it first"
                  return 1
                  ;;
          esac
      
          local dir="$FR_BASE_DIR/$repo" gitdir
          gitdir=$(fr_resolve_gitdir "$dir") || { echo "FAIL $repo: no checkout at $dir"; return 1; }
          fr_check_origin "$gitdir" "$(host_origin_suffix "$repo")" || { echo "FAIL $repo: origin mismatch"; return 1; }
      
          # Idempotent re-run: a tag already on the remote means this repo released
          # (or a human tagged it) — verify the Release and stop, instead of failing
          # on "tip moved past the tag" after later commits landed on the default
          # branch.
          local pre_rc=0 pre_existing
          pre_existing=$(fr_remote_tag_commit "$gitdir" "$tag") || pre_rc=$?
          if [[ "$pre_rc" -eq 2 ]]; then
              echo "FAIL $repo: cannot read remote tags — not proceeding blind"
              return 1
          fi
          if [[ -n "$pre_existing" ]]; then
              echo "tag $tag already on the remote at $pre_existing — verifying its Release"
              host_release_verify "$repo" "$tag" < /dev/null \
                  || { echo "FAIL $repo: tag $tag exists but has no verified Release (post-tag — halting the sweep)"; return 2; }
              fr_cleanup_release_branch "$dir" "$gitdir" "$branch" "$version" "$(fr_opened_lookup "$repo")"
              echo "OK $repo $tag released (verified on re-run)"
              return 0
          fi
      
          local prid
          prid=$(fr_opened_lookup "$repo")
          if [[ -z "$prid" ]]; then
              # No PR recorded: legitimate for TAG-ONLY / already-at-version repos.
              # An open release PR by SELF from a crashed earlier workdir is adopted;
              # anyone else's PR is never touched.
              local pr prurl prstate prauthor
              pr=$(host_find_release_pr "$repo" "$branch" < /dev/null || true)
              if [[ -n "$pr" ]]; then
                  read -r prid prurl prstate prauthor <<< "$pr"
                  case "$prstate" in
                      open|OPEN|opened)
                          if [[ "$prauthor" == "$FR_SELF_LOGIN" ]]; then
                              echo "adopting own open PR $prurl (crashed earlier sweep)"
                              fr_record_opened "$repo" "$prid" "$prurl" "$branch"
                          else
                              echo "FAIL $repo: open release PR $prurl by @$prauthor — needs their go-ahead, not this sweep's"
                              return 1
                          fi
                          ;;
                      *) prid="" ;;
                  esac
              fi
          fi
          local target i
          if [[ -n "$prid" ]]; then
              host_merge_gate "$repo" "$prid" < /dev/null || { echo "FAIL $repo: merge gate"; return 1; }
              echo "merged."
              # Tag the MERGE COMMIT, not whatever the tip is by now: between merge
              # and tag a colleague's push can land, and tagging the live tip would
              # release unsurveyed commits. The merge commit is exactly what the
              # sweep's approval covered.
              target=$(host_merge_commit "$repo" "$prid" < /dev/null || true)
              [[ -n "$target" ]] || { echo "FAIL $repo: cannot resolve the merge commit of the bump PR"; return 1; }
              for i in 1 2 3 4 5; do
                  fr_fetch_branches "$gitdir" > /dev/null
                  git -C "$gitdir" cat-file -e "$target^{commit}" 2> /dev/null && break
                  [[ "$i" -lt 5 ]] && sleep "$FR_POLL_SECONDS"
              done
              git -C "$gitdir" cat-file -e "$target^{commit}" 2> /dev/null \
                  || { echo "FAIL $repo: merge commit $target never became fetchable"; return 1; }
              git -C "$gitdir" merge-base --is-ancestor "$target" "origin/$default" \
                  || { echo "FAIL $repo: merge commit $target is not on origin/$default"; return 1; }
          else
              echo "no bump PR for this repo (${cls}) — tagging the committed version directly"
              target=$(host_remote_tip "$repo" "$default" < /dev/null) \
                  || { echo "FAIL $repo: cannot read remote tip"; return 1; }
              for i in 1 2 3 4 5; do
                  fr_fetch_branches "$gitdir" > /dev/null
                  [[ "$(git -C "$gitdir" rev-parse "origin/$default")" == "$target" ]] && break
                  [[ "$i" -lt 5 ]] && sleep "$FR_POLL_SECONDS"
              done
              [[ "$(git -C "$gitdir" rev-parse "origin/$default")" == "$target" ]] \
                  || { echo "FAIL $repo: origin/$default never reached the remote tip $target"; return 1; }
              # Freshness: the sweep's approval covered the SURVEYED tip. Commits
              # since then are acceptable only if they touch nothing but the version
              # surfaces (i.e. the bump itself landed in between) — anything else on
              # the branch was never surveyed and must not ride into this tag.
              if [[ -n "$head" && "$target" != "$head" ]]; then
                  local drift
                  if ! drift=$(git -C "$gitdir" diff --name-only "$head" "$target" 2> /dev/null); then
                      echo "FAIL $repo: surveyed head $head is unknown here (force-push?) — re-run survey"
                      return 1
                  fi
                  drift=$(grep -vE "$FR_ALLOWLIST_RE" <<< "$drift" || true)
                  if [[ -n "$drift" ]]; then
                      echo "FAIL $repo: $default moved past the surveyed head with non-bump changes ($(tr '\n' ' ' <<< "$drift")) — re-run survey and re-plan"
                      return 1
                  fi
                  echo "tip moved past the surveyed head by version-surface changes only — acceptable"
              elif [[ -z "$head" ]]; then
                  echo "WARN: plan row carries no surveyed head — tagging the current tip unverified against the survey"
              fi
          fi
      
          fr_parity_from_ref "$gitdir" "$target" "$version" \
              || { echo "FAIL $repo: version parity at $target"; return 1; }
          echo "parity OK ($version at $target)"
      
          fr_tag_on_tip "$gitdir" "$tag" "$target" || { echo "FAIL $repo: tag"; return 1; }
      
          # Past this point a failure means the tag is on the remote but the release
          # machinery did not produce a verified Release object — that class is
          # contagious (a broken shared workflow breaks every following repo too).
          host_release_verify "$repo" "$tag" < /dev/null \
              || { echo "FAIL $repo: no verified Release for $tag (post-tag — halting the sweep)"; return 2; }
      
          fr_cleanup_release_branch "$dir" "$gitdir" "$branch" "$version" "$prid"
          echo "OK $repo $tag released"
      }
      
      fr_cleanup_release_branch() { # DIR GITDIR BRANCH VERSION PRID — tolerant;
          # the REMOTE branch is deleted only when this sweep owns a PR for it —
          # deleting a branch the sweep never created is not cleanup, it is damage
          local dir="$1" gitdir="$2" branch="$3" version="$4" prid="$5" wt
          wt=$(fr_worktree_path "$dir" "$version")
          if [[ -e "$wt" ]]; then
              git -C "$gitdir" worktree remove "$wt" 2> /dev/null && echo "worktree removed"
          fi
          git -C "$gitdir" branch -D "$branch" > /dev/null 2>&1 || true
          if [[ -n "$prid" ]]; then
              git -C "$gitdir" push origin --delete "$branch" 2> /dev/null \
                  && echo "remote branch deleted" || echo "remote branch already gone"
          fi
      }
      
      fr_finish() { # PLAN [--continue]
          local plan="$1" cont="${2:-}" logdir="$FR_WORKDIR/logs/finish" row repo n=0 rc=0 ec
          fr_plan_validate "$plan"
          find "$logdir" -maxdepth 1 -name '*.log' -delete 2> /dev/null || true
          while IFS= read -u 3 -r row || [[ -n "$row" ]]; do
              [[ -n "$row" ]] || continue
              repo=$(jq -r '.repo' <<< "$row")
              n=$((n + 1))
              fr_note "finish: $repo ..."
              ec=0
              # Tested context on purpose: it suppresses errexit inside the subshell
              # so the explicit FAIL/return handling governs, not set -e.
              ( fr_finish_one "$row" ) > "$logdir/$repo.log" 2>&1 || ec=$?
              tail -1 "$logdir/$repo.log" 2> /dev/null || rc=1
              if [[ "$ec" -eq 2 && "$cont" != "--continue" ]]; then
                  rc=1
                  fr_err "post-tag failure in $repo — halting (release-discipline: halt all further releases if one fails). Diagnose, then re-run finish; --continue overrides."
                  break
              fi
              [[ "$ec" -eq 0 ]] || rc=1
          done 3< "$plan"
          fr_summarize_logs "$logdir" || rc=1
          fr_note ""
          fr_note "logs: $logdir"
          return "$rc"
      }
      
      # ---------------------------------------------------------------------------
      # Repo-list resolution
      # ---------------------------------------------------------------------------
      fr_repos_from() { # REPOS_INLINE REPOS_FILE DEFAULT_FILE — echoes the list;
          # '#' comment lines and blanks in files are skipped
          local inline="$1" file="$2" fallback="$3"
          if [[ -n "$inline" ]]; then
              printf '%s\n' "$inline" | tr ' ' '\n' | grep -v '^$' || true
          elif [[ -n "$file" ]]; then
              grep -vE '^\s*(#|$)' "$file" || true
          elif [[ -f "$fallback" ]]; then
              grep -vE '^\s*(#|$)' "$fallback" || true
          fi
      }
      
    • fleet-release-github.sh 18 KB
      #!/usr/bin/env bash
      #
      # fleet-release-github.sh — fleet release driver for the PUBLIC GitHub skill
      # repos (github.com/<org>, default netresearch).
      #
      # Mechanizes references/release-discipline.md so a sweep stops hand-writing
      # driver scripts and re-earning their bugs (issue #218). Fleets on other hosts
      # ship their own driver in their own (private) infrastructure — the API shapes
      # differ enough that one script serving several hosts is how jq filters
      # silently return empty. Everything host-independent lives in
      # fleet-release-common.sh; a foreign-host driver vendors that engine and
      # defines the host_* callbacks documented in its header.
      #
      # Usage:
      #   fleet-release-github.sh survey   --workdir DIR [--repos "r1 r2"|--repos-file F] [--org ORG]
      #   fleet-release-github.sh manifest --workdir DIR
      #   fleet-release-github.sh bump     --workdir DIR [--plan F]
      #   fleet-release-github.sh finish   --workdir DIR [--plan F] [--continue]
      #
      # Flow: survey (read-only API) -> manifest (offline classification; STOPS for
      # human approval) -> bump (worktree, bump-version.sh, changelog roll, PR,
      # auto-merge arm) -> finish (merge gate, parity from origin/<default>, signed
      # tag on the remote tip, Release verification, cleanup). Re-running any phase
      # is resumable: remote reality, not local state, decides what is left to do.
      #
      # The driver never merges a PR this sweep did not open, never unarchives,
      # never clones, never invents versions or PR bodies, never touches CI files,
      # and only releases the default-branch tip.
      #
      # Environment:
      #   FR_COMMIT_TRAILERS  Newline-separated `Key: value` lines appended to every
      #                       bump commit as git trailers (requires git >= 2.32).
      #                       Where a policy requires agent/tool disclosure on every
      #                       commit, set it — the driver appends the lines verbatim
      #                       and interprets none of them, so which keys a fleet uses
      #                       stays the operator's decision:
      #                         export FR_COMMIT_TRAILERS='Assisted-by: <agent>:<model>
      #                         Agent-Session: <url>'
      #                       Unset, bump commits keep `--signoff` alone.
      
      set -euo pipefail
      
      SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
      # shellcheck source=fleet-release-common.sh
      # shellcheck disable=SC1091  # resolved at runtime relative to this script
      source "$SCRIPT_DIR/fleet-release-common.sh"
      
      FR_HOST=github
      # shellcheck disable=SC2034  # consumed by the sourced engine's epilogues
      FR_DRIVER="$0"
      FR_ORG="${FR_ORG:-netresearch}"
      
      usage() {
          sed -n '2,32p' "$0" | sed 's/^# \{0,1\}//'
          exit "${1:-1}"
      }
      
      # ---------------------------------------------------------------------------
      # Host callbacks (see the contract in fleet-release-common.sh)
      # ---------------------------------------------------------------------------
      host_display() { echo "$FR_ORG/$1"; }
      host_origin_suffix() { echo "github.com/$FR_ORG/$1"; }
      
      host_remote_tip() { # REPO BRANCH
          gh api "repos/$FR_ORG/$1/commits/$2" --jq .sha
      }
      
      gh_json() { # ARGS... — gh api that NEVER leaks an HTTP error body as data:
          # on 403/404/5xx gh prints the JSON error to STDOUT (verified with gh
          # 2.97), so every `$(gh api ... || echo "")` capture is poisoned. Output
          # only reaches the caller when the call succeeded.
          local out
          if out=$(gh api "$@" 2> /dev/null); then
              printf '%s\n' "$out"
              return 0
          fi
          return 1
      }
      
      fr_resolve_self_login() { # set FR_SELF_LOGIN, or die — never leave it poisoned.
          # `gh api user --jq .login` applies the filter to the ERROR body too, so an
          # expired token yields the string "null". The author gate then compares
          # every PR author against "null" and classifies the sweep's own PRs as
          # BLOCKED-FOREIGN-PR. Failing here says why; a statement, not a $( ), so
          # fr_die ends the run instead of a subshell.
          [[ -n "$FR_SELF_LOGIN" ]] && return 0
          FR_SELF_LOGIN=$(gh_json user | jq -r '.login // empty') || FR_SELF_LOGIN=""
          [[ -n "$FR_SELF_LOGIN" ]] \
              || fr_die "cannot resolve the authenticated gh login — check 'gh auth status'"
      }
      
      host_survey_repo() { # REPO — one normalized row on stdout, or exit 1
          local r="$1" meta resolved archived def
          meta=$(gh_json "repos/$FR_ORG/$r") || return 1
          resolved=$(jq -r '.full_name' <<< "$meta")
          archived=$(jq -r '.archived' <<< "$meta")
          def=$(jq -r '.default_branch' <<< "$meta")
      
          local head
          head=$(gh_json "repos/$FR_ORG/$r/commits/$def" | jq -r '.sha // empty' || echo "")
      
          local rel last_tag tags
          rel=$(gh_json "repos/$FR_ORG/$r/releases/latest" | jq -r '.tag_name // empty' || echo "")
          # releases/latest 404s for never-released AND for tag-without-Release repos
          # (a release workflow that ran and died). The tags list tells them apart —
          # and it is NOT date-ordered, so pick the max by version, never .[0].
          # --paginate: one page would hide the newest tag of a >100-tag repo.
          # Same poisoned-capture class as host_release_verify: on 403/404 gh prints
          # the error body to STDOUT, and `[.[][]]` happily flattens {"message":…}
          # into ["Not Found", …]. The map(.name) below then fails, last_tag comes
          # back empty, and a rate-limited repo classifies FIRST-RELEASE — the survey
          # would offer to tag v1.0.0 over an existing history.
          local tags_failed=false
          if tags=$(gh_json --paginate "repos/$FR_ORG/$r/tags?per_page=100"); then
              tags=$(jq -s '[.[][]]' <<< "$tags" 2> /dev/null) || { tags='[]'; tags_failed=true; }
          else
              tags='[]'
              tags_failed=true
          fi
          # A failed query is NOT an empty tag list. Conflating them is how a
          # rate-limited repo reaches FIRST-RELEASE, so the failure travels in the row
          # and fr_classify turns it into SURVEY-INCOMPLETE — the same convention the
          # negative ahead/nonci sentinels already use for a failed compare.
          last_tag=$(jq -r "$FR_JQ_VPARSE"' map(.name) | max_by(vparse) // empty' <<< "$tags")
      
          local rootv cpv composer_has_version release_gate
          rootv=$(gh_json -H "Accept: application/vnd.github.raw" \
              "repos/$FR_ORG/$r/contents/plugin.json?ref=$def" \
              | jq -r '.version // empty' 2> /dev/null || echo "")
          cpv=$(gh_json -H "Accept: application/vnd.github.raw" \
              "repos/$FR_ORG/$r/contents/.claude-plugin/plugin.json?ref=$def" \
              | jq -r '.version // empty' 2> /dev/null || echo "")
          # bump-version.sh refuses a composer version field mid-bump with a
          # half-written tree — surface it as a manifest row instead.
          if gh_json -H "Accept: application/vnd.github.raw" \
              "repos/$FR_ORG/$r/contents/composer.json?ref=$def" \
              | jq -e 'has("version")' > /dev/null 2>&1; then
              composer_has_version=true
          else
              composer_has_version=false
          fi
          if gh_json "repos/$FR_ORG/$r/contents/.github/workflows/release.yml?ref=$def" \
              > /dev/null; then
              release_gate=true
          else
              release_gate=false
          fi
      
          # Unreleased-section state, by the same rule the bump-time roll enforces —
          # an empty section becomes a manifest warning instead of a mid-bump failure.
          local cl_raw cl_state
          cl_raw=$(gh_json -H "Accept: application/vnd.github.raw" \
              "repos/$FR_ORG/$r/contents/CHANGELOG.md?ref=$def" || echo "")
          cl_state=$(fr_changelog_state "$cl_raw")
      
          # Measured CI state of the default branch — the manifest's precondition
          # column must show a value, never an assumed checkmark.
          local ci_status
          ci_status=$(gh_json "repos/$FR_ORG/$r/commits/$def/check-runs?per_page=100" \
              | jq -r '[.check_runs[]] as $r
                  | ([$r[] | select(.status != "completed")] | length) as $pending
                  | ([$r[] | select(.status == "completed"
                      and (.conclusion | IN("success","skipped","neutral") | not)) | .name]) as $failed
                  | if ($r | length) == 0 then "none"
                    elif .total_count > 100 then "truncated(\(.total_count) runs)"
                    elif ($failed | length) > 0 then "failed: " + ($failed | join(","))
                    elif $pending > 0 then "pending(\($pending))"
                    else "success" end' 2> /dev/null || echo "unknown")
      
          # --limit 100: the default of 30 would let a foreign release PR beyond the
          # first page slip past the BLOCKED-FOREIGN-PR gate.
          local prs
          prs=$(gh pr list --repo "$FR_ORG/$r" --state open --limit 100 \
              --json number,title,author,headRefName,url 2> /dev/null \
              | jq '[.[] | select((.headRefName | startswith("release/"))
                              or (.title | startswith("chore(release)"))
                              or (.title | startswith("chore: release")))
                     | {id: (.number | tostring), url, author: .author.login,
                        branch: .headRefName, title}]' || echo '[]')
      
          local base ahead=-1 nonci=-1 subjects='[]' files='[]' files_truncated=false cmp
          base="$rel"
          [[ -n "$base" ]] || base="$last_tag"
          if [[ -z "$base" ]]; then
              ahead=0   # first release: there is no base, and -1 must mean only
              nonci=0   # "measurement FAILED" (SURVEY-INCOMPLETE), never "no base"
          else
              # compare runs for archived repos too: the unarchive decision needs to
              # see whether anything (e.g. a deprecation banner) is unshipped.
              # A FAILED call keeps the -1 sentinel -> SURVEY-INCOMPLETE, because a
              # transient 403 must never read as "no delta".
              cmp=$(gh_json "repos/$FR_ORG/$r/compare/$base...$def" || echo "")
              if [[ -n "$cmp" ]]; then
                  ahead=$(jq -r '.ahead_by // 0' <<< "$cmp")
                  nonci=$(jq --arg f "$FR_CI_ONLY_RE" \
                      '[.files[].filename] | map(select(test($f) | not)) | length' <<< "$cmp" 2> /dev/null || echo -1)
                  subjects=$(jq '[.commits[].commit.message | split("\n")[0]]' <<< "$cmp" 2> /dev/null || echo '[]')
                  files=$(jq --arg f "$FR_CI_ONLY_RE" \
                      '[.files[].filename | select(test($f) | not)]' <<< "$cmp" 2> /dev/null || echo '[]')
                  # the compare API caps .files at 300 — 0-vs-nonzero stays valid,
                  # the file list itself may be partial
                  [[ "$(jq '.files | length' <<< "$cmp")" -ge 300 ]] && files_truncated=true
              fi
          fi
      
          jq -nc \
              --arg host "$FR_HOST" --arg repo "$r" --arg resolved "$resolved" \
              --arg def "$def" --arg head "$head" --argjson archived "$archived" \
              --arg rel "$rel" --arg last_tag "$last_tag" \
              --arg rootv "$rootv" --arg cpv "$cpv" \
              --argjson chv "$composer_has_version" --argjson gate "$release_gate" \
              --arg ci "$ci_status" --argjson ahead "$ahead" --argjson nonci "$nonci" \
              --argjson subjects "$subjects" --argjson files "$files" \
              --argjson trunc "$files_truncated" --argjson prs "$prs" \
              --arg clstate "$cl_state" --argjson tagsfailed "$tags_failed" \
              '{host: $host, repo: $repo, resolved: $resolved, default: $def, head: $head,
                archived: $archived, empty: false, unreachable: false, duplicate_of: "",
                last_release: $rel, last_tag: $last_tag,
                root_plugin: $rootv, claude_plugin: $cpv,
                composer_has_version: $chv, release_gate: $gate, ci_status: $ci,
                ahead: $ahead, nonci: $nonci, files_truncated: $trunc,
                tags_failed: $tagsfailed,
                changelog_unreleased: $clstate,
                subjects: $subjects, files: $files,
                open_release_prs: $prs, ci_tag_rules: "", notes: ""}'
      }
      
      host_create_pr() { # REPO BRANCH TITLE BODY TARGET -> "<number> <url>"
          local r="$1" branch="$2" title="$3" body="$4" target="$5" url num
          url=$(gh pr create --repo "$FR_ORG/$r" --head "$branch" --base "$target" \
              --title "$title" --body "$body") || return 1
          num=$(gh pr view "$branch" --repo "$FR_ORG/$r" --json number --jq .number) || return 1
          echo "$num $url"
      }
      
      host_find_release_pr() { # REPO BRANCH -> "<number> <url> <state> <author>"
          gh pr list --repo "$FR_ORG/$1" --head "$2" --state all \
              --json number,url,state,author --limit 1 \
              --jq '.[0] | select(. != null)
                    | "\(.number) \(.url) \(.state | ascii_downcase) \(.author.login)"'
      }
      
      host_arm_automerge() { # REPO NUMBER
          if gh pr merge --auto --merge "$2" --repo "$FR_ORG/$1" 2>&1; then
              echo "auto-merge armed"
          else
              echo "auto-merge not armed (finish will merge plainly when green)"
          fi
      }
      
      host_merge_gate() { # REPO NUMBER — completion-first: a QUEUED check reports
          # conclusion "" (not null) and must read as pending, never as failure
          local r="$1" num="$2" deadline j state mss failed pending automerge merr
          deadline=$(($(date +%s) + FR_MERGE_TIMEOUT))
          while true; do
              j=$(gh pr view "$num" --repo "$FR_ORG/$r" \
                  --json state,mergeStateStatus,statusCheckRollup,autoMergeRequest 2> /dev/null) || j=""
              if [[ -z "$j" ]]; then
                  echo "cannot read PR #$num"
                  return 1
              fi
              state=$(jq -r '.state' <<< "$j")
              [[ "$state" == "MERGED" ]] && return 0
              if [[ "$state" == "CLOSED" ]]; then
                  echo "PR #$num was closed without merge"
                  return 1
              fi
              mss=$(jq -r '.mergeStateStatus // "UNKNOWN"' <<< "$j")
              failed=$(jq '[.statusCheckRollup[]
                  | select(.__typename == "CheckRun" and .status == "COMPLETED"
                      and (.conclusion | IN("SUCCESS","SKIPPED","NEUTRAL") | not))
                  | .name]
                  + [.statusCheckRollup[]
                  | select(.__typename == "StatusContext"
                      and (.state | IN("SUCCESS","PENDING","EXPECTED") | not))
                  | .context]' <<< "$j")
              if [[ "$(jq length <<< "$failed")" -gt 0 ]]; then
                  echo "red checks on PR #$num: $(jq -c . <<< "$failed") — PR left open"
                  return 1
              fi
              pending=$(jq '[.statusCheckRollup[]
                  | select((.__typename == "CheckRun" and .status != "COMPLETED")
                        or (.__typename == "StatusContext" and (.state | IN("PENDING","EXPECTED"))))]
                  | length' <<< "$j")
              automerge=$(jq -r '.autoMergeRequest // empty | tostring' <<< "$j")
              echo "poll: state=$state mergeState=$mss pending=$pending automerge=${automerge:+armed}${automerge:-no}"
              # Merge plainly ONLY on GitHub's own CLEAN verdict — an empty or
              # not-yet-reported rollup is NOT green (queued workflows are invisible
              # to the rollup; pending==0 never meant "all reported").
              if [[ "$mss" == "CLEAN" && -z "$automerge" ]] \
                  && ! merr=$(gh pr merge --merge "$num" --repo "$FR_ORG/$r" 2>&1); then
                  echo "$merr"
                  case "$merr" in
                      *"not allowed"*|*"protected branch"*)
                          echo "merge method rejected for PR #$num — needs a human (merge it per repo policy)"
                          return 1
                          ;;
                      *) : ;;
                  esac
              fi
              if [[ "$(date +%s)" -ge "$deadline" ]]; then
                  echo "TIMEOUT waiting for PR #$num to merge (state=$state mergeState=$mss pending=$pending)"
                  return 1
              fi
              sleep "$FR_POLL_SECONDS"
          done
      }
      
      host_merge_commit() { # REPO NUMBER — the SHA the merged PR produced
          gh pr view "$2" --repo "$FR_ORG/$1" --json mergeCommit \
              --jq '.mergeCommit.oid // empty'
      }
      
      host_release_verify() { # REPO TAG — a pushed tag is not a release; assert the
          # Release object and its assets, then link it
          local r="$1" tag="$2" deadline rel assets
          deadline=$(($(date +%s) + FR_RELEASE_TIMEOUT))
          while true; do
              # gh_json, not a bare capture: while the Release workflow is queued
              # this 404s, and gh prints the error body to STDOUT, so `|| echo ""`
              # leaves {"message":"Not Found"} in $rel — non-empty, and the jq below
              # then errors on every poll of every repo (19 batch logs, 2026-09-03).
              rel=$(gh_json "repos/$FR_ORG/$r/releases/tags/$tag") || rel=""
              if [[ -n "$rel" ]]; then
                  assets=$(jq '[.assets[].name] | length' <<< "$rel")
                  if [[ "$assets" -gt 0 ]]; then
                      echo "RELEASE: $(jq -r '.html_url' <<< "$rel") assets: $(jq -c '[.assets[].name]' <<< "$rel")"
                      return 0
                  fi
              fi
              if [[ "$(date +%s)" -ge "$deadline" ]]; then
                  echo "no Release with assets for $tag — last workflow runs:"
                  gh run list --repo "$FR_ORG/$r" --limit 3 2>&1 || true
                  return 1
              fi
              sleep "$FR_POLL_SECONDS"
          done
      }
      
      # ---------------------------------------------------------------------------
      # Dispatch
      # ---------------------------------------------------------------------------
      [[ $# -ge 1 ]] || usage
      CMD="$1"
      shift
      case "$CMD" in -h|--help) usage 0 ;; esac
      REPOS_INLINE=""
      REPOS_FILE=""
      PLAN=""
      CONTINUE=""
      FR_WORKDIR=""
      # shellcheck disable=SC2034  # FR_BASE_DIR is read by the sourced engine
      while [[ $# -gt 0 ]]; do
          case "$1" in
              --workdir)    FR_WORKDIR="${2:?}"; shift 2 ;;
              --repos)      REPOS_INLINE="${2:?}"; shift 2 ;;
              --repos-file) REPOS_FILE="${2:?}"; shift 2 ;;
              --org)        FR_ORG="${2:?}"; shift 2 ;;
              --plan)       PLAN="${2:?}"; shift 2 ;;
              --base-dir)   FR_BASE_DIR="${2:?}"; shift 2 ;;
              --continue)   CONTINUE="--continue"; shift ;;
              -h|--help)    usage 0 ;;
              *)            fr_err "unknown option: $1"; usage ;;
          esac
      done
      [[ -n "$FR_WORKDIR" ]] || { fr_err "--workdir is required"; usage; }
      mkdir -p "$FR_WORKDIR"
      PLAN="${PLAN:-$FR_WORKDIR/plan.jsonl}"
      
      case "$CMD" in
          survey)
              fr_preflight survey gh
              REPOS=$(fr_repos_from "$REPOS_INLINE" "$REPOS_FILE" "$SCRIPT_DIR/fleet-repos-github.txt")
              [[ -n "$REPOS" ]] || fr_die "no repos: pass --repos/--repos-file or ship fleet-repos-github.txt"
              # shellcheck disable=SC2086  # word splitting is the point
              fr_survey $REPOS
              ;;
          manifest)
              fr_resolve_self_login
              fr_preflight manifest gh
              fr_manifest
              ;;
          bump)
              fr_resolve_self_login
              fr_preflight bump gh
              fr_bump "$PLAN"
              ;;
          finish)
              fr_resolve_self_login
              fr_preflight finish gh
              fr_finish "$PLAN" "$CONTINUE"
              ;;
          *)
              fr_err "unknown subcommand: $CMD"
              usage
              ;;
      esac
      
    • fleet-repos-github.txt 2.3 KB
      # Default repo list for fleet-release-github.sh survey (--repos/--repos-file
      # override it). One repo per line, org-relative; '#' and blank lines ignored.
      #
      # Snapshot of the public GitHub skill fleet as of the 2026-09-17 sweep.
      # Naming conventions do NOT enumerate this fleet (plugins and oddly-named
      # repos exist), and a name filter over the org returns repos that are not
      # skills at all — composer-agent-skill-plugin and composer-patches-plugin are
      # composer plugins with no plugin.json. The marketplace manifest is what
      # defines the fleet: a skill the marketplace ships is one the fleet releases.
      # Refresh against it, and diff both ways:
      #   gh api repos/netresearch/claude-code-marketplace/contents/.claude-plugin/marketplace.json \
      #     --jq .content | base64 -d | jq -r '.plugins[].source.repo' \
      #     | sed 's|netresearch/||' | sort
      # A repo in the org but not in the manifest is either not a skill or a
      # marketplace gap; a repo in the manifest but not here is one the sweep never
      # surveys. The GitLab counterpart carries the same rule and was short three
      # repos by it on 2026-09-17 (coding-ai/gitlab-skill!142).
      # The survey dedupes rename redirects by resolved full_name, so a stale name
      # here becomes a DUPLICATE row, never a second bump PR.
      # One deliberate exception to "the manifest defines the fleet":
      # coding_agent_cli_toolset ships the cli-tools plugin but is versioned by
      # commit (no plugin.json version, no release tags), so a bump would be wrong.
      # It replaced cli-tools-skill, archived 2026-09.
      agent-harness-skill
      agent-rules-skill
      automated-assessment-skill
      concourse-ci-skill
      context7-skill
      data-tools-skill
      docker-development-skill
      enterprise-readiness-skill
      file-search-skill
      german-technical-writing-skill
      git-workflow-skill
      github-project-skill
      github-release-skill
      go-development-skill
      jira-skill
      jujutsu-workflow-skill
      markdown-to-pdf-skill
      matrix-skill
      netresearch-branding-skill
      orocommerce-skill
      peer-qa-review-skill
      php-ast-edit-skill
      php-modernization-skill
      retro-skill
      security-audit-skill
      skill-repo-skill
      typo3-a11y-skill
      typo3-ckeditor5-skill
      typo3-conformance-skill
      typo3-core-contributions-skill
      typo3-ddev-skill
      typo3-docs-skill
      typo3-extension-upgrade-skill
      typo3-project-upgrade-skill
      typo3-site-conformance-skill
      typo3-testing-skill
      typo3-typoscript-ref-skill
      typo3-upgrade-effort-model-skill
      typo3-vite-skill
      
    • migrate-licensing.sh 9 KB
      #!/bin/bash
      # migrate-licensing.sh - Migrate a skill repo from single LICENSE to split licensing
      # Usage: ./migrate-licensing.sh [<repo-root-path>]
      #        ./migrate-licensing.sh --help
      set -euo pipefail
      
      usage() { sed -n '2,4p' "$0" | sed 's/^# \{0,1\}//'; }
      
      case "${1:-}" in
          -h|--help) usage; exit 0 ;;
          -*) echo "unknown option: $1" >&2; usage >&2; exit 2 ;;
      esac
      
      REPO_DIR="${1:-.}"
      if [[ ! -d "$REPO_DIR" ]]; then
          echo "not a directory: $REPO_DIR" >&2
          exit 2
      fi
      
      # The copyright years come from the repository, not from this file: the year of
      # its first commit through the current year, or the current year alone when the
      # two are the same or the directory has no history.
      #
      # A shallow clone's oldest reachable commit is not the repository's first, so
      # the year would be wrong without anything saying so. Refuse rather than guess;
      # fetching the history is the caller's decision, not a side effect of this one.
      CURRENT_YEAR="$(date +%Y)"
      if [[ "$(git -C "$REPO_DIR" rev-parse --is-shallow-repository 2>/dev/null || true)" == "true" ]]; then
          echo "shallow clone: the first commit is not available, so the copyright years cannot be derived." >&2
          echo "Run: git -C $REPO_DIR fetch --unshallow" >&2
          exit 2
      fi
      FIRST_YEAR="$(git -C "$REPO_DIR" log --reverse --format=%ad --date=format:%Y 2>/dev/null | head -n 1 || true)"
      FIRST_YEAR="${FIRST_YEAR:-$CURRENT_YEAR}"
      if [[ "$FIRST_YEAR" == "$CURRENT_YEAR" ]]; then
          YEAR="$CURRENT_YEAR"
      else
          YEAR="$FIRST_YEAR-$CURRENT_YEAR"
      fi
      
      echo "Migrating licensing in: $REPO_DIR"
      
      # 1. Create LICENSE-MIT
      if [[ -f "$REPO_DIR/LICENSE" ]]; then
          if grep -q "GNU GENERAL PUBLIC LICENSE" "$REPO_DIR/LICENSE"; then
              echo "INFO: Repo has GPL license — creating MIT from scratch"
              cat > "$REPO_DIR/LICENSE-MIT" << MITEOF
      MIT License
      
      Copyright (c) $YEAR Netresearch DTT GmbH
      
      Permission is hereby granted, free of charge, to any person obtaining a copy
      of this software and associated documentation files (the "Software"), to deal
      in the Software without restriction, including without limitation the rights
      to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
      copies of the Software, and to permit persons to whom the Software is
      furnished to do so, subject to the following conditions:
      
      The above copyright notice and this permission notice shall be included in all
      copies or substantial portions of the Software.
      
      THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
      IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
      FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
      AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
      LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
      OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
      SOFTWARE.
      MITEOF
          else
              # Existing MIT — copy it and extend its own year range to the current
              # year. The holder and the start year are the licence's, never ours. A
              # range already present is extended, not nested: 2024-2025 becomes
              # 2024-<current>, not 2024-<current>-2025.
              cp "$REPO_DIR/LICENSE" "$REPO_DIR/LICENSE-MIT"
              # Each notice on its own: a licence can carry several, and one that is
              # already current must not stop the others from being extended. A
              # notice that starts in the current year stays a single year.
              python3 - "$REPO_DIR/LICENSE-MIT" "$CURRENT_YEAR" << 'PYEOF'
      import re, sys
      path, now = sys.argv[1], sys.argv[2]
      def extend(m):
          start = m.group(1)
          return m.group(0) if start == now else f"Copyright (c) {start}-{now}"
      with open(path) as f:
          text = f.read()
      with open(path, "w") as f:
          f.write(re.sub(r"Copyright \(c\) (\d{4})(?:-\d{4})?", extend, text))
      PYEOF
          fi
          # Stage removal of old LICENSE
          git -C "$REPO_DIR" rm -f LICENSE 2>/dev/null || rm -f "$REPO_DIR/LICENSE"
      elif [[ -f "$REPO_DIR/LICENSE-MIT" ]]; then
          # Already migrated. Writing it again from scratch would replace a copyright
          # holder the first run deliberately preserved with Netresearch's.
          echo "INFO: LICENSE-MIT already exists — leaving it untouched"
      else
          echo "INFO: No LICENSE found — creating LICENSE-MIT from scratch"
          cat > "$REPO_DIR/LICENSE-MIT" << MITEOF
      MIT License
      
      Copyright (c) $YEAR Netresearch DTT GmbH
      
      Permission is hereby granted, free of charge, to any person obtaining a copy
      of this software and associated documentation files (the "Software"), to deal
      in the Software without restriction, including without limitation the rights
      to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
      copies of the Software, and to permit persons to whom the Software is
      furnished to do so, subject to the following conditions:
      
      The above copyright notice and this permission notice shall be included in all
      copies or substantial portions of the Software.
      
      THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
      IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
      FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
      AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
      LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
      OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
      SOFTWARE.
      MITEOF
      fi
      
      # 2. Create LICENSE-CC-BY-SA-4.0
      cat > "$REPO_DIR/LICENSE-CC-BY-SA-4.0" << CCEOF
      Creative Commons Attribution-ShareAlike 4.0 International
      
      Copyright (c) $YEAR Netresearch DTT GmbH
      
      This work is licensed under the Creative Commons Attribution-ShareAlike 4.0
      International License. To view a copy of this license, visit
      https://creativecommons.org/licenses/by-sa/4.0/ or send a letter to
      Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.
      
      You are free to:
      - Share: copy and redistribute the material in any medium or format
      - Adapt: remix, transform, and build upon the material for any purpose,
        even commercially
      
      Under the following terms:
      - Attribution: You must give appropriate credit, provide a link to the
        license, and indicate if changes were made.
      - ShareAlike: If you remix, transform, or build upon the material, you
        must distribute your contributions under the same license as the original.
      CCEOF
      
      # 3. Update composer.json and the plugin manifest.
      #
      # The root plugin.json is the source of truth; .claude-plugin/plugin.json is
      # generated from it by sync-plugin-manifest.sh, which would overwrite an edit
      # made to the copy. So the root is edited and the copy regenerated. Only a
      # repository that has no root manifest yet gets the copy edited directly.
      # Each file keeps its own indentation, so the edit is one line and not a
      # reformat of the whole document.
      python3 - "$REPO_DIR" << 'PYEOF'
      import json, os, re, sys
      repo_dir = sys.argv[1]
      
      manifests = ["composer.json"]
      if os.path.isfile(os.path.join(repo_dir, "plugin.json")):
          manifests.append("plugin.json")
      elif os.path.isfile(os.path.join(repo_dir, ".claude-plugin", "plugin.json")):
          manifests.append(os.path.join(".claude-plugin", "plugin.json"))
      
      for rel_path in manifests:
          full_path = os.path.join(repo_dir, rel_path)
          if not os.path.isfile(full_path):
              continue
          with open(full_path) as f:
              text = f.read()
          m = re.search(r'^([ \t]+)"', text, re.M)
          indent = m.group(1) if m else "  "
          data = json.loads(text)
          data['license'] = '(MIT AND CC-BY-SA-4.0)'
          with open(full_path, 'w') as f:
              json.dump(data, f, indent=indent, ensure_ascii=False)
              f.write('\n')
          print(f"Updated {rel_path} license")
      PYEOF
      
      SYNC="$(dirname "$0")/sync-plugin-manifest.sh"
      if [[ -f "$REPO_DIR/plugin.json" && -f "$REPO_DIR/.claude-plugin/plugin.json" ]]; then
          if [[ -f "$SYNC" ]]; then
              bash "$SYNC" --repo "$REPO_DIR"
          else
              echo "WARN: .claude-plugin/plugin.json not regenerated — run sync-plugin-manifest.sh"
          fi
      fi
      
      # 5. Update README.md license section
      if [[ -f "$REPO_DIR/README.md" ]]; then
          python3 - "$REPO_DIR" << 'PYEOF'
      import re, sys
      
      repo_dir = sys.argv[1] if len(sys.argv) > 1 else "."
      readme_path = f"{repo_dir}/README.md"
      
      with open(readme_path, 'r') as f:
          content = f.read()
      
      # Replace license section
      new_license = """## License
      
      This project uses split licensing:
      
      - **Code** (scripts, workflows, configs): [MIT](LICENSE-MIT)
      - **Content** (skill definitions, documentation, references): [CC-BY-SA-4.0](LICENSE-CC-BY-SA-4.0)
      
      See the individual license files for full terms."""
      
      # Match ## License section until next ## heading or --- or end of file
      pattern = r'## License\n.*?(?=\n## |\n---|\Z)'
      content = re.sub(pattern, new_license, content, flags=re.DOTALL)
      
      # Fix structure diagrams
      content = re.sub(
          r'├── LICENSE\s+# (?:MIT|GPL[^\n]*)',
          '├── LICENSE-MIT           # Code license (MIT)\n├── LICENSE-CC-BY-SA-4.0  # Content license (CC-BY-SA-4.0)',
          content
      )
      
      with open(readme_path, 'w') as f:
          f.write(content)
      PYEOF
          echo "Updated README.md license section"
      fi
      
      echo "Done! Review changes with: git -C $REPO_DIR diff"
      
    • roll-changelog.py 14.8 KB
      #!/usr/bin/env python3
      """roll-changelog.py — move the [Unreleased] section of a CHANGELOG.md under a
      new released heading, reproducing the repo's own heading shape.
      
      Usage:
          roll-changelog.py FILE VERSION [--date YYYY-MM-DD] [--dry-run]
      
      Why this exists (references/release-discipline.md, "Changelog Rollover"):
      five released-heading shapes coexist in the fleet, and a roll anchored on one
      shape matches nothing in the others and exits 0 — the release then ships with
      its content still under Unreleased and nobody sees an error. This script
      detects the shape from the newest *released* heading and reproduces it,
      including the dash character and, for the linked form, the tag inside the URL.
      
      Shapes:
          bracketed dash      ## [1.2.3] - 2026-08-08
          bare paren          ## 1.2.3 (2026-08-08)
          bracketed em dash   ## [0.3.23] — 2026-07-02
          linked              ## [v2.6.0](https://.../releases/tag/v2.6.0) — 2026-02-28
          none yet            first release: default to the keep-a-changelog
                              bracketed-dash form the [Unreleased] bracket implies
      
      Guards:
          * exactly one Unreleased heading outside code fences, or abort;
          * the Unreleased section must have content — rolling nothing is an error,
            never a silent success;
          * heading-looking lines inside fenced code blocks are ignored (a ^##
            anchored scan is fence-blind and splices sections into examples);
          * output keeps a blank line after each heading so the rolled file stays
            markdownlint-clean (MD022/MD032 red-lit 16 of 63 repos mid-sweep once).
      
      Exit codes: 0 = rolled (or clean --dry-run), 1 = refused; the file is never
      partially written (temp file + atomic replace).
      """
      
      import argparse
      import datetime
      import os
      import re
      import sys
      import tempfile
      
      UNRELEASED_RE = re.compile(r"^##\s+\[?unreleased\]?\s*$", re.IGNORECASE)
      FENCE_RE = re.compile(r"^ {0,3}(```|~~~)")
      # Newest released heading, one regex per shape. Order matters: the linked form
      # also starts with "## [" and must win over plain bracketed.
      LINKED_RE = re.compile(
          r"^##\s+\[(?P<tag>[^\]]+)\]\((?P<url>[^)]+)\)\s+(?P<dash>[-—–])\s+(?P<date>.+?)\s*$"
      )
      BRACKETED_RE = re.compile(
          r"^##\s+\[(?P<tag>[^\]]+)\]\s+(?P<dash>[-—–])\s+(?P<date>.+?)\s*$"
      )
      BARE_PAREN_RE = re.compile(r"^##\s+(?P<tag>\S+)\s+\((?P<date>[^)]+)\)\s*$")
      
      
      def fail(msg: str) -> None:
          print(f"ERROR: {msg}", file=sys.stderr)
          sys.exit(1)
      
      
      def scan(lines):
          """Yield (index, line, in_fence) tracking fenced code blocks."""
          fence = None
          for i, line in enumerate(lines):
              m = FENCE_RE.match(line)
              if m:
                  marker = m.group(1)
                  if fence is None:
                      fence = marker
                  elif fence == marker:
                      fence = None
                  yield i, line, True
                  continue
              yield i, line, fence is not None
      
      
      def released_heading(version: str, sample: str, date: str) -> str:
          """Reproduce the shape of `sample` (a released heading) for `version`."""
          m = LINKED_RE.match(sample)
          if m:
              old_tag = m.group("tag")
              new_tag = f"v{version}" if old_tag.startswith("v") else version
              url = m.group("url")
              if old_tag not in url:
                  fail(
                      f"linked heading URL does not contain its own tag "
                      f"('{old_tag}' not in '{url}') — refusing to guess the rewrite"
                  )
              new_url = url.replace(old_tag, new_tag)
              return f"## [{new_tag}]({new_url}) {m.group('dash')} {date}"
          m = BRACKETED_RE.match(sample)
          if m:
              old_tag = m.group("tag")
              new_tag = f"v{version}" if old_tag.startswith("v") else version
              return f"## [{new_tag}] {m.group('dash')} {date}"
          m = BARE_PAREN_RE.match(sample)
          if m:
              old_tag = m.group("tag")
              new_tag = f"v{version}" if old_tag.startswith("v") else version
              return f"## {new_tag} ({date})"
          fail(f"unrecognized released-heading shape: '{sample.rstrip()}'")
          raise AssertionError  # unreachable; fail() exits
      
      
      def parse_args():
          parser = argparse.ArgumentParser(
              description="Roll the [Unreleased] CHANGELOG section into a release."
          )
          parser.add_argument("file", help="path to CHANGELOG.md")
          parser.add_argument(
              "version",
              nargs="?",
              help="release version (leading v stripped); not needed with --check-unreleased",
          )
          parser.add_argument(
              "--date",
              default=datetime.datetime.now(datetime.timezone.utc)
              .astimezone()
              .date()
              .isoformat(),
              help="release date (default: today, ISO 8601)",
          )
          parser.add_argument(
              "--dry-run",
              action="store_true",
              help="print the new heading and moved-section size, write nothing",
          )
          parser.add_argument(
              "--check-unreleased",
              action="store_true",
              help="report the Unreleased section's state (no-unreleased | empty | "
              "has-content) with the SAME emptiness rule the roll enforces, write "
              "nothing; surveys use this so an empty section surfaces before the "
              "bump fails on it",
          )
          args = parser.parse_args()
          if args.check_unreleased:
              return args
          if args.version is None:
              fail("version is required (only --check-unreleased goes without one)")
          args.version = args.version.lstrip("v")
          if not re.fullmatch(r"\d+\.\d+\.\d+(?:[-+][0-9A-Za-z.-]+)?", args.version):
              fail(f"'{args.version}' is not a semantic version")
          if not re.fullmatch(r"\d{4}-\d{2}-\d{2}", args.date):
              fail(f"--date '{args.date}' is not YYYY-MM-DD")
          return args
      
      
      def read_changelog(path):
          """Return (lines, eol, trailing_newline), preserving the file's own line
          endings — rewriting a CRLF changelog wholesale to LF buries the roll in a
          whole-file diff."""
          try:
              with open(path, encoding="utf-8", newline="") as fh:
                  text = fh.read()
          except OSError as exc:
              fail(str(exc))
          eol = "\r\n" if text.count("\r\n") > text.count("\n") - text.count("\r\n") else "\n"
          return (
              text.replace("\r\n", "\n").splitlines(),
              eol,
              text.endswith(("\n", "\r\n")),
          )
      
      
      def locate_headings(lines):
          """Return (unreleased_index, first_released) or refuse: exactly one
          Unreleased heading outside code fences, and no released heading above it."""
          unreleased_at = []
          first_released = None  # (index, line)
          for i, line, in_fence in scan(lines):
              if in_fence:
                  continue
              if UNRELEASED_RE.match(line):
                  unreleased_at.append(i)
              elif first_released is None and line.startswith("## "):
                  first_released = (i, line)
          if len(unreleased_at) != 1:
              fail(
                  f"{len(unreleased_at)} Unreleased headings outside code fences "
                  f"(need exactly 1)"
              )
          idx = unreleased_at[0]
          if first_released is not None and first_released[0] < idx:
              fail(
                  f"a released heading (line {first_released[0] + 1}) precedes the "
                  f"Unreleased heading (line {idx + 1}) — refusing to roll a "
                  f"non-standard layout"
              )
          return idx, first_released
      
      
      def splice_roll(lines, idx, end, new_heading):
          """Return (new_lines, moved_count): the Unreleased heading stays as found,
          one blank line, the new released heading, one blank line, then the old
          section content — blank-line hygiene keeps markdownlint green."""
          content = lines[idx + 1 : end]
          while content and not content[0].strip():
              content.pop(0)
          # Exactly one blank line between the moved content and the next heading.
          while content and not content[-1].strip():
              content.pop()
          tail = lines[end:]
          new_lines = lines[:idx] + [lines[idx], "", new_heading, ""] + content
          if tail:
              new_lines += [""] + tail
          return new_lines, len(content)
      
      
      def write_atomic(path, out):
          directory = os.path.dirname(os.path.abspath(path))
          fd, tmp_path = tempfile.mkstemp(dir=directory, prefix=".roll-changelog.")
          try:
              with os.fdopen(fd, "w", encoding="utf-8") as fh:
                  fh.write(out)
              os.replace(tmp_path, path)
          except OSError as exc:
              os.unlink(tmp_path)
              fail(str(exc))
      
      
      def unreleased_state(lines):
          """One of 'no-unreleased' | 'empty' | 'has-content', by the same section
          boundaries and comments-are-not-content rule the roll itself enforces."""
          unreleased_at = []
          heading_at = []
          for i, line, in_fence in scan(lines):
              if in_fence:
                  continue
              if UNRELEASED_RE.match(line):
                  unreleased_at.append(i)
              elif line.startswith("## "):
                  heading_at.append(i)
          if not unreleased_at:
              return "no-unreleased"
          if len(unreleased_at) > 1:
              fail(
                  f"{len(unreleased_at)} Unreleased headings outside code fences "
                  f"(need exactly 1)"
              )
          idx = unreleased_at[0]
          end = next((h for h in heading_at if h > idx), len(lines))
          section = "\n".join(lines[idx + 1 : end])
          if re.sub(r"<!--.*?-->", "", section, flags=re.DOTALL).strip():
              return "has-content"
          return "empty"
      
      
      def main() -> None:
          args = parse_args()
          if args.check_unreleased:
              lines, _eol, _tn = read_changelog(args.file)
              print(unreleased_state(lines))
              return
          version = args.version
          lines, eol, trailing_newline = read_changelog(args.file)
          idx, first_released = locate_headings(lines)
      
          if first_released is None:
              # First release: no shape to copy. The [Unreleased] bracket implies
              # keep-a-changelog, so default to its bracketed-dash form.
              new_heading = f"## [{version}] - {args.date}"
          else:
              new_heading = released_heading(version, first_released[1], args.date)
      
          # The section to be rolled must contain content — and HTML comments are
          # not content: rolling a comment-only section would ship a visually empty
          # release.
          end = first_released[0] if first_released else len(lines)
          section = "\n".join(lines[idx + 1 : end])
          if not re.sub(r"<!--.*?-->", "", section, flags=re.DOTALL).strip():
              fail(
                  "the Unreleased section is empty (comments are not content) — nothing to release"
              )
      
          new_lines, moved = splice_roll(lines, idx, end, new_heading)
          if new_lines == lines:
              fail("the roll changed no lines — refusing to report success")
      
          defs_note = update_link_reference_definitions(new_lines, new_heading, version)
      
          if args.dry_run:
              print(f"would insert: {new_heading}")
              print(f"would move: {moved} lines out of Unreleased")
          else:
              write_atomic(args.file, eol.join(new_lines) + (eol if trailing_newline else ""))
              print(f"rolled Unreleased -> {new_heading}")
          if defs_note:
              print(defs_note)
      
      
      def released_labels(lines):
          """Tag labels of every released heading, document order, fences skipped.
      
          The compare base for a new release is the heading directly below it, not
          whatever [Unreleased] happened to point at — see update_link_reference_
          definitions.
          """
          labels = []
          for _, line, in_fence in scan(lines):
              if in_fence:
                  continue
              for rx in (LINKED_RE, BRACKETED_RE, BARE_PAREN_RE):
                  m = rx.match(line)
                  if m:
                      labels.append(m.group("tag"))
                      break
          return labels
      
      
      def styled_tag(label, v_prefixed):
          """Render `label` in the tag style the link definitions already use."""
          bare = label.removeprefix("v")
          return f"v{bare}" if v_prefixed else bare
      
      
      def update_link_reference_definitions(new_lines, new_heading, version):
          """Keep-a-changelog ref-style links: `[Unreleased]: …/compare/vOLD...HEAD`
          at the bottom, one definition per released heading. After a roll the
          Unreleased range must restart at the new tag and the new heading needs its
          own definition — otherwise the new release renders as a dead link and the
          Unreleased diff keeps showing released commits. Mutates new_lines in
          place; returns a note (updated/warning) or None when no definitions exist.
          """
          def_re = re.compile(r"^\[(?P<label>[Uu]nreleased)\]:\s*(?P<url>\S+)\s*$")
          for i, line in enumerate(new_lines):
              m = def_re.match(line)
              if not m:
                  continue
              url = m.group("url")
              # Parsed from the right — a lazy .*? prefix before the version would
              # backtrack super-linearly on crafted input. The range URL must end in
              # <old-tag>...HEAD.
              stem = url[: -len("...HEAD")] if url.endswith("...HEAD") else None
              old_m = re.search(r"v?\d+\.\d+\.\d+$", stem) if stem else None
              head_m = re.match(r"^##\s+\[(?P<label>[^\]]+)\]", new_heading)
              if not old_m or not head_m:
                  return (
                      "WARNING: link-reference definitions found but not in the "
                      "known compare-range shape — update [Unreleased] and add "
                      f"[{version}] manually"
                  )
              old = old_m.group(0)
              v_prefixed = old.startswith("v")
              new_tag = f"v{version}" if v_prefixed else version
              base = stem[: old_m.start()]
      
              # The compare base comes from the HEADINGS, not from [Unreleased].
              # [Unreleased] can be stale — an earlier roll that skipped a definition
              # leaves it pointing several releases back, and inheriting it spans the
              # new release across all of them (matrix-skill shipped
              # v3.1.1...v3.1.3 that way).
              labels = released_labels(new_lines)
              head_label = head_m.group("label")
              prev_label = labels[1] if len(labels) > 1 else None
              prev_tag = styled_tag(prev_label, v_prefixed) if prev_label else old
      
              new_lines[i] = f"[{m.group('label')}]: {base}{new_tag}...HEAD"
              new_lines.insert(i + 1, f"[{head_label}]: {base}{prev_tag}...{new_tag}")
      
              # A previous release with no definition of its own is what let the
              # stale base survive in the first place; write it while we can still
              # see the version below it.
              note = f"updated link-reference definitions ([Unreleased] -> {new_tag}...HEAD)"
              if prev_label is not None and len(labels) > 2:
                  # Compared without the v prefix: a file may head its sections
                  # `## [v3.1.2]` while defining `[3.1.2]:`. A raw match would call
                  # the existing definition missing and add a second one under the
                  # other spelling.
                  want = prev_label.removeprefix("v")
                  have_prev = any(
                      ln.split("]:", 1)[0][1:].removeprefix("v") == want
                      for ln in new_lines
                      if ln.startswith("[") and "]:" in ln
                  )
                  if not have_prev:
                      prev2_tag = styled_tag(labels[2], v_prefixed)
                      new_lines.insert(
                          i + 2, f"[{prev_label}]: {base}{prev2_tag}...{prev_tag}"
                      )
                      note += f"; backfilled the missing [{prev_label}] definition"
              return note
          return None
      
      
      if __name__ == "__main__":
          main()
      
    • sync-plugin-manifest.sh 4.5 KB
      #!/usr/bin/env bash
      #
      # sync-plugin-manifest.sh — project the portable Agent Plugins manifest
      # (./plugin.json) into the Claude Code manifest (.claude-plugin/plugin.json).
      #
      # Usage:
      #   sync-plugin-manifest.sh                 # write .claude-plugin/plugin.json
      #   sync-plugin-manifest.sh --check         # verify it is in sync, write nothing
      #   sync-plugin-manifest.sh --repo DIR ...  # operate on DIR instead of cwd
      #
      # Root ./plugin.json is the source of truth for the shared metadata: name,
      # version, description, author, homepage, repository, license, keywords.
      # .claude-plugin/plugin.json is generated from it and keeps its Claude-only
      # keys (skills, agents, commands, outputStyles, hooks, mcpServers, metadata, …)
      # untouched — those fields have no place in the portable manifest, whose schema
      # is closed. `support` is not among them: it is no Claude Code field either, so
      # this script preserving it is a carry-over, not an endorsement — see
      # references/agent-plugins-compat.md.
      #
      # Exit codes: 0 = written / in sync (or no portable manifest to sync from),
      #             1 = out of sync (--check) or invalid input.
      
      set -euo pipefail
      
      REPO_DIR="."
      CHECK=0
      
      while [[ $# -gt 0 ]]; do
        case "$1" in
          --check) CHECK=1; shift ;;
          --repo)
            REPO_DIR="${2:-}"
            [[ -n "$REPO_DIR" ]] || { echo "ERROR: --repo needs a directory" >&2; exit 1; }
            shift 2
            ;;
          --repo=*) REPO_DIR="${1#--repo=}"; shift ;;
          -h|--help)
            grep -E '^#' "$0" | sed -e '1d' -e 's/^# \{0,1\}//'
            exit 0
            ;;
          *) echo "ERROR: unknown argument: $1" >&2; exit 1 ;;
        esac
      done
      
      command -v python3 >/dev/null 2>&1 || { echo "ERROR: python3 is required" >&2; exit 1; }
      
      PORTABLE="$REPO_DIR/plugin.json"
      CLAUDE="$REPO_DIR/.claude-plugin/plugin.json"
      
      if [[ ! -f "$PORTABLE" ]]; then
        echo "SKIP: $PORTABLE not found — nothing to sync from"
        exit 0
      fi
      
      CHECK="$CHECK" PORTABLE="$PORTABLE" CLAUDE="$CLAUDE" python3 <<'PYEOF'
      import json
      import os
      import sys
      
      check = os.environ["CHECK"] == "1"
      portable_path = os.environ["PORTABLE"]
      claude_path = os.environ["CLAUDE"]
      
      # Shared metadata, in the order the generated manifest lists it.
      SHARED = ["name", "version", "description", "author",
                "homepage", "repository", "license", "keywords"]
      
      try:
          with open(portable_path, encoding="utf-8") as fh:
              portable = json.load(fh)
      except (OSError, ValueError) as exc:
          print(f"ERROR: cannot read {portable_path}: {exc}", file=sys.stderr)
          sys.exit(1)
      
      if not isinstance(portable, dict):
          print(f"ERROR: {portable_path} must contain a JSON object", file=sys.stderr)
          sys.exit(1)
      
      current = {}
      if os.path.isfile(claude_path):
          try:
              with open(claude_path, encoding="utf-8") as fh:
                  current = json.load(fh)
          except (OSError, ValueError) as exc:
              print(f"ERROR: cannot read {claude_path}: {exc}", file=sys.stderr)
              sys.exit(1)
          if not isinstance(current, dict):
              print(f"ERROR: {claude_path} must contain a JSON object", file=sys.stderr)
              sys.exit(1)
      
      generated = {}
      for key in SHARED:
          if key in portable:
              generated[key] = portable[key]
      
      # Claude-only keys survive: everything the portable manifest does not own.
      # `$schema` and `extensions` are portable-manifest concerns and never copied.
      for key, value in current.items():
          if key in SHARED or key in ("$schema", "extensions"):
              continue
          generated[key] = value
      
      rendered = json.dumps(generated, indent=2, ensure_ascii=False) + "\n"
      
      if check:
          # Value parity, not byte identity: key order and formatting in the Claude
          # manifest are nobody's business, drift in the shared metadata is.
          drift = [k for k in SHARED if portable.get(k) != current.get(k)]
          if not drift:
              print(f"OK: {claude_path} is in sync with {portable_path}")
              sys.exit(0)
          print(f"ERROR: {claude_path} is out of sync with {portable_path}", file=sys.stderr)
          print("       Run: bash skills/skill-repo/scripts/sync-plugin-manifest.sh", file=sys.stderr)
          for key in drift:
              print(f"       {key}: portable={portable.get(key)!r} claude={current.get(key)!r}",
                    file=sys.stderr)
          sys.exit(1)
      
      try:
          with open(claude_path, encoding="utf-8") as fh:
              on_disk = fh.read()
      except OSError:
          on_disk = None
      
      if on_disk == rendered:
          print(f"OK: {claude_path} is in sync with {portable_path}")
          sys.exit(0)
      
      os.makedirs(os.path.dirname(claude_path) or ".", exist_ok=True)
      with open(claude_path, "w", encoding="utf-8") as fh:
          fh.write(rendered)
      print(f"WROTE: {claude_path}")
      PYEOF
      
    • validate-evals.sh 28.6 KB
      #!/usr/bin/env bash
      # validate-evals.sh - Structural validation of evals.json files
      # Supports three formats:
      #   Unified (recommended): {"skill_name": "...", "evals": [{id, eval_name, prompt,
      #       expected_output?, expectations?: ["..."], assertions?: [{type, pattern}]}]}
      #     - Evals may use expectations (string[], LLM-as-judge), assertions (object[],
      #       regex matching), or both. At least one grading mechanism is required.
      #   Legacy A: {"skill_name": "...", "evals": [{id, eval_name, prompt, assertions: [...]}]}
      #   Legacy B: [{name, prompt, assertions: [{type, value/pattern, description?}]}]
      #
      # Any eval may carry an optional self-check:
      #   "samples": {"passing": "<answer every assertion must match>",
      #               "failing": ["<answer at least one assertion must reject>"]}
      # Patterns are checked with `grep -E` and matched with `grep -qiE`, the same
      # binary and flags run-ab-evals.sh grades with, so the check cannot certify an
      # eval the grader would fail. Every assertion carrying value/pattern counts,
      # whatever its type, because the grader greps every one of them; `must_not` is
      # checked inverted here and graded inverted there.
      #
      # `samples` are OPTIONAL on an eval that already exists and is left alone, and
      # REQUIRED on one that is new or whose assertions changed — the half of
      # netresearch/retro-skill#92 that has to live here, because this is where the
      # fleet's eval gate runs. The comparison needs the same file at the base
      # revision, which the caller supplies:
      #
      #   EVALS_BASE_FILE=<path to the base copy of THIS evals.json>
      #
      # Unset or empty (every local run, every push build, every consumer that has
      # not updated its workflow): no comparison, no requirement, verdict byte for
      # byte what it was before. Set: every eval that is new, or whose `assertions`
      # value differs from the base copy, must carry `samples.passing`. Untouched
      # evals are never looked at — the 485 fleet evals without samples are not
      # retrofitted. An eval whose assertions carry no pattern this validator would
      # run against a sample is exempt, because it FAILS samples that no assertion
      # backs: that covers an eval graded by `expectations` alone, one whose
      # assertions are plain strings rather than {type, pattern} objects, and one
      # whose only pattern is a `*_contains` literal grep cannot parse. Plain-string
      # assertions ARE graded at run time (run-ab-evals.sh falls back to `str(a)`),
      # so that exemption follows this validator's samples machinery rather than the
      # grader's reach — 210 of the 518 evals installed on one host are in that
      # shape, and closing it would change what samples mean for evals that already
      # carry them.
      #
      # Usage: bash validate-evals.sh [path-to-evals.json] [--require-evals]
      #   If no path given, searches skills/*/evals/evals.json then evals/evals.json
      #
      #   --require-evals: only takes effect when no evals.json is found. Instead of
      #     the plain "no evals.json found" error, checks whether any SKILL.md in
      #     the repo exceeds the breadth threshold (body >300 words or >3 files in
      #     references/). Small pointer-skills stay exempt. Broad skills without an
      #     evals.json fail with a ::error:: annotation naming the skill.
      
      set -euo pipefail
      
      PASS=0
      FAIL=0
      WARN=0
      
      pass() { PASS=$((PASS + 1)); echo "  PASS: $1"; }
      fail() { FAIL=$((FAIL + 1)); echo "  FAIL: $1"; }
      warn() { WARN=$((WARN + 1)); echo "  WARN: $1"; }
      
      # --- Breadth check (--require-evals mode) ---
      # Only called when no evals.json was found anywhere in the repo. Determines
      # whether any skill is broad enough (body >300 words or >3 reference files)
      # that it should have had one. Small pointer-skills are exempt. Exits 1 with
      # a ::error:: annotation per offending skill; otherwise returns 0.
      check_required_evals() {
        local skill_md skill_dir skill_name body_words ref_count
        local -a skill_files=() broad_skills=() ref_files=()
      
        [[ -f "SKILL.md" ]] && skill_files+=("SKILL.md")
        for skill_md in skills/*/SKILL.md; do
          [[ -f "$skill_md" ]] && skill_files+=("$skill_md")
        done
      
        if [[ ${#skill_files[@]} -eq 0 ]]; then
          echo "WARN: --require-evals set but no SKILL.md found, skipping breadth check"
          return 0
        fi
      
        for skill_md in "${skill_files[@]}"; do
          skill_dir=$(dirname "$skill_md")
          skill_name=$(basename "$skill_dir")
          [[ "$skill_dir" == "." ]] && skill_name=$(basename "$(pwd)")
      
          # Body word count: SKILL.md content after the frontmatter closing '---'.
          body_words=$(awk '
            BEGIN { has_fm = 0; fm_closed = 0; total_words = 0; body_words = 0 }
            NR == 1 { if (/^---$/) { has_fm = 1; next } }
            /^---$/ && has_fm && !fm_closed { fm_closed = 1; next }
            {
              total_words += NF
              if (fm_closed) { body_words += NF }
            }
            END { print (has_fm && fm_closed ? body_words : total_words) }
          ' "$skill_md")
      
          ref_count=0
          if [[ -d "$skill_dir/references" ]]; then
            ref_files=("$skill_dir/references"/*.md)
            if [[ -e "${ref_files[0]}" ]]; then
              ref_count=${#ref_files[@]}
            fi
          fi
      
          if [[ "$body_words" -gt 300 ]] || [[ "$ref_count" -gt 3 ]]; then
            broad_skills+=("$skill_name (body: ${body_words} words, references: ${ref_count} files)")
          fi
        done
      
        if [[ ${#broad_skills[@]} -eq 0 ]]; then
          echo "PASS: no skill exceeds the breadth threshold (body >300 words or >3 reference files); evals.json not required"
          return 0
        fi
      
        local entry
        for entry in "${broad_skills[@]}"; do
          echo "::error::require_evals: $entry exceeds the breadth threshold but has no evals.json — add skills/<name>/evals/evals.json (or evals/evals.json) with structural evals, or keep require_evals: false for this repo"
        done
        echo ""
        echo "Results: 0 passed, ${#broad_skills[@]} failed, 0 warnings"
        exit 1
      }
      
      # --- Locate evals.json ---
      REQUIRE_EVALS=0
      EVALS_FILE=""
      for arg in "$@"; do
        case "$arg" in
          --require-evals) REQUIRE_EVALS=1 ;;
          *) EVALS_FILE="$arg" ;;
        esac
      done
      
      EVALS_BASE_FILE="${EVALS_BASE_FILE:-}"
      if [[ -z "$EVALS_FILE" && -n "$EVALS_BASE_FILE" ]]; then
        # A base copy belongs to one named file. In discovery mode the run can cover
        # several evals.json, and the sub-invocations below inherit the environment,
        # so an unqualified base would be compared against all of them.
        echo "WARN: EVALS_BASE_FILE is set but no evals.json was named — the base"
        echo "      copy belongs to one file, so the samples requirement is off for"
        echo "      this run. Pass the path explicitly to enforce it."
        EVALS_BASE_FILE=""
        export EVALS_BASE_FILE
      fi
      
      if [[ -z "$EVALS_FILE" ]]; then
        # Every candidate, not the first: a repo shipping several skills ships several
        # evals.json, and stopping at the first left the rest unchecked — six repos in
        # the fleet carry more than one, matrix-skill three. Same hole the skill
        # validator had (issue #214).
        EVALS_FILES=()
        for candidate in skills/*/evals/evals.json evals/evals.json; do
          [[ -f "$candidate" ]] && EVALS_FILES+=("$candidate")
        done
      
        if [[ ${#EVALS_FILES[@]} -gt 1 ]]; then
          # One run per file so each gets its own counters and summary, with the exit
          # code covering all of them.
          MULTI_RC=0
          for candidate in "${EVALS_FILES[@]}"; do
            echo "=============================================="
            bash "$0" "$candidate" || MULTI_RC=1
            echo ""
          done
          if [[ "$MULTI_RC" -ne 0 ]]; then
            echo "::error::at least one evals.json failed validation"
          else
            echo "All ${#EVALS_FILES[@]} evals.json files validate"
          fi
          exit "$MULTI_RC"
        fi
      
        EVALS_FILE="${EVALS_FILES[0]:-}"
      fi
      
      if [[ -z "$EVALS_FILE" ]] || [[ ! -f "$EVALS_FILE" ]]; then
        if [[ "$REQUIRE_EVALS" -eq 1 ]]; then
          check_required_evals
          exit 0
        fi
        echo "ERROR: No evals.json found"
        echo "Searched: skills/*/evals/evals.json, evals/evals.json"
        exit 1
      fi
      
      echo "Validating: $EVALS_FILE"
      echo "---"
      
      # --- Valid JSON ---
      if ! python3 -c "import json, sys; json.load(open(sys.argv[1]))" "$EVALS_FILE" 2>/dev/null; then
        fail "Invalid JSON"
        echo ""
        echo "Results: $PASS passed, $FAIL failed, $WARN warnings"
        exit 1
      fi
      pass "Valid JSON"
      
      # --- Run all structural checks via Python ---
      RESULT=$(python3 - "$EVALS_FILE" "$EVALS_BASE_FILE" <<'PYEOF'
      import json
      import subprocess
      import sys
      
      # Assertions are graded by run-ab-evals.sh with `grep -qiE` after stripping a
      # leading (?i) — case-insensitive POSIX ERE, not Python's re. Validating them
      # with a different engine measures a different thing: `[^\n]` means one thing
      # in Python and another in ERE, and case-sensitive matching rejects patterns
      # the grader accepts. So the checks below call the same binary with the same
      # flags rather than approximating it.
      def _grep(args, text):
          try:
              return subprocess.run(
                  ["grep", *args], input=text, text=True,
                  stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL,
              ).returncode
          except (OSError, subprocess.SubprocessError):
              return None
      
      
      def _ere(pattern):
          return pattern[4:] if pattern.startswith("(?i)") else pattern
      
      
      def pattern_compiles(pattern):
          # grep exits 2 on a pattern it cannot parse and 1 on "no match".
          rc = _grep(["-qE", "--", _ere(pattern)], "")
          return rc is None or rc != 2
      
      
      def grader_matches(pattern, text):
          return _grep(["-qiE", "--", _ere(pattern)], text) == 0
      
      
      def identity(record, index):
          # Same key as retro-skill's check-eval-samples.py, so both halves of the
          # rule name an eval the same way. Not the position: format A
          # requires sequential ids, so inserting an eval renumbers every later one
          # and an index- or id-first key would report the whole file as changed.
          for key in ("eval_name", "name", "id"):
              value = record.get(key)
              if isinstance(value, (str, int)) and str(value).strip():
                  return f"{key}={value}"
          return f"index={index}"
      
      
      def has_samples(record):
          samples = record.get("samples")
          if not isinstance(samples, dict):
              return False
          passing = samples.get("passing")
          return isinstance(passing, str) and bool(passing.strip())
      
      
      def load_base(path):
          """Index the base copy by eval identity, or return a reason it is unusable.
      
          Returns (index, error). A base that cannot be read is an error, never an
          empty index: "the base could not be read" and "the base had no evals" would
          otherwise be the same state, and the second one silently demands samples
          from every eval in the file.
          """
          try:
              with open(path) as fh:
                  base_raw = json.load(fh)
          except OSError as exc:
              return None, f"cannot read base copy {path}: {exc.strerror or exc}"
          except ValueError as exc:
              return None, f"base copy {path} is not valid JSON: {exc}"
          if isinstance(base_raw, dict) and isinstance(base_raw.get("evals"), list):
              base_evals = base_raw["evals"]
          elif isinstance(base_raw, list):
              base_evals = base_raw
          else:
              return None, f"base copy {path} is neither an array nor an object with 'evals'"
          return {
              identity(r, i): r for i, r in enumerate(base_evals) if isinstance(r, dict)
          }, None
      
      
      with open(sys.argv[1]) as f:
          raw = json.load(f)
      
      base_path = sys.argv[2] if len(sys.argv) > 2 else ""
      base_index, base_error = (None, None)
      if base_path:
          base_index, base_error = load_base(base_path)
      
      # Detect format and normalize to list of evals
      if isinstance(raw, dict) and "evals" in raw:
          # Format A: {skill_name, evals: [...]}
          evals = raw["evals"]
          fmt = "A"
      elif isinstance(raw, list):
          # Format B: [...]
          evals = raw
          fmt = "B"
      else:
          print("FAIL|Top-level structure must be an array or object with 'evals' key")
          sys.exit(0)
      
      print(f"INFO|Detected format {'A (object with evals key)' if fmt == 'A' else 'B (top-level array)'}")
      print(f"INFO|Total evals: {len(evals)}")
      
      if base_error:
          print(f"FAIL|{base_error}")
      elif base_index is not None:
          print(
              f"INFO|Base comparison against {base_path} ({len(base_index)} evals): "
              "samples required on evals that are new or whose assertions changed"
          )
      else:
          print("INFO|No base copy (EVALS_BASE_FILE unset): samples stay optional")
      
      if not isinstance(evals, list):
          print("FAIL|'evals' must be an array")
          sys.exit(0)
      
      if len(evals) == 0:
          print("FAIL|No evals found")
          sys.exit(0)
      
      # Check eval count thresholds
      if len(evals) < 10:
          print(f"FAIL|Eval count {len(evals)} < 10 minimum")
      elif len(evals) < 15:
          print(f"WARN|Eval count {len(evals)} < 15 recommended")
      else:
          print(f"PASS|Eval count {len(evals)} >= 15")
      
      # Track names for duplicate check
      names = []
      ids_found = []
      has_ids = False
      
      for i, ev in enumerate(evals):
          label = f"eval[{i}]"
      
          if not isinstance(ev, dict):
              print(f"FAIL|{label}: not an object")
              continue
      
          # Name/ID check: accept 'name', 'eval_name', or 'id' (integer) as identifier
          name = str(ev.get("name") or ev.get("eval_name") or "").strip()
          if not name:
              # Fall back to id as identifier for Anthropic format
              eid = ev.get("id")
              if isinstance(eid, int):
                  name = f"id={eid}"
              else:
                  print(f"FAIL|{label}: missing or empty name/eval_name/id")
          if name:
              names.append(name)
      
          # Prompt check (accept 'prompt' or legacy 'input')
          prompt = ev.get("prompt") or ev.get("input") or ""
          if not prompt or not str(prompt).strip():
              print(f"FAIL|{label} ({name}): missing or empty prompt/input")
          else:
              print(f"PASS|{label} ({name}): has prompt")
      
          # ID check (optional but validated if present; must be integer)
          if "id" in ev:
              if not isinstance(ev["id"], int):
                  print(f"FAIL|{label} ({name}): id must be an integer")
              else:
                  has_ids = True
                  ids_found.append(ev["id"])
      
          # expected_output check (recommended for unified format, skip for legacy)
          is_unified = any(key in ev for key in ("eval_name", "id", "expectations", "expected_output"))
          if is_unified and "expected_output" not in ev:
              print(f"WARN|{label} ({name}): missing expected_output (recommended)")
      
          # Grading check: must have expectations OR assertions (or both)
          expectations = ev.get("expectations")
          assertions = ev.get("assertions")
          has_expectations = False
          has_assertions = False
      
          # Validate expectations (string[], 2+ items)
          if expectations is not None:
              if not isinstance(expectations, list):
                  print(f"FAIL|{label} ({name}): expectations must be an array")
              elif len(expectations) < 2:
                  print(f"FAIL|{label} ({name}): has {len(expectations)} expectations, need >= 2")
              else:
                  invalid_exp = 0
                  for j, e in enumerate(expectations):
                      if not isinstance(e, str):
                          invalid_exp += 1
                          print(f"FAIL|{label} ({name}): expectations[{j}] must be a string")
                      elif not e.strip():
                          invalid_exp += 1
                          print(f"FAIL|{label} ({name}): expectations[{j}] is empty")
                  if invalid_exp == 0:
                      has_expectations = True
                      print(f"PASS|{label} ({name}): {len(expectations)} valid expectations")
      
          # Assertions the grader can actually execute against a sample. Defined
          # before the block that fills it: the samples requirement below reads it
          # whatever path validation took, and an eval with no usable pattern is
          # exempt from that requirement.
          patterns = []
      
          # Validate assertions (object[] or string[], 2+ items)
          if assertions is not None:
              if not isinstance(assertions, list):
                  print(f"FAIL|{label} ({name}): assertions must be an array")
              elif len(assertions) < 2:
                  print(f"FAIL|{label} ({name}): has {len(assertions)} assertions, need >= 2")
              else:
                  invalid_assertions = 0
                  for j, a in enumerate(assertions):
                      if isinstance(a, str):
                          if not a.strip():
                              invalid_assertions += 1
                              print(f"FAIL|{label} ({name}): assertion[{j}] is empty")
                      elif isinstance(a, dict):
                          if "type" not in a:
                              invalid_assertions += 1
                              print(f"FAIL|{label} ({name}): assertion[{j}] missing 'type'")
                          val = a.get("value") or a.get("pattern") or ""
                          if not str(val).strip():
                              invalid_assertions += 1
                              print(f"FAIL|{label} ({name}): assertion[{j}] missing 'value' or 'pattern'")
                      else:
                          invalid_assertions += 1
                          print(f"FAIL|{label} ({name}): assertion[{j}] invalid type (not string or object)")
      
                  # A pattern only had to be a non-empty string until now, so an
                  # unbalanced group passed validation and first misbehaved wherever
                  # the eval was actually graded. Which assertions count is decided by
                  # the grader, not by a type name: run-ab-evals.sh greps
                  # `value or pattern` from EVERY assertion, so a broken pattern under
                  # type "tool_use", "content_regex" or any other label is graded and
                  # must be validated the same way. The type decides only the
                  # DIRECTION of the verdict -- `must_not` passes when the pattern is
                  # absent -- not whether the assertion is graded at all.
                  patterns = []
                  for j, a in enumerate(assertions):
                      if not isinstance(a, dict):
                          continue
                      pat = str(a.get("pattern") or a.get("value") or "")
                      if not pat.strip():
                          continue
                      if not pattern_compiles(pat):
                          # A *_contains assertion states a literal, and the value is
                          # usually not meant as a regex at all — `((` in a Concourse
                          # eval, `public function __construct(` in a PHP one. The
                          # grader still greps it as an ERE, so it never matches, but
                          # the defect is the grader's literal handling rather than
                          # the eval's, and failing here would turn CI red in repos
                          # this change does not fix. It warns instead.
                          literal = "contains" in str(a.get("type") or "")
                          if literal:
                              print(
                                  f"WARN|{label} ({name}): assertion[{j}] is not a valid POSIX ERE. "
                                  "run-ab-evals.sh greps every assertion with -qiE, including this "
                                  "one, so it can never match — escape the value or make it a regex"
                              )
                          else:
                              invalid_assertions += 1
                              print(
                                  f"FAIL|{label} ({name}): assertion[{j}] is not a valid POSIX ERE "
                                  "— grep rejects it, so the grader scores it as never matching"
                              )
                          continue
                      patterns.append((j, pat, a.get("type") == "must_not"))
      
                  # Optional self-check. Without it nothing distinguishes an assertion
                  # that discriminates from one that is inverted or vacuous: both look
                  # like a non-empty string. With it, the eval carries one answer that
                  # must satisfy every assertion and answers that must not.
                  samples = ev.get("samples")
                  if samples is not None and not isinstance(samples, dict):
                      invalid_assertions += 1
                      print(f"FAIL|{label} ({name}): samples must be an object, got {type(samples).__name__}")
                  elif isinstance(samples, dict):
                      unknown = sorted(set(samples) - {"passing", "failing"})
                      if unknown:
                          invalid_assertions += 1
                          print(
                              f"FAIL|{label} ({name}): samples has unknown key(s) {', '.join(unknown)} "
                              "— only 'passing' and 'failing' are read, so a typo would check nothing"
                          )
                      if not patterns:
                          invalid_assertions += 1
                          print(
                              f"FAIL|{label} ({name}): samples present but no assertion carries a "
                              "pattern — the self-check would verify nothing"
                          )
                      passing = samples.get("passing")
                      if passing is not None and (not isinstance(passing, str) or not passing.strip()):
                          invalid_assertions += 1
                          print(f"FAIL|{label} ({name}): samples.passing must be a non-empty string")
                      elif isinstance(passing, str) and passing.strip():
                          for j, pat, negated in patterns:
                              hit = grader_matches(pat, passing)
                              if hit == negated:
                                  problem = (
                                      "matches its own passing sample although it is a must_not "
                                      "assertion" if negated else
                                      "does not match its own passing sample — the eval would "
                                      "reject a correct answer"
                                  )
                                  invalid_assertions += 1
                                  print(f"FAIL|{label} ({name}): assertion[{j}] {problem}")
                      failing = samples.get("failing")
                      if isinstance(failing, str):
                          failing = [failing]
                      if failing is not None and not isinstance(failing, list):
                          invalid_assertions += 1
                          print(f"FAIL|{label} ({name}): samples.failing must be a string or an array")
                          failing = []
                      for k, bad in enumerate(failing or []):
                          if not isinstance(bad, str) or not bad.strip():
                              invalid_assertions += 1
                              print(f"FAIL|{label} ({name}): samples.failing[{k}] must be a non-empty string")
                              continue
                          if patterns and all(
                              grader_matches(pat, bad) != negated for _, pat, negated in patterns
                          ):
                              invalid_assertions += 1
                              print(
                                  f"FAIL|{label} ({name}): failing sample[{k}] satisfies every "
                                  "assertion — the eval would accept an answer it calls wrong"
                              )
      
                  if invalid_assertions == 0:
                      has_assertions = True
                      print(f"PASS|{label} ({name}): {len(assertions)} valid assertions")
      
          # Samples required on a new or tightened eval (retro-skill#92). Only
          # reached when a base copy was supplied, and only for an eval carrying a
          # pattern the grader can run against a sample.
          if base_index is not None and patterns and not has_samples(ev):
              key = identity(ev, i)
              previous = base_index.get(key)
              state = None
              if previous is None:
                  state = "is new"
              elif previous.get("assertions") != ev.get("assertions"):
                  state = "has changed assertions"
              if state:
                  print(
                      f"FAIL|{label} ({name}): {state} and carries no samples — add "
                      "samples.passing (and at least one samples.failing) so the "
                      "assertions are graded both ways here, not only at run time "
                      "(retro-skill#92). Untouched evals are unaffected."
                  )
      
          # Must have at least one grading mechanism
          if not has_expectations and not has_assertions:
              if expectations is None and assertions is None:
                  print(f"FAIL|{label} ({name}): missing grading (need expectations or assertions)")
              # else: already reported specific validation errors above
      
      # Duplicate names
      seen = set()
      dupes = set()
      for n in names:
          if n in seen:
              dupes.add(n)
          seen.add(n)
      
      if dupes:
          print(f"FAIL|Duplicate eval names: {', '.join(sorted(dupes))}")
      else:
          print(f"PASS|No duplicate eval names")
      
      # ID validation (if IDs are present)
      if has_ids:
          # Check for duplicates
          id_counts = {}
          for eid in ids_found:
              id_counts[eid] = id_counts.get(eid, 0) + 1
          dupe_ids = [k for k, v in id_counts.items() if v > 1]
          if dupe_ids:
              print(f"FAIL|Duplicate IDs: {dupe_ids}")
          else:
              print(f"PASS|No duplicate IDs")
      
          # Check sequential (1-based)
          numeric_ids = sorted([x for x in ids_found if isinstance(x, int)])
          if numeric_ids:
              expected = list(range(1, len(numeric_ids) + 1))
              if numeric_ids != expected:
                  gaps = set(expected) - set(numeric_ids)
                  extra = set(numeric_ids) - set(expected)
                  msg = ""
                  if gaps:
                      msg += f"missing: {sorted(gaps)}"
                  if extra:
                      if msg:
                          msg += ", "
                      msg += f"unexpected: {sorted(extra)}"
                  print(f"FAIL|IDs not sequential: {msg}")
              else:
                  print(f"PASS|IDs sequential (1-{len(numeric_ids)})")
      PYEOF
      )
      
      # --- Parse Python output ---
      while IFS='|' read -r level msg; do
        case "$level" in
          PASS) pass "$msg" ;;
          FAIL) fail "$msg" ;;
          WARN) warn "$msg" ;;
          INFO) echo "  INFO: $msg" ;;
        esac
      done <<< "$RESULT"
      
      # --- Trigger evals -----------------------------------------------------------
      # evals.json cannot detect a description that never routes: run-ab-evals.sh
      # pastes the whole SKILL.md into the system prompt, so the skill is force-fed
      # and routing never happens. A skill whose description omits the words its users
      # say passes every eval and is reached by nobody -- measured at 1 of 25 trials
      # for one fleet skill whose 30 evals were green throughout. See
      # references/materialization-contract.md, Rule 7.
      #
      # A warning, not an error: nothing in this repository runs these queries yet,
      # and failing a build for a missing file whose runner does not exist would be a
      # gate nobody can pass.
      TRIGGER_FILE="$(dirname "$EVALS_FILE")/eval_queries.json"
      if [[ -f "$TRIGGER_FILE" ]]; then
        # One reader, one verdict. Counting positives in a second call meant a
        # malformed entry raised there and `|| echo 0` turned the failure into
        # "0 positives", so the validator printed PASS for a file it could not read.
        TRIGGER_READ=$(python3 -c "
      import json, re, sys
      try:
          d = json.load(open(sys.argv[1]))
      except Exception as exc:
          print('ERR|not valid JSON: %s' % exc); raise SystemExit
      q = d.get('queries') if isinstance(d, dict) else d
      if not isinstance(q, list):
          print('ERR|no list of queries (expected a top-level list, or a queries key)')
          raise SystemExit
      bad = [i for i, x in enumerate(q) if not isinstance(x, dict) or 'query' not in x
             or not isinstance(x.get('should_trigger'), bool)]
      if bad:
          print('ERR|entries %s lack a query string or a boolean should_trigger'
                % ', '.join(str(i) for i in bad[:5]))
          raise SystemExit
      # A benchmark case identifier is a shape no user request carries, and it is
      # how the first trigger file written under this rule leaked: queries copied
      # from the benchmark that was measuring the skill, its case ids beside them.
      # Shaped like OFR-TYPO3-EXT-001: uppercase words, then a zero-padded
      # three-digit ordinal. Deliberately not four digits, so CVE-2024-1234 and
      # OWASP-A01-2021 -- which a security skill may legitimately quote -- do not
      # trip it.
      marks = sorted(set(re.findall(r'\\b[A-Z]{2,}(?:-[A-Z][A-Z0-9]*)+-\\d{3}\\b',
                                    json.dumps(q))))
      print('OK|%d|%d|%s' % (len(q), sum(1 for x in q if x['should_trigger']),
                             ','.join(marks[:5])))
      " "$TRIGGER_FILE" 2>&1)
      
        case "$TRIGGER_READ" in
          OK\|*)
            n_q=$(echo "$TRIGGER_READ" | cut -d'|' -f2)
            n_pos=$(echo "$TRIGGER_READ" | cut -d'|' -f3)
            case_ids=$(echo "$TRIGGER_READ" | cut -d'|' -f4)
            if [[ -n "$case_ids" ]]; then
              fail "$(basename "$TRIGGER_FILE") names benchmark case identifier(s): ${case_ids} - a trigger query is a user's request, and a case id is the answer key to the run measuring this skill (materialization-contract.md, Rule 7)"
            fi
            if [[ "$n_q" -eq 0 ]]; then
              warn "$(basename "$TRIGGER_FILE") carries no queries - an empty file is not a trigger test"
            else
              pass "$(basename "$TRIGGER_FILE"): $n_q trigger queries, $n_pos labelled should_trigger"
              if [[ "$n_pos" -eq "$n_q" ]]; then
                warn "every trigger query is a positive - without negatives the file cannot catch a description that fires on everything"
              fi
            fi
            ;;
          *)
            fail "$(basename "$TRIGGER_FILE"): ${TRIGGER_READ#ERR|}"
            ;;
        esac
      else
        warn "no eval_queries.json beside this file - evals.json measures what the skill does once loaded, never whether a real request reaches it (materialization-contract.md, Rule 7)"
      fi
      
      # --- Summary ---
      echo ""
      echo "---"
      echo "Results: $PASS passed, $FAIL failed, $WARN warnings"
      
      if [[ $FAIL -gt 0 ]]; then
        exit 1
      fi
      
      exit 0
      
    • validate-skill.sh 49.4 KB
      #!/usr/bin/env bash
      # validate-skill.sh - Validate Netresearch skill repository structure
      # Usage: ./validate-skill.sh [repo-root-path]
      #
      # Checks: SKILL.md frontmatter, word count, composer.json, plugin.json,
      #          cross-file consistency, required files
      # Env:    STRICT_README=1 (also true/yes, case-insensitive) promotes README heading misses from warnings to errors
      # Exit: 0 = valid, 1 = errors found
      
      set -euo pipefail
      
      REPO_DIR="${1:-.}"
      ERRORS=0
      WARNINGS=0
      NAME=""
      
      RED='\033[0;31m'
      GREEN='\033[0;32m'
      YELLOW='\033[1;33m'
      NC='\033[0m'
      
      error() { echo -e "${RED}ERROR:${NC} $1"; ((ERRORS++)) || true; }
      warning() { echo -e "${YELLOW}WARNING:${NC} $1"; ((WARNINGS++)) || true; }
      success() { echo -e "${GREEN}OK:${NC} $1"; }
      
      # Check python3 availability (required for JSON parsing)
      if ! command -v python3 &>/dev/null; then
          echo -e "${RED}ERROR:${NC} python3 is required for JSON parsing but not found in PATH"
          exit 1
      fi
      
      echo "Validating skill repository: $REPO_DIR"
      echo "========================================"
      
      # --- Discover every SKILL.md ---
      # A repo may ship several skills. Validating only the first one meant the rest
      # had no frontmatter check, no "Use when" check and no word count: matrix-skill
      # reported a single 498-word line for skills/matrix-administration while
      # matrix-communication sat at 896 words, over the cap, for as long as the repo
      # existed (issue #214). Every skill found is validated and every finding counts
      # toward the exit code.
      SKILL_FILES=()
      if [[ -f "$REPO_DIR/SKILL.md" ]]; then
          SKILL_FILES+=("$REPO_DIR/SKILL.md")
      fi
      for f in "$REPO_DIR"/skills/*/SKILL.md; do
          if [[ -f "$f" ]]; then
              # A skills/<name>/SKILL.md that is the root file (symlink or hardlink)
              # would otherwise be reported twice under two names.
              if [[ ${#SKILL_FILES[@]} -gt 0 && "$f" -ef "${SKILL_FILES[0]}" ]]; then
                  continue
              fi
              SKILL_FILES+=("$f")
          fi
      done
      
      # --- Per-skill checks ---
      # Every message names the file it is about, so a finding in a multi-skill repo
      # is attributable without counting output lines.
      validate_skill_md() {
          local skill_file="$1"
          local rel="${skill_file#"$REPO_DIR"/}"
          local skill_name=""
          # Declared local so nothing carries over from the previous skill in the loop.
          local closing_line frontmatter extra_fields field_names desc compat
          local desc_chars body_lines skill_body base unnamed linked_from ref_lines
          local relative_paths count skill_dir skill_dir_rel checkpoints_justified
          local untested misclassified f s base dangling token rel
          # After the local declarations, not before them: `local skill_dir` resets a
          # value assigned above it, so an earlier assignment silently becomes unset
          # and `set -u` then aborts the run mid-way -- which prints no summary and
          # reads as a repository that passed.
          skill_dir="$(dirname "$skill_file")"
          skill_dir_rel="${skill_dir#"$REPO_DIR"}"
          skill_dir_rel="${skill_dir_rel#/}"
          skill_dir_rel="${skill_dir_rel:-.}"
          success "SKILL.md found: $rel"
      
          # Frontmatter delimiter
          if head -1 "$skill_file" | grep -q "^---$"; then
              # Verify closing --- delimiter exists (within first 30 lines)
              closing_line=$(sed -n '2,30{/^---$/=}' "$skill_file" | head -1)
              if [[ -z "$closing_line" ]]; then
                  error "$rel frontmatter has opening --- but no closing --- delimiter"
              else
                  success "$rel has frontmatter"
              fi
      
              # Extract frontmatter fields (between first two --- lines)
              frontmatter=$(sed -n '2,/^---$/{ /^---$/d; p; }' "$skill_file")
      
              # Check frontmatter fields match Agent Skills spec
              # Allowed: name, description, license, compatibility, metadata, allowed-tools
              extra_fields=$(echo "$frontmatter" | grep -E "^[a-z_-]+:" | grep -vE "^(name|description|license|compatibility|metadata|allowed-tools):" || true)
              if [[ -z "$extra_fields" ]]; then
                  success "$rel frontmatter fields are valid per Agent Skills spec"
              else
                  field_names=$(echo "$extra_fields" | sed 's/:.*//' | tr '\n' ', ' | sed 's/,$//')
                  error "$rel frontmatter has non-spec fields: $field_names (allowed: name, description, license, compatibility, metadata, allowed-tools)"
              fi
      
              # allowed-tools: a skill that ships scripts should name the scripts, not
              # the interpreter. Bash(bash:*) pre-approves every bash command the skill
              # issues for the whole turn; a bare Bash covers every shell command there
              # is. ${CLAUDE_SKILL_DIR} is substituted in the frontmatter Bash rules as
              # well as in the body, so a rule can name the shipped scripts without
              # pinning an install path. See references/repository-quality-rules.md,
              # section "allowed-tools".
              # Warning, not error: the field is optional, the narrower form is a
              # judgement call for tools the agent runs itself, and a skill may have
              # reasons to keep an interpreter listed.
              if [[ -d "$(dirname "$skill_file")/scripts" ]]; then
                  # The spec accepts a plain scalar, a folded or literal block and a
                  # YAML list, in block or flow style, quoted or not, space- or
                  # comma-separated. Rather than widen a pattern per shape -- each
                  # review round found another legal spelling -- take the whole value
                  # and split it into entries.
                  #
                  # Collect: key line plus every continuation. Only a new top-level key
                  # ends the value, so a blank line inside a folded scalar keeps it
                  # open. Comment lines and trailing comments are dropped; they are not
                  # part of the value.
                  #
                  # Split: on whitespace, comma, bracket and quote, but only at
                  # parenthesis depth 0 -- Bash(git:*,make:*,bash:*) is one entry with
                  # commas inside it, and Bash(bash ${CLAUDE_SKILL_DIR}/scripts/*)
                  # carries a space.
                  at_tokens=$(echo "$frontmatter" | awk '
                      /^allowed-tools:/ { found = 1; sub(/^allowed-tools:/, ""); }
                      !found { next }
                      found && NR > 1 && !/^[[:space:]]/ && !/^-[[:space:]]/ && !/^$/ && !/^allowed-tools:/ { exit }
                      {
                          sub(/^[[:space:]]*#.*$/, "")
                          # A " #" outside quotes starts a YAML comment; the rest
                          # of the line is not part of the value, parentheses included.
                          sub(/[[:space:]]#.*$/, "")
                          depth = 0
                          for (i = 1; i <= length($0); i++) {
                              c = substr($0, i, 1)
                              if (c == "(") depth++
                              else if (c == ")") depth--
                              if (depth == 0 && (c == " " || c == "\t" || c == "," || \
                                                 c == "[" || c == "]" || c == "\"" || c == "'"'"'"))
                                  printf "\n"
                              else
                                  printf "%s", c
                          }
                          printf "\n"
                      }
                  ' | sed -e 's/^-$//' -e '/^$/d')
                  if [[ -n "$at_tokens" ]] && echo "$at_tokens" \
                      | grep -qE '^(Bash$|Bash\([^)]*\b(bash|sh|python3?|uv|node|perl|ruby):\*)'; then
                      warning "$rel ships scripts/ but allowed-tools grants an interpreter (or bare Bash) - name the scripts instead, e.g. Bash(\${CLAUDE_SKILL_DIR}/scripts/*); see repository-quality-rules.md"
                  fi
              fi
      
              # Check name field
              if echo "$frontmatter" | grep -q "^name:"; then
                  skill_name=$(echo "$frontmatter" | grep "^name:" | head -1 | sed 's/name: *//' | tr -d '"')
                  # The plugin.json comparison further down uses one name. Single-skill
                  # repos keep the behaviour they had; multi-skill repos skip that
                  # comparison anyway, since plugin.json names the plugin, not a skill.
                  if [[ -z "$NAME" ]]; then
                      NAME="$skill_name"
                  fi
                  # The spec forbids more than the character class does: no leading
                  # or trailing hyphen, and no consecutive hyphens. A name that only
                  # passes the class still fails a spec-conformant loader.
                  if [[ ! "$skill_name" =~ ^[a-z0-9-]{1,64}$ ]]; then
                      error "$rel name invalid (lowercase, hyphens, max 64): $skill_name"
                  elif [[ "$skill_name" == -* || "$skill_name" == *- ]]; then
                      error "$rel name must not start or end with a hyphen: $skill_name"
                  elif [[ "$skill_name" == *--* ]]; then
                      error "$rel name must not contain consecutive hyphens: $skill_name"
                  else
                      success "$rel name valid: $skill_name"
                  fi
              else
                  error "$rel missing 'name' field"
              fi
      
              # Check description field and prefix
              if echo "$frontmatter" | grep -q "^description:"; then
                  # Parse the *YAML value* of description so every valid scalar style
                  # (plain, single/double-quoted, block) is accepted as long as the
                  # parsed value starts with "Use when". Uses PyYAML when available,
                  # otherwise a stdlib-only fallback covering the common scalar styles,
                  # so the script keeps running with just python3 (no yq/PyYAML needed).
                  # When PyYAML is present it is authoritative: invalid YAML is
                  # reported (sentinel __PARSE_ERROR__), not silently re-parsed by the
                  # fallback. The stdlib-only fallback runs solely when PyYAML is
                  # absent, so the script still works with just python3.
                  desc=$(FRONTMATTER="$frontmatter" python3 <<'PYEOF' 2>/dev/null || echo "__PARSE_ERROR__"
      import os, re, sys
      
      fm = os.environ["FRONTMATTER"]
      
      try:
          import yaml
      except Exception:
          yaml = None
      
      if yaml is not None:
          # PyYAML available: trust it fully so semantics match CI exactly.
          try:
              data = yaml.safe_load(fm)
          except Exception:
              print("__PARSE_ERROR__")
              sys.exit(0)
          desc = data.get("description") if isinstance(data, dict) else None
          print(desc if desc is not None else "")
          sys.exit(0)
      
      # Fallback without PyYAML: best-effort for the common scalar styles
      # (plain, single/double-quoted, block). description: is a column-0 key.
      desc = None
      lines = fm.splitlines()
      for i, line in enumerate(lines):
          m = re.match(r"description:[ \t]*(.*)$", line)
          if not m:
              continue
          val = m.group(1).strip()
          if val[:1] in ("|", ">"):
              # Block scalar: first non-blank line that is indented into the block.
              # A column-0 (non-indented) line is the next sibling key -> empty body.
              for nxt in lines[i + 1:]:
                  if not nxt.strip():
                      continue
                  if not nxt[:1].isspace():
                      break
                  desc = nxt.strip()
                  break
          else:
              dq = re.match(r'"((?:[^"\\]|\\.)*)"[ \t]*(?:#.*)?$', val)
              sq = re.match(r"'((?:[^']|'')*)'[ \t]*(?:#.*)?$", val)
              if dq:
                  desc = dq.group(1)
              elif sq:
                  desc = sq.group(1).replace("''", "'")
              else:
                  # Plain scalar: strip a trailing ' #' comment (YAML needs the space).
                  desc = re.sub(r"[ \t]+#.*$", "", val)
          break
      
      print(desc if desc is not None else "")
      PYEOF
      )
                  if [[ "$desc" == "__PARSE_ERROR__" ]]; then
                      error "$rel frontmatter is not valid YAML (could not parse 'description')"
                  elif [[ "$desc" == Use\ when* ]]; then
                      success "$rel description starts with 'Use when'"
                  else
                      error "$rel description must start with 'Use when': ${desc:0:60}..."
                  fi
      
                  # The description is a ROUTER, not documentation. It is the only
                  # thing loaded at startup for every skill, so it decides whether
                  # this skill is ever consulted -- a gap here is the one failure
                  # that cannot be recovered later. The spec sets a hard 1024.
                  #
                  # The soft 500 is a nudge, not a defect: the official guidance is
                  # "a few sentences to a short paragraph" and explicitly says to
                  # err on the side of being pushy about listing contexts. Long is
                  # not automatically wrong; long AND full of workflow steps is.
                  if [[ "$desc" != "__PARSE_ERROR__" ]]; then
                      desc_chars=${#desc}
                      if (( desc_chars > 1024 )); then
                          error "$rel description is $desc_chars chars (spec hard limit 1024) - cut process detail, keep capability and trigger"
                      elif (( desc_chars > 500 )); then
                          warning "$rel description is $desc_chars chars - past 500 it is usually workflow narration; the description should say WHAT and WHEN, not HOW"
                      else
                          success "$rel description is $desc_chars chars"
                      fi
                  fi
              else
                  error "$rel missing 'description' field"
              fi
      
              # compatibility: spec caps it at 500 characters, and most skills should
              # not carry the field at all.
              if echo "$frontmatter" | grep -q "^compatibility:"; then
                  # Octal escapes for the two quote characters: a literal quote here
                  # would terminate the string it lives in (the same trap the awk
                  # programs in this file avoid the same way).
                  compat=$(echo "$frontmatter" | grep "^compatibility:" | head -1 \
                           | sed 's/compatibility: *//' | tr -d '\42\47')
                  if (( ${#compat} > 500 )); then
                      error "$rel compatibility is ${#compat} chars (spec max 500)"
                  fi
              fi
          else
              error "$rel missing frontmatter (must start with ---)"
          fi
      
          # Body size. The spec recommends "Keep your main SKILL.md under 500 lines"
          # and "< 5000 tokens recommended" for the instructions loaded on activation.
          # This used to count 500 WORDS over the WHOLE file -- a much tighter and
          # differently shaped budget, and one that charged the frontmatter to the
          # body. That inversion is the expensive part: the description is the
          # routing surface, so making it compete with the instructions for one
          # allowance buys a shorter description at the price of an undocumented
          # capability. Lines, body only.
          #
          # 300 is the WARN, not the target: past it, ask which lines are control
          # flow and which are reference material that belongs in references/.
          body_lines=$(awk 'BEGIN{d=0} /^---$/{d++; next} d>=2{print}' "$skill_file" | wc -l)
          if (( body_lines > 500 )); then
              error "$rel body is $body_lines lines (spec recommends under 500) - move reference material into references/"
          elif (( body_lines > 300 )); then
              warning "$rel body is $body_lines lines - past 300, split reference material out and keep SKILL.md the control plane"
          else
              success "$rel body is $body_lines lines"
          fi
          # Check for relative script paths that should use ${CLAUDE_SKILL_DIR}
          # Matches: uv run scripts/, python3 scripts/, python scripts/, bash scripts/, ./scripts/, sh scripts/
          # But ignores lines already using ${CLAUDE_SKILL_DIR}
          relative_paths=$(grep -nE '(uv run|python3?|bash|sh|\./)([ ]+)scripts/' "$skill_file" | grep -v 'CLAUDE_SKILL_DIR' || true)
          if [[ -n "$relative_paths" ]]; then
              count=$(echo "$relative_paths" | wc -l)
              warning "$rel has $count script reference(s) using relative paths instead of \${CLAUDE_SKILL_DIR}/scripts/"
          fi
      
          # --- Flat discovery ------------------------------------------------------
          # Agent Skills spec: "Keep file references one level deep from SKILL.md.
          # Avoid deeply nested reference chains." The rule exists because each hop is
          # a decision the agent may not make. SKILL.md is read in full on activation;
          # a reference is read only if SKILL.md said what it holds and when to open
          # it. A file reachable only through a second hop sits behind an unmarked
          # door. These are warnings, not errors: a chain is a smell, not a breach.
          if [[ -n "$skill_dir" && -d "$skill_dir" ]]; then
              skill_body=$(awk 'BEGIN{d=0} /^---$/{d++; next} d>=2{print}' "$skill_file")
      
              # Scripts are executed, never loaded into context, so naming one costs a
              # line and is the only chance the agent has of knowing it exists.
              if [[ -d "$skill_dir/scripts" ]]; then
                  unnamed=""
                  for f in "$skill_dir"/scripts/*; do
                      [[ -f "$f" ]] || continue
                      base="$(basename "$f")"
                      case "$base" in *.md|*.txt|README*) continue ;; esac
                      # Non-executable files are sourced libraries, not capabilities
                      # the agent invokes; their caller is what belongs in SKILL.md.
                      [[ -x "$f" ]] || continue
                      grep -qF "$base" <<<"$skill_body" || unnamed="$unnamed $base"
                  done
                  if [[ -n "$unnamed" ]]; then
                      warning "${skill_dir_rel}: script(s) not named in SKILL.md:${unnamed} - a script reachable only through a reference is found only if that reference is opened"
                  fi
              fi
      
              if [[ -d "$skill_dir/references" ]]; then
                  for f in "$skill_dir"/references/*.md; do
                      [[ -f "$f" ]] || continue
                      base="$(basename "$f")"
      
                      if ! grep -qF "$base" <<<"$skill_body"; then
                          linked_from=""
                          for other in "$skill_dir"/references/*.md; do
                              [[ -f "$other" && "$other" != "$f" ]] || continue
                              if grep -qF "$base" "$other"; then
                                  linked_from="$(basename "$other")"
                                  break
                              fi
                          done
                          if [[ -n "$linked_from" ]]; then
                              warning "${skill_dir_rel}: references/${base} is reachable only via references/${linked_from} - the spec asks for one level; link it from SKILL.md too"
                          else
                              warning "${skill_dir_rel}: references/${base} is not named in SKILL.md - nothing tells the agent it exists or when to read it"
                          fi
                      fi
      
                      # Agents preview long files rather than reading them whole, so a
                      # contents list is what makes the rest of a long reference
                      # visible at all.
                      ref_lines=$(grep -c "" "$f")
                      if (( ref_lines > 100 )) && ! grep -qiE '^#{1,3} +(contents|table of contents|overview|in this (file|document))' "$f"; then
                          warning "${skill_dir_rel}: references/${base} is $ref_lines lines with no Contents section - agents preview long files; add one so the rest is discoverable"
                      fi
                  done
              fi
          fi
      
          # --- Dangling references -------------------------------------------------
          # The mirror image of flat discovery: a path SKILL.md names must exist under
          # the skill directory. A reference that resolves to nothing is a door the
          # agent opens onto a wall -- it costs a tool call and returns an error where
          # the skill promised content. retro-skill#19 hit this in production: the
          # skill shipped with scripts/ and references/ outside the declared skill
          # directory, so every path in SKILL.md was dead on a faithful install.
          # Skill-relative directories per the Agent Skills spec (scripts/,
          # references/, assets/) plus evals/; `${CLAUDE_SKILL_DIR}/` resolves to the
          # skill directory and is accepted as a prefix. Globs, placeholders and
          # other variables are not paths and are skipped. An error since the fleet
          # was measured clean (skill-repo-skill#267): 40 GitHub and 37 GitLab skill
          # repositories, every finding fixed at its source first.
          #
          # Both spellings count: a backticked span and a Markdown link target. A
          # SKILL.md that names its references as links only -- the common shape --
          # would otherwise pass this check without a single path being looked at.
          if [[ -n "$skill_dir" && -d "$skill_dir" ]]; then
              dangling=""
              while IFS= read -r token; do
                  token="${token#\`}"
                  token="${token%\`}"
                  token="${token#](}"
                  token="${token%)}"
                  # `[x](<references/y.md>)` is a valid link. Only a wholly wrapped
                  # target is unwrapped, so the `references/<topic>.md` placeholder
                  # keeps its angle brackets and stays filtered out below.
                  if [[ "$token" == "<"*">" ]]; then
                      token="${token#<}"
                      token="${token%>}"
                  fi
                  token="${token%[.,;:)]}"
                  rel="${token#\$\{CLAUDE_SKILL_DIR\}/}"
                  # A link may point into a file: references/x.md#section is x.md.
                  rel="${rel%%#*}"
                  case "$rel" in
                      scripts/*|references/*|assets/*|evals/*) ;;
                      *) continue ;;
                  esac
                  # A bare directory name is prose about a kind of place ("keep
                  # fixtures under `assets/`", a project's own `assets/`), not a file
                  # the agent is told to open; dxp-frontend-license was flagged for a
                  # project directory this way.
                  case "$rel" in scripts/|references/|assets/|evals/) continue ;; esac
                  case "$rel" in *'*'*|*'<'*|*'>'*|*'{'*|*'}'*|*'$'*|*' '*|*'?'*) continue ;; esac
                  [[ -e "$skill_dir/$rel" ]] && continue
                  case " $dangling " in *" $rel "*) continue ;; esac
                  dangling="$dangling $rel"
              done < <(
                  grep -oE "\`[^\`]+\`" <<<"$skill_body" || true
                  grep -oE '\]\([^) ]+\)' <<<"$skill_body" || true
              )
              if [[ -n "$dangling" ]]; then
                  error "${skill_dir_rel}: SKILL.md names path(s) that do not exist under the skill directory:${dangling} - fix the path or ship the file"
              fi
          fi
      
          # checkpoints.yaml presence (warning only — many skills legitimately lack
          # one; a documented justification marker suppresses the warning per
          # add-checkpoints' suitability criteria, e.g. purely conceptual skills)
          checkpoints_justified=0
          for f in "$skill_file" "$REPO_DIR/README.md"; do
              if [[ -f "$f" ]] && grep -qiE "^[[:space:]]*([*-][[:space:]]+)?checkpoints:[[:space:]]*none[[:space:]]*\(justified" "$f"; then
                  checkpoints_justified=1
                  break
              fi
          done
          if [[ -f "$skill_dir/checkpoints.yaml" ]]; then
              success "checkpoints.yaml exists"
          elif [[ $checkpoints_justified -eq 1 ]]; then
              success "checkpoints.yaml absence is justified"
          else
              warning "checkpoints.yaml not found in ${skill_dir_rel} — add checkpoints (see add-checkpoints skill) or document opt-out with 'Checkpoints: none (justified — <reason>)' in SKILL.md or README.md"
          fi
      
          # --- Shipped scripts without a test ---
          # A skill's scripts are its executable surface, and until this check existed
          # nothing noticed when they had no test: across the fleet, 27 of 33 repos
          # that ship scripts had no test file at all. Referenced-by-name is a coarse
          # signal on purpose — it costs nothing and catches the "no test whatsoever"
          # case, which is the one that actually occurs.
          if [[ -d "$skill_dir/scripts" ]]; then
              untested=$(
                  shopt -s nullglob
                  for s in "$skill_dir"/scripts/*; do
                      [[ -f "$s" ]] || continue
                      base="$(basename "$s")"
                      if [[ -d "$REPO_DIR/tests" ]] && grep -rqF -- "$base" "$REPO_DIR/tests" 2>/dev/null; then
                          continue
                      fi
                      printf '%s ' "$base"
                  done
              )
              untested="${untested% }"
              if [[ -z "$untested" ]]; then
                  success "every script under ${skill_dir_rel}/scripts is referenced by a test"
              else
                  warning "no test references these script(s): ${untested} — add a test under tests/ (run by the tests.yml reusable) or the script ships unexercised"
              fi
          fi
      
          # --- LLM checkpoints that are mechanically verifiable ---
          # add-checkpoints reserves llm_reviews for "subjective requirements that
          # can't be mechanically verified". A prompt that opens a line with a
          # runnable command is describing a mechanical check in prose — GW-21 in
          # git-workflow shipped a whole shell pipeline inside an llm_review, so the
          # rule was never executable and never regressed visibly. Prose that merely
          # mentions a command inline ("Use `git log …` to analyze") does not match:
          # the command must start the line.
          #
          # An entry that legitimately keeps both halves — a mechanical checkpoint for
          # the decidable part, an LLM prompt for the judgement — declares it with a
          # `# mechanical-counterpart: <ID>` comment anywhere in its block, and is
          # then exempt.
          if [[ -f "$skill_dir/checkpoints.yaml" ]]; then
              misclassified=$(awk '
                  /^llm_reviews:/ { in_llm = 1; next }
                  /^[a-z_]+:/     { in_llm = 0 }
                  !in_llm         { next }
                  # A marker above the entry (2-space comment) belongs to the entry
                  # that follows; one inside the block belongs to the current entry.
                  /^  #.*mechanical-counterpart:/     { pending = 1; next }
                  /^ {3,}#.*mechanical-counterpart:/  { if (id != "") exempt[id] = 1; next }
                  /^  - id:/ { id = $3; if (pending) { exempt[id] = 1; pending = 0 } ; next }
                  /^[[:space:]]+(git|gh|grep|sed|awk|test|jq|yq|find|ls|python3?|composer|npm|curl)[[:space:]]/ {
                      if (id != "") hit[id] = 1
                  }
                  END { n = 0; for (i in hit) if (!(i in exempt)) ids[n++] = i
                        for (a = 0; a < n; a++) for (b = a + 1; b < n; b++)
                            if (ids[b] < ids[a]) { t = ids[a]; ids[a] = ids[b]; ids[b] = t }
                        for (a = 0; a < n; a++) printf "%s ", ids[a] }
              ' "$skill_dir/checkpoints.yaml")
              misclassified="${misclassified% }"
              if [[ -n "$misclassified" ]]; then
                  warning "${skill_dir_rel}/checkpoints.yaml: llm_reviews checkpoint(s) contain a runnable command: ${misclassified} — if the command decides the outcome, move it to mechanical (type: command); keep the LLM entry only for the judgement the command cannot make"
              fi
          fi
      }
      
      if [[ ${#SKILL_FILES[@]} -gt 0 ]]; then
          for skill_file in "${SKILL_FILES[@]}"; do
              validate_skill_md "$skill_file"
          done
      else
          error "SKILL.md not found (checked root and skills/*/)"
      fi
      
      # --- Shebang without the committed executable bit ---
      # Repo-wide, so it runs once rather than once per skill.
      #
      # ruff's EXE001 covers the Python case, but it does not fire on every developer
      # machine: the same pinned ruff, same command, same mode-0644 file passes
      # locally and fails on the runner (issue #235, mechanism unestablished). A local
      # "clean" is therefore not evidence, and the first signal is a red CI job on
      # someone else's push.
      #
      # This reads the INDEX rather than the working tree, so it answers the same
      # everywhere regardless of what the filesystem reports. Severity follows what CI
      # already enforces: an error for *.py, because ruff fails the build on exactly
      # these; a warning for *.sh, where nothing fails today and a shebang on a file
      # only ever invoked as `bash file` is merely decorative.
      if git -C "$REPO_DIR" rev-parse --git-dir >/dev/null 2>&1; then
          shebang_no_exec() { # shebang_no_exec <pathspec...>
              git -C "$REPO_DIR" ls-files -s -- "$@" 2>/dev/null \
                  | awk '$1=="100644"{ sub(/^[0-9]+ [0-9a-f]+ [0-9]+\t/, ""); print }' \
                  | while IFS= read -r f; do
                      [[ -n "$f" ]] || continue
                      case "$(head -c 2 "$REPO_DIR/$f" 2>/dev/null)" in
                          '#!') printf '%s ' "$f" ;;
                      esac
                  done
          }
      
          PY_NOT_EXEC="$(shebang_no_exec '*.py')"; PY_NOT_EXEC="${PY_NOT_EXEC% }"
          SH_NOT_EXEC="$(shebang_no_exec '*.sh')"; SH_NOT_EXEC="${SH_NOT_EXEC% }"
      
          if [[ -n "$PY_NOT_EXEC" ]]; then
              error "committed 100644 but carries a shebang: ${PY_NOT_EXEC} — ruff EXE001 fails the build on this, and a local ruff run does not reproduce it (issue #235). Fix: chmod +x <file> && git update-index --chmod=+x <file>, or drop the shebang if the file is only ever imported — a module is not a script, and making it executable settles the mismatch from the wrong side"
          elif [[ -z "$SH_NOT_EXEC" ]]; then
              success "every committed script with a shebang is mode 100755"
          fi
          if [[ -n "$SH_NOT_EXEC" ]]; then
              warning "committed 100644 but carries a shebang: ${SH_NOT_EXEC} — either make it executable (chmod +x && git update-index --chmod=+x) or drop the shebang if it is only ever run as \`bash <file>\`"
          fi
      fi
      
      # --- Required files ---
      for file in README.md LICENSE-MIT LICENSE-CC-BY-SA-4.0 .gitignore; do
          if [[ -f "$REPO_DIR/$file" ]]; then
              success "$file exists"
          else
              error "$file not found"
          fi
      done
      
      # Warn about old single LICENSE file
      if [[ -f "$REPO_DIR/LICENSE" ]] && [[ -f "$REPO_DIR/LICENSE-MIT" ]]; then
          warning "Old LICENSE file still exists alongside LICENSE-MIT — remove it"
      elif [[ -f "$REPO_DIR/LICENSE" ]] && [[ ! -f "$REPO_DIR/LICENSE-MIT" ]]; then
          warning "Single LICENSE file found — migrate to LICENSE-MIT + LICENSE-CC-BY-SA-4.0"
      fi
      
      # Release path. A GitHub repository releases through .github/workflows/release.yml.
      # A GitLab repository has no such file and must not have one: the claude-code-skill
      # CI component included from .gitlab-ci.yml creates the Release from the tag
      # pipeline (references/release-discipline.md, "GitLab (git.netresearch.de) skill
      # repos release on tag too").
      # Demanding release.yml there is an error nobody can act on.
      if [[ -f "$REPO_DIR/.gitlab-ci.yml" ]]; then
          # A `component:` entry naming it, outside a comment. A bare text match
          # would accept `# claude-code-skill component removed`.
          if grep -qE '^[^#]*component:[^#]*claude-code-skill' "$REPO_DIR/.gitlab-ci.yml"; then
              success "release path: the claude-code-skill CI component creates the Release from the tag pipeline"
          else
              error ".gitlab-ci.yml does not include the claude-code-skill CI component — nothing creates a Release when a tag is pushed"
          fi
      elif [[ -f "$REPO_DIR/.github/workflows/release.yml" ]]; then
          success "release.yml exists"
      else
          error ".github/workflows/release.yml not found"
      fi
      
      # No composer.lock
      if [[ -f "$REPO_DIR/composer.lock" ]]; then
          error "composer.lock must not exist in skill repos"
      else
          success "No composer.lock"
      fi
      
      # --- composer.json checks ---
      if [[ -f "$REPO_DIR/composer.json" ]]; then
          success "composer.json exists"
      
          # Type
          if grep -q '"type".*"ai-agent-skill"' "$REPO_DIR/composer.json"; then
              success "composer.json type is ai-agent-skill"
          else
              error "composer.json type must be 'ai-agent-skill'"
          fi
      
          # License SPDX expression
          COMP_LICENSE=$(python3 - "$REPO_DIR" <<'PYEOF' 2>/dev/null || echo ""
      import json, sys
      with open(f'{sys.argv[1]}/composer.json', 'r') as f:
          print(json.load(f).get('license', ''))
      PYEOF
      )
          if [[ "$COMP_LICENSE" == "(MIT AND CC-BY-SA-4.0)" ]]; then
              success "composer.json license is correct SPDX expression"
          else
              warning "composer.json license should be '(MIT AND CC-BY-SA-4.0)', got: $COMP_LICENSE"
          fi
      
          # Name must match GitHub repo name (netresearch/{repo-name})
          COMP_NAME=$(python3 - "$REPO_DIR" <<'PYEOF' 2>/dev/null || echo ""
      import json, sys
      with open(f'{sys.argv[1]}/composer.json', 'r') as f:
          print(json.load(f).get('name', ''))
      PYEOF
      )
          REPO_NAME=""
          if [[ -n "${GITHUB_REPOSITORY:-}" ]]; then
              REPO_NAME="${GITHUB_REPOSITORY#*/}"
          elif git -C "$REPO_DIR" remote get-url origin &>/dev/null; then
              REMOTE_URL=$(git -C "$REPO_DIR" remote get-url origin 2>/dev/null)
              REPO_NAME=$(basename "$REMOTE_URL" .git)
          fi
          if [[ -n "$REPO_NAME" ]]; then
              EXPECTED_NAME="netresearch/$REPO_NAME"
              if [[ "$COMP_NAME" == "$EXPECTED_NAME" ]]; then
                  success "composer.json name matches repo: $COMP_NAME"
              else
                  error "composer.json name must match repo name: expected '$EXPECTED_NAME', got '$COMP_NAME'"
              fi
          elif [[ "$COMP_NAME" =~ ^netresearch/.*-skill$ ]]; then
              success "composer.json name: $COMP_NAME (repo name check skipped - no git remote)"
          else
              error "composer.json name must match netresearch/{repo-name}: $COMP_NAME"
          fi
      
          # Plugin dependency
          if grep -q "composer-agent-skill-plugin" "$REPO_DIR/composer.json"; then
              success "composer.json requires skill plugin"
          else
              warning "composer.json should require netresearch/composer-agent-skill-plugin"
          fi
      
          # ai-agent-skill extra path(s) exist (supports both string and array values)
          SKILL_PATH_ERRORS=$(python3 - "$REPO_DIR" <<'PYEOF' 2>/dev/null || echo "ERROR"
      import json, os, sys
      repo_dir = sys.argv[1]
      data = json.load(open(os.path.join(repo_dir, 'composer.json')))
      val = data.get('extra', {}).get('ai-agent-skill', '')
      paths = val if isinstance(val, list) else [val] if val else []
      if not paths:
          print('MISSING')
      else:
          for p in paths:
              if not os.path.isfile(os.path.join(repo_dir, p)):
                  print('NOTFOUND:' + p)
              else:
                  print('OK:' + p)
      PYEOF
      )
          if [[ "$SKILL_PATH_ERRORS" == "MISSING" ]]; then
              error "composer.json missing extra.ai-agent-skill"
          elif [[ "$SKILL_PATH_ERRORS" == "ERROR" ]]; then
              error "composer.json extra.ai-agent-skill could not be parsed"
          else
              while IFS= read -r line; do
                  case "$line" in
                      OK:*) success "composer.json skill path exists: ${line#OK:}" ;;
                      NOTFOUND:*) error "composer.json skill path missing: ${line#NOTFOUND:}" ;;
                  esac
              done <<< "$SKILL_PATH_ERRORS"
          fi
      else
          error "composer.json not found"
      fi
      
      # --- plugin.json checks ---
      PLUGIN_FILE="$REPO_DIR/.claude-plugin/plugin.json"
      if [[ -f "$PLUGIN_FILE" ]]; then
          success "plugin.json exists"
      
          # Name matches SKILL.md name (only for single-skill repos)
          PLUGIN_NAME=$(python3 - "$PLUGIN_FILE" <<'PYEOF' 2>/dev/null || echo ""
      import json, sys
      with open(sys.argv[1], 'r') as f:
          print(json.load(f).get('name', ''))
      PYEOF
      )
          SKILL_COUNT=$(python3 - "$PLUGIN_FILE" <<'PYEOF' 2>/dev/null || echo "1"
      import json, sys
      with open(sys.argv[1], 'r') as f:
          print(len(json.load(f).get('skills', [])))
      PYEOF
      )
          if [[ "$SKILL_COUNT" -le 1 ]]; then
              if [[ -n "$NAME" ]] && [[ "$PLUGIN_NAME" == "$NAME" ]]; then
                  success "plugin.json name matches SKILL.md: $PLUGIN_NAME"
              elif [[ -n "$NAME" ]]; then
                  error "plugin.json name '$PLUGIN_NAME' does not match SKILL.md name '$NAME'"
              fi
          else
              success "plugin.json is multi-skill ($SKILL_COUNT skills), name check skipped"
          fi
      
          # Skills is array
          SKILLS_TYPE=$(python3 - "$PLUGIN_FILE" <<'PYEOF' 2>/dev/null || echo "unknown"
      import json, sys
      with open(sys.argv[1], 'r') as f:
          s = json.load(f).get('skills')
      print('array' if isinstance(s, list) else type(s).__name__)
      PYEOF
      )
          if [[ "$SKILLS_TYPE" == "array" ]]; then
              success "plugin.json skills is array"
      
              # Check each skill path exists as directory
              MISSING_PATHS=$(python3 - "$PLUGIN_FILE" "$REPO_DIR" <<'PYEOF' 2>/dev/null || true
      import json, os, sys
      with open(sys.argv[1], 'r') as f:
          data = json.load(f)
      for path in data.get('skills', []):
          full = os.path.join(sys.argv[2], path)
          if not os.path.isdir(full):
              print(path)
      PYEOF
      )
              if [[ -z "$MISSING_PATHS" ]]; then
                  success "All plugin.json skill paths exist"
              else
                  while IFS= read -r p; do
                      error "plugin.json skill path missing: $p"
                  done <<< "$MISSING_PATHS"
              fi
          else
              error "plugin.json skills must be an array (got: $SKILLS_TYPE)"
          fi
      
          # Author URL
          AUTHOR_URL=$(python3 - "$PLUGIN_FILE" <<'PYEOF' 2>/dev/null || echo ""
      import json, sys
      with open(sys.argv[1], 'r') as f:
          print(json.load(f).get('author', {}).get('url', ''))
      PYEOF
      )
          if [[ -z "$AUTHOR_URL" ]]; then
              error "plugin.json author.url is missing or empty; it must be https://www.netresearch.de"
          else
              AUTHOR_URL_CLEAN="${AUTHOR_URL%/}"
              if [[ "$AUTHOR_URL_CLEAN" == "https://www.netresearch.de" ]]; then
                  success "plugin.json author.url is correct"
              else
                  error "plugin.json author.url must be https://www.netresearch.de (got: $AUTHOR_URL)"
              fi
          fi
      else
          error ".claude-plugin/plugin.json not found"
      fi
      
      # --- Portable manifest checks (Agent Plugins 1.0.0) ---
      # Root ./plugin.json is the portable manifest every Agent Plugins client reads.
      # It is the source of truth for shared metadata; .claude-plugin/plugin.json is
      # generated from it by sync-plugin-manifest.sh. Its absence is an error: the
      # fleet finished adopting it on 2026-08-07, so a repo without one is a new gap,
      # not a repo still waiting its turn.
      PORTABLE_FILE="$REPO_DIR/plugin.json"
      if [[ -f "$PORTABLE_FILE" ]]; then
          PORTABLE_REPORT=$(python3 - "$PORTABLE_FILE" "$PLUGIN_FILE" "$REPO_DIR" <<'PYEOF' 2>&1 || true
      import json
      import os
      import re
      import sys
      
      portable_path, claude_path, repo_dir = sys.argv[1], sys.argv[2], sys.argv[3]
      
      SCHEMA_URL = "https://agent-plugins.org/schemas/1.0.0/plugin.schema.json"
      # Closed schema: agent-plugins.org/schemas/1.0.0/plugin.schema.json
      ALLOWED = ["$schema", "name", "version", "description", "author",
                 "homepage", "repository", "license", "keywords", "extensions"]
      SHARED = ["name", "version", "description", "author",
                "homepage", "repository", "license", "keywords"]
      NAME_RE = re.compile(r"^(?!.*(?:--|\.\.))[a-z0-9](?:[a-z0-9.-]*[a-z0-9])?$")
      
      out = []
      
      
      def err(msg):
          out.append("ERROR:" + msg)
      
      
      def ok(msg):
          out.append("OK:" + msg)
      
      
      try:
          with open(portable_path, encoding="utf-8") as fh:
              data = json.load(fh)
      except ValueError as exc:
          err(f"plugin.json is not valid JSON: {exc}")
          data = None
      except OSError as exc:
          err(f"plugin.json could not be read: {exc}")
          data = None
      
      if isinstance(data, dict):
          ok("plugin.json (portable Agent Plugins manifest) exists")
      
          if data.get("$schema") != SCHEMA_URL:
              err(f'plugin.json $schema must be "{SCHEMA_URL}" (got: {data.get("$schema")!r})')
          else:
              ok("plugin.json targets Agent Plugins 1.0.0")
      
          name = data.get("name")
          if not isinstance(name, str) or not name:
              err("plugin.json is missing the required 'name' field")
          elif len(name) > 64 or not NAME_RE.match(name):
              err(f"plugin.json name is invalid: {name!r} "
                  "(1-64 chars, lowercase a-z0-9 . -, must start and end alphanumeric, no -- or ..)")
          else:
              ok(f"plugin.json name valid: {name}")
      
          unknown = [k for k in data if k not in ALLOWED]
          if unknown:
              err("plugin.json has fields outside the Agent Plugins schema: "
                  + ", ".join(sorted(unknown))
                  + " (client-specific data belongs in 'extensions' or in .claude-plugin/plugin.json)")
          else:
              ok("plugin.json has no fields outside the Agent Plugins schema")
      
          for key in ("version", "description", "homepage", "repository", "license"):
              if key in data and not isinstance(data[key], str):
                  err(f"plugin.json {key} must be a string")
          if "keywords" in data and (not isinstance(data["keywords"], list)
                                     or not all(isinstance(k, str) for k in data["keywords"])):
              err("plugin.json keywords must be an array of strings")
          if "author" in data:
              author = data["author"]
              if not isinstance(author, dict):
                  err("plugin.json author must be an object with name/email/url")
              else:
                  extra = [k for k in author if k not in ("name", "email", "url")]
                  if extra:
                      err("plugin.json author allows only name, email, url — got: "
                          + ", ".join(sorted(extra)))
                  for k, v in author.items():
                      if not isinstance(v, str):
                          err(f"plugin.json author.{k} must be a string")
          if "extensions" in data:
              ext = data["extensions"]
              if not isinstance(ext, dict) or not all(isinstance(v, dict) for v in ext.values()):
                  err("plugin.json extensions must be an object of reverse-domain "
                      "namespace keys mapping to objects")
      
          # Parity with the generated Claude Code manifest.
          if os.path.isfile(claude_path):
              try:
                  with open(claude_path, encoding="utf-8") as fh:
                      claude = json.load(fh)
              except (OSError, ValueError):
                  claude = None
              if isinstance(claude, dict):
                  drift = [k for k in SHARED if k in data and claude.get(k) != data[k]]
                  if drift:
                      err(".claude-plugin/plugin.json is out of sync with plugin.json on: "
                          + ", ".join(drift)
                          + " — run skills/skill-repo/scripts/sync-plugin-manifest.sh")
                  else:
                      ok(".claude-plugin/plugin.json is in sync with plugin.json")
      
          # Agent Plugins clients discover skills only under skills/<name>/SKILL.md.
          skills_dir = os.path.join(repo_dir, "skills")
          found = []
          if os.path.isdir(skills_dir):
              found = sorted(d for d in os.listdir(skills_dir)
                             if os.path.isfile(os.path.join(skills_dir, d, "SKILL.md")))
          if found:
              ok(f"skills/ holds {len(found)} portable skill(s): " + ", ".join(found))
          else:
              err("no skills/<name>/SKILL.md found — Agent Plugins clients do not "
                  "discover a SKILL.md at the repository root")
      
      print("\n".join(out))
      PYEOF
      )
          while IFS= read -r line; do
              [[ -n "$line" ]] || continue
              case "$line" in
                  OK:*) success "${line#OK:}" ;;
                  ERROR:*) error "${line#ERROR:}" ;;
                  *) error "portable manifest check produced unexpected output: $line" ;;
              esac
          done <<< "$PORTABLE_REPORT"
      else
          error "plugin.json (portable Agent Plugins 1.0.0 manifest) not found at repo root — Agent Plugins clients (Cursor, Copilot, …) cannot load this plugin; see skill-repo skills/skill-repo/references/agent-plugins-compat.md"
      fi
      
      # --- README.md quality checks (warnings by default) ---
      # Heading misses are warnings unless STRICT_README=1 promotes them to errors.
      # The default must stay warnings-only: consumer repos run this script from
      # main via the reusable validate.yml, so flipping the default would break
      # their CI. Opt in per repo (or org-wide, later) by exporting STRICT_README=1.
      readme_heading_miss() { case "${STRICT_README:-0}" in 1|[Tt][Rr][Uu][Ee]|[Yy][Ee][Ss]) error "$1" ;; *) warning "$1" ;; esac; }
      
      # Required level-2 headings (whole-line match) per skills/skill-repo/references/readme-template.md
      README_REQUIRED_HEADINGS=(
          "What this skill solves"
          "Why this is a skill (model delta)"
          "Use when"
          "Expected outputs"
          "Context requirements"
          "Example prompts"
          "Related skills"
          "Installation"
          "Contributing"
          "License"
      )
      
      if [[ -f "$REPO_DIR/README.md" ]]; then
          if grep -q "Netresearch" "$REPO_DIR/README.md"; then
              success "README.md contains Netresearch reference"
          else
              warning "README.md should contain Netresearch credits"
          fi
      
          for heading in "${README_REQUIRED_HEADINGS[@]}"; do
              # Whole-line match only (avoids substring hits inside ### headings or prose)
              if grep -Fxq "## ${heading}" "$REPO_DIR/README.md"; then
                  success "README.md has ## ${heading}"
              else
                  readme_heading_miss "README.md missing section (exact line ## ${heading}) — see skill-repo-skill skills/skill-repo/references/readme-template.md"
              fi
          done
      
          # --- Hand-typed version badge ------------------------------------------
          # No release step reads README.md (bump-version.sh, check-version-parity.sh),
          # so a static shields.io version badge is never updated. Measured
          # 2026-09-26: 2 of 39 netresearch/*-skill repos carried one, and
          # typo3-testing-skill showed 3.0.0 while the latest release was v5.22.1.
          # Warning only: a stale badge misleads readers but breaks nothing.
          if grep -qF -- 'img.shields.io/badge/version-' "$REPO_DIR/README.md"; then
              warning "README.md has a static version badge (img.shields.io/badge/version-…) that no release step updates — use the live form https://img.shields.io/github/v/release/netresearch/<repo>?sort=semver linked to https://github.com/netresearch/<repo>/releases (see readme-template.md)"
          fi
      
          # Every documented slash-command should be enumerated in the README so the
          # command (and mode) list does not silently drift when one is added. Warning
          # only: the name match is heuristic and a skill may intentionally omit one.
          if [[ -d "$REPO_DIR/commands" ]]; then
              for cmd_file in "$REPO_DIR"/commands/*.md; do
                  [[ -e "$cmd_file" ]] || continue
                  cmd_name="$(basename "$cmd_file" .md)"
                  # -F: match the command name literally (a filename with regex
                  # metacharacters can't break or mis-match). -w: require word
                  # boundaries, so `/work-update` does not match inside
                  # `commands/work-update.md` or a URL like `netresearch/work-update`.
                  if grep -qFw -- "/${cmd_name}" "$REPO_DIR/README.md"; then
                      success "README.md references /${cmd_name}"
                  else
                      warning "README.md does not mention command /${cmd_name} (commands/${cmd_name}.md) — keep the README command/mode list in sync when adding commands or modes"
                  fi
              done
          fi
      
          # --- Install instructions must be runnable -------------------------------
          # A README documented steps that fail on the first command for as long as it
          # existed, and no gate could see it: an outside user reported it
          # (netresearch/retro-skill#90). Only FENCED command lines are considered — a
          # README may legitimately quote the wrong form in prose to warn against it,
          # and this one does.
          README_CMDS=$(awk '/^ {0,3}(```|~~~)/ { infence = !infence; next } infence' \
              "$REPO_DIR/README.md")
          MP_ADD=$(grep -E '^[[:space:]]*/plugin[[:space:]]+marketplace[[:space:]]+add[[:space:]]' \
              <<< "$README_CMDS" || true)
          if [[ -n "$MP_ADD" ]]; then
              # The ARGUMENT, not the line: a substring test makes
              # `add netresearch/foo-catalog` match the slug `netresearch/foo`, and
              # splitting on a single space mis-parses `install  <two spaces>`.
              MP_ADD_TARGETS=$(awk '{ for (i = 1; i <= NF; i++) if ($i == "add") { print $(i + 1); break } }' \
                  <<< "$MP_ADD")
              SELF_SLUG=""
              if [[ -n "${GITHUB_REPOSITORY:-}" ]]; then
                  SELF_SLUG="$GITHUB_REPOSITORY"
              elif SELF_URL=$(git -C "$REPO_DIR" remote get-url origin 2> /dev/null); then
                  # owner/repo from any remote form, including the scp-style
                  # git@host:group/repo.git and a GitLab subgroup path.
                  SELF_URL="${SELF_URL%.git}"
                  SELF_SLUG=$(awk -F'[:/]' '{ print $(NF - 1) "/" $NF }' <<< "$SELF_URL")
              fi
              # `marketplace add` resolves .claude-plugin/marketplace.json in the
              # TARGET repo. Pointing it at a repo that ships only plugin.json fails
              # with "Marketplace file not found" — provable locally when the target
              # is this repo.
              if [[ -n "$SELF_SLUG" && ! -f "$REPO_DIR/.claude-plugin/marketplace.json" ]] \
                  && grep -Fxq -- "$SELF_SLUG" <<< "$MP_ADD_TARGETS"; then
                  error "README.md: '/plugin marketplace add $SELF_SLUG' cannot work — the target must be a catalog repo shipping .claude-plugin/marketplace.json, and this repo ships .claude-plugin/plugin.json. Point it at netresearch/claude-code-marketplace (its entry names this repo as the source, so the code still comes from here)."
              else
                  success "README.md marketplace-add target is not this repo"
              fi
              # The add line alone registers a marketplace and installs nothing.
              MP_INSTALL=$(grep -E '^[[:space:]]*/plugin[[:space:]]+install[[:space:]]' \
                  <<< "$README_CMDS" || true)
              if [[ -z "$MP_INSTALL" ]]; then
                  warning "README.md documents '/plugin marketplace add' but no '/plugin install <plugin>@<marketplace>' line — following it leaves the reader with a registered marketplace and no plugin"
              else
                  MP_INSTALL_TARGETS=$(awk '{ for (i = 1; i <= NF; i++) if ($i == "install") { print $(i + 1); break } }' \
                      <<< "$MP_INSTALL")
                  while IFS= read -r target; do
                      [[ -n "$target" ]] || continue
                      case "$target" in
                          *@*/* | */*@*)
                              error "README.md: '/plugin install $target' — the part after @ is the marketplace NAME (e.g. netresearch-claude-code-marketplace), not owner/repo"
                              ;;
                          *@*) success "README.md install target names a marketplace: $target" ;;
                          *)
                              warning "README.md: '/plugin install $target' has no @<marketplace> suffix — ambiguous once more than one marketplace is configured"
                              ;;
                      esac
                  done <<< "$MP_INSTALL_TARGETS"
              fi
              # Not checkable here: whether the named catalog actually lists this
              # plugin. That needs the catalog fetched at validation time — see the
              # scheduled-job option in netresearch/skill-repo-skill#292.
          fi
      fi
      
      # --- Summary ---
      echo ""
      echo "========================================"
      echo "Validation Summary"
      echo "========================================"
      echo -e "Errors:   ${RED}$ERRORS${NC}"
      echo -e "Warnings: ${YELLOW}$WARNINGS${NC}"
      
      if [[ $ERRORS -eq 0 ]]; then
          echo -e "${GREEN}Skill repository is valid!${NC}"
          exit 0
      else
          echo -e "${RED}Skill repository has $ERRORS error(s) that must be fixed.${NC}"
          exit 1
      fi
      
  • templates
    • .github
      • workflows
        • npm-pack-smoke.yml.template 1.2 KB · in bundle
        • validate.yml.template 540 B · in bundle
    • auto-merge-deps.yml.template 237 B · in bundle
    • composer.json.template 417 B · in bundle
    • LICENSE-CC-BY-SA-4.0.template 839 B · in bundle
    • LICENSE-MIT.template 1.1 KB · in bundle
    • package.json.template 1.8 KB · in bundle
    • plugin.json.template 362 B · in bundle
    • pr-quality.yml.template 834 B · in bundle
    • pre-commit.template 434 B · in bundle
    • README.md.template 6.3 KB · in bundle
    • release.yml.template 1.8 KB · in bundle
  • checkpoints.yaml 15.7 KB
    # Checkpoints for skill-repo
    # Validates Netresearch skill repository structure and conventions
    
    version: 1
    skill_id: skill-repo
    
    preconditions:
      - type: file_exists
        target: .claude-plugin/plugin.json
      - type: command
        pattern: "find . -path '*/SKILL.md' -not -path './.skill-repo-tools/*' -not -path './node_modules/*' | head -1 | grep -q ."
    
    mechanical:
      # License files
      - id: SR-01
        type: file_exists
        target: LICENSE-MIT
        severity: error
        desc: "LICENSE-MIT must exist"
    
      - id: SR-02
        type: file_exists
        target: LICENSE-CC-BY-SA-4.0
        severity: error
        desc: "LICENSE-CC-BY-SA-4.0 must exist"
    
      - id: SR-03
        type: file_not_exists
        target: LICENSE
        severity: error
        desc: "Bare LICENSE file must not exist (use split licensing)"
    
      - id: SR-04
        type: contains
        target: LICENSE-MIT
        pattern: "Netresearch DTT GmbH"
        severity: error
        desc: "LICENSE-MIT must use correct entity name"
    
      - id: SR-05
        type: not_contains
        target: LICENSE-MIT
        pattern: "GmbH & Co. KG"
        severity: error
        desc: "LICENSE-MIT must not use old entity name"
    
      # Metadata
      - id: SR-06
        type: json_path
        target: composer.json
        pattern: ".license"
        severity: error
        desc: "composer.json must have license field"
    
      - id: SR-07
        type: contains
        target: composer.json
        pattern: "(MIT AND CC-BY-SA-4.0)"
        severity: error
        desc: "composer.json license must be SPDX compound expression"
    
      - id: SR-08
        type: contains
        target: .claude-plugin/plugin.json
        pattern: "(MIT AND CC-BY-SA-4.0)"
        severity: error
        desc: "plugin.json license must be SPDX compound expression"
    
      - id: SR-09
        type: json_path
        target: .claude-plugin/plugin.json
        pattern: ".version"
        severity: error
        desc: "plugin.json must have version field"
    
      - id: SR-09b
        type: not_contains
        target: composer.json
        pattern: '"version":'
        severity: error
        desc: "composer.json must NOT have a version field — Composer derives version from git tags. Adding it creates maintenance drift."
    
      - id: SR-10
        type: contains
        target: composer.json
        pattern: "ai-agent-skill"
        severity: error
        desc: "composer.json type must be ai-agent-skill"
    
      - id: SR-11
        type: contains
        target: composer.json
        pattern: "composer-agent-skill-plugin"
        severity: warning
        desc: "composer.json should require composer-agent-skill-plugin"
    
      # Required files
      - id: SR-12
        type: file_exists
        target: README.md
        severity: error
        desc: "README.md must exist"
    
      - id: SR-13
        type: file_exists
        target: .gitignore
        severity: error
        desc: ".gitignore must exist"
    
      - id: SR-14
        type: file_exists
        target: renovate.json
        severity: warning
        desc: "renovate.json should exist for dependency updates"
    
      - id: SR-15
        type: file_exists
        target: .github/workflows/release.yml
        severity: error
        desc: "Release workflow must exist"
    
      - id: SR-16
        type: file_exists
        target: .github/workflows/auto-merge-deps.yml
        severity: warning
        desc: "Auto-merge caller workflow should exist"
    
      - id: SR-17
        type: file_not_exists
        target: composer.lock
        severity: error
        desc: "composer.lock must not be committed in skill repos"
    
      # README quality
      - id: SR-18
        type: contains
        target: README.md
        pattern: "## License"
        severity: warning
        desc: "README should have License section"
    
      - id: SR-19
        type: contains
        target: README.md
        pattern: "## Installation"
        severity: warning
        desc: "README should have Installation section"
    
      - id: SR-20
        type: contains
        target: README.md
        pattern: "Netresearch"
        severity: warning
        desc: "README should reference Netresearch"
    
      # SKILL.md checks
      - id: SR-21
        type: command
        # Rewritten to satisfy the runner's command allowlist (no `;`, no `$()`):
        # pipe `find` → `head` → `xargs wc -w` → `awk` for the threshold check.
        # awk uses {if ($1<=500) ok=1} END {exit !ok} so empty input (no SKILL.md
        # found, or one filtered out by `-not -path`) exits non-zero — preventing
        # the previous vacuous-pass when no SKILL.md existed.
        # `-not -path './node_modules/*'` matches the precondition filter so SKILL.md
        # files inside dependencies don't trigger false positives.
        pattern: "find . -path '*/SKILL.md' -not -path './.skill-repo-tools/*' -not -path './node_modules/*' | head -1 | xargs -r wc -w | awk '{if ($1<=500) ok=1} END {exit !ok}'"
        severity: error
        desc: "SKILL.md must be under 500 words"
    
      # SKILL.md frontmatter checks
      - id: SR-22
        type: regex
        target: skills/*/SKILL.md
        pattern: "^name: [a-z][a-z0-9-]{0,63}$"
        severity: error
        desc: "SKILL.md frontmatter name must be lowercase, hyphens only, max 64 chars"
    
      - id: SR-23
        type: regex
        target: skills/*/SKILL.md
        pattern: 'description: "Use when '
        severity: error
        desc: "SKILL.md frontmatter description must start with 'Use when'"
    
      # Composer package checks
      - id: SR-24
        type: contains
        target: composer.json
        pattern: "ai-agent-skill"
        severity: error
        desc: "composer.json extra must reference ai-agent-skill entry point"
    
      - id: SR-25
        type: json_path
        target: .claude-plugin/plugin.json
        pattern: ".name"
        severity: error
        desc: "plugin.json must have name field"
    
      - id: SR-26
        type: json_path
        target: .claude-plugin/plugin.json
        pattern: ".skills"
        severity: error
        desc: "plugin.json must have skills array"
    
      - id: SR-27
        type: json_path
        target: .claude-plugin/plugin.json
        pattern: ".author"
        severity: warning
        desc: "plugin.json should have author field"
    
      # === WORKFLOW QUALITY CHECKS ===
      - id: SR-28
        type: regex
        target: .github/workflows/auto-merge-deps.yml
        pattern: 'pull_request_target'
        severity: error
        desc: "Auto-merge workflow must use pull_request_target trigger (not pull_request) for bot PR write permissions. LLM review SR-34 validates full correctness"
    
      # === CROSS-PLATFORM COMPATIBILITY ===
      - id: SR-29
        type: command
        pattern: "! grep -qr --include='*.sh' --exclude-dir=node_modules --exclude-dir=.git 'grep.*-P ' scripts/ .github/workflows/ 2>/dev/null"
        severity: warning
        desc: "Shell scripts must not use grep -P (Perl regex) — not available on macOS BSD grep. Use grep -E (extended regex) instead"
    
      # Caller release workflow checks
      - id: SR-30
        type: regex
        target: .github/workflows/release.yml
        # Use POSIX bracket class instead of `\s`. The runner's YAML parser
        # captures the raw text between double quotes, so `\\s` reaches grep
        # as literal backslash-s and never matches. `[[:space:]]` works under
        # both grep -E and grep -P regardless of YAML escape semantics.
        pattern: "^[[:space:]]*tags:"
        severity: error
        desc: "Release workflow must trigger on tag push (releases require a signed annotated tag pushed locally)"
    
      - id: SR-31
        type: command
        # Strip full-line comments before scanning so the docstring describing
        # the historical anti-pattern doesn't trip the check. The previous
        # `awk 'NF && $0 !~ /…/'` form contained `&&`, which the runner's
        # allowlist rejects as a command-chaining metacharacter — replaced
        # with a `grep -v` comment-strip pipeline that uses only `|`.
        # Policy enforced (per skill-repo-skill PR #15, which removed the
        # bump-job): no `workflow_dispatch:` trigger and no `--admin` merge
        # in the caller release workflow. Manual dispatch on the wrapper
        # would re-introduce the unsigned-tag bug the bump-job removal fixed.
        pattern: "! grep -v '^[[:space:]]*#' .github/workflows/release.yml | grep -qE '^[[:space:]]*workflow_dispatch:|--admin'"
        severity: warning
        desc: "Release workflow should not declare workflow_dispatch or use --admin merges — the auto-bump path produced unsigned tags. Use signed tag-push only (see skill-repo-skill PR #15)."
    
      - id: SR-37
        type: command
        # Mirrors scripts/check-version-parity.sh mechanically — the runner's
        # command allowlist forbids invoking repo scripts (only vendor/bin/*),
        # `bash` as base command, and `;`/`&&` anywhere in the pattern; hence
        # a single-line awk with one statement per pattern-action block.
        # Reads plugin.json .version, then compares every metadata.version
        # declared in skills/*/SKILL.md FRONTMATTER against it (the fm gate
        # keeps JSON examples in SKILL.md bodies from matching). Fails on any
        # mismatch or when plugin.json declares no version.
        # Not covered here (script-only): composer.json must-not-have-version,
        # tag-argument parity. Run the script for exact diagnostics.
        # The runner parses checkpoints line-by-line (no yq), so the pattern
        # cannot be folded across lines — exempt it from the 200-char rule.
        # yamllint disable-line rule:line-length
        pattern: awk 'FNR==1{fm=0} /^---$/{fm=1-fm} FILENAME~/json$/{if ($0 ~ /"version"/) gsub(/[",]/,"")} FILENAME~/json$/{if ($1 == "version:") pv=$2} fm{if ($1 == "version:") sv=$2} fm{if ($1 == "version:") gsub(/["\047]/,"",sv)} fm{if ($1 == "version:") if (sv != pv) bad=1} END{if (pv == "") bad=1} END{exit bad}' .claude-plugin/plugin.json skills/*/SKILL.md
        severity: error
        desc: "plugin.json version must match every SKILL.md metadata.version (run skills/skill-repo/scripts/check-version-parity.sh to verify before tagging)"
        tags: [release, versioning, parity]
    
    llm_reviews:
      - id: SR-32
        domain: repo-health
        prompt: |
          Review the README.md for a Netresearch skill repository:
          1. Does it clearly explain what the skill does?
          2. Does it have Installation section with marketplace/release/composer options?
          3. Does the License section correctly describe split licensing (MIT + CC-BY-SA-4.0)?
          4. Is the structure diagram accurate (shows LICENSE-MIT, LICENSE-CC-BY-SA-4.0)?
        severity: warning
        desc: "README structure and content quality"
    
      - id: SR-33
        domain: repo-health
        prompt: |
          Review the SKILL.md for a Netresearch skill:
          1. Does frontmatter have only name + description fields?
          2. Does description start with "Use when"?
          3. Is the content clear and actionable for an AI agent?
          4. Does it reference extended docs in references/ where appropriate?
        severity: warning
        desc: "SKILL.md clarity and completeness"
    
      # === REUSABLE WORKFLOW USAGE ===
      - id: SR-34
        domain: ci
        prompt: |
          Check if the skill repo uses reusable workflows from skill-repo-skill:
          1. Verify .github/workflows/ contains callers that delegate to
             netresearch/skill-repo-skill/.github/workflows/<name>.yml@main.
          2. Baseline expectation: at minimum a caller for validate.yml and
             release.yml. For the canonical list of all reusable workflows
             hosted in skill-repo-skill, fetch
             https://raw.githubusercontent.com/netresearch/skill-repo-skill/main/docs/ARCHITECTURE.md
             and read the "Reusable CI Workflows" section — do NOT rely on
             a list embedded in this prompt, since enumerating it here has
             caused drift before.
          3. Auto-merge for Dependabot/Renovate PRs is NOT hosted in
             skill-repo-skill. The reusable workflow lives at
             netresearch/.github/.github/workflows/auto-merge-deps.yml@main
             and is called via a thin local caller (commonly named
             .github/workflows/auto-merge-deps.yml in consuming repos).
             Flag any caller that points at netresearch/skill-repo-skill for
             auto-merge — that path is wrong.
          4. Flag any workflow that reimplements logic available in
             skill-repo-skill's reusable workflows (e.g., inline validation).
          5. Exception: repo-specific test workflows (e.g., validate-agents.yml
             in agent-rules-skill) that test repo-specific logic may use
             reusable workflows from skill-repo-skill but also contain inline
             steps.
          Report missing reusable workflow usage, inline reimplementations,
          and any caller that targets the wrong source repo.
        severity: warning
        desc: "Skill repos should use reusable workflows from skill-repo-skill for CI standardization"
    
      - id: SR-35
        domain: release
        severity: error
        desc: "Caller release workflow permissions must satisfy all shared workflow job requirements"
        prompt: |
          Review .github/workflows/release.yml as a caller of a reusable workflow.
          After the bump-job removal in skill-repo-skill PR #15, the shared workflow
          no longer creates release PRs, so `pull-requests: write` is NOT required.
          The release job needs:
            - contents: write       (for publishing the GitHub Release)
            - id-token: write       (OIDC for sigstore: cosign sign-blob + attest)
            - attestations: write   (GitHub native attestation API for SLSA build provenance)
          GitHub validates ALL job-level permissions at startup, even for skipped jobs.
          Check: Does the caller grant all three? Flag missing permissions as error
          (missing id-token: write or attestations: write causes the cosign / attest
          steps to fail; missing contents: write blocks release creation).
          Note: An older revision of this checkpoint required `pull-requests: write`
          for the now-removed auto-bump path. Do not flag its absence.
    
      - id: SR-36
        domain: release
        severity: warning
        desc: "plugin.json version must stay in sync with latest git tag"
        prompt: |
          Check version synchronization:
          1. Read version from .claude-plugin/plugin.json
          2. Get latest git tag (strip 'v' prefix)
          3. If plugin.json version is BEHIND the latest tag, flag as warning (version drift)
          4. If AHEAD, flag as info (pending release)
    
      # === SKILL CONTENT VALUE ===
      # Advisory (severity: warning) for two release cycles while thresholds
      # calibrate against the corpus; promotion to error is planned for NEW
      # skills only (netresearch/skill-repo-skill issue #144).
      - id: SR-38
        domain: skill-value
        severity: warning
        desc: "SKILL.md and reference content must earn its token cost under the six-category content value rubric (advisory during calibration)"
        prompt: |
          Classify the content value of the SKILL.md body and every
          references/*.md file in this skill repo:
          1. Read the canonical rubric — do NOT rely on a summary embedded
             in this prompt (an embedded copy would drift from the owning
             doc). Fetch
             https://raw.githubusercontent.com/netresearch/skill-repo-skill/main/skills/skill-repo/references/skill-quality.md
             and read the "Content value rubric" section. It defines six
             value categories: (1) org/project-specific knowledge,
             (2) version/ecosystem facts models get wrong, (3) retro-born
             failure patterns, (4) executable scripts/validators,
             (5) inference suppression, (6) anti-rationalization guards.
          2. Classify each section (heading-delimited block) of each file
             against the six categories. A section earns its place when it
             provides at least one category.
          3. A section that provides none is generic bloat. Apply the
             rubric's generate-it-with-a-prompt test: if a one-line prompt
             would regenerate the passage, the skill does not need to
             carry it.
          4. PROTECTED — never classify as generic: category 3 (failure
             patterns: symptom, cause, required behavior, verification,
             encoded from a real incident) and category 5 (inference
             suppression: "read file X, never guess Y" rules). Per the
             rubric these are first-class value even when they read as
             generic prose advice. Do not flag them.
          5. Report per file: the share of sections classified as generic
             bloat, quoting each flagged section so a human can verify the
             classification.
          This checkpoint is advisory while thresholds calibrate — report
          findings, do not fail the repo on them.
    
  • SKILL.md 5.7 KB
    ---
    name: skill-repo
    description: "Use when creating skill repositories, standardizing or validating skill repo structure, setting up composer/release workflows, configuring split licensing (MIT + CC-BY-SA-4.0), fixing plugin.json / SKILL.md validation or version-parity errors, or releasing a skill version (version bump, tagging)."
    license: "(MIT AND CC-BY-SA-4.0). See LICENSE-MIT and LICENSE-CC-BY-SA-4.0"
    compatibility: "Requires bash 4.3+, python3."
    metadata:
      author: Netresearch DTT GmbH
      version: "2.3.3"
      repository: https://github.com/netresearch/skill-repo-skill
    allowed-tools: Bash(${CLAUDE_SKILL_DIR}/scripts/*) Bash(bash ${CLAUDE_SKILL_DIR}/scripts/*) Bash(skills/skill-repo/scripts/*) Bash(bash skills/skill-repo/scripts/*) Read Write Glob Grep
    ---
    
    # Skill Repository Structure Guide
    
    ## Repository Structure
    
    ```
    {repo-name}/
    ├── plugin.json                  # portable manifest
    ├── .claude-plugin/plugin.json   # generated
    ├── skills/{name}/SKILL.md       # the control plane
    ├── README.md                    # human docs
    ├── LICENSE-MIT                  # code
    ├── LICENSE-CC-BY-SA-4.0         # content
    ├── composer.json                # PHP distribution
    ├── references/                  # detail, loaded on demand
    ├── scripts/                     # executables, never loaded
    └── .github/workflows/
        ├── release.yml              # tag-triggered
        ├── validate.yml             # validation caller
        └── auto-merge-deps.yml      # dep auto-merge caller
    ```
    
    ## Licensing (Split Model)
    
    | Path pattern | License |
    |---|---|
    | `skills/**/*.md`, `references/**`, `README.md`, `docs/**` | CC-BY-SA-4.0 |
    | `scripts/**`, `.github/workflows/**`, `*.sh`, `*.py`, `*.php` | MIT |
    | `composer.json`, `plugin.json`, config files | MIT |
    
    SPDX: `(MIT AND CC-BY-SA-4.0)`. Copyright: `Netresearch DTT GmbH`. No bare `LICENSE` — split files only.
    
    ## SKILL.md Frontmatter
    
    ```yaml
    ---
    name: skill-name
    description: "Use when <trigger conditions>"
    ---
    ```
    
    **Budgets** (spec): `name` ≤64, no doubled/edge hyphen, matches its directory. `description` ≤1024, warn past 500 — a router, not documentation. Body ≤500 lines, warn at 300. `compatibility` ≤500, usually omit.
    
    **Flat discovery**: references one level deep; SKILL.md names every `references/*.md` and every `scripts/` executable; `## Contents` past 100 lines. Rationale and trigger evals: [skill-architecture](references/skill-architecture.md). Audit: `audit-skills.sh` in the repository's top-level `scripts/` (not shipped with the skill).
    
    ## Manifests
    
    Root `plugin.json` ([Agent Plugins 1.0.0](https://agent-plugins.org), closed field set) is the source of truth; `sync-plugin-manifest.sh` generates `.claude-plugin/plugin.json` plus Claude-only keys — [agent-plugins-compat](references/agent-plugins-compat.md).
    
    ```json
    {
      "$schema": "https://agent-plugins.org/schemas/1.0.0/plugin.schema.json",
      "name": "skill-name",
      "version": "1.0.0",
      "license": "(MIT AND CC-BY-SA-4.0)",
      "author": {"name": "Netresearch DTT GmbH", "url": "https://www.netresearch.de"}
    }
    ```
    
    ## composer.json
    
    Name **must match GitHub repo**. Type `ai-agent-skill`. No `version` field (from git tags). No `composer.lock`.
    
    ```json
    {
      "name": "netresearch/{repo-name}",
      "type": "ai-agent-skill",
      "license": "(MIT AND CC-BY-SA-4.0)",
      "require": {"netresearch/composer-agent-skill-plugin": "*"},
      "extra": {"ai-agent-skill": "skills/{name}/SKILL.md"}
    }
    ```
    
    ## Reusable Workflow Callers
    
    Skill repos MUST delegate CI to skill-repo-skill reusable workflows:
    
    ```yaml
    # .github/workflows/validate.yml
    uses: netresearch/skill-repo-skill/.github/workflows/validate.yml@main
    ```
    
    Callers: `validate.yml`, `release.yml` (here); `auto-merge-deps.yml` (`netresearch/.github`). Auto-merge/pr-quality use `pull_request_target`. No inline Actions. Domain reusables: `docs/ARCHITECTURE.md`.
    
    ## Releasing
    
    Bump root `plugin.json` → sync → PR → merge → pull main → verify parity → signed tag → push → monitor Release. **Tag only after bump PR merges.** Multi-repo (>3) needs dry-run + approval. Never edit installed paths — [release-discipline](references/release-discipline.md).
    
    ## Installation
    
    Marketplace, release download, Composer, npm — commands and the npm `files` default in [installation-methods](references/installation-methods.md).
    
    ## Validation
    
    `scripts/validate-skill.sh` (layout, manifests, budgets, flat discovery). Shell portability: [authoring-ci-gotchas](references/authoring-ci-gotchas.md).
    
    Named here so they are findable: `bump-version.sh`, `check-version-parity.sh`, `sync-plugin-manifest.sh`, `roll-changelog.py`, `fleet-release-github.sh`, `migrate-licensing.sh`, `validate-evals.sh` — each with `--help`.
    
    ## References (`references/`)
    
    [agent-plugins-compat](references/agent-plugins-compat.md) · [installation-methods](references/installation-methods.md) · [plugin-hooks](references/plugin-hooks.md) · [composer-setup](references/composer-setup.md) · [release-discipline](references/release-discipline.md) · [review-replies](references/review-replies.md) · [skill-quality](references/skill-quality.md) · [repository-quality-rules](references/repository-quality-rules.md) · [readme-template](references/readme-template.md) · [skill-discovery-metadata](references/skill-discovery-metadata.md) · [validation-checklist](references/validation-checklist.md) · [marketplace-integration](references/marketplace-integration.md) · [materialization-contract](references/materialization-contract.md) · [authoring-ci-gotchas](references/authoring-ci-gotchas.md) · [skill-retirement](references/skill-retirement.md)
    
    
    ---
    
    > **Contributing:** <https://github.com/netresearch/skill-repo-skill>
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related