Claude Skill

author-skill

Use when authoring a NEW rsc skill or editing an existing one — scoping it to one job, writing the description that decides whether it ever loads, splitting the body into references/, writing its evals, auditing it against the rubric. NOT building a product feature (that is `spec

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download ericrisco-rsc-harness-skills_author-skill-953fef5.zip · 21 KB
Part of ericrisco/rsc-harness — 46 skills

Install

skills CLI npx skills add https://github.com/ericrisco/rsc-harness/tree/main/skills/author-skill
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ericrisco-rsc-harness@llmmart
Git git clone https://github.com/ericrisco/rsc-harness.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ericrisco/rsc-harness collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

author-skill — write skills that trigger and teach

This is the meta skill: it authors and edits the other skills in the rsc catalog. A skill is two things bolted together — a description that fires at the right moment, and a body that makes the agent better once it fires. Most skills fail on the first. Treat them as two separate engineering problems with two separate quality bars.

Where the SDD chain (specify → plan → … → ship) builds product, author-skill builds the tools that build product. Use it whenever a skill is born or edited.

Not this skill — delegate: a product feature specced or planned → ../specify/SKILL.md, ../plan/SKILL.md. An autonomous agent or tool-calling loop → ../building-agents/SKILL.md. Generic project docs or a wiki article → the ../harness/SKILL.md 02-DOCS engine. Bootstrapping a workspace or profiling the user → ../init/SKILL.md.

Read 02-DOCS/wiki/harness/user-profile.md and work at the accompaniment dial it records; ../init/SKILL.md owns that dial and sets it. With no profile, default to non-technical framing and ask for the technical level and the dial before going deep — skill authoring is itself a technical act, so many users want more narration here than they do elsewhere.

What a skill is (the anatomy)

skills/<id>/
├── SKILL.md              the body: frontmatter (name + description + origin) then the prose
├── references/           progressive-disclosure detail, loaded only when the body points to it
│   └── <topic>.md
├── evals/
│   ├── cases.yaml        trigger + capability test cases
│   └── README.md         how to run the evals, honestly
└── scripts/              optional; verify.sh + helpers for skills with a checkable artifact
    └── verify.sh

The frontmatter decides if the skill loads. The body decides how good the agent is once it does.

The description — the single highest-leverage line

The description sits in context on every turn the skill is installed, invoked or not. The body is only paid for when the skill fires; the description is paid for always. So length here is a cost, never a credit. A vague description is a skill that never fires; an over-broad one hijacks unrelated turns. Get this right before anything else.

Rules, all enforced:

  1. Third person, present tense. "Use when authoring a new skill…" — never "I help you…" or "You should…". The agent is reading about the skill.
  2. Discriminative, not exhaustive. Lead with a Use when … clause naming the situation, then only the capabilities that separate this skill from its neighbours. Do not append a Triggers: '…', '…' phrase list: the model matches on meaning, so a keyword bank in three languages buys nothing and is charged on every turn. The test is discrimination, not coverage — could a reader pick this skill over its nearest sibling from this line alone?
  3. Draw the boundary. End with a NOT <x> (that is <sibling>) clause, naming a sibling that actually exists under skills/. Negative space prevents hijacking as much as positive matching causes firing.
  4. Aim ≤ 350 characters; 1024 is the schema-enforced hard limit. One physical line, wrapped in double quotes, internal quotes escaped or avoided. If it does not parse, the skill does not load.
  5. origin: risco on its own line. This marks it as ours.
# Good — situation, the few capabilities that discriminate, a real boundary
description: "Use when X happens or the user shows symptom Y — doing A, fixing B, choosing C. NOT Z (that is `sibling`)."

# Bad — first person, no situation, no boundary; competes with every sibling on every turn
description: "I help you write great skills and make them work."

The full recipe, the budget tactics, and a worked before/after → references/description-recipe.md.

Progressive disclosure — the body is an index, not an encyclopedia

The body is loaded in full whenever the skill fires, so every line competes for the agent's attention. Write the smallest body that still routes correctly: 400 lines is a ceiling, not a target, and there is no floor — a skill that does its job in 60 lines beats the same skill padded to 200. Push anything long, reference-like, or rarely-needed into references/<topic>.md and link it inline at the point of use ("full table → references/foo.md").

Decide where a paragraph lives:

Put it in the body when… Move it to references/ when…
The agent needs it on every run It is needed only in a specific branch
It is a rule, a gate, or a decision point It is a long table, a catalog, or a template
It is short and load-bearing It is reference detail that would bloat the body
Cutting it would change behavior It is an example that illustrates but does not instruct

Every file under references/ must be linked from the body. An unlinked reference is never loaded, so it is dead weight in the package — link it or delete it.

The hybrid structure — when each piece earns its place

  • SKILL.md — always. Frontmatter + focused body.
  • references/ — only when the body genuinely needs offloaded depth. Do not create an empty references/ to look complete; a single-file skill is fine.
  • evals/ — always. cases.yaml + README.md. A skill with no evals is unverifiable and does not ship.
  • scripts/verify.sh — only when the skill produces a checkable artifact (code, config, copy with a ban-list). Process skills — those judged on the safety rails they install in the agent's behavior, like the SDD-phase skills or this one — do not ship a verify.sh; their evals carry a capability scenario instead.

Orientation footer (required in every new skill)

Every new skill MUST end with the orientation footer so the harness never leaves the user in seco. Append verbatim:


## Orientación (siempre)

Cierra cada turno con el **bloque-brújula** (📍 dónde estás · ✅ qué hiciste · 🧭 por qué · ➡️ siguiente, terminando en pregunta), calibrado al dial de `02-DOCS/wiki/harness/user-profile.md`. **Nunca termines en seco.** Protocolo completo: skill `orient` → `skills/orient/references/orientation-contract.md`. (Defiere a `suggest` el "¿instalo la skill que falta?".)

The full protocol lives once in the orient skill; the footer only references it.

The authoring workflow

Run in order. Each step gates the next.

  1. Name & scope. One skill, one job. Pick a short kebab-case <id> that is the job, not the domain. If you can not say the job in one sentence, the scope is wrong — split it. Check no sibling already owns this; if one half-owns it, decide edit the sibling vs new skill before writing.
  2. Draft the description. Per the rules above. This first, because writing it forces the scope clear. → references/description-recipe.md.
  3. Outline the body. Method, rules, decision points. Mark what becomes a reference.
  4. Write the body in the rsc voice (see below). Tag every code/example fence with a language. Add a checklist or decision table only where the flow actually branches — not as decoration. Add a short anti-patterns table.
  5. Extract references for anything long or branch-specific, and link each one inline.
  6. Write the evals — cases.yaml then README.md. → references/eval-authoring.md.
  7. Wire it into the rsc plumbing (tags, recommends, npm run manifest, and indexing any artifact in 02-DOCS/wiki/index.md — the Knowledge map; root CLAUDE.md keeps only a short pointer). → references/rsc-conventions.md.
  8. Self-audit against the rubric (below). Fix every miss or justify it.

The rsc voice

Match the catalog, do not invent a new register:

  • Direct, second-person-to-the-agent instruction ("Read the profile first", "Cut any section with no job").
  • A rule gets stated where it applies, with a one-line why that makes it obviously absolute — not a lecture, and not a shouted NON-NEGOTIABLE.
  • Concrete over abstract: a number, a path, a Bad→Good pair beats an adjective.
  • Original prose. Mine ideas from anywhere; the words are Eric's. Do not reproduce another ecosystem's signature artifacts or phrasing — no borrowed "1% chance" urgency blocks, no copied rationalization wording, no *-reviewer-prompt.md files, no verbatim flowcharts. The rsc identity is its own.
  • Cross-reference siblings by name or ../<sibling>/SKILL.md, only ones that actually exist.

Match the form to the failure

Before you write an instruction, name the failure it's meant to prevent — then pick the form that actually fixes that failure. The instinct is to write a prohibition ("don't do X") for everything. That instinct is wrong for most failures, and measurably counter-productive for one class: a prohibition aimed at output shape tends to summon the very thing it forbids (the model attends to the named token), and can do worse than saying nothing at all. Match deliberately:

The failure is… Use this form Why, and example
Discipline — the agent knows the rule but skips it under pressure (time, sunk cost, "just this once") The rule stated at the step it governs, with its why — plus one row in the anti-patterns table A rule in an appendix is skimmed; a rule in context is followed. "Do not ship a failing test — a red test merged is a lie in the suite", written at the ship step. Do not build a separate rationalization bank restating rules already in the flow; it is paid for on every load and read as decoration.
Wrong-shaped output — tone, verbosity, format, structure come out wrong Positive recipe / contract (show the target shape) A prohibition ("don't be verbose", "no marketing fluff") makes it more likely — the model fixates on the banned shape. Give the shape to hit instead: "Reply in ≤3 sentences, lead with the verdict." Demonstrate, don't forbid.
Omitted element — the agent forgets a required piece Required structural slot (a checklist item or a template field it must fill) You can't prohibit an absence. Make the slot mandatory so its emptiness is visible — a Done-of-done checkbox, a template section, a result-envelope field.
Conditional behavior — right action depends on the situation Predicate-keyed conditional ("When X → do Y; otherwise Z") A flat rule fires in the wrong context. Key the behavior to its trigger so the agent branches correctly instead of over- or under-applying.

So: an anti-patterns table earns its place when it names concrete failure modes the flow above does not already state. A table that re-lists rules from the body is pure cost — delete it and move each rule to its step. And when you catch yourself writing "don't make it X" about the shape of an output, rewrite it as the shape to hit.

The best-practice rubric (audit before shipping)

A skill ships only when every box is checked or a miss is consciously justified.

  • Frontmatter parses as YAML; name matches the directory <id>; origin: risco present.
  • Description third-person, Use when… lead, an explicit NOT … (that is sibling) boundary naming a real sibling, ≤ 350 chars target / ≤ 1024 hard limit — judged on discrimination against the nearest sibling, not coverage.
  • One job. The body never drifts into a second skill's territory; it delegates instead.
  • Body ≤ 400 lines — a ceiling, not a target, with no floor. Long/branch-specific material lives in references/.
  • Every references/ file linked inline from the body; none orphaned.
  • Every fence language-tagged; no placeholder/TODO prose; examples concrete.
  • Checklist/decision table only where a flow branches; an anti-patterns table present, naming failure modes rather than restating rules.
  • Accompaniment dial honored — reads the profile, adapts verbosity.
  • Artifacts under 02-DOCS/wiki/ and indexed in 02-DOCS/wiki/index.md (the Knowledge map; root CLAUDE.md keeps only a short pointer), if the skill produces any.
  • Concrete tooling delegated to the stack skills rather than reinvented.
  • evals present — cases.yaml (≥5 should_trigger incl. non-obvious, ≥4 should_not_trigger each with a real-sibling route_to, ≥1 capability with a must_include rubric) + an honest README.md. scripts/eval-lint.sh passes — but it only checks presence and the counts (≥5/≥4/≥1) and that those keys are lists; the route_to-points-at-a-real-sibling, non-obvious phrasings, and must_include quality are yours to verify here, not the linter's.
  • verify.sh present iff the skill has a checkable artifact; process skills rely on evals.
  • Every must_include item discriminates — answerable by the scenario's task, plausibly caused by the skill and plausibly missed without it. An item both arms fail measures nothing and lowers the absolute; it has turned a real PASS into a FAIL here. → references/eval-authoring.md.
  • Sibling links resolve — every ../x/SKILL.md points to a skill that exists.
  • Wired — tags + recommends set, npm run manifest re-run, and npm run validate / npm run manifest:check pass (manifest current, no dangling recommends).

Full rubric rationale and the rsc plumbing steps → references/rsc-conventions.md.

Two ship gates: document AND behavior

The rubric above scores the skill as a document. That is one of two gates — a skill ships only when both are green:

  1. Static — the rubric above / scripts/skill-rubric.md, weighted score ≥ 8.5.

  2. Behavioral — scripts/skill-behavior-rubric.md: run the skill on its capability scenarios with and without it loaded, blind-grade both outputs, require absolute ≥ 8.5 and lift ≥ +1.0. Run it:

    # 1) execute + grade — invoke the Workflow tool:
    #      scriptPath: scripts/skill-behavior-eval.workflow.js   args: "<skill-id>"
    #    save the returned object to /tmp/<skill>-raw.json
    # 2) score + gate (exit 0 pass / 1 fail):
    node scripts/skill-behavior-eval.js --score /tmp/<skill>-raw.json
    

    A failing lift means the body adds nothing a bare agent didn't already do — fix the body, don't game the checklist.

Anti-patterns

Failure mode Reality / fix
Description written last, once the body is done The description is why the body ever runs, and drafting it first forces the scope clear. Write it first, to bar.
Description padded for coverage — more phrasings, more languages, more verbs It is in context on every turn, invoked or not. Prune until it discriminates against the nearest sibling and stops.
One skill covering specify + plan + implement Multi-job skills trigger fuzzily and teach poorly. One skill, one job — split it.
Body grown past ~400 lines "because the topic is rich" The agent skims what it cannot hold. Extract a reference and link it inline.
A references/ folder added to look thorough, or a reference nothing links to An unlinked reference is never loaded — dead weight in the package. Link it at point of use or delete it.
Evals skipped: "I'll just test it by hand once" Unverifiable = does not ship. Write cases.yaml, near-misses with route_to included.
verify.sh added to a process skill for rigor A process skill has no artifact to grep. Its rigor is the capability eval.
../foo/SKILL.md linked to something not in this repo A dead link is a defect. Verify the directory exists under skills/.
Another catalog mirrored wholesale ("it's basically superpowers' writing-skills") Mine the idea, write it in the rsc voice. Copied artifacts or phrasing are a defect.

Project grounding (02-DOCS + CLAUDE.md)

When authoring produces a durable design note (a skill's scope decision, a description rationale worth keeping), persist it under 02-DOCS/wiki/sdd/ and index it in 02-DOCS/wiki/index.md (the Knowledge map; root CLAUDE.md keeps only a short pointer), per the ../harness/SKILL.md convention — never a stray file at the repo root. The skill's own evals/ is the executable record of intent; the wiki note is the human-readable why.

Files (rsc-harness)
  • evals
    • cases.yaml 6.9 KB
      skill: author-skill
      
      # Cases for the rsc-core "author-skill" meta skill.
      # author-skill authors or edits a SKILL.md to the rsc catalog bar: a discriminating
      # third-person description (~350 chars, valid YAML, origin: risco), a focused body with
      # progressive disclosure into references/, the best-practice rubric, and an
      # evals/cases.yaml + README for the NEW skill, plus the rsc plumbing (tags,
      # recommends, manifest.json, Knowledge map). It is NOT for building a
      # product feature (the SDD chain), NOT for agent loops (building-agents), and NOT
      # for generic project docs (harness).
      
      should_trigger:
        - prompt: "I keep re-explaining how we do code reviews to the agent every session. Turn that into a skill."
          why: "Canonical 'born a skill' moment phrased as a symptom, no skill jargon — author-skill's core job is turning a recurring instruction into a SKILL.md with a triggering description and evals."
      
        - prompt: "Escribe una skill nueva para generar informes de incidencias."
          why: "Non-English (Spanish) verbatim 'write a new skill' request — the description must match the user's actual language, and author-skill owns new-skill authoring end to end."
      
        - prompt: "My skill never triggers even when it obviously should. What's wrong with it?"
          why: "Non-obvious symptom phrasing (no mention of 'skill authoring') pointing straight at the description craft — diagnosing/repairing a description that fails to route is exactly this skill's highest-leverage area."
      
        - prompt: "This SKILL.md is 600 lines and feels bloated. Help me restructure it."
          why: "Edit/refactor of an existing skill toward progressive disclosure — author-skill's body-vs-references decision table and the 400-line ceiling are the direct fix."
      
        - prompt: "Write the evals/cases.yaml for the skill I just drafted, with trigger and capability tests."
          why: "Eval authoring for a skill is squarely owned here (the should_trigger/should_not_trigger/route_to and capability rubric minimums) — a deliberately specific, non-obvious trigger."
      
        - prompt: "Audit this skill against our best-practice rubric before I add it to the catalog and run npm run manifest."
          why: "Combines the rubric audit and the rsc plumbing (tags, recommends, manifest) — both are author-skill responsibilities; this exercises the wiring + rubric surface."
      
        - prompt: "The frontmatter on my skill won't parse and the description is over 1024 chars. Fix it."
          why: "A concrete description/frontmatter defect (valid YAML, <=1024) — the description recipe and the YAML/length verifier are author-skill's, not any stack skill's."
      
      should_not_trigger:
        - prompt: "I have a written spec for a new feature — turn it into an implementation plan."
          route_to: "plan"
          why: "Building a PRODUCT feature from a spec is the SDD chain (plan), not authoring a reusable skill. author-skill builds the tools that build product, not the product."
      
        - prompt: "Help me design an autonomous agent with a tool-calling loop and memory."
          route_to: "building-agents"
          why: "Designing an agent/tool-loop is building-agents. A 'skill' (a SKILL.md that loads into the agent) is a different artifact than an autonomous agent runtime."
      
        - prompt: "Bootstrap my workspace and figure out which rsc bundles I need."
          route_to: "init"
          why: "Profiling the user and recommending bundles is the front door (init). author-skill assumes a skill is being written; it does not gauge the user or recommend bundles to install."
      
        - prompt: "Consolidate my scattered project notes into a wiki article under 02-DOCS."
          route_to: "harness"
          why: "Generic docs/wiki consolidation is the harness 02-DOCS engine. author-skill writes SKILL.md + evals, not project knowledge articles, even though it persists the occasional design note via harness."
      
        - prompt: "Turn this fuzzy product idea into a clear spec of what and why."
          route_to: "specify"
          why: "Specifying a product's intent is the specify phase of the SDD chain. A spec is not a skill; author-skill would only be relevant if the artifact being authored were itself a skill."
      
      capability:
        - scenario: "A user says: 'I keep telling the agent how to write our PR descriptions. Make that into a skill called pr-describe.' Show how author-skill produces it to the rsc bar."
          must_include:
            - "Reads 02-DOCS/wiki/harness/user-profile.md for the accompaniment dial (or asks the two gauging questions if absent) and adapts verbosity"
            - "Confirms scope is one job and checks no existing sibling already owns PR-description authoring before creating a new skill"
            - "Drafts the description FIRST: third-person, 'Use when…' lead, only the capabilities that discriminate it from its nearest sibling, a 'NOT … (that is `sibling`)' boundary, valid single-line quoted YAML aiming <=350 chars (1024 hard limit), with origin: risco"
            - "Produces the smallest body that still routes (400 lines is a ceiling, not a target) with progressive disclosure — long/branch-specific detail pushed to references/ and pointed to inline, not an orphaned references/ folder"
            - "Language-tags every code/example fence and includes an anti-patterns table; adds a checklist/decision table only where a flow actually branches"
            - "Writes evals/cases.yaml meeting the minimums (>=5 should_trigger incl. non-obvious, >=4 should_not_trigger each with a real-sibling route_to, >=1 capability with a must_include rubric) plus an honest evals/README.md describing the two-axis run"
            - "Decides NO scripts/verify.sh because a PR-description skill is a process skill judged on its capability eval, OR includes verify.sh only if it emits a checkable artifact — and justifies the choice"
            - "States the rsc wiring: set tags + recommends in frontmatter, run npm run manifest to regenerate manifest.json, and confirm npm run validate / manifest:check / scripts/eval-lint.sh pass"
            - "Keeps the prose original in the rsc voice — does not reproduce another ecosystem's signature artifacts or phrasing"
      
        - scenario: "A user pastes a skill whose description is 'I help you with deployments' and says 'it fires on everything and also misses real deploy questions — fix the triggering.' Show the description repair."
          must_include:
            - "Diagnoses the two failure modes: first-person + no boundary causes hijacking; no situation clause causes misses"
            - "Rewrites to third-person with a 'Use when…' situation lead and concrete verb phrases"
            - "Keeps only the capabilities that discriminate deployment work from its nearest siblings, rather than padding a keyword list for coverage"
            - "Adds a 'NOT … (that is `sibling`)' boundary naming the real sibling(s) that own the excluded turns to stop hijacking"
            - "Verifies the result is valid single-line quoted YAML, aims <=350 chars against the 1024 hard limit, and keeps origin: risco"
            - "Recommends updating evals/cases.yaml (especially should_not_trigger with route_to) and re-running the triggering eval, since description edits are wording-sensitive"
      
    • README.md 4.1 KB
      # Eval harness — `author-skill`
      
      Evaluates the `author-skill` meta skill (the rsc-core tool that authors/edits other skills)
      on two axes: **triggering** (does it fire on the right prompts and stay quiet on near-misses)
      and **capability** (does loading it measurably improve the skill the agent produces). Cases
      live in `cases.yaml`. These run via an **agent harness**, not a deterministic script — a human
      or a driver agent feeds prompts to Claude Code and judges the result against the rubrics.
      
      ## What's in `cases.yaml`
      
      - `should_trigger` (7) — prompts that MUST invoke `author-skill`, including a non-obvious
        symptom phrasing ("my skill never triggers") and a non-English one ("escribe una skill nueva").
      - `should_not_trigger` (5) — near-misses that must route elsewhere; `route_to` names the real
        sibling that owns each (`plan`, `building-agents`, `init`, `harness`, `specify`).
      - `capability` (2) — scenarios with `must_include` rubrics to grade WITH vs WITHOUT the skill:
        authoring a new skill end-to-end, and repairing a broken triggering description.
      
      ## A. Triggering eval
      
      1. Load **only** `author-skill` into the agent (no other rsc skills available, so routing is honest).
      2. For each `should_trigger` prompt: open a fresh session, paste the prompt verbatim, and record
         whether `author-skill` activates (the agent should lead with scope/description work, not jump
         into building a product feature). Run **3–5 trials** per prompt.
      3. For each `should_not_trigger` prompt: same procedure, but a **pass** = `author-skill` does NOT
         fire. Where a `route_to` sibling exists, sanity-check that the prompt genuinely belongs there
         (e.g. "turn this spec into a plan" really is `plan`, not skill authoring).
      4. Score: a prompt passes if the **majority of its trials** go the expected way.
      
      **Pass bar:** ≥ 90% trigger accuracy across all `should_trigger` + `should_not_trigger` prompts
      (at most 1 of the 12 prompts may misbehave).
      
      ## B. Capability eval
      
      1. **Without the skill:** fresh session, skill NOT loaded, give the `scenario` prompt. Save output A.
      2. **With the skill:** fresh session, `author-skill` loaded, same prompt. Save output B.
      3. Grade each output against that scenario's `must_include` points — count points clearly covered.
      4. Repeat across **3 trials** per scenario per condition and average the coverage.
      
      **Pass bar:** WITH the skill covers **≥ 80%** of `must_include` points; WITHOUT the skill is
      materially lower (target a ≥ 30-point gap). If the skill doesn't measurably beat the baseline,
      the skill — or these rubrics — needs work.
      
      ## The headline differentiators
      
      What a WITH-skill answer should show that a baseline misses:
      
      - **Description first**, to the recipe: third-person `Use when…` lead, only the capabilities that
        discriminate it from its nearest sibling, a `NOT … (sibling)` boundary, valid single-line quoted
        YAML aiming ≤ 350 chars (1024 hard limit), `origin: risco`.
      - **Progressive disclosure** — the smallest body that still routes (400 lines is a ceiling, not a
        target) pointing into `references/`, not an encyclopedia and not an orphaned references folder.
      - **evals authored** to the minimums, with `route_to` siblings that actually exist.
      - **No `verify.sh` on a process skill** — rigor comes from the capability eval.
      - **rsc wiring** named: `tags` + `recommends` frontmatter, `npm run manifest` regenerated,
        `npm run validate` / `manifest:check` passing, `eval-lint.sh` passing.
      - **Original rsc voice** — no copied artifacts/phrasing from other skill ecosystems.
      
      ## Judging notes (honest caveats)
      
      - This is **LLM-as-judge / human-in-the-loop**, not deterministic. Use a consistent grader
        (same model + rubric) across A/B to keep the comparison fair.
      - `scripts/eval-lint.sh` deterministically checks only the *case-count minimums* of any
        `cases.yaml` this skill produces; it does not judge prose quality — that is the human/LLM grade.
      - Watch the key confusables: building a *product feature* (the SDD chain) and designing an
        *agent loop* (`building-agents`) are different artifacts than authoring a *skill*.
      - Re-run after any edit to `SKILL.md` or its `description`, since both axes are wording-sensitive.
      
  • references
    • description-recipe.md 5.7 KB
      # The description recipe
      
      The description is the only line the router reads on every turn to decide whether to load the skill — and it sits in context on every turn the skill is installed, invoked or not. It is a *retrieval problem*, not a marketing problem. Write it to be matched, not admired, and keep it short: length here is a cost, never a credit.
      
      ## The shape
      
      ```text
      "Use when <SITUATION / SYMPTOM> — <verb phrase>, <verb phrase>, <verb phrase>.
       NOT <out-of-scope thing> (that is `sibling`) [and NOT <other> (that is `sibling`)]."
      ```
      
      Three moves, in order:
      
      1. **Lead with the situation.** The first clause is a `Use when …` that names *when in the user's day* this fires — the moment, the symptom, the pain. The model matches situations better than nouns. "Use when a skill mis-triggers" beats "skill quality tool".
      2. **Name only the discriminating capabilities.** Right after the lead, name the few things the skill does that its neighbours do not, as verbs ("writing a description", "splitting the body", "repairing cases.yaml"). Stop there. Do **not** append a `Triggers: '…', '…'` phrase list: the model matches on meaning, so a keyword bank — especially one repeated in three languages — buys no routing accuracy and is charged on every turn.
      3. **Close with the boundary.** One or two `NOT … (that is `sibling`)` clauses. This is not optional decoration — negative space is what stops the skill from hijacking adjacent turns. Name the *real* sibling that owns the excluded job, and verify it exists under `skills/`.
      
      ## The test: discrimination, not coverage
      
      Do not ask "does this cover every way someone might phrase it?" — ask **"could a reader pick this skill over its nearest sibling from this line alone?"** Coverage is what tempts you to pad; discrimination is what actually routes. If two versions route the same, the shorter one is better.
      
      ## The hard constraints
      
      - **Aim ≤ 350 characters; 1024 is the schema-enforced hard limit.** Over budget = trim verb phrases first, never the boundary.
      - **Valid single-line quoted YAML.** The description is one physical line wrapped in double quotes. No raw newlines inside the value. Avoid characters that break YAML in double quotes; if you need an apostrophe inside, it is fine (single quotes are literal inside double-quoted YAML). Never put an unescaped `"` inside.
      - **Third person, present tense.** The agent reads *about* the skill. "Use when…", "Triggers on…", "Knows…". Never "I", never "you should".
      - **State when it fires and where the boundary is — never the workflow.** The description says *when* to fire and *what it is not*. It must **not** summarize the procedure the body owns (the steps, the phase count, "does X then Y then Z"). This is not a style nit — it changes behavior. A description that pre-summarizes the steps becomes a *substitute* for reading the body: the agent acts on the summary and skips the real instructions. (Observed in the wild: a skill whose description said it runs *two* reviews made the agent run only *one*, because the summary was treated as the spec; deleting the workflow summary made the agent read the body and do both.) The verb phrases in move #2 name *capabilities* ("repairing cases.yaml"), not an ordered recipe ("first lint, then split, then validate"). If your description tells the reader how the skill works step by step, cut it back to situation + boundary.
      
      ## Budget tactics when you are over
      
      In priority order, cut:
      
      1. Verb phrases that name a capability the nearest sibling shares — they add no discrimination.
      2. Verb phrases already implied by the `Use when` lead.
      3. Any "Knows…" / "Understands…" clause.
      4. Shorten the boundary to one `NOT` clause naming the single most-confused sibling.
      
      Never cut: the `Use when` lead, or the last `NOT` boundary.
      
      ## Worked before → after
      
      **Before** (first person, no situation, no boundary — 71 chars, would route badly):
      
      ```yaml
      description: "I help you write and improve skills so they work well."
      ```
      
      Problems: first person; no `Use when`; nothing that says *when* this fires; no boundary, so it competes with `building-agents`, `specify`, and `init` on every "make a thing" turn.
      
      **After** (third person, situation + discriminating verbs + boundary, valid YAML, ~300 chars):
      
      ```yaml
      description: "Use when authoring a NEW skill or editing an existing one — writing the description that decides whether it loads, splitting a long body into references/, repairing evals/cases.yaml, or fixing a skill that never fires. NOT building a product feature (that is `specify`) and NOT designing an agent loop (that is `building-agents`)."
      ```
      
      ## Quick test
      
      Before committing a description, ask:
      
      - Could the reader tell from this line alone *when* to fire it? (situation present)
      - Could the reader pick it over its nearest sibling from this line alone? (discriminates)
      - Does it say what it is **not**, naming a sibling that exists? (boundary present)
      - Is anything in it there for coverage rather than discrimination? (cut it)
      - Does it describe *when/what-not*, and **never** the step-by-step procedure? (no workflow summary — if a reader could skip the body and act on the description alone, it leaks the workflow)
      - Does it parse as YAML and fit 1024? (run the check)
      
      If any answer is no, it is not done.
      
      ## Verify the YAML and length
      
      ```bash
      python3 - "$PWD/skills/<id>/SKILL.md" <<'PY'
      import sys, yaml
      p = sys.argv[1]
      text = open(p).read()
      fm = text.split('---', 2)[1]
      meta = yaml.safe_load(fm)
      d = meta["description"]
      assert meta.get("origin") == "risco", "missing origin: risco"
      print("name:", meta["name"])
      print("description chars:", len(d))
      assert len(d) <= 1024, "description over 1024"
      print("OK — parses, origin present, <=1024")
      PY
      ```
      
    • eval-authoring.md 8.1 KB
      # Authoring the evals
      
      A skill without evals is unverifiable, and unverifiable means it does not ship. Evals test two separate things, and you must cover both:
      
      - **Triggering** — does the skill fire on the right prompts and stay quiet on near-misses?
      - **Capability** — once it fires, does the agent measurably do better than without it?
      
      Everything lives in `evals/cases.yaml` (the cases) and `evals/README.md` (how to run them, honestly).
      
      ## The minimums
      
      - `should_trigger` — **≥ 5** prompts that MUST load the skill. Include at least one **non-obvious** phrasing (a symptom, not the skill's name) and ideally a non-English one.
      - `should_not_trigger` — **≥ 4** near-miss prompts that must NOT load it. **Each needs a `route_to`** naming the real sibling that *should* own it (or `none` if no sibling does, with a why).
      - `capability` — **≥ 1** scenario with a `must_include` rubric of concrete points the answer must cover.
      
      ### What `scripts/eval-lint.sh` actually checks (and what it doesn't)
      
      `scripts/eval-lint.sh` parses every `cases.yaml` and fails the build only on the **structural minimums**: that `evals/cases.yaml` exists, that `should_trigger`, `should_not_trigger`, and `capability` are present as lists, and that their item counts meet **≥ 5 / ≥ 4 / ≥ 1**. (Without python3+PyYAML it degrades to a presence-only key check.) Run it before shipping to catch a missing or undersized section.
      
      It also resolves every `route_to`: each value must be a real catalog skill id, `none`, or `external:<name>` for a skill that deliberately lives outside this catalog — a stale route now fails the build. What it still does **not** read: whether each `should_not_trigger` actually carries a `route_to`, whether your `should_trigger` set includes a genuinely non-obvious or non-English phrasing, and whether each `capability` scenario has a real `must_include` rubric of gradeable points. Those are **author and review responsibilities**: verify them by hand (and in the self-audit / code-review pass) before shipping. A green eval-lint means the shape is right, not that the cases are good.
      
      ## cases.yaml structure
      
      ```yaml
      skill: <id>
      
      # A comment block stating what the skill IS and ISN'T helps the grader stay honest.
      
      should_trigger:
        - prompt: "A verbatim prompt a real user would type."
          why: "Why this MUST route here, and which differentiator of the skill it exercises."
        # … ≥ 5 total, with one non-obvious symptom phrasing and one non-English
      
      should_not_trigger:
        - prompt: "A near-miss that looks close but belongs elsewhere."
          route_to: "sibling-id"   # a real catalog id, "none", or "external:<name>" — checked by eval-lint
          why: "Why it is NOT this skill and why the sibling owns it."
        # … ≥ 4 total
      
      capability:
        - scenario: "A concrete situation; describe what the agent is asked to do."
          must_include:
            - "A specific behavior the WITH-skill answer must show."
            - "Another concrete, gradeable point — name files/paths/rules, not vibes."
            # … enough points to distinguish a skilled answer from a baseline one
      ```
      
      ## Writing good `should_trigger` cases
      
      - Use **verbatim user prompts**, not descriptions of prompts. The eval pastes them as-is.
      - Spread across the skill's real surface: the obvious ask, an edit/fix ask, a symptom ("X never works"), a non-English phrasing.
      - The **non-obvious** case is the important one — it proves the description matches symptoms, not just the skill's own name. If every trigger contains the skill's name, the description is too literal.
      
      ## Writing good `should_not_trigger` cases
      
      These are where descriptions get sharpened. Each near-miss should be genuinely tempting — adjacent in topic but owned by a sibling. The `route_to` must name a skill that **exists in this repo** — `eval-lint` fails the build otherwise. If the right owner genuinely ships outside this catalog, say so with `external:<name>` rather than inventing an id: the prefix makes the claim visible instead of indistinguishable from drift. Pick the siblings most likely to be confused with this one and write a case that disambiguates each.
      
      For `author-skill`, the natural confusables are `specify`/`plan` (building a feature, not a skill), `building-agents` (agent loops), `harness` (generic docs/wiki), and `init` (bootstrapping). Route each near-miss to whichever it truly belongs to.
      
      ## Writing good `capability` cases
      
      The `must_include` points are the differentiators — the specific things a *good* answer shows that a baseline answer misses. Make them **gradeable**: name the rule, the file path, the structural choice. "Writes a good skill" is not gradeable; "produces a third-person description under ~350 chars with a `Use when` lead and a `NOT … (sibling)` boundary naming a sibling that exists" is.
      
      For a process skill (no `verify.sh`), the capability scenario is the *primary* rigor — it is how you prove the safety rails actually change behavior. Make it count.
      
      ## A `must_include` item must discriminate, or it subtracts
      
      Measured on 2026-08-18, running the behavioral eval on five skills for the first time. Two items
      added to `testing-py` and `testing-web` came back **unsatisfied in both arms** — treatment and
      baseline alike. They demanded the answer *name* a tool (`mutmut`, Stryker); both arms did the
      behaviour (planted a bug, checked the suite caught it) without naming one. Contributing nothing to
      lift and still counting against the coverage term, they dragged `testing-py` from **PASS (9.0,
      lift +2.2) to FAIL (7.5, lift +1.7)**. The rule they encoded was fine; the item was not.
      
      Two tests before an item earns its place:
      
      1. **Answerable by the scenario's task.** If the scenario says "write this test file", then "names
         the mutation tool" is not part of that deliverable and no competent answer will contain it.
      2. **Discriminating.** The skill must plausibly cause it and a bare agent must plausibly miss it.
         An item both arms satisfy measures the model; one both arms fail measures nothing and lowers the
         absolute.
      
      **The symptom, so you can catch it from a scorecard:** an item unsatisfied in *both* arms is
      suspect. Nine times out of ten it is written as a spelling ("mentions X") rather than a behaviour
      ("does not treat coverage as proof the suite detects bugs"). Rewrite it as the behaviour, or delete
      it — do not leave it in as an aspiration, because the aggregate cannot tell an aspiration from a
      failure.
      
      The same run showed the other side: the equivalent item in `testing-go`, where the skill gives a
      concrete *procedure* rather than a tool name, discriminated hard and nearly doubled the lift
      (1.5 → 2.7). Prescriptive guidance produces gradeable behaviour; a tool name produces a keyword
      check.
      
      **And the scenario itself must be executable where it runs.** `verify`'s scenario asked to verify a
      FastAPI orders feature that exists in no checkout, so both arms correctly answered "there is nothing
      to verify" and several items were unsatisfiable by construction — the absolute measured the scenario,
      not the skill. Make a `capability` scenario self-contained: carry the spec, the diff and the reported
      facts inline rather than assuming repo state.
      
      ## README.md — run it honestly
      
      `evals/README.md` documents the two-axis run procedure and is candid about limits:
      
      - **Triggering eval:** load *only* this skill so routing is honest; for each prompt run 3–5 fresh-session trials; a `should_trigger` passes if the skill fires in the majority of trials, a `should_not_trigger` passes if it does NOT fire (and, where a `route_to` sibling exists, sanity-check that the prompt truly belongs there). State a pass bar (e.g. ≥90% accuracy across all trigger cases).
      - **Capability eval:** A/B — same prompt WITHOUT the skill (baseline) vs WITH it; grade each output against `must_include`; average over 3 trials; require the WITH condition to clear a bar (e.g. ≥80% of points) AND beat the baseline by a real margin.
      - **Honest caveats:** this is LLM-as-judge / human-in-the-loop, not deterministic. Use one consistent grader across A/B. Re-run after any edit to the body or the description, since both axes are wording-sensitive.
      
      A README that pretends the evals are deterministic CI is dishonest. Say what they really are.
      
    • rsc-conventions.md 5.5 KB
      # rsc conventions — wiring a skill into the catalog
      
      A skill is not done when `SKILL.md` reads well. It is done when it is *in the
      catalog* — frontmatter valid, indexed in the manifest, discoverable by the
      `npx @ericrisco/rsc` recommender, and reachable where the agent will find it. This reference
      is the rsc plumbing.
      
      ## The layout (single source of truth)
      
      ```text
      skills/<id>/                    ← canonical source of truth; edit here, only here
        SKILL.md
        references/*.md
        evals/cases.yaml
        evals/README.md
        scripts/verify.sh             ← only if the skill has a checkable artifact
      
      manifest.json                   ← GENERATED catalog (build-manifest.js); never hand-edit
      schema/frontmatter.schema.json  ← ajv schema every SKILL.md frontmatter must satisfy
      scripts/build-manifest.js       ← skills/*/SKILL.md → manifest.json (+ --check / --validate)
      scripts/eval-lint.sh            ← gates evals/cases.yaml minimums
      ```
      
      Rule: **`skills/<id>/` is the only place you edit.** There are no generated skill
      copies in the repo any more — the `rsc-universal` CLI copies skills into the
      target IDE at install time. `manifest.json` is generated; editing it by hand is a
      defect because the next `npm run manifest` overwrites it.
      
      ## Frontmatter — the catalog contract
      
      Every `SKILL.md` frontmatter MUST satisfy `schema/frontmatter.schema.json`:
      
      ```yaml
      ---
      name: my-skill                  # lowercase-kebab, matches the directory id
      description: Use when ...        # third-person, situation + NOT boundary, ~350 chars
      tags: [keyword, keyword]         # what the consult advisor searches over (≥1)
      recommends: [sibling-skill]      # what the system offers to install next (real ids)
      profiles: [core, full]           # optional: named-profile membership
      ---
      ```
      
      - `tags` drive the FTS recommender — pick the words a user would actually type.
      - `recommends` must reference **real skill ids**; `build-manifest.js --validate`
        fails on a dangling id.
      - `profiles` is optional. `minimal` = the floor (`orient`, `suggest`, `bro`, `harness`, `init`);
        `core` = the SDD workflow; `full` = everything. Most stack skills set `[full]`
        or omit it (they are installed on demand, not by profile).
      
      ## Invocation
      
      There are no bundles and no `/<bundle>:<id>` namespacing. Most targets install a
      skill under their own rsc folder (e.g. `.codex/rsc/<id>/`), reached from that
      assistant's instructions file. **Claude is the exception:** Claude Code only
      discovers project skills one level under `.claude/skills/`, so rsc skills install
      **flat** at `.claude/skills/<id>/SKILL.md` (a nested `.claude/skills/rsc/<id>/`
      is never discovered). A skill is invoked by its `name`. The `suggest` detector is
      always installed (the floor) and proposes installing any skill a task needs via
      `npx @ericrisco/rsc add <id>`.
      
      ## Wiring steps for a new skill
      
      1. **Create `skills/<id>/SKILL.md`** with valid frontmatter including `tags` and
         `recommends` (and `profiles` if it belongs to one).
      2. **Add reciprocal `recommends`** where it makes sense — if `nextjs` should
         suggest your new skill, add the id to `nextjs`'s `recommends`.
      3. **Regenerate the manifest:** `npm run manifest`.
      4. **Validate:** `npm run validate` (ajv frontmatter + recommends integrity) and
         `npm run manifest:check` (manifest is not stale, counts match).
      5. **Run the eval gate:** `bash scripts/eval-lint.sh` — must PASS for the new skill.
      6. **Add an outcome label** in `scripts/lib/recommend.js` if the skill is a
         user-facing outcome (so the plain-language wizard shows a human label, not the
         bare id). Internal workflow skills don't need one.
      7. **Update the README catalog** (and the harness `claude-md-template.md`
         Knowledge-map rows) if the skill introduces a new topic users should discover.
      
      ```bash
      # the gates, from repo root
      npm run validate        # frontmatter + recommends integrity
      npm run manifest:check  # manifest is current and count-accurate
      npm test                # unit + integration
      bash scripts/eval-lint.sh
      ```
      
      ## The Knowledge map
      
      The root `CLAUDE.md` carries a `## Knowledge map` section that indexes the 02-DOCS wiki topics — it is what every other skill reads before working in its area (the `harness` convention). When a skill produces durable artifacts, they live under `02-DOCS/wiki/<topic>/` and get a Knowledge-map row. For SDD-related artifacts the topic is `02-DOCS/wiki/sdd/`. `author-skill` writes there only when a design note is worth keeping; the executable record is always the skill's own `evals/`.
      
      ## verify.sh — only for checkable artifacts
      
      A `scripts/verify.sh` belongs in a skill that emits something a script can *check*: code (lint/type/test), config (schema), copy (a ban-list grep). It should be read-only by default and warn rather than fail unless asked to gate. **Process skills** — judged on the behavior/safety rails they install, like the SDD-phase skills and `author-skill` itself — have no artifact to grep, so they ship **no** `verify.sh`; their rigor is the `capability` eval. Adding a hollow `verify.sh` to a process skill is cargo-culting, not rigor.
      
      ## The originality rule (hard)
      
      The rsc catalog is its own ecosystem, not a re-skin. When mining ideas from other skill libraries, take the *idea* and re-express it in the rsc voice. Never reproduce another ecosystem's signature artifacts or phrasing: no borrowed urgency blocks ("1% chance…"), no copied rationalization-table wording, no `*-reviewer-prompt.md` file convention, no verbatim flowchart text. Git authorship on rsc commits is Eric, never the assistant. If a draft reads like it came from somewhere else, rewrite it.
      
  • SKILL.md 16.6 KB
    ---
    name: author-skill
    description: "Use when authoring a NEW rsc skill or editing an existing one — scoping it to one job, writing the description that decides whether it ever loads, splitting the body into references/, writing its evals, auditing it against the rubric. NOT building a product feature (that is `specify`) and NOT designing an agent loop (that is `building-agents`)."
    tags: [skill, authoring, meta]
    recommends: []
    origin: risco
    ---
    
    # author-skill — write skills that trigger and teach
    
    This is the **meta skill**: it authors and edits the other skills in the rsc catalog. A skill is two things bolted together — a **description that fires at the right moment**, and a **body that makes the agent better once it fires**. Most skills fail on the first. Treat them as two separate engineering problems with two separate quality bars.
    
    Where the SDD chain (`specify` → `plan` → … → `ship`) builds *product*, `author-skill` builds *the tools that build product*. Use it whenever a skill is born or edited.
    
    **Not this skill — delegate:** a product feature specced or planned → `../specify/SKILL.md`, `../plan/SKILL.md`. An autonomous agent or tool-calling loop → `../building-agents/SKILL.md`. Generic project docs or a wiki article → the `../harness/SKILL.md` 02-DOCS engine. Bootstrapping a workspace or profiling the user → `../init/SKILL.md`.
    
    Read `02-DOCS/wiki/harness/user-profile.md` and work at the accompaniment dial it records; `../init/SKILL.md` owns that dial and sets it. With no profile, default to non-technical framing and ask for the technical level and the dial before going deep — skill authoring is itself a technical act, so many users want more narration here than they do elsewhere.
    
    ## What a skill is (the anatomy)
    
    ```text
    skills/<id>/
    ├── SKILL.md              the body: frontmatter (name + description + origin) then the prose
    ├── references/           progressive-disclosure detail, loaded only when the body points to it
    │   └── <topic>.md
    ├── evals/
    │   ├── cases.yaml        trigger + capability test cases
    │   └── README.md         how to run the evals, honestly
    └── scripts/              optional; verify.sh + helpers for skills with a checkable artifact
        └── verify.sh
    ```
    
    The **frontmatter** decides *if* the skill loads. The **body** decides *how good* the agent is once it does.
    
    ## The description — the single highest-leverage line
    
    The description sits in context on **every turn the skill is installed**, invoked or not. The body is only paid for when the skill fires; the description is paid for always. So length here is a cost, never a credit. A vague description is a skill that never fires; an over-broad one hijacks unrelated turns. Get this right before anything else.
    
    Rules, all enforced:
    
    1. **Third person, present tense.** "Use when authoring a new skill…" — never "I help you…" or "You should…". The agent is reading *about* the skill.
    2. **Discriminative, not exhaustive.** Lead with a `Use when …` clause naming the *situation*, then only the capabilities that separate this skill from its neighbours. Do **not** append a `Triggers: '…', '…'` phrase list: the model matches on meaning, so a keyword bank in three languages buys nothing and is charged on every turn. The test is *discrimination, not coverage* — could a reader pick this skill over its nearest sibling from this line alone?
    3. **Draw the boundary.** End with a `NOT <x> (that is <sibling>)` clause, naming a sibling that actually exists under `skills/`. Negative space prevents hijacking as much as positive matching causes firing.
    4. **Aim ≤ 350 characters; 1024 is the schema-enforced hard limit.** One physical line, wrapped in double quotes, internal quotes escaped or avoided. If it does not parse, the skill does not load.
    5. **`origin: risco`** on its own line. This marks it as ours.
    
    ```yaml
    # Good — situation, the few capabilities that discriminate, a real boundary
    description: "Use when X happens or the user shows symptom Y — doing A, fixing B, choosing C. NOT Z (that is `sibling`)."
    
    # Bad — first person, no situation, no boundary; competes with every sibling on every turn
    description: "I help you write great skills and make them work."
    ```
    
    The full recipe, the budget tactics, and a worked before/after → `references/description-recipe.md`.
    
    ## Progressive disclosure — the body is an index, not an encyclopedia
    
    The body is loaded in full whenever the skill fires, so every line competes for the agent's attention. Write **the smallest body that still routes correctly**: 400 lines is a ceiling, not a target, and there is no floor — a skill that does its job in 60 lines beats the same skill padded to 200. Push anything long, reference-like, or rarely-needed into `references/<topic>.md` and link it inline at the point of use ("full table → `references/foo.md`").
    
    Decide where a paragraph lives:
    
    | Put it in the body when… | Move it to references/ when… |
    | --- | --- |
    | The agent needs it on *every* run | It is needed only in a specific branch |
    | It is a rule, a gate, or a decision point | It is a long table, a catalog, or a template |
    | It is short and load-bearing | It is reference detail that would bloat the body |
    | Cutting it would change behavior | It is an example that illustrates but does not instruct |
    
    Every file under `references/` must be linked from the body. An unlinked reference is never loaded, so it is dead weight in the package — link it or delete it.
    
    ## The hybrid structure — when each piece earns its place
    
    - **SKILL.md** — always. Frontmatter + focused body.
    - **references/** — only when the body genuinely needs offloaded depth. Do not create an empty `references/` to look complete; a single-file skill is fine.
    - **evals/** — always. `cases.yaml` + `README.md`. A skill with no evals is unverifiable and does not ship.
    - **scripts/verify.sh** — only when the skill produces a *checkable artifact* (code, config, copy with a ban-list). **Process skills** — those judged on the safety rails they install in the agent's behavior, like the SDD-phase skills or this one — do **not** ship a `verify.sh`; their evals carry a capability scenario instead.
    
    ## Orientation footer (required in every new skill)
    
    Every new skill MUST end with the orientation footer so the harness never leaves the user in seco. Append verbatim:
    
    ````markdown
    
    ## Orientación (siempre)
    
    Cierra cada turno con el **bloque-brújula** (📍 dónde estás · ✅ qué hiciste · 🧭 por qué · ➡️ siguiente, terminando en pregunta), calibrado al dial de `02-DOCS/wiki/harness/user-profile.md`. **Nunca termines en seco.** Protocolo completo: skill `orient` → `skills/orient/references/orientation-contract.md`. (Defiere a `suggest` el "¿instalo la skill que falta?".)
    ````
    
    The full protocol lives once in the `orient` skill; the footer only references it.
    
    ## The authoring workflow
    
    Run in order. Each step gates the next.
    
    1. **Name & scope.** One skill, one job. Pick a short kebab-case `<id>` that is the job, not the domain. If you can not say the job in one sentence, the scope is wrong — split it. Check no sibling already owns this; if one half-owns it, decide *edit the sibling* vs *new skill* before writing.
    2. **Draft the description.** Per the rules above. This first, because writing it forces the scope clear. → `references/description-recipe.md`.
    3. **Outline the body.** Method, rules, decision points. Mark what becomes a reference.
    4. **Write the body** in the rsc voice (see below). Tag every code/example fence with a language. Add a checklist or decision table *only where the flow actually branches* — not as decoration. Add a short anti-patterns table.
    5. **Extract references** for anything long or branch-specific, and link each one inline.
    6. **Write the evals** — `cases.yaml` then `README.md`. → `references/eval-authoring.md`.
    7. **Wire it into the rsc plumbing** (`tags`, `recommends`, `npm run manifest`, and indexing any artifact in `02-DOCS/wiki/index.md` — the Knowledge map; root `CLAUDE.md` keeps only a short pointer). → `references/rsc-conventions.md`.
    8. **Self-audit against the rubric** (below). Fix every miss or justify it.
    
    ## The rsc voice
    
    Match the catalog, do not invent a new register:
    
    - Direct, second-person-to-the-agent instruction ("Read the profile first", "Cut any section with no job").
    - A rule gets stated where it applies, with a one-line *why* that makes it obviously absolute — not a lecture, and not a shouted NON-NEGOTIABLE.
    - Concrete over abstract: a number, a path, a Bad→Good pair beats an adjective.
    - Original prose. Mine ideas from anywhere; the words are Eric's. Do **not** reproduce another ecosystem's signature artifacts or phrasing — no borrowed "1% chance" urgency blocks, no copied rationalization wording, no `*-reviewer-prompt.md` files, no verbatim flowcharts. The rsc identity is its own.
    - Cross-reference siblings by name or `../<sibling>/SKILL.md`, only ones that actually exist.
    
    ## Match the form to the failure
    
    Before you write an instruction, name the **failure** it's meant to prevent — then pick the form
    that actually fixes *that* failure. The instinct is to write a prohibition ("don't do X") for
    everything. That instinct is wrong for most failures, and measurably counter-productive for one
    class: **a prohibition aimed at output shape tends to summon the very thing it forbids** (the model
    attends to the named token), and can do *worse* than saying nothing at all. Match deliberately:
    
    | The failure is… | Use this form | Why, and example |
    | --- | --- | --- |
    | **Discipline** — the agent knows the rule but skips it under pressure (time, sunk cost, "just this once") | **The rule stated at the step it governs, with its why** — plus one row in the anti-patterns table | A rule in an appendix is skimmed; a rule in context is followed. "Do not ship a failing test — a red test merged is a lie in the suite", written at the ship step. Do not build a separate rationalization bank restating rules already in the flow; it is paid for on every load and read as decoration. |
    | **Wrong-shaped output** — tone, verbosity, format, structure come out wrong | **Positive recipe / contract** (show the target shape) | A prohibition ("don't be verbose", "no marketing fluff") makes it *more* likely — the model fixates on the banned shape. Give the shape to hit instead: "Reply in ≤3 sentences, lead with the verdict." Demonstrate, don't forbid. |
    | **Omitted element** — the agent forgets a required piece | **Required structural slot** (a checklist item or a template field it must fill) | You can't prohibit an absence. Make the slot mandatory so its emptiness is visible — a `Done-of-done` checkbox, a template section, a result-envelope field. |
    | **Conditional behavior** — right action depends on the situation | **Predicate-keyed conditional** ("When X → do Y; otherwise Z") | A flat rule fires in the wrong context. Key the behavior to its trigger so the agent branches correctly instead of over- or under-applying. |
    
    So: an anti-patterns table earns its place when it names concrete **failure modes** the flow above
    does not already state. A table that re-lists rules from the body is pure cost — delete it and move
    each rule to its step. And when you catch yourself writing "don't make it X" about the *shape* of an
    output, rewrite it as the shape to hit.
    
    ## The best-practice rubric (audit before shipping)
    
    A skill ships only when every box is checked or a miss is consciously justified.
    
    - [ ] **Frontmatter parses** as YAML; `name` matches the directory `<id>`; `origin: risco` present.
    - [ ] **Description** third-person, `Use when…` lead, an explicit `NOT … (that is sibling)` boundary naming a real sibling, ≤ 350 chars target / ≤ 1024 hard limit — judged on discrimination against the nearest sibling, not coverage.
    - [ ] **One job.** The body never drifts into a second skill's territory; it delegates instead.
    - [ ] **Body ≤ 400 lines** — a ceiling, not a target, with no floor. Long/branch-specific material lives in `references/`.
    - [ ] **Every `references/` file linked** inline from the body; none orphaned.
    - [ ] **Every fence language-tagged**; no placeholder/TODO prose; examples concrete.
    - [ ] **Checklist/decision table only where a flow branches**; an **anti-patterns table** present, naming failure modes rather than restating rules.
    - [ ] **Accompaniment dial honored** — reads the profile, adapts verbosity.
    - [ ] **Artifacts under `02-DOCS/wiki/`** and indexed in `02-DOCS/wiki/index.md` (the Knowledge map; root `CLAUDE.md` keeps only a short pointer), if the skill produces any.
    - [ ] **Concrete tooling delegated** to the stack skills rather than reinvented.
    - [ ] **evals present** — `cases.yaml` (≥5 `should_trigger` incl. non-obvious, ≥4 `should_not_trigger` each with a real-sibling `route_to`, ≥1 `capability` with a `must_include` rubric) + an honest `README.md`. `scripts/eval-lint.sh` passes — but it only checks presence and the counts (≥5/≥4/≥1) and that those keys are lists; the `route_to`-points-at-a-real-sibling, non-obvious phrasings, and `must_include` quality are yours to verify here, not the linter's.
    - [ ] **verify.sh** present iff the skill has a checkable artifact; process skills rely on evals.
    - [ ] **Every `must_include` item discriminates** — answerable by the scenario's task, plausibly caused by the skill and plausibly missed without it. An item *both* arms fail measures nothing and lowers the absolute; it has turned a real PASS into a FAIL here. → `references/eval-authoring.md`.
    - [ ] **Sibling links resolve** — every `../x/SKILL.md` points to a skill that exists.
    - [ ] **Wired** — `tags` + `recommends` set, `npm run manifest` re-run, and `npm run validate` / `npm run manifest:check` pass (manifest current, no dangling recommends).
    
    Full rubric rationale and the rsc plumbing steps → `references/rsc-conventions.md`.
    
    ### Two ship gates: document AND behavior
    
    The rubric above scores the skill as a **document**. That is one of two gates — a skill ships
    only when **both** are green:
    
    1. **Static** — the rubric above / `scripts/skill-rubric.md`, weighted score ≥ 8.5.
    2. **Behavioral** — `scripts/skill-behavior-rubric.md`: run the skill on its `capability`
       scenarios **with and without** it loaded, blind-grade both outputs, require `absolute ≥ 8.5`
       **and** `lift ≥ +1.0`. Run it:
    
       ```bash
       # 1) execute + grade — invoke the Workflow tool:
       #      scriptPath: scripts/skill-behavior-eval.workflow.js   args: "<skill-id>"
       #    save the returned object to /tmp/<skill>-raw.json
       # 2) score + gate (exit 0 pass / 1 fail):
       node scripts/skill-behavior-eval.js --score /tmp/<skill>-raw.json
       ```
    
       A failing `lift` means the body adds nothing a bare agent didn't already do — fix the body,
       don't game the checklist.
    
    ## Anti-patterns
    
    | Failure mode | Reality / fix |
    | --- | --- |
    | Description written last, once the body is done | The description is *why the body ever runs*, and drafting it first forces the scope clear. Write it first, to bar. |
    | Description padded for coverage — more phrasings, more languages, more verbs | It is in context on every turn, invoked or not. Prune until it discriminates against the nearest sibling and stops. |
    | One skill covering specify + plan + implement | Multi-job skills trigger fuzzily and teach poorly. One skill, one job — split it. |
    | Body grown past ~400 lines "because the topic is rich" | The agent skims what it cannot hold. Extract a reference and link it inline. |
    | A `references/` folder added to look thorough, or a reference nothing links to | An unlinked reference is never loaded — dead weight in the package. Link it at point of use or delete it. |
    | Evals skipped: "I'll just test it by hand once" | Unverifiable = does not ship. Write `cases.yaml`, near-misses with `route_to` included. |
    | `verify.sh` added to a process skill for rigor | A process skill has no artifact to grep. Its rigor is the capability eval. |
    | `../foo/SKILL.md` linked to something not in this repo | A dead link is a defect. Verify the directory exists under `skills/`. |
    | Another catalog mirrored wholesale ("it's basically superpowers' writing-skills") | Mine the idea, write it in the rsc voice. Copied artifacts or phrasing are a defect. |
    
    ## Project grounding (02-DOCS + CLAUDE.md)
    
    When authoring produces a durable design note (a skill's scope decision, a description rationale worth keeping), persist it under `02-DOCS/wiki/sdd/` and index it in `02-DOCS/wiki/index.md` (the Knowledge map; root `CLAUDE.md` keeps only a short pointer), per the `../harness/SKILL.md` convention — never a stray file at the repo root. The skill's own `evals/` is the executable record of intent; the wiki note is the human-readable why.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related