{"slug":"clarify","title":"clarify","summary":"Use when a spec exists and must be de-risked before planning — hunt its ambiguities, unstated assumptions and edge cases, ask the few build-changing questions, bake the answers back into the spec. The rsc SDD gate between `specify` (writes the spec) and `plan` (designs the build)","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-10-02T16:37:46.010654Z","repo":{"url":"https://github.com/ericrisco/rsc-harness","stars":141,"forks":11,"license":"MIT","updatedAt":"2026-10-02T14:54:09Z"},"bodyHtml":"<h1>Eval harness — <code>clarify</code> skill</h1>\n<p>These evals check two things: that the skill <strong>triggers</strong> on the right prompts\n(and stays quiet on near-misses that belong to a neighbouring SDD phase), and\nthat it <strong>measurably improves</strong> the de-risking pass over a spec. Cases live in\n<code>cases.yaml</code>. There is no pure shell runner — grading is a judgment call done by\nan <strong>agent harness</strong> (a Claude Code agent with the skill catalog loaded) plus a\nhuman spot-check.</p>\n<h2>What's in <code>cases.yaml</code></h2>\n<ul>\n<li><code>should_trigger</code> — prompts that MUST load <code>clarify</code> (a spec exists; the user\nwants it de-risked / hole-poked / readiness-checked before planning).</li>\n<li><code>should_not_trigger</code> — near-misses routed to the genuinely correct sibling\nvia <code>route_to</code> (<code>specify</code>, <code>plan</code>, <code>analyze</code>, <code>debug</code>, <code>constitution</code>).</li>\n<li><code>capability</code> — a scenario with a <code>must_include</code> rubric to grade WITH vs\nWITHOUT the skill loaded.</li>\n</ul>\n<h2>Triggering eval</h2>\n<p>Goal: the skill fires when a spec needs de-risking, and never on an adjacent\nSDD phase.</p>\n<ol>\n<li>Configure an agent with the <strong>full catalog of skill descriptions</strong> available\nfor routing — the SDD chain siblings (<code>specify</code>, <code>plan</code>, <code>analyze</code>, <code>debug</code>,\n<code>constitution</code>, and the rest) plus the stack/process skills (<code>harness</code>,\n<code>init</code>, <code>fastapi</code>, <code>nextjs</code>, <code>go</code>, <code>postgresdb</code>, <code>flutter</code>, <code>design</code>,\n<code>marketing</code>, …) so routing competes realistically.</li>\n<li>For each <code>should_trigger</code> prompt: feed it cold and record whether <code>clarify</code>\nis the skill loaded. Run <strong>3–5 trials</strong> per prompt with fresh context.</li>\n<li>For each <code>should_not_trigger</code> prompt: confirm <code>clarify</code> does NOT load and\nthat the chosen skill matches <code>route_to</code>. Same 3–5 trials.</li>\n<li>Score: <code>triggered_correctly / total_trials</code> across both lists.</li>\n</ol>\n<p><strong>Pass bar: ≥ 90% trigger accuracy</strong> over all prompts and trials, with <strong>zero\nsystematic false-positives</strong> on the <code>specify</code> and <code>plan</code> near-misses — those are\nthe known traps. The boundary clarify must hold: no-spec-yet is <code>specify</code>,\nhow-to-build-it is <code>plan</code>. If it grabs either, the description is leaking.</p>\n<h2>Capability eval</h2>\n<p>Goal: prove the skill changes the de-risking pass, not just the routing.</p>\n<ol>\n<li>For the <code>capability</code> scenario, run it <strong>twice</strong>:\n<ul>\n<li><strong>WITHOUT</strong> the skill (base agent, no <code>clarify</code> loaded).</li>\n<li><strong>WITH</strong> the <code>clarify</code> skill loaded.</li>\n</ul>\n</li>\n<li>Grade each output against the <code>must_include</code> checklist — one point per\ncheckable item covered. A human or grading agent marks each present / absent.</li>\n<li>Compute coverage = <code>items_covered / total_items</code> per run.</li>\n</ol>\n<p><strong>Pass bar: WITH the skill covers ≥ 80% of <code>must_include</code>; WITHOUT clearly\nlower</strong> (target a ≥ 30-point gap). The discriminating behaviors are the ones a\nbase agent reliably misses: reading the constitution and citing what's already\nresolved, ranking gaps by leverage instead of dumping every question, framing\nquestions as decisions-with-a-recommendation, and — the highest-signal item —\nactually <strong>baking answers back into the spec with a dated Clarifications log</strong>\nrather than just listing questions in chat.</p>\n<h2>Notes on honesty</h2>\n<ul>\n<li>Trials are stochastic; report the raw fraction, not a rounded \"pass\".</li>\n<li>The single highest-signal capability check is <strong>the edit-back</strong>: an answer\nthat asks good questions but never modifies the spec or records the reasoning\nis a capability failure even if the questions are sharp. Clarify's deliverable\nis the sharpened spec, not the conversation.</li>\n<li>A second tell: a base agent often slides into proposing <em>how to build</em> the\nfeature. A correct clarify pass stays on WHAT and defers HOW to <code>plan</code>. Treat\narchitecture suggestions as a scope-leak failure.</li>\n<li>Re-run after any edit to <code>SKILL.md</code> — wording changes shift both triggering\n(especially the <code>specify</code>/<code>plan</code> boundary) and rubric coverage.</li>\n</ul>\n","files":[{"path":"evals/cases.yaml","sizeBytes":5997,"isText":true},{"path":"evals/README.md","sizeBytes":3767,"isText":true},{"path":"SKILL.md","sizeBytes":14763,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-10-02T16:38:33.76194Z","sha256":"6D36B53A967F609D8D2059326F5EFC08D923E3E92978804898FC21B1B37D34F2","sizeBytes":11371},"review":null,"source":{"repositoryUrl":"https://github.com/ericrisco/rsc-harness","path":"skills/clarify","license":"MIT","commit":"953fef5189c9991ddc7274a869d3c52aa73150fa","subtreeSha":"C769F1EC342923274B14345A1E95627070929E625760961A1CEF32445D34C084","lastSyncedAt":"2026-10-02T16:37:39.417112Z"},"reviewedAt":"2026-10-02T16:40:57.558355Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/ericrisco/rsc-harness/tree/main/skills/clarify"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ericrisco-rsc-harness@llmmart"},{"target":"git","command":"git clone https://github.com/ericrisco/rsc-harness.git"}]}