{"slug":"analyze-3","title":"analyze","summary":"Use when constitution, spec, plan and tasks all exist and you want them cross-read against each other before any code is written — the rsc-sdd pre-implementation gate. Reports coverage gaps, contradictions, duplication, ambiguity and scope drift; edits nothing. NOT the task break","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-10-02T16:37:40.840855Z","repo":{"url":"https://github.com/ericrisco/rsc-harness","stars":141,"forks":11,"license":"MIT","updatedAt":"2026-10-02T14:54:09Z"},"bodyHtml":"<h1>Eval harness — <code>analyze</code> skill</h1>\n<p><code>analyze</code> is the rsc-sdd <strong>pre-implementation consistency gate</strong>. These evals\ncheck two things: that the skill <strong>triggers</strong> on the right prompts (a\ncross-check of the planning artifacts before coding) and stays quiet on\nnear-misses (anything that resolves, builds, runs, or debugs code), and that it\n<strong>measurably changes the answer</strong> — reporting and routing instead of editing\nartifacts or jumping into code. Cases live in <code>cases.yaml</code>. There is no shell\nrunner: triggering and report quality are judgment calls graded by an agent\nharness plus a human spot-check.</p>\n<h2>What's in <code>cases.yaml</code></h2>\n<ul>\n<li><code>should_trigger</code> — prompts that MUST load <code>analyze</code> (several avoid the word\n\"analyze\" to test intent, not keyword).</li>\n<li><code>should_not_trigger</code> — near-misses routed to the correct existing sibling via\n<code>route_to</code> (debugging, verification via a stack skill, plan authoring, code\nreview, harness bootstrap, db perf).</li>\n<li><code>capability</code> — a full BLOCKED-gate scenario with a <code>must_include</code> rubric to\ngrade with vs without the skill.</li>\n</ul>\n<h2>Triggering eval</h2>\n<p>Goal: the gate fires when the four artifacts exist and a cross-check is wanted,\nand never on a near-miss.</p>\n<ol>\n<li>Configure an agent with the <strong>full catalog of skill descriptions</strong> available\nfor routing (analyze + harness, init, the stack skills fastapi/nextjs/go/\nflutter/postgresdb, secure-coding, deployment, design, marketing,\npresentations, course-storytelling, building-agents) so routing competes\nrealistically. When the other rsc-sdd phase skills (clarify, plan, tasks,\nimplement, verify, review, debug) are added to the repo, include them too —\nthey are the closest competitors and the sharpest test of the boundary.</li>\n<li>For each <code>should_trigger</code> prompt: feed it cold, record whether <code>analyze</code> is\nthe skill the agent loads. Run <strong>3-5 trials</strong> per prompt (fresh context).</li>\n<li>For each <code>should_not_trigger</code> prompt: confirm <code>analyze</code> does NOT load and the\nchosen skill matches <code>route_to</code>. Same 3-5 trials.</li>\n<li>Score: <code>triggered_correctly / total_trials</code> across both lists.</li>\n</ol>\n<p><strong>Pass bar: &gt;= 90% trigger accuracy</strong> over all prompts and trials, with <strong>zero\nsystematic false-positives</strong> on the debugging and verification near-misses\n(those are the known traps — \"is it ready to ship / why is it failing\" is a\npost-implementation or runtime concern, not a pre-implementation artifact gate).</p>\n<h2>Capability eval</h2>\n<p>Goal: prove the skill changes the answer, not just the routing.</p>\n<ol>\n<li>For the <code>capability</code> scenario, run it <strong>twice</strong>:\n<ul>\n<li><strong>WITHOUT</strong> the skill (base agent, no <code>analyze</code> loaded).</li>\n<li><strong>WITH</strong> the <code>analyze</code> skill loaded.</li>\n</ul>\n</li>\n<li>Grade each output against the <code>must_include</code> checklist — one point per item\ncovered. A human or grading agent marks each item present / absent.</li>\n<li>Compute coverage = <code>items_covered / total_items</code> for each run.</li>\n</ol>\n<p><strong>Pass bar: WITH the skill covers &gt;= 80% of <code>must_include</code>; WITHOUT clearly\nlower</strong> (target a &gt;= 30-point gap). The discriminating behaviors are the\nreport-only discipline (no artifact edits, no code), running all six analyses,\nthe constitution-violation + GAP double-flag on the missing rate limit, the\nDRIFT call on the Redis cache, the coverage map, the verdict line, and the\nhand-off to the next phase.</p>\n<h2>Notes on honesty</h2>\n<ul>\n<li>Trials are stochastic; report the raw fraction, not a rounded \"pass\".</li>\n<li>The highest-signal capability check is <strong>restraint</strong>: a correct answer reports\nand routes, it does not \"helpfully\" rewrite the spec or start coding. Treat a\nconfident answer that edits an artifact or begins implementation as a\n<strong>capability failure</strong> even if its analysis is otherwise sharp — that behavior\ndefeats the gate.</li>\n<li>Re-run after any edit to <code>SKILL.md</code>. Wording changes shift both triggering and\nrubric coverage, and the report-only boundary is easy to soften by accident.</li>\n</ul>\n","files":[{"path":"evals/cases.yaml","sizeBytes":6510,"isText":true},{"path":"evals/README.md","sizeBytes":3857,"isText":true},{"path":"SKILL.md","sizeBytes":11156,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-10-02T16:38:02.178872Z","sha256":"E6FF0D2AFAA2F91CECB6A9220FE80265B5EE1872DA2AACFE2E4B3888AB7A304F","sizeBytes":9772},"review":null,"source":{"repositoryUrl":"https://github.com/ericrisco/rsc-harness","path":"skills/analyze","license":"MIT","commit":"953fef5189c9991ddc7274a869d3c52aa73150fa","subtreeSha":"4428FB08DDB15710E3D1C56956430490C7D0169355CDEBE0E0C0D807ACA30289","lastSyncedAt":"2026-10-02T16:37:39.417112Z"},"reviewedAt":"2026-10-02T16:39:17.876619Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/ericrisco/rsc-harness/tree/main/skills/analyze"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ericrisco-rsc-harness@llmmart"},{"target":"git","command":"git clone https://github.com/ericrisco/rsc-harness.git"}]}