{"slug":"skill-eval-3","title":"skill-eval","summary":"Measure whether a skill helps a named task or needs revision or removal. Use when: a bounded routing or coding evaluation is requested; conformance alone cannot show benefit.","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-11T17:28:25.722199Z","repo":{"url":"https://github.com/boshu2/agentops","stars":445,"forks":41,"license":"Apache-2.0","updatedAt":"2026-09-24T01:09:16Z"},"bodyHtml":"<hr>\n<p>name: skill-eval\ndescription: 'Measure whether a skill helps a named task or needs revision or removal. Use when: a bounded routing or coding evaluation is requested; conformance alone cannot show benefit.'\npractices:</p>\n<ul>\n<li>measurement-over-assertion</li>\n<li>ab-testing\nskill_api_version: 1\nhexagonal_role: supporting\nconsumes:</li>\n<li>skill-source-package\nproduces:</li>\n<li>probe-package</li>\n<li>probe-result.v1\ncontext_rel:</li>\n<li>kind: supplier-to\nwith: skill-builder\nuser-invocable: true\nmetadata:\ntier: meta\ndependencies: []\ncapabilities: [\"author_seeded_probe\",\"run_probe_tier\",\"evaluate_skill_decision\"]\neffects: [\"write_probe_package\",\"dispatch_probe_producer\"]\ncanonical_status: canonical\ndisposition: keep_specialist\nstability: experimental</li>\n</ul>\n<hr>\n<h1>/skill-eval</h1>\n<p>Answer one named maintenance decision: <strong>retain, revise, remove, or insufficient\nevidence</strong>. Choose the measurement that can answer that decision, use the caller's\naccepted cases and resource envelope, make one scoped recommendation, and stop.\nA completed evaluation does not require a positive difference.</p>\n<p>This is an optional specialist. The repository's selected runner owns execution\nand bounds; native results own measurements; BD and Git retain their authority.\nDo not add a core skill, AO evaluation command, scheduler, dashboard, second\ntracker, or mandatory review merely to run an experiment.</p>\n<h2>Choose the question</h2>\n<table>\n<thead>\n<tr>\n<th>Caller decision</th>\n<th>Measurement</th>\n<th>What it can establish</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Does loading this skill change a specific observable act?</td>\n<td>Behavioral probe with <code>scripts/probe-skill.sh</code></td>\n<td>Behavior change on that scenario; not correct code or productivity</td>\n</tr>\n<tr>\n<td>Does this package or version improve engineering outcomes at acceptable cost?</td>\n<td>Repository-selected controlled coding comparison, such as <code>evals/skills-rpi</code></td>\n<td>Endpoint outcomes and cost on selected tasks; independent completion only when required exact-subject evidence exists</td>\n</tr>\n<tr>\n<td>Does a qualified memory update help later work?</td>\n<td>Separate frozen-versus-updated memory transfer test</td>\n<td>Narrow later-task reuse evidence with skill and runtime held fixed</td>\n</tr>\n<tr>\n<td>What happened in ordinary runs?</td>\n<td>Existing native accounting and acceptance evidence</td>\n<td>Observational failures, repairs and cost; not causal skill benefit</td>\n</tr>\n</tbody>\n</table>\n<p>Start from the caller's intended decision, not a mandatory quiz. Reuse an\nexisting accepted decision and scope. For a behavioral question, name one\nobservable action (a file written, tool used, criterion rejected); a belief such\nas “understands validation” needs translation into an action. For coding or\nmemory questions, name unchanged task acceptance and the maintenance choice.</p>\n<h2>Procedure</h2>\n<ol>\n<li><strong>Fix the decision and bounds.</strong> Name the subject package/version or qualified\nmemory update, relevant cases, allowed runtime and existing aggregate time,\ntrial and cost limits. Do not infer billing enforcement from token counters.\nSmoke runs, infrastructure retries, interrupted attempts and inner review\nconsume the same declared envelope. A new configuration or context does not\nrenew it. Do not launch live work without caller authorization and bounds.</li>\n<li><strong>Choose the smallest relevant measurement.</strong> Use behavioral probes for acts,\ncoding tasks for engineering outcomes, and separate later sessions for memory.\nThere is no universal two-effort requirement. Keep the deployed model and\neffort unless the caller's decision concerns effort. Retain easy regression\nand cost controls; do not weaken the producer to manufacture separation.</li>\n<li><strong>Freeze and calibrate.</strong> Fix task, acceptance, package, model/runtime,\nenvironment and grader identities before trials. Executable oracles must\naccept the intended solution and reject plausible incorrect/no-op solutions.\nInclude genuinely correct and incomplete cases when evaluating judgment.\nExposed incidents are development cases, never unseen holdouts by renaming.\nBroken or leaked cases invalidate affected comparisons; preserve their\nhistorical disposition when versioning a correction.</li>\n<li><strong>Run within the selected consumer's bounds.</strong> Equalize instructions, tools\nand environment across arms apart from the intended variable. Coding trials\nexpose the actual selected package and required resources. A worktree or a\nprompt prohibition is not runtime isolation. Exclude operator home, production\ntracker, session history, sibling output and solutions; capture launched\nconfiguration and final artifacts outside the worker. Report an incompatible\nadapter as such; do not build a replacement platform to rescue a result.</li>\n<li><strong>Read all attempts.</strong> Use native runner results and existing accounting;\ncollection must not require another model call or handwritten evaluation.\nKeep failed, abandoned, blocked, interrupted and missing attempts visible.\nWrong identity, changed acceptance, contamination or ambiguous pairing cannot\nestablish comparison proof even when a deterministic check passed.</li>\n<li><strong>Compare only supported facts.</strong> Pair by task and repetition; preserve\nrepetitions within task clusters. Report case outcomes, denominators,\nuncertainty and failure disposition. Endpoint reward, worker done claim,\nin-workflow validator PASS and independent acceptance are different facts.\nMissing review, usage, billing, phase or feasibility evidence stays unknown.</li>\n<li><strong>Recommend once and stop.</strong> State retain, revise, remove or insufficient\nevidence, the scope and supporting facts, and what remains unproven. A\nconcrete reproduced defect with clean controls can support a provisional\nnarrow repair; general improvement needs held-out comparison. Do not add\ntrials until green, require a positive result, or automatically publish a\nlesson. Do not remove losing observations or relax acceptance.</li>\n</ol>\n<h2>Coding and memory readout</h2>\n<p>Use the development adapter documented in\n<a href=\"../../evals/skills-rpi/readout.md\"><code>evals/skills-rpi/readout.md</code></a>, or the caller's\nexisting equivalent. Its report is a rebuildable view, not work authority.\nThe pilot's default <code>insufficient-evidence</code> recommendation is an honest limit;\nthe specialist may make a narrower supported maintenance recommendation and\nmust state its evidence and provisional scope.</p>\n<ul>\n<li>Report endpoint success against <strong>all assigned/observed attempts</strong> alongside\nany feasible-task rate. Retain infrastructure invalidity, infeasibility and\nunknown coverage separately; do not hide them by dropping the denominator.</li>\n<li>Report false completion, false acceptance and needless blocking separately\nwhen independent evidence measures them. Clean cases and abstentions are\ndenominators, not opportunities to reward finding-count spray. Unknown is not\nzero. Deterministic code truth may settle an experimental criterion, while a\nrequired native handoff or exact-subject judgment remains unproven.</li>\n<li>Report raw time/cost distributions and total cost of all attempts per accepted\noutcome. Zero accepted outcomes makes that ratio undefined. Partial Harbor\ncost is not total billing. Native input includes cached input; native output\nincludes reasoning. Keep counters distinct and never add native totals to\nHarbor totals or assume parents exclude children. Split producer, in-workflow\nvalidation, orchestration and grading only where native identity supports it.\nState the measurement window and excluded setup/analysis overhead.</li>\n<li>Use <code>evals/_stats</code> for paired task-cluster uncertainty after verifying its\ndependencies and semantics. A pilot is descriptive unless sample size and\ndecision thresholds were justified and fixed in advance. A zero-crossing\ninterval or <code>no_change</code> is <strong>not equivalence</strong>; equivalence needs its own margin\nand test. Same numeric repetitions/seeds do not prove controlled provider\nrandomness. Do not extrapolate local results across libraries or models.</li>\n<li>For memory, hold skill/runtime fixed and compare frozen with independently\nqualified updated memory in fresh later sessions, using an unseen transfer\ntask and an unrelated or invalidating control. Count acquisition, qualification,\nretrieval and downstream trial cost separately. Package available, content\ndelivered, relevant action and later outcome are separate facts. Saving a page\nearns no benefit credit; coding-pilot completion does not establish compounding.</li>\n</ul>\n<p>Raw trials and new proof belong in caller-selected protected external non-Git\nstorage. Only public/sanitized fixtures cleared for that destination belong in\nGit. Preserve legacy <code>.agents/</code> evidence. Existing independent support and\ndisclosure review precedes memory import; this skill does not auto-publish\ntranscripts or mutate knowledge from aggregate scores (ADR-0016).</p>\n<h2>Behavioral probes: preserve their existing meaning</h2>\n<p><code>scripts/probe-skill.sh</code> remains the runner for small behavioral regression\nprobes and immutable replay. It exposes an empty workspace and one injected\nSKILL.md, not a complete installed-package coding trial. Its verdict measures\n<strong>behavior change</strong>, never quality uplift or productive engineering completion.\nExisting ledger entries retain that meaning and their recorded limitations.</p>\n<table>\n<thead>\n<tr>\n<th>Probe form</th>\n<th>Use when</th>\n<th>Discriminator</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Tier 1 — quiz</td>\n<td>A decision rule is the caller's behavioral question</td>\n<td>The answer/action on the scenario</td>\n</tr>\n<tr>\n<td>Tier 2 — seeded task</td>\n<td>Applying a discipline in work is the question</td>\n<td>Whether the agent acted on a realistic planted defect</td>\n</tr>\n</tbody>\n</table>\n<p>Either form may be the starting point. Use\n<a href=\"references/seeding.md\"><code>references/seeding.md</code></a> for seeded tasks. Grade the act,\nnever vocabulary copied from the treatment. A floor probe detects at least one\nact; a multi-defect band needs both lower and upper bounds to catch omission and\nfinding spray. Calibrate against a transcript performing the act without the\nprelude's wording and one repeating the wording without the act.</p>\n<p>The declared <code>treatment_source</code> remains the only arm variable: <code>canonical-skill</code>\nuses exact SKILL.md bytes and is the mode the coverage gate counts;\n<code>injected-prelude</code> establishes prelude-only evidence. Live runs use the selected\nauthorized native producer with equal scenario and repetitions. Effort levels\nare a declared experimental choice, not a prerequisite for every question.</p>\n<pre><code>bash scripts/probe-skill.sh --probe &lt;id&gt; --replay\n# Only within an already authorized live envelope:\nbash scripts/probe-skill.sh --probe &lt;id&gt; --live --capture --reps 3 --output out.json\nbash scripts/check-skill-probe-headroom.sh\n</code></pre>\n<p>The existing <code>skill.probe-headroom</code> gate in <code>cli/internal/probeheadroom</code> owns\nclassification and thresholds. Its multi-effort saturation rule remains the\nlegacy gate contract; do not fabricate enough runs to satisfy it or rederive\nthe rule in a new report. Read and report the actual answer:</p>\n<ul>\n<li><strong>SATURATED:</strong> the probe cannot distinguish the targeted act. Preserve the\nobservation as a scenario limitation in the RUNBOOK; do not append a skill\nverdict to the legacy ledger. Do not infer skill value or lack of value.</li>\n<li><strong>FLOOR:</strong> treatment did not act. Check the discriminator on a known passing\ntranscript. The result alone does not prove the skill cannot help elsewhere.</li>\n<li><strong>UNMEASURED:</strong> no usable measurement, not INERT.</li>\n<li><strong>SEPARATED:</strong> the gate found usable headroom. This classification itself does\nnot establish positive treatment benefit; retain the actual probe verdict.</li>\n</ul>\n<p>Legacy behavioral ledger rows cite the headroom result, model, effort and\nsample size. Append one row only under that ledger's existing admissibility\nrules; preserve a valid INERT or losing result. Small samples remain\ndirectional. If producer failure or truncation makes a rep <code>infra</code>\n(discriminator exit 2), exclude it from the legacy <strong>usable behavioral rate</strong>\nand report its count in the all-attempt accounting. Zero usable treatment reps\nis UNMEASURED, never INERT. This rate convention does not authorize dropping\ninfrastructure attempts from coding-cohort accounting.</p>\n<h2>Output and completion</h2>\n<p>One scoped recommendation with the decision, cases, all attempts/coverage,\npaired outcomes when valid, uncertainty, cost/unknowns and failure disposition.\nFor behavioral authoring, also supply the existing probe package (<code>probe.json</code>,\n<code>question.md</code>, <code>discriminator.sh</code>, <code>fixtures/</code>, and a prelude only in\n<code>injected-prelude</code> mode) and its replay result. Use the legacy ledger/RUNBOOK\nonly for their existing consumers. No new per-run worksheet is required.</p>\n<p>Done when the requested measurement has reached its accepted stop, the relevant\nreplay/oracle checks discriminate, missing coverage is explicit, and one\nrecommendation answers the named maintenance decision. Insufficient evidence,\nan adverse result or an incompatible runtime can complete this evaluation;\nnone counts as demonstrated skill benefit.</p>\n<h2>References</h2>\n<ul>\n<li>Behavioral runner and conventions: <a href=\"../../scripts/probe-skill.sh\"><code>scripts/probe-skill.sh</code></a>, <a href=\"../../evals/skill-probes/README.md\"><code>evals/skill-probes/README.md</code></a>.</li>\n<li>Behavioral verdicts and non-verdict incidents: <a href=\"../../evals/skill-probes/LEDGER.md\"><code>LEDGER.md</code></a>, <a href=\"../../evals/skill-probes/RUNBOOK.md\"><code>RUNBOOK.md</code></a>.</li>\n<li>Existing coverage and headroom gates: <a href=\"../../scripts/check-skill-probe-coverage.sh\"><code>check-skill-probe-coverage.sh</code></a>, <a href=\"../../scripts/check-skill-probe-headroom.sh\"><code>check-skill-probe-headroom.sh</code></a>.</li>\n<li>Evidence and overclaim limits: ADR-0011, ADR-0016 and <a href=\"../../docs/architecture/rpi-traversal.md\"><code>RPI traversal</code></a>.</li>\n</ul>\n","files":[{"path":"references/seeding.md","sizeBytes":5861,"isText":true},{"path":"SKILL.md","sizeBytes":13925,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-16T15:57:01.252968Z","sha256":"21AB5CF89F614F35FFAC3D1BD5A6EF46B57BADAC8BDC25D867A45083C696B6E3","sizeBytes":8797},"review":null,"source":{"repositoryUrl":"https://github.com/boshu2/agentops","path":"images/gemini/skills/skill-eval","license":"Apache-2.0","commit":"c3fe161dce0b85d1e0490df757bbb841d22e4ea1","subtreeSha":"31828C7F4BC4033172DD1B8F4674CF9EB3EBDE272C9CA4FF9B31E45CA6CB6440","lastSyncedAt":"2026-09-24T06:48:55.360254Z"},"reviewedAt":"2026-09-16T16:04:54.226188Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/boshu2/agentops/tree/main/images/gemini/skills/skill-eval"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install boshu2-agentops@llmmart"},{"target":"git","command":"git clone https://github.com/boshu2/agentops.git"}]}