{"slug":"evaluation-methodology","title":"evaluation-methodology","summary":"PluginEval quality methodology — dimensions, rubrics, statistical methods, and scoring formulas. Use this skill when understanding how plugin quality is measured, when interpreting a low score on a specific dimension, when deciding how to improve a skill's triggering accuracy or ","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-01T18:59:55.818471Z","repo":{"url":"https://github.com/wshobson/agents","stars":40003,"forks":4267,"license":"MIT","updatedAt":"2026-09-26T19:54:17Z"},"bodyHtml":"<hr>\n<h2>name: evaluation-methodology\ndescription: \"PluginEval quality methodology — dimensions, rubrics, statistical methods, and scoring formulas. Use this skill when understanding how plugin quality is measured, when interpreting a low score on a specific dimension, when deciding how to improve a skill's triggering accuracy or orchestration fitness, when calibrating scoring thresholds for your marketplace, or when explaining quality badges to external partners like Neon.\"</h2>\n<h1>Evaluation Methodology</h1>\n<p>This document is the authoritative reference for how PluginEval measures plugin and skill quality.\nIt covers the three evaluation layers, all ten scoring dimensions, the composite formula, badge\nthresholds, anti-pattern flags, Elo ranking, and actionable improvement tips.</p>\n<p>Related: <a href=\"references/rubrics.md\">Full rubric anchors</a></p>\n<hr>\n<h2>The Three Evaluation Layers</h2>\n<p>PluginEval stacks three complementary layers. Each layer produces a score between 0.0 and 1.0 for\neach applicable dimension, and later layers override or blend with earlier ones according to\nper-dimension blend weights.</p>\n<h3>Layer 1 — Static Analysis</h3>\n<p><strong>Speed:</strong> &lt; 2 seconds. No LLM calls. Deterministic.</p>\n<p>The static analyzer (<code>layers/static.py</code>) runs six sub-checks directly against the parsed SKILL.md:</p>\n<table>\n<thead>\n<tr>\n<th>Sub-check</th>\n<th>What it measures</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>frontmatter_quality</code></td>\n<td>Name presence, description length, trigger-phrase quality</td>\n</tr>\n<tr>\n<td><code>orchestration_wiring</code></td>\n<td>Output/input documentation, code block count, orchestrator anti-pattern</td>\n</tr>\n<tr>\n<td><code>progressive_disclosure</code></td>\n<td>Line count vs. sweet-spot (200–600 lines), references/ and assets/ bonuses</td>\n</tr>\n<tr>\n<td><code>structural_completeness</code></td>\n<td>Heading density, code blocks, examples section, troubleshooting section</td>\n</tr>\n<tr>\n<td><code>token_efficiency</code></td>\n<td>MUST/NEVER/ALWAYS density, duplicate-line repetition ratio</td>\n</tr>\n<tr>\n<td><code>ecosystem_coherence</code></td>\n<td>Cross-references to other skills/agents, \"related\"/\"see also\" mentions</td>\n</tr>\n</tbody>\n</table>\n<p>These six sub-checks feed directly into six of the ten final dimensions (via <code>STATIC_TO_DIMENSION</code>\nmapping). The remaining four dimensions — <code>output_quality</code>, <code>scope_calibration</code>,\n<code>robustness</code>, and part of <code>triggering_accuracy</code> — receive no static contribution and rely\nentirely on Layer 2 and/or Layer 3.</p>\n<p><strong>Anti-pattern penalty</strong> is applied multiplicatively to the Layer 1 score:</p>\n<pre><code>penalty = max(0.5, 1.0 − 0.05 × anti_pattern_count)\n</code></pre>\n<p>Each additional detected anti-pattern reduces the score by 5%, flooring at 50%.</p>\n<h3>Layer 2 — LLM Judge</h3>\n<p><strong>Speed:</strong> 30–90 seconds. One or more LLM calls (Sonnet by default). Non-deterministic.</p>\n<p>The <code>eval-judge</code> agent reads the SKILL.md and any <code>references/</code> files, then scores four\ndimensions using anchored rubrics (see <a href=\"references/rubrics.md\">references/rubrics.md</a>):</p>\n<ol>\n<li><strong>Triggering accuracy</strong> — F1 score derived from 10 mental test prompts</li>\n<li><strong>Orchestration fitness</strong> — Worker purity assessment (0–1 rubric)</li>\n<li><strong>Output quality</strong> — Simulates 3 realistic tasks; assesses instruction quality</li>\n<li><strong>Scope calibration</strong> — Judges depth and breadth relative to the skill's category</li>\n</ol>\n<p>The judge returns a structured JSON object (no markdown fences) that the eval engine merges\ninto the composite. When <code>judges &gt; 1</code>, scores are averaged and Cohen's kappa is reported as\nan inter-judge agreement metric.</p>\n<h3>Layer 3 — Monte Carlo Simulation</h3>\n<p><strong>Speed:</strong> 5–20 minutes. N=50 simulated Agent SDK invocations (default). Statistical.</p>\n<p>Monte Carlo runs <code>N</code> real prompts through the skill and records:</p>\n<ul>\n<li><strong>Activation rate</strong> — Fraction of prompts that triggered the skill</li>\n<li><strong>Output consistency</strong> — Coefficient of variation (CV) across quality scores</li>\n<li><strong>Failure rate</strong> — Error/crash fraction with Clopper-Pearson exact CIs</li>\n<li><strong>Token efficiency</strong> — Median token count, IQR, outlier count</li>\n</ul>\n<p>The Layer 3 composite formula:</p>\n<pre><code>mc_score = 0.40 × activation_rate\n         + 0.30 × (1 − min(1.0, CV))\n         + 0.20 × (1 − failure_rate)\n         + 0.10 × efficiency_norm\n</code></pre>\n<p>where <code>efficiency_norm = max(0, 1 − median_tokens / 8000)</code>.</p>\n<hr>\n<h2>Composite Scoring Formula</h2>\n<p>The final score is a weighted blend across all three layers for each dimension, then summed:</p>\n<pre><code>composite = Σ(dimension_weight × blended_dimension_score) × 100 × anti_pattern_penalty\n</code></pre>\n<h3>Dimension Weights</h3>\n<table>\n<thead>\n<tr>\n<th>Dimension</th>\n<th>Weight</th>\n<th>Why it matters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>triggering_accuracy</code></td>\n<td>0.25</td>\n<td>A skill that never fires — or fires incorrectly — has no value</td>\n</tr>\n<tr>\n<td><code>orchestration_fitness</code></td>\n<td>0.20</td>\n<td>Skills must be pure workers; supervisor logic belongs in agents</td>\n</tr>\n<tr>\n<td><code>output_quality</code></td>\n<td>0.15</td>\n<td>Correct, complete output is the primary deliverable</td>\n</tr>\n<tr>\n<td><code>scope_calibration</code></td>\n<td>0.12</td>\n<td>Neither a stub nor a bloated monster</td>\n</tr>\n<tr>\n<td><code>progressive_disclosure</code></td>\n<td>0.10</td>\n<td>SKILL.md is lean; detail lives in references/</td>\n</tr>\n<tr>\n<td><code>token_efficiency</code></td>\n<td>0.06</td>\n<td>Minimal context waste per invocation</td>\n</tr>\n<tr>\n<td><code>robustness</code></td>\n<td>0.05</td>\n<td>Handles edge cases without crashing</td>\n</tr>\n<tr>\n<td><code>structural_completeness</code></td>\n<td>0.03</td>\n<td>Correct sections in the right order</td>\n</tr>\n<tr>\n<td><code>code_template_quality</code></td>\n<td>0.02</td>\n<td>Working, copy-paste-ready examples</td>\n</tr>\n<tr>\n<td><code>ecosystem_coherence</code></td>\n<td>0.02</td>\n<td>Cross-references; no duplication with siblings</td>\n</tr>\n</tbody>\n</table>\n<h3>Layer Blend Weights</h3>\n<p>Each dimension draws from different layers at different ratios. With all three layers active\n(<code>--depth deep</code> or <code>certify</code>):</p>\n<table>\n<thead>\n<tr>\n<th>Dimension</th>\n<th>Static</th>\n<th>Judge</th>\n<th>Monte Carlo</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>triggering_accuracy</code></td>\n<td>0.15</td>\n<td>0.25</td>\n<td>0.60</td>\n</tr>\n<tr>\n<td><code>orchestration_fitness</code></td>\n<td>0.10</td>\n<td>0.70</td>\n<td>0.20</td>\n</tr>\n<tr>\n<td><code>output_quality</code></td>\n<td>0.00</td>\n<td>0.40</td>\n<td>0.60</td>\n</tr>\n<tr>\n<td><code>scope_calibration</code></td>\n<td>0.30</td>\n<td>0.55</td>\n<td>0.15</td>\n</tr>\n<tr>\n<td><code>progressive_disclosure</code></td>\n<td>0.80</td>\n<td>0.20</td>\n<td>0.00</td>\n</tr>\n<tr>\n<td><code>token_efficiency</code></td>\n<td>0.40</td>\n<td>0.10</td>\n<td>0.50</td>\n</tr>\n<tr>\n<td><code>robustness</code></td>\n<td>0.00</td>\n<td>0.20</td>\n<td>0.80</td>\n</tr>\n<tr>\n<td><code>structural_completeness</code></td>\n<td>0.90</td>\n<td>0.10</td>\n<td>0.00</td>\n</tr>\n<tr>\n<td><code>code_template_quality</code></td>\n<td>0.30</td>\n<td>0.70</td>\n<td>0.00</td>\n</tr>\n<tr>\n<td><code>ecosystem_coherence</code></td>\n<td>0.85</td>\n<td>0.15</td>\n<td>0.00</td>\n</tr>\n</tbody>\n</table>\n<p>At <code>--depth standard</code> (static + judge only), blends are renormalized to drop the Monte Carlo\ncolumn. At <code>--depth quick</code> (static only), all weight falls on Layer 1.</p>\n<h3>Blended Score Calculation</h3>\n<p>For a given depth, the blended score for dimension <code>d</code> is:</p>\n<pre><code>blended[d] = Σ( layer_weight[d][layer] × layer_score[d][layer] )\n             ─────────────────────────────────────────────────────\n             Σ( layer_weight[d][layer] for available layers )\n</code></pre>\n<p>This normalization ensures that skipping Monte Carlo at standard depth doesn't artificially\ndeflate scores.</p>\n<hr>\n<h2>Interpreting Dimension Scores</h2>\n<p>Each dimension score is a float in <code>[0.0, 1.0]</code>. The CLI converts it to a letter grade:</p>\n<table>\n<thead>\n<tr>\n<th>Grade</th>\n<th>Score range</th>\n<th>Meaning</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A</td>\n<td>0.90 – 1.00</td>\n<td>Excellent — no meaningful improvement needed</td>\n</tr>\n<tr>\n<td>B</td>\n<td>0.80 – 0.89</td>\n<td>Good — minor gaps only</td>\n</tr>\n<tr>\n<td>C</td>\n<td>0.70 – 0.79</td>\n<td>Adequate — one or two clear improvement areas</td>\n</tr>\n<tr>\n<td>D</td>\n<td>0.60 – 0.69</td>\n<td>Marginal — needs targeted work</td>\n</tr>\n<tr>\n<td>F</td>\n<td>&lt; 0.60</td>\n<td>Failing — significant remediation required</td>\n</tr>\n</tbody>\n</table>\n<p>When reading a report, focus first on the lowest-graded dimension that has the highest weight.\nA D in <code>triggering_accuracy</code> (weight 0.25) costs far more than a D in <code>ecosystem_coherence</code>\n(weight 0.02).</p>\n<p><strong>Confidence intervals</strong> appear in the report when Layer 2 or Layer 3 ran. Narrow CIs (± &lt; 5\npoints) indicate stable scores. Wide CIs suggest inconsistency — often caused by an ambiguous\ndescription or instructions that work for some prompt styles but not others.</p>\n<hr>\n<h2>Quality Badges</h2>\n<p>Badges require both a composite score threshold AND an Elo threshold (when Elo is available).\nThe <code>Badge.from_scores()</code> logic checks composite first, then Elo if provided:</p>\n<table>\n<thead>\n<tr>\n<th>Badge</th>\n<th>Composite</th>\n<th>Elo</th>\n<th>Meaning</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Platinum ★★★★★</td>\n<td>≥ 90</td>\n<td>≥ 1600</td>\n<td>Reference quality — suitable for gold corpus</td>\n</tr>\n<tr>\n<td>Gold ★★★★</td>\n<td>≥ 80</td>\n<td>≥ 1500</td>\n<td>Production ready</td>\n</tr>\n<tr>\n<td>Silver ★★★</td>\n<td>≥ 70</td>\n<td>≥ 1400</td>\n<td>Functional, has improvement opportunities</td>\n</tr>\n<tr>\n<td>Bronze ★★</td>\n<td>≥ 60</td>\n<td>≥ 1300</td>\n<td>Minimum viable — not yet recommended for users</td>\n</tr>\n<tr>\n<td>—</td>\n<td>&lt; 60</td>\n<td>any</td>\n<td>Does not meet minimum bar</td>\n</tr>\n</tbody>\n</table>\n<p>The Elo threshold is skipped when Elo has not been computed (i.e., at quick or standard depth\nwithout <code>certify</code>). A skill can earn a badge on composite score alone in those cases.</p>\n<hr>\n<h2>Anti-Pattern Flags</h2>\n<p>The static analyzer detects five anti-patterns. Each carries a severity multiplier that feeds\ninto the penalty formula.</p>\n<h3>OVER_CONSTRAINED</h3>\n<p><strong>Trigger:</strong> More than 15 occurrences of MUST, ALWAYS, or NEVER in the SKILL.md.</p>\n<p><strong>Problem:</strong> Overly prescriptive instructions reduce model flexibility, increase token overhead,\nand signal that the author is trying to micromanage every output rather than providing\nprincipled guidance.</p>\n<p><strong>Fix:</strong> Audit every MUST/ALWAYS/NEVER. Replace directive language with explanatory framing\nwhere possible. Reserve hard constraints for genuine safety or correctness requirements. Target\nfewer than 10 such directives per 100 lines.</p>\n<h3>EMPTY_DESCRIPTION</h3>\n<p><strong>Trigger:</strong> The frontmatter <code>description</code> field is fewer than 20 characters after stripping.</p>\n<p><strong>Problem:</strong> Without a meaningful description, the Claude Code plugin system cannot determine\nwhen to invoke the skill. The skill becomes invisible to autonomous invocation.</p>\n<p><strong>Fix:</strong> Write a description of at least 60–120 characters that includes:</p>\n<ul>\n<li>A \"Use this skill when...\" or \"Use when...\" trigger clause</li>\n<li>Two or more concrete contexts separated by commas or \"or\"</li>\n</ul>\n<h3>MISSING_TRIGGER</h3>\n<p><strong>Trigger:</strong> The description does not contain \"use when\", \"use this skill when\",\n\"use proactively\", or \"trigger when\" (case-insensitive).</p>\n<p><strong>Problem:</strong> Even a long description is useless for autonomous invocation if it doesn't\ninclude a clear trigger signal. The system's routing model needs an explicit cue.</p>\n<p><strong>Fix:</strong> Prepend \"Use this skill when...\" to the description, followed by specific scenarios.\nExample: \"Use this skill when measuring plugin quality, interpreting score reports, or\nexplaining badge thresholds to a team.\"</p>\n<h3>BLOATED_SKILL</h3>\n<p><strong>Trigger:</strong> SKILL.md exceeds 800 lines AND the skill has no <code>references/</code> directory.</p>\n<p><strong>Problem:</strong> A monolithic SKILL.md forces the entire document into context on every invocation,\nwasting tokens on content only needed in edge cases.</p>\n<p><strong>Fix:</strong> Create a <code>references/</code> directory and move supporting material there:</p>\n<ul>\n<li>Detailed rubrics → <code>references/rubrics.md</code></li>\n<li>Extended examples → <code>references/examples.md</code></li>\n<li>Configuration reference → <code>references/config.md</code></li>\n</ul>\n<p>The SKILL.md should link to these files with <code>[text](references/filename.md)</code> so the model\ncan fetch them on demand.</p>\n<h3>ORPHAN_REFERENCE</h3>\n<p><strong>Trigger:</strong> SKILL.md contains a markdown link <code>[text](references/filename)</code> where\n<code>filename</code> does not exist in the <code>references/</code> directory.</p>\n<p><strong>Problem:</strong> Dead links waste tokens on context that will never resolve and confuse the model.</p>\n<p><strong>Fix:</strong> Either create the missing reference file or remove the dead link.</p>\n<h3>DEAD_CROSS_REF</h3>\n<p><strong>Trigger:</strong> SKILL.md references another skill or agent by relative path and that path\ncannot be resolved from the skills/ directory.</p>\n<p><strong>Problem:</strong> Broken ecosystem links undermine the plugin's coherence score and may cause\nthe model to attempt navigation to non-existent files.</p>\n<p><strong>Fix:</strong> Verify the referenced skill exists. Update the path or remove the reference.</p>\n<hr>\n<h2>Elo Ranking</h2>\n<p>PluginEval uses an Elo/Bradley-Terry rating system to rank a skill against the gold corpus.</p>\n<p><strong>Starting rating:</strong> 1500 (the corpus median by convention).</p>\n<p><strong>K-factor:</strong> 32 (standard for moderate-stakes ratings).</p>\n<p><strong>Expected score formula</strong> (standard Elo):</p>\n<pre><code>E(A vs B) = 1 / (1 + 10^((B_rating − A_rating) / 400))\n</code></pre>\n<p><strong>Rating update after each matchup:</strong></p>\n<pre><code>new_rating = old_rating + 32 × (actual_score − expected_score)\n</code></pre>\n<p>where <code>actual_score</code> is 1.0 for a win, 0.5 for a draw, 0.0 for a loss.</p>\n<p><strong>Confidence intervals</strong> are computed via 500-sample bootstrap, reported as 95% CI.\n<strong>Corpus percentile</strong> reflects pairwise win rate against the gold corpus.\n<strong>Position bias check:</strong> Pairs are evaluated in both orders; disagreements are flagged.</p>\n<p>The <code>plugin-eval init</code> command builds the corpus index from a plugins directory:</p>\n<pre><code>plugin-eval init ./plugins --corpus-dir ~/.plugineval/corpus\n</code></pre>\n<hr>\n<h2>CLI Reference</h2>\n<h3>Score a skill (quick static analysis only)</h3>\n<pre><code>plugin-eval score ./path/to/skill --depth quick\n</code></pre>\n<p>Returns Layer 1 results in &lt; 2 seconds. Useful for fast feedback during authoring.</p>\n<h3>Score with LLM judge (default)</h3>\n<pre><code>plugin-eval score ./path/to/skill\n</code></pre>\n<p>Runs static + LLM judge (standard depth). Takes 30–90 seconds.</p>\n<h3>Score with full output as JSON</h3>\n<pre><code>plugin-eval score ./path/to/skill --output json\n</code></pre>\n<p>Emits structured JSON including <code>composite.score</code>, <code>composite.dimensions</code>, and\n<code>layers[0].anti_patterns</code>. Suitable for CI integration:</p>\n<pre><code>plugin-eval score ./path/to/skill --depth quick --output json --threshold 70\n# exits with code 1 if score &lt; 70\n</code></pre>\n<h3>Full certification (all three layers + Elo)</h3>\n<pre><code>plugin-eval certify ./path/to/skill\n</code></pre>\n<p>Runs static + LLM judge + Monte Carlo (50 simulations) + Elo ranking. Takes 15–20 minutes.\nAssigns a quality badge. Use before publishing a skill to the marketplace.</p>\n<h3>Head-to-head comparison</h3>\n<pre><code>plugin-eval compare ./skill-a ./skill-b\n</code></pre>\n<p>Evaluates both skills at quick depth and prints a dimension-by-dimension comparison table.\nUseful for deciding between two implementations or measuring improvement before/after a\nrewrite.</p>\n<h3>Initialize corpus for Elo</h3>\n<pre><code>plugin-eval init ./plugins\n</code></pre>\n<p>Builds the local corpus index at <code>~/.plugineval/corpus</code>. Required before Elo ranking works.</p>\n<h3>Scripting the Composite Formula</h3>\n<p>Reproduce the composite score offline (pre-commit hook, CI gate):</p>\n<pre><code>def composite_score(dimension_scores: dict, anti_pattern_count: int = 0) -&gt; float:\n    \"\"\"Replicate the PluginEval composite formula.\"\"\"\n    WEIGHTS = {\n        \"triggering_accuracy\":    0.25,\n        \"orchestration_fitness\":  0.20,\n        \"output_quality\":         0.15,\n        \"scope_calibration\":      0.12,\n        \"progressive_disclosure\": 0.10,\n        \"token_efficiency\":       0.06,\n        \"robustness\":             0.05,\n        \"structural_completeness\":0.03,\n        \"code_template_quality\":  0.02,\n        \"ecosystem_coherence\":    0.02,\n    }\n    raw = sum(WEIGHTS[d] * s for d, s in dimension_scores.items())\n    penalty = max(0.5, 1.0 - 0.05 * anti_pattern_count)\n    return round(raw * 100 * penalty, 2)\n\n# Example: a skill with a weak triggering score\nscores = {\n    \"triggering_accuracy\":    0.65,  # D — needs description work\n    \"orchestration_fitness\":  0.85,\n    \"output_quality\":         0.80,\n    # … fill in remaining 7 dimensions …\n}\n# composite_score(scores, anti_pattern_count=1) → ~76.5\n</code></pre>\n<h3>JSON Output Format</h3>\n<p>Top-level shape of <code>--output json</code>:</p>\n<pre><code>{\n  \"composite\": { \"score\": 76.5, \"badge\": \"Silver\", \"elo\": null },\n  \"dimensions\": {\n    \"triggering_accuracy\": { \"score\": 0.65, \"grade\": \"D\", \"ci_low\": 0.60, \"ci_high\": 0.70 },\n    \"orchestration_fitness\": { \"score\": 0.85, \"grade\": \"B\", \"ci_low\": 0.80, \"ci_high\": 0.90 }\n  },\n  \"layers\": [\n    { \"name\": \"static\", \"duration_ms\": 1243, \"anti_patterns\": [\"OVER_CONSTRAINED\"] },\n    { \"name\": \"judge\", \"duration_ms\": 48200, \"judges\": 1, \"kappa\": null }\n  ]\n}\n</code></pre>\n<p>Parse <code>composite.score</code> in CI to gate deployments:</p>\n<pre><code>score=$(plugin-eval score ./my-skill --output json | python3 -c \"import sys,json; print(json.load(sys.stdin)['composite']['score'])\")\nif (( $(echo \"$score &lt; 70\" | bc -l) )); then\n  echo \"Quality gate failed: score $score &lt; 70\"\n  exit 1\nfi\n</code></pre>\n<hr>\n<h2>Tips for Improving a Skill's Score</h2>\n<p>Work through dimensions in weight order. The largest gains come from fixing the top-weighted\ndimensions first.</p>\n<h3>Which Dimension to Improve First</h3>\n<p>Use this table when a score report shows multiple D/F grades and you need to prioritize effort.</p>\n<table>\n<thead>\n<tr>\n<th>Dimension</th>\n<th>Weight</th>\n<th>Typical fix effort</th>\n<th>Score impact / hour</th>\n<th>Fix first if…</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>triggering_accuracy</code></td>\n<td>0.25</td>\n<td>Low — description rewrite</td>\n<td>High</td>\n<td>Score &lt; 70 overall</td>\n</tr>\n<tr>\n<td><code>orchestration_fitness</code></td>\n<td>0.20</td>\n<td>Medium — restructure sections</td>\n<td>High</td>\n<td>Skill mixes worker + supervisor logic</td>\n</tr>\n<tr>\n<td><code>output_quality</code></td>\n<td>0.15</td>\n<td>Medium — add examples</td>\n<td>Medium</td>\n<td>Judge score &lt; 0.70</td>\n</tr>\n<tr>\n<td><code>scope_calibration</code></td>\n<td>0.12</td>\n<td>Low — move content to references/</td>\n<td>Medium</td>\n<td>File is &lt; 100 or &gt; 800 lines</td>\n</tr>\n<tr>\n<td><code>progressive_disclosure</code></td>\n<td>0.10</td>\n<td>Low — create references/ dir</td>\n<td>Medium</td>\n<td>No references/ directory exists</td>\n</tr>\n<tr>\n<td><code>token_efficiency</code></td>\n<td>0.06</td>\n<td>Low — reduce MUST/ALWAYS/NEVER</td>\n<td>Low</td>\n<td>Anti-pattern count ≥ 3</td>\n</tr>\n<tr>\n<td><code>robustness</code></td>\n<td>0.05</td>\n<td>Low — add Troubleshooting section</td>\n<td>Low</td>\n<td>No edge-case handling documented</td>\n</tr>\n<tr>\n<td><code>structural_completeness</code></td>\n<td>0.03</td>\n<td>Very low — add headings/code blocks</td>\n<td>Low</td>\n<td>Fewer than 4 H2 headings</td>\n</tr>\n<tr>\n<td><code>code_template_quality</code></td>\n<td>0.02</td>\n<td>Very low — add language tags</td>\n<td>Very low</td>\n<td>Code blocks missing language tags</td>\n</tr>\n<tr>\n<td><code>ecosystem_coherence</code></td>\n<td>0.02</td>\n<td>Very low — add Related section</td>\n<td>Very low</td>\n<td>No cross-references at all</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Rule of thumb:</strong> Fix <code>triggering_accuracy</code> before anything else — at weight 0.25 it delivers\nmore composite-score gain per hour than all low-weight dimensions combined.</p>\n<h3>Triggering Accuracy (weight 0.25)</h3>\n<ul>\n<li>Include \"Use this skill when...\" followed by 3–4 comma-separated specific contexts.</li>\n<li>Add \"proactively\" if the skill should auto-activate without an explicit user request.</li>\n<li>Mental test: write 5 prompts that should trigger it and 5 that should not — does\nyour description discriminate? If not, add or tighten the context phrases.</li>\n</ul>\n<h3>Orchestration Fitness (weight 0.20)</h3>\n<ul>\n<li>Document what the skill <em>receives</em> and what it <em>returns</em> — not what it orchestrates.</li>\n<li>Avoid \"orchestrate\", \"coordinate\", \"dispatch\", \"manage workflow\" in SKILL.md.</li>\n<li>Include an \"Output format\" section and 2+ code blocks showing concrete worker behavior.</li>\n</ul>\n<h3>Output Quality (weight 0.15)</h3>\n<ul>\n<li>Give specific, actionable instructions — not just goals.</li>\n<li>Cover at least one edge case explicitly (empty input, malformed data, etc.).</li>\n<li>Include an examples section showing representative inputs and expected outputs.</li>\n<li>The more concrete the instructions, the higher the judge will score this dimension.</li>\n</ul>\n<h3>Scope Calibration (weight 0.12)</h3>\n<ul>\n<li>Target 200–600 lines. Below 100 is a stub; above 800 without <code>references/</code> is bloat.</li>\n<li>Move background reading, extended examples, and reference tables to <code>references/</code>.</li>\n<li>Very narrow skills should be merged with a sibling; very broad ones should be split.</li>\n</ul>\n<h3>Progressive Disclosure (weight 0.10)</h3>\n<ul>\n<li>Add a <code>references/</code> directory (earns 0.15–0.25 bonus) and keep SKILL.md focused on\nthe execution path. An <code>assets/</code> directory adds a further bonus.</li>\n</ul>\n<h3>Token Efficiency (weight 0.06)</h3>\n<ul>\n<li>Audit MUST/ALWAYS/NEVER count. Target &lt; 1 per 10 lines.</li>\n<li>Consolidate near-duplicate bullet points and repeated-structure tables.</li>\n</ul>\n<h3>Robustness (weight 0.05)</h3>\n<ul>\n<li>Add a \"Troubleshooting\" or \"Edge Cases\" section covering at least 3 failure modes.</li>\n<li>State what the skill returns when it cannot complete its task.</li>\n</ul>\n<h3>Structural Completeness (weight 0.03)</h3>\n<ul>\n<li>Ensure at least 4 H2/H3 headings, 3 code blocks, an Examples section, and a Troubleshooting section.</li>\n</ul>\n<h3>Code Template Quality (weight 0.02)</h3>\n<ul>\n<li>All code blocks must be syntactically valid and copy-paste ready with language tags.</li>\n</ul>\n<h3>Ecosystem Coherence (weight 0.02)</h3>\n<ul>\n<li>Add a \"## Related\" section listing sibling skills or agents with relative paths.</li>\n<li>Avoid duplicating content that already exists in another skill — link to it instead.</li>\n</ul>\n<hr>\n<h2>Troubleshooting</h2>\n<h3>\"Score is much lower than expected after adding content\"</h3>\n<p>The anti-pattern penalty compounds. Run with <code>--output json</code> and inspect\n<code>layers[0].anti_patterns</code>. If you have 5+ anti-patterns, the multiplier can reduce your\nscore to 75% of its raw value regardless of how good the content is. Fix the flags first.</p>\n<h3>\"triggering_accuracy is low despite a detailed description\"</h3>\n<p>The <code>_description_pushiness</code> scorer looks for specific syntactic patterns, not just length.\nVerify your description contains the phrase \"Use this skill when\" or \"Use when\" (exact\nphrasing matters — it's a regex match). Also check that you have multiple use cases separated\nby commas or \"or\" to earn the specificity bonus.</p>\n<h3>\"LLM judge scores vary significantly between runs\"</h3>\n<p>This is expected for ambiguous skills. The judge generates 10 mental test prompts\nnon-deterministically. Improve score stability by tightening the description and adding\nconcrete examples. When <code>judges &gt; 1</code>, averaged scores will be more stable. Use\n<code>--depth deep</code> with <code>certify</code> which runs Monte Carlo to get statistically-bounded scores.</p>\n<h3>\"progressive_disclosure score is low even though the file is the right length\"</h3>\n<p>Check whether the file is in the 200–600 line sweet spot. Files shorter than 100 lines\nscore only 0.20 on this sub-check. Also confirm that <code>references/</code> files are not empty —\nthe scorer checks for non-empty reference files, not just the directory.</p>\n<h3>\"compare shows my rewrite scores lower than the original\"</h3>\n<p>Quick depth (<code>--depth quick</code>) only runs static analysis. If the rewrite moved content to\n<code>references/</code> and shortened SKILL.md significantly, static scores for structural completeness\nmay drop even though overall quality improved. Run <code>--depth standard</code> for a fairer comparison\nthat includes the LLM judge's assessment of content quality.</p>\n<hr>\n<h2>References</h2>\n<ul>\n<li><a href=\"references/rubrics.md\">Full Rubric Anchors — all 4 judge dimensions</a></li>\n</ul>\n<h3>Related Agents</h3>\n<ul>\n<li><strong>eval-judge</strong> (<code>../../agents/eval-judge.md</code>) — the LLM judge that scores Layer 2 dimensions\n(<code>triggering_accuracy</code>, <code>orchestration_fitness</code>, <code>output_quality</code>, <code>scope_calibration</code>).\nInvoke directly when you need to re-run only the judge layer or inspect its reasoning.</li>\n<li><strong>eval-orchestrator</strong> (<code>../../agents/eval-orchestrator.md</code>) — the top-level orchestrator that\nsequences all three layers, merges results, assigns badges, and writes the final report.\nInvoke when running a full certification pass or comparing two skills head-to-head.</li>\n</ul>\n","files":[{"path":"references/improving-scores.md","sizeBytes":5602,"isText":true},{"path":"references/rubrics.md","sizeBytes":22738,"isText":true},{"path":"SKILL.md","sizeBytes":7908,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-26T23:12:27.147006Z","sha256":"165ED0DB6FFABC318BB1DEC03AC8007F7E10CE0EACA4FDB52D362A86452C63D1","sizeBytes":14311},"review":null,"source":{"repositoryUrl":"https://github.com/wshobson/agents","path":"plugins/plugin-eval/skills/evaluation-methodology","license":"MIT","commit":"9b15b34b0bfc13a815cbfc2366e14ea549e09422","subtreeSha":"0D3FCCD271BFE8FE60F2C50E53A07EA79E50643163728C74D704152939E25756","lastSyncedAt":"2026-09-26T23:12:03.520842Z"},"reviewedAt":"2026-09-26T23:13:07.653287Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/wshobson/agents/tree/main/plugins/plugin-eval/skills/evaluation-methodology"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install wshobson-agents@llmmart"},{"target":"git","command":"git clone https://github.com/wshobson/agents.git"}]}