{"slug":"evals-analyze","title":"evals-analyze","summary":"Analyze evaluation results and close the loop. Specification failures create local CDRs to fix agent rules; generalization failures go to evaluator backlog.","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-06T17:21:36.520425Z","repo":{"url":"https://github.com/tikalk/adlc-team-skills","stars":137,"forks":1,"license":"MIT","updatedAt":"2026-09-22T21:09:55Z"},"bodyHtml":"<hr>\n<h2>name: evals-analyze\ndescription: Analyze evaluation results and close the loop. Specification failures create local CDRs to fix agent rules; generalization failures go to evaluator backlog.\ndisable-model-invocation: true</h2>\n<h1>evals-analyze</h1>\n<h2>What this skill does</h2>\n<p>Provides <strong>cross-functional team elevation</strong> and <strong>closed-loop feedback</strong> following <strong>EDD Principle VIII</strong> (Close the Production Loop) by deep-analyzing trajectory failure traces and routing them to correct resolution pathways.</p>\n<p><strong>Output</strong>:</p>\n<ol>\n<li><strong>Trajectory Analysis</strong> - Full multi-turn trace analysis with tool calls and context preservation (EDD Principle V)</li>\n<li><strong>Failure Routing</strong>:\n<ul>\n<li><strong>Specification Failures</strong> (agent logic missing/ambiguous) → Automatically triggers a local call to <code>levelup-specify</code> to propose new context rules in <code>.adlc/drafts/cdr/</code> to fix agent behavior.</li>\n<li><strong>Generalization Failures</strong> (grader flawed or lacks edge-case coverage) → Appends evaluator backlog items to the project backlog for ongoing monitoring.</li>\n</ul>\n</li>\n<li><strong>Cross-Functional PR</strong> - Creates a team-ai-directives PR with insights and rule updates (EDD Principle X)</li>\n</ol>\n<p><strong>Key EDD Principles Applied</strong>:</p>\n<ul>\n<li><strong>Principle VIII</strong>: Close Production Loop - Spec failures → fix directives; Gen failures → evaluator backlog</li>\n<li><strong>Principle V</strong>: Trajectory Observability - Full multi-turn traces, not just outputs</li>\n<li><strong>Principle X</strong>: Cross-Functional Observability - PMs, domain experts, and AI engineers collaborate</li>\n</ul>\n<h2>When to use</h2>\n<ul>\n<li><strong>After <code>/evals-validate</code></strong>: Analyze failures and resolve them</li>\n<li><strong>Closing a development loop</strong>: Translate evaluation failure insights into rule or evaluator fixes</li>\n<li><strong>Reporting to stakeholders</strong>: Generate readable summaries for PMs and domain experts</li>\n</ul>\n<h2>When NOT to use</h2>\n<ul>\n<li><strong>Evals not yet executed</strong>: Run <code>/evals-validate</code> first to generate results in <code>evals/results/</code></li>\n<li><strong>Trivial tasks</strong>: Closed-loop analysis is overhead for simple features</li>\n</ul>\n<h2>Process</h2>\n<h3>User Input</h3>\n<pre><code>$ARGUMENTS\n</code></pre>\n<ul>\n<li><code>--focus AREA</code> — Focus analysis on specific areas (e.g., security, quality, performance)</li>\n<li><code>--dry-run</code> — Analyze results and print report, but skip PR creation and local skill triggers</li>\n</ul>\n<h3>Execution Steps</h3>\n<h4>Phase 1: Load Evaluation Results</h4>\n<ul>\n<li>Reads results JSON from <code>evals/results/</code>.</li>\n<li>Extracts failure cases and full multi-turn conversation traces (including tool calls).</li>\n</ul>\n<h4>Phase 2: Failure Classification</h4>\n<p>Categorizes each failure trace:</p>\n<ul>\n<li><strong>Specification Failure</strong>: The agent was correct relative to its context, but the rule/directive was missing, ambiguous, or incorrect.</li>\n<li><strong>Generalization Failure</strong>: The rule was correct, but the agent made a mistake anyway (hallucinated, missed a constraint, or grader lacked edge-case coverage).</li>\n</ul>\n<h4>Phase 3: Action Routing (Close the Loop)</h4>\n<ul>\n<li><strong>For Specification Failures</strong>: Automatically triggers local skill <code>/levelup-specify</code> with the failure trace as input. This creates new rule/persona/example CDRs in <code>.adlc/drafts/cdr/</code> to fix the agent's behavior.</li>\n<li><strong>For Generalization Failures</strong>: Appends an evaluator backlog item to <code>evals/results/evaluator_backlog.md</code> detailing the needed grader edge-case updates.</li>\n</ul>\n<h4>Phase 4: Cross-Functional Insights &amp; PR</h4>\n<ul>\n<li>Generates a stakeholder-specific report in <code>evals/results/team_insights.md</code> (tailored for PMs, domain experts, and AI engineers).</li>\n<li>If git remote and gh CLI are available, commits rule/eval changes in <code>team-ai-directives</code> and opens a draft PR (uses <code>levelup-publish</code> logic under the hood).</li>\n</ul>\n<h2>Verification</h2>\n<ul>\n<li>Trajectory failure traces analyzed and classified</li>\n<li>Specification failures successfully routed to <code>/levelup-specify</code> (proposes CDRs in <code>.adlc/drafts/cdr/</code>)</li>\n<li>Generalization failures written to <code>evals/results/evaluator_backlog.md</code></li>\n<li>Stakeholder report <code>evals/results/team_insights.md</code> generated</li>\n<li>Draft PR created in team-ai-directives (if applicable)</li>\n<li>Final report summary presented with PR link and backlog details</li>\n</ul>\n","files":[{"path":"scripts/bash/setup-evals-analyze.sh","sizeBytes":1272,"isText":true},{"path":"scripts/powershell/setup-evals-analyze.ps1","sizeBytes":1359,"isText":false},{"path":"SKILL.md","sizeBytes":4791,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-23T13:51:10.391126Z","sha256":"8D4E7013EA09AC3247A499A0A93A83682A170781330D22D723A75B2395EF4825","sizeBytes":3652},"review":null,"source":{"repositoryUrl":"https://github.com/tikalk/adlc-team-skills","path":"skills/evals/evals-analyze","license":"MIT","commit":"3035db246f5f39088f26504d5a019b8397dafcf4","subtreeSha":"6CA5BC75C06B668B0053E9238BED66BAE6007D63895FD193F61CBF4E3CE6B55E","lastSyncedAt":"2026-09-23T13:50:41.913881Z"},"reviewedAt":"2026-09-23T13:59:50.612677Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/tikalk/adlc-team-skills/tree/main/skills/evals/evals-analyze"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install tikalk-adlc-team-skills@llmmart"},{"target":"git","command":"git clone https://github.com/tikalk/adlc-team-skills.git"}]}