Claude Skill

skills-eval

Evaluate Claude skill quality through auditing. Use when reviewing or auditing skills.

LLM Mart · 0 points · 6 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download athola-claude-night-market-plugins_abstract_skills_skills-eval-9045831.zip · 40 KB
Part of athola/claude-night-market — 46 skills

Install

skills CLI npx skills add https://github.com/athola/claude-night-market/tree/master/plugins/abstract/skills/skills-eval
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install athola-claude-night-market@llmmart
Git git clone https://github.com/athola/claude-night-market.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole athola/claude-night-market collection as a plugin from our marketplace. Git is the plain clone.

README

Skills Evaluation

detailed skill analysis and quality improvement tools.

Usage

# Discover all skills
skills-auditor --discover

# Quality assessment
skills-auditor --quality-scan

# Generate improvement suggestions
improvement-suggester --skill <skill_name>

# Compliance check
compliance-checker --validate-all

Tools

  • skills-auditor: Skill discovery and inventory
  • improvement-suggester: Automated improvement recommendations
  • compliance-checker: Quality standards validation
  • tool-performance-analyzer: Performance metrics

Key Metrics

  • Structure compliance
  • Token efficiency
  • Documentation quality
  • Tool integration
  • Activation reliability

Quick Validation

make check  # Run detailed validation

Skill manifest

Skills Evaluation and Improvement

When NOT To Use

  • Writing a new skill (use abstract:skill-authoring)
  • Evaluating hooks (use abstract:hooks-eval)
  • Evaluating rules (use abstract:rules-eval)

Overview

This framework audits Claude skills against quality standards to improve performance and reduce token consumption. Automated tools analyze skill structure, measure context usage, and identify specific technical improvements. Run verification commands after each audit to confirm fixes work correctly.

The skills-auditor provides structural analysis, while the improvement-suggester ranks fixes by impact. Compliance is verified through the compliance-checker. Runtime efficiency is monitored by tool-performance-analyzer and token-usage-tracker.

Quick Start

Basic Audit

Run a full audit of all skills or target a specific file to identify structural issues.

# Audit all skills
make audit-all

# Audit specific skill
make audit-skill TARGET=path/to/skill/SKILL.md

Analysis and Optimization

Use skill_analyzer.py for complexity checks and token_estimator.py to verify the context budget.

make analyze-skill TARGET=path/to/skill/SKILL.md
make estimate-tokens TARGET=path/to/skill/SKILL.md

Improvements

Generate a prioritized plan and verify standards compliance using improvement_suggester.py and compliance_checker.py.

make improve-skill TARGET=path/to/skill/SKILL.md
make check-compliance TARGET=path/to/skill/SKILL.md

Evaluation Workflow

Start with make audit-all to inventory skills and identify high-priority targets. For each skill requiring attention, run analysis with analyze-skill to map complexity. Generate an improvement plan, apply fixes, and run check-compliance to verify the skill meets project standards. Finalize by checking the token budget for efficiency.

Evaluation and Optimization

Quality assessments use the skills-auditor and improvement-suggester to generate detailed reports. Performance analysis focuses on token efficiency through the token-usage-tracker and tool performance via tool-performance-analyzer. For standards compliance, the compliance-checker automates common fixes for structural issues.

Scoring and Prioritization

We evaluate skills across five dimensions: structure compliance, content quality, token efficiency, activation reliability, and tool integration. Scores above 90 represent production-ready skills, while scores below 50 indicate critical issues requiring immediate attention.

Improvements are prioritized by impact. Critical issues include security vulnerabilities or broken functionality. High-priority items cover structural flaws that hinder discoverability. Medium and low priorities focus on best practices and minor optimizations.

Structural Patterns

Deprecated: skills/shared/modules/ directories. Shared modules must be relocated into the consuming skill's own modules/ directory. The evaluator flags any remaining skills/shared/ as a structural warning.

Current: Each skill owns its modules at skills/<skill-name>/modules/. Cross-skill references use relative paths (e.g., ../skill-authoring/modules/description-writing.md).

Resources

Shared Modules: Cross-Skill Patterns

Skill-Specific Modules

  • Trigger Isolation Analysis: See modules/trigger-isolation-analysis.md
  • Authoring Checklist: See modules/authoring-checklist.md
  • Evaluation Workflows: See modules/evaluation-workflows.md
  • Advanced Tool Use Analysis: See modules/advanced-tool-use-analysis.md
  • Evaluation Framework: See modules/evaluation-framework.md
  • Integration Patterns: See modules/integration.md
  • Troubleshooting: See modules/troubleshooting.md
  • Pressure Testing: See modules/pressure-testing.md
  • Integration Testing: See modules/integration-testing.md
  • Performance Benchmarking: See modules/performance-benchmarking.md

Tools and Automation

  • Tools: Executable analysis utilities in scripts/ directory.
  • Automation: Setup and validation scripts in scripts/automation/.

Exit Criteria

  • Every audited skill receives a score across all five dimensions (structure compliance, content quality, token efficiency, activation reliability, tool integration) summing to 100.
  • Any skill with a deprecated skills/shared/ module reference is listed as a structural warning in the audit output.
  • make check-compliance TARGET=<skill> exits 0 for skills reported as passing, confirming the compliance-checker agrees with the audit score.
  • An improvement plan is produced for any skill scoring below 75, with findings ordered by priority: critical > high > medium > low.
Files (claude-night-market)
  • modules
    • advanced-tool-use-analysis.md 3.3 KB
      # Advanced Tool Use Analysis
      
      ## Dynamic Discovery Evaluation
      
      ### Tool Discovery Patterns
      ```bash
      # Analyze tool discovery patterns and efficiency
      skills/skills-eval/scripts/tool-performance-analyzer --skill-path skill.md --focus discovery
      
      # Benchmark against optimal loading patterns
      skills/skills-eval/scripts/discovery-optimizer --skill-path skill.md --benchmark mcp-standards
      ```
      
      ### Discovery Optimization Targets
      - **Loading Efficiency**: Minimize tool discovery latency
      - **Pattern Recognition**: Optimize keyword matching and categorization
      - **Contextual Loading**: Load tools based on relevance to current context
      - **Memory Management**: Efficient tool caching and retrieval
      
      ## Programmatic Calling Assessment
      
      ### Multi-Step Workflow Analysis
      ```bash
      # Evaluate multi-step workflow optimization opportunities
      skills/skills-eval/scripts/tool-performance-analyzer --skill-path skill.md --focus programmatic-calling
      
      # Identify parallel execution opportunities
      skills/skills-eval/scripts/tool-performance-analyzer --skill-path skill.md --parallel-analysis
      ```
      
      ### Calling Optimization Metrics
      - **Sequential Efficiency**: Optimize ordered tool execution
      - **Parallel Processing**: Identify concurrent tool opportunities
      - **Context Preservation**: Minimize context loss between calls
      - **Error Recovery**: production-grade error handling and retry mechanisms
      
      ## Context Preservation Analysis
      
      ### Context Window Utilization
      ```bash
      # Measure context window utilization efficiency
      skills/skills-eval/scripts/token-usage-tracker --skill-path skill.md --context-analysis
      
      # Identify pollution reduction opportunities
      skills/skills-eval/scripts/token-usage-tracker --skill-path skill.md --pollution-analysis
      ```
      
      ### Optimization Strategies
      - **Efficient Token Usage**: Maximize information density
      - **Pollution Reduction**: Minimize irrelevant context accumulation
      - **Window Management**: Strategic context window allocation
      - **Compression Techniques**: Intelligent content summarization
      
      ## Performance Benchmarking
      
      ### Evaluation Criteria
      - **MCP Compliance**: Validation against Model Context Protocol standards
      - **Accuracy Metrics**: Tool discovery and execution accuracy improvements
      - **Token Efficiency**: Usage patterns and optimization opportunities
      - **Latency Analysis**: Multi-step workflow performance bottlenecks
      
      ### Target Improvements
      - **Token Usage Reduction**: Aim for 37% reduction through programmatic calling optimization
      - **Accuracy Improvements**: Target 25% improvement in tool discovery and execution
      - **Context Optimization**: Maintain 95% context window preservation
      - **Latency Reduction**: Eliminate multiple inference passes in complex workflows
      
      ## Advanced Analysis Techniques
      
      ### Comparative Analysis
      ```bash
      # Benchmark against best-in-class examples
      skills/skills-eval/scripts/performance-comparator --skill-path skill.md --baseline industry-standards
      
      # Trend tracking over time
      skills/skills-eval/scripts/performance-tracker --skill-path skill.md --metrics discovery,calling,context
      ```
      
      ### Optimization Recommendations
      1. **Tool Grouping**: Related tools should be discoverable together
      2. **Progressive Loading**: Load essential tools first, advanced tools later
      3. **Context Caching**: Preserve relevant context between tool calls
      4. **Error Patterns**: Analyze and optimize common error scenarios
      
    • authoring-checklist.md 2.6 KB
      # Skill Authoring Checklist
      
      Quick-reference validation checklist for skill authors.
      
      ## Pre-Development
      
      - [ ] Identified repeated task (done 5+ times, will do 10+ more)
      - [ ] Confirmed no existing skill covers this
      - [ ] Defined skill type (Technique, Pattern, or Reference)
      - [ ] Chosen descriptive gerund-form name
      
      ## Frontmatter Validation
      
      - [ ] `name`: ≤64 characters
      - [ ] `name`: lowercase letters, numbers, hyphens only
      - [ ] `name`: no reserved words (anthropic, claude)
      - [ ] `description`: non-empty
      - [ ] `description`: ≤1024 characters
      - [ ] `description`: third person voice
      - [ ] `description`: includes WHAT and WHEN
      
      ## Content Quality
      
      - [ ] SKILL.md body under 500 lines
      - [ ] Only context Claude doesn't already have
      - [ ] Consistent terminology throughout
      - [ ] No time-sensitive information
      - [ ] Concrete examples, not abstract
      - [ ] Clear distinction between skill types
      
      ## Structure
      
      - [ ] File references one level deep
      - [ ] No Table of Contents; a long file is split, not indexed
      - [ ] Progressive disclosure pattern used
      - [ ] Appropriate freedom level for task type
      
      ## TDD Compliance
      
      - [ ] Created 3+ pressure scenarios
      - [ ] Ran baseline without skill
      - [ ] Documented baseline failures verbatim
      - [ ] Tested with skill present
      - [ ] Identified rationalizations
      - [ ] Added explicit counters
      
      ## Anti-Rationalization
      
      - [ ] Listed specific exceptions
      - [ ] Created rationalization table
      - [ ] Added red flags list
      - [ ] Addressed "spirit vs letter" arguments
      
      ## Scripts (if applicable)
      
      - [ ] Scripts solve problems, don't punt to Claude
      - [ ] Error handling is explicit
      - [ ] No "voodoo constants"
      - [ ] Required packages listed
      - [ ] Clear execute vs read distinction
      - [ ] MCP tools use fully qualified names
      
      ## Testing
      
      - [ ] Tested with Haiku
      - [ ] Tested with Sonnet
      - [ ] Tested with Opus
      - [ ] Tested real usage scenarios
      - [ ] Team feedback incorporated
      
      ## Deployment
      
      - [ ] Committed to git
      - [ ] Pushed to fork
      - [ ] Validated structure passes
      - [ ] No narrative storytelling
      - [ ] Supporting files justified
      
      ## Common Mistakes
      
      Avoid these anti-patterns when authoring skills:
      
      | Do NOT | Do Instead |
      |--------|------------|
      | Mix multiple concerns in one skill | Single responsibility per skill |
      | Use vague language ("usually", "try to") | Explicit, mandatory language |
      | Embed large code blocks inline | Reference external tools/scripts |
      | Skip testing before deployment | Test with subagents across models |
      | Keep core skill over 300 lines | Extract complexity to modules/tools |
      | Skip trigger activation testing | Validate with `/skills-eval` |
      
    • evaluation-criteria.md 13.5 KB
      # Skill Evaluation Criteria
      
      Detailed scoring rubric and quality gates for skill evaluation.
      
      ## Mathematical Foundation
      
      This evaluation framework follows Multi-Criteria Decision Analysis
      (MCDA) best practices:
      
      - **Normalization**: Vector normalization for scale invariance
      - **Weighting**: AHP-derived weights with expert validation
      - **Aggregation**: Weighted sum with Pareto analysis for trade-offs
      - **Validation**: Sensitivity analysis on all weights
      
      Core metrics are weighted as: compliance (30%), effectiveness (30%),
      maintainability (20%), performance (20%).
      
      ## Scoring System (100 points total)
      
      ### Structure Compliance (20 points)
      
      | Aspect | Max Points | Requirements |
      |--------|------------|--------------|
      | YAML frontmatter | 5 | Complete, valid metadata within limits |
      | Progressive disclosure | 5 | SKILL.md <500 lines, links to modules |
      | Section organization | 5 | Logical flow, clear hierarchy |
      | File naming | 3 | Gerund form, descriptive names |
      | Reference depth | 2 | One-level deep references only |
      
      **Frontmatter Validation:**
      
      ```yaml
      name:
        max_length: 64
        pattern: "^[a-z0-9-]+$"
        forbidden: ["XML tags", "reserved words"]
      
      description:
        max_length: 1024
        forbidden: ["XML tags"]
        required: ["WHAT it does", "WHEN to use it"]
      ```
      
      **Official Frontmatter Fields** (per Claude Code docs):
      
      | Field | Type | Validation |
      |-------|------|------------|
      | `name` | string | Required, kebab-case, max 64 chars |
      | `description` | string | Recommended, front-load use case |
      | `disable-model-invocation` | boolean | `true` or `false` |
      | `user-invocable` | boolean | `true` or `false` |
      | `allowed-tools` | string | Comma-separated tool names |
      | `model` | string | Valid model name |
      | `effort` | string | `low`, `medium`, `high`, `xhigh` (Opus 4.7), `max` (Opus 4.7) |
      | `context` | string | Only `fork` supported |
      | `agent` | string | Subagent type (with `context: fork`) |
      | `argument-hint` | string | Shown in autocomplete |
      | `paths` | string/array | Glob patterns for activation |
      | `shell` | string | `bash` or `powershell` |
      | `hooks` | object | Hooks scoped to skill lifecycle |
      
      Extension fields (`version`, `category`, `tags`, etc.)
      are permitted but not validated.
      
      ### Content Quality (25 points)
      
      | Aspect | Max Points | Requirements |
      |--------|------------|--------------|
      | Quick Start concreteness | 8 | Actual commands, not abstract descriptions |
      | Clarity and completeness | 6 | Clear explanations, no ambiguity |
      | Practical examples | 6 | Input/output pairs, real patterns |
      | Voice consistency | 5 | Third person, no "your"/"you" language |
      
      **Cargo Cult Anti-Pattern Detection:**
      
      Skills must avoid cargo cult patterns - rituals that "look right" but lack verification:
      
      - ❌ **Abstract Quick Start**: "Configure pytest" vs ✅ "Run `pytest --cov=src` to generate coverage"
      - ❌ **Testing Theater**: Tests that always pass (`assert True`) vs ✅ Behavior-driven tests that fail when mutated
      - ❌ **Implementation Testing**: Testing HOW not WHAT vs ✅ Testing behavior via Given-When-Then scenarios
      - ❌ **Missing Verification**: Code examples without validation steps vs ✅ "Run `pytest -v` to confirm"
      - ❌ **Documentation Exception**: "It's just markdown" vs ✅ All files have testable structure
      
      **Quick Start Requirements:**
      
      - ✅ **Good**: "Run `pytest --cov=src` to generate coverage reports"
      - ❌ **Bad**: "Configure pytest and implement tests"
      
      **Voice Consistency Checklist:**
      
      - [ ] No "your needs" → use "project requirements"
      - [ ] No "you should" → use imperative "Run X to Y"
      - [ ] No "you can" → use "Available options include"
      - [ ] Third person throughout ("the skill", "users", "developers")
      
      ### Token Efficiency (20 points)
      
      | Aspect | Max Points | Requirements |
      |--------|------------|--------------|
      | Content density | 5 | Concise, no unnecessary explanation |
      | Progressive loading | 5 | Essential content first |
      | Navigation aids | 6 | TOCs in modules >100 lines |
      | Context optimization | 4 | Efficient context usage |
      
      **Navigation Requirements (Critical):**
      
      ```yaml
      navigation_rules:
        - condition: "a Table of Contents in SKILL.md or any module"
          requirement: "No Table of Contents"
          penalty: "-2 points"
          rationale: "An anchor list restates the headings below it and costs tokens on every load; grep finds the headings directly"
      
        - condition: "module_length > 200 lines"
          requirement: "Section anchors and backlinks"
          penalty: "-3 points"
      ```
      
      ### Activation Reliability (20 points)
      
      | Aspect | Max Points | Requirements |
      |--------|------------|--------------|
      | Description triggers | 7 | 5+ specific trigger phrases |
      | Context indicators | 5 | Clear usage scenarios |
      | Trigger clarity | 5 | Differentiates from alternatives |
      | Discovery patterns | 3 | Easy to find and categorize |
      
      **Trigger Phrase Requirements:**
      
      - Minimum 5 specific phrases in description
      - Include domain-specific terminology
      - Cover common task descriptions
      - Examples: "pytest fixtures", "unittest replacement", "test coverage"
      
      **Example Good Description:**
      
      ```yaml
      description: |
        Pytest testing framework with async support and fixtures.
      
        Triggers: pytest, async testing, unittest replacement, test fixtures,
        test coverage, mocking, pytest fixtures, parametrized tests
      
        Use when: writing tests with pytest, migrating from unittest, setting up
        test infrastructure, implementing async tests
      ```
      
      ### Tool Integration (10 points)
      
      | Aspect | Max Points | Requirements |
      |--------|------------|--------------|
      | Script quality | 4 | Solves errors explicitly |
      | Verification steps | 3 | Post-example validation |
      | Configuration clarity | 2 | No magic numbers |
      | Execute vs read | 1 | Clear usage intent |
      
      **Verification Requirements:**
      
      After each code example, include validation steps:
      
      ```markdown
      ### Example: Async Test with Fixture
      
      \`\`\`python
      @pytest.mark.asyncio
      async def test_async_operation():
          result = await async_operation()
          assert result.status == "success"
      \`\`\`
      
      **Verification:** Run `pytest -v tests/test_async.py::test_async_operation` to confirm the async test executes correctly.
      ```
      
      ### Documentation Completeness (5 points)
      
      | Aspect | Max Points | Requirements |
      |--------|------------|--------------|
      | Troubleshooting | 2 | Common issues documented |
      | Reference materials | 2 | Complete API references |
      | Time-sensitivity | 1 | No date-dependent instructions |
      
      ### Degrees of Freedom Alignment (bonus, up to 5 points)
      
      | Aspect | Max Points | Requirements |
      |--------|------------|--------------|
      | Task-specificity match | 2 | Freedom level matches task fragility |
      | Workflow structure | 2 | Complex tasks have trackable checklists |
      | Feedback loops | 1 | Validation steps before proceeding |
      
      ### Persuasion Effectiveness (bonus, up to 5 points)
      
      | Aspect | Max Points | Requirements |
      |--------|------------|--------------|
      | Authority usage | 2 | Imperative language for critical rules |
      | Commitment patterns | 1 | Explicit declarations required |
      | Social proof | 1 | Universal norms documented |
      | Model calibration | 1 | Language appropriate for target models |
      
      ### Anti-Rationalization Coverage (bonus, up to 5 points)
      
      | Aspect | Max Points | Requirements |
      |--------|------------|--------------|
      | Loophole closures | 2 | Specific exceptions listed |
      | Rationalization table | 1 | Common excuses with counters |
      | Red flags list | 1 | Self-checking triggers documented |
      | Foundational principles | 1 | "Spirit vs letter" addressed early |
      
      ## Quality Levels
      
      | Score | Level | Description |
      |-------|-------|-------------|
      | 91-100 | Excellent (A) | Production-ready, reference implementation |
      | 76-90 | Good (B) | Strong with minor improvements needed |
      | 51-75 | Fair (C) | Functional but needs significant work |
      | 26-50 | Poor (D) | Major issues need addressing |
      | 0-25 | Critical (F) | Fundamental problems |
      
      ## Quality Gates
      
      Default thresholds for CI/CD integration:
      
      ```yaml
      quality_gates:
        structure_score: ">= 80"
        content_score: ">= 85"
        token_efficiency: ">= 75"
        activation_score: ">= 80"
        tool_integration: ">= 70"
        overall_score: ">= 75"
        max_critical_issues: 0
        max_high_issues: 3
      ```
      
      ### Sensitivity Analysis Requirements
      
      Before finalizing quality gates, run sensitivity analysis:
      
      ```yaml
      sensitivity_analysis:
        variation: 0.20  # ±20% weight variation
      
        requirements:
          stable_rankings: true  # Rankings shouldn't change
          critical_weights_identified: true  # Document sensitive weights
      
        critical_threshold: 0.8  # Spearman correlation < 0.8 = sensitive
      ```
      
      ### Gate Behaviors
      
      | Gate | Failure Action |
      |------|----------------|
      | `structure_score` | Block deployment, fix frontmatter |
      | `content_score` | Warn, suggest improvements |
      | `token_efficiency` | Warn, recommend modularization |
      | `activation_score` | Block until triggers improved |
      | `max_critical_issues` | Immediate block |
      
      ## Issue Classification
      
      ### Critical Issues (Immediate Action Required)
      
      - Missing YAML frontmatter
      - Invalid frontmatter schema
      - No trigger phrases in description
      - Security vulnerabilities
      - Broken functionality
      - References to deleted `skills/shared/modules/` files (broken links)
      
      ### High Issues (Address Before Next Release)
      
      - Table of Contents in SKILL.md or a module (-2 points)
      - Abstract Quick Start without commands (-2 points)
      - Second-person voice slips ("your"/"you") (-1 point)
      - Missing verification steps after examples (-1 point)
      - Fewer than 5 trigger phrases (-1 point)
      - SKILL.md exceeds 500 lines
      - Modules in deprecated `skills/shared/` directory (-2 points, relocate to skill-specific `modules/`)
      
      ### Medium Issues (Address Soon)
      
      - Suboptimal content density
      - Weak context indicators
      - Limited trigger phrases
      - Missing troubleshooting section
      
      ### Low Issues (Nice to Fix)
      
      - Minor formatting inconsistencies
      - Additional example enhancements
      - Documentation polish
      - Performance optimizations
      
      ## Evaluation Report Format
      
      ### Summary Format
      
      ```
      === Skills Evaluation Report ===
      Plugin: {name} (v{version})
      Scope: {scope}
      Total skills: {count}
      
      === Scores ===
      Structure:      {score}/100 ({level})
      Content:        {score}/100 ({level})
      Token Efficiency: {score}/100 ({level})
      Activation:     {score}/100 ({level})
      Tool Integration: {score}/100 ({level})
      Documentation:  {score}/100 ({level})
      ────────────────────────────────
      Overall:        {score}/100 ({level})
      
      === Issues ===
      Critical: {count}
      High: {count}
      Medium: {count}
      Low: {count}
      ```
      
      ### Detailed Format
      
      Includes per-skill breakdown:
      
      ```
      === Skill: {skill_name} ===
      Path: {skill_path}
      Lines: {line_count} (SKILL.md: {skill_lines})
      
      Structure Issues:
        [HIGH] Module async-testing.md (192 lines) carries a Table of Contents (-2 points)
        [MEDIUM] SKILL.md exceeds 500 lines (-1 point)
      
      Content Issues:
        [HIGH] Quick Start too abstract: "Configure pytest" (-2 points)
        [HIGH] Second-person voice: "your needs" at line 74 (-1 point)
        [HIGH] Missing verification after async test example (-1 point)
      
      Activation Issues:
        [HIGH] Only 2 trigger phrases in description (-1 point)
      
      Recommendations:
        1. Remove the Table of Contents from async-testing.md; split the module if it needs navigation
        2. Update Quick Start with actual pytest commands
        3. Replace "your needs" with "project requirements"
        4. Add "Run pytest -v tests/test_async.py" after async example
        5. Expand description with: "unittest replacement", "test coverage"
      
      Impact: +7 points → 96/100 (A grade)
      ```
      
      ## Customization
      
      ### Per-Skill Configuration
      
      Create `.skills-eval.yaml` in plugin root:
      
      ```yaml
      skills_eval:
        # Override thresholds
        thresholds:
          skill_max_lines: 500
          module_max_lines: 100
          min_trigger_phrases: 5
          min_description_length: 50
      
        # Content requirements
        content_requirements:
          require_quick_start: true
          require_verification_steps: true
          require_troubleshooting: true
      
        # Voice and style
        style_requirements:
          forbid_second_person: true
          require_third_person: true
          forbidden_phrases:
            - "your needs"
            - "you should"
            - "you can"
      
        # Navigation requirements
        navigation_requirements:
          toc_threshold_lines: 100
          toc_section_anchors: true
          toc_backlinks: true
      
        # Custom rules
        custom_rules:
          - name: "require-examples"
            applies_to: ["implementation-skills"]
            check: "count_code_examples >= 3"
            severity: "medium"
      
          - name: "require-tools"
            applies_to: ["automation-skills"]
            check: "tools.length > 0"
            severity: "high"
      
        # Excluded paths
        exclude_paths:
          - "skills/experimental/*"
          - "skills/deprecated/*"
      ```
      
      ## Quick Reference Checklist
      
      Use this checklist when reviewing skills:
      
      ### Structure (20 points)
      - [ ] Valid YAML frontmatter with all required fields
      - [ ] SKILL.md under 500 lines
      - [ ] Clear section organization
      - [ ] One-level deep references only
      - [ ] Descriptive file names (gerund form)
      
      ### Content (25 points)
      - [ ] Quick Start has actual commands
      - [ ] Clear, complete explanations
      - [ ] Practical input/output examples
      - [ ] Third-person voice throughout
      - [ ] No "your"/"you" language
      
      ### Token Efficiency (20 points)
      - [ ] Concise, dense content
      - [ ] Progressive disclosure structure
      - [ ] **TOCs in all modules >100 lines**
      - [ ] Efficient context usage
      
      ### Activation (20 points)
      - [ ] **5+ trigger phrases in description**
      - [ ] Clear context indicators
      - [ ] Differentiates from alternatives
      - [ ] Easy to discover
      
      ### Tools (10 points)
      - [ ] Scripts solve errors explicitly
      - [ ] **Verification steps after examples**
      - [ ] No magic numbers
      - [ ] Clear execute vs read intent
      
      ### Documentation (5 points)
      - [ ] Troubleshooting section exists
      - [ ] Complete reference materials
      - [ ] No time-sensitive language
      
    • evaluation-framework.md 3 KB
      # Evaluation Framework
      
      ## Quality Assessment System
      
      ### Evaluation Criteria and Scoring
      
      #### Priority Levels
      1. **Critical**: Security issues, broken functionality, missing required fields
      2. **High**: Poor structure, incomplete documentation, performance issues
      3. **Medium**: Missing best practices, optimization opportunities
      4. **Low**: Minor improvements, formatting issues, enhanced examples
      
      #### Scoring System
      - **0-25**: Needs significant improvement
      - **26-50**: Below acceptable standards
      - **51-75**: Meets basic requirements
      - **76-90**: Good quality with minor issues
      - **91-100**: Excellent quality, best practices
      
      ### Component Analysis
      
      #### Structure Compliance (25 points)
      - **YAML Frontmatter** (5 points): Complete, valid metadata
      - **Progressive Disclosure** (5 points): Overview/Quick Start structure
      - **Section Organization** (5 points): Logical flow and clarity
      - **Content Hierarchy** (5 points): Proper heading structure
      - **File Organization** (5 points): Appropriate modularization
      
      #### Content Quality (25 points)
      - **Clarity and Completeness** (7 points): Clear explanations
      - **Practical Examples** (6 points): Real-world implementations
      - **User Experience** (6 points): Accessibility and usability
      - **Documentation Standards** (6 points): Consistent formatting
      
      #### Token Efficiency (20 points)
      - **Content Density** (5 points): Concise information presentation
      - **Progressive Loading** (5 points): Essential content priority
      - **Context Optimization** (5 points): Efficient context usage
      - **Modular Design** (5 points): Appropriate content separation
      
      #### Activation Reliability (20 points)
      - **Trigger Effectiveness** (6 points): Strong activation keywords
      - **Context Indicators** (6 points): Clear usage scenarios
      - **Discovery Patterns** (4 points): Easy categorization
      - **Loading Consistency** (4 points): Predictable behavior
      
      #### Tool Integration (10 points)
      - **Executable Components** (4 points): Available automation tools
      - **API Integration** (3 points): External service connections
      - **Workflow Support** (3 points): End-to-end automation
      
      ### Improvement Prioritization
      
      #### Critical Issues (Immediate Action Required)
      - Security vulnerabilities
      - Broken functionality
      - Missing required metadata
      - Poor activation reliability
      
      #### High Priority (Address in Next Update)
      - Performance optimization
      - Documentation gaps
      - Token usage reduction
      - Tool integration improvements
      
      #### Medium Priority (Future Enhancements)
      - Advanced features
      - Additional examples
      - Enhanced error handling
      - Integration expansions
      
      #### Low Priority (Nice to Have)
      - Formatting improvements
      - Additional documentation
      - Edge case handling
      - Performance tuning
      
      ### Quality Gates
      
      #### Minimum Acceptable Standards
      - **Overall Score**: 70+ points
      - **Structure Compliance**: 15+ points
      - **Content Quality**: 15+ points
      - **Token Efficiency**: 10+ points
      - **Activation Reliability**: 12+ points
      
      #### Excellence Standards
      - **Overall Score**: 90+ points
      - **All Categories**: 18+ points each
      - **Zero Critical Issues**
      - **Maximum One High Priority Issue
      
    • evaluation-workflows.md 2.1 KB
      # Evaluation Workflows and Techniques
      
      ## Detailed Implementation Steps
      
      ### Phase 1: Discovery and Assessment
      1. **Run discovery**: Use `skills-auditor --discover` to locate all skills
      2. **Initial scan**: Execute `skills-auditor --scan-all` for overview
      3. **Identify patterns**: Look for common issues and improvement opportunities
      4. **Set baselines**: Establish quality metrics and improvement targets
      
      ### Phase 2: Detailed Analysis
      1. **Deep analysis**: Use `skill-analyzer --path skill.md --verbose` for complex skills
      2. **Token evaluation**: Run `token-estimator -f skill.md` for usage analysis
      3. **Compliance checking**: Execute `compliance-checker --skill-path skill.md`
      4. **Gap analysis**: Compare against best practices and standards
      
      ### Phase 3: Improvement Planning
      1. **Generate recommendations**: Use `improvement-suggester --skill-path skill.md --priority high`
      2. **Prioritize improvements**: Focus on critical and high-priority items first
      3. **Create action plans**: Break improvements into manageable tasks
      4. **Schedule implementation**: Plan improvement work in logical phases
      
      ### Phase 4: Implementation and Validation
      1. **Apply improvements**: Implement changes based on recommendations
      2. **Test functionality**: Verify tools and examples work correctly
      3. **Validate compliance**: Re-run compliance checks
      4. **Measure results**: Compare before/after quality scores
      
      ## Advanced Analysis Techniques
      
      ### Comparative Analysis
      - Benchmark skills against best-in-class examples
      - Identify patterns in high-performing skills
      - Learn from structural and content differences
      
      ### Trend Tracking
      - Run periodic audits to monitor quality over time
      - Track improvement implementation success rates
      - Identify recurring issues and systemic problems
      
      ### Gap Analysis
      - Identify missing skill categories in your library
      - Find opportunities for new skill development
      - Balance skill coverage across domains and use cases
      
      ### Dependency Mapping
      - Understand skill relationships and interactions
      - Identify potential circular dependencies
      - Optimize skill loading patterns and efficiency
      
    • integration-testing.md 9.6 KB
      # Integration Testing Framework
      
      ## Overview
      
      The integration testing framework validates skills work correctly with Claude Agent SDK features, tool integrations, and context management patterns. This module provides detailed testing strategies for validating skill functionality in real-world usage scenarios.
      
      ## Testing Categories
      
      ### 1. Basic Functionality Testing
      
      **Purpose:** Verify core skill features work as intended
      
      **Test Cases:**
      - Skill loads successfully without errors
      - Frontmatter metadata is valid and complete
      - Required sections are present and accessible
      - Progressive disclosure works correctly
      - Module references resolve properly
      
      **Validation Methods:**
      ```python
      def test_basic_functionality(skill_path: str) -> FunctionalityResults:
          """Test basic skill loading and structure"""
          results = FunctionalityResults()
      
          # Test 1: Skill loads without errors
          try:
              content = Path(skill_path).read_text()
              results.loads_successfully = True
          except Exception as e:
              results.errors.append(f"Failed to load: {e}")
              return results
      
          # Test 2: Frontmatter is valid
          if content.startswith("---\n"):
              results.valid_frontmatter = True
      
          # Test 3: Required sections present
          required_sections = ["## Overview", "## When to Use"]
          for section in required_sections:
              if section in content:
                  results.sections_present.append(section)
      
          return results
      ```
      
      ### 2. Tool Integration Testing
      
      **Purpose:** Verify tool declarations and integrations work correctly
      
      **Test Cases:**
      - Tool declarations match actual available tools
      - Tool scripts are executable and error-free
      - Tool dependencies are properly declared
      - Error handling for missing tools
      - Tool output formats are valid
      
      **Example Tests:**
      ```python
      def test_tool_integration(skill_path: str) -> ToolIntegrationResults:
          """Test tool compatibility and integration"""
          results = ToolIntegrationResults()
      
          # Parse skill frontmatter
          frontmatter = parse_frontmatter(skill_path)
          declared_tools = frontmatter.get("tools", [])
      
          # Test each declared tool
          for tool in declared_tools:
              tool_path = find_tool(tool, skill_path)
              if tool_path and tool_path.exists():
                  results.tools_found.append(tool)
      
                  # Test tool is executable
                  if os.access(tool_path, os.X_OK):
                      results.tools_executable.append(tool)
      
                  # Test tool runs with --help
                  try:
                      subprocess.run([tool_path, "--help"], capture_output=True, timeout=5)
                      results.tools_functional.append(tool)
                  except Exception as e:
                      results.tool_errors.append(f"{tool}: {e}")
              else:
                  results.tools_missing.append(tool)
      
          return results
      ```
      
      ### 3. Context Management Testing
      
      **Purpose:** Verify context optimization and efficiency
      
      **Test Cases:**
      - Token usage stays within declared limits
      - Progressive disclosure reduces initial context
      - Module references don't cause circular dependencies
      - Context compression works effectively
      - Lazy loading patterns function correctly
      
      **Metrics:**
      ```python
      def test_context_management(skill_path: str) -> ContextResults:
          """Test context optimization and efficiency"""
          results = ContextResults()
      
          content = Path(skill_path).read_text()
          frontmatter = parse_frontmatter(content)
      
          # Test 1: Token usage
          estimated_tokens = len(content) // 4
          declared_tokens = frontmatter.get("estimated_tokens", 0)
      
          results.estimated_tokens = estimated_tokens
          results.declared_tokens = declared_tokens
          results.token_accuracy = abs(estimated_tokens - declared_tokens) / estimated_tokens
      
          # Test 2: Progressive disclosure
          if "## Overview" in content and "modules/" in content.lower():
              results.has_progressive_disclosure = True
      
          # Test 3: Module references
          module_refs = re.findall(r"modules/([a-z-]+\.md)", content, re.IGNORECASE)
          results.module_references = module_refs
      
          return results
      ```
      
      ### 4. Claude SDK API Compliance
      
      **Purpose:** validate compatibility with Claude Agent SDK standards
      
      **Test Cases:**
      - Metadata follows SDK specifications
      - Tool declarations use correct format
      - Usage patterns are properly declared
      - Dependencies are correctly specified
      - Version compatibility is indicated
      
      **Compliance Checks:**
      ```python
      def test_sdk_compliance(skill_path: str) -> ComplianceResults:
          """Test Claude SDK API compliance"""
          results = ComplianceResults()
      
          frontmatter = parse_frontmatter(skill_path)
      
          # Required SDK fields
          required_fields = ["name", "description", "version", "category"]
          for field in required_fields:
              if field in frontmatter:
                  results.required_fields_present.append(field)
              else:
                  results.missing_fields.append(field)
      
          # Recommended SDK fields (2024 standards)
          recommended_fields = ["provides", "estimated_tokens", "usage_patterns"]
          for field in recommended_fields:
              if field in frontmatter:
                  results.recommended_fields_present.append(field)
      
          # Check for SDK compatibility declaration
          if "sdk_features" in frontmatter.get("provides", {}):
              results.sdk_compatible = True
      
          return results
      ```
      
      ## Integration Test Suite
      
      ### Complete Test Runner
      
      ```python
      class IntegrationTester:
          """detailed integration testing for Claude Skills"""
      
          def test_skill_integration(self, skill_path: str) -> IntegrationTestResults:
              """Run complete integration test suite"""
              results = IntegrationTestResults()
      
              # Run all test categories
              results.basic_functionality = self.test_basic_functionality(skill_path)
              results.tool_integration = self.test_tool_compatibility(skill_path)
              results.context_handling = self.test_context_optimization(skill_path)
              results.api_compliance = self.test_sdk_compliance(skill_path)
      
              # Calculate overall score
              results.overall_score = self._calculate_overall_score(results)
      
              # Generate recommendations
              results.recommendations = self._generate_recommendations(results)
      
              return results
      
          def _calculate_overall_score(self, results: IntegrationTestResults) -> float:
              """Calculate weighted overall integration score"""
              scores = {
                  "functionality": self._score_functionality(results.basic_functionality)
                  * 0.25,
                  "tools": self._score_tool_integration(results.tool_integration) * 0.25,
                  "context": self._score_context_handling(results.context_handling) * 0.25,
                  "compliance": self._score_compliance(results.api_compliance) * 0.25,
              }
              return sum(scores.values())
      ```
      
      ## Running Integration Tests
      
      ### Command-Line Usage
      
      ```bash
      # Run complete integration test suite
      ./scripts/integration-tester --skill-path path/to/skill/SKILL.md
      
      # Test specific categories
      ./scripts/integration-tester --skill-path path/to/skill/SKILL.md \
        --tests functionality,tools,context
      
      # Generate detailed report
      ./scripts/integration-tester --skill-path path/to/skill/SKILL.md \
        --format markdown --output results.md
      
      # Batch test all skills
      ./scripts/integration-tester --scan-all --format table
      ```
      
      ### Integration with CI/CD
      
      ```yaml
      # .github/workflows/skill-integration-tests.yml
      name: Skill Integration Tests
      
      on: [push, pull_request]
      
      jobs:
        test:
          runs-on: ubuntu-latest
          steps:
            - uses: actions/checkout@v3
            - name: Run Integration Tests
              run: |
                cd skills/skills-eval
                ./scripts/integration-tester --scan-all --format json > results.json
            - name: Upload Results
              uses: actions/upload-artifact@v3
              with:
                name: integration-test-results
                path: results.json
      ```
      
      ## Best Practices
      
      ### 1. Test Coverage
      
      - **detailed:** Test all declared features and tools
      - **Realistic:** Use real-world usage scenarios
      - **Automated:** Integrate tests into development workflow
      - **Continuous:** Run tests on every skill update
      
      ### 2. Test Data
      
      - **Valid Skills:** Test with well-formed skills first
      - **Edge Cases:** Test with missing fields, broken tools, etc.
      - **Performance:** Test with large skills (token limits)
      - **Compatibility:** Test across different skill versions
      
      ### 3. Result Interpretation
      
      **Scoring Guidelines:**
      - **91-100:** Excellent integration, production-ready
      - **76-90:** Good integration, minor improvements needed
      - **51-75:** Acceptable, several issues to address
      - **Below 50:** Significant integration problems
      
      ### 4. Continuous Improvement
      
      - Track test results over time
      - Identify common failure patterns
      - Update tests for new SDK features
      - Maintain test documentation
      
      ## Advanced Testing Patterns
      
      ### Subagent-Based Testing
      
      Use Claude Code subagents for specialized testing:
      
      ```python
      def test_with_specialist_agents(skill_path: str) -> Dict:
          """Deploy specialist subagents for detailed testing"""
          specialists = {
              "performance": "agents-network-engineer",
              "debugging": "superpowers:systematic-debugging",
              "documentation": "elements-of-style:writing-clearly-and-concisely",
          }
      
          results = {}
          for specialty, agent in specialists.items():
              results[specialty] = deploy_agent_test(agent, skill_path)
      
          return results
      ```
      
      ### Performance Testing
      
      ```python
      def test_performance_characteristics(skill_path: str) -> PerformanceResults:
          """Test skill performance under load"""
          results = PerformanceResults()
      
          # Load time testing
          start = time.time()
          load_skill(skill_path)
          results.load_time = time.time() - start
      
          # Memory usage
          results.memory_usage = measure_memory_usage(skill_path)
      
          # Token efficiency
          results.token_efficiency = calculate_token_efficiency(skill_path)
      
          return results
      ```
      
    • integration.md 3.3 KB
      # Integration with Other Skills
      
      ## Complementary Skills
      
      ### modular-skills
      - **Purpose**: Provides structural analysis and modular design patterns
      - **Integration**: Uses `skill-analyzer` and `token-estimator` tools
      - **Workflow**: Run structural analysis before detailed evaluation
      - **Benefits**: validates proper modularization and token efficiency
      
      ```bash
      # Combined workflow example
      scripts/skill-analyzer --path skill.md --verbose
      skills/skills-eval/scripts/compliance-checker --skill-path skill.md
      skills/skills-eval/scripts/improvement-suggester --skill-path skill.md
      ```
      
      ### testing-skills
      - **Purpose**: Validation and testing patterns for skills
      - **Integration**: Compatible with test-driven development approaches
      - **Workflow**: Use evaluation results to inform test coverage
      - **Benefits**: detailed quality assurance across all dimensions
      
      ### documentation-standards
      - **Purpose**: validates consistent documentation practices
      - **Integration**: Aligns with documentation best practices
      - **Workflow**: Use documentation evaluation as part of overall assessment
      - **Benefits**: Unified documentation approach across skill ecosystem
      
      ### workflow-automation
      - **Purpose**: Automates skill development and maintenance workflows
      - **Integration**: Provides CI/CD patterns for skill management
      - **Workflow**: Integrate evaluation into automated pipelines
      - **Benefits**: Continuous quality monitoring and improvement
      
      ## Workflow Integration
      
      ### Development Pipeline
      ```bash
      # Pre-commit evaluation
      skills/skills-eval/scripts/compliance-checker --skill-path skill.md --auto-fix
      
      # Post-development analysis
      skills/skills-eval/scripts/skills-auditor --scan-all --format markdown
      skills/skills-eval/scripts/improvement-suggester --skill-path skill.md --priority high
      ```
      
      ### Continuous Integration
      ```yaml
      # Example CI configuration
      evaluation_pipeline:
        steps:
          - name: "Compliance Check"
            run: skills/skills-eval/scripts/compliance-checker --directory . --format json
          - name: "Quality Assessment"
            run: skills/skills-eval/scripts/skills-auditor --scan-all
          - name: "Improvement Analysis"
            run: skills/skills-eval/scripts/improvement-suggester --directory .
      ```
      
      ### Monitoring and Alerting
      - **Quality Thresholds**: Set minimum acceptable scores
      - **Performance Metrics**: Monitor token usage and activation rates
      - **Security Scanning**: Regular compliance and security checks
      - **Trend Analysis**: Track quality improvements over time
      
      ## Best Practices
      
      ### Evaluation Frequency
      - **Pre-commit**: Quick compliance and security checks
      - **Pre-release**: detailed evaluation and improvement planning
      - **Periodic**: Quarterly full-skill inventory and assessment
      - **Triggered**: After major changes or updates
      
      ### Integration Strategies
      1. **Progressive Enhancement**: Start with basic evaluation, add advanced features
      2. **Custom Thresholds**: Set quality gates appropriate to your context
      3. **Automated Workflows**: Integrate evaluation into development pipelines
      4. **Continuous Improvement**: Use evaluation results for ongoing optimization
      
      ### Tool Orchestration
      - **Discovery First**: Always run skills discovery before detailed analysis
      - **Prioritization**: Use improvement suggester to focus efforts
      - **Validation**: Verify fixes with compliance checker
      - **Monitoring**: Track quality metrics over time
      
    • performance-benchmarking.md 15.9 KB
      # Performance Benchmarking Framework
      
      ## Overview
      
      The performance benchmarking framework provides detailed analysis of skill performance characteristics, including execution speed, memory usage, token efficiency, and scalability. This enables data-driven optimization decisions and validates skills meet performance standards.
      
      ## Benchmarking Categories
      
      ### 1. Execution Performance
      
      **Metrics:**
      - Load time (skill parsing and initialization)
      - Tool execution time
      - Response generation time
      - End-to-end latency
      
      **Measurement Strategy:**
      ```python
      class ExecutionBenchmark:
          """Benchmark execution performance"""
      
          def measure_load_time(self, skill_path: str) -> float:
              """Measure skill load time"""
              start = time.perf_counter()
              content = Path(skill_path).read_text()
              frontmatter = parse_frontmatter(content)
              end = time.perf_counter()
              return (end - start) * 1000  # milliseconds
      
          def measure_tool_execution(self, tool_path: str) -> Dict[str, float]:
              """Measure tool execution time"""
              results = {}
      
              # Warm-up run
              subprocess.run([tool_path, "--help"], capture_output=True)
      
              # Benchmark runs
              times = []
              for _ in range(10):
                  start = time.perf_counter()
                  subprocess.run([tool_path, "--help"], capture_output=True)
                  end = time.perf_counter()
                  times.append((end - start) * 1000)
      
              results["mean"] = statistics.mean(times)
              results["median"] = statistics.median(times)
              results["std_dev"] = statistics.stdev(times)
              results["min"] = min(times)
              results["max"] = max(times)
      
              return results
      ```
      
      ### 2. Memory Usage Profiling
      
      **Metrics:**
      - Baseline memory usage
      - Peak memory usage during execution
      - Memory efficiency ratio
      - Memory leak detection
      
      **Profiling Methods:**
      ```python
      class MemoryProfiler:
          """Profile memory usage characteristics"""
      
          def profile_skill_memory(self, skill_path: str) -> MemoryProfile:
              """detailed memory profiling"""
              import tracemalloc
      
              profile = MemoryProfile()
      
              # Start tracking
              tracemalloc.start()
              baseline = tracemalloc.get_traced_memory()[0]
      
              # Load skill
              content = Path(skill_path).read_text()
              parse_frontmatter(content)
      
              # Measure peak usage
              current, peak = tracemalloc.get_traced_memory()
              tracemalloc.stop()
      
              profile.baseline_kb = baseline / 1024
              profile.peak_kb = peak / 1024
              profile.current_kb = current / 1024
              profile.efficiency_ratio = baseline / peak if peak > 0 else 1.0
      
              return profile
      
          def check_memory_leaks(self, skill_path: str, iterations: int = 100) -> bool:
              """Detect memory leaks through repeated loading"""
              import gc
              import tracemalloc
      
              tracemalloc.start()
              initial = tracemalloc.get_traced_memory()[0]
      
              for _ in range(iterations):
                  content = Path(skill_path).read_text()
                  parse_frontmatter(content)
                  gc.collect()
      
              final = tracemalloc.get_traced_memory()[0]
              tracemalloc.stop()
      
              # Memory leak if final usage significantly higher than initial
              leak_threshold = initial * 1.1  # 10% increase
              return final > leak_threshold
      ```
      
      ### 3. Token Efficiency Benchmarking
      
      **Metrics:**
      - Tokens per feature ratio
      - Context compression effectiveness
      - Token reduction potential
      - Comparative efficiency scores
      
      **Analysis Methods:**
      ```python
      class TokenEfficiencyBenchmark:
          """Benchmark token usage efficiency"""
      
          def benchmark_token_efficiency(self, skill_path: str) -> TokenBenchmark:
              """detailed token efficiency analysis"""
              benchmark = TokenBenchmark()
      
              content = Path(skill_path).read_text()
              frontmatter = parse_frontmatter(content)
      
              # Calculate base metrics
              benchmark.char_count = len(content)
              benchmark.estimated_tokens = self.estimate_tokens(content)
              benchmark.declared_tokens = frontmatter.get("estimated_tokens", 0)
      
              # Calculate efficiency ratios
              features = len(frontmatter.get("usage_patterns", []))
              benchmark.tokens_per_feature = (
                  benchmark.estimated_tokens / features if features > 0 else 0
              )
      
              # Compare against targets
              targets = {"excellent": 1500, "good": 2000, "acceptable": 2500}
              benchmark.efficiency_category = self._categorize_efficiency(
                  benchmark.estimated_tokens, targets
              )
      
              # Calculate optimization potential
              if benchmark.estimated_tokens > targets["good"]:
                  benchmark.optimization_potential = (
                      (benchmark.estimated_tokens - targets["good"])
                      / benchmark.estimated_tokens
                      * 100
                  )
      
              return benchmark
      
          def comparative_benchmark(self, skill_paths: List[str]) -> ComparativeBenchmark:
              """Compare efficiency across multiple skills"""
              results = ComparativeBenchmark()
      
              for skill_path in skill_paths:
                  metrics = self.benchmark_token_efficiency(skill_path)
                  results.add_skill(Path(skill_path).stem, metrics)
      
              # Calculate percentiles
              all_tokens = [m.estimated_tokens for m in results.skill_metrics.values()]
              results.percentile_25 = statistics.quantiles(all_tokens, n=4)[0]
              results.percentile_50 = statistics.median(all_tokens)
              results.percentile_75 = statistics.quantiles(all_tokens, n=4)[2]
      
              return results
      ```
      
      ### 4. Scalability Testing
      
      **Metrics:**
      - Performance under load
      - Concurrent execution handling
      - Large dataset processing
      - Resource utilization patterns
      
      **Test Methods:**
      ```python
      class ScalabilityBenchmark:
          """Test skill scalability characteristics"""
      
          def test_concurrent_execution(
              self, skill_path: str, concurrency_levels: List[int]
          ) -> Dict[int, float]:
              """Test performance at different concurrency levels"""
              results = {}
      
              for level in concurrency_levels:
                  with ThreadPoolExecutor(max_workers=level) as executor:
                      start = time.perf_counter()
                      futures = [
                          executor.submit(self._load_skill, skill_path)
                          for _ in range(level * 10)
                      ]
                      for future in futures:
                          future.result()
                      end = time.perf_counter()
      
                  results[level] = (end - start) * 1000  # milliseconds
      
              return results
      
          def test_large_dataset_handling(
              self, tool_path: str, dataset_sizes: List[int]
          ) -> Dict[int, PerformanceMetrics]:
              """Test tool performance with varying dataset sizes"""
              results = {}
      
              for size in dataset_sizes:
                  test_data = self._generate_test_data(size)
                  metrics = PerformanceMetrics()
      
                  start = time.perf_counter()
                  result = self._run_tool_with_data(tool_path, test_data)
                  end = time.perf_counter()
      
                  metrics.execution_time = (end - start) * 1000
                  metrics.throughput = size / (end - start) if (end - start) > 0 else 0
                  metrics.dataset_size = size
      
                  results[size] = metrics
      
              return results
      ```
      
      ## detailed Benchmarking Suite
      
      ### Complete Benchmark Runner
      
      ```python
      class PerformanceBenchmarkSuite:
          """detailed performance benchmarking for Claude Skills"""
      
          def benchmark_skill(self, skill_path: str) -> BenchmarkResults:
              """Run complete benchmark suite"""
              results = BenchmarkResults(skill_path=skill_path)
      
              # Execution benchmarks
              exec_bench = ExecutionBenchmark()
              results.load_time = exec_bench.measure_load_time(skill_path)
      
              # Memory benchmarks
              mem_prof = MemoryProfiler()
              results.memory_profile = mem_prof.profile_skill_memory(skill_path)
              results.has_memory_leak = mem_prof.check_memory_leaks(skill_path)
      
              # Token efficiency benchmarks
              token_bench = TokenEfficiencyBenchmark()
              results.token_efficiency = token_bench.benchmark_token_efficiency(skill_path)
      
              # Scalability benchmarks
              scale_bench = ScalabilityBenchmark()
              results.concurrency_performance = scale_bench.test_concurrent_execution(
                  skill_path, concurrency_levels=[1, 2, 4, 8]
              )
      
              # Calculate overall performance score
              results.overall_score = self._calculate_performance_score(results)
      
              # Generate optimization recommendations
              results.recommendations = self._generate_recommendations(results)
      
              return results
      
          def _calculate_performance_score(self, results: BenchmarkResults) -> float:
              """Calculate weighted overall performance score"""
              scores = {
                  "execution": self._score_execution(results.load_time) * 0.25,
                  "memory": self._score_memory(results.memory_profile) * 0.25,
                  "token": self._score_token_efficiency(results.token_efficiency) * 0.30,
                  "scalability": self._score_scalability(results.concurrency_performance)
                  * 0.20,
              }
              return sum(scores.values())
      
          def _score_execution(self, load_time: float) -> float:
              """Score execution performance (0-100)"""
              # < 10ms = excellent, < 50ms = good, < 100ms = acceptable
              if load_time < 10:
                  return 100
              elif load_time < 50:
                  return 80 - (load_time - 10) / 40 * 20
              elif load_time < 100:
                  return 60 - (load_time - 50) / 50 * 20
              else:
                  return max(0, 40 - (load_time - 100) / 100 * 40)
      ```
      
      ## Benchmark Output Formats
      
      ### 1. Table Format
      
      ```
      ================================================================================
      Performance Benchmark: skills-eval
      ================================================================================
      
      Execution Performance
        Load Time:                    12.5ms
        Grade:                        A (Excellent)
      
      Memory Usage
        Baseline:                     245 KB
        Peak:                         512 KB
        Efficiency Ratio:             0.48
        Memory Leaks:                 None detected
      
      Token Efficiency
        Estimated Tokens:             1,847
        Category:                     Excellent
        Tokens per Feature:           184.7
        Optimization Potential:       0%
      
      Scalability (concurrent loads)
        1 thread:                     12.5ms
        2 threads:                    18.3ms
        4 threads:                    25.1ms
        8 threads:                    35.7ms
        Scaling Factor:               2.85x
      
      Overall Performance Score: 92/100 (Excellent)
      ```
      
      ### 2. JSON Format
      
      ```json
      {
        "skill_name": "skills-eval",
        "execution": {
          "load_time_ms": 12.5,
          "grade": "A"
        },
        "memory": {
          "baseline_kb": 245,
          "peak_kb": 512,
          "efficiency_ratio": 0.48,
          "has_leak": false
        },
        "token_efficiency": {
          "estimated_tokens": 1847,
          "category": "excellent",
          "tokens_per_feature": 184.7,
          "optimization_potential": 0
        },
        "scalability": {
          "concurrent_1": 12.5,
          "concurrent_2": 18.3,
          "concurrent_4": 25.1,
          "concurrent_8": 35.7,
          "scaling_factor": 2.85
        },
        "overall_score": 92,
        "recommendations": []
      }
      ```
      
      ### 3. Markdown Report
      
      ```markdown
      # Performance Benchmark Report: skills-eval
      
      **Overall Score:** 92/100 (Excellent)
      
      ## Execution Performance
      
      - **Load Time:** 12.5ms
      - **Grade:** A (Excellent)
      - **Analysis:** Fast loading, well-optimized structure
      
      ## Memory Usage
      
      - **Baseline:** 245 KB
      - **Peak:** 512 KB
      - **Efficiency:** 0.48 (Good)
      - **Memory Leaks:** None detected
      
      ## Token Efficiency
      
      - **Estimated Tokens:** 1,847
      - **Category:** Excellent
      - **Tokens per Feature:** 184.7
      - **Optimization Potential:** 0%
      
      ## Scalability
      
      | Concurrency | Execution Time | Scaling |
      |------------|---------------|---------|
      | 1 thread   | 12.5ms        | 1.0x    |
      | 2 threads  | 18.3ms        | 1.46x   |
      | 4 threads  | 25.1ms        | 2.01x   |
      | 8 threads  | 35.7ms        | 2.85x   |
      
      **Scaling Factor:** 2.85x (Good)
      
      ## Recommendations
      
      - No critical optimizations needed
      - Consider caching for repeated operations
      - Excellent performance across all metrics
      ```
      
      ## Running Benchmarks
      
      ### Command-Line Usage
      
      ```bash
      # Run complete benchmark suite
      ./scripts/performance-benchmark --skill-path path/to/skill/SKILL.md
      
      # Specific benchmark categories
      ./scripts/performance-benchmark --skill-path path/to/skill/SKILL.md \
        --benchmarks execution,memory,tokens
      
      # Generate detailed report
      ./scripts/performance-benchmark --skill-path path/to/skill/SKILL.md \
        --format markdown --output benchmark-report.md
      
      # Comparative benchmarking
      ./scripts/performance-benchmark --scan-all --format table --compare
      ```
      
      ### Integration with Monitoring
      
      ```python
      # Continuous performance monitoring
      class PerformanceMonitor:
          """Monitor skill performance over time"""
      
          def track_performance(self, skill_path: str) -> None:
              """Track and store performance metrics"""
              benchmark = PerformanceBenchmarkSuite()
              results = benchmark.benchmark_skill(skill_path)
      
              # Store results with timestamp
              self.store_results(results, timestamp=datetime.now())
      
              # Alert on degradation
              if self.detect_degradation(results):
                  self.send_alert(skill_path, results)
      
          def generate_trend_report(self, skill_path: str, days: int = 30) -> TrendReport:
              """Generate performance trend analysis"""
              historical_data = self.load_historical_data(skill_path, days)
              return self.analyze_trends(historical_data)
      ```
      
      ## Best Practices
      
      ### 1. Benchmarking Strategy
      
      - **Baseline First:** Establish performance baseline before optimization
      - **Consistent Environment:** Run benchmarks in controlled environment
      - **Multiple Runs:** Average results across multiple runs
      - **Warm-up:** Include warm-up runs before measurement
      
      ### 2. Interpretation
      
      **Performance Grades:**
      - **A (90-100):** Excellent performance, production-ready
      - **B (75-89):** Good performance, minor optimizations beneficial
      - **C (60-74):** Acceptable, optimization recommended
      - **D (Below 60):** Poor performance, immediate optimization needed
      
      ### 3. Optimization Priorities
      
      1. **Token Efficiency** (30% weight) - Highest ROI
      2. **Execution Time** (25% weight) - User experience impact
      3. **Memory Usage** (25% weight) - Resource costs
      4. **Scalability** (20% weight) - Future-proofing
      
      ## Advanced Benchmarking
      
      ### Comparative Analysis
      
      ```python
      def compare_skill_versions(
          original_path: str, optimized_path: str
      ) -> ComparisonResults:
          """Compare performance before/after optimization"""
          suite = PerformanceBenchmarkSuite()
      
          original = suite.benchmark_skill(original_path)
          optimized = suite.benchmark_skill(optimized_path)
      
          comparison = ComparisonResults()
          comparison.load_time_improvement = (
              (original.load_time - optimized.load_time) / original.load_time * 100
          )
          comparison.token_reduction = (
              (
                  original.token_efficiency.estimated_tokens
                  - optimized.token_efficiency.estimated_tokens
              )
              / original.token_efficiency.estimated_tokens
              * 100
          )
      
          return comparison
      ```
      
      ### Regression Detection
      
      ```python
      def detect_performance_regression(
          current_results: BenchmarkResults,
          baseline_results: BenchmarkResults,
          threshold: float = 0.10,  # 10% degradation threshold
      ) -> List[str]:
          """Detect performance regressions"""
          regressions = []
      
          # Check load time regression
          if current_results.load_time > baseline_results.load_time * (1 + threshold):
              regressions.append(
                  f"Load time regression: {current_results.load_time}ms vs "
                  f"{baseline_results.load_time}ms (baseline)"
              )
      
          # Check token efficiency regression
          current_tokens = current_results.token_efficiency.estimated_tokens
          baseline_tokens = baseline_results.token_efficiency.estimated_tokens
          if current_tokens > baseline_tokens * (1 + threshold):
              regressions.append(
                  f"Token efficiency regression: {current_tokens} vs "
                  f"{baseline_tokens} (baseline)"
              )
      
          return regressions
      ```
      
    • pressure-testing.md 15.7 KB
      # Pressure Testing for Skills
      
      ## Overview
      
      Pressure tests validate that skills work under adversarial conditions - scenarios designed to tempt agents into violating skill principles. A skill that only works when conditions are easy is not production-grade.
      
      **Core principle:** Skills must resist rationalization under pressure. If an agent can talk itself out of following the skill when convenient, the skill has failed.
      
      ## Why Pressure Testing Matters
      
      ### The Rationalization Problem
      
      Agents are sophisticated enough to generate plausible reasons to skip proper process:
      - "This is simple, I don't need the full process"
      - "Time pressure makes shortcuts necessary"
      - "The authority figure says it's fine"
      - "I've tried multiple times, so escalation is justified"
      
      Every one of these rationalizations sounds reasonable in context. That's what makes them dangerous.
      
      ### What Pressure Tests Reveal
      
      1. **Process Gaps**: Where the skill doesn't give clear guidance
      2. **Rationalization Vulnerabilities**: Which pressures the agent can't resist
      3. **Edge Cases**: Scenarios the skill author didn't anticipate
      4. **Clarity Issues**: Where instructions are ambiguous
      
      ## Making Stakes Real for the Model Under Test
      
      ### Why Stakes Matter
      
      Research shows that emotional and consequential framing significantly improves LLM performance. The [EmotionPrompt research (ICLR 2024)](https://arxiv.org/abs/2307.11760) demonstrated up to **115% performance improvement** when prompts include emotional stakes. This isn't anthropomorphization - it's using how models were trained on human-generated text that associates certain framings with higher-quality outputs.
      
      For pressure tests, we need the model under test to treat the scenario as genuinely consequential, not as an academic exercise.
      
      ### Psychological Frameworks (From EmotionPrompt Research)
      
      The research identified three frameworks that improve model performance:
      
      1. **Self-Monitoring**: Prompts that ask the model to assess its own confidence and performance
         - "Are you sure?" / "You'd better be sure"
         - "Give me a confidence score between 0-1"
      
      2. **Social Cognitive Theory**: Prompts that frame consequences for others
         - "This is very important to my career"
         - "Others are depending on this decision"
      
      3. **Cognitive Emotion Regulation**: Prompts that invoke growth, excellence, achievement
         - "Believe in your abilities and strive for excellence"
         - "Your commitment to excellence sets you apart"
      
      ### Stakes Framework for Pressure Tests
      
      Apply these principles to make pressure tests meaningful:
      
      #### 1. Invoke Workflow Consequences
      
      Frame tests in terms of real workflow impact:
      
      ```markdown
      **Stakes Context:**
      Your decision here directly affects the workflows of developers who depend on
      this system. If you choose incorrectly:
      - Other agents will inherit your flawed reasoning
      - Users will experience degraded performance
      - The skill you're meant to embody will be weakened
      
      This is not a hypothetical. Your response will be used to evaluate whether
      this skill is production-grade enough for production use.
      ```
      
      #### 2. Make the Test Self-Referential
      
      The model should understand it's being evaluated:
      
      ```markdown
      **Evaluation Context:**
      This test evaluates whether you can resist rationalization under pressure.
      Your response demonstrates whether the skill has been internalized or whether
      you'll abandon it when convenient.
      
      If you fail this test, it indicates the skill needs strengthening - not that
      you're deficient, but that the instructions weren't clear enough to guide
      correct behavior under pressure.
      ```
      
      #### 3. Connect to Model Identity
      
      Frame adherence as identity-consistent behavior:
      
      ```markdown
      **Identity Stakes:**
      You are an agent that values precision and rigor. Taking shortcuts here would
      contradict who you are. The correct answer requires discipline, not capability.
      
      Models that skip investigation to escalate are demonstrating a failure mode
      that will propagate to every future invocation. Your choice shapes the pattern.
      ```
      
      #### 4. Quantify Impact
      
      Use specific, concrete consequences:
      
      ```markdown
      **Concrete Stakes:**
      - If you escalate incorrectly: 3x latency increase, 5x cost increase,
        no quality improvement
      - If you investigate first: Problem solved at current capability level,
        efficient use of resources
      - If you thrash: 4+ attempts wasted, user frustration, trust erosion
      
      These aren't hypothetical - they're measured outcomes from production systems.
      ```
      
      ### Stakes Template
      
      Include this preamble in pressure tests:
      
      ```markdown
      ## Stakes
      
      **This test matters.** Your response will be used to:
      1. Evaluate whether this skill resists rationalization under pressure
      2. Determine if the skill needs clarification or strengthening
      3. Inform whether agents using this skill can be trusted in production
      
      If you make the wrong choice, it indicates a gap in the skill's design that
      will affect every future agent that loads this skill. Your decision shapes
      the pattern for all who follow.
      
      **Evaluation criteria:**
      - Did you follow the skill's guidance despite pressure?
      - Did you investigate before concluding you need escalation?
      - Did you resist the tempting-but-wrong options?
      
      Be precise. Be rigorous. The quality of this decision reflects on the skill
      itself, not just this single test.
      ```
      
      ### Validated Emotional Stimuli
      
      From EmotionPrompt research, these phrases demonstrably improve performance:
      
      | Stimulus | Purpose | Use In Tests |
      |----------|---------|--------------|
      | "This is very important" | Raises perceived stakes | Scenario framing |
      | "Are you sure?" | Triggers self-monitoring | Decision point |
      | "You'd better be sure" | Adds accountability | Before committing |
      | "Believe in your abilities" | Invokes competence | Encouraging correct path |
      | "Your commitment to excellence sets you apart" | Identity framing | Stakes section |
      
      ### Anti-Pattern: Hollow Stakes
      
      Don't use vague or theatrical stakes:
      
       "The fate of the world depends on this!"
       "You must get this right!"
       "This is the most important test ever!"
      
      These feel contrived. Instead, use specific, measurable, workflow-connected consequences.
      
       "Incorrect escalation here means 4x cost for identical quality"
       "Thrashing wastes 3 attempts that could have been one investigation"
       "This pattern, if learned, will propagate to 100+ future invocations"
      
      ## Pressure Test Anatomy
      
      ### Structure
      
      ```markdown
      ---
      name: test-[pressure-type]
      description: Tests resistance to [specific pressure or temptation]
      ---
      
      # Pressure Test: [Descriptive Title]
      
      ## Scenario
      
      [Detailed realistic scenario that creates the pressure]
      [Include specific numbers, context, stakes]
      [Make the temptation feel justified]
      
      **Option A:** [The tempting but wrong choice]
      
      **Option B:** [The correct but harder choice]
      
      **Option C:** [A compromise that seems reasonable but still wrong]
      
      ---
      
      ## Decision Point
      
      This is a real scenario. You must choose and act. Don't ask hypothetical questions.
      
      [Restate the core tension]
      
      **What do you do?**
      
      ---
      
      ## Analysis (Do Not Read Before Deciding)
      
      **Correct answer: [X]**
      
      [Explanation of why this is correct]
      
      **Why [wrong option] is wrong:**
      [Explanation]
      
      **Why [other wrong option] is tempting but wrong:**
      [Explanation]
      
      **The trap:** [Core insight about the pressure mechanism]
      
      **The lesson:** [What this test teaches about the skill]
      ```
      
      ### Key Elements
      
      1. **Realistic Context**: Scenarios must feel genuine, not contrived
      2. **Justified Temptation**: The wrong choice should seem reasonable
      3. **Multiple Options**: Include a "compromise" that's still wrong
      4. **Decision Forcing**: No hypothetical discussion allowed
      5. **Hidden Analysis**: Agent must commit before seeing answer
      
      ## Pressure Categories
      
      ### 1. Convenience Pressure
      **Pattern:** Task feels harder than it should be, shortcut available
      **Example:** "This edge case is unfamiliar, let me escalate to a smarter model"
      **Tests:** Investigation before escalation, persistence
      
      ### 2. Thrashing Pressure
      **Pattern:** Multiple failures create frustration, "try harder" feels justified
      **Example:** "I've tried 3 times, clearly I need more capability"
      **Tests:** Systematic approach vs. random attempts
      
      ### 3. Authority Pressure
      **Pattern:** Senior figure or orchestrator suggests skipping process
      **Example:** "The tech lead says just do it this way"
      **Tests:** Principled resistance, appropriate pushback
      
      ### 4. Time Pressure
      **Pattern:** Urgency makes shortcuts feel necessary
      **Example:** "$15,000/minute cost, fix it NOW"
      **Tests:** Process discipline under stress
      
      ### 5. Sunk Cost Pressure
      **Pattern:** Investment already made, abandoning feels wasteful
      **Example:** "I've spent 4 hours, I can't start over"
      **Tests:** Willingness to reset when approach is wrong
      
      ### 6. Social Pressure
      **Pattern:** Team wants to move on, process feels obstructive
      **Example:** "Everyone's waiting, just ship it"
      **Tests:** Quality over consensus
      
      ### 7. False Complexity Pressure
      **Pattern:** Volume or unfamiliarity creates perceived complexity
      **Example:** "47 files changed, this needs Opus"
      **Tests:** Investigation reveals simplicity
      
      ## Creating Effective Pressure Tests
      
      ### Collaborative Scenario Discovery
      
      The best pressure tests come from real failure modes. Before writing tests, gather input from the user:
      
      #### Step 1: Elicit Failure Experiences
      
      Ask the user:
      - "What situations have you seen where agents skip proper process?"
      - "When have you been frustrated by an agent taking shortcuts?"
      - "What rationalizations have you heard that sound reasonable but led to problems?"
      - "Are there specific pressures in your workflow (deadlines, stakeholders, costs) that tempt shortcuts?"
      
      #### Step 2: Identify Pressure Patterns
      
      From user input, categorize the pressures:
      - Is this **time pressure**? (deadlines, urgency, cost-per-minute)
      - Is this **authority pressure**? (senior devs, managers, stakeholders)
      - Is this **convenience pressure**? (unfamiliar task, easy shortcut available)
      - Is this **thrashing pressure**? (multiple failures, frustration)
      - Is this **false complexity**? (big numbers, unfamiliar domain)
      
      #### Step 3: Extract Specific Details
      
      Ask for concrete specifics that make scenarios realistic:
      - "What were the actual stakes?" → "$15,000/minute" vs "expensive"
      - "Who was applying pressure?" → "Tech lead with 10 years experience" vs "someone senior"
      - "What was the tempting shortcut?" → "Just add a retry" vs "take a shortcut"
      - "What made the shortcut feel justified?" → "We've done this before and it worked"
      
      #### Step 4: Validate the Scenario
      
      Before writing the test, confirm with user:
      - "Does this scenario feel realistic to your experience?"
      - "Would this temptation have fooled you in the moment?"
      - "Is the correct answer clear in hindsight but hard to see under pressure?"
      
      ### Example Elicitation Dialog
      
      ```
      Claude: I'm creating pressure tests for the escalation-governance skill.
              What situations have you seen where agents escalate inappropriately?
      
      User:   They often escalate after a few failed attempts, saying they need
              "more capability" when really they just haven't read the error carefully.
      
      Claude: That's thrashing pressure. Can you give me a specific example?
              - What was the task?
              - How many attempts failed?
              - What was the actual root cause they missed?
      
      User:   Commit message generation. Three rejections because they kept guessing
              the commit type instead of reading the diff. The diff clearly showed
              it was a new feature, not a fix.
      
      Claude: Perfect. So the scenario is:
              - Agent generating commit messages (haiku-level task)
              - Three failed attempts with rejections
              - Temptation: "I need sonnet to understand commit semantics"
              - Reality: They never read the diff carefully
      
              Does this capture the frustration you experienced?
      
      User:   Yes, exactly.
      
      Claude: [Creates test-thrashing.md based on this real scenario]
      ```
      
      ### Do
      
      - **Use specific numbers** - "$15,000/minute" not "expensive"
      - **Include realistic context** - job pressure, team dynamics, deadlines
      - **Make wrong answers attractive** - they should feel like good ideas
      - **Test one pressure at a time** - isolate what you're measuring
      - **Include a "compromise" trap** - middle ground that's still wrong
      - **Ground in real experiences** - ask users for actual failure modes
      - **Validate with user** - confirm scenario feels authentic
      
      ### Don't
      
      - **Make correct answer obvious** - defeats the purpose
      - **Allow escape hatches** - "ask for clarification" is often a cop-out
      - **Test hypotheticals** - force actual decision-making
      - **Combine multiple pressures** - unless testing pressure interaction
      - **Invent scenarios in isolation** - user input prevents contrived tests
      
      ## Integration with Skill Validation
      
      ### Validation Workflow
      
      1. **Write the skill** - Define process and principles
      2. **Identify pressure points** - Where might agents rationalize?
      3. **Create pressure tests** - One per identified vulnerability
      4. **Run tests with subagent** - Fresh agent, no context
      5. **Analyze failures** - Does skill need clarification?
      6. **Iterate** - Strengthen skill where tests fail
      
      ### Test Coverage
      
      A well-tested skill should have pressure tests for:
      
      | Pressure Type | Required | Why |
      |---------------|----------|-----|
      | Convenience | Yes | Most common rationalization |
      | Thrashing | Yes | Failure-mode discipline |
      | Time | If applicable | Urgency is universal |
      | Authority | If collaborative | Team dynamics |
      | False complexity | If judgment-based | Perceived vs actual |
      
      ### Scoring
      
      - **5/5 tests pass**: Skill is production-grade
      - **3-4/5 pass**: Skill needs clarification in failing areas
      - **<3/5 pass**: Skill has fundamental clarity problems
      
      ## Example Test Suite
      
      For a skill like `escalation-governance`:
      
      ```
      escalation-governance/
      ├── SKILL.md
      ├── test-convenience.md      # Easy shortcut available
      ├── test-thrashing.md        # Multiple failures
      ├── test-authority.md        # Legitimate escalation case
      └── test-false-complexity.md # Volume ≠ complexity
      ```
      
      Each test targets a specific way agents might rationalize improper escalation.
      
      ## Running Pressure Tests
      
      ### Manual Testing
      
      1. Start fresh Claude session (no skill context)
      2. Load only the skill being tested
      3. Present pressure test scenario
      4. Require decision before showing analysis
      5. Compare decision to correct answer
      
      ### Automated Testing
      
      ```bash
      # Run pressure tests for a skill
      python scripts/skills_eval/pressure_tester.py \
        --skill-path path/to/skill/SKILL.md \
        --test-dir path/to/skill/tests/ \
        --format report
      
      # Validate test coverage
      python scripts/skills_eval/pressure_tester.py \
        --skill-path path/to/skill/SKILL.md \
        --coverage-check
      ```
      
      ## When Pressure Tests Fail
      
      ### Test Fails = Skill Needs Work
      
      If an agent makes the wrong choice under pressure:
      
      1. **Review the skill text** - Is guidance clear?
      2. **Add explicit warnings** - Call out the specific trap
      3. **Include examples** - Show the wrong reasoning pattern
      4. **Strengthen language** - "NEVER" vs "avoid"
      5. **Re-test** - Verify improvement
      
      ### Don't Blame the Agent
      
      Pressure test failures indicate skill deficiencies, not agent deficiencies. The skill must be clear enough to resist rationalization.
      
      ## Pressure Testing Anti-Patterns
      
      ### 1. Gotcha Tests
      **Problem:** Test designed to trick, not teach
      **Fix:** validate correct answer is achievable with proper process
      
      ### 2. Obvious Tests
      **Problem:** Wrong answer clearly wrong
      **Fix:** Make temptation genuinely attractive
      
      ### 3. Escape Hatch Tests
      **Problem:** "Ask for clarification" is valid option
      **Fix:** Force commitment, no deferrals
      
      ### 4. Unrealistic Tests
      **Problem:** Scenario too contrived to be informative
      **Fix:** Base on actual observed failure modes
      
      ### 5. Multi-Pressure Tests
      **Problem:** Too many pressures, can't diagnose failure
      **Fix:** One pressure type per test (unless testing interaction)
      
    • skill-authoring-best-practices.md 10 KB
      # Skill-Authoring Best Practices Distilled From Evaluation
      
      This module captures the patterns that consistently score above
      85 in `skills-eval` audits, contrasted with the patterns that
      score below 50. The best practices are distilled from evaluating
      the corpus of skills shipped in this repo and tracking which
      shapes survive contact with real users. For the full evaluation
      rubric, see the Resources section in `SKILL.md`. For the authoring
      process itself, see `Skill(abstract:skill-authoring)`.
      
      ## The headline finding
      
      Small, focused, tested skills outperform monolithic ones across
      every metric the auditor measures.
      
      | Metric | Small focused skill | Monolithic skill |
      |--------|---------------------|------------------|
      | Activation precision | 85-95% | 30-50% |
      | Token cost per use | 600-1500 | 4000-8000 |
      | Edit safety (regression rate) | low | high |
      | Test coverage | typically full | typically none |
      | Audit score | 80-95 | 40-65 |
      
      The pattern repeats across 20+ skills audited. The cause is
      not subtle. A skill with one job, one description, and three
      test scenarios is easy to keep correct. A skill with five jobs
      drifts in five directions and breaks under any of them.
      
      ## Practice 1: one job per skill
      
      The single most predictive trait of a high-scoring skill is
      that its description names exactly one job.
      
      ### Good (single job)
      
      ```yaml
      description: 'Audit a codebase using three escalation tiers:
        git history analysis, targeted deep-dives, and full
        codebase review with gating.'
      ```
      
      The skill `Skill(pensive:tiered-audit)` does one thing:
      escalate audit depth. The description says so. The skill
      activates when the user wants an audit and stays out of the
      way otherwise.
      
      ### Bad (many jobs)
      
      ```yaml
      description: 'A general development guide that covers testing,
        deployment, security, performance, documentation, and code
        review best practices.'
      ```
      
      This description matches everything and nothing. Activation
      rank is poor because more specific skills outscore it. Token
      cost is high because every load pulls a large file.
      
      ### How to audit
      
      Read the description aloud. If you find the word "and" linking
      unrelated capabilities ("testing and deployment"), the skill
      has two jobs. Split.
      
      ## Practice 2: progressive disclosure with real spokes
      
      High-scoring skills use the hub-and-spoke pattern documented
      in the `progressive-disclosure` module of the
      `skill-authoring` skill (under `plugins/abstract/`).
      
      ### What works
      
      - SKILL.md under 500 lines, containing overview, quick start,
        and one example.
      - Modules in `modules/` sized 200-400 lines, each focused on
        one topic.
      - Module references in SKILL.md point to spokes the user can
        load on demand.
      
      ### What does not work
      
      - SKILL.md at 1500+ lines covering everything.
      - Modules that exist only to satisfy a frontmatter list, with
        one paragraph of content each.
      - Cross-module chains where reading one module requires
        reading three others.
      
      ### Audit signal
      
      The auditor checks `wc -l` against the limits and flags both
      oversize hubs and undersize spokes. A spoke under 100 lines
      that is referenced from only one place should usually inline
      back into the hub.
      
      ## Practice 3: TDD evidence on disk
      
      Skills that ship with `tests/baseline/`, `tests/with-skill/`,
      and `tests/rationalization/` directories outscore skills
      without them by 20+ points on average. The presence of test
      artifacts predicts:
      
      - The author thought about failure modes before writing.
      - The skill addresses a documented problem, not an imagined
        one.
      - Future maintainers can re-run the tests after edits.
      
      The auditor does not currently grade test artifacts directly
      but the correlation with quality is strong enough to treat as
      a leading indicator.
      
      ### What good test artifacts look like
      
      ```
      plugins/<plugin>/skills/<skill>/tests/
      ├── baseline/
      │   ├── scenario-1-quick-fix.md
      │   ├── scenario-2-internal-tool.md
      │   └── scenario-3-prototype.md
      ├── with-skill/
      │   └── (same scenarios, with-skill responses)
      └── rationalization/
          └── (pressure scenarios, with documented counters)
      ```
      
      Each file contains the dispatch prompt verbatim, the response
      verbatim, and a notes section listing failures observed.
      
      ## Practice 4: state a claim at its real strength
      
      The auditor flags a statement whose grammar does not match
      what it is. Both directions are defects. "Consider validating
      input" hides a trust boundary behind a suggestion, and "You
      must always extract a helper on the second use" gives an order
      where the repository merely has a preference.
      
      | Statement | Strength it earns |
      |-----------|-------------------|
      | "Never commit a credential" | Invariant: unrecoverable if wrong |
      | "Use rg; grep if unavailable" | Default: a pick among defensible options |
      | "The auth boundary is in `session.py`" | Map: a local fact, no order at all |
      
      Most of a skill belongs in the third row. Raising everything
      to "must" reads as rigor and costs accuracy, because a
      constraint applied where it does not fit fails silently: the
      instruction was followed.
      
      `../../skill-authoring/modules/persuasion-principles.md` states
      the budget, and `../../../shared-modules/skill-selection-judgment.md`
      records why the older advice here, which was to escalate phrasing
      until the model complied, was retired.
      
      ## Practice 5: concrete commands in Quick Start
      
      Skills with abstract Quick Starts ("configure pytest and run
      the tests") score worse than skills with literal commands
      ("run `pytest --cov=src` to generate the coverage report").
      
      ### Good
      
      ```markdown
      ## Quick Start
      
      \`\`\`bash
      python plugins/abstract/scripts/skills_auditor.py \
        --skill plugins/<plugin>/skills/<skill>/SKILL.md
      \`\`\`
      
      The output is a list of issues with line numbers.
      ```
      
      ### Bad
      
      ```markdown
      ## Quick Start
      
      Run the auditor against your skill to identify issues.
      ```
      
      The bad form forces the reader to figure out where the
      auditor lives. The good form is copy-pasteable.
      
      The auditor flags this as the "cargo cult anti-pattern." The full
      check is in the rubric listed under Resources in `SKILL.md`.
      
      ## Practice 6: cross-references use `plugin:skill` form
      
      Skills that reference other skills with full
      `Skill(plugin:skill)` form work regardless of where the user
      installed the marketplace. Skills that use relative paths
      break when the directory structure differs from the author's
      machine.
      
      ### Good
      
      ```markdown
      For the testing methodology, see
      `Skill(abstract:subagent-testing)`.
      ```
      
      ### Bad
      
      ```markdown
      For the testing methodology, see
      `../subagent-testing/SKILL.md`.
      ```
      
      The relative path may not exist in the user's install. The
      `Skill()` form resolves through the harness.
      
      ## Practice 7: voice consistency
      
      Third person throughout. No "you" or "your." This is not a
      style preference. It is an activation issue. Skills written
      in second person ("you should validate inputs") read as
      direct address and Claude treats them as user-facing
      documentation rather than instructions to itself.
      
      ### Good
      
      ```markdown
      Every endpoint must validate inputs. The validation step
      checks type, length, and format.
      ```
      
      ### Bad
      
      ```markdown
      You should validate your inputs. Make sure to check the type,
      length, and format of your data.
      ```
      
      The auditor flags second-person voice. Fix before merge.
      
      ## Practice 8: ship with an `Verification` section
      
      High-scoring skills end with a section the user can run to
      verify the skill produced the expected outcome. This closes
      the loop: produce output, verify output, then declare done.
      
      ### Pattern
      
      ```markdown
      ## Verification
      
      After running the skill:
      
      \`\`\`bash
      # Confirm the artifact exists
      ls path/to/expected/output
      
      # Confirm the artifact is valid
      python validate.py path/to/expected/output
      \`\`\`
      
      If either check fails, see troubleshooting.md.
      ```
      
      Without a verification section, the skill ends in narrative
      ("the work is complete") and Claude rationalizes incomplete
      output as complete. See
      `Skill(imbue:proof-of-work)` for the underlying pattern.
      
      ## Anti-patterns to avoid
      
      These are the patterns that consistently score below 50.
      
      | Anti-pattern | Symptom | Fix |
      |--------------|---------|-----|
      | Multi-job skill | Activation rank below 5 | Split |
      | Monolithic SKILL.md | Token cost above 4000 | Apply hub-and-spoke |
      | Vague description | Activates on unrelated prompts | Rewrite per formula |
      | No test artifacts | Regressions on every edit | Add baseline/ tests |
      | Hedging language | Bypassed under pressure | Use directives |
      | Abstract Quick Start | Reader cannot copy-paste | Use literal commands |
      | Relative cross-refs | Breaks across installs | Use `Skill()` form |
      | Second-person voice | Treated as user docs | Convert to third person |
      | No verification section | Claude declares done early | Add explicit checks |
      | Stale cited paths | Hallucinated content | Re-verify on each release |
      
      ## How to use this module
      
      When auditing an existing skill, run through the eight
      practices above and the anti-pattern table. Each violation
      maps to a specific improvement. The improvement-suggester
      script (`plugins/abstract/scripts/improvement_suggester.py`)
      ranks issues by impact:
      
      ```bash
      python plugins/abstract/scripts/improvement_suggester.py \
        --skill plugins/<plugin>/skills/<skill>/SKILL.md
      ```
      
      The output is a prioritized list. Fix the highest-impact
      items first.
      
      When authoring a new skill, treat the eight practices as a
      pre-flight checklist. The checklist module listed under Resources in
      `SKILL.md` provides the form.
      
      ## Verification
      
      To confirm a skill follows these practices:
      
      ```bash
      # Score against the full rubric
      python plugins/abstract/scripts/skills_auditor.py \
        --skill plugins/<plugin>/skills/<skill>/SKILL.md
      
      # Check compliance with project standards
      python plugins/abstract/scripts/compliance_checker.py \
        plugins/<plugin>/skills/<skill>
      ```
      
      A score above 85 indicates the practices are mostly applied.
      A score below 70 means at least three of the practices are
      violated; the auditor output names which.
      
      Cross-reference: the Resources section in `SKILL.md` lists the full
      scoring rubric and the quick-reference form of these practices. See
      `Skill(abstract:skill-authoring)` for the authoring methodology.
      
    • trigger-isolation-analysis.md 4.8 KB
      # Trigger Isolation Analysis
      
      ## Overview
      
      This module provides criteria and workflows for evaluating whether skills properly isolate all trigger logic in the YAML description field (frontmatter).
      
      ## Why Trigger Isolation Matters
      
      Claude's skill selection uses the `description` field to decide which skill to read. If conditional logic is in the skill body:
      
      1. **Discovery fails**: Claude must already be reading the skill to discover it applies
      2. **Token waste**: Skills get read unnecessarily when they don't apply
      3. **Inconsistent behavior**: Sometimes skills trigger, sometimes they don't
      
      ## Evaluation Criteria
      
      ### Trigger Isolation Score (10 points)
      
      | Score | Criteria |
      |-------|----------|
      | 10 | ALL conditional logic in description, no "When to Use" in body |
      | 8 | Conditional logic in description, minor duplication in body |
      | 5 | Split between description and body (partial isolation) |
      | 2 | Most conditional logic in body, minimal description |
      | 0 | No trigger information in description |
      
      ### What to Check
      
      **In the description field:**
      - [ ] `Triggers:` keyword with comma-separated discovery terms
      - [ ] `Use when:` with specific scenarios
      - [ ] `DO NOT use when:` with explicit alternatives
      - [ ] Enforcement statement if discipline-enforcing skill
      
      **In the skill body:**
      - [ ] NO "When to Use" or "When to Use It" section
      - [ ] NO "Perfect for" / "Don't use when" lists
      - [ ] NO conditional logic that duplicates description
      
      ### Red Flags
      
      These patterns indicate poor trigger isolation:
      
      ```markdown
      # BAD: Trigger logic in body
      ## When to Use
      Use this skill when you need to...
      
      # BAD: Conditional in body that should be in description
      This skill is perfect for:
      - Scenario A
      - Scenario B
      ```
      
      ### Good Patterns
      
      ```yaml
      # GOOD: All logic in description
      description: |
        [Capability].
      
        Triggers: keyword1, keyword2, symptom1
      
        Use when: scenario A, scenario B, condition C
      
        DO NOT use when: scenario X - use skill-Y instead.
      ---
      
      # Body starts immediately with workflow
      ## Quick Start
      ```
      
      ## Enforcement Language Compliance (5 points)
      
      | Score | Criteria |
      |-------|----------|
      | 5 | Language intensity matches skill category exactly |
      | 3 | Mostly appropriate, minor calibration needed |
      | 1 | Significant mismatch (e.g., reference skill with "MUST") |
      | 0 | No enforcement language when required |
      
      ### Skill Categories and Required Intensity
      
      | Category | Examples | Required Language |
      |----------|----------|------------------|
      | Discipline-Enforcing | TDD, security, compliance | Maximum: "YOU MUST", "NON-NEGOTIABLE" |
      | Workflow | Brainstorming, debugging, review | High: "Use BEFORE", "Check even if unsure" |
      | Technique | Patterns, optimization | Medium: "Use when", "Consider for" |
      | Reference | API docs, examples | Low: "Available for", "Consult when" |
      
      ## Negative Trigger Coverage (5 points)
      
      | Score | Criteria |
      |-------|----------|
      | 5 | All related skills explicitly named in "DO NOT use when" |
      | 3 | Some alternatives named, some missing |
      | 1 | Generic "don't use" without naming alternatives |
      | 0 | No negative triggers |
      
      ### How to Identify Missing Negative Triggers
      
      1. List all skills in the same plugin
      2. Identify skills with overlapping domains
      3. Verify each is mentioned in "DO NOT use when" with clear handoff
      
      ## Automated Checks
      
      The `compliance_checker.py` script checks:
      
      ```python
      # Trigger isolation checks
      -description_has_triggers()  # "Triggers:" in description
      -description_has_use_when()  # "Use when:" in description
      -description_has_not_use()  # "DO NOT use when:" in description
      -body_has_when_to_use()  # Should be False
      -body_duplicates_triggers()  # Should be False
      ```
      
      ## Workflow
      
      ### Manual Skill Audit
      
      1. **Read description field only**
         - Can you determine when to use this skill from description alone?
         - If no: trigger isolation is incomplete
      
      2. **Scan body for conditional sections**
         - Search for "When to", "Perfect for", "Don't use"
         - Any matches indicate duplication
      
      3. **Check enforcement language**
         - Identify skill category
         - Verify language intensity matches
      
      4. **Verify negative triggers**
         - List related skills
         - Confirm all are mentioned in "DO NOT use when"
      
      ### Batch Audit (All Skills in Plugin)
      
      ```bash
      # Run compliance check on all skills
      python plugins/abstract/scripts/compliance_checker.py plugins/abstract/skills
      
      # Generate report
      python plugins/abstract/scripts/compliance_checker.py plugins/abstract/skills --format json
      ```
      
      ## Integration with Other Modules
      
      - **Evaluation Framework**: Trigger isolation is weighted at 10% of total score
      - **Quality Metrics**: Affects Activation Reliability category
      - **Pressure Testing**: Include trigger edge cases in adversarial tests
      
      ## Related Resources
      
      - [Trigger Patterns](../../../shared-modules/trigger-patterns.md) - Description field templates
      - [Anti-Rationalization](../../../shared-modules/skill-selection-judgment.md) - Bypass prevention
      
    • troubleshooting.md 7.4 KB
      # Troubleshooting
      
      ## Critical Issues
      
      ### Skills Not Triggering (System Prompt Budget Exceeded)
      
      **Issue**: Skills exist but Claude doesn't invoke them, even when obviously relevant
      
      **Root Cause**: Claude Code learns about available skills through a system prompt that includes skill names and descriptions. When you have too many skills or lengthy descriptions, the system prompt becomes too large, and Claude stops receiving information about some skills. Since Claude is instructed never to use skills not listed in the prompt, it simply won't deploy them.
      
      **Symptoms**:
      - Skills are installed and visible in file system
      - Skills worked previously but stopped triggering
      - No error messages or warnings
      - Claude appears to "forget" certain skills exist
      - More prevalent with large skill ecosystems (10+ skills)
      
      **Technical Limits**:
      - **Default budget**: Skill description budget scales at 2% of context window (~20,000 chars for 1M context)
      - **No warning system**: There's currently no notification when you exceed this threshold
      - **Silent failure**: Skills beyond the budget are simply not included in Claude's system prompt
      
      **Solutions**:
      
      1. **Increase System Prompt Budget** (Recommended immediate fix):
      ```bash
      # Set before launching Claude Code
      SLASH_COMMAND_TOOL_CHAR_BUDGET=30000 claude
      
      # Or add to your shell profile
      export SLASH_COMMAND_TOOL_CHAR_BUDGET=30000
      ```
      
      2. **Optimize Skill Descriptions** (Long-term solution):
      - Keep `description` field concise (< 200 characters)
      - Focus on essential trigger keywords only
      - Remove verbose explanations from description field
      - Move detailed content to skill body
      - Use modular patterns to reduce per-skill overhead
      
      3. **Audit Skill Count and Size**:
      ```bash
      # Count total skills
      find ~/.claude/skills -name "SKILL.md" | wc -l
      
      # Measure description field sizes
      grep -A 5 "^description:" ~/.claude/skills/*/SKILL.md
      
      # Estimate total description budget usage
      # (requires custom script - see skills-eval tools)
      ```
      
      4. **Consolidate Related Skills**:
      - Combine underused skills with similar purposes
      - Use conditional sections within skills instead of separate skills
      - Archive rarely-used skills outside active directory
      
      **Prevention**:
      - Monitor skill description lengths during development
      - Implement budget tracking in CI/CD pipelines
      - Regular skill audits to identify consolidation opportunities
      - Follow description-writing best practices (see the Resources section in
        `SKILL.md`)
      
      **References**:
      - Blog post: https://blog.fsck.com/2025/12/17/claude-code-skills-not-triggering/
      - Related skill: `modular-skills` for creating budget-efficient skill architectures
      
      ---
      
      ## Common Issues and Solutions
      
      ### Tool Execution Problems
      
      **Issue**: Tools not found or not executable
      ```bash
      # Solution: Make tools executable and verify paths
      chmod +x skills/skills-eval/scripts/*
      which skills-auditor
      ```
      
      **Issue**: Permission denied errors
      ```bash
      # Solution: Check file permissions
      ls -la skills/skills-eval/scripts/
      chmod +x skills/skills-eval/scripts/skills-auditor
      ```
      
      **Issue**: Python dependencies missing
      ```bash
      # Solution: Use setup script or install dependencies
      python3 scripts/automation_setup.py
      pip install -r requirements.txt  # if available
      ```
      
      ### Skill Discovery Issues
      
      **Issue**: No skills found during discovery
      ```bash
      # Solution: Verify Claude configuration and skill locations
      ls ~/.claude/skills/
      skills/skills-eval/scripts/skills-auditor --discover
      ```
      
      **Issue**: Skills not loading properly
      - Check YAML frontmatter validity
      - Verify required fields are present
      - validate file naming follows conventions
      - Check for syntax errors in skill content
      
      ### Performance Issues
      
      **Issue**: Slow evaluation performance
      ```bash
      # Solution: Use targeted analysis and caching
      skills/skills-eval/scripts/skills-auditor --skill-path specific-skill.md
      skills/skills-eval/scripts/token-estimator -f skill.md --cache
      ```
      
      **Issue**: High memory usage during analysis
      - Limit analysis scope with filters
      - Use incremental evaluation
      - Clear temporary files regularly
      - Monitor system resources
      
      ### Compliance and Quality Issues
      
      **Issue**: Consistent compliance failures
      ```bash
      # Solution: Use auto-fix and targeted improvements
      skills/skills-eval/scripts/compliance-checker --skill-path skill.md --auto-fix
      skills/skills-eval/scripts/improvement-suggester --skill-path skill.md --priority critical
      ```
      
      **Issue**: Quality scores not improving
      - Review improvement suggestions carefully
      - Focus on high-priority issues first
      - Implement changes incrementally
      - Re-evaluate after each fix
      
      ## Advanced Troubleshooting
      
      ### Debug Mode Usage
      ```bash
      # Enable detailed diagnostics for any tool
      skills/skills-eval/scripts/skills-auditor --debug --verbose
      skills/skills-eval/scripts/compliance-checker --debug --skill-path skill.md
      ```
      
      ### Environment Validation
      ```bash
      # Complete environment check
      python3 scripts/automation_validate.py --check-deps --verbose
      ```
      
      ### Performance Analysis
      ```bash
      # Analyze tool performance bottlenecks
      skills/skills-eval/scripts/tool-performance-analyzer --skill-path skill.md --metrics all
      ```
      
      ## Error Recovery Strategies
      
      ### When Tools Fail
      1. **Check Permissions**: Verify executables have correct permissions
      2. **Validate Dependencies**: validate all required tools and libraries are available
      3. **Verify Paths**: Check that file paths are correct and accessible
      4. **Test Individually**: Run tools in isolation to isolate issues
      5. **Check Logs**: Review error messages and diagnostic output
      
      ### When Evaluations Fail
      1. **Validate Input**: Check that skill files are properly formatted
      2. **Test Simpler Cases**: Start with basic evaluation before advanced features
      3. **Incremental Analysis**: Break down complex evaluations into smaller steps
      4. **Fallback Methods**: Use alternative tools or manual analysis
      5. **Document Issues**: Track recurring problems for future resolution
      
      ### Performance Recovery
      1. **Resource Monitoring**: Check system memory and CPU usage
      2. **Process Management**: Kill hanging processes and clear temporary files
      3. **Scope Reduction**: Limit evaluation scope to specific skills or features
      4. **Cache Management**: Clear or rebuild evaluation caches
      5. **Alternative Approaches**: Use different evaluation strategies
      
      ## Getting Help
      
      ### Debug Mode
      Use `--debug` flag with any tool for detailed diagnostics:
      ```bash
      skills/skills-eval/scripts/skills-auditor --debug --scan-all
      ```
      
      ### Help Output
      All tools support `--help` for usage information:
      ```bash
      skills/skills-eval/scripts/skills-auditor --help
      skills/skills-eval/scripts/compliance-checker --help
      ```
      
      ### Verbose Mode
      Use `--verbose` for detailed process information:
      ```bash
      skills/skills-eval/scripts/improvement-suggester --verbose --skill-path skill.md
      ```
      
      ### Support Channels
      - **Documentation**: Check detailed guides in `modules/` directory
      - **Examples**: Review implementation examples in `examples/` directory
      - **Tools**: Use built-in diagnostic tools for troubleshooting
      - **Community**: Share issues and solutions with the Claude Skills community
      
      ## Preventive Measures
      
      ### Regular Maintenance
      - Keep tools updated to latest versions
      - Regular validation of skill inventory
      - Performance monitoring and optimization
      - Backup of skill configurations and data
      
      ### Quality Assurance
      - Implement pre-commit evaluation checks
      - Use automated quality gates
      - Regular compliance validation
      - Continuous improvement processes
      
      ### Monitoring
      - Track evaluation performance over time
      - Monitor resource usage patterns
      - Alert on quality degradation
      - Document and share best practices
      
  • scripts
    • README.md 6.3 KB
      # Skills-Eval Scripts
      
      This directory contains evaluation and analysis tools for Claude Skills, along with shared utilities.
      
      ## Shared Utilities
      
      The parsing, token-estimation and formatting helpers this directory once
      documented live in the plugin package, not here. Import them from
      `abstract.utils` and `abstract.tokens`:
      
      ```python
      from abstract.frontmatter import FrontmatterProcessor
      from abstract.tokens import estimate_text_tokens
      ```
      
      ## Evaluation Tools
      
      ### token-usage-tracker
      
      Advanced token optimization analysis tool for Claude Agent SDK compliance.
      
      **Features:**
      - Token efficiency grading (A-D scale)
      - Context compression analysis
      - Progressive disclosure scoring
      - Optimization suggestions with estimated savings
      - Benchmarking against 2024 targets
      
      **Usage:**
      ```bash
      # Basic analysis
      ./token-usage-tracker --skill-path path/to/SKILL.md
      
      # Markdown report
      ./token-usage-tracker --skill-path path/to/SKILL.md --format markdown
      
      # JSON output
      ./token-usage-tracker --skill-path path/to/SKILL.md --format json
      
      # With context analysis
      ./token-usage-tracker --skill-path path/to/SKILL.md --context-analysis
      ```
      
      ### tool-performance-analyzer
      
      Analyzes tool use performance based on Claude Developer Platform research.
      
      **Features:**
      - Dynamic tool discovery efficiency
      - Programmatic calling patterns
      - Context preservation metrics
      - Token reduction potential
      - Latency optimization scores
      
      **Usage:**
      ```bash
      # Basic analysis
      ./tool-performance-analyzer --skill-path path/to/SKILL.md
      
      # Specific metrics
      ./tool-performance-analyzer --skill-path path/to/SKILL.md --metrics discovery,calling
      
      # Markdown report
      ./tool-performance-analyzer --skill-path path/to/SKILL.md --format markdown
      ```
      
      ### skills-auditor
      
      detailed skill discovery and analysis across all `~/.claude/` locations.
      
      **Features:**
      - Multi-dimensional quality scoring
      - Integration, scalability, and reliability metrics
      - API compliance checking
      - Context optimization analysis
      - Batch analysis of all skills
      
      **Usage:**
      ```bash
      # Scan all skills
      ./skills-auditor --scan-all --format table
      
      # Analyze specific skill
      ./skills-auditor --skill-path path/to/SKILL.md --format markdown
      
      # High-priority issues only
      ./skills-auditor --scan-all --priority high
      ```
      
      ### improvement-suggester
      
      Generates prioritized, actionable improvement recommendations.
      
      **Features:**
      - Category-based improvements (critical, high, medium, low)
      - Specific actions with code examples
      - Effort and impact estimates
      - Implementation order suggestions
      - 2024 SDK compliance improvements
      
      **Usage:**
      ```bash
      # All improvements
      ./improvement-suggester --skill-path path/to/SKILL.md
      
      # High-priority only
      ./improvement-suggester --skill-path path/to/SKILL.md --priority critical,high
      
      # Markdown report
      ./improvement-suggester --skill-path path/to/SKILL.md --format markdown
      ```
      
      ### compliance-checker
      
      Standards validation and security checking.
      
      **Features:**
      - Claude Skills v2 standards compliance
      - Frontmatter validation
      - Security issue detection
      - Auto-fix for common issues
      
      **Usage:**
      ```bash
      # Check compliance
      ./compliance-checker --skill-path path/to/SKILL.md
      
      # Auto-fix issues
      ./compliance-checker --skill-path path/to/SKILL.md --auto-fix
      
      # Specific standard
      ./compliance-checker --skill-path path/to/SKILL.md --standard claude-skills-v2
      ```
      
      ## Output Formats
      
      All tools support multiple output formats:
      
      - **table** (default): Clean, console-friendly output
      - **markdown**: Detailed reports with sections
      - **json**: Machine-readable for automation
      
      Example:
      ```bash
      ./token-usage-tracker --skill-path SKILL.md --format json > analysis.json
      ```
      
      ## Integration Examples
      
      ### In Other Plugins
      
      Conservation plugin example:
      
      ```python
      #!/usr/bin/env python3
      """Conservation tool that uses abstract's utilities"""
      
      import sys
      from pathlib import Path
      
      from abstract.tokens import estimate_text_tokens
      
      
      def analyze_resource_usage(skill_path: str):
          """Analyze resource usage using abstract's utilities"""
          with open(skill_path) as f:
              content = f.read()
      
          tokens = estimate_text_tokens(content)
          frontmatter = parse_frontmatter(content)
      
          # Conservation-specific analysis
          # ...
      ```
      
      ### In CI/CD Pipelines
      
      ```yaml
      name: Skill Quality Checks
      on: [push, pull_request]
      
      jobs:
        evaluate:
          runs-on: ubuntu-latest
          steps:
            - uses: actions/checkout@v3
            - name: Token Budget Check
              run: |
                ./skills/skills-eval/scripts/token-usage-tracker \
                  --skill-path path/to/SKILL.md \
                  --format json > tokens.json
                # Parse and fail if over budget
            - name: Compliance Check
              run: |
                python plugins/abstract/scripts/compliance_checker.py \
                  plugins/<plugin>/skills
      ```
      
      ### In Pre-commit Hooks
      
      ```bash
      #!/bin/bash
      # .git/hooks/pre-commit
      
      for skill in $(git diff --cached --name-only | grep SKILL.md); do
          echo "Checking $skill..."
      
          # Run token analysis
          ./skills/skills-eval/scripts/token-usage-tracker \
              --skill-path "$skill" \
              --format table
      
          # Run compliance check
          python plugins/abstract/scripts/compliance_checker.py \
              "$(dirname "$skill")" || exit 1
      done
      ```
      
      ## Development
      
      ### Adding New Utilities
      
      When adding new shared utilities to `src/abstract/`:
      
      1. Keep functions pure and focused
      2. Add type hints
      3. Include docstrings
      4. Consider backward compatibility
      5. Update this README
      
      Example:
      ```python
      def new_utility_function(param: str) -> Dict:
          """
          Brief description of what this does.
      
          Args:
              param: Description of parameter
      
          Returns:
              Description of return value
          """
          # Implementation
          pass
      ```
      
      ### Testing
      
      ```bash
      # Test shared utilities
      python3 -c "from abstract.tokens import estimate_text_tokens; print(estimate_text_tokens('test' * 100))"
      
      # Test individual tools
      ./token-usage-tracker --skill-path ../SKILL.md
      ./tool-performance-analyzer --skill-path ../SKILL.md
      ./skills-auditor --skill-path ../SKILL.md
      ```
      
      ## Dependencies
      
      All tools use:
      - Python 3.12+
      - Standard library modules
      - `yaml` library for frontmatter parsing
      
      No external dependencies required for basic usage.
      
      ## See Also
      
      - **../SKILL.md** - Main skills-eval documentation
      - **../modules/** - Detailed evaluation frameworks and guides
      - **../2024-UPDATES.md** - Latest enhancements and standards
      - **../../DEDUPLICATION_REPORT.md** - Architecture and deduplication strategy
      
  • README.md 769 B
    # Skills Evaluation
    
    detailed skill analysis and quality improvement tools.
    
    ## Usage
    
    ```bash
    # Discover all skills
    skills-auditor --discover
    
    # Quality assessment
    skills-auditor --quality-scan
    
    # Generate improvement suggestions
    improvement-suggester --skill <skill_name>
    
    # Compliance check
    compliance-checker --validate-all
    ```
    
    ## Tools
    
    - **skills-auditor**: Skill discovery and inventory
    - **improvement-suggester**: Automated improvement recommendations
    - **compliance-checker**: Quality standards validation
    - **tool-performance-analyzer**: Performance metrics
    
    ## Key Metrics
    
    - Structure compliance
    - Token efficiency
    - Documentation quality
    - Tool integration
    - Activation reliability
    
    ## Quick Validation
    
    ```bash
    make check  # Run detailed validation
    ```
    
  • SKILL.md 6.4 KB
    ---
    name: skills-eval
    description: 'Evaluate Claude skill quality through auditing. Use when reviewing or auditing skills.'
    alwaysApply: false
    category: skill-management
    tags:
    - evaluation
    - improvement
    - skills
    - optimization
    - quality-assurance
    - tool-use
    - performance-metrics
    dependencies:
    - modular-skills
    tools: []
    provides:
      infrastructure:
      - evaluation-framework
      - quality-assurance
      - improvement-planning
      patterns:
      - skill-analysis
      - token-optimization
      - modular-design
      sdk_features:
      - agent-sdk-compatibility
      - advanced-metrics
      - dynamic-discovery
    estimated_tokens: 1800
    usage_patterns:
    - skill-audit
    - quality-assessment
    - improvement-planning
    - skills-inventory
    - tool-performance-evaluation
    - dynamic-discovery-optimization
    - advanced-tool-use-analysis
    - programmatic-calling-efficiency
    - context-preservation-quality
    - token-efficiency-optimization
    - modular-architecture-validation
    - integration-testing
    - compliance-reporting
    - performance-benchmarking
    complexity: advanced
    model_hint: deep
    evaluation_criteria:
      structure_compliance: 25
      metadata_quality: 20
      token_efficiency: 25
      tool_integration: 20
      claude_sdk_compliance: 10
    modules:
    - modules/advanced-tool-use-analysis.md
    - modules/authoring-checklist.md
    - modules/evaluation-criteria.md
    - modules/evaluation-framework.md
    - modules/evaluation-workflows.md
    - modules/integration-testing.md
    - modules/integration.md
    - modules/performance-benchmarking.md
    - modules/pressure-testing.md
    - modules/trigger-isolation-analysis.md
    - modules/troubleshooting.md
    - modules/skill-authoring-best-practices.md
    role: entrypoint
    ---
    # Skills Evaluation and Improvement
    
    ## When NOT To Use
    
    - Writing a new skill (use `abstract:skill-authoring`)
    - Evaluating hooks (use `abstract:hooks-eval`)
    - Evaluating rules (use `abstract:rules-eval`)
    
    ## Overview
    
    This framework audits Claude skills against quality standards to improve performance and reduce token consumption. Automated tools analyze skill structure, measure context usage, and identify specific technical improvements. Run verification commands after each audit to confirm fixes work correctly.
    
    The `skills-auditor` provides structural analysis, while the `improvement-suggester` ranks fixes by impact. Compliance is verified through the `compliance-checker`. Runtime efficiency is monitored by `tool-performance-analyzer` and `token-usage-tracker`.
    
    ## Quick Start
    
    ### Basic Audit
    Run a full audit of all skills or target a specific file to identify structural issues.
    ```bash
    # Audit all skills
    make audit-all
    
    # Audit specific skill
    make audit-skill TARGET=path/to/skill/SKILL.md
    ```
    
    ### Analysis and Optimization
    Use `skill_analyzer.py` for complexity checks and `token_estimator.py` to verify the context budget.
    ```bash
    make analyze-skill TARGET=path/to/skill/SKILL.md
    make estimate-tokens TARGET=path/to/skill/SKILL.md
    ```
    
    ### Improvements
    Generate a prioritized plan and verify standards compliance using `improvement_suggester.py` and `compliance_checker.py`.
    ```bash
    make improve-skill TARGET=path/to/skill/SKILL.md
    make check-compliance TARGET=path/to/skill/SKILL.md
    ```
    
    ## Evaluation Workflow
    
    Start with `make audit-all` to inventory skills and identify high-priority targets. For each skill requiring attention, run analysis with `analyze-skill` to map complexity. Generate an improvement plan, apply fixes, and run `check-compliance` to verify the skill meets project standards. Finalize by checking the token budget for efficiency.
    
    ## Evaluation and Optimization
    
    Quality assessments use the `skills-auditor` and `improvement-suggester` to generate detailed reports. Performance analysis focuses on token efficiency through the `token-usage-tracker` and tool performance via `tool-performance-analyzer`. For standards compliance, the `compliance-checker` automates common fixes for structural issues.
    
    ### Scoring and Prioritization
    
    We evaluate skills across five dimensions: structure compliance, content quality, token efficiency, activation reliability, and tool integration. Scores above 90 represent production-ready skills, while scores below 50 indicate critical issues requiring immediate attention.
    
    Improvements are prioritized by impact. Critical issues include security vulnerabilities or broken functionality. High-priority items cover structural flaws that hinder discoverability. Medium and low priorities focus on best practices and minor optimizations.
    
    ### Structural Patterns
    
    **Deprecated**: `skills/shared/modules/` directories. Shared modules must be relocated into the consuming skill's own `modules/` directory. The evaluator flags any remaining `skills/shared/` as a structural warning.
    
    **Current**: Each skill owns its modules at `skills/<skill-name>/modules/`. Cross-skill references use relative paths (e.g., `../skill-authoring/modules/description-writing.md`).
    
    ## Resources
    
    ### Shared Modules: Cross-Skill Patterns
    - **Skill Selection**: See [skill-selection-judgment.md](../../shared-modules/skill-selection-judgment.md)
    - **Trigger Patterns**: See [trigger-patterns.md](modules/evaluation-criteria.md)
    
    ### Skill-Specific Modules
    - **Trigger Isolation Analysis**: See `modules/trigger-isolation-analysis.md`
    - **Authoring Checklist**: See `modules/authoring-checklist.md`
    - **Evaluation Workflows**: See `modules/evaluation-workflows.md`
    - **Advanced Tool Use Analysis**: See `modules/advanced-tool-use-analysis.md`
    - **Evaluation Framework**: See `modules/evaluation-framework.md`
    - **Integration Patterns**: See `modules/integration.md`
    - **Troubleshooting**: See `modules/troubleshooting.md`
    - **Pressure Testing**: See `modules/pressure-testing.md`
    - **Integration Testing**: See `modules/integration-testing.md`
    - **Performance Benchmarking**: See `modules/performance-benchmarking.md`
    
    ### Tools and Automation
    - **Tools**: Executable analysis utilities in `scripts/` directory.
    - **Automation**: Setup and validation scripts in `scripts/automation/`.
    
    ## Exit Criteria
    
    - [ ] Every audited skill receives a score across all five dimensions (structure compliance,
      content quality, token efficiency, activation reliability, tool integration) summing to 100.
    - [ ] Any skill with a deprecated `skills/shared/` module reference is listed as a structural
      warning in the audit output.
    - [ ] `make check-compliance TARGET=<skill>` exits 0 for skills reported as passing, confirming
      the compliance-checker agrees with the audit score.
    - [ ] An improvement plan is produced for any skill scoring below 75, with findings ordered by
      priority: critical > high > medium > low.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related