Claude Skill

skill-monitor

Analyze skill effectiveness across sessions. Computes per-skill metrics (action rate, friction, outcomes), identifies degrading skills, and generates improvement recommendations. Requires session-scan data in metrics.jsonl.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download oliver-kriska-claude-elixir-phoenix-.claude_skills_skill-monitor-9767a82.zip · 9 KB
Part of oliver-kriska/claude-elixir-phoenix — 93 skills

Install

skills CLI npx skills add https://github.com/oliver-kriska/claude-elixir-phoenix/tree/main/.claude/skills/skill-monitor
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oliver-kriska-claude-elixir-phoenix@llmmart
Git git clone https://github.com/oliver-kriska/claude-elixir-phoenix.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oliver-kriska/claude-elixir-phoenix collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Skill Monitor

Closed-loop skill effectiveness monitoring. Reads session metrics, computes per-skill signals, identifies what's working and what needs improvement.

Inspired by the deploy-monitor-evaluate-improve feedback loop: skills get better over time instead of staying static.

Requirements

Requires .claude/session-metrics/metrics.jsonl from /session-scan. If no data: suggest running /session-scan first.

Usage

/skill-monitor                     # Dashboard: all skills
/skill-monitor --skill review      # Deep-dive on one skill
/skill-monitor --improve           # Generate improvement recommendations
/skill-monitor --window 30d        # Change comparison window (default: 7d)

What Main Context Does

Step 1: Parse Arguments

Extract from $ARGUMENTS:

  • --skill NAME: Focus on one skill (e.g., review, plan, investigate)
  • --improve: Spawn analysis agent for improvement recommendations
  • --window PERIOD: Comparison window (7d, 30d, all; default: 7d)

Step 2: Load Metrics

Read .claude/session-metrics/metrics.jsonl. For each entry, extract the skill_effectiveness field (added by compute-metrics.py v2).

Filter by window period. Count sessions with and without skill usage.

If no skill_effectiveness data exists in metrics: "Metrics were computed before skill tracking was added. Run /session-scan --rescan to recompute."

OTel invocation_trigger (CC v2.1.126+): when compute-metrics.py ingests claude_code.skill_activated events, each invocation carries an invocation_trigger of "user-slash", "claude-proactive", or "nested-skill". If absent (older sessions), default to "unknown" — do NOT assume "user-slash".

Step 3: Compute Per-Skill Aggregates

For each skill found across all sessions, aggregate:

| Metric                  | Computation                                    |
|-------------------------|------------------------------------------------|
| Total invocations       | Sum of invocation_count across sessions        |
| Sessions used in        | Count of sessions containing this skill        |
| Action rate             | Weighted avg of per-session action_rate         |
| Avg post-errors         | Weighted avg of avg_post_errors                |
| Avg post-corrections    | Weighted avg of avg_post_corrections           |
| Outcome distribution    | Count of effective/friction/no_action/mixed    |
| Effectiveness score     | action_rate - (0.3 * avg_post_corrections)     |
| Adjusted score          | For analysis/check skills, use lower thresholds |
| Trigger distribution    | Counts of user-slash / claude-proactive / nested-skill / unknown |
| Proactive trigger rate  | claude-proactive / (user-slash + claude-proactive + nested-skill) |
| Auto-load gap           | Skills with 0 claude-proactive invocations across window |

Auto-load gap detection (CC v2.1.126+): Skills with auto-loaded behavior in their description (i.e., not disable-model-invocation: true) are EXPECTED to fire as claude-proactive. A skill that is ONLY ever invoked via user-slash is failing its description's routing intent. Flag any auto-loadable skill where proactive_trigger_rate == 0 over the window. This is the structural answer to the "zero skill auto-loading" gap from the 137-session analysis (see MEMORY.md). Confidence floor: only flag if total invocations >= 5 in window.

Skill type weighting: Analysis and check skills (verify, triage, perf, boundaries, pr-review, audit) have low action rates BY DESIGN — their success is "found issues" or "confirmed things pass". Apply adjusted thresholds:

Skill Type Flag Threshold Expected Action Rate
Execution (work, quick, full) < 0.5 > 0.7
Analysis (perf, boundaries, audit, pr-review) < 0.3 0.3-0.5
Check (verify, triage) < 0.1 0.0-0.3
Knowledge (compound, learn, brief) < 0.5 > 0.5

Also compute baseline friction (avg friction of sessions WITHOUT any skill usage) vs skill friction (avg friction of sessions WITH skill usage). Delta = skill_friction - baseline_friction. Negative delta = skills reduce friction (good).

Step 4: Display Dashboard

Dashboard mode (no --skill):

## Skill Effectiveness Dashboard (last {window})

Baseline friction (no skills): 0.32 | With skills: 0.18 | Delta: -0.14

| Skill           | Uses | Sessions | Slash/Proactive/Nested | Action% | Errors | Corr | Outcome   | Score |
|-----------------|------|----------|------------------------|---------|--------|------|-----------|-------|
| /phx:review     | 12   | 8        |    8 /  3 /  1         | 92%     | 0.5    | 0.1  | effective | 0.89  |
| /phx:plan       | 9    | 7        |    9 /  0 /  0         | 100%    | 0.2    | 0.0  | effective | 1.00  |
| /phx:investigate| 5    | 5        |    5 /  0 /  0         | 80%     | 1.2    | 0.4  | mixed     | 0.68  |

Skills needing attention:
- /phx:investigate (high post-errors)
- /phx:plan (auto-load gap — 0/9 proactive; description not routing)

Flag skills using type-adjusted thresholds (see weighting table above). Also flag if avg_post_corrections > 1 or outcome is predominantly "friction". Also flag auto-load gap: auto-loadable skills (without disable-model-invocation: true) with proactive_trigger_rate == 0 and total invocations >= 5. This is a description/routing problem — the skill exists but Claude isn't loading it on its own.

When displaying flagged skills, note if the flag is "expected" for the skill type (e.g., verify at 0.24 is normal for a check skill).

Skill deep-dive (--skill NAME):

Show per-session breakdown for that skill, including session IDs, dates, individual outcome signals, AND invocation_trigger per invocation. If a skill is dominated by user-slash triggers, surface which 1-3 description keywords might unlock proactive routing — cross-reference against the skill's current description in plugins/elixir-phoenix/skills/{name}/SKILL.md. If session reports exist in .claude/session-analysis/, reference them.

Step 5: Improvement Mode (--improve)

Spawn skill-effectiveness-analyzer agent:

Agent(subagent_type="skill-effectiveness-analyzer", model="sonnet", prompt="""
Analyze skill effectiveness data and recommend improvements.

Metrics data: {aggregated_metrics_json}

Sessions with friction outcomes: {session_ids}

For each underperforming skill:
1. Identify failure patterns from outcome signals
2. Propose specific skill/agent changes
3. Suggest new Iron Laws if patterns are systematic

Write recommendations to: .claude/skill-metrics/recommendations-{date}.md
""")

Step 6: Write Output

Write aggregated metrics to .claude/skill-metrics/dashboard-{date}.json:

{
  "computed_at": "2026-03-03T14:00:00Z",
  "window": "7d",
  "baseline_friction": 0.32,
  "skill_friction": 0.18,
  "friction_delta": -0.14,
  "skills": {
    "/phx:plan": {
      "invocations": 9,
      "trigger_distribution": {
        "user-slash": 9,
        "claude-proactive": 0,
        "nested-skill": 0,
        "unknown": 0
      },
      "proactive_trigger_rate": 0.0,
      "auto_load_gap": true
    }
  },
  "flagged_skills": ["investigate", "plan:auto-load-gap"]
}

Append-only: never modify previous dashboard files.

Iron Laws

  1. NEVER modify metrics.jsonl — read-only from this skill
  2. Baseline comparison is mandatory — raw numbers without baseline are meaningless
  3. Flag, don't judge — surface data, let the human decide what to fix
  4. Evidence tags on recommendations — every suggestion needs session citations
  5. Trigger source must not be inferred — only treat invocations as user-slash / claude-proactive / nested-skill when the OTel invocation_trigger attribute is present (CC v2.1.126+). Older sessions use "unknown"; never silently bucket them as user-slash — it would hide the auto-load gap.

Integration

/session-scan → metrics.jsonl (with skill_effectiveness)
       ↓
/skill-monitor → dashboard + flagged skills
       ↓
/skill-monitor --improve → recommendations
       ↓
Developer updates skills/agents → deploy → repeat

References

  • references/effectiveness-metrics.md — Full metrics schema and evaluation criteria
  • references/improvement-template.md — Template for improvement recommendations
Files (claude-elixir-phoenix)
  • references
    • effectiveness-metrics.md 8.8 KB
      # Skill Effectiveness Metrics
      
      Defines the measurement framework for evaluating plugin skill
      effectiveness across sessions.
      
      ## Core Principle
      
      The OpenHands feedback loop: **deploy - monitor - evaluate - improve**.
      
      The "correctness" signal comes from user behavior — not explicit
      ratings. When a skill produces useful output, the developer acts
      on it (edits files, runs tests). When it doesn't help, the
      developer corrects, ignores, or abandons the approach.
      
      ## Per-Session Skill Signals
      
      Extracted by `compute-metrics.py` in the `skill_effectiveness` field:
      
      ### Raw Signals
      
      | Signal | Source | Meaning |
      |--------|--------|---------|
      | `invocation_count` | User messages | How many times skill was invoked |
      | `total_post_edits` | Tool calls after invocation | Edits made following skill |
      | `total_post_reads` | Tool calls after invocation | Research following skill |
      | `total_post_test_runs` | Bash calls with `mix test` | Verification after skill |
      | `total_post_errors` | Error patterns in messages | Failures after skill |
      | `total_post_corrections` | Correction patterns in user messages | User redirections |
      | `led_to_action_count` | post_edits > 0 or post_test_runs > 0 | Skill produced action |
      | `trigger_user_slash` | OTel `invocation_trigger == "user-slash"` (CC ≥ 2.1.126) | User typed `/phx:foo` |
      | `trigger_proactive` | OTel `invocation_trigger == "claude-proactive"` (CC ≥ 2.1.126) | Claude auto-loaded the skill |
      | `trigger_nested` | OTel `invocation_trigger == "nested-skill"` (CC ≥ 2.1.126) | Skill invoked from inside another skill |
      | `trigger_unknown` | Pre-2.1.126 sessions or missing OTel | Source not recoverable |
      
      ### OTel `claude_code.skill_activated` Schema
      
      Available CC ≥ 2.1.126. Each event carries:
      
      | Attribute | Values | Notes |
      |-----------|--------|-------|
      | `skill.name` | e.g. `"phx:plan"` | Plugin-prefixed skills include the plugin name |
      | `invocation_trigger` | `"user-slash"`, `"claude-proactive"`, `"nested-skill"` | Use this verbatim — do NOT infer for older events |
      | `session_id` | UUID | Joins to session metrics |
      
      `compute-metrics.py` should join OTel events on `session_id`,
      matching `skill.name` against invocation timestamps in the session
      transcript. When the join is missing (no OTel collector configured,
      or sessions older than 2.1.126), set `trigger_unknown` to
      `invocation_count` and skip auto-load gap analysis for that skill.
      
      ### Computed Signals
      
      | Signal | Formula | Range | Good |
      |--------|---------|-------|------|
      | `action_rate` | led_to_action_count / invocation_count | 0-1 | > 0.7 |
      | `avg_post_errors` | total_post_errors / invocation_count | 0+ | < 1.0 |
      | `avg_post_corrections` | total_post_corrections / invocation_count | 0+ | < 0.5 |
      | `dominant_outcome` | Most common outcome classification | enum | "effective" |
      | `proactive_trigger_rate` | trigger_proactive / (trigger_user_slash + trigger_proactive + trigger_nested) | 0-1 | > 0.2 (auto-loadable skills) |
      | `auto_load_gap` | proactive_trigger_rate == 0 AND not disable-model-invocation AND invocation_count >= 5 | bool | false |
      
      ### Outcome Classification
      
      Each invocation is classified into one of:
      
      | Outcome | Criteria | Interpretation |
      |---------|----------|----------------|
      | `effective` | No errors, no corrections, led to action | Skill worked well |
      | `friction` | Corrections > 0 or errors > 3 | Skill caused problems |
      | `no_action` | No edits, no test runs | Skill output was ignored |
      | `mixed` | Some errors but also action | Partial success |
      
      ## Cross-Session Aggregates
      
      The `/skill-monitor` command computes these from metrics.jsonl:
      
      ### Per-Skill Metrics
      
      | Aggregate | Computation | Threshold |
      |-----------|-------------|-----------|
      | Total invocations | Sum across sessions | Informational |
      | Session count | Distinct sessions | Informational |
      | Weighted action rate | Σ(action_rate × n) / Σ(n) | Flag if < 0.5 |
      | Weighted avg errors | Σ(avg_errors × n) / Σ(n) | Flag if > 2.0 |
      | Weighted avg corrections | Σ(avg_corr × n) / Σ(n) | Flag if > 1.0 |
      | Outcome distribution | Counts of each outcome type | Flag if friction > 30% |
      | Effectiveness score | action_rate - (0.3 × corrections) | Flag if < 0.5 |
      
      ### Baseline Comparison
      
      Critical for meaningful interpretation. Without baseline, raw
      numbers are uninterpretable.
      
      **Baseline group**: Sessions with zero skill invocations.
      **Skill group**: Sessions with at least one skill invocation.
      
      | Comparison | Good Signal | Bad Signal |
      |------------|-------------|------------|
      | Friction delta (skill - baseline) | Negative (skills reduce friction) | Positive (skills add friction) |
      | Error rate delta | Negative | Positive |
      | Duration delta | Negative (faster) | Positive (slower) |
      
      ### Trend Detection
      
      Compare across time windows to detect degradation:
      
      ```
      7d_effectiveness vs 30d_effectiveness → trend direction
      ```
      
      | Trend | Interpretation | Action |
      |-------|----------------|--------|
      | Rising effectiveness | Skill improvements working | Continue monitoring |
      | Stable | No change | Check if already good enough |
      | Declining | Skill degrading | Investigate: model changes? code changes? |
      | Insufficient data | < 3 invocations in window | Extend window or wait |
      
      ## Per-Skill Evaluation Criteria
      
      Different skills have different "correctness" proxies:
      
      ### `/phx:review`
      
      | Proxy | Signal | How to measure |
      |-------|--------|----------------|
      | Suggestion acceptance | Edits to files flagged in review | post_edits to review-mentioned files |
      | False positive rate | Corrections after review | post_corrections |
      | Completeness | No new issues found later | absence of subsequent /phx:review |
      
      ### `/phx:plan`
      
      | Proxy | Signal | How to measure |
      |-------|--------|----------------|
      | Task completion | Checkboxes completed | Requires plan.md parsing |
      | Scope accuracy | No corrections during /phx:work | post_corrections in work sessions |
      | Rework rate | Plan re-done via --existing | Subsequent /phx:plan --existing |
      
      ### `/phx:investigate`
      
      | Proxy | Signal | How to measure |
      |-------|--------|----------------|
      | Root cause found | Edits follow investigation | post_edits > 0 |
      | Fix success | Tests pass after fix | post_test_runs with no post_errors |
      | Debugging loop break | No retry loops after | absence of retry_loop friction signal |
      
      ### `/phx:compound`
      
      | Proxy | Signal | How to measure |
      |-------|--------|----------------|
      | Solution created | File written to .claude/solutions/ | post_edits to solutions dir |
      | Reuse | Solution referenced in future sessions | Cross-session grep (deep-dive only) |
      | Quality | No corrections during creation | post_corrections == 0 |
      
      ### `/phx:verify`
      
      | Proxy | Signal | How to measure |
      |-------|--------|----------------|
      | Pass rate | Tests/compile/credo pass | post_errors == 0 |
      | Issues found | Led to fixes | post_edits > 0 after errors found |
      | False alarms | Corrections about irrelevant failures | post_corrections |
      
      ### `/phx:quick`
      
      | Proxy | Signal | How to measure |
      |-------|--------|----------------|
      | One-shot success | Task done without follow-up | no subsequent corrections |
      | Scope containment | Small change count | post_edits < 5 |
      | Speed | Low tool count after invocation | total window tools < 20 |
      
      ## Dashboard JSON Schema
      
      Output format for `.claude/skill-metrics/dashboard-{date}.json`:
      
      ```json
      {
        "computed_at": "ISO8601",
        "window": "7d|30d|all",
        "session_count": 50,
        "sessions_with_skills": 18,
        "sessions_without_skills": 32,
        "baseline_friction": 0.32,
        "skill_friction": 0.18,
        "friction_delta": -0.14,
        "skills": {
          "/phx:review": {
            "invocations": 12,
            "sessions": 8,
            "action_rate": 0.92,
            "avg_post_errors": 0.5,
            "avg_post_corrections": 0.1,
            "effectiveness_score": 0.89,
            "outcome_distribution": {
              "effective": 9,
              "mixed": 2,
              "friction": 1,
              "no_action": 0
            },
            "trigger_distribution": {
              "user-slash": 8,
              "claude-proactive": 3,
              "nested-skill": 1,
              "unknown": 0
            },
            "proactive_trigger_rate": 0.25,
            "auto_load_gap": false,
            "trend": "stable"
          }
        },
        "flagged_skills": [
          {
            "skill": "/phx:investigate",
            "reason": "effectiveness_score < 0.5",
            "score": 0.42,
            "recommendation": "Review error handling patterns"
          },
          {
            "skill": "/phx:plan",
            "reason": "auto_load_gap (proactive_trigger_rate == 0 over 9 invocations)",
            "proactive_trigger_rate": 0.0,
            "recommendation": "Tune description keywords — Claude routes only on explicit slash"
          }
        ]
      }
      ```
      
      ## Minimum Data Requirements
      
      | Analysis Level | Minimum Sessions | Minimum Invocations |
      |----------------|-----------------|---------------------|
      | Dashboard | 5 | 3 per skill |
      | Trend comparison | 10 (across both windows) | 5 per skill |
      | Improvement recs | 10 | 5 per flagged skill |
      | Statistical confidence | 30 | 15 per skill |
      
      Below minimums: show data with "LOW CONFIDENCE" warning.
      
    • improvement-template.md 4.2 KB
      # Skill Improvement Analysis Template
      
      Template for the skill-effectiveness-analyzer agent. Produces
      actionable recommendations for improving underperforming skills.
      
      ## Inputs
      
      1. **Aggregated metrics** — per-skill dashboard data
      2. **Flagged skills** — skills below effectiveness thresholds
      3. **Session IDs** — sessions where skills had friction outcomes
      4. **Session reports** (if available) — qualitative analysis from deep-dive
      
      ## Analysis Sections
      
      ### 1. Executive Summary
      
      One paragraph: how many skills analyzed, how many flagged, overall
      plugin health trend.
      
      ### 2. Per-Skill Analysis
      
      For each flagged skill:
      
      ```markdown
      ## /phx:{skill} — Effectiveness: {score}
      
      **Status**: {Needs improvement | Critical | Watch}
      **Evidence strength**: {STRONG | MODERATE | WEAK}
      
      ### Failure Pattern
      
      Describe the dominant failure mode:
      - What goes wrong when this skill underperforms?
      - At what stage does friction occur?
      - Is it a skill design issue, agent behavior issue, or user expectation mismatch?
      
      ### Supporting Data
      
      | Metric | Value | Threshold | Status |
      |--------|-------|-----------|--------|
      | Action rate | 0.45 | > 0.7 | BELOW |
      | Avg post-errors | 2.3 | < 1.0 | ABOVE |
      | Avg post-corrections | 1.1 | < 0.5 | ABOVE |
      | Friction outcome % | 40% | < 30% | ABOVE |
      
      ### Session Evidence
      
      | Session | Date | Project | Outcome | Key Signal |
      |---------|------|---------|---------|------------|
      | abc123 | 2026-02-28 | myapp | friction | 4 corrections after review |
      | def456 | 2026-03-01 | myapp | no_action | Review output ignored |
      
      ### Root Cause Hypothesis
      
      {Specific hypothesis about WHY the skill underperforms}
      
      Examples:
      - "Review agent flags too many low-severity issues, causing fatigue"
      - "Investigation skill doesn't check compound solutions first"
      - "Plan tasks are too granular for simple features"
      - "Skill prompt doesn't handle edge case X"
      
      ### Recommended Changes
      
      1. **{Change type}**: {Description}
         - File: `{path to skill or agent file}`
         - What: {specific change}
         - Expected impact: {how this improves effectiveness}
         - Evidence: {session IDs supporting this}
      
      Change types:
      - SKILL_PROMPT — Modify skill SKILL.md instructions
      - AGENT_PROMPT — Modify agent system prompt
      - IRON_LAW — Add new Iron Law to prevent pattern
      - REFERENCE — Update or add reference documentation
      - HOOK — Add/modify hook behavior
      - WORKFLOW — Change skill interaction/ordering
      ```
      
      ### 3. Cross-Skill Patterns
      
      Patterns that affect multiple skills:
      
      | Pattern | Affected Skills | Impact | Fix |
      |---------|----------------|--------|-----|
      | Over-verbose output | review, plan | Information fatigue | Add output length limits |
      | Missing context handoff | plan → work | Rework on transition | Add scratchpad passing |
      
      ### 4. Positive Patterns
      
      What's working well — don't break these:
      
      | Skill | Strength | Why It Works |
      |-------|----------|--------------|
      | /phx:compound | High action rate | Clear schema, minimal friction |
      | /phx:plan | Good completion | Checkbox format enables tracking |
      
      ### 5. Priority Ranking
      
      Ordered by: evidence_strength × impact × ease_of_fix
      
      | # | Skill | Change | Evidence | Impact | Effort | Priority |
      |---|-------|--------|----------|--------|--------|----------|
      | 1 | /phx:review | Reduce false positives | STRONG (5 sessions) | High | Low | P0 |
      | 2 | /phx:investigate | Add solution search | MODERATE (3 sessions) | High | Medium | P1 |
      
      ### 6. Tracking Plan
      
      How to verify improvements after changes:
      
      ```
      1. Apply recommended changes
      2. Run /session-scan --rescan after 5+ sessions
      3. Run /skill-monitor --window 7d
      4. Compare effectiveness scores to this report's baseline
      5. If improved: continue monitoring
         If not: run /skill-monitor --improve for new analysis
      ```
      
      ## Output Format
      
      Write as structured markdown to:
      `.claude/skill-metrics/recommendations-{date}.md`
      
      Keep under 200 lines. Every recommendation must cite specific
      sessions and metrics. Include tracking plan so improvements
      can be verified.
      
      ## Anti-Patterns
      
      - **Don't recommend adding complexity** to fix simplicity issues
      - **Don't suggest new skills** when existing ones need fixing
      - **Don't blame the user** — if corrections are high, the skill
        is misleading, not the user
      - **Don't recommend changes without evidence** — every change
        needs session citations
      
  • SKILL.md 8.6 KB
    ---
    name: skill-monitor
    description: Analyze skill effectiveness across sessions. Computes per-skill metrics (action rate, friction, outcomes), identifies degrading skills, and generates improvement recommendations. Requires session-scan data in metrics.jsonl.
    argument-hint: "[--skill NAME] [--improve] [--window 7d|30d|all]"
    disable-model-invocation: true
    ---
    
    # Skill Monitor
    
    Closed-loop skill effectiveness monitoring. Reads session metrics,
    computes per-skill signals, identifies what's working and what needs
    improvement.
    
    Inspired by the deploy-monitor-evaluate-improve feedback loop:
    skills get better over time instead of staying static.
    
    ## Requirements
    
    Requires `.claude/session-metrics/metrics.jsonl` from `/session-scan`.
    If no data: suggest running `/session-scan` first.
    
    ## Usage
    
    ```
    /skill-monitor                     # Dashboard: all skills
    /skill-monitor --skill review      # Deep-dive on one skill
    /skill-monitor --improve           # Generate improvement recommendations
    /skill-monitor --window 30d        # Change comparison window (default: 7d)
    ```
    
    ## What Main Context Does
    
    ### Step 1: Parse Arguments
    
    Extract from `$ARGUMENTS`:
    
    - **`--skill NAME`**: Focus on one skill (e.g., `review`, `plan`, `investigate`)
    - **`--improve`**: Spawn analysis agent for improvement recommendations
    - **`--window PERIOD`**: Comparison window (`7d`, `30d`, `all`; default: `7d`)
    
    ### Step 2: Load Metrics
    
    Read `.claude/session-metrics/metrics.jsonl`. For each entry, extract
    the `skill_effectiveness` field (added by compute-metrics.py v2).
    
    Filter by window period. Count sessions with and without skill usage.
    
    If no `skill_effectiveness` data exists in metrics: "Metrics were
    computed before skill tracking was added. Run `/session-scan --rescan`
    to recompute."
    
    **OTel `invocation_trigger` (CC v2.1.126+)**: when `compute-metrics.py`
    ingests `claude_code.skill_activated` events, each invocation carries
    an `invocation_trigger` of `"user-slash"`, `"claude-proactive"`, or
    `"nested-skill"`. If absent (older sessions), default to
    `"unknown"` — do NOT assume `"user-slash"`.
    
    ### Step 3: Compute Per-Skill Aggregates
    
    For each skill found across all sessions, aggregate:
    
    ```
    | Metric                  | Computation                                    |
    |-------------------------|------------------------------------------------|
    | Total invocations       | Sum of invocation_count across sessions        |
    | Sessions used in        | Count of sessions containing this skill        |
    | Action rate             | Weighted avg of per-session action_rate         |
    | Avg post-errors         | Weighted avg of avg_post_errors                |
    | Avg post-corrections    | Weighted avg of avg_post_corrections           |
    | Outcome distribution    | Count of effective/friction/no_action/mixed    |
    | Effectiveness score     | action_rate - (0.3 * avg_post_corrections)     |
    | Adjusted score          | For analysis/check skills, use lower thresholds |
    | Trigger distribution    | Counts of user-slash / claude-proactive / nested-skill / unknown |
    | Proactive trigger rate  | claude-proactive / (user-slash + claude-proactive + nested-skill) |
    | Auto-load gap           | Skills with 0 claude-proactive invocations across window |
    ```
    
    **Auto-load gap detection (CC v2.1.126+)**: Skills with `auto-loaded`
    behavior in their description (i.e., not `disable-model-invocation: true`)
    are EXPECTED to fire as `claude-proactive`. A skill that is ONLY ever
    invoked via `user-slash` is failing its description's routing intent.
    Flag any auto-loadable skill where `proactive_trigger_rate == 0` over
    the window. This is the structural answer to the "zero skill
    auto-loading" gap from the 137-session analysis (see MEMORY.md).
    **Confidence floor**: only flag if total invocations >= 5 in window.
    
    **Skill type weighting**: Analysis and check skills (verify, triage,
    perf, boundaries, pr-review, audit) have low action rates BY DESIGN —
    their success is "found issues" or "confirmed things pass". Apply
    adjusted thresholds:
    
    | Skill Type | Flag Threshold | Expected Action Rate |
    |------------|---------------|---------------------|
    | Execution (work, quick, full) | < 0.5 | > 0.7 |
    | Analysis (perf, boundaries, audit, pr-review) | < 0.3 | 0.3-0.5 |
    | Check (verify, triage) | < 0.1 | 0.0-0.3 |
    | Knowledge (compound, learn, brief) | < 0.5 | > 0.5 |
    
    Also compute **baseline friction** (avg friction of sessions WITHOUT
    any skill usage) vs **skill friction** (avg friction of sessions
    WITH skill usage). Delta = skill_friction - baseline_friction.
    Negative delta = skills reduce friction (good).
    
    ### Step 4: Display Dashboard
    
    **Dashboard mode** (no `--skill`):
    
    ```
    ## Skill Effectiveness Dashboard (last {window})
    
    Baseline friction (no skills): 0.32 | With skills: 0.18 | Delta: -0.14
    
    | Skill           | Uses | Sessions | Slash/Proactive/Nested | Action% | Errors | Corr | Outcome   | Score |
    |-----------------|------|----------|------------------------|---------|--------|------|-----------|-------|
    | /phx:review     | 12   | 8        |    8 /  3 /  1         | 92%     | 0.5    | 0.1  | effective | 0.89  |
    | /phx:plan       | 9    | 7        |    9 /  0 /  0         | 100%    | 0.2    | 0.0  | effective | 1.00  |
    | /phx:investigate| 5    | 5        |    5 /  0 /  0         | 80%     | 1.2    | 0.4  | mixed     | 0.68  |
    
    Skills needing attention:
    - /phx:investigate (high post-errors)
    - /phx:plan (auto-load gap — 0/9 proactive; description not routing)
    ```
    
    Flag skills using type-adjusted thresholds (see weighting table above).
    Also flag if avg_post_corrections > 1 or outcome is predominantly "friction".
    **Also flag auto-load gap**: auto-loadable skills (without
    `disable-model-invocation: true`) with proactive_trigger_rate == 0 and
    total invocations >= 5. This is a description/routing problem — the skill
    exists but Claude isn't loading it on its own.
    
    When displaying flagged skills, note if the flag is "expected" for the
    skill type (e.g., verify at 0.24 is normal for a check skill).
    
    **Skill deep-dive** (`--skill NAME`):
    
    Show per-session breakdown for that skill, including session IDs,
    dates, individual outcome signals, AND `invocation_trigger` per
    invocation. If a skill is dominated by `user-slash` triggers, surface
    which 1-3 description keywords might unlock proactive routing —
    cross-reference against the skill's current description in
    `plugins/elixir-phoenix/skills/{name}/SKILL.md`. If session reports
    exist in `.claude/session-analysis/`, reference them.
    
    ### Step 5: Improvement Mode (--improve)
    
    Spawn `skill-effectiveness-analyzer` agent:
    
    ```
    Agent(subagent_type="skill-effectiveness-analyzer", model="sonnet", prompt="""
    Analyze skill effectiveness data and recommend improvements.
    
    Metrics data: {aggregated_metrics_json}
    
    Sessions with friction outcomes: {session_ids}
    
    For each underperforming skill:
    1. Identify failure patterns from outcome signals
    2. Propose specific skill/agent changes
    3. Suggest new Iron Laws if patterns are systematic
    
    Write recommendations to: .claude/skill-metrics/recommendations-{date}.md
    """)
    ```
    
    ### Step 6: Write Output
    
    Write aggregated metrics to `.claude/skill-metrics/dashboard-{date}.json`:
    
    ```json
    {
      "computed_at": "2026-03-03T14:00:00Z",
      "window": "7d",
      "baseline_friction": 0.32,
      "skill_friction": 0.18,
      "friction_delta": -0.14,
      "skills": {
        "/phx:plan": {
          "invocations": 9,
          "trigger_distribution": {
            "user-slash": 9,
            "claude-proactive": 0,
            "nested-skill": 0,
            "unknown": 0
          },
          "proactive_trigger_rate": 0.0,
          "auto_load_gap": true
        }
      },
      "flagged_skills": ["investigate", "plan:auto-load-gap"]
    }
    ```
    
    Append-only: never modify previous dashboard files.
    
    ## Iron Laws
    
    1. **NEVER modify metrics.jsonl** — read-only from this skill
    2. **Baseline comparison is mandatory** — raw numbers without baseline are meaningless
    3. **Flag, don't judge** — surface data, let the human decide what to fix
    4. **Evidence tags on recommendations** — every suggestion needs session citations
    5. **Trigger source must not be inferred** — only treat invocations as
       `user-slash` / `claude-proactive` / `nested-skill` when the OTel
       `invocation_trigger` attribute is present (CC v2.1.126+). Older
       sessions use `"unknown"`; never silently bucket them as user-slash —
       it would hide the auto-load gap.
    
    ## Integration
    
    ```
    /session-scan → metrics.jsonl (with skill_effectiveness)
           ↓
    /skill-monitor → dashboard + flagged skills
           ↓
    /skill-monitor --improve → recommendations
           ↓
    Developer updates skills/agents → deploy → repeat
    ```
    
    ## References
    
    - `references/effectiveness-metrics.md` — Full metrics schema and evaluation criteria
    - `references/improvement-template.md` — Template for improvement recommendations
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related