skill-monitor
Analyze skill effectiveness across sessions. Computes per-skill metrics (action rate, friction, outcomes), identifies degrading skills, and generates improvement recommendations. Requires session-scan data in metrics.jsonl.
Install
npx skills add https://github.com/oliver-kriska/claude-elixir-phoenix/tree/main/.claude/skills/skill-monitor
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oliver-kriska-claude-elixir-phoenix@llmmart
git clone https://github.com/oliver-kriska/claude-elixir-phoenix.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oliver-kriska/claude-elixir-phoenix collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Skill Monitor
Closed-loop skill effectiveness monitoring. Reads session metrics, computes per-skill signals, identifies what's working and what needs improvement.
Inspired by the deploy-monitor-evaluate-improve feedback loop: skills get better over time instead of staying static.
Requirements
Requires .claude/session-metrics/metrics.jsonl from /session-scan.
If no data: suggest running /session-scan first.
Usage
/skill-monitor # Dashboard: all skills
/skill-monitor --skill review # Deep-dive on one skill
/skill-monitor --improve # Generate improvement recommendations
/skill-monitor --window 30d # Change comparison window (default: 7d)
What Main Context Does
Step 1: Parse Arguments
Extract from $ARGUMENTS:
--skill NAME: Focus on one skill (e.g.,review,plan,investigate)--improve: Spawn analysis agent for improvement recommendations--window PERIOD: Comparison window (7d,30d,all; default:7d)
Step 2: Load Metrics
Read .claude/session-metrics/metrics.jsonl. For each entry, extract
the skill_effectiveness field (added by compute-metrics.py v2).
Filter by window period. Count sessions with and without skill usage.
If no skill_effectiveness data exists in metrics: "Metrics were
computed before skill tracking was added. Run /session-scan --rescan
to recompute."
OTel invocation_trigger (CC v2.1.126+): when compute-metrics.py
ingests claude_code.skill_activated events, each invocation carries
an invocation_trigger of "user-slash", "claude-proactive", or
"nested-skill". If absent (older sessions), default to
"unknown" — do NOT assume "user-slash".
Step 3: Compute Per-Skill Aggregates
For each skill found across all sessions, aggregate:
| Metric | Computation |
|-------------------------|------------------------------------------------|
| Total invocations | Sum of invocation_count across sessions |
| Sessions used in | Count of sessions containing this skill |
| Action rate | Weighted avg of per-session action_rate |
| Avg post-errors | Weighted avg of avg_post_errors |
| Avg post-corrections | Weighted avg of avg_post_corrections |
| Outcome distribution | Count of effective/friction/no_action/mixed |
| Effectiveness score | action_rate - (0.3 * avg_post_corrections) |
| Adjusted score | For analysis/check skills, use lower thresholds |
| Trigger distribution | Counts of user-slash / claude-proactive / nested-skill / unknown |
| Proactive trigger rate | claude-proactive / (user-slash + claude-proactive + nested-skill) |
| Auto-load gap | Skills with 0 claude-proactive invocations across window |
Auto-load gap detection (CC v2.1.126+): Skills with auto-loaded
behavior in their description (i.e., not disable-model-invocation: true)
are EXPECTED to fire as claude-proactive. A skill that is ONLY ever
invoked via user-slash is failing its description's routing intent.
Flag any auto-loadable skill where proactive_trigger_rate == 0 over
the window. This is the structural answer to the "zero skill
auto-loading" gap from the 137-session analysis (see MEMORY.md).
Confidence floor: only flag if total invocations >= 5 in window.
Skill type weighting: Analysis and check skills (verify, triage, perf, boundaries, pr-review, audit) have low action rates BY DESIGN — their success is "found issues" or "confirmed things pass". Apply adjusted thresholds:
| Skill Type | Flag Threshold | Expected Action Rate |
|---|---|---|
| Execution (work, quick, full) | < 0.5 | > 0.7 |
| Analysis (perf, boundaries, audit, pr-review) | < 0.3 | 0.3-0.5 |
| Check (verify, triage) | < 0.1 | 0.0-0.3 |
| Knowledge (compound, learn, brief) | < 0.5 | > 0.5 |
Also compute baseline friction (avg friction of sessions WITHOUT any skill usage) vs skill friction (avg friction of sessions WITH skill usage). Delta = skill_friction - baseline_friction. Negative delta = skills reduce friction (good).
Step 4: Display Dashboard
Dashboard mode (no --skill):
## Skill Effectiveness Dashboard (last {window})
Baseline friction (no skills): 0.32 | With skills: 0.18 | Delta: -0.14
| Skill | Uses | Sessions | Slash/Proactive/Nested | Action% | Errors | Corr | Outcome | Score |
|-----------------|------|----------|------------------------|---------|--------|------|-----------|-------|
| /phx:review | 12 | 8 | 8 / 3 / 1 | 92% | 0.5 | 0.1 | effective | 0.89 |
| /phx:plan | 9 | 7 | 9 / 0 / 0 | 100% | 0.2 | 0.0 | effective | 1.00 |
| /phx:investigate| 5 | 5 | 5 / 0 / 0 | 80% | 1.2 | 0.4 | mixed | 0.68 |
Skills needing attention:
- /phx:investigate (high post-errors)
- /phx:plan (auto-load gap — 0/9 proactive; description not routing)
Flag skills using type-adjusted thresholds (see weighting table above).
Also flag if avg_post_corrections > 1 or outcome is predominantly "friction".
Also flag auto-load gap: auto-loadable skills (without
disable-model-invocation: true) with proactive_trigger_rate == 0 and
total invocations >= 5. This is a description/routing problem — the skill
exists but Claude isn't loading it on its own.
When displaying flagged skills, note if the flag is "expected" for the skill type (e.g., verify at 0.24 is normal for a check skill).
Skill deep-dive (--skill NAME):
Show per-session breakdown for that skill, including session IDs,
dates, individual outcome signals, AND invocation_trigger per
invocation. If a skill is dominated by user-slash triggers, surface
which 1-3 description keywords might unlock proactive routing —
cross-reference against the skill's current description in
plugins/elixir-phoenix/skills/{name}/SKILL.md. If session reports
exist in .claude/session-analysis/, reference them.
Step 5: Improvement Mode (--improve)
Spawn skill-effectiveness-analyzer agent:
Agent(subagent_type="skill-effectiveness-analyzer", model="sonnet", prompt="""
Analyze skill effectiveness data and recommend improvements.
Metrics data: {aggregated_metrics_json}
Sessions with friction outcomes: {session_ids}
For each underperforming skill:
1. Identify failure patterns from outcome signals
2. Propose specific skill/agent changes
3. Suggest new Iron Laws if patterns are systematic
Write recommendations to: .claude/skill-metrics/recommendations-{date}.md
""")
Step 6: Write Output
Write aggregated metrics to .claude/skill-metrics/dashboard-{date}.json:
{
"computed_at": "2026-03-03T14:00:00Z",
"window": "7d",
"baseline_friction": 0.32,
"skill_friction": 0.18,
"friction_delta": -0.14,
"skills": {
"/phx:plan": {
"invocations": 9,
"trigger_distribution": {
"user-slash": 9,
"claude-proactive": 0,
"nested-skill": 0,
"unknown": 0
},
"proactive_trigger_rate": 0.0,
"auto_load_gap": true
}
},
"flagged_skills": ["investigate", "plan:auto-load-gap"]
}
Append-only: never modify previous dashboard files.
Iron Laws
- NEVER modify metrics.jsonl — read-only from this skill
- Baseline comparison is mandatory — raw numbers without baseline are meaningless
- Flag, don't judge — surface data, let the human decide what to fix
- Evidence tags on recommendations — every suggestion needs session citations
- Trigger source must not be inferred — only treat invocations as
user-slash/claude-proactive/nested-skillwhen the OTelinvocation_triggerattribute is present (CC v2.1.126+). Older sessions use"unknown"; never silently bucket them as user-slash — it would hide the auto-load gap.
Integration
/session-scan → metrics.jsonl (with skill_effectiveness)
↓
/skill-monitor → dashboard + flagged skills
↓
/skill-monitor --improve → recommendations
↓
Developer updates skills/agents → deploy → repeat
References
references/effectiveness-metrics.md— Full metrics schema and evaluation criteriareferences/improvement-template.md— Template for improvement recommendations
Files (claude-elixir-phoenix)
-
references
-
effectiveness-metrics.md 8.8 KB
# Skill Effectiveness Metrics Defines the measurement framework for evaluating plugin skill effectiveness across sessions. ## Core Principle The OpenHands feedback loop: **deploy - monitor - evaluate - improve**. The "correctness" signal comes from user behavior — not explicit ratings. When a skill produces useful output, the developer acts on it (edits files, runs tests). When it doesn't help, the developer corrects, ignores, or abandons the approach. ## Per-Session Skill Signals Extracted by `compute-metrics.py` in the `skill_effectiveness` field: ### Raw Signals | Signal | Source | Meaning | |--------|--------|---------| | `invocation_count` | User messages | How many times skill was invoked | | `total_post_edits` | Tool calls after invocation | Edits made following skill | | `total_post_reads` | Tool calls after invocation | Research following skill | | `total_post_test_runs` | Bash calls with `mix test` | Verification after skill | | `total_post_errors` | Error patterns in messages | Failures after skill | | `total_post_corrections` | Correction patterns in user messages | User redirections | | `led_to_action_count` | post_edits > 0 or post_test_runs > 0 | Skill produced action | | `trigger_user_slash` | OTel `invocation_trigger == "user-slash"` (CC ≥ 2.1.126) | User typed `/phx:foo` | | `trigger_proactive` | OTel `invocation_trigger == "claude-proactive"` (CC ≥ 2.1.126) | Claude auto-loaded the skill | | `trigger_nested` | OTel `invocation_trigger == "nested-skill"` (CC ≥ 2.1.126) | Skill invoked from inside another skill | | `trigger_unknown` | Pre-2.1.126 sessions or missing OTel | Source not recoverable | ### OTel `claude_code.skill_activated` Schema Available CC ≥ 2.1.126. Each event carries: | Attribute | Values | Notes | |-----------|--------|-------| | `skill.name` | e.g. `"phx:plan"` | Plugin-prefixed skills include the plugin name | | `invocation_trigger` | `"user-slash"`, `"claude-proactive"`, `"nested-skill"` | Use this verbatim — do NOT infer for older events | | `session_id` | UUID | Joins to session metrics | `compute-metrics.py` should join OTel events on `session_id`, matching `skill.name` against invocation timestamps in the session transcript. When the join is missing (no OTel collector configured, or sessions older than 2.1.126), set `trigger_unknown` to `invocation_count` and skip auto-load gap analysis for that skill. ### Computed Signals | Signal | Formula | Range | Good | |--------|---------|-------|------| | `action_rate` | led_to_action_count / invocation_count | 0-1 | > 0.7 | | `avg_post_errors` | total_post_errors / invocation_count | 0+ | < 1.0 | | `avg_post_corrections` | total_post_corrections / invocation_count | 0+ | < 0.5 | | `dominant_outcome` | Most common outcome classification | enum | "effective" | | `proactive_trigger_rate` | trigger_proactive / (trigger_user_slash + trigger_proactive + trigger_nested) | 0-1 | > 0.2 (auto-loadable skills) | | `auto_load_gap` | proactive_trigger_rate == 0 AND not disable-model-invocation AND invocation_count >= 5 | bool | false | ### Outcome Classification Each invocation is classified into one of: | Outcome | Criteria | Interpretation | |---------|----------|----------------| | `effective` | No errors, no corrections, led to action | Skill worked well | | `friction` | Corrections > 0 or errors > 3 | Skill caused problems | | `no_action` | No edits, no test runs | Skill output was ignored | | `mixed` | Some errors but also action | Partial success | ## Cross-Session Aggregates The `/skill-monitor` command computes these from metrics.jsonl: ### Per-Skill Metrics | Aggregate | Computation | Threshold | |-----------|-------------|-----------| | Total invocations | Sum across sessions | Informational | | Session count | Distinct sessions | Informational | | Weighted action rate | Σ(action_rate × n) / Σ(n) | Flag if < 0.5 | | Weighted avg errors | Σ(avg_errors × n) / Σ(n) | Flag if > 2.0 | | Weighted avg corrections | Σ(avg_corr × n) / Σ(n) | Flag if > 1.0 | | Outcome distribution | Counts of each outcome type | Flag if friction > 30% | | Effectiveness score | action_rate - (0.3 × corrections) | Flag if < 0.5 | ### Baseline Comparison Critical for meaningful interpretation. Without baseline, raw numbers are uninterpretable. **Baseline group**: Sessions with zero skill invocations. **Skill group**: Sessions with at least one skill invocation. | Comparison | Good Signal | Bad Signal | |------------|-------------|------------| | Friction delta (skill - baseline) | Negative (skills reduce friction) | Positive (skills add friction) | | Error rate delta | Negative | Positive | | Duration delta | Negative (faster) | Positive (slower) | ### Trend Detection Compare across time windows to detect degradation: ``` 7d_effectiveness vs 30d_effectiveness → trend direction ``` | Trend | Interpretation | Action | |-------|----------------|--------| | Rising effectiveness | Skill improvements working | Continue monitoring | | Stable | No change | Check if already good enough | | Declining | Skill degrading | Investigate: model changes? code changes? | | Insufficient data | < 3 invocations in window | Extend window or wait | ## Per-Skill Evaluation Criteria Different skills have different "correctness" proxies: ### `/phx:review` | Proxy | Signal | How to measure | |-------|--------|----------------| | Suggestion acceptance | Edits to files flagged in review | post_edits to review-mentioned files | | False positive rate | Corrections after review | post_corrections | | Completeness | No new issues found later | absence of subsequent /phx:review | ### `/phx:plan` | Proxy | Signal | How to measure | |-------|--------|----------------| | Task completion | Checkboxes completed | Requires plan.md parsing | | Scope accuracy | No corrections during /phx:work | post_corrections in work sessions | | Rework rate | Plan re-done via --existing | Subsequent /phx:plan --existing | ### `/phx:investigate` | Proxy | Signal | How to measure | |-------|--------|----------------| | Root cause found | Edits follow investigation | post_edits > 0 | | Fix success | Tests pass after fix | post_test_runs with no post_errors | | Debugging loop break | No retry loops after | absence of retry_loop friction signal | ### `/phx:compound` | Proxy | Signal | How to measure | |-------|--------|----------------| | Solution created | File written to .claude/solutions/ | post_edits to solutions dir | | Reuse | Solution referenced in future sessions | Cross-session grep (deep-dive only) | | Quality | No corrections during creation | post_corrections == 0 | ### `/phx:verify` | Proxy | Signal | How to measure | |-------|--------|----------------| | Pass rate | Tests/compile/credo pass | post_errors == 0 | | Issues found | Led to fixes | post_edits > 0 after errors found | | False alarms | Corrections about irrelevant failures | post_corrections | ### `/phx:quick` | Proxy | Signal | How to measure | |-------|--------|----------------| | One-shot success | Task done without follow-up | no subsequent corrections | | Scope containment | Small change count | post_edits < 5 | | Speed | Low tool count after invocation | total window tools < 20 | ## Dashboard JSON Schema Output format for `.claude/skill-metrics/dashboard-{date}.json`: ```json { "computed_at": "ISO8601", "window": "7d|30d|all", "session_count": 50, "sessions_with_skills": 18, "sessions_without_skills": 32, "baseline_friction": 0.32, "skill_friction": 0.18, "friction_delta": -0.14, "skills": { "/phx:review": { "invocations": 12, "sessions": 8, "action_rate": 0.92, "avg_post_errors": 0.5, "avg_post_corrections": 0.1, "effectiveness_score": 0.89, "outcome_distribution": { "effective": 9, "mixed": 2, "friction": 1, "no_action": 0 }, "trigger_distribution": { "user-slash": 8, "claude-proactive": 3, "nested-skill": 1, "unknown": 0 }, "proactive_trigger_rate": 0.25, "auto_load_gap": false, "trend": "stable" } }, "flagged_skills": [ { "skill": "/phx:investigate", "reason": "effectiveness_score < 0.5", "score": 0.42, "recommendation": "Review error handling patterns" }, { "skill": "/phx:plan", "reason": "auto_load_gap (proactive_trigger_rate == 0 over 9 invocations)", "proactive_trigger_rate": 0.0, "recommendation": "Tune description keywords — Claude routes only on explicit slash" } ] } ``` ## Minimum Data Requirements | Analysis Level | Minimum Sessions | Minimum Invocations | |----------------|-----------------|---------------------| | Dashboard | 5 | 3 per skill | | Trend comparison | 10 (across both windows) | 5 per skill | | Improvement recs | 10 | 5 per flagged skill | | Statistical confidence | 30 | 15 per skill | Below minimums: show data with "LOW CONFIDENCE" warning. -
improvement-template.md 4.2 KB
# Skill Improvement Analysis Template Template for the skill-effectiveness-analyzer agent. Produces actionable recommendations for improving underperforming skills. ## Inputs 1. **Aggregated metrics** — per-skill dashboard data 2. **Flagged skills** — skills below effectiveness thresholds 3. **Session IDs** — sessions where skills had friction outcomes 4. **Session reports** (if available) — qualitative analysis from deep-dive ## Analysis Sections ### 1. Executive Summary One paragraph: how many skills analyzed, how many flagged, overall plugin health trend. ### 2. Per-Skill Analysis For each flagged skill: ```markdown ## /phx:{skill} — Effectiveness: {score} **Status**: {Needs improvement | Critical | Watch} **Evidence strength**: {STRONG | MODERATE | WEAK} ### Failure Pattern Describe the dominant failure mode: - What goes wrong when this skill underperforms? - At what stage does friction occur? - Is it a skill design issue, agent behavior issue, or user expectation mismatch? ### Supporting Data | Metric | Value | Threshold | Status | |--------|-------|-----------|--------| | Action rate | 0.45 | > 0.7 | BELOW | | Avg post-errors | 2.3 | < 1.0 | ABOVE | | Avg post-corrections | 1.1 | < 0.5 | ABOVE | | Friction outcome % | 40% | < 30% | ABOVE | ### Session Evidence | Session | Date | Project | Outcome | Key Signal | |---------|------|---------|---------|------------| | abc123 | 2026-02-28 | myapp | friction | 4 corrections after review | | def456 | 2026-03-01 | myapp | no_action | Review output ignored | ### Root Cause Hypothesis {Specific hypothesis about WHY the skill underperforms} Examples: - "Review agent flags too many low-severity issues, causing fatigue" - "Investigation skill doesn't check compound solutions first" - "Plan tasks are too granular for simple features" - "Skill prompt doesn't handle edge case X" ### Recommended Changes 1. **{Change type}**: {Description} - File: `{path to skill or agent file}` - What: {specific change} - Expected impact: {how this improves effectiveness} - Evidence: {session IDs supporting this} Change types: - SKILL_PROMPT — Modify skill SKILL.md instructions - AGENT_PROMPT — Modify agent system prompt - IRON_LAW — Add new Iron Law to prevent pattern - REFERENCE — Update or add reference documentation - HOOK — Add/modify hook behavior - WORKFLOW — Change skill interaction/ordering ``` ### 3. Cross-Skill Patterns Patterns that affect multiple skills: | Pattern | Affected Skills | Impact | Fix | |---------|----------------|--------|-----| | Over-verbose output | review, plan | Information fatigue | Add output length limits | | Missing context handoff | plan → work | Rework on transition | Add scratchpad passing | ### 4. Positive Patterns What's working well — don't break these: | Skill | Strength | Why It Works | |-------|----------|--------------| | /phx:compound | High action rate | Clear schema, minimal friction | | /phx:plan | Good completion | Checkbox format enables tracking | ### 5. Priority Ranking Ordered by: evidence_strength × impact × ease_of_fix | # | Skill | Change | Evidence | Impact | Effort | Priority | |---|-------|--------|----------|--------|--------|----------| | 1 | /phx:review | Reduce false positives | STRONG (5 sessions) | High | Low | P0 | | 2 | /phx:investigate | Add solution search | MODERATE (3 sessions) | High | Medium | P1 | ### 6. Tracking Plan How to verify improvements after changes: ``` 1. Apply recommended changes 2. Run /session-scan --rescan after 5+ sessions 3. Run /skill-monitor --window 7d 4. Compare effectiveness scores to this report's baseline 5. If improved: continue monitoring If not: run /skill-monitor --improve for new analysis ``` ## Output Format Write as structured markdown to: `.claude/skill-metrics/recommendations-{date}.md` Keep under 200 lines. Every recommendation must cite specific sessions and metrics. Include tracking plan so improvements can be verified. ## Anti-Patterns - **Don't recommend adding complexity** to fix simplicity issues - **Don't suggest new skills** when existing ones need fixing - **Don't blame the user** — if corrections are high, the skill is misleading, not the user - **Don't recommend changes without evidence** — every change needs session citations
-
-
SKILL.md 8.6 KB
--- name: skill-monitor description: Analyze skill effectiveness across sessions. Computes per-skill metrics (action rate, friction, outcomes), identifies degrading skills, and generates improvement recommendations. Requires session-scan data in metrics.jsonl. argument-hint: "[--skill NAME] [--improve] [--window 7d|30d|all]" disable-model-invocation: true --- # Skill Monitor Closed-loop skill effectiveness monitoring. Reads session metrics, computes per-skill signals, identifies what's working and what needs improvement. Inspired by the deploy-monitor-evaluate-improve feedback loop: skills get better over time instead of staying static. ## Requirements Requires `.claude/session-metrics/metrics.jsonl` from `/session-scan`. If no data: suggest running `/session-scan` first. ## Usage ``` /skill-monitor # Dashboard: all skills /skill-monitor --skill review # Deep-dive on one skill /skill-monitor --improve # Generate improvement recommendations /skill-monitor --window 30d # Change comparison window (default: 7d) ``` ## What Main Context Does ### Step 1: Parse Arguments Extract from `$ARGUMENTS`: - **`--skill NAME`**: Focus on one skill (e.g., `review`, `plan`, `investigate`) - **`--improve`**: Spawn analysis agent for improvement recommendations - **`--window PERIOD`**: Comparison window (`7d`, `30d`, `all`; default: `7d`) ### Step 2: Load Metrics Read `.claude/session-metrics/metrics.jsonl`. For each entry, extract the `skill_effectiveness` field (added by compute-metrics.py v2). Filter by window period. Count sessions with and without skill usage. If no `skill_effectiveness` data exists in metrics: "Metrics were computed before skill tracking was added. Run `/session-scan --rescan` to recompute." **OTel `invocation_trigger` (CC v2.1.126+)**: when `compute-metrics.py` ingests `claude_code.skill_activated` events, each invocation carries an `invocation_trigger` of `"user-slash"`, `"claude-proactive"`, or `"nested-skill"`. If absent (older sessions), default to `"unknown"` — do NOT assume `"user-slash"`. ### Step 3: Compute Per-Skill Aggregates For each skill found across all sessions, aggregate: ``` | Metric | Computation | |-------------------------|------------------------------------------------| | Total invocations | Sum of invocation_count across sessions | | Sessions used in | Count of sessions containing this skill | | Action rate | Weighted avg of per-session action_rate | | Avg post-errors | Weighted avg of avg_post_errors | | Avg post-corrections | Weighted avg of avg_post_corrections | | Outcome distribution | Count of effective/friction/no_action/mixed | | Effectiveness score | action_rate - (0.3 * avg_post_corrections) | | Adjusted score | For analysis/check skills, use lower thresholds | | Trigger distribution | Counts of user-slash / claude-proactive / nested-skill / unknown | | Proactive trigger rate | claude-proactive / (user-slash + claude-proactive + nested-skill) | | Auto-load gap | Skills with 0 claude-proactive invocations across window | ``` **Auto-load gap detection (CC v2.1.126+)**: Skills with `auto-loaded` behavior in their description (i.e., not `disable-model-invocation: true`) are EXPECTED to fire as `claude-proactive`. A skill that is ONLY ever invoked via `user-slash` is failing its description's routing intent. Flag any auto-loadable skill where `proactive_trigger_rate == 0` over the window. This is the structural answer to the "zero skill auto-loading" gap from the 137-session analysis (see MEMORY.md). **Confidence floor**: only flag if total invocations >= 5 in window. **Skill type weighting**: Analysis and check skills (verify, triage, perf, boundaries, pr-review, audit) have low action rates BY DESIGN — their success is "found issues" or "confirmed things pass". Apply adjusted thresholds: | Skill Type | Flag Threshold | Expected Action Rate | |------------|---------------|---------------------| | Execution (work, quick, full) | < 0.5 | > 0.7 | | Analysis (perf, boundaries, audit, pr-review) | < 0.3 | 0.3-0.5 | | Check (verify, triage) | < 0.1 | 0.0-0.3 | | Knowledge (compound, learn, brief) | < 0.5 | > 0.5 | Also compute **baseline friction** (avg friction of sessions WITHOUT any skill usage) vs **skill friction** (avg friction of sessions WITH skill usage). Delta = skill_friction - baseline_friction. Negative delta = skills reduce friction (good). ### Step 4: Display Dashboard **Dashboard mode** (no `--skill`): ``` ## Skill Effectiveness Dashboard (last {window}) Baseline friction (no skills): 0.32 | With skills: 0.18 | Delta: -0.14 | Skill | Uses | Sessions | Slash/Proactive/Nested | Action% | Errors | Corr | Outcome | Score | |-----------------|------|----------|------------------------|---------|--------|------|-----------|-------| | /phx:review | 12 | 8 | 8 / 3 / 1 | 92% | 0.5 | 0.1 | effective | 0.89 | | /phx:plan | 9 | 7 | 9 / 0 / 0 | 100% | 0.2 | 0.0 | effective | 1.00 | | /phx:investigate| 5 | 5 | 5 / 0 / 0 | 80% | 1.2 | 0.4 | mixed | 0.68 | Skills needing attention: - /phx:investigate (high post-errors) - /phx:plan (auto-load gap — 0/9 proactive; description not routing) ``` Flag skills using type-adjusted thresholds (see weighting table above). Also flag if avg_post_corrections > 1 or outcome is predominantly "friction". **Also flag auto-load gap**: auto-loadable skills (without `disable-model-invocation: true`) with proactive_trigger_rate == 0 and total invocations >= 5. This is a description/routing problem — the skill exists but Claude isn't loading it on its own. When displaying flagged skills, note if the flag is "expected" for the skill type (e.g., verify at 0.24 is normal for a check skill). **Skill deep-dive** (`--skill NAME`): Show per-session breakdown for that skill, including session IDs, dates, individual outcome signals, AND `invocation_trigger` per invocation. If a skill is dominated by `user-slash` triggers, surface which 1-3 description keywords might unlock proactive routing — cross-reference against the skill's current description in `plugins/elixir-phoenix/skills/{name}/SKILL.md`. If session reports exist in `.claude/session-analysis/`, reference them. ### Step 5: Improvement Mode (--improve) Spawn `skill-effectiveness-analyzer` agent: ``` Agent(subagent_type="skill-effectiveness-analyzer", model="sonnet", prompt=""" Analyze skill effectiveness data and recommend improvements. Metrics data: {aggregated_metrics_json} Sessions with friction outcomes: {session_ids} For each underperforming skill: 1. Identify failure patterns from outcome signals 2. Propose specific skill/agent changes 3. Suggest new Iron Laws if patterns are systematic Write recommendations to: .claude/skill-metrics/recommendations-{date}.md """) ``` ### Step 6: Write Output Write aggregated metrics to `.claude/skill-metrics/dashboard-{date}.json`: ```json { "computed_at": "2026-03-03T14:00:00Z", "window": "7d", "baseline_friction": 0.32, "skill_friction": 0.18, "friction_delta": -0.14, "skills": { "/phx:plan": { "invocations": 9, "trigger_distribution": { "user-slash": 9, "claude-proactive": 0, "nested-skill": 0, "unknown": 0 }, "proactive_trigger_rate": 0.0, "auto_load_gap": true } }, "flagged_skills": ["investigate", "plan:auto-load-gap"] } ``` Append-only: never modify previous dashboard files. ## Iron Laws 1. **NEVER modify metrics.jsonl** — read-only from this skill 2. **Baseline comparison is mandatory** — raw numbers without baseline are meaningless 3. **Flag, don't judge** — surface data, let the human decide what to fix 4. **Evidence tags on recommendations** — every suggestion needs session citations 5. **Trigger source must not be inferred** — only treat invocations as `user-slash` / `claude-proactive` / `nested-skill` when the OTel `invocation_trigger` attribute is present (CC v2.1.126+). Older sessions use `"unknown"`; never silently bucket them as user-slash — it would hide the auto-load gap. ## Integration ``` /session-scan → metrics.jsonl (with skill_effectiveness) ↓ /skill-monitor → dashboard + flagged skills ↓ /skill-monitor --improve → recommendations ↓ Developer updates skills/agents → deploy → repeat ``` ## References - `references/effectiveness-metrics.md` — Full metrics schema and evaluation criteria - `references/improvement-template.md` — Template for improvement recommendations
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.