session-trends
Analyze trends across session metrics. Computes windowed aggregates, deltas, and compares against MEMORY.md findings. Use periodically for progress tracking.
Install
npx skills add https://github.com/oliver-kriska/claude-elixir-phoenix/tree/main/.claude/skills/session-trends
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oliver-kriska-claude-elixir-phoenix@llmmart
git clone https://github.com/oliver-kriska/claude-elixir-phoenix.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oliver-kriska/claude-elixir-phoenix collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Session Trends
Analyze trends from the metrics ledger. Computes windowed aggregates, fingerprint distributions, and compares against MEMORY.md baselines.
Requirements
Requires .claude/session-metrics/metrics.jsonl from /session-scan.
Usage
/session-trends # All windows (7d, 30d, all)
/session-trends --window 30d # Specific window only
/session-trends --project enaia # Filter by project
/session-trends --compare MEMORY.md # Compare against memory baseline
/session-trends --html out.html # Write HTML report with ASCII bars
For pure context-window stats (max prompt tokens, ctx %, compaction rate)
across raw Claude Code JSONL files, see the --scan-jsonl mode of
compute-metrics.py (inspired by badlogic / earendil-works/pi).
Pipeline
Step 1: Parse Arguments
Extract from $ARGUMENTS:
--window WINDOW: Time window —7d,30d, orall(default: show all three)--project NAME: Filter metrics by project name--compare PATH: Path to MEMORY.md for baseline comparison (default: auto-detect from.claude/project memory)
Step 2: Read Metrics Ledger
Read .claude/session-metrics/metrics.jsonl.
If empty or missing:
No metrics found. Run
/session-scanfirst.
If --project specified, filter entries by project field.
Step 3: Compute Trends via Python
python3 .claude/skills/session-scan/references/compute-metrics.py \
--trends .claude/session-metrics/metrics.jsonl \
--memory {MEMORY_PATH}
Capture the JSON output.
Step 4: Display Trend Report
Format the JSON output as a readable report:
Overview
Total sessions: {N} ({backfilled} backfilled from v1)
Date range: {earliest} to {latest}
Window Comparison
| Metric | 7 days | 30 days | All time |
|-------------------------|--------|---------|----------|
| Sessions | 12 | 45 | 165 |
| Avg friction | 0.28 | 0.24 | 0.22 |
| Max friction | 0.72 | 0.72 | 0.89 |
| Avg opportunity | 0.35 | 0.30 | 0.28 |
| Tier 2 eligible | 40% | 33% | 30% |
| Plugin adoption | 12% | 10% | 8% |
Fingerprint Distribution
| Type | 7d | 30d | All |
|---------------|-----|-----|------|
| bug-fix | 4 | 15 | 52 |
| feature | 3 | 12 | 48 |
| exploration | 2 | 8 | 30 |
| maintenance | 1 | 5 | 18 |
| review | 1 | 3 | 10 |
| refactoring | 1 | 2 | 7 |
MEMORY.md Comparison (if --compare)
Compare measured values against MEMORY.md claims:
| MEMORY.md Claim | Measured | Match? |
|------------------------------|-------------|--------|
| Plugin adoption: 8-12% | 10.2% | Yes |
| Minimal friction in 40+ of 74| 68% smooth | Yes |
Step 5: Write trends.json
Write computed trends to .claude/session-metrics/trends.json.
Step 6: Suggest Actions
Based on trends:
- If friction is increasing: "Friction trending up — run
/session-deep-dive --from-scanto investigate" - If plugin adoption is growing: "Plugin adoption growing — check which commands drive value"
- If many Tier 2 eligible: " sessions need deep analysis"
Output Files
| File | Purpose |
|---|---|
.claude/session-metrics/trends.json |
Computed trend data |
Common Queries
See references/trend-queries.md for interpreting specific trend patterns.
Iron Laws
- ALWAYS use Python for computation — no manual aggregation
- NEVER modify metrics.jsonl — read-only for trends
- ALWAYS show window comparison — single numbers lack context
Acknowledgements
The HTML report layout (preformatted text + ASCII bar charts via █/░)
and per-model + threshold-bucket breakdown (>=80%, >=90%, >=100%,
compaction_rate) were borrowed from
badlogic / earendil-works/pi session-context-stats.mjs.
Our pipeline's qualitative metrics (friction, fingerprint, plugin
opportunity, skill effectiveness) are additive on top.
Files (claude-elixir-phoenix)
-
references
-
trend-queries.md 4.4 KB
# Trend Queries Reference Common questions you can answer with `/session-trends` data, and how to interpret the results. ## Adoption & Usage ### "Is plugin adoption increasing?" Look at `plugin_adoption_rate` across windows: ``` 7d: 15% → 30d: 10% → all: 8% ``` **Interpretation**: If 7d > 30d > all, adoption is accelerating. If 7d < all, recent sessions aren't using the plugin. **Action**: If declining, check which session types (fingerprints) are least likely to use plugin commands. These are automation targets. ### "Which commands are most used?" Not directly in trends — run `/session-scan --list` and grep `phx_commands_used` from metrics.jsonl: ```bash grep -o '"phx_commands_used":\[[^]]*\]' .claude/session-metrics/metrics.jsonl | sort | uniq -c | sort -rn ``` ### "Are sessions getting longer or shorter?" Compare `duration_minutes` averages across windows. Longer sessions may indicate harder problems or more friction. ## Friction Analysis ### "Which session types have most friction?" Cross-reference fingerprint with friction in metrics.jsonl: ```bash python3 -c " import json from collections import defaultdict data = defaultdict(list) for line in open('.claude/session-metrics/metrics.jsonl'): e = json.loads(line) data[e.get('fingerprint','?')].append(e.get('friction_score',0)) for k,v in sorted(data.items(), key=lambda x: -sum(x[1])/len(x[1])): print(f'{k}: avg={sum(v)/len(v):.2f} (n={len(v)})') " ``` **Typical findings**: `bug-fix` and `refactoring` tend to have higher friction than `exploration` or `maintenance`. ### "Are our fixes working?" Compare friction trends over time. If friction is decreasing after plugin improvements: ``` 30d avg: 0.28 → 7d avg: 0.22 (improvement) ``` Also check: Tier 2 eligible percentage declining = fewer high-friction sessions. ### "What are the biggest friction sources?" Aggregate friction signals across sessions: ```bash python3 -c " import json from collections import Counter signals = Counter() for line in open('.claude/session-metrics/metrics.jsonl'): e = json.loads(line) for k,v in e.get('friction_signals',{}).items(): if isinstance(v,(int,float)) and v > 0: signals[k] += v for k,v in signals.most_common(): print(f'{k}: {v}') " ``` ## Plugin Opportunities ### "What commands are most frequently missed?" Aggregate `could_use` from `plugin_signals`: ```bash python3 -c " import json from collections import Counter missed = Counter() for line in open('.claude/session-metrics/metrics.jsonl'): e = json.loads(line) for cmd in e.get('plugin_signals',{}).get('could_use',[]): missed[cmd] += 1 for k,v in missed.most_common(): print(f'/phx:{k}: missed in {v} sessions') " ``` ### "Is Tidewave being utilized?" Check `tidewave_pct` in tool profiles and `tidewave_available` vs `tidewave_used` in plugin signals. ## File & Code Patterns ### "What files are hotspots?" Aggregate `file_hotspots` across sessions to find frequently touched files: ```bash python3 -c " import json from collections import Counter files = Counter() for line in open('.claude/session-metrics/metrics.jsonl'): e = json.loads(line) for h in e.get('file_hotspots',[]): files[h['path']] += h.get('reads',0) + h.get('edits',0) for k,v in files.most_common(20): print(f'{v:4d} {k}') " ``` ### "What domains get the most work?" Aggregate `file_categories`: ```bash python3 -c " import json from collections import Counter cats = Counter() for line in open('.claude/session-metrics/metrics.jsonl'): e = json.loads(line) for k,v in e.get('file_categories',{}).items(): cats[k] += v for k,v in cats.most_common(): print(f'{k}: {v} edits') " ``` ## Session Chaining ### "Is session chaining decreasing?" Track `chain_length` distribution. High chaining = related sessions not completing in one sitting. Note: Session chaining detection requires the scan to identify same-project sessions within 2 hours. Currently set to basic detection — chain_length defaults to 1. ## Backfill Quality ### "How reliable are backfilled metrics?" Backfilled sessions (`"backfilled": true`) have limited signals: - `retry_loops`: always 0 (can't detect from v1 extracts) - `approach_changes`: always 0 - `context_compactions`: always 0 - `tool_bigrams`: empty - `file_hotspots`: empty Use `backfilled_count` in trends to know what percentage of data is lower-quality. For trend analysis, consider filtering to non-backfilled sessions only.
-
-
SKILL.md 4.5 KB
--- name: session-trends description: Analyze trends across session metrics. Computes windowed aggregates, deltas, and compares against MEMORY.md findings. Use periodically for progress tracking. argument-hint: "[--window 7d|30d|all] [--project NAME] [--compare MEMORY.md]" disable-model-invocation: true --- # Session Trends Analyze trends from the metrics ledger. Computes windowed aggregates, fingerprint distributions, and compares against MEMORY.md baselines. ## Requirements Requires `.claude/session-metrics/metrics.jsonl` from `/session-scan`. ## Usage ``` /session-trends # All windows (7d, 30d, all) /session-trends --window 30d # Specific window only /session-trends --project enaia # Filter by project /session-trends --compare MEMORY.md # Compare against memory baseline /session-trends --html out.html # Write HTML report with ASCII bars ``` For pure context-window stats (max prompt tokens, ctx %, compaction rate) across raw Claude Code JSONL files, see the `--scan-jsonl` mode of `compute-metrics.py` (inspired by badlogic / earendil-works/pi). ## Pipeline ### Step 1: Parse Arguments Extract from `$ARGUMENTS`: - **`--window WINDOW`**: Time window — `7d`, `30d`, or `all` (default: show all three) - **`--project NAME`**: Filter metrics by project name - **`--compare PATH`**: Path to MEMORY.md for baseline comparison (default: auto-detect from `.claude/` project memory) ### Step 2: Read Metrics Ledger Read `.claude/session-metrics/metrics.jsonl`. If empty or missing: > No metrics found. Run `/session-scan` first. If `--project` specified, filter entries by project field. ### Step 3: Compute Trends via Python ```bash python3 .claude/skills/session-scan/references/compute-metrics.py \ --trends .claude/session-metrics/metrics.jsonl \ --memory {MEMORY_PATH} ``` Capture the JSON output. ### Step 4: Display Trend Report Format the JSON output as a readable report: #### Overview ``` Total sessions: {N} ({backfilled} backfilled from v1) Date range: {earliest} to {latest} ``` #### Window Comparison ``` | Metric | 7 days | 30 days | All time | |-------------------------|--------|---------|----------| | Sessions | 12 | 45 | 165 | | Avg friction | 0.28 | 0.24 | 0.22 | | Max friction | 0.72 | 0.72 | 0.89 | | Avg opportunity | 0.35 | 0.30 | 0.28 | | Tier 2 eligible | 40% | 33% | 30% | | Plugin adoption | 12% | 10% | 8% | ``` #### Fingerprint Distribution ``` | Type | 7d | 30d | All | |---------------|-----|-----|------| | bug-fix | 4 | 15 | 52 | | feature | 3 | 12 | 48 | | exploration | 2 | 8 | 30 | | maintenance | 1 | 5 | 18 | | review | 1 | 3 | 10 | | refactoring | 1 | 2 | 7 | ``` #### MEMORY.md Comparison (if --compare) Compare measured values against MEMORY.md claims: ``` | MEMORY.md Claim | Measured | Match? | |------------------------------|-------------|--------| | Plugin adoption: 8-12% | 10.2% | Yes | | Minimal friction in 40+ of 74| 68% smooth | Yes | ``` ### Step 5: Write trends.json Write computed trends to `.claude/session-metrics/trends.json`. ### Step 6: Suggest Actions Based on trends: - If friction is **increasing**: "Friction trending up — run `/session-deep-dive --from-scan` to investigate" - If plugin adoption is **growing**: "Plugin adoption growing — check which commands drive value" - If many Tier 2 eligible: "{N} sessions need deep analysis" ## Output Files | File | Purpose | |------|---------| | `.claude/session-metrics/trends.json` | Computed trend data | ## Common Queries See `references/trend-queries.md` for interpreting specific trend patterns. ## Iron Laws 1. **ALWAYS use Python for computation** — no manual aggregation 2. **NEVER modify metrics.jsonl** — read-only for trends 3. **ALWAYS show window comparison** — single numbers lack context ## Acknowledgements The HTML report layout (preformatted text + ASCII bar charts via `█`/`░`) and per-model + threshold-bucket breakdown (`>=80%`, `>=90%`, `>=100%`, `compaction_rate`) were borrowed from [badlogic / earendil-works/pi `session-context-stats.mjs`](https://github.com/earendil-works/pi/blob/main/scripts/session-context-stats.mjs). Our pipeline's qualitative metrics (friction, fingerprint, plugin opportunity, skill effectiveness) are additive on top.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.