session-deep-dive
Deep qualitative analysis of high-signal sessions. Spawns subagents with v2 template, synthesizes patterns, compares against known findings. Use after /session-scan.
Install
npx skills add https://github.com/oliver-kriska/claude-elixir-phoenix/tree/main/.claude/skills/session-deep-dive
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oliver-kriska-claude-elixir-phoenix@llmmart
git clone https://github.com/oliver-kriska/claude-elixir-phoenix.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oliver-kriska/claude-elixir-phoenix collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Session Deep Dive (Tier 2)
Qualitative analysis of high-signal sessions identified by /session-scan.
Spawns subagents with pre-computed metrics context for focused analysis.
Requirements
Requires ccrider MCP. If not available:
ccrider MCP is required. See: https://github.com/neilberkman/ccrider
Usage
/session-deep-dive ffa155ee-ed8a-492c-8797-878fcbec4d9e
/session-deep-dive --last # Most recent Tier 2 eligible
/session-deep-dive --from-scan # All Tier 2 eligible from last scan
/session-deep-dive --from-scan --compare .claude/UPDATED_PLUGIN_REPORT_160_SESSIONS.md
Pipeline
Step 1: Resolve Target Sessions
From $ARGUMENTS:
- Session ID: Single session to analyze
--last: Most recent Tier 2 eligible session from metrics.jsonl--from-scan: All sessions wheretier2_eligible: trueANDtier2_completed: falsein.claude/session-metrics/metrics.jsonl--compare REPORT.md: Previous report to compare against (default: most recent.claude/session-analysis/insights-*.md)
If no metrics.jsonl exists, tell the user:
No metrics found. Run
/session-scanfirst to discover and score sessions.
Step 2: Load Pre-computed Metrics
For each target session, read its entry from metrics.jsonl.
Format the metrics as a context block for subagent prompts:
## Pre-computed Metrics (from /session-scan)
- Friction: 0.42 (retry_loops: 1, user_corrections: 3, approach_changes: 2)
- Fingerprint: bug-fix (confidence: 0.85)
- Plugin opportunity: 0.65 (could use: investigate, quick)
- Tool profile: Read 28.7%, Edit 15.2%, Bash 19.3%, Tidewave 22.8%
- Duration: 78 minutes, 19 user messages, 171 tool calls
Determine PROJECT_ROOT from current working directory.
Step 3: Fetch Transcripts — One Subagent Per Session
CRITICAL: One ccrider call = one subagent. Full transcripts are 5-30KB each. Even 3 per worker floods the worker's context.
For EACH session, spawn a haiku subagent:
Task(subagent_type="general-purpose", model="haiku", mode="bypassPermissions", prompt="""
Fetch one session transcript and save it.
1. mcp__ccrider__get_session_messages(session_id: "{SESSION_ID}")
If > 200 messages: use last_n: 200
2. Write transcript to {PROJECT_ROOT}/.claude/session-analysis/{SHORT_ID}-transcript.md
Format:
# Session: {SHORT_ID}
Project: {PROJECT}
Date: {DATE}
Messages: {COUNT}
## Messages
### User (seq N)
{content}
### Assistant (seq N)
{content}
3. Report: "Wrote {SHORT_ID}-transcript.md ({N} messages)"
""")
Spawn ALL fetch subagents in parallel. Wait for all to complete.
Step 4: Analyze Sessions
Read the analysis template — inline it into subagent prompts:
Glob: **/session-deep-dive/references/analysis-template-v2.md
ALWAYS use subagents — never analyze in main context.
- 1-6 sessions: Spawn sonnet subagents (one per session)
- 7+ sessions: Spawn haiku subagents for speed
Each analysis subagent prompt:
Read the session transcript at . Apply the analysis template below to analyze this session. The pre-computed metrics below give you quantitative context — validate them and add qualitative depth.
Write your report (under 200 lines) to .
Reports go to .claude/session-analysis/{short_id}-report.md.
Step 5: Compress (if 3+ sessions)
If 3+ sessions analyzed, spawn context-supervisor (haiku) to compress:
Read all report files in
.claude/session-analysis/*-report.md. Write a consolidated summary to.claude/session-analysis/summaries/consolidated.md. Preserve: friction patterns, plugin opportunities, evidence strength tags. Remove: per-file details, generic observations, repeated context.
Step 6: Synthesize
Read the synthesis template:
Glob: **/session-deep-dive/references/synthesis-template.md
Read the --compare report (or latest insights file).
Read MEMORY.md for known findings.
If 3+ sessions: read summaries/consolidated.md (NOT individual reports).
If 1-2 sessions: read individual reports directly.
Produce synthesis comparing:
- New findings vs known patterns from MEMORY.md
- Confirmed patterns (seen before, still present)
- New patterns (not in previous reports)
- Resolved patterns (previously noted, no new occurrences)
Step 7: Update Ledger
Use Python to safely update metrics.jsonl — never manually
read/modify/rewrite in the LLM context:
python3 -c "
import json
ids = {SESSION_IDS_SET} # e.g., {'ffa155ee-...', '90a74843-...'}
lines = open('{PROJECT_ROOT}/.claude/session-metrics/metrics.jsonl').readlines()
with open('{PROJECT_ROOT}/.claude/session-metrics/metrics.jsonl', 'w') as f:
for line in lines:
entry = json.loads(line)
if entry.get('session_id') in ids:
entry['tier2_completed'] = True
f.write(json.dumps(entry) + '\n')
"
Step 8: Write Output
Write synthesis to .claude/session-analysis/insights-{date}.md
Present key findings directly in conversation. Tell user:
Full report:
.claude/session-analysis/insights-{date}.mdPer-session reports:.claude/session-analysis/{id}-report.md
Output Files
| File | Purpose |
|---|---|
.claude/session-analysis/{id}-transcript.md |
Raw transcript |
.claude/session-analysis/{id}-report.md |
Per-session analysis |
.claude/session-analysis/summaries/consolidated.md |
Compressed reports |
.claude/session-analysis/insights-{date}.md |
Cross-session synthesis |
Iron Laws
- ONE ccrider call = ONE subagent — never batch multiple fetches
- NEVER fetch or analyze in main context — always subagents
- Absolute paths in subagent prompts — subagents don't inherit skill context
- Python for jsonl updates — never manually rewrite in LLM context
- ALWAYS pass pre-computed metrics to analysis subagents — don't re-derive
- NEVER skip synthesis — cross-session patterns are the real value
- TAG evidence strength — every finding must be STRONG/MODERATE/WEAK
Files (claude-elixir-phoenix)
-
references
-
analysis-template-v2.md 4.1 KB
# Session Analysis Template v2 Analyze this Claude Code session transcript and produce a structured report. You have both pre-computed quantitative metrics AND the full transcript. ## Goals 1. **Validate metrics** — do the pre-computed scores match the qualitative reality? 2. **Add depth** — identify specific friction moments, user preferences, workflow patterns 3. **Assess plugin fit** — which commands/skills would help, which wouldn't? 4. **Tag evidence** — every finding must have a strength tag ## Evidence Strength Tags Tag every finding with one of: - **STRONG**: Direct evidence (user complained, explicit error, 3+ occurrences) - **MODERATE**: Indirect evidence (pattern suggests friction, 1-2 occurrences) - **WEAK**: Inference only (could be interpreted differently) ## Analysis Sections ### 1. Session Summary - What was the developer trying to accomplish? - Single task or multiple tasks? - Success level: fully, partially, not at all? - Elixir/Phoenix domains touched (LiveView, Ecto, Oban, etc.)? - Does the fingerprint from metrics match your assessment? ### 2. User Correction Tracking Enumerate every user correction or redirection: | # | User Said | What Went Wrong | Impact | |---|-----------|-----------------|--------| | 1 | "no, I meant..." | Claude misunderstood scope | Wasted 5 tool calls | This directly validates the `user_corrections` friction signal. ### 3. Decision Preferences Identify code style and workflow preferences: - Pattern matching vs if/else/cond? - `with` chains vs nested `case`? - Test-first vs implementation-first? - Inline vs extracted functions? - Prefers detailed explanations or terse responses? - How they handle review findings (fix all vs selective)? ### 4. How They Worked - Planned before coding or dove straight in? - Iterative cycle pattern (edit → test → fix → test)? - Used subagents or worked solo? - Debugging approach (read-first vs trial-and-error)? - Used Tidewave MCP? (project_eval, browser_eval, execute_sql_query) - Tool mix interpretation (Read-heavy = exploration, Edit-heavy = implementation) ### 5. Friction Points For each friction point found: | # | Type | Description | Evidence | Strength | |---|------|-------------|----------|----------| | 1 | Error loop | mix compile failed 4× | Bash calls 23-27 | STRONG | | 2 | Approach change | Switched from GenServer to Task | Edits 15-20 | MODERATE | Types: error_loop, approach_change, manual_repetition, long_debugging, scope_creep, missing_context, tool_confusion ### 6. Plugin Skills Assessment #### Used Commands If any `/phx:*` commands were used: | Command | Worked Well? | Issues? | |---------|-------------|---------| | `/phx:plan` | Yes — kept scope focused | Plan was too detailed for small task | #### Suggested Commands For each friction point, suggest a specific plugin command: | Friction Point | Suggested Command | Why It Helps | Strength | |----------------|-------------------|-------------|----------| | Error loop (#1) | `/phx:investigate` | Structured 4-track analysis | STRONG | Only suggest commands that genuinely match. Don't force-fit. #### Hook Effectiveness - Did PostToolUse verification fire? (mix compile + format after edits) - Was the security Iron Laws reminder shown for auth files? - Did the developer heed or ignore hook output? ### 7. Plugin Improvement Opportunities Most important section. For each opportunity: ``` **[STRONG/MODERATE/WEAK] {Category}: {Description}** Evidence: {specific messages, commands, patterns from transcript} Session count estimate: {how many other sessions likely have this} Suggested implementation: {concrete suggestion} ``` Categories: - Missing automation - Missing Iron Law - Missing skill/agent - Auto-loading gap - Workflow friction WITH plugin - Tool integration gap ### 8. Efficiency Assessment Rate: **Smooth** / **Some friction** / **High friction** / **Abandoned** Estimate effort savings with right plugin skills: {X}% ## Output Format Write structured markdown with all sections above. Keep under 200 lines. Be concrete — cite actual messages, commands, patterns. Every finding must have an evidence strength tag. -
synthesis-template.md 2.2 KB
# Cross-Session Synthesis Template Synthesize findings from multiple session analysis reports into a trend-aware summary that compares against known patterns. ## Inputs 1. **Per-session reports** (from analysis-template-v2) 2. **Previous synthesis report** (for trend comparison) 3. **MEMORY.md** (for known findings baseline) ## Synthesis Sections ### 1. Confirmed Patterns Patterns seen in previous reports/MEMORY.md that are still present. | Pattern | Previous Count | New Count | Total | Trend | |---------|---------------|-----------|-------|-------| | Zero skill auto-loading | 137 | +5 | 142 | Stable | | PR review workflow demand | 9 | +2 | 11 | Growing | Only include patterns with STRONG or MODERATE evidence in new sessions. ### 2. New Patterns Patterns not found in previous reports or MEMORY.md. | Pattern | Sessions | Evidence | Strength | |---------|----------|----------|----------| | {new finding} | 3 | {citations} | STRONG | Require at least 2 sessions OR 1 session with STRONG evidence. ### 3. Resolved Patterns Previously noted patterns with no new occurrences. | Pattern | Last Seen | Sessions Since | Status | |---------|-----------|----------------|--------| | {old issue} | 2026-01-15 | 12 | Likely resolved | ### 4. Actionable Recommendations Max 5 recommendations, ordered by evidence strength × impact. | # | Recommendation | Evidence | Impact | Effort | |---|----------------|----------|--------|--------| | 1 | {what to do} | {N sessions, strength} | High | Low | Each recommendation must cite specific sessions and evidence. ### 5. Updated Statistics | Metric | Previous | Current | Delta | |--------|----------|---------|-------| | Total sessions analyzed | 160 | 165 | +5 | | Avg friction score | 0.22 | 0.24 | +0.02 | | Plugin adoption rate | 8% | 10% | +2% | | Tier 2 eligible rate | 30% | 28% | -2% | | Most common fingerprint | bug-fix | bug-fix | — | ### 6. MEMORY.md Update Suggestions List specific edits to MEMORY.md based on findings: - **Add**: {new confirmed pattern to add} - **Update**: {existing entry with new data} - **Remove**: {pattern that appears resolved} ## Output Format Write as structured markdown. Every claim must cite sessions. Keep under 150 lines. Focus on actionable, evidence-backed findings.
-
-
SKILL.md 6.4 KB
--- name: session-deep-dive description: Deep qualitative analysis of high-signal sessions. Spawns subagents with v2 template, synthesizes patterns, compares against known findings. Use after /session-scan. argument-hint: "<session-id> | --last | --from-scan [--compare REPORT.md]" disable-model-invocation: true --- # Session Deep Dive (Tier 2) Qualitative analysis of high-signal sessions identified by `/session-scan`. Spawns subagents with pre-computed metrics context for focused analysis. ## Requirements Requires **ccrider MCP**. If not available: > ccrider MCP is required. See: <https://github.com/neilberkman/ccrider> ## Usage ``` /session-deep-dive ffa155ee-ed8a-492c-8797-878fcbec4d9e /session-deep-dive --last # Most recent Tier 2 eligible /session-deep-dive --from-scan # All Tier 2 eligible from last scan /session-deep-dive --from-scan --compare .claude/UPDATED_PLUGIN_REPORT_160_SESSIONS.md ``` ## Pipeline ### Step 1: Resolve Target Sessions From `$ARGUMENTS`: - **Session ID**: Single session to analyze - **`--last`**: Most recent Tier 2 eligible session from metrics.jsonl - **`--from-scan`**: All sessions where `tier2_eligible: true` AND `tier2_completed: false` in `.claude/session-metrics/metrics.jsonl` - **`--compare REPORT.md`**: Previous report to compare against (default: most recent `.claude/session-analysis/insights-*.md`) If no metrics.jsonl exists, tell the user: > No metrics found. Run `/session-scan` first to discover and score sessions. ### Step 2: Load Pre-computed Metrics For each target session, read its entry from `metrics.jsonl`. Format the metrics as a context block for subagent prompts: ``` ## Pre-computed Metrics (from /session-scan) - Friction: 0.42 (retry_loops: 1, user_corrections: 3, approach_changes: 2) - Fingerprint: bug-fix (confidence: 0.85) - Plugin opportunity: 0.65 (could use: investigate, quick) - Tool profile: Read 28.7%, Edit 15.2%, Bash 19.3%, Tidewave 22.8% - Duration: 78 minutes, 19 user messages, 171 tool calls ``` Determine `PROJECT_ROOT` from current working directory. ### Step 3: Fetch Transcripts — One Subagent Per Session **CRITICAL: One ccrider call = one subagent.** Full transcripts are 5-30KB each. Even 3 per worker floods the worker's context. For EACH session, spawn a **haiku** subagent: ``` Task(subagent_type="general-purpose", model="haiku", mode="bypassPermissions", prompt=""" Fetch one session transcript and save it. 1. mcp__ccrider__get_session_messages(session_id: "{SESSION_ID}") If > 200 messages: use last_n: 200 2. Write transcript to {PROJECT_ROOT}/.claude/session-analysis/{SHORT_ID}-transcript.md Format: # Session: {SHORT_ID} Project: {PROJECT} Date: {DATE} Messages: {COUNT} ## Messages ### User (seq N) {content} ### Assistant (seq N) {content} 3. Report: "Wrote {SHORT_ID}-transcript.md ({N} messages)" """) ``` **Spawn ALL fetch subagents in parallel.** Wait for all to complete. ### Step 4: Analyze Sessions Read the analysis template — inline it into subagent prompts: ``` Glob: **/session-deep-dive/references/analysis-template-v2.md ``` **ALWAYS use subagents** — never analyze in main context. - **1-6 sessions**: Spawn **sonnet** subagents (one per session) - **7+ sessions**: Spawn **haiku** subagents for speed Each analysis subagent prompt: > Read the session transcript at {transcript_path}. > Apply the analysis template below to analyze this session. > The pre-computed metrics below give you quantitative context — > validate them and add qualitative depth. > > {metrics_context_block} > > {analysis_template_content} > > Write your report (under 200 lines) to {report_path}. Reports go to `.claude/session-analysis/{short_id}-report.md`. ### Step 5: Compress (if 3+ sessions) If 3+ sessions analyzed, spawn context-supervisor (haiku) to compress: > Read all report files in `.claude/session-analysis/*-report.md`. > Write a consolidated summary to `.claude/session-analysis/summaries/consolidated.md`. > Preserve: friction patterns, plugin opportunities, evidence strength tags. > Remove: per-file details, generic observations, repeated context. ### Step 6: Synthesize Read the synthesis template: ``` Glob: **/session-deep-dive/references/synthesis-template.md ``` Read the `--compare` report (or latest insights file). Read `MEMORY.md` for known findings. If 3+ sessions: read `summaries/consolidated.md` (NOT individual reports). If 1-2 sessions: read individual reports directly. Produce synthesis comparing: - New findings vs known patterns from MEMORY.md - Confirmed patterns (seen before, still present) - New patterns (not in previous reports) - Resolved patterns (previously noted, no new occurrences) ### Step 7: Update Ledger Use Python to safely update `metrics.jsonl` — never manually read/modify/rewrite in the LLM context: ```bash python3 -c " import json ids = {SESSION_IDS_SET} # e.g., {'ffa155ee-...', '90a74843-...'} lines = open('{PROJECT_ROOT}/.claude/session-metrics/metrics.jsonl').readlines() with open('{PROJECT_ROOT}/.claude/session-metrics/metrics.jsonl', 'w') as f: for line in lines: entry = json.loads(line) if entry.get('session_id') in ids: entry['tier2_completed'] = True f.write(json.dumps(entry) + '\n') " ``` ### Step 8: Write Output Write synthesis to `.claude/session-analysis/insights-{date}.md` Present key findings directly in conversation. Tell user: > Full report: `.claude/session-analysis/insights-{date}.md` > Per-session reports: `.claude/session-analysis/{id}-report.md` ## Output Files | File | Purpose | |------|---------| | `.claude/session-analysis/{id}-transcript.md` | Raw transcript | | `.claude/session-analysis/{id}-report.md` | Per-session analysis | | `.claude/session-analysis/summaries/consolidated.md` | Compressed reports | | `.claude/session-analysis/insights-{date}.md` | Cross-session synthesis | ## Iron Laws 1. **ONE ccrider call = ONE subagent** — never batch multiple fetches 2. **NEVER fetch or analyze in main context** — always subagents 3. **Absolute paths in subagent prompts** — subagents don't inherit skill context 4. **Python for jsonl updates** — never manually rewrite in LLM context 5. **ALWAYS pass pre-computed metrics to analysis subagents** — don't re-derive 6. **NEVER skip synthesis** — cross-session patterns are the real value 7. **TAG evidence strength** — every finding must be STRONG/MODERATE/WEAK
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.