skmtc-retro-review
Aggregate friction log files across a time period to identify recurring patterns, classify each cluster by intervention type, produce a prioritized action plan with success criteria, and calculate convergence metrics. Complements `skmtc-retro` (which captures per-session signal)
Install
npx skills add https://github.com/skmtc/skmtc/tree/main/deno/docs/skills/skmtc-retro-review
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install skmtc-skmtc@llmmart
git clone https://github.com/skmtc/skmtc.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole skmtc/skmtc collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
SKMTC retro review
The friction log is a sensor. This skill is the actuator. Its job is to read accumulated observations, detect patterns the per-session view cannot see, classify each pattern by what kind of intervention would eliminate it, and produce an action plan the user can execute.
Without this skill, the friction log is a cemetery: observations accumulate but nothing systematically improves. The review closes the loop.
1. When to invoke
Scheduled cadence (recommended):
- Monthly: review all sessions in the prior month.
- Pre-release: audit open entries against the release's scope to catch unresolved blockers before they reach users.
On-demand:
- After 3+ sessions in the same feature area (e.g., "three sessions on enrichment config in a week — what's the systemic issue?").
- When the user suspects a pattern but can't see it from individual retros.
- After a significant skill or doc update: "did the change actually work?"
Skip if:
- Fewer than 3 session files have accumulated since the last review.
- The friction log is empty or contains only resolved entries.
2. Locating files
Friction log:
<skmtc-root>/skmtc/deno/docs/friction-log/
Review output:
<skmtc-root>/skmtc/deno/docs/friction-log/reviews/
Create the reviews/ subdirectory if it doesn't exist.
Review filename:
<YYYY-MM-DD>-review-<period>.md
Where <period> is a human-readable date range: 2026-05 (monthly),
2026-Q2 (quarterly), or 2026-05-12-to-05-14 (targeted).
Examples:
2026-05-15-review-2026-05.md2026-05-15-review-pre-v0.5.md
3. Reading the friction log
Which files to read
- Glob the friction log directory for
*.mdfiles, excludingCLAUDE.md,README.md,discrepancy-catalog.md, and anything inreviews/. - Filter by date: include only files whose
YYYY-MM-DDprefix falls within the review period. For a monthly review, include all files from that month plus any older files with open entries. - Always include files with open entries regardless of date — unresolved observations accumulate across periods and contribute to the leak metric.
What to extract from each file
For each retro file, extract:
- Session date and topic (from the filename and
# heading) - All entries: number, heading, severity tag, status
- Knowledge acquired rows (from
## Knowledge acquiredtables if present) - Version anchors on friction entries (which version was the pain observed against?)
- Resolution status:
open,resolved <date>,superseded,wontfix
Do not try to re-derive the root cause from entry bodies at this stage — extract the structured fields first, then read bodies only for entries you intend to cluster.
Reading efficiency
With many files, read the ## Index table first (it's a summary of the
whole file). Read entry bodies only for entries that survive initial
filtering. This avoids loading the full text of every historical file.
4. Clustering entries
Clustering is the hardest step and the most error-prone. The goal is to group entries that share a root cause, not just a surface topic. Two entries about "import registration" may have completely different root causes; two entries that look unrelated (e.g., "error message was unclear" and "I had to read the source to understand X") may share the root cause "API behavior is not legible from its interface."
Clustering rules
Group by root cause, not by topic. Ask: "Would fixing X also fix Y?" If yes, they belong in the same cluster. If not, keep them separate even if their surface topic is the same.
A cluster must have at least 2 entries from different sessions (or 1 entry classified as
[blocker]) to warrant action. Single-session single-occurrence friction may be genuinely incidental.One entry can belong to multiple clusters if it has multiple root causes. Prefer the most specific cluster.
Explicitly note your clustering hypothesis in the review. "I'm grouping these because I believe they share root cause X" — this lets the user push back if the grouping is wrong. Do not present clusters as facts.
Unclustered entries — entries that don't fit any cluster — are still worth listing. Single-occurrence blockers in particular get their own cluster even without a second instance, because their severity warrants attention.
Cluster description format (internal working note)
For each cluster before writing the review:
Cluster: <short name>
Root cause hypothesis: <one sentence>
Entries: <file>#<N>, <file>#<N>, ...
Severity range: <lowest> to <highest>
Resolved entries: N of total
5. Intervention taxonomy
For each cluster, classify by intervention type using the decision tree below. Work through the questions in order — higher-ranked interventions are more permanent than lower-ranked ones.
Decision tree (in priority order)
1. Can this be eliminated by invariant enforcement (tooling)?
This is the highest-leverage intervention. A mechanically-checked rule cannot be violated silently; documentation can.
Candidates:
- Rules that must hold for every generator (single-base, location-independence,
no cross-package peers) →
doctorcheck or bundle-time lint rule. - Rules about file structure or import shape → custom
deno lintplugin or pre-bundle validation step. - Invariants currently only documented in skills → consider promoting to runtime assertion or CLI warning.
If tooling is feasible: recommend a specific check (what it tests, what
error it produces, where it runs). Label intervention type: tooling.
2. Is this a footgun in the API design?
A footgun is an API that is easy to misuse in a way that compiles and runs
but produces wrong output. Examples: ImportNameArg object form producing
unexpected aliasing; .isRef() ? resolve() : schema ternary being
redundant because the non-ref variants also implement .resolve().
Footguns cannot be fully fixed by documentation — the fix is normalization, tightening the type to block invalid input, or a better runtime error.
If a footgun: recommend the specific API change (normalize the edge case,
restrict the type, add a runtime guard). Label: skmtc-code.
3. Is this a missing runtime capability (architectural limit)?
When the task genuinely cannot be done with the current API — not a
documentation gap, but a structural absence. Example: the
one-operation-to-many-forms blocker that required the operation-variant
axis in @skmtc/core@0.5.0.
These are the most expensive to fix but also have the highest impact.
If it's an architectural limit: flag for core roadmap with a description
of what the API surface should look like. Label: skmtc-core-feature.
4. Is this a wrong LLM prior about SKMTC behavior?
The LLM's training data includes many frameworks. It applies defaults from those frameworks to SKMTC. When the SKMTC behavior differs from the prior, friction occurs even when the behavior is correctly documented — because the LLM didn't reach for the docs.
Skill updates override priors at task time. They are most effective when they name the prior explicitly: "In TypeScript you would do X. In SKMTC you do Y instead, because Z."
If a wrong prior: add an operational principle to the relevant skill with
the pattern "In <other> you would... In SKMTC you... because...". Label:
skill-update.
5. Is this a knowledge gap in docs?
The agent had to discover correct behavior by trial, reading source, or asking the user — and the answer exists nowhere in docs or skills.
Sub-types:
- API reference gap: a method, shape, or option isn't documented. Fix: add to API reference.
- Principle gap: a design rule or philosophy isn't articulated. Fix: add to a how-to doc or concept doc.
- Discoverability gap: the doc exists but the agent couldn't find it. Fix: improve cross-referencing, add to agent-context surfacing, or move the doc closer to where agents look first.
Also consider adding the knowledge to skmtc agent-context output — if
an agent starts briefed with this fact, the friction never occurs. Label:
doc-update, agent-context, or doc-discoverability.
6. Is this a missing worked example?
Some principles are best taught by a concrete, working example rather than a rule. If the friction repeatedly occurs despite the rule being documented, the rule alone isn't working — add an example.
Examples live in test fixtures, an examples/ directory, or inline in
skill entries. The example must be a complete, real, runnable snippet —
not pseudocode. Label: example-addition.
7. Should this API or generator be removed?
Sometimes the right fix is deletion. If an API creates persistent footguns with no clean fix, or a generator's design is fundamentally incompatible with SKMTC's regeneration contract, removing it is higher leverage than continuing to document around it. The GraphQL thin-wrapper generators were deleted rather than fixed.
Recommend removal only when: (a) the API is a persistent source of friction
across multiple sessions, (b) a skill/doc fix has already been attempted and
didn't eliminate recurrence, and (c) there is an alternative approach.
Label: removal.
Intervention label reference
| Label | What it is |
|---|---|
tooling |
Doctor check, lint rule, bundle-time validation |
skmtc-code |
Bug fix or API normalization in @skmtc/core or a gen-* package |
skmtc-core-feature |
New capability needed in @skmtc/core |
skill-update |
New or modified operational principle in a skill |
doc-update |
New or expanded content in API reference or how-to docs |
agent-context |
Add fact to skmtc agent-context output |
doc-discoverability |
Improve cross-referencing or surfacing of existing doc |
example-addition |
Add a worked example to a test, examples dir, or skill |
removal |
Deprecate or delete the problematic API or generator |
6. Convergence metrics
Calculate these metrics for the review period and append a row to the metrics history table in the review file.
Friction Recurrence Rate (FRR)
The primary convergence signal.
FRR = recurrent entries / total entries (for the period)
A "recurrent entry" is one where the same root cause was logged in a prior
period and the prior instance has status open (no intervention was made
or the intervention didn't work).
- FRR trending down: interventions are working; new friction is mostly novel.
- FRR flat: interventions aren't landing. Check whether recommended actions were actually completed. Check whether the skill/doc change targeted the right root cause.
- FRR trending up: regression — something broke that was previously working, or a new footgun was introduced.
Severity distribution
Blocker% = blocker entries / total entries
Convergence looks like: Blocker% → 0, then Friction% → 0, leaving
only Polish% and eventually silence. If Blocker% is flat or rising,
the system is not converging at the structural level — code changes are
needed, not just docs.
Open entry accumulation
Total open entries (all time) = sum of unresolved entries across all files
This number should not grow unboundedly. If it grows faster than entries are resolved, the action loop is too slow. A target: open entries should halve within two review cycles after a recommendation is acted on.
Knowledge backlog
Knowledge backlog = count of "Knowledge acquired" rows across all retro
files whose doc implication is not "none" and whose
content has not yet been added to docs/skills
This measures how much extracted intelligence remains un-acted. A large backlog means the retro → doc pipeline is blocked somewhere.
Resolution velocity
Resolution velocity = average days from entry creation to resolved status
(for entries resolved in the review period)
Long velocity (30+ days) indicates the action loop is slow — typically because "recommended action" stays in a review doc without being assigned or prioritized. Short velocity (< 7 days) is ideal.
7. Review file format
# Friction Review — <period>
**Reviewed:** <YYYY-MM-DD>
**Period:** <date range>
**Sessions included:** N
**Files read:** <list of filenames>
## Summary
| Metric | Value | vs. prior period |
|--------|-------|-----------------|
| Total entries | N | +N / -N |
| Blocker% | X% | ↑ / ↓ / — |
| Friction Recurrence Rate | X% | ↑ / ↓ / — |
| Total open entries (all time) | N | +N / -N |
| Knowledge backlog | N items | +N / -N |
| Resolution velocity (period) | N days avg | |
## Convergence signal
<2-3 sentences. Is the system converging, flat, or diverging? What's the
dominant pattern? Be direct — "FRR is 60% and flat; skill updates are not
reaching the root cause" is more useful than "mixed results.">
## Clusters
| # | Cluster | Root cause hypothesis | Sessions | Severity | Open | Recommended intervention |
|---|---------|----------------------|----------|----------|------|--------------------------|
| C1 | <name> | <one sentence> | N | blocker/friction/polish | N/total | <label> |
## Action plan
### P1. <Intervention label> — <Cluster name>
**Root cause:** <one sentence>
**Evidence:** <file>#<N>, <file>#<N> (N total instances)
**What to do:** <specific, actionable change — not "improve docs" but "add
a note to §import-registration in the generator skill explaining that the
object form of ImportNameArg always produces `name as alias` output, even
when alias is omitted, and the bare string form is correct for non-type
non-aliased imports">
**Success criterion:** <falsifiable condition — "subsequent retro files
should contain zero entries about ImportNameArg object form misuse">
**Verification:** check the next 2-3 retro files after the change for
recurrence of this pattern
---
### P2. ...
## Knowledge backlog
Items from `## Knowledge acquired` tables not yet reflected in docs/skills:
| Source | Item | Doc implication | Priority |
|--------|------|-----------------|----------|
| <file> K1 | <what was learned> | <skill/doc/agent-context> | high/medium/low |
*If backlog is empty: "All knowledge-acquired items from this period are
reflected in current docs and skills."*
## Metrics history
<!-- append one row per review; do not delete prior rows -->
| Review date | Period | Sessions | FRR | Blocker% | Open (all time) | Knowledge backlog | Actions completed |
|-------------|--------|----------|-----|----------|-----------------|-------------------|-------------------|
| <YYYY-MM-DD> | <period> | N | X% | X% | N | N | N of N prior |
Metrics history
The ## Metrics history table accumulates across reviews — append one row
per review cycle, never delete prior rows. This is the convergence record.
It is the only place where the trend question ("is it getting better?") can
be answered empirically rather than by impression.
If no prior review file exists, create the table with one row. Future reviews append to the same table in the same file, or to a new review file that cross-references the prior one with a link.
Action plan ordering
Order the action plan by:
- Intervention type permanence:
tooling>skmtc-code>skmtc-core-feature>skill-update>doc-update>example-addition - Within the same type: frequency × severity. A cluster of 4
[friction]entries outranks 1[friction]entry; a[blocker]outranks multiple[friction]even at the same type level. - Recurrent clusters (FRR contributors) rank above first-occurrence clusters of the same severity — recurrence means a prior attempt didn't work.
Action plan specificity requirement
Every action plan entry must name:
- The exact file/section to change (not just "update the skill")
- The exact content to add or change (not just "clarify this")
- A falsifiable success criterion
- How to verify it worked (what to look for in the next retro files)
Vague recommendations ("improve documentation of X") are not actionable and
will not close the loop. Specific recommendations ("add the following
paragraph to §3 of skmtc-generator SKILL.md, under the register API
section:...") can be completed in minutes.
8. Prior action follow-up
Before drafting new action items, check the most recent prior review file (if one exists) and answer:
- What was recommended? List the prior P1, P2, P3 items.
- What was completed? For each, check whether the recommended change appears in the skill, doc, or code (read the target file or run a grep).
- Did it work? For completed items, check whether the friction pattern recurred in sessions after the change. If yes: the intervention targeted the wrong root cause — escalate to a higher-permanence intervention.
- What was not completed? List items that were recommended but not acted on. These carry forward to the new action plan with increased priority — a recommendation that's been skipped once needs a specific owner or a lower-effort formulation.
Record this follow-up as the first section after the summary in the new review file.
9. Composing the review
Full flow:
- Glob the friction log for files in scope. List them explicitly.
- Read the Index tables of each file. Extract entry numbers, headings, severities, statuses. Do not read bodies yet.
- Filter: separate open from resolved entries. Separate the review period entries from the all-time-open entries carried forward.
- Read bodies only for open entries you intend to cluster — skip resolved entries unless investigating whether a fix worked.
- Check prior review if one exists: list what was recommended, what was completed, what recurred.
- Cluster open entries by root cause. State each clustering hypothesis explicitly. Assign intervention type per cluster using the decision tree.
- Calculate metrics: FRR, Blocker%, total open, knowledge backlog, resolution velocity.
- Compile the knowledge backlog: extract
## Knowledge acquiredrows from the period's retro files. Check each against current docs/skills to see if it's already been addressed. - Write the review file: summary → convergence signal → prior action follow-up → clusters → action plan → knowledge backlog → metrics history.
- Summarise to the user in one short message:
Review written to <filename>. N clusters, top intervention: <P1 label — cluster name>. FRR: X% (prior: Y%). <convergence signal sentence>.
10. After the review
The user decides which action plan items to execute. The review skill does
not execute them — it produces decisions. Each action item is a unit of
work for the relevant skill (skmtc-generator, skmtc-cli, docs editing,
or filing a SKMTC code issue).
When an action item is completed, update the corresponding retro file
entries' **Status:** lines and the ## Index rows. Also update the
review file's metrics history row for the period (if the resolution
happened within the same period).
The next review cycle opens with §8 "Prior action follow-up" to verify that completed items actually eliminated the friction. This is the feedback loop that drives convergence: observe → cluster → intervene → verify → observe.
The convergence guarantee: the system converges when every friction cluster, on recurrence, triggers escalation to a higher-permanence intervention. Friction that doesn't yield to docs must be taken to skill-update. Skill-updates that don't eliminate recurrence must be taken to code. Code changes that are blocked must be taken to the roadmap. Clusters that cannot be addressed at any level are candidates for removal of the offending API or feature. With this escalation discipline, every class of friction either gets eliminated or gets explicitly accepted as a known cost — neither outcome leaves the system in an unexamined drift state.
Files (skmtc)
-
commands
-
skmtc-retro-review.md 3.6 KB
--- description: Aggregate SKMTC friction log entries into a review — cluster patterns, classify interventions, calculate convergence metrics, produce an action plan argument-hint: "[period]" --- Run a SKMTC friction log review. **If `$ARGUMENTS` is non-empty**, use it as the period description. Accepted forms: - `2026-05` — a specific month - `2026-Q2` — a quarter - `pre-v0.6` — a named milestone - `2026-05-01-to-05-14` — an explicit date range - `since-last-review` — all session files created after the most recent file in `friction-log/reviews/` **If `$ARGUMENTS` is empty**, default to the current calendar month. Then follow the standard review flow (full details in the `skmtc-retro-review` skill's `SKILL.md`): 1. **Locate the friction-log directory.** Walk up from CWD looking for `skmtc/deno/docs/friction-log/`. If not found, ask — do not silently default. 2. **Check for a prior review file** in `friction-log/reviews/`: - If one exists: read its action plan items and metrics history row to prepare the prior-action follow-up section. - If none: note that this is the first review. 3. **Glob and filter session files** for the period. Always include files with open entries regardless of date (carried-forward friction). Exclude `README.md`, `CLAUDE.md`, `discrepancy-catalog.md`, and anything in `reviews/`. 4. **Extract structured data** from each file's `## Index` table (severities, statuses) and `## Knowledge acquired` table if present. Read entry bodies only for entries you intend to cluster. 5. **Check prior action items** — for each P1/P2/P3 from the last review: verify whether the recommended change was made (read the target file or grep), and whether the friction pattern recurred in sessions after the change. 6. **Cluster** open entries by root cause (not surface topic). State each clustering hypothesis explicitly. Use the intervention decision tree (tooling > skmtc-code > skmtc-core-feature > skill-update > doc-update > agent-context > example-addition > removal) to classify each cluster. 7. **Calculate convergence metrics:** - FRR = recurrent entries / total entries - Blocker% = blocker entries / total entries - Total open entries (all time) - Knowledge backlog = `## Knowledge acquired` items not yet in docs/skills - Resolution velocity = avg days from creation to resolved (for entries resolved this period) 8. **Write the review file** to `friction-log/reviews/` named `<YYYY-MM-DD>-review-<period>.md`. Include in order: summary metrics table → convergence signal → prior action follow-up → cluster table → action plan (most specific action first, with falsifiable success criteria) → knowledge backlog → metrics history (append one row). 9. **Summarise to the user** in one short message: `Review written to <filename>. N clusters, top intervention: <P1 label — cluster name>. FRR: X% (prior: Y%). <1-sentence convergence signal>.` **Action plan specificity requirement:** every recommended action must name the exact file/section to change and the exact content to add — not "improve documentation of X" but the precise paragraph or principle. Each item needs a falsifiable success criterion and a verification method (what to check in next retro files). If the session files contain no open entries beyond what prior reviews have already addressed, **say so explicitly** rather than fabricating clusters. An empty review is better than false signal. For full conventions, clustering rules, intervention taxonomy, convergence metric definitions, and worked examples, defer to the `skmtc-retro-review` skill's `SKILL.md`.
-
-
design.md 8.3 KB
# skmtc-retro-review skill — design document > The loadable skill is [`SKILL.md`](SKILL.md); the slash command is > [`commands/skmtc-retro-review.md`](commands/skmtc-retro-review.md). ## Purpose Turn accumulated per-session observations into decisions. The `skmtc-retro` skill is the sensor — it captures what went wrong session by session. The `skmtc-retro-review` skill is the actuator — it aggregates observations across sessions, identifies systemic patterns, classifies each pattern by the best intervention type, and produces a prioritized action plan with falsifiable success criteria. Without this skill, the friction log is a cemetery: observations accumulate but nothing systematically improves. The review is where "is it getting better?" becomes empirically answerable. ## Audience The user (Dmitri), running the review periodically. Not intended for invocation during active SKMTC work — it's a reflection-and-decision tool run at session boundaries or on a monthly/pre-release cadence. ## Triggers - "review retros" / "review friction" - "run a retro review" - "what's the pattern across sessions" - "what should we fix next" - Slash command: `/skmtc-retro-review [period]` - Monthly cadence, or before a release ## Relationship to skmtc-retro | | skmtc-retro | skmtc-retro-review | |---|---|---| | **Timescale** | Per session | Monthly / pre-release | | **Input** | Session content | Friction log files | | **Output** | Session file in `friction-log/` | Review file in `friction-log/reviews/` | | **Primary question** | What happened in this session? | Is the system getting better? | | **Action** | Captures observations | Produces decisions | The two skills form a closed loop: retro → friction log → review → action plan → interventions → reduced future friction → fewer retro entries. ## Scope boundary ### In skill - Reading strategy for the friction log (Index-first, body on demand) - Clustering rules: group by root cause, not surface topic; state hypothesis explicitly; 2+ entries required for a cluster - The intervention decision tree and all 9 intervention types - Convergence metrics: FRR, Blocker%, open entry accumulation, knowledge backlog, resolution velocity - Review file format (summary → prior follow-up → clusters → action plan → knowledge backlog → metrics history) - Specificity requirement for action plan entries - Prior-action follow-up procedure (verify completion + recurrence) ### Deferred - Per-session entry format: [`skmtc-retro SKILL.md`](../skmtc-retro/SKILL.md) - Friction log file conventions: [`../../friction-log/README.md`](../../friction-log/README.md) - Operational principles: [`../../llms.md`](../../llms.md) ## Design decisions ### The convergence guarantee via escalation discipline The core guarantee: no friction cluster can stay at the same intervention level indefinitely. If a skill update doesn't eliminate recurrence, the next review escalates to a code change. If a code change is roadmap-blocked, the API becomes a removal candidate. This escalation discipline is what makes convergence structural rather than aspirational. ### Intervention taxonomy ordered by permanence Nine intervention types, ordered tooling > code > feature > skill > doc > agent-context > discoverability > example > removal. The ordering reflects permanence: tooling enforces mechanically; documentation only informs. A finding that could be addressed by tooling should not be addressed by documentation alone — it would recur. The "something else altogether" types that go beyond code/doc/skill: - **`tooling`**: invariant enforcement via doctor checks or lint rules. The single-base rule, location-independence, no-cross-package-peers are currently only documented. Tooling would catch violations before they enter the friction log. - **`agent-context`**: the knowledge-acquired tables in retros are a queue of facts to add to `skmtc agent-context` output. Fixing the briefing upstream eliminates the friction at source rather than documenting around it after the fact. - **`removal`**: sometimes deletion is higher-leverage than continued documentation. The GraphQL thin-wrapper generators were removed rather than fixed. The decision tree includes removal as a legitimate option after lower-permanence interventions have failed. ### Friction Recurrence Rate as the primary metric FRR = recurrent entries / total entries per period. The choice of FRR as the headline metric reflects the core failure mode: friction that's been logged before but not fixed. A low FRR means novel friction (new API surface, new use case). A high FRR means interventions aren't landing. The trend matters more than the absolute value. ### Clustering is a hypothesis, not a fact The skill requires the agent to state clustering hypotheses explicitly ("I'm grouping these because I believe they share root cause X") rather than presenting clusters as objective groupings. This keeps the user in the loop and surfaces the agent's reasoning so it can be corrected. Two entries about the same surface topic may have completely different root causes; false grouping leads to interventions that don't work. ### No category-of-fix tagging at retro time Inherited from `skmtc-retro`: "Possible fixes" in individual entries are left open-ended. The review is where intervention classification happens, using the full decision tree and cross-session context. This separates observation from prescription and prevents the first obvious fix from being locked in before the pattern is visible. ### Prior action follow-up as a mandatory step Every review opens by checking what was recommended last time and whether it worked. This is what distinguishes a system that converges from one that just accumulates recommendations. Without this step, the review becomes another layer of friction documentation rather than an actuator. ### Specificity requirement for action items Vague recommendations ("improve documentation of X") don't close the loop. The skill requires action items to name the exact file and section, the exact content, a falsifiable success criterion, and a verification method. This makes each item completable in minutes and verifiable in the next review cycle. ## Open design questions ### Tagging for cross-file pattern matching Currently, clustering relies on the reviewing agent reading multiple files and grouping by semantic similarity. A lightweight tag system (e.g., `[import-api]`, `[enrichment]`, `[naming]`) on entries would make grep-based clustering possible and more reliable. Cost: adds per-entry maintenance during retro authoring. Worth considering after 15+ friction files accumulate. ### Integration with agent-context command The skill recommends `agent-context` as an intervention type but doesn't currently guide *how* to add to it — what the command reads, what format its output takes, where that content lives. A follow-up review of `skmtc agent-context` source would complete this loop. ### Automated recurrence detection Currently recurrence is detected manually by the reviewing agent reading prior files. A grep-based check for key phrases or entry headings across all files would be faster and more reliable. Could be a doctor subcommand: `skmtc doctor friction-recurrence`. ### Review file accumulation Over time, the `reviews/` directory will grow. Each review's metrics history table is the running record — cross-review trends require reading multiple review files. A single `metrics.md` that aggregates all reviews' metric rows would make the trend view available without loading every review file. Defer until 5+ review files exist. ### Wins balance monitoring The retro skill now requires a high bar for `[win]` entries (codification candidates only). If the log becomes dominated entirely by friction, the signal that "this approach works and should be prescribed" disappears. The review should monitor: are wins appearing at all? If not, the bar may be too high, or everything worth codifying has been codified. ## Cross-references - Skill: [`SKILL.md`](SKILL.md) - Slash command: [`commands/skmtc-retro-review.md`](commands/skmtc-retro-review.md) - Feeds from: [`skmtc-retro skill`](../skmtc-retro/SKILL.md) - Friction log: [`../../friction-log/README.md`](../../friction-log/README.md) - Review output: [`../../friction-log/reviews/`](../../friction-log/reviews/) - LLM doc: [`../../llms.md`](../../llms.md) -
SKILL.md 20.9 KB
--- name: skmtc-retro-review version: 0.1.0 description: | Aggregate friction log files across a time period to identify recurring patterns, classify each cluster by intervention type, produce a prioritized action plan with success criteria, and calculate convergence metrics. Complements `skmtc-retro` (which captures per-session signal) by acting as the system's actuator: turning accumulated observations into decisions. The primary output is a review document that makes the "is it getting better?" question answerable. Use this skill when the user asks to "review retros", "review friction", "run a retro review", "what's the pattern across sessions", "what should we fix next", or on a periodic cadence (monthly, pre-release). Also run after a cluster of sessions on the same feature area to surface systemic issues before they compound. Distinct from `skmtc-retro` — that skill captures per-session observations; this skill synthesizes them into decisions. Do not run this skill as a substitute for a per-session retro. allowed-tools: - Read - Write - Edit - Glob - Grep - Bash metadata: internal: true --- # SKMTC retro review The friction log is a sensor. This skill is the actuator. Its job is to read accumulated observations, detect patterns the per-session view cannot see, classify each pattern by what kind of intervention would eliminate it, and produce an action plan the user can execute. Without this skill, the friction log is a cemetery: observations accumulate but nothing systematically improves. The review closes the loop. ## 1. When to invoke **Scheduled cadence (recommended):** - Monthly: review all sessions in the prior month. - Pre-release: audit open entries against the release's scope to catch unresolved blockers before they reach users. **On-demand:** - After 3+ sessions in the same feature area (e.g., "three sessions on enrichment config in a week — what's the systemic issue?"). - When the user suspects a pattern but can't see it from individual retros. - After a significant skill or doc update: "did the change actually work?" **Skip if:** - Fewer than 3 session files have accumulated since the last review. - The friction log is empty or contains only resolved entries. ## 2. Locating files **Friction log:** ``` <skmtc-root>/skmtc/deno/docs/friction-log/ ``` **Review output:** ``` <skmtc-root>/skmtc/deno/docs/friction-log/reviews/ ``` Create the `reviews/` subdirectory if it doesn't exist. **Review filename:** ``` <YYYY-MM-DD>-review-<period>.md ``` Where `<period>` is a human-readable date range: `2026-05` (monthly), `2026-Q2` (quarterly), or `2026-05-12-to-05-14` (targeted). Examples: - `2026-05-15-review-2026-05.md` - `2026-05-15-review-pre-v0.5.md` ## 3. Reading the friction log ### Which files to read 1. **Glob** the friction log directory for `*.md` files, excluding `CLAUDE.md`, `README.md`, `discrepancy-catalog.md`, and anything in `reviews/`. 2. **Filter by date**: include only files whose `YYYY-MM-DD` prefix falls within the review period. For a monthly review, include all files from that month plus any older files with open entries. 3. **Always include files with open entries** regardless of date — unresolved observations accumulate across periods and contribute to the leak metric. ### What to extract from each file For each retro file, extract: - **Session date and topic** (from the filename and `# heading`) - **All entries**: number, heading, severity tag, status - **Knowledge acquired rows** (from `## Knowledge acquired` tables if present) - **Version anchors** on friction entries (which version was the pain observed against?) - **Resolution status**: `open`, `resolved <date>`, `superseded`, `wontfix` Do not try to re-derive the root cause from entry bodies at this stage — extract the structured fields first, then read bodies only for entries you intend to cluster. ### Reading efficiency With many files, read the `## Index` table first (it's a summary of the whole file). Read entry bodies only for entries that survive initial filtering. This avoids loading the full text of every historical file. ## 4. Clustering entries Clustering is the hardest step and the most error-prone. The goal is to group entries that share a **root cause**, not just a surface topic. Two entries about "import registration" may have completely different root causes; two entries that look unrelated (e.g., "error message was unclear" and "I had to read the source to understand X") may share the root cause "API behavior is not legible from its interface." ### Clustering rules 1. **Group by root cause, not by topic.** Ask: "Would fixing X also fix Y?" If yes, they belong in the same cluster. If not, keep them separate even if their surface topic is the same. 2. **A cluster must have at least 2 entries** from different sessions (or 1 entry classified as `[blocker]`) to warrant action. Single-session single-occurrence friction may be genuinely incidental. 3. **One entry can belong to multiple clusters** if it has multiple root causes. Prefer the most specific cluster. 4. **Explicitly note your clustering hypothesis** in the review. "I'm grouping these because I believe they share root cause X" — this lets the user push back if the grouping is wrong. Do not present clusters as facts. 5. **Unclustered entries** — entries that don't fit any cluster — are still worth listing. Single-occurrence blockers in particular get their own cluster even without a second instance, because their severity warrants attention. ### Cluster description format (internal working note) For each cluster before writing the review: ``` Cluster: <short name> Root cause hypothesis: <one sentence> Entries: <file>#<N>, <file>#<N>, ... Severity range: <lowest> to <highest> Resolved entries: N of total ``` ## 5. Intervention taxonomy For each cluster, classify by intervention type using the decision tree below. Work through the questions in order — higher-ranked interventions are more permanent than lower-ranked ones. ### Decision tree (in priority order) **1. Can this be eliminated by invariant enforcement (tooling)?** This is the highest-leverage intervention. A mechanically-checked rule cannot be violated silently; documentation can. Candidates: - Rules that must hold for every generator (single-base, location-independence, no cross-package peers) → `doctor` check or bundle-time lint rule. - Rules about file structure or import shape → custom `deno lint` plugin or pre-bundle validation step. - Invariants currently only documented in skills → consider promoting to runtime assertion or CLI warning. If tooling is feasible: recommend a specific check (what it tests, what error it produces, where it runs). Label intervention type: `tooling`. **2. Is this a footgun in the API design?** A footgun is an API that is easy to misuse in a way that compiles and runs but produces wrong output. Examples: `ImportNameArg` object form producing unexpected aliasing; `.isRef() ? resolve() : schema` ternary being redundant because the non-ref variants also implement `.resolve()`. Footguns cannot be fully fixed by documentation — the fix is normalization, tightening the type to block invalid input, or a better runtime error. If a footgun: recommend the specific API change (normalize the edge case, restrict the type, add a runtime guard). Label: `skmtc-code`. **3. Is this a missing runtime capability (architectural limit)?** When the task genuinely cannot be done with the current API — not a documentation gap, but a structural absence. Example: the one-operation-to-many-forms blocker that required the operation-variant axis in `@skmtc/core@0.5.0`. These are the most expensive to fix but also have the highest impact. If it's an architectural limit: flag for core roadmap with a description of what the API surface should look like. Label: `skmtc-core-feature`. **4. Is this a wrong LLM prior about SKMTC behavior?** The LLM's training data includes many frameworks. It applies defaults from those frameworks to SKMTC. When the SKMTC behavior differs from the prior, friction occurs even when the behavior is correctly documented — because the LLM didn't reach for the docs. Skill updates override priors at task time. They are most effective when they name the prior explicitly: "In TypeScript you would do X. In SKMTC you do Y instead, because Z." If a wrong prior: add an operational principle to the relevant skill with the pattern `"In <other> you would... In SKMTC you... because..."`. Label: `skill-update`. **5. Is this a knowledge gap in docs?** The agent had to discover correct behavior by trial, reading source, or asking the user — and the answer exists nowhere in docs or skills. Sub-types: - **API reference gap**: a method, shape, or option isn't documented. Fix: add to API reference. - **Principle gap**: a design rule or philosophy isn't articulated. Fix: add to a how-to doc or concept doc. - **Discoverability gap**: the doc exists but the agent couldn't find it. Fix: improve cross-referencing, add to agent-context surfacing, or move the doc closer to where agents look first. Also consider adding the knowledge to `skmtc agent-context` output — if an agent starts briefed with this fact, the friction never occurs. Label: `doc-update`, `agent-context`, or `doc-discoverability`. **6. Is this a missing worked example?** Some principles are best taught by a concrete, working example rather than a rule. If the friction repeatedly occurs despite the rule being documented, the rule alone isn't working — add an example. Examples live in test fixtures, an `examples/` directory, or inline in skill entries. The example must be a complete, real, runnable snippet — not pseudocode. Label: `example-addition`. **7. Should this API or generator be removed?** Sometimes the right fix is deletion. If an API creates persistent footguns with no clean fix, or a generator's design is fundamentally incompatible with SKMTC's regeneration contract, removing it is higher leverage than continuing to document around it. The GraphQL thin-wrapper generators were deleted rather than fixed. Recommend removal only when: (a) the API is a persistent source of friction across multiple sessions, (b) a skill/doc fix has already been attempted and didn't eliminate recurrence, and (c) there is an alternative approach. Label: `removal`. ### Intervention label reference | Label | What it is | |-------|------------| | `tooling` | Doctor check, lint rule, bundle-time validation | | `skmtc-code` | Bug fix or API normalization in `@skmtc/core` or a gen-* package | | `skmtc-core-feature` | New capability needed in `@skmtc/core` | | `skill-update` | New or modified operational principle in a skill | | `doc-update` | New or expanded content in API reference or how-to docs | | `agent-context` | Add fact to `skmtc agent-context` output | | `doc-discoverability` | Improve cross-referencing or surfacing of existing doc | | `example-addition` | Add a worked example to a test, examples dir, or skill | | `removal` | Deprecate or delete the problematic API or generator | ## 6. Convergence metrics Calculate these metrics for the review period and append a row to the metrics history table in the review file. ### Friction Recurrence Rate (FRR) The primary convergence signal. ``` FRR = recurrent entries / total entries (for the period) ``` A "recurrent entry" is one where the same root cause was logged in a prior period and the prior instance has status `open` (no intervention was made or the intervention didn't work). - **FRR trending down**: interventions are working; new friction is mostly novel. - **FRR flat**: interventions aren't landing. Check whether recommended actions were actually completed. Check whether the skill/doc change targeted the right root cause. - **FRR trending up**: regression — something broke that was previously working, or a new footgun was introduced. ### Severity distribution ``` Blocker% = blocker entries / total entries ``` Convergence looks like: `Blocker% → 0`, then `Friction% → 0`, leaving only `Polish%` and eventually silence. If `Blocker%` is flat or rising, the system is not converging at the structural level — code changes are needed, not just docs. ### Open entry accumulation ``` Total open entries (all time) = sum of unresolved entries across all files ``` This number should not grow unboundedly. If it grows faster than entries are resolved, the action loop is too slow. A target: open entries should halve within two review cycles after a recommendation is acted on. ### Knowledge backlog ``` Knowledge backlog = count of "Knowledge acquired" rows across all retro files whose doc implication is not "none" and whose content has not yet been added to docs/skills ``` This measures how much extracted intelligence remains un-acted. A large backlog means the retro → doc pipeline is blocked somewhere. ### Resolution velocity ``` Resolution velocity = average days from entry creation to resolved status (for entries resolved in the review period) ``` Long velocity (30+ days) indicates the action loop is slow — typically because "recommended action" stays in a review doc without being assigned or prioritized. Short velocity (< 7 days) is ideal. ## 7. Review file format ```markdown # Friction Review — <period> **Reviewed:** <YYYY-MM-DD> **Period:** <date range> **Sessions included:** N **Files read:** <list of filenames> ## Summary | Metric | Value | vs. prior period | |--------|-------|-----------------| | Total entries | N | +N / -N | | Blocker% | X% | ↑ / ↓ / — | | Friction Recurrence Rate | X% | ↑ / ↓ / — | | Total open entries (all time) | N | +N / -N | | Knowledge backlog | N items | +N / -N | | Resolution velocity (period) | N days avg | | ## Convergence signal <2-3 sentences. Is the system converging, flat, or diverging? What's the dominant pattern? Be direct — "FRR is 60% and flat; skill updates are not reaching the root cause" is more useful than "mixed results."> ## Clusters | # | Cluster | Root cause hypothesis | Sessions | Severity | Open | Recommended intervention | |---|---------|----------------------|----------|----------|------|--------------------------| | C1 | <name> | <one sentence> | N | blocker/friction/polish | N/total | <label> | ## Action plan ### P1. <Intervention label> — <Cluster name> **Root cause:** <one sentence> **Evidence:** <file>#<N>, <file>#<N> (N total instances) **What to do:** <specific, actionable change — not "improve docs" but "add a note to §import-registration in the generator skill explaining that the object form of ImportNameArg always produces `name as alias` output, even when alias is omitted, and the bare string form is correct for non-type non-aliased imports"> **Success criterion:** <falsifiable condition — "subsequent retro files should contain zero entries about ImportNameArg object form misuse"> **Verification:** check the next 2-3 retro files after the change for recurrence of this pattern --- ### P2. ... ## Knowledge backlog Items from `## Knowledge acquired` tables not yet reflected in docs/skills: | Source | Item | Doc implication | Priority | |--------|------|-----------------|----------| | <file> K1 | <what was learned> | <skill/doc/agent-context> | high/medium/low | *If backlog is empty: "All knowledge-acquired items from this period are reflected in current docs and skills."* ## Metrics history <!-- append one row per review; do not delete prior rows --> | Review date | Period | Sessions | FRR | Blocker% | Open (all time) | Knowledge backlog | Actions completed | |-------------|--------|----------|-----|----------|-----------------|-------------------|-------------------| | <YYYY-MM-DD> | <period> | N | X% | X% | N | N | N of N prior | ``` ### Metrics history The `## Metrics history` table accumulates across reviews — append one row per review cycle, never delete prior rows. This is the convergence record. It is the only place where the trend question ("is it getting better?") can be answered empirically rather than by impression. If no prior review file exists, create the table with one row. Future reviews append to the same table in the same file, or to a new review file that cross-references the prior one with a link. ### Action plan ordering Order the action plan by: 1. Intervention type permanence: `tooling` > `skmtc-code` > `skmtc-core-feature` > `skill-update` > `doc-update` > `example-addition` 2. Within the same type: frequency × severity. A cluster of 4 `[friction]` entries outranks 1 `[friction]` entry; a `[blocker]` outranks multiple `[friction]` even at the same type level. 3. Recurrent clusters (FRR contributors) rank above first-occurrence clusters of the same severity — recurrence means a prior attempt didn't work. ### Action plan specificity requirement Every action plan entry must name: - The exact file/section to change (not just "update the skill") - The exact content to add or change (not just "clarify this") - A falsifiable success criterion - How to verify it worked (what to look for in the next retro files) Vague recommendations ("improve documentation of X") are not actionable and will not close the loop. Specific recommendations ("add the following paragraph to §3 of skmtc-generator SKILL.md, under the `register` API section:...") can be completed in minutes. ## 8. Prior action follow-up Before drafting new action items, check the most recent prior review file (if one exists) and answer: 1. **What was recommended?** List the prior P1, P2, P3 items. 2. **What was completed?** For each, check whether the recommended change appears in the skill, doc, or code (read the target file or run a grep). 3. **Did it work?** For completed items, check whether the friction pattern recurred in sessions after the change. If yes: the intervention targeted the wrong root cause — escalate to a higher-permanence intervention. 4. **What was not completed?** List items that were recommended but not acted on. These carry forward to the new action plan with increased priority — a recommendation that's been skipped once needs a specific owner or a lower-effort formulation. Record this follow-up as the first section after the summary in the new review file. ## 9. Composing the review Full flow: 1. **Glob** the friction log for files in scope. List them explicitly. 2. **Read the Index tables** of each file. Extract entry numbers, headings, severities, statuses. Do not read bodies yet. 3. **Filter**: separate open from resolved entries. Separate the review period entries from the all-time-open entries carried forward. 4. **Read bodies** only for open entries you intend to cluster — skip resolved entries unless investigating whether a fix worked. 5. **Check prior review** if one exists: list what was recommended, what was completed, what recurred. 6. **Cluster** open entries by root cause. State each clustering hypothesis explicitly. Assign intervention type per cluster using the decision tree. 7. **Calculate metrics**: FRR, Blocker%, total open, knowledge backlog, resolution velocity. 8. **Compile the knowledge backlog**: extract `## Knowledge acquired` rows from the period's retro files. Check each against current docs/skills to see if it's already been addressed. 9. **Write the review file**: summary → convergence signal → prior action follow-up → clusters → action plan → knowledge backlog → metrics history. 10. **Summarise to the user** in one short message: `Review written to <filename>. N clusters, top intervention: <P1 label — cluster name>. FRR: X% (prior: Y%). <convergence signal sentence>.` ## 10. After the review The user decides which action plan items to execute. The review skill does not execute them — it produces decisions. Each action item is a unit of work for the relevant skill (`skmtc-generator`, `skmtc-cli`, docs editing, or filing a SKMTC code issue). When an action item is completed, update the corresponding retro file entries' `**Status:**` lines and the `## Index` rows. Also update the review file's metrics history row for the period (if the resolution happened within the same period). The next review cycle opens with §8 "Prior action follow-up" to verify that completed items actually eliminated the friction. This is the feedback loop that drives convergence: observe → cluster → intervene → verify → observe. **The convergence guarantee**: the system converges when every friction cluster, on recurrence, triggers escalation to a higher-permanence intervention. Friction that doesn't yield to docs must be taken to skill-update. Skill-updates that don't eliminate recurrence must be taken to code. Code changes that are blocked must be taken to the roadmap. Clusters that cannot be addressed at any level are candidates for removal of the offending API or feature. With this escalation discipline, every class of friction either gets eliminated or gets explicitly accepted as a known cost — neither outcome leaves the system in an unexamined drift state.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.