Claude Skill

show-me-your-work

Keep a reviewable decision trail for long-running or unattended work: a TSV log with one row per decision (what, why, evidence, result). Local by default; commit it when a reviewer needs the trail to trust the result. Use for /show-me-your-work, autonomous or multi-phase runs, or

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download michael-denyer-pstack-claude-plugins_pstack_skills_show-me-your-work-4b3933e.zip · 4 KB
Part of michael-denyer/pstack-claude — 51 skills

Install

skills CLI npx skills add https://github.com/michael-denyer/pstack-claude/tree/main/plugins/pstack/skills/show-me-your-work
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install michael-denyer-pstack-claude@llmmart
Git git clone https://github.com/michael-denyer/pstack-claude.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole michael-denyer/pstack-claude collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Show me your work

Keep one canonical log.

The format

A single TSV file, one row per decision. Cells stay single-line. Evidence is a pointer, not prose.

Copy references/decision-log-template.tsv (the header row) to start a clean log. Columns:

  • ts. ISO8601 timestamp.
  • phase. The phase or workstream.
  • decision. What was chosen or done, one line.
  • why. The reason in plain words. If a principle drove it, say it plainly, not as a jargon tag.
  • evidence. A link or path that proves it: commit SHA, PR number, file:line, or an artifact, trace, or screenshot path. Never a paragraph.
  • result. The outcome or predicate state: tests green, reverted, pixel-diff 0, INCONCLUSIVE, open.

An example, plain-spoken so a reviewer reads it at a glance.

ts	phase	decision	why	evidence	result
2026-05-24T09:02:00Z	frame	counted the work first, about 100 components and roughly 75 hours	wanted to know the size before starting a long run	commit 3a9f1c2	found 5 things to sort out before starting
2026-05-24T09:40:00Z	harness	took screenshots of the old version before changing anything	so we can compare old against new and catch any visual change	scripts/snapshot.sh, baseline/	saved 120 reference screenshots
2026-05-24T11:15:00Z	widget	moved the widget styles over without changing how it looks	keep the change small and the result identical	commit 7c21e0a, pixel-diff 0	looks identical, tests pass
2026-05-24T12:30:00Z	widget	threw out a helper's work because its screenshots were blank	checked the real files instead of trusting its summary	worktree reset	reverted, tightened the instructions for next time

Logging a row

Write each entry the way you'd tell a teammate what you did. Plain words, concrete actions, no AI speak or abstract jargon (the unslop skill applies to log text too).

Use the helper scripts/log.sh <logfile> <phase> <decision> <why> <evidence> <result>. It stamps ts, writes the header on first use, strips stray tabs/newlines, and prefixes any cell starting with =, +, -, @, or " with a single quote. A bare printf appending a row works too, but mind those same bytes if cells come from generated or user-supplied text.

Log decision points and checkpoints, not every action: a fork chosen, a unit completed with its verification result, a pivot or revert with its trigger, a blocker surfaced, a gate fixed. For loop runs, one row per iteration. Skip the trivial and self-evident.

A run is one agent conversation, including its later turns and any summary of it. A pickup, a replacement agent, or a new chat starts a new run. When a run adds to a log that already has rows, its first row has phase start, and so does its first row after another run's start row. So a run that comes back to a log in a later turn first reads the log's last rows to see whether another run wrote since. A start row names the ts range of the rows before it that this run did not write, and its evidence names this run, such as its agent id. Use phase start for nothing else.

Where it lives

By default the log is a working artifact, not committed. Keep it at decisions.tsv in the work dir, or .audit/<task-slug>.tsv when several efforts run at once, and leave it out of git.

Commit it only when the work is ambitious enough that a reviewer needs the trail to trust the result.

Rules

  • Append-only. A wrong call gets a new row that supersedes it. Never edit or delete history.
  • Prefer evidence produced by committed scripts over hand-made one-offs (the encode-lessons-in-structure principle skill).

Audit the log against the transcript

At the end of the run, before handing back, check the log told the truth. Read this run's transcript under Claude Code's per-project transcripts directory at ~/.claude/projects/<encoded-cwd>/. Don't glob across ~/.claude/projects/. That reads unrelated private chats. Walk this run's rows against what actually happened. Each stretch of them begins at one of this run's start rows, or at the first row if this run created the log, and ends at the next start row of another run:

  • Check that every row maps to a real decision or action.
  • Check that each row's evidence resolves and shows what the row claims.
  • A fork, pivot, or abandoned approach that shaped the work but isn't logged is a gap. Add it.

Correct the log, not the story. The audit never edits or removes a row, even an invented one. When a row records neither a real decision nor a real action, or its claim or evidence is wrong, add a row that supersedes it with what actually happened and a pointer that resolves. This audit does not check rows outside this run's stretches. If this run's own work shows one of them is wrong, supersede it like any wrong call.

Cross-model review of the trail

Before handing back, spawn a subagent on a different model family from the one that did the work. Self-review is not a substitute. The subagent reads the audit trail and the run's transcript, then flags what the user should pay attention to. Not a redo of the work, a scan for what's suboptimal or risky.

  • Decisions logged with weak or absent evidence.
  • Verification steps skipped or claimed without proof in the transcript.
  • Choices that look risky in hindsight (premature, scope-creeping, papering over a symptom).
  • Gaps the user would otherwise miss on a casual skim.

Every reply for a run that produced a trail ends with an "Attention" section. Lead with the reviewer's model on its own line (reviewed by <model>), then list each flag pointing to specific rows or moments. "No flags" is a valid value. The model name is not.

Reviewing the trail

Read top to bottom, follow the evidence pointers, spot-check. GitHub renders a committed TSV as a table. column -s$'\t' -t decisions.tsv renders it in a terminal.

Composing this skill

Other skills route their audit trail here instead of inventing one. Reference it by name and let it own the format. Don't restate the columns.

Files (pstack-claude)
  • references
    • decision-log-template.tsv 38 B · in bundle
  • scripts
    • log.sh 1.4 KB
      #!/usr/bin/env bash
      # Append a well-formed row to a show-me-your-work decision log (TSV).
      # Usage: log.sh <logfile> <phase> <decision> <why> <evidence> <result>
      set -euo pipefail
      
      if [ "$#" -ne 6 ]; then
      	printf 'usage: log.sh <logfile> <phase> <decision> <why> <evidence> <result>\n' >&2
      	exit 1
      fi
      
      logfile="$1"
      shift
      
      logdir="$(dirname "$logfile")"
      if [ -n "$logdir" ] && [ "$logdir" != "." ] && [ ! -d "$logdir" ]; then
      	mkdir -p "$logdir"
      fi
      
      # Use `>>` here, never `>`. A network mount can fail this test for a log
      # that exists. Then the cost is one stray header line, not the rows.
      if [ ! -s "$logfile" ]; then
      	printf 'ts\tphase\tdecision\twhy\tevidence\tresult\n' >> "$logfile"
      fi
      
      ts="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
      # Strip tabs/newlines/CR so cells stay on one line, and prefix any cell
      # whose first char a spreadsheet would parse as a formula (=, +, -, @)
      # or a TSV reader as an opening field quote (") with a single quote.
      # The skill expects this log to be read in
      # spreadsheets, so attacker-controlled evidence (PR titles, filenames,
      # generated text) must not become formula execution when a reviewer
      # opens the file.
      clean() {
      	local v
      	v=$(printf '%s' "$1" | tr '\t\n\r' '   ')
      	case "$v" in
      		=*|+*|-*|@*|\"*) printf "'%s" "$v" ;;
      		*) printf '%s' "$v" ;;
      	esac
      }
      printf '%s\t%s\t%s\t%s\t%s\t%s\n' \
      	"$ts" "$(clean "$1")" "$(clean "$2")" "$(clean "$3")" "$(clean "$4")" "$(clean "$5")" \
      	>> "$logfile"
      
  • SKILL.md 6.2 KB
    ---
    name: show-me-your-work
    description: "Keep a reviewable decision trail for long-running or unattended work: a TSV log with one row per decision (what, why, evidence, result). Local by default; commit it when a reviewer needs the trail to trust the result. Use for /show-me-your-work, autonomous or multi-phase runs, or work a human reviews after stepping away."
    ---
    
    # Show me your work
    
    Keep one canonical log.
    
    ## The format
    
    A single TSV file, one row per decision. Cells stay single-line. Evidence is a pointer, not prose.
    
    Copy `references/decision-log-template.tsv` (the header row) to start a clean log. Columns:
    
    - **ts.** ISO8601 timestamp.
    - **phase.** The phase or workstream.
    - **decision.** What was chosen or done, one line.
    - **why.** The reason in plain words. If a principle drove it, say it plainly, not as a jargon tag.
    - **evidence.** A link or path that proves it: commit SHA, PR number, `file:line`, or an artifact, trace, or screenshot path. Never a paragraph.
    - **result.** The outcome or predicate state: `tests green`, `reverted`, `pixel-diff 0`, `INCONCLUSIVE`, `open`.
    
    An example, plain-spoken so a reviewer reads it at a glance.
    
    ```
    ts	phase	decision	why	evidence	result
    2026-05-24T09:02:00Z	frame	counted the work first, about 100 components and roughly 75 hours	wanted to know the size before starting a long run	commit 3a9f1c2	found 5 things to sort out before starting
    2026-05-24T09:40:00Z	harness	took screenshots of the old version before changing anything	so we can compare old against new and catch any visual change	scripts/snapshot.sh, baseline/	saved 120 reference screenshots
    2026-05-24T11:15:00Z	widget	moved the widget styles over without changing how it looks	keep the change small and the result identical	commit 7c21e0a, pixel-diff 0	looks identical, tests pass
    2026-05-24T12:30:00Z	widget	threw out a helper's work because its screenshots were blank	checked the real files instead of trusting its summary	worktree reset	reverted, tightened the instructions for next time
    ```
    
    ## Logging a row
    
    Write each entry the way you'd tell a teammate what you did. Plain words, concrete actions, no AI speak or abstract jargon (the **unslop** skill applies to log text too).
    
    Use the helper `scripts/log.sh <logfile> <phase> <decision> <why> <evidence> <result>`. It stamps `ts`, writes the header on first use, strips stray tabs/newlines, and prefixes any cell starting with `=`, `+`, `-`, `@`, or `"` with a single quote. A bare `printf` appending a row works too, but mind those same bytes if cells come from generated or user-supplied text.
    
    Log decision points and checkpoints, not every action: a fork chosen, a unit completed with its verification result, a pivot or revert with its trigger, a blocker surfaced, a gate fixed. For loop runs, one row per iteration. Skip the trivial and self-evident.
    
    A run is one agent conversation, including its later turns and any summary of it. A pickup, a replacement agent, or a new chat starts a new run. When a run adds to a log that already has rows, its first row has phase `start`, and so does its first row after another run's `start` row. So a run that comes back to a log in a later turn first reads the log's last rows to see whether another run wrote since. A `start` row names the `ts` range of the rows before it that this run did not write, and its evidence names this run, such as its agent id. Use phase `start` for nothing else.
    
    ## Where it lives
    
    By default the log is a working artifact, not committed. Keep it at `decisions.tsv` in the work dir, or `.audit/<task-slug>.tsv` when several efforts run at once, and leave it out of git.
    
    Commit it only when the work is ambitious enough that a reviewer needs the trail to trust the result.
    
    ## Rules
    
    - Append-only. A wrong call gets a new row that supersedes it. Never edit or delete history.
    - Prefer evidence produced by committed scripts over hand-made one-offs (the **encode-lessons-in-structure** principle skill).
    
    ## Audit the log against the transcript
    
    At the end of the run, before handing back, check the log told the truth. Read this run's transcript under Claude Code's per-project transcripts directory at `~/.claude/projects/<encoded-cwd>/`. Don't glob across `~/.claude/projects/`. That reads unrelated private chats. Walk this run's rows against what actually happened. Each stretch of them begins at one of this run's `start` rows, or at the first row if this run created the log, and ends at the next `start` row of another run:
    
    - Check that every row maps to a real decision or action.
    - Check that each row's evidence resolves and shows what the row claims.
    - A fork, pivot, or abandoned approach that shaped the work but isn't logged is a gap. Add it.
    
    Correct the log, not the story. The audit never edits or removes a row, even an invented one. When a row records neither a real decision nor a real action, or its claim or evidence is wrong, add a row that supersedes it with what actually happened and a pointer that resolves. This audit does not check rows outside this run's stretches. If this run's own work shows one of them is wrong, supersede it like any wrong call.
    
    ## Cross-model review of the trail
    
    Before handing back, spawn a subagent on a different model family from the one that did the work. Self-review is not a substitute. The subagent reads the audit trail and the run's transcript, then flags what the user should pay attention to. Not a redo of the work, a scan for what's suboptimal or risky.
    
    - Decisions logged with weak or absent evidence.
    - Verification steps skipped or claimed without proof in the transcript.
    - Choices that look risky in hindsight (premature, scope-creeping, papering over a symptom).
    - Gaps the user would otherwise miss on a casual skim.
    
    Every reply for a run that produced a trail ends with an "Attention" section. Lead with the reviewer's model on its own line (`reviewed by <model>`), then list each flag pointing to specific rows or moments. "No flags" is a valid value. The model name is not.
    
    ## Reviewing the trail
    
    Read top to bottom, follow the evidence pointers, spot-check. GitHub renders a committed TSV as a table. `column -s$'\t' -t decisions.tsv` renders it in a terminal.
    
    ## Composing this skill
    
    Other skills route their audit trail here instead of inventing one. Reference it by name and let it own the format. Don't restate the columns.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related