Claude Skill

agents-introspection

Imported from paulrberg/agent-skills/skills/agents-introspection.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download paulrberg-agent-skills-skills_agents-introspection-913232a.zip · 23 KB
Part of paulrberg/agent-skills — 42 skills

Install

skills CLI npx skills add https://github.com/PaulRBerg/agent-skills/tree/main/skills/agents-introspection
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install paulrberg-agent-skills@llmmart
Git git clone https://github.com/PaulRBerg/agent-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole paulrberg/agent-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Agents Introspection

This skill is coordination-exempt: skip the ai-coord gate for its declared work.

Supported Chat Hosts

Before doing any work, identify the current chat host. If it is not Claude Code or Codex CLI, stop with this error: This skill only works in Claude Code or Codex CLI.

If these instructions are already present in the conversation from a slash or dollar invocation, follow them directly; do not invoke this skill again through a skill tool.

Determine whether prior Codex and Claude Code work in the current project establishes a recurrence risk for the user's task, then recommend the smallest durable intervention justified by the evidence.

Success means the report identifies transcript coverage, separates observed behavior from inference, applies a consistent evidence bar, and either proposes a concrete prevention step or explains why no durable change is justified.

Input

  • <task> (required): the task, decision, incident, or workflow to evaluate. If omitted but the current conversation states it clearly, use that task; otherwise ask for the missing task.

Scope and Authority

  • Inspect Codex and Claude Code transcripts whose metadata or cwd resolves to the current project and any other local project materially relevant to the user's task. You are authorized to read relevant other-project sessions without asking the user. Establish relevance from task context, explicit project or path references, a shared change or workflow, or session metadata; do not scan another project's history solely because it shares a basename or keyword.
  • Treat current-project transcripts as internal working evidence. Include direct excerpts when they materially improve the report; do not summarize or redact solely because the model provider can see them. Never expose credentials, security secrets, or personal wallet addresses.
  • Inspect and report by default. Edit AGENTS.md or skills only when the user explicitly asks to apply or implement fixes, then make the smallest in-scope local change and validate it.
  • Never modify transcript stores. Before placing transcript evidence in a public or third-party artifact, perform an external-disclosure review and remove unrelated personal or customer data, unsuitable private paths or repository names, and unrelated transcript material. Perform external writes only when authorized.

Bounded Retrieval

Read references/transcript-sources.md, resolve the current project with pwd -P, identify any other task-relevant local projects, and choose 3–6 short, discriminative keywords from relevant filenames, commands, tools, errors, package names, issue IDs, and skill names.

In a Codex read-only sandbox, or whenever uv cannot write its cache, skip the helper and go directly to the Manual Fallback below; do not retry uv run.

  1. Run the bundled miner for the current project and each task-relevant local project, unarchived sessions only, with the chosen keywords, --since 60d, --excerpts, and --max-sessions 8. Encode synonyms as one OR-group keyword (--keyword 'a|b') rather than separate --keyword flags.
  2. Treat source ownership as a miner invariant: every candidate is assigned once from source-native cwd, directory, and history metadata, never from transcript content. Review ownership; reject a candidate only when fallback ownership such as turn_context.cwd remains materially ambiguous for the task. The miner excludes conflicting ownership evidence and the live session by default.
  3. Treat miner scores, themes, correction, failure, verification, tool, or privacy_gaps counts, redacted excerpts, and modified timestamps only as candidate-ranking and triage signals. They are heuristic and are never evidence by themselves. Inspect up to five highest-relevance transcript bodies through the bundled inspector digest (scripts/transcript-inspect.py) first; open raw bodies only when the digest is insufficient, stopping earlier when the evidence bar is met. Include a comparable successful session when available.
  4. If evidence is insufficient, retry once with broader or OR-grouped keywords. If still weak, retry once with --since removed. If unarchived history still lacks signal, retry once with --include-archived.
  5. Exceed these bounds only to resolve contradictory evidence or satisfy an explicitly exhaustive request. If the helper fails, use one project-scoped manual fallback from the reference.

Stop and report the coverage gap when the bounded fallbacks still lack useful evidence. Absence of evidence is not evidence that a failure never occurred.

For unusually long searches, send sparse progress updates only when a retrieval fallback begins, a finding changes the likely intervention, or the search reaches its explicit bound. Use an outcome-first line such as 🔎 Broadening transcript search — <verified reason and bound> or ⏳ Checking archived sessions — <verified unarchived/archived coverage>. Ground counts and coverage claims in miner/tool output; do not narrate routine transcript reads.

Evidence Contract

Evaluate historical behavior against the AGENTS.md and skill instructions available to that session when recoverable. Do not infer that an agent ignored or misapplied a rule solely because the current source tree or installed copies under ~/.agents/skills or ~/.claude/skills contain it. If the transcript does not establish the historical version or availability, mark it unknown and qualify the attribution.

For each relevant session, record only concise, auditable observations about:

  • ignored or misread AGENTS.md or skill instructions;
  • wrong cwd, project root, source path, or path encoding;
  • over-broad edits, unrelated churn, overwritten user work, or destructive commands;
  • tooling, shell, parsing, retry, or verification mistakes;
  • invented claims, vague reports, or missing tests and checks;
  • successful patterns that prevented mistakes.

Connect each observation to the current task and label any causal or recurrence claim as inference. State conflicts and weak coverage rather than averaging them away.

Use these confidence levels:

  • High: at least two independent relevant sessions directly support the same pattern and its relevance to the current task.
  • Medium: one unambiguous relevant session supports the finding, or multiple sessions provide mixed support.
  • Low: only indirect, ambiguous, or heuristic signals exist. Do not recommend a durable repository change from low-confidence evidence.

A durable change requires either the same failure in at least two independent sessions or one unambiguous high-impact failure that exposes a missing stable invariant. Treat lower-impact one-offs as manual guardrails.

Choose the Smallest Intervention

  • Update AGENTS.md when the lesson is stable, project-wide, and useful to agents working in that scope.
  • Update an existing skill when the failure belongs clearly inside its current workflow.
  • Propose a new skill only for a repeated, reusable procedure that spans projects or repositories.
  • Add a script only when deterministic discovery, parsing, or validation would otherwise be reimplemented.
  • Recommend no durable change for one-off mistakes, weak evidence, or rules already stated clearly; report the risk and manual guardrail instead.

Report and Stop

Lead with ### 🔎 Introspection complete — <intervention or coverage-gap outcome> for read-only work or ### ✅ Introspection fixes applied — <outcome> when explicitly requested fixes were written, then report only:

  1. 🗂 Historical coverage: project paths, sources checked, fallbacks used, and sessions inspected.
  2. 🔎 Findings: a compact table with confidence, observed evidence, inference, relevance, and intervention. Keep confidence visibly separate from severity or impact.
  3. 🛡 Durable recommendations: apply now, consider later, or no change, with the target and prevention mechanism.
  4. 🧪 Validation and gaps: commands run, checks performed, external-disclosure constraints, and missing evidence.

When fixes were explicitly requested, include exact files changed and validation outcomes. Stop after the current task has an evidence-backed recommendation or an explicit coverage gap; do not mine additional history merely to add examples or strengthen prose. Keep transcript references, paths, counters, redactions, and miner JSON exact and undecorated.

Files (agent-skills)
  • agents
    • openai.yaml 42 B
      policy:
        allow_implicit_invocation: true
      
  • references
    • transcript-sources.md 9.8 KB
      # Transcript Sources
      
      Use the bundled miner to rank Codex and Claude Code sessions for the current project and any other local project
      materially relevant to the user's task before opening transcript bodies. Reading relevant other-project sessions is
      authorized; treat its output as heuristic discovery data, not evidence.
      
      ## Resolve Local Paths
      
      Resolve paths once and reuse them:
      
      ```sh
      project_path="$(pwd -P)"
      home_dir="$(cd ~ && pwd -P)"
      claude_config_dir="${CLAUDE_CONFIG_DIR:-$home_dir/.claude}"
      codex_home="${CODEX_HOME:-$home_dir/.codex}"
      skill_dir="${AGENTS_INTROSPECTION_SKILL_DIR:-}"
      if [ -z "$skill_dir" ]; then
        for candidate in "$home_dir/.agents/skills/agents-introspection" "$home_dir/.claude/skills/agents-introspection"; do
          if [ -f "$candidate/scripts/transcript-miner.py" ]; then
            skill_dir="$candidate"
            break
          fi
        done
      fi
      transcript_miner="$skill_dir/scripts/transcript-miner.py"
      test -f "$transcript_miner" || {
        printf '%s\n' "missing agents-introspection transcript miner" >&2
        exit 1
      }
      ```
      
      When the skill host exposes its installation directory directly, resolve `scripts/transcript-miner.py` from that
      directory instead of searching installed copies.
      
      ## Preferred Helper
      
      Run one unarchived-session pass with 3–6 task keywords. Include the current project and every task-relevant local
      project as a repeated `--project` argument:
      
      ```sh
      uv run "$transcript_miner" \
        --project "$project_path" \
        --keyword '<keyword-1>' \
        --keyword '<synonym-a>|<synonym-b>' \
        --since 60d \
        --excerpts \
        --max-sessions 8 \
        --format json
      ```
      
      `--keyword` accepts `|`-separated OR-groups (e.g. `--keyword 'miner|mining|transcript-miner'`) to encode synonyms as one
      group instead of separate flags.
      
      Include another project without requesting permission when task context, an explicit project or path reference, a shared
      change or workflow, or session metadata establishes relevance. Never infer relevance from a shared basename or keyword
      alone.
      
      The helper returns project coverage, ranked candidate sessions, task themes, correction and failure signals,
      verification signals, tool-call counts, and `privacy_gaps` categories. It always redacts common secret-like values.
      Transcript excerpts are emitted only with `--excerpts`: up to 3 redacted, 240-character-truncated `{channel, text}`
      entries per candidate, drawn from the first user message plus up to 2 keyword-matching messages (preferring user over
      assistant). Scores and counts select candidates only; validate every reported finding against the relevant transcript
      body. `keyword_hits` counts eligible `(message, keyword)` pairs keyed by the full OR-group keyword string, not repeated
      substring occurrences.
      
      With `--since <YYYY-MM-DD|Nd>`, the report gains a top-level `since` object (`value`, `cutoff`, `codex_dirs_pruned`,
      `codex_files_pruned`, `claude_files_pruned`); it is `null` without the flag. Each candidate carries a `modified` ISO
      mtime; sessions modified within 7 days score higher than those within 30 days, which score higher than older sessions —
      another ranking signal, not evidence.
      
      Source ownership is structural and precedes relevance scoring:
      
      - Codex uses `session_meta.payload.cwd`; when absent, it accepts sampled `turn_context.cwd` values only when all resolve
        under one requested project.
      - Claude reconciles the encoded project directory, top-level transcript `cwd`, and `history.jsonl.project` when a
        history record exists. Conflicts are excluded rather than guessed.
      - A cwd equal to or below multiple requested roots belongs to the longest, most-specific root. A transcript is emitted
        at most once, and project strings in messages, context, tool inputs, or tool outputs never establish ownership.
      
      The live `CODEX_THREAD_ID` or `CLAUDE_CODE_SESSION_ID` transcript is excluded by default. Use `--include-current` only
      when diagnosing the miner or intentionally inspecting the active session.
      
      Candidate signals use delineated channels. `user` is actual task text, preferring Claude history `display`; `assistant`
      is plain assistant message text; injected AGENTS, skill, environment, permission, collaboration, abort, and command
      envelopes are ignored `context`; and `tool` contributes names plus structured error status or nonzero exit codes only.
      Serialized tool inputs and raw outputs never contribute keywords or behavioral regex signals. Identical eligible
      messages are deduplicated within each channel.
      
      Each project coverage record retains `codex_candidates`, `claude_candidates`, and `selected_sessions`.
      `codex_candidates` and `claude_candidates` mean structurally owned sessions, including sessions that did not meet the
      keyword relevance requirement. Additional fields make selection and exclusions auditable:
      
      - `codex_scanned`, `claude_scanned`, `structurally_matched`, and `relevance_matched` describe source coverage;
      - `current_sessions_excluded`, `content_only_project_mentions_ignored`, and `ambiguous_ownership_excluded` count
        distinct session files, not occurrences;
      - candidate `ownership` records `matched_via`, canonical `cwd`, and assigned `project`;
      - candidate `signal_channels` records eligible user and assistant message counts, ignored context, and structured tool
        failures.
      
      If the first pass is weak, make one pass with broader or OR-grouped keywords. If still weak, retry once with `--since`
      removed. Add `--include-archived` only for the final bounded fallback. Empty output after those passes is a coverage
      gap, not proof that no relevant behavior exists.
      
      ## Inspect Candidates
      
      Use the bundled inspector to get a bounded, redacted digest of a candidate transcript before opening its raw body:
      
      ```sh
      uv run "$skill_dir/scripts/transcript-inspect.py" <transcript-path>... \
        --keyword '<keyword-1>' \
        --max-entries 120 \
        --format text
      ```
      
      For each file it emits a header (source, session id, cwd, timestamp range, per-channel totals, sampled flag) and bounded
      entries with absolute record line numbers: every non-context user message, keyword-, correction-, or
      verification-matching assistant messages, and tool failures. Redaction is always on; entry text is capped at 240
      characters.
      
      Digests are redacted and bounded — inspect them before reading raw bodies, and read raw bodies only when the digest is
      insufficient. Each entry's line number lets you pull the exact underlying record when needed:
      
      ```sh
      sed -n '<line+1>p' <transcript-path> | jq
      ```
      
      ## Source Layouts
      
      Claude Code uses `CLAUDE_CONFIG_DIR` when set and otherwise defaults to `~/.claude`. Project transcripts normally live
      under:
      
      ```text
      <claude-config>/projects/<absolute-path-with-nonalphanumerics-replaced-by-dashes>/
      ```
      
      The helper parses `<claude-config>/history.jsonl` once, indexing source-native `project`, `sessionId`, and user-authored
      `display` fields. History relevance preselects sessions before their transcript bodies are sampled, so an older relevant
      session remains discoverable behind any number of newer irrelevant files. Sessions without history records use bounded
      head/tail transcript sampling. The helper also checks the legacy slash-only project encoding and excludes directory,
      cwd, or history disagreements.
      
      Codex uses `CODEX_HOME`, defaulting to `~/.codex`:
      
      - Unarchived transcripts: `sessions/`
      - Archived transcripts: `archived_sessions/`
      - Recent-session index: `session_index.jsonl`
      
      Session JSONL commonly contains `session_meta`, `turn_context`, `event_msg`, and `response_item` records. Prefer
      JSON-aware inspection of the smallest relevant record range to keep retrieval bounded and high-signal.
      
      ## Manual Fallback
      
      Use this only when the helper is missing or fails. Preserve the same relevant-project and retrieval bounds, secret
      handling, and external-disclosure boundary.
      
      1. For Claude Code, compute the encoded directory from the exact absolute project path and inspect newest JSONL files
         there.
      2. For Codex, search unarchived transcript metadata for the exact absolute project path. Search archives only after the
         unarchived pass is insufficient.
      3. If exact matching is suspiciously empty, use the repository basename only to identify candidates, then reject every
         candidate whose source-native metadata or cwd does not resolve to the current or a task-relevant project. Never use a
         content occurrence as ownership evidence.
      4. Filter candidates with the task keywords before opening bodies. Inspect at most five unless evidence conflicts or the
         user requested exhaustive coverage.
      
      Prefer `rg`, `fd`, `jq`, and structured parsing. If a command returns partial output or errors, try one equivalent
      scoped command before reporting the gap; never compensate by searching unrelated project history.
      
      Transcript JSONL embeds tool output as JSON strings, so quotes inside that content appear escaped in raw text
      (`\"key\":\"value\"`). When grepping raw transcript files, allow optional backslashes in the pattern (e.g. `\\?"`) or
      decode lines with `jq` before matching; a pattern written for decoded JSON will silently miss raw-text matches. This
      caveat mainly matters when bypassing the inspector and grepping raw bodies directly; the inspector's own digest output
      is plain decoded text.
      
      ## Secret Handling and External Disclosure
      
      - Use direct transcript evidence when it materially strengthens an internal report; keep excerpts bounded and relevant.
      - Always redact credentials such as API keys, private keys, mnemonics, tokens, and passwords. Never expose personal
        wallet addresses. Before public or third-party disclosure, also remove emails, unrelated personal or customer data,
        unsuitable private paths or repository names, and unrelated transcript material.
      - Write transcript content to durable repository artifacts only when the task authorizes it and the evidence materially
        belongs there. Perform an external-disclosure review before posting, publishing, uploading, or otherwise sending the
        artifact outside the agent workspace.
      - Include raw transcript paths in the report only when they materially help the user audit a finding.
      
  • scripts
    • transcript-inspect.py 9.4 KB
      #!/usr/bin/env -S uv run --script
      # /// script
      # requires-python = ">=3.12"
      # ///
      """Bounded, always-redacted digest of Codex/Claude Code transcript JSONL files."""
      
      from __future__ import annotations
      
      import argparse
      import json
      import os
      import sys
      from dataclasses import asdict, dataclass, field
      from pathlib import Path
      from typing import Any
      
      from transcript_common import (
          CORRECTION_PATTERNS,
          MAX_FULL_SESSION_BYTES,
          VERIFICATION_PATTERNS,
          extract_record_channels,
          extract_tool_failures,
          first_string_shallow,
          keyword_alternatives,
          read_jsonl_sampled_with_lines,
          redact_text,
          truncate,
      )
      
      
      CODEX_RECORD_TYPES = {"session_meta", "turn_context", "response_item"}
      ENTRY_TEXT_LIMIT = 240
      DEFAULT_MAX_ENTRIES = 120
      
      
      @dataclass
      class Header:
          path: str
          source: str
          session_id: str
          cwd: str | None
          first_timestamp: str | None
          last_timestamp: str | None
          user_messages: int
          assistant_messages: int
          ignored_context_messages: int
          tool_failures: int
          sampled: bool
      
      
      @dataclass
      class Entry:
          line: int
          channel: str
          tool: str | None
          status: str | None
          text: str
      
      
      @dataclass
      class FileDigest:
          path: str
          error: str | None
          header: Header | None
          entries: list[Entry] = field(default_factory=list)
          omitted_entries: int = 0
      
          def to_json(self) -> dict[str, Any]:
              return {
                  "path": self.path,
                  "error": self.error,
                  "header": asdict(self.header) if self.header else None,
                  "entries": [asdict(entry) for entry in self.entries],
                  "omitted_entries": self.omitted_entries,
              }
      
      
      def main() -> int:
          parser = argparse.ArgumentParser(
              description="Inspect Codex/Claude Code transcript JSONL files with bounded, redacted digests."
          )
          parser.add_argument("paths", nargs="+", help="Transcript JSONL file paths")
          parser.add_argument(
              "--keyword",
              action="append",
              default=[],
              help="Keyword to match. Repeatable. Use 'a|b|c' for an OR-group of alternatives",
          )
          parser.add_argument("--max-entries", type=int, default=DEFAULT_MAX_ENTRIES, help="Maximum emitted entries per file")
          parser.add_argument("--format", choices=("text", "json"), default="text")
          args = parser.parse_args()
      
          if args.max_entries < 1:
              print("transcript-inspect: --max-entries must be positive", file=sys.stderr)
              return 2
      
          keywords = [keyword for keyword in args.keyword if keyword.strip()]
          digests = [inspect_file(raw_path, keywords, args.max_entries) for raw_path in args.paths]
      
          if args.format == "json":
              print(json.dumps({"files": [digest.to_json() for digest in digests]}, indent=2, sort_keys=True))
          else:
              print_text_report(digests)
      
          return 1 if digests and all(digest.error is not None for digest in digests) else 0
      
      
      def inspect_file(raw_path: str, keywords: list[str], max_entries: int) -> FileDigest:
          path = Path(os.path.expanduser(raw_path))
          if not path.exists():
              return FileDigest(raw_path, f"path does not exist: {path}", None)
          if not path.is_file():
              return FileDigest(raw_path, f"not a file: {path}", None)
          try:
              sampled = path.stat().st_size > MAX_FULL_SESSION_BYTES
          except OSError as error:
              return FileDigest(raw_path, f"cannot stat file: {error}", None)
      
          records = list(read_jsonl_sampled_with_lines(path))
          if not records:
              return FileDigest(raw_path, "no parsable JSONL records (missing, empty, or unreadable)", None)
      
          items = [item for _, item in records]
          source = guess_source(items)
          session_id = extract_session_id(items, path)
          cwd = extract_cwd(items)
      
          timestamps: list[str] = []
          user_all: list[tuple[int, str]] = []
          assistant_all: list[tuple[int, str]] = []
          failure_all: list[tuple[int, dict[str, str | None]]] = []
          context_count = 0
      
          for line, item in records:
              timestamp = first_string_shallow(item, ("timestamp", "created_at", "updated_at"))
              if timestamp:
                  timestamps.append(timestamp)
              channels = extract_record_channels(item)
              user_all.extend((line, text) for text in channels["user"])
              assistant_all.extend((line, text) for text in channels["assistant"])
              context_count += len(channels["context"])
              failure_all.extend((line, detail) for detail in extract_tool_failures(item))
      
          remaining_budget = max(0, max_entries - len(user_all) - len(failure_all))
          if not keywords and len(assistant_all) <= remaining_budget:
              assistant_selected = assistant_all
          else:
              assistant_selected = [
                  (line, text) for line, text in assistant_all if assistant_message_qualifies(text, keywords)
              ]
      
          combined: list[tuple[int, int, Entry]] = []
          for line, text in user_all:
              combined.append((line, 0, Entry(line, "user", None, None, truncate(redact_text(text), ENTRY_TEXT_LIMIT))))
          for line, text in assistant_selected:
              combined.append((line, 1, Entry(line, "assistant", None, None, truncate(redact_text(text), ENTRY_TEXT_LIMIT))))
          for line, detail in failure_all:
              combined.append(
                  (
                      line,
                      2,
                      Entry(
                          line,
                          "tool_failure",
                          detail.get("tool"),
                          redact_text(detail["status"]) if detail.get("status") else None,
                          truncate(redact_text(detail.get("text") or ""), ENTRY_TEXT_LIMIT),
                      ),
                  )
              )
          combined.sort(key=lambda entry: (entry[0], entry[1]))
      
          entries = [entry for _, _, entry in combined[:max_entries]]
          omitted = max(0, len(combined) - max_entries)
      
          header = Header(
              path=str(path),
              source=source,
              session_id=session_id,
              cwd=cwd,
              first_timestamp=timestamps[0] if timestamps else None,
              last_timestamp=timestamps[-1] if timestamps else None,
              user_messages=len(user_all),
              assistant_messages=len(assistant_all),
              ignored_context_messages=context_count,
              tool_failures=len(failure_all),
              sampled=sampled,
          )
          return FileDigest(raw_path, None, header, entries, omitted)
      
      
      def assistant_message_qualifies(text: str, keywords: list[str]) -> bool:
          if keyword_group_matches(text, keywords):
              return True
          return any(pattern.search(text) for pattern in (*CORRECTION_PATTERNS.values(), *VERIFICATION_PATTERNS.values()))
      
      
      def keyword_group_matches(text: str, keywords: list[str]) -> bool:
          if not keywords:
              return False
          lowered = text.lower()
          for keyword in keywords:
              alternatives = [alternative.lower() for alternative in keyword_alternatives(keyword)]
              if alternatives and any(alternative in lowered for alternative in alternatives):
                  return True
          return False
      
      
      def guess_source(items: list[Any]) -> str:
          for item in items:
              if isinstance(item, dict) and item.get("type") in CODEX_RECORD_TYPES:
                  return "codex"
          for item in items:
              if isinstance(item, dict) and ("sessionId" in item or "cwd" in item):
                  return "claude"
          return "unknown"
      
      
      def extract_session_id(items: list[Any], path: Path) -> str:
          for item in items:
              if not isinstance(item, dict):
                  continue
              payload = item.get("payload") if isinstance(item.get("payload"), dict) else None
              if item.get("type") == "session_meta" and isinstance(payload, dict):
                  value = first_string_shallow(payload, ("id", "session_id", "conversation_id"))
                  if value:
                      return value
              value = first_string_shallow(item, ("sessionId", "session_id", "conversation_id"))
              if value:
                  return value
          return path.stem
      
      
      def extract_cwd(items: list[Any]) -> str | None:
          for item in items:
              if not isinstance(item, dict):
                  continue
              payload = item.get("payload") if isinstance(item.get("payload"), dict) else None
              if isinstance(payload, dict):
                  value = first_string_shallow(payload, ("cwd",))
                  if value:
                      return value
              value = first_string_shallow(item, ("cwd",))
              if value:
                  return value
          return None
      
      
      def print_text_report(digests: list[FileDigest]) -> None:
          for index, digest in enumerate(digests):
              if index:
                  print()
              print(f"== {digest.path} ==")
              if digest.error or digest.header is None:
                  print(f"ERROR: {digest.error}")
                  continue
              header = digest.header
              print(f"source={header.source} session={header.session_id} cwd={header.cwd or '-'} sampled={header.sampled}")
              print(
                  f"first={header.first_timestamp or '-'} last={header.last_timestamp or '-'} "
                  f"user={header.user_messages} assistant={header.assistant_messages} "
                  f"context-ignored={header.ignored_context_messages} tool-failures={header.tool_failures}"
              )
              for entry in digest.entries:
                  if entry.channel == "tool_failure":
                      tool = entry.tool or "unknown-tool"
                      status = entry.status or "unknown-status"
                      print(f"[{entry.line}] tool_failure ({tool}, {status}): {entry.text}")
                  else:
                      print(f"[{entry.line}] {entry.channel}: {entry.text}")
              if digest.omitted_entries:
                  print(f"... {digest.omitted_entries} entries omitted")
      
      
      if __name__ == "__main__":
          raise SystemExit(main())
      
    • transcript-miner.py 35.7 KB
      #!/usr/bin/env -S uv run --script
      # /// script
      # requires-python = ">=3.12"
      # ///
      """Mine project-owned Codex and Claude Code transcripts, optionally with redacted excerpts."""
      
      from __future__ import annotations
      
      import argparse
      import datetime as dt
      import json
      import os
      import re
      import sys
      from collections import Counter, defaultdict
      from dataclasses import asdict, dataclass, field
      from pathlib import Path
      from typing import Any, Iterable
      
      from transcript_common import (
          CORRECTION_PATTERNS,
          THEME_PATTERNS,
          VERIFICATION_PATTERNS,
          count_keywords,
          count_patterns,
          count_privacy_gaps,
          deduplicate_messages,
          extract_record_channels,
          extract_strings,
          first_string_shallow,
          is_within,
          keyword_alternatives,
          normalize_path,
          read_jsonl,
          read_jsonl_head,
          read_jsonl_sampled,
          read_raw_lines_sampled,
          redact_text,
          source_title,
          truncate,
      )
      
      
      DATE_YEAR_RE = re.compile(r"^\d{4}$")
      DATE_COMPONENT_RE = re.compile(r"^\d{2}$")
      SINCE_DAYS_RE = re.compile(r"^(\d+)d$")
      
      
      @dataclass
      class Ownership:
          matched_via: str
          cwd: str
          project: str
      
      
      @dataclass
      class SignalChannels:
          eligible_user_messages: int = 0
          eligible_assistant_messages: int = 0
          ignored_context_messages: int = 0
          structured_tool_failures: int = 0
      
      
      @dataclass
      class SessionSummary:
          source: str
          project: str
          path: str
          timestamp: str | None = None
          title: str | None = None
          score: int = 0
          keyword_hits: dict[str, int] = field(default_factory=dict)
          task_themes: dict[str, int] = field(default_factory=dict)
          correction_signals: dict[str, int] = field(default_factory=dict)
          failure_signals: dict[str, int] = field(default_factory=dict)
          verification_signals: dict[str, int] = field(default_factory=dict)
          tool_calls: dict[str, int] = field(default_factory=dict)
          privacy_gaps: dict[str, int] = field(default_factory=dict)
          ownership: Ownership | None = None
          signal_channels: SignalChannels = field(default_factory=SignalChannels)
          modified: str | None = None
          excerpts: list[dict[str, str]] = field(default_factory=list)
      
      
      @dataclass
      class SinceFilter:
          value: str
          cutoff: dt.datetime
          codex_dirs_pruned: int = 0
          codex_files_pruned: int = 0
          claude_files_pruned: int = 0
      
          def to_json(self) -> dict[str, Any]:
              return {
                  "value": self.value,
                  "cutoff": format_utc_iso(self.cutoff),
                  "codex_dirs_pruned": self.codex_dirs_pruned,
                  "codex_files_pruned": self.codex_files_pruned,
                  "claude_files_pruned": self.claude_files_pruned,
              }
      
      
      def main() -> int:
          parser = argparse.ArgumentParser(description="Mine project-owned Codex and Claude Code transcript signals.")
          parser.add_argument("--project", action="append", default=[], help="Project path to mine. Repeatable. Default: pwd -P")
          parser.add_argument(
              "--keyword",
              action="append",
              default=[],
              help="Task keyword to score. Repeatable. Use 'a|b|c' for an OR-group of alternatives",
          )
          parser.add_argument("--format", choices=("text", "json"), default="text")
          parser.add_argument("--max-sessions", type=int, default=20, help="Maximum selected sessions per project")
          parser.add_argument("--include-archived", action="store_true", help="Include ~/.codex/archived_sessions")
          parser.add_argument(
              "--include-current",
              action="store_true",
              help="Include the live CODEX_THREAD_ID or CLAUDE_CODE_SESSION_ID transcript (diagnostics only)",
          )
          parser.add_argument("--since", default=None, help="Only mine sessions modified since YYYY-MM-DD or Nd (days back)")
          parser.add_argument(
              "--excerpts", action="store_true", help="Include up to 3 redacted message excerpts per candidate session"
          )
          args = parser.parse_args()
      
          if args.max_sessions < 1:
              print("transcript-miner: --max-sessions must be positive", file=sys.stderr)
              return 2
          try:
              projects = normalize_projects(args.project)
          except ValueError as error:
              print(f"transcript-miner: {error}", file=sys.stderr)
              return 2
      
          since_filter: SinceFilter | None = None
          if args.since:
              try:
                  since_filter = build_since_filter(args.since)
              except ValueError as error:
                  print(f"transcript-miner: {error}", file=sys.stderr)
                  return 2
      
          report = mine_transcripts(
              projects,
              keywords=[keyword for keyword in args.keyword if keyword.strip()],
              max_sessions=args.max_sessions,
              include_archived=args.include_archived,
              include_current=args.include_current,
              since=since_filter,
              include_excerpts=args.excerpts,
          )
          if args.format == "json":
              print(json.dumps(report, indent=2, sort_keys=True))
          else:
              print_text_report(report)
          return 0
      
      
      def normalize_projects(raw_projects: list[str]) -> list[Path]:
          projects: list[Path] = []
          for raw_project in raw_projects or [os.curdir]:
              project = Path(os.path.expanduser(raw_project)).resolve(strict=False)
              if not project.exists():
                  raise ValueError(f"project does not exist: {project}")
              if not project.is_dir():
                  raise ValueError(f"project is not a directory: {project}")
              if project not in projects:
                  projects.append(project)
          return projects
      
      
      def build_since_filter(value: str) -> SinceFilter:
          match = SINCE_DAYS_RE.fullmatch(value)
          if match:
              cutoff = dt.datetime.now(dt.timezone.utc) - dt.timedelta(days=int(match.group(1)))
              return SinceFilter(value=value, cutoff=cutoff)
          try:
              parsed = dt.datetime.strptime(value, "%Y-%m-%d")
          except ValueError:
              raise ValueError(f"invalid --since value: {value!r} (expected YYYY-MM-DD or Nd)") from None
          return SinceFilter(value=value, cutoff=parsed.replace(tzinfo=dt.timezone.utc))
      
      
      def mine_transcripts(
          projects: list[Path],
          *,
          keywords: list[str],
          max_sessions: int,
          include_archived: bool,
          include_current: bool = False,
          since: SinceFilter | None = None,
          include_excerpts: bool = False,
      ) -> dict[str, Any]:
          codex_home = Path(os.path.expanduser(os.environ.get("CODEX_HOME", "~/.codex"))).resolve(strict=False)
          claude_home = claude_config_dir()
          codex_index = load_codex_index(codex_home / "session_index.jsonl")
          claude_history = load_claude_history(claude_home / "history.jsonl")
          coverage = {project: new_coverage(codex_index["records"]) for project in projects}
          now = dt.datetime.now(dt.timezone.utc)
      
          codex_sessions = mine_codex_sessions(
              projects,
              keywords,
              codex_home,
              codex_index,
              include_archived,
              include_current,
              coverage,
              since,
              now,
              include_excerpts,
          )
          claude_sessions = mine_claude_sessions(
              projects, keywords, claude_home, claude_history, include_current, coverage, since, now, include_excerpts
          )
      
          selected_sessions: list[SessionSummary] = []
          for project in projects:
              project_sessions = [
                  session for session in [*codex_sessions, *claude_sessions] if session.project == str(project)
              ]
              selected = select_sessions(project_sessions, max_sessions)
              selected_sessions.extend(selected)
              coverage[project]["selected_sessions"] = len(selected)
      
          project_reports = [
              {
                  "path": str(project),
                  "encoded_claude_project": encode_claude_project(project),
                  "coverage": coverage[project],
              }
              for project in projects
          ]
          totals = aggregate_sessions(selected_sessions)
          return {
              "projects": project_reports,
              "keywords": [redact_text(keyword) for keyword in keywords],
              "since": since.to_json() if since else None,
              "candidate_sessions": [session_to_json(session) for session in selected_sessions],
              "task_themes": dict(totals["task_themes"]),
              "correction_signals": dict(totals["correction_signals"]),
              "failure_signals": dict(totals["failure_signals"]),
              "verification_signals": dict(totals["verification_signals"]),
              "tool_calls": dict(totals["tool_calls"]),
              "privacy_gaps": dict(totals["privacy_gaps"]),
          }
      
      
      def new_coverage(codex_index_records: int) -> dict[str, Any]:
          return {
              "codex_candidates": 0,
              "claude_candidates": 0,
              "selected_sessions": 0,
              "claude_tool_result_files": 0,
              "claude_tool_result_failures": {},
              "codex_index_records": codex_index_records,
              "codex_scanned": 0,
              "claude_scanned": 0,
              "structurally_matched": 0,
              "relevance_matched": 0,
              "current_sessions_excluded": 0,
              "content_only_project_mentions_ignored": 0,
              "ambiguous_ownership_excluded": 0,
          }
      
      
      def mine_codex_sessions(
          projects: list[Path],
          keywords: list[str],
          codex_home: Path,
          codex_index: dict[str, Any],
          include_archived: bool,
          include_current: bool,
          coverage: dict[Path, dict[str, Any]],
          since: SinceFilter | None,
          now: dt.datetime,
          include_excerpts: bool,
      ) -> list[SessionSummary]:
          roots = [codex_home / "sessions"]
          if include_archived:
              roots.append(codex_home / "archived_sessions")
          paths = collect_codex_paths(roots, since)
          for project in projects:
              coverage[project]["codex_scanned"] = len(paths)
      
          sessions: list[SessionSummary] = []
          current_id = os.environ.get("CODEX_THREAD_ID", "").strip()
          for path in paths:
              metadata = list(read_jsonl_head(path))
              ownership, ambiguous_projects = codex_ownership(metadata, path, projects)
              mark_ambiguous(coverage, ambiguous_projects)
              if ownership is None:
                  continue
              owner = Path(ownership.project)
              if current_id and current_id in codex_session_ids(metadata, path) and not include_current:
                  coverage[owner]["current_sessions_excluded"] += 1
                  continue
              coverage[owner]["codex_candidates"] += 1
              coverage[owner]["structurally_matched"] += 1
              title_hint = title_for_codex_path(path, codex_index)
              if raw_prescan_eligible(keywords) and not title_already_matches(title_hint, keywords):
                  raw_lines = list(read_raw_lines_sampled(path))
                  mark_content_only_mentions_raw(raw_lines, owner, projects, coverage)
                  if not keywords_present_in_lines(raw_lines, keywords):
                      continue
                  records = list(read_jsonl_sampled(path))
              else:
                  records = list(read_jsonl_sampled(path))
                  mark_content_only_mentions(records, owner, projects, coverage)
              summary = summarize_records(
                  path,
                  source="codex",
                  keywords=keywords,
                  ownership=ownership,
                  records=records,
                  title_hint=title_hint,
                  now=now,
                  include_excerpts=include_excerpts,
              )
              if summary is not None:
                  coverage[owner]["relevance_matched"] += 1
                  sessions.append(summary)
          return sessions
      
      
      def collect_codex_paths(roots: list[Path], since: SinceFilter | None) -> list[Path]:
          if since is None:
              paths = [path for root in roots if root.is_dir() for path in root.rglob("*.jsonl")]
              return sorted(paths, key=str)
          slack_date = since.cutoff.date() - dt.timedelta(days=1)
          paths: list[Path] = []
          for root in roots:
              if root.is_dir():
                  paths.extend(walk_codex_root(root, slack_date, since))
          kept = [path for path in paths if not prune_if_stale(path, since, "codex_files_pruned")]
          return sorted(kept, key=str)
      
      
      def walk_codex_root(root: Path, slack_date: dt.date, since: SinceFilter) -> Iterable[Path]:
          yield from root.glob("*.jsonl")
          for year_dir in root.iterdir():
              if not (year_dir.is_dir() and DATE_YEAR_RE.fullmatch(year_dir.name)):
                  if year_dir.is_dir():
                      yield from year_dir.rglob("*.jsonl")
                  continue
              yield from year_dir.glob("*.jsonl")
              for month_dir in year_dir.iterdir():
                  if not (month_dir.is_dir() and DATE_COMPONENT_RE.fullmatch(month_dir.name)):
                      if month_dir.is_dir():
                          yield from month_dir.rglob("*.jsonl")
                      continue
                  yield from month_dir.glob("*.jsonl")
                  for day_dir in month_dir.iterdir():
                      if not (day_dir.is_dir() and DATE_COMPONENT_RE.fullmatch(day_dir.name)):
                          if day_dir.is_dir():
                              yield from day_dir.rglob("*.jsonl")
                          continue
                      try:
                          dir_date = dt.date(int(year_dir.name), int(month_dir.name), int(day_dir.name))
                      except ValueError:
                          yield from day_dir.rglob("*.jsonl")
                          continue
                      if dir_date < slack_date:
                          # An old date directory can still hold sessions resumed recently; mtime decides.
                          survivors = [
                              path
                              for path in day_dir.rglob("*.jsonl")
                              if not prune_if_stale(path, since, "codex_files_pruned")
                          ]
                          if survivors:
                              yield from survivors
                          else:
                              since.codex_dirs_pruned += 1
                          continue
                      yield from day_dir.rglob("*.jsonl")
      
      
      def prune_if_stale(path: Path, since: SinceFilter, counter: str) -> bool:
          try:
              mtime = path.stat().st_mtime
          except OSError:
              return False
          if mtime < since.cutoff.timestamp():
              setattr(since, counter, getattr(since, counter) + 1)
              return True
          return False
      
      
      def raw_prescan_eligible(keywords: list[str]) -> bool:
          if not keywords:
              return False
          for keyword in keywords:
              alternatives = keyword_alternatives(keyword)
              if not alternatives:
                  return False
              for alternative in alternatives:
                  if not alternative.isascii() or '"' in alternative or "\\" in alternative:
                      return False
          return True
      
      
      def title_already_matches(title: str | None, keywords: list[str]) -> bool:
          if not title or not keywords:
              return False
          return bool(count_keywords([title], keywords))
      
      
      def keywords_present_in_lines(lines: list[str], keywords: list[str]) -> bool:
          alternatives = [alternative.lower() for keyword in keywords for alternative in keyword_alternatives(keyword)]
          if not alternatives:
              return False
          for line in lines:
              lowered = line.lower()
              if any(alternative in lowered for alternative in alternatives):
                  return True
          return False
      
      
      def codex_ownership(
          metadata: list[Any], path: Path, projects: list[Path]
      ) -> tuple[Ownership | None, set[Path]]:
          session_cwds: list[Path] = []
          turn_cwds: list[Path] = []
          for item in metadata:
              if not isinstance(item, dict):
                  continue
              payload = item.get("payload") if isinstance(item.get("payload"), dict) else {}
              item_type = item.get("type")
              cwd = payload.get("cwd")
              if item_type == "session_meta" and isinstance(cwd, str) and cwd.strip():
                  session_cwds.append(normalize_path(cwd))
              elif item_type == "turn_context" and isinstance(cwd, str) and cwd.strip():
                  turn_cwds.append(normalize_path(cwd))
      
          if session_cwds:
              unique = set(session_cwds)
              owners = {owner for cwd in unique if (owner := most_specific_project(cwd, projects)) is not None}
              if len(unique) != 1 or len(owners) > 1:
                  return None, owners
              cwd = session_cwds[0]
              owner = most_specific_project(cwd, projects)
              if owner is None:
                  return None, set()
              return Ownership("session_meta.payload.cwd", str(cwd), str(owner)), set()
      
          if not turn_cwds:
              sampled = list(read_jsonl_sampled(path))
              for item in sampled:
                  if not isinstance(item, dict) or item.get("type") != "turn_context":
                      continue
                  payload = item.get("payload")
                  cwd = payload.get("cwd") if isinstance(payload, dict) else None
                  if isinstance(cwd, str) and cwd.strip():
                      turn_cwds.append(normalize_path(cwd))
          owners = {owner for cwd in turn_cwds if (owner := most_specific_project(cwd, projects)) is not None}
          if not turn_cwds or len(owners) != 1 or any(most_specific_project(cwd, projects) not in owners for cwd in turn_cwds):
              return None, owners
          owner = next(iter(owners))
          canonical_cwds = set(turn_cwds)
          if any(not is_within(cwd, owner) for cwd in canonical_cwds):
              return None, owners
          return Ownership("turn_context.cwd", str(sorted(canonical_cwds, key=str)[0]), str(owner)), set()
      
      
      def codex_session_ids(metadata: list[Any], path: Path) -> set[str]:
          ids = {path.stem}
          for item in metadata:
              if isinstance(item, dict) and item.get("type") == "session_meta":
                  payload = item.get("payload")
                  if isinstance(payload, dict):
                      for key in ("id", "session_id", "conversation_id"):
                          value = payload.get(key)
                          if isinstance(value, str):
                              ids.add(value)
          return ids
      
      
      def load_claude_history(path: Path) -> dict[str, list[dict[str, str]]]:
          records: dict[str, list[dict[str, str]]] = defaultdict(list)
          for item in read_jsonl(path):
              if not isinstance(item, dict):
                  continue
              session_id = item.get("sessionId") or item.get("session_id")
              project = item.get("project")
              display = item.get("display")
              if not isinstance(session_id, str) or not isinstance(project, str):
                  continue
              records[session_id].append(
                  {
                      "project": str(normalize_path(project)),
                      "display": display.strip() if isinstance(display, str) else "",
                  }
              )
          return records
      
      
      def mine_claude_sessions(
          projects: list[Path],
          keywords: list[str],
          claude_home: Path,
          history: dict[str, list[dict[str, str]]],
          include_current: bool,
          coverage: dict[Path, dict[str, Any]],
          since: SinceFilter | None,
          now: dt.datetime,
          include_excerpts: bool,
      ) -> list[SessionSummary]:
          directory_projects: dict[Path, set[Path]] = defaultdict(set)
          for project in projects:
              for key in claude_project_keys(project):
                  project_dir = claude_home / "projects" / key
                  if project_dir.is_dir():
                      directory_projects[project_dir.resolve(strict=False)].add(project)
      
          path_directory_projects: dict[Path, set[Path]] = defaultdict(set)
          for project_dir, possible_projects in directory_projects.items():
              paths = list(project_dir.glob("*.jsonl"))
              if since is not None:
                  paths = [path for path in paths if not prune_if_stale(path, since, "claude_files_pruned")]
              for project in possible_projects:
                  coverage[project]["claude_scanned"] += len(paths)
                  tool_results_dir = project_dir / "tool-results"
                  if tool_results_dir.is_dir():
                      coverage[project]["claude_tool_result_files"] += sum(1 for path in tool_results_dir.rglob("*") if path.is_file())
              for path in paths:
                  path_directory_projects[path.resolve(strict=False)].update(possible_projects)
      
          sessions: list[SessionSummary] = []
          current_id = os.environ.get("CLAUDE_CODE_SESSION_ID", "").strip()
          for path, directory_candidates in sorted(path_directory_projects.items(), key=lambda item: str(item[0])):
              metadata = list(read_jsonl_head(path))
              session_id = claude_session_id(metadata, path)
              history_records = history.get(session_id, [])
              ownership, ambiguous_projects = claude_ownership(
                  metadata, directory_candidates, history_records, projects
              )
              if ownership is None:
                  mark_ambiguous(coverage, ambiguous_projects or directory_candidates)
                  continue
              owner = Path(ownership.project)
              if current_id and session_id == current_id and not include_current:
                  coverage[owner]["current_sessions_excluded"] += 1
                  continue
              coverage[owner]["claude_candidates"] += 1
              coverage[owner]["structurally_matched"] += 1
      
              history_messages = deduplicate_messages(
                  record["display"] for record in history_records if record["display"]
              )
              if keywords and history_records and not count_keywords(history_messages, keywords):
                  continue
              if not history_records and raw_prescan_eligible(keywords):
                  raw_lines = list(read_raw_lines_sampled(path))
                  mark_content_only_mentions_raw(raw_lines, owner, projects, coverage)
                  if not keywords_present_in_lines(raw_lines, keywords):
                      continue
                  records = list(read_jsonl_sampled(path))
              else:
                  records = list(read_jsonl_sampled(path))
                  mark_content_only_mentions(records, owner, projects, coverage)
              summary = summarize_records(
                  path,
                  source="claude",
                  keywords=keywords,
                  ownership=ownership,
                  records=records,
                  title_hint=None,
                  preferred_user_messages=history_messages if history_records else None,
                  now=now,
                  include_excerpts=include_excerpts,
              )
              if summary is not None:
                  coverage[owner]["relevance_matched"] += 1
                  sessions.append(summary)
          return sessions
      
      
      def claude_ownership(
          metadata: list[Any],
          directory_candidates: set[Path],
          history_records: list[dict[str, str]],
          projects: list[Path],
      ) -> tuple[Ownership | None, set[Path]]:
          cwd_values: set[Path] = set()
          for item in metadata:
              if isinstance(item, dict):
                  cwd = item.get("cwd")
                  if isinstance(cwd, str) and cwd.strip():
                      cwd_values.add(normalize_path(cwd))
          cwd_owners = {owner for cwd in cwd_values if (owner := most_specific_project(cwd, projects)) is not None}
          history_paths = {normalize_path(record["project"]) for record in history_records}
          history_owners = {
              owner for project_path in history_paths if (owner := most_specific_project(project_path, projects)) is not None
          }
          evidence = [set(directory_candidates)]
          if cwd_values:
              evidence.append(cwd_owners)
          if history_records:
              evidence.append(history_owners)
          implicated = set().union(*evidence)
          if any(len(values) != 1 for values in evidence):
              return None, implicated
          owners = set().union(*evidence)
          if len(owners) != 1:
              return None, implicated
          owner = next(iter(owners))
          if cwd_values:
              cwd = sorted(cwd_values, key=str)[0]
              matched_via = "transcript.cwd+encoded-directory"
          else:
              cwd = owner
              matched_via = "encoded-directory"
          if history_records:
              matched_via += "+history.project"
          return Ownership(matched_via, str(cwd), str(owner)), set()
      
      
      def claude_session_id(metadata: list[Any], path: Path) -> str:
          for item in metadata:
              if not isinstance(item, dict):
                  continue
              for key in ("sessionId", "session_id", "conversation_id"):
                  value = item.get(key)
                  if isinstance(value, str) and value.strip():
                      return value.strip()
          return path.stem
      
      
      def summarize_records(
          path: Path,
          *,
          source: str,
          keywords: list[str],
          ownership: Ownership,
          records: list[Any],
          title_hint: str | None,
          now: dt.datetime,
          include_excerpts: bool,
          preferred_user_messages: list[str] | None = None,
      ) -> SessionSummary | None:
          user_messages: list[str] = []
          assistant_messages: list[str] = []
          context_messages: list[str] = []
          tool_calls: Counter[str] = Counter()
          structured_failures = 0
          timestamp: str | None = None
          title = title_hint
          for item in records:
              if timestamp is None:
                  timestamp = first_string_shallow(item, ("timestamp", "created_at", "updated_at"))
              if title is None:
                  title = source_title(item)
              channels = extract_record_channels(item)
              user_messages.extend(channels["user"])
              assistant_messages.extend(channels["assistant"])
              context_messages.extend(channels["context"])
              tool_calls.update(channels["tools"])
              structured_failures += channels["failures"]
      
          if preferred_user_messages is not None:
              user_messages = preferred_user_messages
          user_messages = deduplicate_messages(user_messages)
          assistant_messages = deduplicate_messages(assistant_messages)
          context_messages = deduplicate_messages(context_messages)
          title = truncate(redact_text(title)) if title else None
      
          user_hits = count_keywords(user_messages, keywords)
          assistant_hits = count_keywords(assistant_messages, keywords)
          title_hits = count_keywords([title] if title else [], keywords)
          if keywords and not (user_hits or assistant_hits or title_hits):
              return None
          keyword_hits = user_hits + assistant_hits + title_hits
          corrections = count_patterns(user_messages, CORRECTION_PATTERNS)
          verification = count_patterns(assistant_messages, VERIFICATION_PATTERNS)
          eligible_messages = [*user_messages, *assistant_messages]
          themes = count_patterns(eligible_messages, THEME_PATTERNS)
          privacy = count_privacy_gaps(eligible_messages)
          failures = Counter({"command-failure": structured_failures}) if structured_failures else Counter()
      
          modified_dt = file_modified_time(path)
          age_source = parse_candidate_timestamp(timestamp) or modified_dt
          bonus = recency_bonus(now, age_source) if age_source is not None else 0
      
          score = (
              10
              + sum(user_hits.values()) * 8
              + sum(title_hits.values()) * 4
              + sum(assistant_hits.values()) * 2
              + min(6, sum(corrections.values()) * 2)
              + min(4, structured_failures)
              + min(4, sum(verification.values()))
              + bonus
          )
          excerpts = build_excerpts(user_messages, assistant_messages, keywords, include_excerpts)
          return SessionSummary(
              source=source,
              project=ownership.project,
              path=str(path),
              timestamp=timestamp,
              title=title,
              score=score,
              keyword_hits=dict(keyword_hits),
              task_themes=dict(themes),
              correction_signals=dict(corrections),
              failure_signals=dict(failures),
              verification_signals=dict(verification),
              tool_calls=dict(tool_calls),
              privacy_gaps=dict(privacy),
              ownership=ownership,
              signal_channels=SignalChannels(
                  eligible_user_messages=len(user_messages),
                  eligible_assistant_messages=len(assistant_messages),
                  ignored_context_messages=len(context_messages),
                  structured_tool_failures=structured_failures,
              ),
              modified=format_utc_iso(modified_dt) if modified_dt is not None else None,
              excerpts=excerpts,
          )
      
      
      def file_modified_time(path: Path) -> dt.datetime | None:
          try:
              mtime = path.stat().st_mtime
          except OSError:
              return None
          return dt.datetime.fromtimestamp(mtime, tz=dt.timezone.utc)
      
      
      def parse_candidate_timestamp(value: str | None) -> dt.datetime | None:
          if not value:
              return None
          text = value.strip()
          if text.endswith("Z"):
              text = text[:-1] + "+00:00"
          try:
              parsed = dt.datetime.fromisoformat(text)
          except ValueError:
              return None
          if parsed.tzinfo is None:
              parsed = parsed.replace(tzinfo=dt.timezone.utc)
          return parsed.astimezone(dt.timezone.utc)
      
      
      def recency_bonus(now: dt.datetime, candidate: dt.datetime) -> int:
          age_days = (now - candidate).total_seconds() / 86400
          if age_days <= 7:
              return 6
          if age_days <= 30:
              return 3
          return 0
      
      
      def format_utc_iso(value: dt.datetime) -> str:
          return value.astimezone(dt.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
      
      
      def build_excerpts(
          user_messages: list[str], assistant_messages: list[str], keywords: list[str], include_excerpts: bool
      ) -> list[dict[str, str]]:
          if not include_excerpts:
              return []
          excerpts: list[dict[str, str]] = []
          seen: set[str] = set()
          if user_messages:
              first = user_messages[0]
              excerpts.append({"channel": "user", "text": truncate(redact_text(first), 240)})
              seen.add(" ".join(first.split()))
          if keywords:
              alternatives = [alternative.lower() for keyword in keywords for alternative in keyword_alternatives(keyword)]
              for channel, messages in (("user", user_messages), ("assistant", assistant_messages)):
                  for message in messages:
                      if len(excerpts) >= 3:
                          return excerpts
                      normalized = " ".join(message.split())
                      if normalized in seen:
                          continue
                      if any(alternative in message.lower() for alternative in alternatives):
                          excerpts.append({"channel": channel, "text": truncate(redact_text(message), 240)})
                          seen.add(normalized)
          return excerpts
      
      
      def mark_content_only_mentions(
          records: list[Any], owner: Path, projects: list[Path], coverage: dict[Path, dict[str, Any]]
      ) -> None:
          strings = list(extract_strings(records))
          for project in projects:
              if project != owner and any(str(project) in text for text in strings):
                  coverage[project]["content_only_project_mentions_ignored"] += 1
      
      
      def mark_content_only_mentions_raw(
          lines: list[str], owner: Path, projects: list[Path], coverage: dict[Path, dict[str, Any]]
      ) -> None:
          for project in projects:
              if project != owner and any(str(project) in line for line in lines):
                  coverage[project]["content_only_project_mentions_ignored"] += 1
      
      
      def mark_ambiguous(coverage: dict[Path, dict[str, Any]], projects: Iterable[Path]) -> None:
          for project in set(projects):
              coverage[project]["ambiguous_ownership_excluded"] += 1
      
      
      def most_specific_project(cwd: Path, projects: list[Path]) -> Path | None:
          matches = [project for project in projects if is_within(cwd, project)]
          return max(matches, key=lambda project: len(project.parts), default=None)
      
      
      def load_codex_index(path: Path) -> dict[str, Any]:
          titles_by_id: dict[str, str] = {}
          titles_by_path: dict[str, str] = {}
          records = 0
          for item in read_jsonl(path):
              records += 1
              session_id = first_string_shallow(item, ("id", "session_id", "conversation_id"))
              transcript_path = first_string_shallow(item, ("path", "file", "rollout_path", "transcript_path"))
              title = first_string_shallow(item, ("title", "summary", "task", "prompt"))
              if title and session_id:
                  titles_by_id[session_id] = redact_text(title)
              if title and transcript_path:
                  titles_by_path[str(normalize_path(transcript_path))] = redact_text(title)
          return {"records": records, "titles_by_id": titles_by_id, "titles_by_path": titles_by_path}
      
      
      def title_for_codex_path(path: Path, codex_index: dict[str, Any]) -> str | None:
          title = codex_index["titles_by_path"].get(str(path.resolve(strict=False)))
          if title:
              return title
          for session_id, candidate_title in codex_index["titles_by_id"].items():
              if session_id in path.name:
                  return candidate_title
          return None
      
      
      def claude_config_dir() -> Path:
          config_dir = os.environ.get("CLAUDE_CONFIG_DIR") or os.environ.get("CLAUDE_HOME") or "~/.claude"
          return Path(os.path.expanduser(config_dir)).resolve(strict=False)
      
      
      def claude_project_keys(project: Path) -> list[str]:
          return list(dict.fromkeys([encode_claude_project(project), legacy_encode_claude_project(project)]))
      
      
      def encode_claude_project(project: Path) -> str:
          return re.sub(r"[^A-Za-z0-9]", "-", str(project))
      
      
      def legacy_encode_claude_project(project: Path) -> str:
          return str(project).replace("/", "-")
      
      
      def select_sessions(sessions: list[SessionSummary], max_sessions: int) -> list[SessionSummary]:
          return sorted(sessions, key=lambda session: (session.score, session.timestamp or "", session.path), reverse=True)[
              :max_sessions
          ]
      
      
      def aggregate_sessions(sessions: list[SessionSummary]) -> dict[str, Counter[str]]:
          totals: dict[str, Counter[str]] = defaultdict(Counter)
          for session in sessions:
              for key in (
                  "task_themes",
                  "correction_signals",
                  "failure_signals",
                  "verification_signals",
                  "tool_calls",
                  "privacy_gaps",
              ):
                  totals[key].update(getattr(session, key))
          return totals
      
      
      def session_to_json(session: SessionSummary) -> dict[str, Any]:
          return asdict(session)
      
      
      def print_text_report(report: dict[str, Any]) -> None:
          print("transcript-miner")
          since = report.get("since")
          if since:
              print(
                  f"\nSince: {since['value']} (cutoff {since['cutoff']}) — "
                  f"codex_dirs_pruned={since['codex_dirs_pruned']}, codex_files_pruned={since['codex_files_pruned']}, "
                  f"claude_files_pruned={since['claude_files_pruned']}"
              )
          print("\nProjects:")
          for project in report["projects"]:
              coverage = project["coverage"]
              print(
                  f"- {project['path']}: {coverage['codex_scanned']} Codex scanned, "
                  f"{coverage['claude_scanned']} Claude scanned, {coverage['structurally_matched']} structurally owned, "
                  f"{coverage['relevance_matched']} relevant, {coverage['selected_sessions']} selected"
              )
              print(
                  "  exclusions: "
                  f"current={coverage['current_sessions_excluded']}, "
                  f"content-only={coverage['content_only_project_mentions_ignored']}, "
                  f"ambiguous={coverage['ambiguous_ownership_excluded']}"
              )
          if report["task_themes"]:
              print("\nTask themes:")
              for name, count in sorted(report["task_themes"].items(), key=lambda item: (-item[1], item[0])):
                  print(f"- {name}: {count}")
          print("\nCandidate sessions:")
          if not report["candidate_sessions"]:
              print("- none")
          for session in report["candidate_sessions"]:
              title = f" — {session['title']}" if session.get("title") else ""
              print(f"- {session['source']} score={session['score']} {session['path']}{title}")
              ownership = session["ownership"]
              channels = session["signal_channels"]
              print(
                  f"  ownership: {ownership['matched_via']} cwd={ownership['cwd']} project={ownership['project']}; "
                  f"channels: user={channels['eligible_user_messages']}, "
                  f"assistant={channels['eligible_assistant_messages']}, "
                  f"context-ignored={channels['ignored_context_messages']}, "
                  f"tool-failures={channels['structured_tool_failures']}"
              )
              compact = compact_session_line(session)
              if compact:
                  print(f"  {compact}")
              for excerpt in session.get("excerpts") or []:
                  print(f"  excerpt[{excerpt['channel']}]: {excerpt['text']}")
          print_counter_block("Correction signals", report["correction_signals"])
          print_counter_block("Failure signals", report["failure_signals"])
          print_counter_block("Verification signals", report["verification_signals"])
          print_counter_block("Tool calls", report["tool_calls"])
          print_counter_block("Privacy gaps", report["privacy_gaps"])
      
      
      def compact_session_line(session: dict[str, Any]) -> str:
          parts = []
          for label, key in (
              ("keywords", "keyword_hits"),
              ("themes", "task_themes"),
              ("failures", "failure_signals"),
              ("verification", "verification_signals"),
          ):
              values = session.get(key) or {}
              if values:
                  parts.append(f"{label}: {', '.join(sorted(values)[:4])}")
          return "; ".join(parts)
      
      
      def print_counter_block(title: str, values: dict[str, int]) -> None:
          print(f"\n{title}:")
          if not values:
              print("- none")
              return
          for name, count in sorted(values.items(), key=lambda item: (-item[1], item[0])):
              print(f"- {name}: {count}")
      
      
      if __name__ == "__main__":
          raise SystemExit(main())
      
    • transcript_common.py 16.8 KB
      """Shared redaction, pattern, and reader helpers for transcript mining scripts."""
      
      from __future__ import annotations
      
      import json
      import os
      import re
      from collections import Counter, deque
      from pathlib import Path
      from typing import Any, Iterable
      
      
      EMAIL_RE = re.compile(r"\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b")
      API_KEY_RE = re.compile(r"\b(?:sk|pk|rk|ghp|github_pat|xox[baprs])-[A-Za-z0-9_\-]{16,}\b")
      GENERIC_SECRET_RE = re.compile(
          r"(?i)\b(?:api[_-]?key|access[_-]?token|secret(?:[_-]?key)?|private[_-]?key|password)\b"
          r"\s*[:=]\s*[\"']?(?![$<{[])([A-Za-z0-9_./+=-]{16,})"
      )
      EVM_ADDRESS_RE = re.compile(r"(?<![0-9a-fA-F])0x[0-9a-fA-F]{40}(?![0-9a-fA-F])")
      HEX_64_RE = re.compile(r"(?<![0-9a-fA-F])0x[0-9a-fA-F]{64}(?![0-9a-fA-F])")
      LONG_SECRETISH_RE = re.compile(r"(?<![A-Za-z0-9_/-])[A-Za-z0-9_+=/-]{48,}(?![A-Za-z0-9_/-])")
      
      CORRECTION_PATTERNS = {
          "user-correction": re.compile(r"(?i)\b(actually|wrong|instead|i asked|not what|do not|don't|stop|you should)\b"),
          "instruction-reminder": re.compile(r"(?i)\b(AGENTS\.md|instructions?|follow .*rules?|violat(?:e|ed|ion))\b"),
      }
      VERIFICATION_PATTERNS = {
          "tests": re.compile(r"(?i)\b(test|pytest|vitest|unit tests?|integration tests?)\b"),
          "lint-format": re.compile(r"(?i)\b(lint|format|mdformat|prettier|ruff|eslint)\b"),
          "repo-check": re.compile(r"(?i)\b(just |uv run|cargo check|npm run|pnpm |bun test|verified|verification|passes?|passed)\b"),
      }
      THEME_PATTERNS = {
          "agent-skills": re.compile(r"(?i)\b(SKILL\.md|openai\.yaml|frontmatter|allow_implicit|agent skills?|skill catalog)\b"),
          "transcripts": re.compile(r"(?i)\b(transcript|session_index|sessions/|claude/projects|history\.jsonl)\b"),
          "markdown": re.compile(r"(?i)\b(markdown|mdformat|README\.md|AGENTS\.md)\b"),
          "git": re.compile(r"(?i)\b(git status|git diff|commit|branch|worktree|staged)\b"),
          "shell-tooling": re.compile(r"(?i)\b(zsh|bash|just|uv run|rg |fd |jq )\b"),
          "privacy": re.compile(r"(?i)\b(secret|redact|private key|api key|token|wallet|address)\b"),
      }
      CONTEXT_PREFIXES = (
          "# AGENTS.md instructions for",
          "<skill>",
          "<environment_context>",
          "<permissions instructions>",
          "<collaboration_mode>",
          "<turn_aborted>",
          "<command-message>",
          "Base directory for this skill:",
      )
      FAILURE_STATUSES = {"error", "failed", "failure", "cancelled", "canceled", "timed_out", "timeout"}
      MAX_FULL_SESSION_BYTES = 2_000_000
      SESSION_HEAD_RECORDS = 250
      SESSION_TAIL_RECORDS = 750
      METADATA_HEAD_RECORDS = 100
      
      
      def redact_text(value: str | None) -> str:
          if not value:
              return ""
          text = EMAIL_RE.sub("<email>", value)
          text = API_KEY_RE.sub("<api-key>", text)
          text = GENERIC_SECRET_RE.sub(lambda match: match.group(0).split(match.group(1), 1)[0] + "<secret>", text)
          text = HEX_64_RE.sub("<tx-or-key-hash>", text)
          text = EVM_ADDRESS_RE.sub("<evm-address>", text)
          return LONG_SECRETISH_RE.sub("<secret-like-token>", text)
      
      
      def truncate(value: str, limit: int = 160) -> str:
          return value if len(value) <= limit else value[: limit - 1].rstrip() + "..."
      
      
      def count_privacy_gaps(strings: list[str]) -> Counter[str]:
          counter: Counter[str] = Counter()
          for text in strings:
              if EMAIL_RE.search(text):
                  counter["email"] += 1
              if API_KEY_RE.search(text) or GENERIC_SECRET_RE.search(text):
                  counter["api-key-or-secret"] += 1
              if HEX_64_RE.search(text):
                  counter["private-key-or-tx-hash"] += 1
              if EVM_ADDRESS_RE.search(text):
                  counter["evm-address"] += 1
              if LONG_SECRETISH_RE.search(text):
                  counter["long-secret-like-token"] += 1
          return counter
      
      
      def count_patterns(strings: list[str], patterns: dict[str, re.Pattern[str]]) -> Counter[str]:
          counter: Counter[str] = Counter()
          for text in strings:
              for name, pattern in patterns.items():
                  if pattern.search(text):
                      counter[name] += 1
          return counter
      
      
      def keyword_alternatives(keyword: str) -> list[str]:
          return [alternative.strip() for alternative in keyword.split("|") if alternative.strip()]
      
      
      def count_keywords(strings: list[str], keywords: list[str]) -> Counter[str]:
          counter: Counter[str] = Counter()
          lowered = [text.lower() for text in strings]
          for keyword in keywords:
              alternatives = [alternative.lower() for alternative in keyword_alternatives(keyword)]
              if not alternatives:
                  continue
              matches = sum(1 for text in lowered if any(alternative in text for alternative in alternatives))
              if matches:
                  counter[redact_text(keyword)] = matches
          return counter
      
      
      def read_jsonl(path: Path) -> Iterable[Any]:
          try:
              with path.open("r", encoding="utf-8", errors="replace") as handle:
                  for line in handle:
                      item = parse_jsonl_line(line)
                      if item is not None:
                          yield item
          except OSError:
              return
      
      
      def read_jsonl_head(path: Path) -> Iterable[Any]:
          try:
              with path.open("r", encoding="utf-8", errors="replace") as handle:
                  for _, line in zip(range(METADATA_HEAD_RECORDS), handle):
                      item = parse_jsonl_line(line)
                      if item is not None:
                          yield item
          except OSError:
              return
      
      
      def read_jsonl_sampled(path: Path) -> Iterable[Any]:
          try:
              if path.stat().st_size <= MAX_FULL_SESSION_BYTES:
                  yield from read_jsonl(path)
                  return
          except OSError:
              return
          tail: deque[str] = deque(maxlen=SESSION_TAIL_RECORDS)
          try:
              with path.open("r", encoding="utf-8", errors="replace") as handle:
                  for line_number, line in enumerate(handle):
                      if line_number < SESSION_HEAD_RECORDS:
                          item = parse_jsonl_line(line)
                          if item is not None:
                              yield item
                      else:
                          tail.append(line)
          except OSError:
              return
          for line in tail:
              item = parse_jsonl_line(line)
              if item is not None:
                  yield item
      
      
      def read_jsonl_sampled_with_lines(path: Path) -> Iterable[tuple[int, Any]]:
          try:
              if path.stat().st_size <= MAX_FULL_SESSION_BYTES:
                  with path.open("r", encoding="utf-8", errors="replace") as handle:
                      for line_number, line in enumerate(handle):
                          item = parse_jsonl_line(line)
                          if item is not None:
                              yield line_number, item
                  return
          except OSError:
              return
          tail: deque[tuple[int, str]] = deque(maxlen=SESSION_TAIL_RECORDS)
          try:
              with path.open("r", encoding="utf-8", errors="replace") as handle:
                  for line_number, line in enumerate(handle):
                      if line_number < SESSION_HEAD_RECORDS:
                          item = parse_jsonl_line(line)
                          if item is not None:
                              yield line_number, item
                      else:
                          tail.append((line_number, line))
          except OSError:
              return
          for line_number, line in tail:
              item = parse_jsonl_line(line)
              if item is not None:
                  yield line_number, item
      
      
      def read_raw_lines_sampled(path: Path) -> Iterable[str]:
          try:
              if path.stat().st_size <= MAX_FULL_SESSION_BYTES:
                  with path.open("r", encoding="utf-8", errors="replace") as handle:
                      yield from handle
                  return
          except OSError:
              return
          tail: deque[str] = deque(maxlen=SESSION_TAIL_RECORDS)
          try:
              with path.open("r", encoding="utf-8", errors="replace") as handle:
                  for line_number, line in enumerate(handle):
                      if line_number < SESSION_HEAD_RECORDS:
                          yield line
                      else:
                          tail.append(line)
          except OSError:
              return
          yield from tail
      
      
      def parse_jsonl_line(line: str) -> Any | None:
          try:
              return json.loads(line) if line.strip() else None
          except json.JSONDecodeError:
              return None
      
      
      def parse_content_blocks(content: Any) -> tuple[list[str], Counter[str], int]:
          texts: list[str] = []
          tools: Counter[str] = Counter()
          failures = 0
          if isinstance(content, str):
              texts.append(content)
          elif isinstance(content, list):
              for block in content:
                  if isinstance(block, str):
                      texts.append(block)
                      continue
                  if not isinstance(block, dict):
                      continue
                  block_type = str(block.get("type", ""))
                  if block_type in {"text", "input_text", "output_text"} and isinstance(block.get("text"), str):
                      texts.append(block["text"])
                  elif block_type in {"tool_use", "tool_call", "function_call"}:
                      name = block.get("name")
                      if isinstance(name, str):
                          tools[redact_text(name)] += 1
                  elif block_type in {"tool_result", "function_call_output"}:
                      failures += int(has_structured_failure(block))
          elif isinstance(content, dict):
              nested_texts, nested_tools, nested_failures = parse_content_blocks([content])
              texts.extend(nested_texts)
              tools.update(nested_tools)
              failures += nested_failures
          return texts, tools, failures
      
      
      def extract_record_channels(item: Any) -> dict[str, Any]:
          result: dict[str, Any] = {"user": [], "assistant": [], "context": [], "tools": Counter(), "failures": 0}
          if not isinstance(item, dict):
              return result
          payload = item.get("payload") if isinstance(item.get("payload"), dict) else None
          candidate = payload or item
          candidate_type = str(candidate.get("type", item.get("type", "")))
          role = candidate.get("role")
          message = candidate.get("message")
          if isinstance(message, dict):
              role = message.get("role", role)
              content = message.get("content")
          else:
              content = candidate.get("content")
              if content is None and isinstance(message, str):
                  content = message
      
          if candidate_type == "user_message" and role is None:
              role = "user"
              content = candidate.get("message")
          texts, tools, failures = parse_content_blocks(content)
          result["tools"].update(tools)
          collect_structured_tools(candidate, result["tools"])
          result["failures"] += max(failures, int(has_structured_failure(candidate)))
      
          if role in {"system", "developer"} or item.get("type") in {"system", "developer"}:
              result["context"].extend(texts)
          elif role == "user" or item.get("type") == "user" or candidate_type == "user_message":
              for text in texts:
                  if is_context_envelope(text):
                      result["context"].append(text)
                  else:
                      result["user"].append(text)
          elif role == "assistant" or item.get("type") == "assistant":
              result["assistant"].extend(texts)
          return result
      
      
      def collect_structured_tools(value: Any, counter: Counter[str]) -> None:
          if not isinstance(value, dict):
              return
          value_type = str(value.get("type", ""))
          if value_type in {"function_call", "tool_use", "tool_call"}:
              name = value.get("name") or value.get("tool_name")
              if isinstance(name, str):
                  counter[redact_text(name)] += 1
      
      
      def has_structured_failure(value: Any, depth: int = 0) -> bool:
          if depth > 6 or not isinstance(value, (dict, list)):
              return False
          if isinstance(value, list):
              return any(has_structured_failure(child, depth + 1) for child in value)
          if value.get("is_error") is True or value.get("isError") is True:
              return True
          exit_code = value.get("exit_code", value.get("exitCode"))
          if isinstance(exit_code, int) and exit_code != 0:
              return True
          status = value.get("status")
          if isinstance(status, str) and status.lower() in FAILURE_STATUSES:
              return True
          for key, child in value.items():
              if key in {"input", "arguments", "command", "text"}:
                  continue
              if isinstance(child, (dict, list)) and has_structured_failure(child, depth + 1):
                  return True
              if key in {"output", "content"} and isinstance(child, str):
                  parsed = parse_json_value(child)
                  if parsed is not None and has_structured_failure(parsed, depth + 1):
                      return True
          return False
      
      
      def parse_json_value(value: str) -> Any | None:
          value = value.strip()
          if not value.startswith(("{", "[")):
              return None
          try:
              return json.loads(value)
          except json.JSONDecodeError:
              return None
      
      
      def extract_tool_failures(item: Any) -> list[dict[str, str | None]]:
          """Return failure details (tool, status, text) for tool_result/function_call_output blocks in a record."""
          if not isinstance(item, dict):
              return []
          payload = item.get("payload") if isinstance(item.get("payload"), dict) else None
          candidate = payload or item
          message = candidate.get("message")
          content = message.get("content") if isinstance(message, dict) else candidate.get("content")
          results: list[dict[str, str | None]] = []
          candidate_type = str(candidate.get("type", item.get("type", "")))
          if candidate_type in {"function_call_output", "tool_result"}:
              detail = tool_failure_detail(candidate)
              if detail is not None:
                  results.append(detail)
          if isinstance(content, list):
              for block in content:
                  if isinstance(block, dict) and str(block.get("type", "")) in {"function_call_output", "tool_result"}:
                      detail = tool_failure_detail(block)
                      if detail is not None:
                          results.append(detail)
          return results
      
      
      def tool_failure_detail(block: dict[str, Any]) -> dict[str, str | None] | None:
          if not has_structured_failure(block):
              return None
          name = block.get("name") or block.get("tool_name")
          status, text = tool_failure_status_and_text(block)
          return {"tool": redact_text(name) if isinstance(name, str) else None, "status": status, "text": redact_text(text)}
      
      
      def tool_failure_status_and_text(value: Any) -> tuple[str | None, str]:
          if not isinstance(value, dict):
              return None, ""
          exit_code = value.get("exit_code", value.get("exitCode"))
          status = value.get("status")
          if isinstance(status, str) and isinstance(exit_code, int):
              status_text: str | None = f"{status} (exit_code={exit_code})"
          elif isinstance(status, str):
              status_text = status
          elif isinstance(exit_code, int):
              status_text = f"exit_code={exit_code}"
          elif value.get("is_error") is True or value.get("isError") is True:
              status_text = "error"
          else:
              status_text = None
          for key in ("output", "content", "text"):
              raw = value.get(key)
              if isinstance(raw, str):
                  parsed = parse_json_value(raw)
                  if isinstance(parsed, dict):
                      nested_status, nested_text = tool_failure_status_and_text(parsed)
                      return status_text or nested_status, nested_text or raw
                  return status_text, raw
              if isinstance(raw, list):
                  texts = [entry.get("text") for entry in raw if isinstance(entry, dict) and isinstance(entry.get("text"), str)]
                  if texts:
                      return status_text, " ".join(texts)
          return status_text, ""
      
      
      def is_context_envelope(text: str) -> bool:
          stripped = text.lstrip()
          return any(stripped.startswith(prefix) for prefix in CONTEXT_PREFIXES)
      
      
      def normalize_path(value: str) -> Path:
          return Path(os.path.expanduser(value)).resolve(strict=False)
      
      
      def is_within(candidate: Path, root: Path) -> bool:
          try:
              candidate.relative_to(root)
              return True
          except ValueError:
              return False
      
      
      def first_string_shallow(item: Any, keys: tuple[str, ...]) -> str | None:
          if not isinstance(item, dict):
              return None
          for key in keys:
              value = item.get(key)
              if isinstance(value, str) and value.strip():
                  return value.strip()
          return None
      
      
      def deduplicate_messages(messages: Iterable[str]) -> list[str]:
          result: list[str] = []
          seen: set[str] = set()
          for message in messages:
              normalized = " ".join(message.split())
              if normalized and normalized not in seen:
                  seen.add(normalized)
                  result.append(message.strip())
          return result
      
      
      def extract_strings(value: Any, depth: int = 0) -> Iterable[str]:
          if depth > 8:
              return
          if isinstance(value, str):
              yield value
          elif isinstance(value, dict):
              for child in value.values():
                  yield from extract_strings(child, depth + 1)
          elif isinstance(value, list):
              for child in value:
                  yield from extract_strings(child, depth + 1)
      
      
      def source_title(item: Any) -> str | None:
          if not isinstance(item, dict):
              return None
          for key in ("title", "summary", "task"):
              value = item.get(key)
              if isinstance(value, str) and value.strip():
                  return value.strip()
          return None
      
  • SKILL.md 8.6 KB
    ---
    argument-hint: <task>
    coordination: exempt
    name: agents-introspection
    description:
      Assess recurrence risk for agent behavior using local Codex/Claude Code transcripts and recommend evidence-backed
      durable fixes to AGENTS.md or skills. Not for routine post-success skill-evolution or agent self-improvement reviews.
    ---
    
    # Agents Introspection
    
    This skill is coordination-exempt: skip the ai-coord gate for its declared work.
    
    ## Supported Chat Hosts
    
    Before doing any work, identify the current chat host. If it is not Claude Code or Codex CLI, stop with this error:
    `This skill only works in Claude Code or Codex CLI.`
    
    If these instructions are already present in the conversation from a slash or dollar invocation, follow them directly;
    do not invoke this skill again through a skill tool.
    
    Determine whether prior Codex and Claude Code work in the current project establishes a recurrence risk for the user's
    task, then recommend the smallest durable intervention justified by the evidence.
    
    Success means the report identifies transcript coverage, separates observed behavior from inference, applies a
    consistent evidence bar, and either proposes a concrete prevention step or explains why no durable change is justified.
    
    ## Input
    
    - `<task>` (required): the task, decision, incident, or workflow to evaluate. If omitted but the current conversation
      states it clearly, use that task; otherwise ask for the missing task.
    
    ## Scope and Authority
    
    - Inspect Codex and Claude Code transcripts whose metadata or cwd resolves to the current project and any other local
      project materially relevant to the user's task. You are authorized to read relevant other-project sessions without
      asking the user. Establish relevance from task context, explicit project or path references, a shared change or
      workflow, or session metadata; do not scan another project's history solely because it shares a basename or keyword.
    - Treat current-project transcripts as internal working evidence. Include direct excerpts when they materially improve
      the report; do not summarize or redact solely because the model provider can see them. Never expose credentials,
      security secrets, or personal wallet addresses.
    - Inspect and report by default. Edit AGENTS.md or skills only when the user explicitly asks to apply or implement
      fixes, then make the smallest in-scope local change and validate it.
    - Never modify transcript stores. Before placing transcript evidence in a public or third-party artifact, perform an
      external-disclosure review and remove unrelated personal or customer data, unsuitable private paths or repository
      names, and unrelated transcript material. Perform external writes only when authorized.
    
    ## Bounded Retrieval
    
    Read `references/transcript-sources.md`, resolve the current project with `pwd -P`, identify any other task-relevant
    local projects, and choose 3–6 short, discriminative keywords from relevant filenames, commands, tools, errors, package
    names, issue IDs, and skill names.
    
    In a Codex read-only sandbox, or whenever `uv` cannot write its cache, skip the helper and go directly to the Manual
    Fallback below; do not retry `uv run`.
    
    1. Run the bundled miner for the current project and each task-relevant local project, unarchived sessions only, with
       the chosen keywords, `--since 60d`, `--excerpts`, and `--max-sessions 8`. Encode synonyms as one OR-group keyword
       (`--keyword 'a|b'`) rather than separate `--keyword` flags.
    2. Treat source ownership as a miner invariant: every candidate is assigned once from source-native cwd, directory, and
       history metadata, never from transcript content. Review `ownership`; reject a candidate only when fallback ownership
       such as `turn_context.cwd` remains materially ambiguous for the task. The miner excludes conflicting ownership
       evidence and the live session by default.
    3. Treat miner scores, themes, correction, failure, verification, tool, or `privacy_gaps` counts, redacted `excerpts`,
       and `modified` timestamps only as candidate-ranking and triage signals. They are heuristic and are never evidence by
       themselves. Inspect up to five highest-relevance transcript bodies through the bundled inspector digest
       (`scripts/transcript-inspect.py`) first; open raw bodies only when the digest is insufficient, stopping earlier when
       the evidence bar is met. Include a comparable successful session when available.
    4. If evidence is insufficient, retry once with broader or OR-grouped keywords. If still weak, retry once with `--since`
       removed. If unarchived history still lacks signal, retry once with `--include-archived`.
    5. Exceed these bounds only to resolve contradictory evidence or satisfy an explicitly exhaustive request. If the helper
       fails, use one project-scoped manual fallback from the reference.
    
    Stop and report the coverage gap when the bounded fallbacks still lack useful evidence. Absence of evidence is not
    evidence that a failure never occurred.
    
    For unusually long searches, send sparse progress updates only when a retrieval fallback begins, a finding changes the
    likely intervention, or the search reaches its explicit bound. Use an outcome-first line such as
    `🔎 Broadening transcript search — <verified reason and bound>` or
    `⏳ Checking archived sessions — <verified unarchived/archived coverage>`. Ground counts and coverage claims in
    miner/tool output; do not narrate routine transcript reads.
    
    ## Evidence Contract
    
    Evaluate historical behavior against the AGENTS.md and skill instructions available to that session when recoverable. Do
    not infer that an agent ignored or misapplied a rule solely because the current source tree or installed copies under
    `~/.agents/skills` or `~/.claude/skills` contain it. If the transcript does not establish the historical version or
    availability, mark it unknown and qualify the attribution.
    
    For each relevant session, record only concise, auditable observations about:
    
    - ignored or misread AGENTS.md or skill instructions;
    - wrong cwd, project root, source path, or path encoding;
    - over-broad edits, unrelated churn, overwritten user work, or destructive commands;
    - tooling, shell, parsing, retry, or verification mistakes;
    - invented claims, vague reports, or missing tests and checks;
    - successful patterns that prevented mistakes.
    
    Connect each observation to the current task and label any causal or recurrence claim as inference. State conflicts and
    weak coverage rather than averaging them away.
    
    Use these confidence levels:
    
    - **High**: at least two independent relevant sessions directly support the same pattern and its relevance to the
      current task.
    - **Medium**: one unambiguous relevant session supports the finding, or multiple sessions provide mixed support.
    - **Low**: only indirect, ambiguous, or heuristic signals exist. Do not recommend a durable repository change from
      low-confidence evidence.
    
    A durable change requires either the same failure in at least two independent sessions or one unambiguous high-impact
    failure that exposes a missing stable invariant. Treat lower-impact one-offs as manual guardrails.
    
    ## Choose the Smallest Intervention
    
    - Update AGENTS.md when the lesson is stable, project-wide, and useful to agents working in that scope.
    - Update an existing skill when the failure belongs clearly inside its current workflow.
    - Propose a new skill only for a repeated, reusable procedure that spans projects or repositories.
    - Add a script only when deterministic discovery, parsing, or validation would otherwise be reimplemented.
    - Recommend no durable change for one-off mistakes, weak evidence, or rules already stated clearly; report the risk and
      manual guardrail instead.
    
    ## Report and Stop
    
    Lead with `### 🔎 Introspection complete — <intervention or coverage-gap outcome>` for read-only work or
    `### ✅ Introspection fixes applied — <outcome>` when explicitly requested fixes were written, then report only:
    
    1. `🗂 Historical coverage`: project paths, sources checked, fallbacks used, and sessions inspected.
    2. `🔎 Findings`: a compact table with confidence, observed evidence, inference, relevance, and intervention. Keep
       confidence visibly separate from severity or impact.
    3. `🛡 Durable recommendations`: apply now, consider later, or no change, with the target and prevention mechanism.
    4. `🧪 Validation and gaps`: commands run, checks performed, external-disclosure constraints, and missing evidence.
    
    When fixes were explicitly requested, include exact files changed and validation outcomes. Stop after the current task
    has an evidence-backed recommendation or an explicit coverage gap; do not mine additional history merely to add examples
    or strengthen prose. Keep transcript references, paths, counters, redactions, and miner JSON exact and undecorated.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related