Claude Skill

featuring

Generate hierarchical _FEATURES.md files that describe what a codebase DOES from a user/consumer perspective, anchored to source symbols via tree-sitting. Supports large complex codebases through feature-driven decomposition into sub-feature files. Uses a multi-pass synthesis: or

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download oaustegard-claude-skills-plugins_code-intelligence_skills_featuring-e39c726.zip · 21 KB
Part of oaustegard/claude-skills — 39 skills

Install

skills CLI npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/code-intelligence/skills/featuring
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
Git git clone https://github.com/oaustegard/claude-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.

README

featuring

Generate hierarchical _FEATURES.md files that describe what a codebase does from a user/consumer perspective, anchored to source symbols via tree-sitting.

Features

  • Top-down feature documentation — organized by capability, not by file or directory structure
  • Hierarchical decomposition — complex capability areas get their own sub-feature files with progressive disclosure
  • Multi-pass synthesis — orientation scan, detailed extraction, then overview rewrite (written last, not first)
  • Symbol anchoring — every feature references key symbols via file#symbol notation
  • Drift detection — check.py validates all symbol references against the live codebase, catching renames, deletions, and uncovered new APIs
  • CI-ready — check script returns exit codes suitable for GitHub Actions or pre-commit hooks

Dependency

Requires the tree-sitting skill for AST-based code scanning.

Relationship to Other Skills

  • tree-sitting provides the structural inventory (what symbols exist)
  • featuring adds the semantic layer (why they exist, what they accomplish together)
  • generating-lattice offers stricter bidirectional traceability when needed

Skill manifest

Featuring

Generate _FEATURES.md files — top-down documentation of what a codebase does, organized by feature/capability, anchored to specific source symbols.

tree-sitting tells you WHAT symbols exist. _FEATURES.md tells you WHY they exist and what they accomplish together.

For large codebases, the root _FEATURES.md decomposes into sub-feature files linked by capability area — not by folder structure. An agent starts at the root and is drawn into sub-files only when working on a relevant area.

Dependency

Requires tree-sitting skill. Uses its engine for AST scanning.

uv venv /home/claude/.venv 2>/dev/null
uv pip install tree-sitter-language-pack --python /home/claude/.venv/bin/python

tree-sitting caches its scan to /tmp/treesit-cache, keyed by repo path plus skip set. That cache persists symbols and imports but NOT file source, so a cache HIT returns source=None. gather.py re-reads those files from disk; do not assume entry.source is populated if you write against the engine directly. (Before this was handled, gather crashed on its second run against a repo while the first succeeded — diagnosed 2026-08-22.)

For quick structural orientation before running gather.py, use tree-sitting's CLI:

TREESIT=/mnt/skills/user/tree-sitting/scripts/treesit.py

# Complete tree, sparse detail — see the full shape
/home/claude/.venv/bin/python $TREESIT /path/to/repo --depth=-1 --detail=sparse

Workflow: Multi-Pass Synthesis

Feature documentation is built in three passes. The overview is written LAST, after all features are understood — not first.

Pass 1: Orientation (quick scan)

/home/claude/.venv/bin/python /mnt/skills/user/featuring/scripts/gather.py /path/to/repo \
  --skip tests,.github,node_modules --source-budget 8000

--orient when you are not writing the file. The full output is a complete symbol inventory — 5,697 lines on a 71k-line repo — and it exists so a _FEATURES.md can cite every symbol. When the deliverable is your own understanding (a review, an orientation read), pass --orient: complexity assessment, decomposition ranking, directory tree and entry points, and nothing else. ~115 lines. Reaching for head on the full output means --orient was the right mode.

Pass 1 of THIS skill is the case that wants the full output — you are about to write the inventory down.

Read the gather output. Before writing anything, form a hypothesis:

"This codebase appears to be a [what it is] that provides [capability A], [capability B], and [capability C]."

Write this down as a DRAFT overview. It will be wrong or incomplete — that's fine. The point is to orient before diving into detail.

How to identify capability areas:

  1. What can a user/consumer DO with this? (commands, API endpoints, UI actions)
  2. What problems does it solve? (the WHY behind the code)
  3. What are the main workflows? (how features compose)
  4. What are the constraints/invariants? (rules the code enforces)

Pass 2: Detailed feature extraction

For each capability area identified in Pass 1:

  1. Gather the symbols that implement it (from gather output + targeted get_source())
  2. Understand how they collaborate — the workflow
  3. Identify constraints and invariants
  4. Write the feature section

During this pass, you'll discover:

  • Capabilities you missed in Pass 1
  • Features that are more complex than expected (decomposition candidates)
  • Features that are simpler than expected (merge candidates)
  • Cross-cutting concerns that span multiple capability areas

Hierarchy decision (per feature, during this pass):

Signal Action
≤6 key symbols, self-contained Inline in root _FEATURES.md
>6 key symbols OR clear sub-capabilities Own _FEATURES.md sub-file
Spans many files but is ONE capability Inline (breadth ≠ complexity)
Has sub-features that are independently useful Own sub-file
Is infrastructure (logging, DB layer) Inline briefly, unless it IS the product

Pass 3: Overview rewrite

NOW — after all features are documented — rewrite the overview. The Pass 1 draft was a hypothesis. Pass 3 replaces it with a proper progressive-disclosure overview that:

  1. States what the codebase is in one sentence
  2. Lists the top-level capability areas (3-8 items)
  3. For each area that has a sub-file: one sentence + link + "read when" guidance
  4. For inline features: just the list entry (detail is below in the same file)

This is the most important part. The overview IS the entry point for every agent session. It must be accurate, complete, and fast to scan.

_FEATURES.md Format

Root file

# Features: {project-name}

> One-sentence description of what this codebase is and does.

**Capability areas:**
- **[Area A]** — one-sentence summary
- **[Area B]** — one-sentence summary → [details](path/to/_FEATURES.md)
- **[Area C]** — one-sentence summary

## {Inline Feature Name}

{2-3 sentences: what this feature does from a user perspective.}

**Key symbols:**
- `file.py#function_name` — role in this feature
- `file.py#ClassName` — role in this feature

**Workflow:** {How a user exercises this feature or how symbols collaborate.}

**Constraints:** {Invariants, limits, rules.}

---

## {Complex Feature Area}

> One-sentence summary of what this area covers.

This area is documented in detail in [{area-name}/_FEATURES.md]({path}).
Read it when working on {specific trigger — e.g., "the memory retrieval pipeline",
"adding a new API endpoint", "modifying the build system"}.

At a glance, this area provides:
- {sub-capability 1} — one line
- {sub-capability 2} — one line
- {sub-capability 3} — one line

Sub-feature files

Sub-feature files follow the SAME format as the root, recursively. They can contain inline features and further sub-file references. Each sub-file:

  • Has its own # Features: {area-name} header
  • Has its own overview paragraph
  • Is self-contained — an agent reading only this file understands the area
  • Links back to the root: ← [Root features](../_FEATURES.md)

Format rules

  • Organized by capability, not by file/directory
  • Symbol references use file#symbol notation (relative to repo root)
  • Leading paragraph per feature: what a user gets, not implementation details
  • Key symbols: the 2-6 most important symbols, with their role explained
  • Workflow: how the feature works end-to-end (include when non-obvious)
  • Constraints: rules/invariants (include when they exist)
  • No source code in _FEATURES.md — it's a map, not a mirror
  • "Read when" guidance on every sub-file link — tells agents WHEN to drill in

What makes a good feature entry

Good: "Memory Storage — Persist observations across sessions. Stores typed, tagged memories to a Turso database with BM25 full-text search. Memories have priority levels that affect retrieval ranking."

Bad: "memory.py — Contains remember(), recall(), forget(), and supersede() functions."

The first tells you WHAT you can do. The second describes file contents — tree-sitting already gives you that.

Hierarchy design principles

The hierarchy is feature-driven, not folder-driven. Folders are natural candidates for decomposition boundaries, but the decision is based on:

  1. Does this capability area have enough complexity to warrant its own file? (>6 key symbols, multiple sub-workflows, or independently useful sub-features)
  2. Would an agent working on this area benefit from focused context? (if yes, a sub-file saves them from parsing unrelated features)
  3. Is this area likely to be read independently of the rest? (if yes, it should be self-contained in its own file)

Counter-examples — do NOT split just because:

  • The code lives in a separate folder (folder ≠ feature)
  • There are many files (files ≠ complexity)
  • A class has many methods (one class = one feature unless methods serve distinct user-facing purposes)

Identifying features

Heuristics for finding feature boundaries:

  • Entry points (main, CLI commands, route handlers) often map 1:1 to features
  • Public API functions that aren't helpers are usually feature surfaces
  • Type hierarchies (class + methods) often represent a cohesive feature
  • Config/constants clusters sometimes reveal features (e.g., a group of timeout constants → a retry feature)
  • Import clusters — files that import each other heavily are likely co-implementing a feature

Features to SKIP in _FEATURES.md:

  • Pure infrastructure (logging, error handling) unless it's the project's purpose
  • Internal utilities that only serve other features
  • Test code (unless the testing approach IS a feature, e.g., a testing framework)

Keeping _FEATURES.md in Sync

Three mechanisms, layered:

1. Check script (detect drift)

/home/claude/.venv/bin/python /mnt/skills/user/featuring/scripts/check.py /path/to/repo \
  [--features _FEATURES.md] [--skip tests,.github]

Parses file#symbol references from ALL _FEATURES.md files (root + sub-files), resolves them against the live codebase via tree-sitting, and reports:

  • Broken refs — symbol deleted or renamed (exit code 1)
  • Moved symbols — symbol exists but in a different file than referenced
  • Dead features — ALL key symbols in a feature section are gone
  • Uncovered symbols — new public API not mentioned in any feature
  • Orphan sub-files — sub-feature files not linked from any parent

Exit code 0 = clean, 1 = drift detected. Suitable for CI or pre-commit hooks.

2. Agent instructions (prevent drift)

Add to CLAUDE.md or equivalent:

## Feature Documentation

- `_FEATURES.md` documents what this codebase does, organized by capability.
- Start here when orienting to the codebase. Follow sub-file links as needed.
- After changing behavior (new feature, renamed API, deleted functionality):
  run `python featuring/scripts/check.py .` and fix any broken refs.
- After adding a new public API surface: add it to the appropriate feature
  section, or create a new feature section if it's a new capability.
- Run check before committing. Broken refs = broken documentation.

3. Targeted regeneration (fix drift)

When check reports broken refs, the fix is usually surgical: update the file#symbol reference to the new name/location. For dead features (all refs gone), either delete the section or regenerate it.

Full regeneration (re-running all three passes) is the nuclear option. Prefer targeted updates — they're cheaper and preserve hand-written narrative.

CI Integration

# .github/workflows/features-check.yml
name: Check _FEATURES.md
on: [push, pull_request]
jobs:
  check:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: astral-sh/setup-uv@v4
      - run: uv pip install tree-sitter-language-pack
      - run: python featuring/scripts/check.py . --skip tests

Claude Code Integration

In Claude Code, use tree-sitting's CLI or engine directly. The agent should:

  1. Run treesit.py /path --depth=-1 --detail=sparse for full structural overview
  2. Pass 1: Form a hypothesis about what the codebase does
  3. Run treesit.py /path --path=DIR --detail=full for each capability area
  4. Run treesit.py /path --no-tree 'source:symbol_name' where intent isn't clear
  5. Pass 2: Write detailed feature sections, deciding hierarchy per-feature
  6. Pass 3: Rewrite the overview now that all features are documented

Add to CLAUDE.md:

## Codebase Understanding

Read `_FEATURES.md` for top-down feature orientation before modifying code.
Follow links to sub-feature files when working on a specific area.
Use tree-sitting MCP tools for structural queries (symbol lookup, source retrieval).
After adding new features or changing behavior, update the relevant _FEATURES.md.

Example: Small Codebase (flat)

A CLI tool with 15 public symbols → single _FEATURES.md, all features inline. No sub-files needed.

Example: Large Codebase (hierarchical)

The remembering skill (memory system for an AI agent) has ~60 public symbols across 8 files. Hierarchical decomposition:

_FEATURES.md              (root — overview + 3 inline features + 3 sub-file refs)
├── scripts/_FEATURES.md  (memory operations — storage, retrieval, lifecycle, maintenance)
└── utils/_FEATURES.md    (utility modules — therapy, reminders, blog publishing)

Root _FEATURES.md would contain:

  • Overview: "Persistent memory system for Muninn. Stores, retrieves, and maintains typed memories across sessions via Turso."
  • Inline: Boot Sequence, Configuration, Task Tracking (simple, ≤4 symbols each)
  • Sub-file ref: Memory Operations → scripts/_FEATURES.md ("Read when working on storage, retrieval, or memory lifecycle")
  • Sub-file ref: Utility Modules → utils/_FEATURES.md ("Read when working on therapy sessions, reminders, or blog publishing")

Relationship to Other Skills

Skill What it provides Drift detection
tree-sitting Structural inventory (symbols, signatures) N/A (live queries)
featuring Feature documentation (what/why), hierarchical check.py — docs → code
generating-lattice Bidirectional knowledge graph lat check — docs ↔ code
mapping-webapp Web app behavioral docs (pages, flows) None

featuring's check is lighter than lattice's: no source code annotations needed, no @lat: comments, just reference resolution. The trade-off is that new code without docs is only flagged as "uncovered symbols" — it's advisory, not enforced. Use lattice when you need strict bidirectional traceability; use featuring when you need good-enough orientation docs that catch renames and deletions.

Files (claude-skills)
  • scripts
    • check.py 13.3 KB
      #!/usr/bin/env python3
      """
      check.py — Validate _FEATURES.md hierarchy against current codebase state.
      
      Discovers all _FEATURES.md files (root + sub-files linked from parents),
      parses symbol references, resolves them via tree-sitting, and reports:
        - Broken refs (symbol renamed, deleted, or moved)
        - Dead features (ALL key symbols gone)
        - Uncovered symbols (new public API not mentioned in any feature)
        - Orphan sub-files (_FEATURES.md files not linked from any parent)
        - Broken sub-file links (links to _FEATURES.md that don't exist)
      
      Usage:
          python check.py /path/to/repo [--features _FEATURES.md] [--skip tests]
      
      Exit codes: 0 = clean, 1 = drift detected
      """
      
      import argparse
      import re
      import sys
      from collections import defaultdict
      from pathlib import Path
      
      # Reuse gather's engine discovery
      from gather import _find_treesit_engine
      
      
      def discover_features_files(root_path: Path, repo: Path) -> list:
          """Discover all _FEATURES.md files by following links from the root.
      
          Returns list of (path, linked_from) tuples.
          """
          # Start with the root
          discovered = [(root_path, None)]
          visited = {root_path.resolve()}
      
          # Sub-file link pattern: [text](path/to/_FEATURES.md) or (path/_FEATURES.md)
          link_pattern = re.compile(r'\]\(([^)]*_FEATURES\.md)\)')
      
          queue = [root_path]
          while queue:
              current = queue.pop(0)
              if not current.exists():
                  continue
      
              text = current.read_text()
              parent_dir = current.parent
      
              for match in link_pattern.finditer(text):
                  rel_link = match.group(1)
                  linked_path = (parent_dir / rel_link).resolve()
      
                  if linked_path not in visited:
                      visited.add(linked_path)
                      discovered.append((linked_path, str(current.relative_to(repo))))
                      if linked_path.exists():
                          queue.append(linked_path)
      
          return discovered
      
      
      def find_all_features_files(repo: Path) -> list:
          """Find ALL _FEATURES.md files in the repo, including unlinked ones."""
          return list(repo.rglob('_FEATURES.md'))
      
      
      def parse_features_file(path: Path) -> dict:
          """Parse _FEATURES.md and extract structure.
      
          Returns:
              {
                  'path': str,
                  'features': [
                      {
                          'name': str,
                          'line': int,
                          'refs': [{'file': str, 'symbol': str, 'line': int}, ...],
                      },
                      ...
                  ],
                  'all_refs': [{'file': str, 'symbol': str, 'feature': str, 'line': int, 'source_file': str}, ...],
                  'sub_links': [{'target': str, 'line': int}, ...]
              }
          """
          text = path.read_text()
          lines = text.split('\n')
      
          ref_pattern = re.compile(r'`([^`]+?)#([^`]+?)`')
          link_pattern = re.compile(r'\]\(([^)]*_FEATURES\.md)\)')
      
          features = []
          all_refs = []
          sub_links = []
          current_feature = None
      
          for i, line in enumerate(lines, 1):
              # Feature headers are ## level
              if line.startswith('## ') and not line.startswith('### '):
                  feature_name = line[3:].strip()
                  if feature_name.lower() in ('feature inventory', 'status summary'):
                      continue
                  current_feature = {
                      'name': feature_name,
                      'line': i,
                      'refs': [],
                  }
                  features.append(current_feature)
      
              # Extract symbol references
              for match in ref_pattern.finditer(line):
                  filepath = match.group(1)
                  symbol = match.group(2)
                  ref = {
                      'file': filepath,
                      'symbol': symbol,
                      'line': i,
                      'feature': current_feature['name'] if current_feature else '(preamble)',
                      'source_file': str(path),
                  }
                  if current_feature:
                      current_feature['refs'].append(ref)
                  all_refs.append(ref)
      
              # Extract sub-file links
              for match in link_pattern.finditer(line):
                  sub_links.append({'target': match.group(1), 'line': i})
      
          return {
              'path': str(path),
              'features': features,
              'all_refs': all_refs,
              'sub_links': sub_links,
          }
      
      
      def resolve_refs(cache, refs: list) -> tuple[list, list]:
          """Check each ref against the codebase.
      
          Returns (resolved, broken) where each is a list of ref dicts
          with an added 'match' key for resolved refs.
          """
          resolved = []
          broken = []
      
          for ref in refs:
              filepath = ref['file']
              symbol_parts = ref['symbol'].split('#')
              symbol_name = symbol_parts[-1]
      
              file_syms = cache.file_symbols(filepath)
              found = False
      
              if file_syms:
                  for sym in file_syms:
                      if sym.name == symbol_name:
                          found = True
                          ref['match'] = sym
                          break
                      for child in sym.children:
                          if child.name == symbol_name:
                              found = True
                              ref['match'] = child
                              break
                      if found:
                          break
      
              if not found:
                  global_matches = cache.find_symbol(symbol_name, limit=3)
                  if global_matches:
                      ref['moved_to'] = global_matches[0].file
                      ref['match'] = global_matches[0]
                      found = True
      
              if found:
                  resolved.append(ref)
              else:
                  broken.append(ref)
      
          return resolved, broken
      
      
      def find_uncovered(cache, all_refs: list, skip_patterns: set = None) -> list:
          """Find public symbols not referenced in any feature."""
          skip_patterns = skip_patterns or set()
      
          referenced = set()
          for ref in all_refs:
              parts = ref['symbol'].split('#')
              referenced.add(parts[-1])
              if len(parts) > 1:
                  referenced.add(parts[0])
      
          uncovered = []
          for relpath, entry in cache.files.items():
              if entry.lang == 'markdown':
                  continue
              if any(p in relpath.lower() for p in ('test', 'spec', '__tests__', 'vendor')):
                  continue
              if any(p in relpath for p in skip_patterns):
                  continue
      
              for sym in entry.symbols:
                  if sym.name.startswith('_'):
                      continue
                  if sym.name not in referenced:
                      uncovered.append(sym)
                  for child in sym.children:
                      if child.name.startswith('_'):
                          continue
                      if child.name not in referenced:
                          if sym.name in referenced:
                              uncovered.append(child)
      
          return uncovered
      
      
      def check(repo_path: str, features_path: str = None,
                skip: set = None, verbose: bool = False) -> dict:
          """Run all checks across the full _FEATURES.md hierarchy."""
      
          engine_path = _find_treesit_engine()
          if engine_path is None:
              print("ERROR: tree-sitting skill not found.", file=sys.stderr)
              sys.exit(1)
          sys.path.insert(0, str(engine_path))
          from engine import CodeCache
      
          repo = Path(repo_path).resolve()
      
          # Find root _FEATURES.md
          if features_path:
              root_fpath = Path(features_path)
          else:
              root_fpath = repo / '_FEATURES.md'
          if not root_fpath.exists():
              return {'error': f'_FEATURES.md not found at {root_fpath}'}
      
          # Discover linked hierarchy
          linked_files = discover_features_files(root_fpath, repo)
      
          # Find ALL _FEATURES.md files in repo (for orphan detection)
          all_files_on_disk = find_all_features_files(repo)
          linked_paths = {p.resolve() for p, _ in linked_files}
      
          # Parse all linked feature files
          all_features = []
          all_refs = []
          all_sub_links = []
          broken_sub_links = []
          file_count = 0
      
          for fpath, linked_from in linked_files:
              if not fpath.exists():
                  broken_sub_links.append({
                      'target': str(fpath.relative_to(repo)),
                      'linked_from': linked_from or '(root)',
                  })
                  continue
              file_count += 1
              parsed = parse_features_file(fpath)
              all_features.extend(parsed['features'])
              all_refs.extend(parsed['all_refs'])
      
              # Check sub-links resolve
              for link in parsed['sub_links']:
                  target = (fpath.parent / link['target']).resolve()
                  if not target.exists():
                      broken_sub_links.append({
                          'target': link['target'],
                          'linked_from': str(fpath.relative_to(repo)),
                          'line': link['line'],
                      })
      
          # Detect orphan _FEATURES.md files
          orphans = [
              str(f.relative_to(repo))
              for f in all_files_on_disk
              if f.resolve() not in linked_paths
          ]
      
          # Scan codebase
          cache = CodeCache()
          cache.scan(str(repo), skip=skip)
      
          # Resolve refs
          resolved, broken = resolve_refs(cache, all_refs)
      
          # Find moved symbols
          moved = [r for r in resolved if 'moved_to' in r]
      
          # Find dead features
          dead_features = []
          for feat in all_features:
              if feat['refs'] and all(r in broken for r in feat['refs']):
                  dead_features.append(feat)
      
          # Find uncovered symbols
          uncovered = find_uncovered(cache, all_refs)
      
          return {
              'features_files': file_count,
              'total_features': len(all_features),
              'total_refs': len(all_refs),
              'resolved': len(resolved),
              'broken': broken,
              'moved': moved,
              'dead_features': dead_features,
              'uncovered': uncovered,
              'orphan_files': orphans,
              'broken_sub_links': broken_sub_links,
              'clean': (len(broken) == 0 and len(broken_sub_links) == 0),
          }
      
      
      def format_report(results: dict) -> str:
          """Format check results as readable report."""
          lines = []
      
          if 'error' in results:
              return f"ERROR: {results['error']}"
      
          lines.append("# _FEATURES.md Check")
          lines.append(f"Feature files: {results['features_files']} | "
                       f"Features: {results['total_features']} | "
                       f"Refs: {results['total_refs']} | "
                       f"Resolved: {results['resolved']} | "
                       f"Broken: {len(results['broken'])}")
          lines.append("")
      
          if results['clean'] and not results['moved'] and not results['orphan_files']:
              lines.append("✓ All symbol references resolve. No drift detected.")
          else:
              if results['broken']:
                  lines.append(f"## Broken References ({len(results['broken'])})")
                  for ref in results['broken']:
                      source = ref.get('source_file', '')
                      lines.append(f"  ✗ `{ref['file']}#{ref['symbol']}` "
                                   f"(line {ref['line']}, feature: {ref['feature']})")
                  lines.append("")
      
              if results['moved']:
                  lines.append(f"## Moved Symbols ({len(results['moved'])})")
                  for ref in results['moved']:
                      lines.append(f"  → `{ref['file']}#{ref['symbol']}` "
                                   f"moved to `{ref['moved_to']}` "
                                   f"(line {ref['line']}, feature: {ref['feature']})")
                  lines.append("")
      
              if results['dead_features']:
                  lines.append(f"## Dead Features ({len(results['dead_features'])})")
                  for feat in results['dead_features']:
                      lines.append(f"  ☠ **{feat['name']}** (line {feat['line']}) "
                                   f"— all {len(feat['refs'])} refs broken")
                  lines.append("")
      
              if results['broken_sub_links']:
                  lines.append(f"## Broken Sub-File Links ({len(results['broken_sub_links'])})")
                  for link in results['broken_sub_links']:
                      from_str = link.get('linked_from', '?')
                      lines.append(f"  ✗ `{link['target']}` linked from {from_str}")
                  lines.append("")
      
          if results['orphan_files']:
              lines.append(f"## Orphan Feature Files ({len(results['orphan_files'])})")
              for orphan in results['orphan_files']:
                  lines.append(f"  ? `{orphan}` — not linked from any parent _FEATURES.md")
              lines.append("")
      
          if results['uncovered']:
              lines.append(f"## Uncovered Public Symbols ({len(results['uncovered'])})")
              by_file = defaultdict(list)
              for sym in results['uncovered']:
                  by_file[sym.file].append(sym)
              for filepath in sorted(by_file.keys()):
                  syms = by_file[filepath]
                  names = ', '.join(s.name for s in syms[:8])
                  if len(syms) > 8:
                      names += f', ... +{len(syms) - 8}'
                  lines.append(f"  {filepath}: {names}")
              lines.append("")
              lines.append("These symbols appear in the public API but aren't referenced "
                           "in any _FEATURES.md file.")
      
          return '\n'.join(lines)
      
      
      def main():
          parser = argparse.ArgumentParser(description='Check _FEATURES.md hierarchy against codebase')
          parser.add_argument('repo', help='Path to codebase root')
          parser.add_argument('--features', default=None, help='Path to root _FEATURES.md')
          parser.add_argument('--skip', default='', help='Comma-separated dirs to skip')
          parser.add_argument('-v', '--verbose', action='store_true')
          args = parser.parse_args()
      
          skip = set(args.skip.split(',')) if args.skip else None
          results = check(args.repo, features_path=args.features, skip=skip, verbose=args.verbose)
          print(format_report(results))
          sys.exit(0 if results.get('clean', False) else 1)
      
      
      if __name__ == '__main__':
          main()
      
    • gather.py 15.2 KB
      #!/usr/bin/env python3
      """
      gather.py — Collect structural data from a codebase via tree-sitting for feature synthesis.
      
      Scans a codebase with tree-sitting's engine, then outputs a structured summary
      optimized for LLM consumption: entry points, public APIs, symbol clusters by
      directory, and selective source excerpts for key files.
      
      Supports two modes:
        - Full scan: entire codebase → root _FEATURES.md synthesis
        - Area scan: specific subdirectory or file set → sub-feature file synthesis
      
      Usage:
          # Full repo scan
          python gather.py /path/to/repo [--skip tests,.github] [--source-budget 8000]
      
          # Area scan (specific capability area)
          python gather.py /path/to/repo --area src/memory --source-budget 4000
      
      Output: structured markdown to stdout, suitable as LLM prompt input.
      """
      
      import argparse
      import sys
      from pathlib import Path
      
      
      def _find_treesit_engine():
          """Locate tree-sitting engine across known skill directories."""
          candidates = [
              Path(__file__).resolve().parent.parent.parent / 'tree-sitting' / 'scripts',
              Path('/mnt/skills/user/tree-sitting/scripts'),
              Path('/mnt/skills/public/tree-sitting/scripts'),
          ]
          for p in candidates:
              if (p / 'engine.py').exists():
                  return p
          return None
      
      
      def setup_engine():
          """Import tree-sitting engine."""
          engine_path = _find_treesit_engine()
          if engine_path is None:
              print("ERROR: tree-sitting skill not found. Install it first.", file=sys.stderr)
              sys.exit(1)
          sys.path.insert(0, str(engine_path))
          try:
              from engine import CodeCache
              return CodeCache()
          except ImportError as e:
              print(f"ERROR: tree-sitting engine import failed: {e}\n"
                    "Install deps: uv pip install tree-sitter-language-pack", file=sys.stderr)
              sys.exit(1)
      
      
      def classify_symbols(cache, area_prefix=None) -> dict:
          """Classify symbols into feature-relevant categories.
      
          If area_prefix is set, only include symbols from files under that path.
          """
          categories = {
              'entry_points': [],
              'public_api': [],
              'types': [],
              'constants': [],
              'tests': [],
              'internal': [],
          }
      
          entry_patterns = {'main', 'cli', 'run', 'serve', 'start', 'app', 'init', 'setup', 'boot'}
          test_patterns = {'test_', 'test', 'spec_', 'it_'}
      
          for relpath, entry in cache.files.items():
              # Area filter
              if area_prefix and not relpath.startswith(area_prefix):
                  continue
              if entry.lang == 'markdown':
                  continue
              is_test_file = any(p in relpath.lower() for p in ('test', 'spec', '__tests__'))
              for sym in entry.symbols:
                  name_lower = sym.name.lower()
      
                  if is_test_file or any(name_lower.startswith(p) for p in test_patterns):
                      categories['tests'].append(sym)
                  elif name_lower in entry_patterns or (name_lower == '__main__'):
                      categories['entry_points'].append(sym)
                  elif sym.kind in ('class', 'struct', 'enum', 'interface', 'trait', 'type'):
                      categories['types'].append(sym)
                  elif sym.kind in ('constant', 'define', 'static'):
                      categories['constants'].append(sym)
                  elif sym.name.startswith('_'):
                      categories['internal'].append(sym)
                  else:
                      categories['public_api'].append(sym)
      
          return categories
      
      
      def _source_bytes(repo_path, relpath, entry) -> bytes:
          """Raw bytes of one scanned file.
      
          ``entry.source`` is populated only on a COLD scan: tree-sitting's on-disk
          cache (/tmp/treesit-cache) persists lang/symbols/imports and deliberately
          drops source, so every cache HIT hands back ``source=None``. Both call sites
          here used it unguarded, which made gather.py crash on its SECOND run against
          the same repo+skip set while the first succeeded -- diagnosed 2026-08-22,
          reproduced A/B/C: cold rc=0 (5697 lines), warm rc=1 (TypeError).
          Re-read from disk on the miss; return b"" if the file is gone.
          """
          if entry is not None and entry.source is not None:
              return entry.source
          try:
              return (Path(repo_path) / relpath).read_bytes()
          except OSError:
              return b""
      
      
      def identify_key_files(cache, source_budget: int, area_prefix=None, repo_path=None) -> list:
          """Pick files worth reading source from, within a token budget."""
          file_scores = []
          for relpath, entry in cache.files.items():
              if area_prefix and not relpath.startswith(area_prefix):
                  continue
              if any(p in relpath.lower() for p in ('test', 'spec', '__tests__', 'vendor', 'node_modules')):
                  continue
              if entry.lang in ('json', 'yaml', 'toml', 'css', 'html', 'markdown'):
                  continue
      
              score = 0
              for sym in entry.symbols:
                  name_lower = sym.name.lower()
                  if name_lower in ('main', 'cli', 'run', 'serve', 'boot', 'app'):
                      score += 5
                  elif sym.kind in ('class', 'struct', 'trait', 'interface'):
                      score += 3 + len(sym.children)
                  elif not sym.name.startswith('_'):
                      score += 1
      
              if score > 0:
                  source_len = len(_source_bytes(repo_path, relpath, entry))
                  file_scores.append((relpath, score, source_len))
      
          file_scores.sort(key=lambda x: x[1], reverse=True)
          selected = []
          remaining = source_budget
          for relpath, score, size in file_scores:
              char_estimate = size
              if char_estimate <= remaining:
                  selected.append(relpath)
                  remaining -= char_estimate
              elif remaining > 2000:
                  selected.append(relpath)
                  break
          return selected
      
      
      def compute_complexity(cache, area_prefix=None) -> dict:
          """Compute complexity metrics to help decide hierarchy.
      
          Returns dict with counts and a suggested decomposition.
          """
          public_count = 0
          file_count = 0
          type_count = 0
          dir_symbols = {}  # dir → count of public symbols
      
          for relpath, entry in cache.files.items():
              if area_prefix and not relpath.startswith(area_prefix):
                  continue
              if entry.lang in ('json', 'yaml', 'toml', 'css', 'html', 'markdown'):
                  continue
              if any(p in relpath.lower() for p in ('test', 'spec', '__tests__', 'vendor')):
                  continue
      
              file_count += 1
              parent = str(Path(relpath).parent) if '/' in relpath else '.'
      
              for sym in entry.symbols:
                  if sym.name.startswith('_'):
                      continue
                  public_count += 1
                  dir_symbols[parent] = dir_symbols.get(parent, 0) + 1
                  if sym.kind in ('class', 'struct', 'enum', 'interface', 'trait', 'type'):
                      type_count += 1
      
          # Suggest decomposition if complex enough
          suggestions = []
          if public_count > 30:
              # Find directory clusters with significant symbol counts
              for d, count in sorted(dir_symbols.items(), key=lambda x: -x[1]):
                  if count >= 6:
                      suggestions.append((d, count))
      
          return {
              'public_symbols': public_count,
              'files': file_count,
              'types': type_count,
              'dir_clusters': dir_symbols,
              'decomposition_candidates': suggestions,
              'needs_hierarchy': public_count > 20 or len(suggestions) > 2,
          }
      
      
      def format_symbol_brief(sym, indent=0) -> str:
          """One-line symbol summary."""
          prefix = '  ' * indent
          parts = [f"{prefix}- **{sym.name}** ({sym.kind})"]
          if sym.signature:
              parts.append(f"`{sym.signature}`")
          parts.append(f"@ {sym.file}:{sym.line}")
          if sym.doc:
              parts.append(f"— {sym.doc}")
          return ' '.join(parts)
      
      
      def gather(repo_path: str, skip: set = None, source_budget: int = 8000,
                 area: str = None, orient: bool = False) -> str:
          """Scan a codebase and produce structured output for feature synthesis.
      
          Args:
              repo_path: Root of the codebase to scan
              skip: Directory names to skip
              source_budget: Approximate char budget for source excerpts
              area: Optional subdirectory to focus on (for sub-feature generation)
              orient: Stop after the orientation header (complexity assessment,
                  decomposition ranking, directory tree, entry points) and omit the
                  full symbol inventory. ~115 lines instead of thousands. Use when the
                  deliverable is your own understanding; use the full output only when
                  you are writing a _FEATURES.md that must cite every symbol.
          """
          cache = setup_engine()
      
          stats = cache.scan(repo_path, skip=skip)
          if stats['files'] == 0:
              return f"No parseable files found in {repo_path}"
      
          # Normalize area prefix
          area_prefix = area.rstrip('/') + '/' if area else None
          if area_prefix == './':
              area_prefix = None
      
          categories = classify_symbols(cache, area_prefix)
          key_files = identify_key_files(cache, source_budget, area_prefix, repo_path)
          complexity = compute_complexity(cache, area_prefix)
      
          lines = []
      
          # ── Header
          root_name = Path(repo_path).name
          if area:
              lines.append(f"# Structural Scan: {root_name}/{area}")
          else:
              lines.append(f"# Structural Scan: {root_name}")
          lines.append(f"Files: {stats['files']} | Symbols: {stats['symbols']} | "
                       f"Languages: {', '.join(stats['languages'])}")
          lines.append("")
      
          # ── Complexity assessment (helps LLM decide hierarchy)
          lines.append("## Complexity Assessment")
          lines.append(f"Public symbols: {complexity['public_symbols']} | "
                       f"Source files: {complexity['files']} | "
                       f"Types: {complexity['types']}")
          if complexity['needs_hierarchy']:
              lines.append("**Hierarchy recommended.** This codebase has enough complexity "
                           "for sub-feature files.")
              if complexity['decomposition_candidates']:
                  lines.append("\nCandidate areas for sub-files (by symbol density):")
                  for d, count in complexity['decomposition_candidates']:
                      lines.append(f"  - `{d}/` — {count} public symbols")
          else:
              lines.append("**Flat structure sufficient.** A single _FEATURES.md should work.")
          lines.append("")
      
          # ── Directory structure
          lines.append("## Directory Structure")
          lines.append(cache.tree_overview())
          lines.append("")
      
          # ── Entry points
          if categories['entry_points']:
              lines.append("## Entry Points")
              for sym in categories['entry_points']:
                  lines.append(format_symbol_brief(sym))
              lines.append("")
      
          # Everything above is the orientation header: what this repo is, where it is
          # dense, and where execution starts. Everything below is the full symbol
          # inventory, which exists to be CITED in a _FEATURES.md -- not read end to end.
          if orient:
              lines.append(f"_Orientation only ({len(categories['public_api'])} public symbols and "
                           f"{len(categories['types'])} types withheld). Re-run without --orient "
                           f"to emit the full inventory._")
              return "\n".join(lines)
      
          # ── Public API (grouped by file)
          if categories['public_api']:
              lines.append("## Public API")
              by_file = {}
              for sym in categories['public_api']:
                  by_file.setdefault(sym.file, []).append(sym)
              for filepath in sorted(by_file.keys()):
                  lines.append(f"\n### {filepath}")
                  for sym in by_file[filepath]:
                      lines.append(format_symbol_brief(sym))
                      for child in sym.children[:5]:
                          lines.append(format_symbol_brief(child, indent=1))
                      if len(sym.children) > 5:
                          lines.append(f"    ... +{len(sym.children) - 5} more methods")
              lines.append("")
      
          # ── Types
          if categories['types']:
              lines.append("## Types & Data Structures")
              for sym in categories['types']:
                  lines.append(format_symbol_brief(sym))
                  for child in sym.children[:5]:
                      lines.append(format_symbol_brief(child, indent=1))
                  if len(sym.children) > 5:
                      lines.append(f"    ... +{len(sym.children) - 5} more")
              lines.append("")
      
          # ── Key file sources
          if key_files:
              lines.append("## Key Source Excerpts")
              lines.append(f"*{len(key_files)} files selected by structural importance*\n")
              for filepath in key_files:
                  entry = cache.files.get(filepath)
                  if not entry:
                      continue
                  source_text = _source_bytes(repo_path, filepath, entry).decode('utf-8', errors='replace')
                  source_lines = source_text.split('\n')
                  if len(source_lines) > 150:
                      excerpt = '\n'.join(source_lines[:150])
                      lines.append(f"### {filepath} (first 150/{len(source_lines)} lines)")
                  else:
                      excerpt = source_text
                      lines.append(f"### {filepath}")
                  lines.append(f"```{entry.lang}")
                  lines.append(excerpt)
                  lines.append("```\n")
      
          # ── Import graph summary
          lines.append("## Import Graph (internal dependencies)")
          repo_modules = set()
          for relpath in cache.files:
              stem = Path(relpath).stem
              if stem != '__init__':
                  repo_modules.add(stem)
              parts = Path(relpath).parts
              for p in parts[:-1]:
                  repo_modules.add(p)
      
          for relpath, entry in sorted(cache.files.items()):
              if area_prefix and not relpath.startswith(area_prefix):
                  continue
              if entry.lang == 'markdown':
                  continue
              internal = [imp for imp in entry.imports
                          if imp.startswith('.') or
                          imp.split('.')[0] in repo_modules]
              if internal:
                  lines.append(f"- {relpath} ← {', '.join(internal[:10])}")
      
          return '\n'.join(lines)
      
      
      def main():
          parser = argparse.ArgumentParser(description='Gather codebase structure for feature synthesis')
          parser.add_argument('repo', help='Path to codebase root')
          parser.add_argument('--skip', default='', help='Comma-separated dirs to skip')
          parser.add_argument('--area', default=None,
                              help='Focus on a specific subdirectory (for sub-feature generation)')
          parser.add_argument('--source-budget', type=int, default=8000,
                              help='Approximate char budget for source excerpts (default: 8000)')
          parser.add_argument('--orient', action='store_true',
                              help='Orientation header only (~115 lines): complexity, decomposition '
                                   'ranking, directory tree, entry points. Omits the symbol inventory. '
                                   'Use this unless you are writing a _FEATURES.md.')
          args = parser.parse_args()
      
          skip = set(args.skip.split(',')) if args.skip else None
          out = gather(args.repo, skip=skip, source_budget=args.source_budget,
                       area=args.area, orient=args.orient)
          if not args.orient and out.count("\n") > 400:
              # Diagnosed 2026-08-22 (FreeToken review): a 5,697-line gather was piped
              # through `head -120` and the rest paid for and discarded. Say the size up
              # front so the cost of NOT using --orient is visible before the scroll.
              print(f"<!-- {out.count(chr(10)) + 1} lines. Reading only the top? "
                    f"Re-run with --orient. -->")
          print(out)
      
      
      if __name__ == '__main__':
          main()
      
  • CHANGELOG.md 3.1 KB
    # Changelog
    
    ## 0.4.0 (2026-08-22)
    
    ### Fixed: gather.py crashed on the second run against a repo
    tree-sitting's `/tmp/treesit-cache` persists symbols and imports but not file
    source, so every cache HIT returned `entry.source = None`. `identify_key_files`
    (`len(entry.source)`) and the Key Source Excerpts emitter
    (`entry.source.decode(...)`) both used it unguarded. A cold run succeeded and
    wrote the cache; the next run with the same repo + skip set exited 1 with a
    TypeError, which read as flakiness. `_source_bytes()` now falls back to reading
    the file from disk. Verified cold/warm output is byte-identical.
    
    ### New: `--orient`
    Emits the orientation header only — complexity assessment, decomposition
    ranking, directory tree, entry points — and stops before the symbol inventory.
    ~115 lines against 5,697 for the full output on a 71k-line repo. Use it whenever
    the deliverable is your own understanding rather than a written `_FEATURES.md`.
    The full mode now prints its own line count first when the output exceeds 400
    lines, so the cost of not using `--orient` is visible before the scroll.
    
    ## 0.2.0 (2026-03-31)
    
    ### Hierarchical features support
    - Root `_FEATURES.md` can link to sub-feature files for complex capability areas
    - Sub-files follow the same format recursively, with back-links to parent
    - Hierarchy is feature/capability driven, not folder driven
    
    ### Multi-pass synthesis
    - Pass 1: Orientation scan → form hypothesis about what codebase does
    - Pass 2: Detailed feature extraction, per-capability hierarchy decisions
    - Pass 3: Overview rewrite with progressive disclosure (written LAST, not first)
    
    ### gather.py
    - New `--area` flag for focused sub-directory scanning
    - New "Complexity Assessment" section in output with hierarchy recommendation
    - `compute_complexity()` function identifies decomposition candidates by symbol density
    
    ### check.py
    - Discovers and validates full _FEATURES.md hierarchy (root + all linked sub-files)
    - Detects orphan _FEATURES.md files not linked from any parent
    - Detects broken sub-file links
    - Reports now show which source file contains broken refs
    
    ### Examples
    - Split example into root (_FEATURES_example_root.md) and sub-file (_FEATURES_example_sub.md)
    - Demonstrates progressive disclosure: root has summaries + links, sub-file has full detail
    
    ## 0.1.0 (2026-03-29)
    
    Initial release.
    - gather.py: AST-based structural scanning via tree-sitting
    - check.py: drift detection (broken refs, dead features, uncovered symbols)
    - Single flat _FEATURES.md format
    - CI integration example
    
    ## [0.4.0] - 2026-08-22
    
    ### Other
    
    - featuring: fix source=None crash on cache hit, add --orient (#770)
    - Deprecate mapping-codebases; adopt ruff 0.16.0 baseline (#747)
    - Remove _MAP.md files, direct agents to tree-sitting for code navigation (#545)
    
    ## [0.3.0] - 2026-04-08
    
    ### Added
    
    - add treesit.py CLI, fix cross-process cache loss, fix Symbol dict bug (#536)
    
    ### Other
    
    - marketplace: restructure as category-based plugins for Claude Code discovery (#530)
    - Add missing READMEs for searching-codebases, featuring, tree-sitting (#521)
    
    ## [0.2.0] - 2026-03-31
    
    ### Added
    
    - v0.2.0 — Hierarchical features + multi-pass synthesis (#515)
  • README.md 1.2 KB
    # featuring
    
    Generate hierarchical `_FEATURES.md` files that describe what a codebase **does** from a user/consumer perspective, anchored to source symbols via tree-sitting.
    
    ## Features
    
    - **Top-down feature documentation** — organized by capability, not by file or directory structure
    - **Hierarchical decomposition** — complex capability areas get their own sub-feature files with progressive disclosure
    - **Multi-pass synthesis** — orientation scan, detailed extraction, then overview rewrite (written last, not first)
    - **Symbol anchoring** — every feature references key symbols via `file#symbol` notation
    - **Drift detection** — `check.py` validates all symbol references against the live codebase, catching renames, deletions, and uncovered new APIs
    - **CI-ready** — check script returns exit codes suitable for GitHub Actions or pre-commit hooks
    
    ## Dependency
    
    Requires the **tree-sitting** skill for AST-based code scanning.
    
    ## Relationship to Other Skills
    
    - **tree-sitting** provides the structural inventory (what symbols exist)
    - **featuring** adds the semantic layer (why they exist, what they accomplish together)
    - **generating-lattice** offers stricter bidirectional traceability when needed
    
  • SKILL.md 14.3 KB
    ---
    name: featuring
    description: >-
      Generate hierarchical _FEATURES.md files that describe what a codebase DOES from a
      user/consumer perspective, anchored to source symbols via tree-sitting. Supports large
      complex codebases through feature-driven decomposition into sub-feature files. Uses a
      multi-pass synthesis: orientation → detail → overview rewrite. Use when someone says
      "what does this do", "document features", "feature inventory", "_FEATURES.md",
      or needs to understand a codebase's purpose before modifying it. Complements
      tree-sitting (structural) with semantic (why/what-for) layer.
    metadata:
      version: 0.4.0
    ---
    
    # Featuring
    
    Generate `_FEATURES.md` files — top-down documentation of what a codebase **does**,
    organized by feature/capability, anchored to specific source symbols.
    
    **tree-sitting** tells you WHAT symbols exist.
    **_FEATURES.md** tells you WHY they exist and what they accomplish together.
    
    For large codebases, the root `_FEATURES.md` decomposes into sub-feature files
    linked by capability area — not by folder structure. An agent starts at the root
    and is drawn into sub-files only when working on a relevant area.
    
    ## Dependency
    
    Requires **tree-sitting** skill. Uses its engine for AST scanning.
    
    ```bash
    uv venv /home/claude/.venv 2>/dev/null
    uv pip install tree-sitter-language-pack --python /home/claude/.venv/bin/python
    ```
    
    tree-sitting caches its scan to `/tmp/treesit-cache`, keyed by repo path plus
    skip set. That cache persists symbols and imports but NOT file source, so a
    cache HIT returns `source=None`. `gather.py` re-reads those files from disk;
    do not assume `entry.source` is populated if you write against the engine
    directly. (Before this was handled, gather crashed on its second run against a
    repo while the first succeeded — diagnosed 2026-08-22.)
    
    For quick structural orientation before running gather.py, use tree-sitting's CLI:
    
    ```bash
    TREESIT=/mnt/skills/user/tree-sitting/scripts/treesit.py
    
    # Complete tree, sparse detail — see the full shape
    /home/claude/.venv/bin/python $TREESIT /path/to/repo --depth=-1 --detail=sparse
    ```
    
    ## Workflow: Multi-Pass Synthesis
    
    Feature documentation is built in three passes. The overview is written LAST,
    after all features are understood — not first.
    
    ### Pass 1: Orientation (quick scan)
    
    ```bash
    /home/claude/.venv/bin/python /mnt/skills/user/featuring/scripts/gather.py /path/to/repo \
      --skip tests,.github,node_modules --source-budget 8000
    ```
    
    **`--orient` when you are not writing the file.** The full output is a complete
    symbol inventory — 5,697 lines on a 71k-line repo — and it exists so a
    `_FEATURES.md` can cite every symbol. When the deliverable is your own
    understanding (a review, an orientation read), pass `--orient`: complexity
    assessment, decomposition ranking, directory tree and entry points, and nothing
    else. ~115 lines. Reaching for `head` on the full output means `--orient` was
    the right mode.
    
    Pass 1 of THIS skill is the case that wants the full output — you are about to
    write the inventory down.
    
    Read the gather output. Before writing anything, form a hypothesis:
    
    > "This codebase appears to be a **[what it is]** that provides **[capability A]**,
    > **[capability B]**, and **[capability C]**."
    
    Write this down as a DRAFT overview. It will be wrong or incomplete — that's fine.
    The point is to orient before diving into detail.
    
    **How to identify capability areas:**
    
    1. **What can a user/consumer DO with this?** (commands, API endpoints, UI actions)
    2. **What problems does it solve?** (the WHY behind the code)
    3. **What are the main workflows?** (how features compose)
    4. **What are the constraints/invariants?** (rules the code enforces)
    
    ### Pass 2: Detailed feature extraction
    
    For each capability area identified in Pass 1:
    
    1. Gather the symbols that implement it (from gather output + targeted `get_source()`)
    2. Understand how they collaborate — the workflow
    3. Identify constraints and invariants
    4. Write the feature section
    
    During this pass, you'll discover:
    - Capabilities you missed in Pass 1
    - Features that are more complex than expected (decomposition candidates)
    - Features that are simpler than expected (merge candidates)
    - Cross-cutting concerns that span multiple capability areas
    
    **Hierarchy decision** (per feature, during this pass):
    
    | Signal | Action |
    |--------|--------|
    | ≤6 key symbols, self-contained | Inline in root `_FEATURES.md` |
    | >6 key symbols OR clear sub-capabilities | Own `_FEATURES.md` sub-file |
    | Spans many files but is ONE capability | Inline (breadth ≠ complexity) |
    | Has sub-features that are independently useful | Own sub-file |
    | Is infrastructure (logging, DB layer) | Inline briefly, unless it IS the product |
    
    ### Pass 3: Overview rewrite
    
    NOW — after all features are documented — rewrite the overview. The Pass 1
    draft was a hypothesis. Pass 3 replaces it with a proper progressive-disclosure
    overview that:
    
    1. States what the codebase is in one sentence
    2. Lists the top-level capability areas (3-8 items)
    3. For each area that has a sub-file: one sentence + link + "read when" guidance
    4. For inline features: just the list entry (detail is below in the same file)
    
    This is the most important part. The overview IS the entry point for every
    agent session. It must be accurate, complete, and fast to scan.
    
    
    ## _FEATURES.md Format
    
    ### Root file
    
    ```markdown
    # Features: {project-name}
    
    > One-sentence description of what this codebase is and does.
    
    **Capability areas:**
    - **[Area A]** — one-sentence summary
    - **[Area B]** — one-sentence summary → [details](path/to/_FEATURES.md)
    - **[Area C]** — one-sentence summary
    
    ## {Inline Feature Name}
    
    {2-3 sentences: what this feature does from a user perspective.}
    
    **Key symbols:**
    - `file.py#function_name` — role in this feature
    - `file.py#ClassName` — role in this feature
    
    **Workflow:** {How a user exercises this feature or how symbols collaborate.}
    
    **Constraints:** {Invariants, limits, rules.}
    
    ---
    
    ## {Complex Feature Area}
    
    > One-sentence summary of what this area covers.
    
    This area is documented in detail in [{area-name}/_FEATURES.md]({path}).
    Read it when working on {specific trigger — e.g., "the memory retrieval pipeline",
    "adding a new API endpoint", "modifying the build system"}.
    
    At a glance, this area provides:
    - {sub-capability 1} — one line
    - {sub-capability 2} — one line
    - {sub-capability 3} — one line
    ```
    
    ### Sub-feature files
    
    Sub-feature files follow the SAME format as the root, recursively. They can
    contain inline features and further sub-file references. Each sub-file:
    
    - Has its own `# Features: {area-name}` header
    - Has its own overview paragraph
    - Is self-contained — an agent reading only this file understands the area
    - Links back to the root: `← [Root features](../_FEATURES.md)`
    
    ### Format rules
    
    - **Organized by capability**, not by file/directory
    - **Symbol references** use `file#symbol` notation (relative to repo root)
    - **Leading paragraph** per feature: what a user gets, not implementation details
    - **Key symbols**: the 2-6 most important symbols, with their role explained
    - **Workflow**: how the feature works end-to-end (include when non-obvious)
    - **Constraints**: rules/invariants (include when they exist)
    - **No source code** in _FEATURES.md — it's a map, not a mirror
    - **"Read when" guidance** on every sub-file link — tells agents WHEN to drill in
    
    ### What makes a good feature entry
    
    Good: "**Memory Storage** — Persist observations across sessions. Stores typed,
    tagged memories to a Turso database with BM25 full-text search. Memories have
    priority levels that affect retrieval ranking."
    
    Bad: "**memory.py** — Contains `remember()`, `recall()`, `forget()`, and
    `supersede()` functions."
    
    The first tells you WHAT you can do. The second describes file contents —
    tree-sitting already gives you that.
    
    ### Hierarchy design principles
    
    The hierarchy is **feature-driven**, not folder-driven. Folders are natural
    candidates for decomposition boundaries, but the decision is based on:
    
    1. **Does this capability area have enough complexity to warrant its own file?**
       (>6 key symbols, multiple sub-workflows, or independently useful sub-features)
    2. **Would an agent working on this area benefit from focused context?**
       (if yes, a sub-file saves them from parsing unrelated features)
    3. **Is this area likely to be read independently of the rest?**
       (if yes, it should be self-contained in its own file)
    
    Counter-examples — do NOT split just because:
    - The code lives in a separate folder (folder ≠ feature)
    - There are many files (files ≠ complexity)
    - A class has many methods (one class = one feature unless methods serve
      distinct user-facing purposes)
    
    
    ## Identifying features
    
    Heuristics for finding feature boundaries:
    
    - **Entry points** (main, CLI commands, route handlers) often map 1:1 to features
    - **Public API functions** that aren't helpers are usually feature surfaces
    - **Type hierarchies** (class + methods) often represent a cohesive feature
    - **Config/constants clusters** sometimes reveal features (e.g., a group of
      timeout constants → a retry feature)
    - **Import clusters** — files that import each other heavily are likely
      co-implementing a feature
    
    Features to SKIP in _FEATURES.md:
    - Pure infrastructure (logging, error handling) unless it's the project's purpose
    - Internal utilities that only serve other features
    - Test code (unless the testing approach IS a feature, e.g., a testing framework)
    
    
    ## Keeping _FEATURES.md in Sync
    
    Three mechanisms, layered:
    
    ### 1. Check script (detect drift)
    
    ```bash
    /home/claude/.venv/bin/python /mnt/skills/user/featuring/scripts/check.py /path/to/repo \
      [--features _FEATURES.md] [--skip tests,.github]
    ```
    
    Parses `file#symbol` references from ALL _FEATURES.md files (root + sub-files),
    resolves them against the live codebase via tree-sitting, and reports:
    
    - **Broken refs** — symbol deleted or renamed (exit code 1)
    - **Moved symbols** — symbol exists but in a different file than referenced
    - **Dead features** — ALL key symbols in a feature section are gone
    - **Uncovered symbols** — new public API not mentioned in any feature
    - **Orphan sub-files** — sub-feature files not linked from any parent
    
    Exit code 0 = clean, 1 = drift detected. Suitable for CI or pre-commit hooks.
    
    ### 2. Agent instructions (prevent drift)
    
    Add to CLAUDE.md or equivalent:
    
    ```markdown
    ## Feature Documentation
    
    - `_FEATURES.md` documents what this codebase does, organized by capability.
    - Start here when orienting to the codebase. Follow sub-file links as needed.
    - After changing behavior (new feature, renamed API, deleted functionality):
      run `python featuring/scripts/check.py .` and fix any broken refs.
    - After adding a new public API surface: add it to the appropriate feature
      section, or create a new feature section if it's a new capability.
    - Run check before committing. Broken refs = broken documentation.
    ```
    
    ### 3. Targeted regeneration (fix drift)
    
    When check reports broken refs, the fix is usually surgical: update the
    `file#symbol` reference to the new name/location. For dead features (all refs
    gone), either delete the section or regenerate it.
    
    Full regeneration (re-running all three passes) is the nuclear option.
    Prefer targeted updates — they're cheaper and preserve hand-written narrative.
    
    ### CI Integration
    
    ```yaml
    # .github/workflows/features-check.yml
    name: Check _FEATURES.md
    on: [push, pull_request]
    jobs:
      check:
        runs-on: ubuntu-latest
        steps:
          - uses: actions/checkout@v4
          - uses: astral-sh/setup-uv@v4
          - run: uv pip install tree-sitter-language-pack
          - run: python featuring/scripts/check.py . --skip tests
    ```
    
    ## Claude Code Integration
    
    In Claude Code, use tree-sitting's CLI or engine directly. The agent should:
    
    1. Run `treesit.py /path --depth=-1 --detail=sparse` for full structural overview
    2. **Pass 1:** Form a hypothesis about what the codebase does
    3. Run `treesit.py /path --path=DIR --detail=full` for each capability area
    4. Run `treesit.py /path --no-tree 'source:symbol_name'` where intent isn't clear
    5. **Pass 2:** Write detailed feature sections, deciding hierarchy per-feature
    6. **Pass 3:** Rewrite the overview now that all features are documented
    
    Add to CLAUDE.md:
    ```markdown
    ## Codebase Understanding
    
    Read `_FEATURES.md` for top-down feature orientation before modifying code.
    Follow links to sub-feature files when working on a specific area.
    Use tree-sitting MCP tools for structural queries (symbol lookup, source retrieval).
    After adding new features or changing behavior, update the relevant _FEATURES.md.
    ```
    
    ## Example: Small Codebase (flat)
    
    A CLI tool with 15 public symbols → single `_FEATURES.md`, all features inline.
    No sub-files needed.
    
    ## Example: Large Codebase (hierarchical)
    
    The `remembering` skill (memory system for an AI agent) has ~60 public symbols
    across 8 files. Hierarchical decomposition:
    
    ```
    _FEATURES.md              (root — overview + 3 inline features + 3 sub-file refs)
    ├── scripts/_FEATURES.md  (memory operations — storage, retrieval, lifecycle, maintenance)
    └── utils/_FEATURES.md    (utility modules — therapy, reminders, blog publishing)
    ```
    
    Root `_FEATURES.md` would contain:
    - **Overview**: "Persistent memory system for Muninn. Stores, retrieves, and
      maintains typed memories across sessions via Turso."
    - **Inline**: Boot Sequence, Configuration, Task Tracking (simple, ≤4 symbols each)
    - **Sub-file ref**: Memory Operations → `scripts/_FEATURES.md`
      ("Read when working on storage, retrieval, or memory lifecycle")
    - **Sub-file ref**: Utility Modules → `utils/_FEATURES.md`
      ("Read when working on therapy sessions, reminders, or blog publishing")
    
    ## Relationship to Other Skills
    
    | Skill | What it provides | Drift detection |
    |-------|-----------------|-----------------|
    | **tree-sitting** | Structural inventory (symbols, signatures) | N/A (live queries) |
    | **featuring** | Feature documentation (what/why), hierarchical | `check.py` — docs → code |
    | **generating-lattice** | Bidirectional knowledge graph | `lat check` — docs ↔ code |
    | **mapping-webapp** | Web app behavioral docs (pages, flows) | None |
    
    featuring's check is lighter than lattice's: no source code annotations needed,
    no `@lat:` comments, just reference resolution. The trade-off is that new code
    without docs is only flagged as "uncovered symbols" — it's advisory, not
    enforced. Use lattice when you need strict bidirectional traceability; use
    featuring when you need good-enough orientation docs that catch renames and
    deletions.
    
  • _FEATURES_example_root.md 4.6 KB
    # Features: remembering
    
    > Persistent memory system for an AI agent (Muninn). Stores typed, tagged, prioritized memories in a Turso database with BM25 full-text search, and loads identity/operational config at conversation start.
    
    **Capability areas:**
    - **[Memory Operations](#memory-operations)** — store, retrieve, evolve, and maintain memories → [details](scripts/_FEATURES.md)
    - **Boot Sequence** — load identity and ops at conversation start
    - **Configuration** — two-table architecture for boot-loaded settings vs searchable memories
    - **Task Tracking** — structural forcing function for multi-step work
    - **Session Management** — save, resume, export, import conversation checkpoints
    
    
    ## Memory Operations
    
    > The core capability: storing observations, querying them back, evolving them over time, and keeping the store healthy.
    
    This area is documented in detail in [scripts/_FEATURES.md](scripts/_FEATURES.md).
    Read it when working on storage, retrieval, memory lifecycle, maintenance, or decision tracing.
    
    At a glance:
    - **Storage** — `remember()`, batch storage, background writes
    - **Retrieval** — BM25 search with tag/type/time filters, proactive hints
    - **Lifecycle** — forget, supersede, reprioritize, strengthen/weaken
    - **Maintenance** — consolidate, curate, prune, diagnostics
    - **Decision Tracing** — structured capture with alternatives and reference chains
    
    ---
    
    ## Boot Sequence
    
    Load identity (profile) and operational instructions (ops) from the config table at conversation start. Groups ops entries by cognitive domain for organized output.
    
    **Key symbols:**
    - `scripts/boot.py#boot` — Main entry point. Loads profile + ops, detects GitHub access, installs utilities, surfaces reminders.
    - `scripts/boot.py#profile` — Load profile config entries.
    - `scripts/boot.py#ops` — Load operational config entries, grouped by topic.
    - `scripts/boot.py#classify_ops_key` — Route an ops key to its cognitive domain.
    - `scripts/utilities.py#install_utilities` — Materialize utility-code memories to importable Python files.
    
    **Workflow:** `boot()` calls `_exec_batch` to load profile and ops in one HTTP request, groups ops by topic, detects environment capabilities (GitHub, env files), installs utilities from memory, and returns formatted context for the conversation window.
    
    ---
    
    ## Configuration
    
    Two-table architecture: `config` stores boot-loaded identity and operational settings; `memories` stores searchable observations. Config entries have categories (profile, ops, journal), boot_load flags, and priority for ordering.
    
    **Key symbols:**
    - `scripts/config.py#config_get` — Retrieve a config value by key.
    - `scripts/config.py#config_set` — Store a config value with category, optional char limit, and read-only flag.
    - `scripts/config.py#config_delete` — Remove a config entry.
    - `scripts/config.py#config_list` — List entries, optionally filtered by category.
    
    **Constraints:** Categories are: profile, ops, journal. Boot_load controls whether an entry appears in the boot context window.
    
    ---
    
    ## Task Tracking
    
    Structural forcing function for multi-step work. Tasks have named steps, type-specific checklists, and a completion gate that prevents finishing without storing results.
    
    **Key symbols:**
    - `scripts/task.py#Task` — Core class with steps, completion tracking, and persistence.
    - `scripts/task.py#task` — Factory function to create a tracked task.
    - `scripts/task.py#task_resume` — Load a persisted task for cross-session continuity.
    
    **Workflow:** `t = task("analyze X", steps=["research", "synthesize", "store"])` creates a Task. Call `t.done("research")` as steps complete. `t.complete()` gates on all required steps (including store).
    
    ---
    
    ## Session Management
    
    Save and resume conversation checkpoints. Export and import full system state.
    
    **Key symbols:**
    - `scripts/boot.py#session_save` — Save a checkpoint with summary and context.
    - `scripts/boot.py#session_resume` — Resume from the most recent checkpoint.
    - `scripts/boot.py#muninn_export` — Export all state (memories + config) as portable JSON.
    - `scripts/boot.py#muninn_import` — Import state from exported JSON, with optional merge mode.
    
    ---
    
    ## Database Layer
    
    All persistence goes through a Turso (libSQL) HTTP API. Memories use FTS5 for full-text search. The schema supports soft-delete, versioning via supersede chains, and batch operations.
    
    **Key symbols:**
    - `scripts/turso.py` — HTTP client for Turso: `_exec()`, `_exec_batch()`, `_fts5_search()`.
    - `scripts/bootstrap.py#create_tables` — Schema creation (memories + config tables).
    - `scripts/bootstrap.py#migrate_schema` — Add columns for version upgrades.
    
  • _FEATURES_example_sub.md 4.2 KB
    # Features: Memory Operations
    
    ← [Root features](../_FEATURES.md)
    
    > The core memory pipeline: storing observations, querying them back with flexible filters, evolving memories over time, and keeping the store healthy as it grows.
    
    ## Memory Storage
    
    Store observations, facts, decisions, and experiences that persist across conversations. Each memory has a type (world, decision, analysis, etc.), tags for retrieval, a confidence score, and a priority that affects ranking.
    
    **Key symbols:**
    - `scripts/memory.py#remember` — Primary storage entry point. Validates type, generates embedding-ready summary, writes to Turso.
    - `scripts/memory.py#remember_batch` — Bulk storage in a single HTTP round-trip for multi-memory operations.
    
    **Workflow:** Caller provides a summary string, a type, and optional tags/refs/priority. The function generates a UUID, timestamps it, writes to the memories table with FTS5 indexing, and returns the ID. Background mode defers the write to a thread.
    
    **Constraints:** Type is required (enforced, not defaulted). Priority defaults to 0; range is -1 to 2. Confidence defaults to 0.9 if omitted.
    
    ---
    
    ## Memory Retrieval
    
    Query stored memories by text search, tags, type, time range, or combination. BM25 full-text search handles fuzzy matching; tag filtering supports any/all modes.
    
    **Key symbols:**
    - `scripts/memory.py#recall` — Primary query interface with flexible filters (search, tags, type, time, session).
    - `scripts/memory.py#recall_batch` — Execute multiple search queries in a single HTTP round-trip.
    - `scripts/memory.py#recall_since` — Time-windowed retrieval for recent memories.
    - `scripts/hints.py#recall_hints` — Proactive memory surfacing based on context terms.
    - `scripts/result.py#MemoryResult` — Type-safe wrapper providing attribute access and field validation.
    
    **Workflow:** `recall("search terms", tags=["topic"], n=10)` queries FTS5 with BM25 ranking, applies tag/type/confidence filters, orders by composite score (BM25 × priority weight), and returns `MemoryResultList`.
    
    **Constraints:** Parameter is `n=` not `limit=`. Tag mode defaults to "any" (OR). Strict mode raises on empty results.
    
    ---
    
    ## Memory Lifecycle
    
    Evolve memories over time: soft-delete, supersede with updated versions, adjust priority up or down.
    
    **Key symbols:**
    - `scripts/memory.py#forget` — Soft-delete by full or partial UUID.
    - `scripts/memory.py#supersede` — Replace a memory with an updated version, preserving lineage via refs.
    - `scripts/memory.py#reprioritize` — Adjust priority directly.
    - `scripts/memory.py#strengthen` — Increment priority (used during therapy and reinforcement).
    - `scripts/memory.py#weaken` — Decrement priority.
    
    **Workflow:** `supersede(old_id, new_summary, type)` creates a new memory with a ref pointing to the original, then soft-deletes the original. The chain is traversable via `get_chain()`.
    
    ---
    
    ## Memory Maintenance
    
    Autonomous curation, consolidation, and pruning to keep the memory store healthy as it grows.
    
    **Key symbols:**
    - `scripts/memory.py#consolidate` — Cluster related memories by tag overlap and merge into summary memories.
    - `scripts/memory.py#curate` — Autonomous pipeline: detect duplicates, stale memories, consolidation opportunities.
    - `scripts/memory.py#prune_by_age` — Remove old low-priority memories (dry_run by default).
    - `scripts/memory.py#memory_histogram` — Distribution of memories by type, priority, and age for diagnostics.
    
    **Constraints:** All destructive operations default to `dry_run=True`. Consolidation requires `min_cluster=3` memories to trigger.
    
    ---
    
    ## Decision Tracing
    
    Structured capture of decisions with context, rationale, alternatives, and trade-offs. Enables post-hoc review of why choices were made.
    
    **Key symbols:**
    - `scripts/memory.py#decision_trace` — Store a formatted decision with choice/context/rationale/alternatives.
    - `scripts/memory.py#get_alternatives` — Extract rejected alternatives from a decision's refs.
    - `scripts/memory.py#get_chain` — Follow reference chains to build a context graph around a memory.
    
    **Workflow:** `decision_trace(choice, context, rationale, alternatives=[...])` creates a decision-type memory with standardized format and "decision-trace" tag. Later, `get_chain()` traverses refs to reconstruct the decision graph.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related