exploring-codebases
First-encounter orientation on a repository nobody here has worked in yet. Runs a fixed five-step workflow — venv setup, tarball fetch, tree-sitting structural scan, featuring synthesis, then reasoning over the two — and yields an account of what the repo contains and how it is a
Install
npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/code-intelligence/skills/exploring-codebases
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
git clone https://github.com/oaustegard/claude-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.
README
exploring-codebases
Semantic search for codebases. Locates matches with ripgrep and expands them into full AST nodes (functions/classes) using tree-sitter. Returns complete, syntactically valid code blocks rather than fragmented lines. Use when looking for specific implementations, examples, or references where full context is needed.
Skill manifest
Exploring Codebases
Exploratory code analysis for unfamiliar repositories. Orchestrates tree-sitting (structural) and featuring (semantic) over a local copy.
Workflow
Five numbered steps, in order. Do not skip step 0.
0. Setup (once per session)
uv venv /home/claude/.venv 2>/dev/null
uv pip install tree-sitter --python /home/claude/.venv/bin/python
export PYTHON=/home/claude/.venv/bin/python
export TREESIT=/mnt/skills/user/tree-sitting/scripts/treesit.py
export GATHER=/mnt/skills/user/featuring/scripts/gather.py
If step 2's --stats reports Symbols: 0 on a repo you know contains code,
the tree-sitter core package isn't installed — come back here and install it
(the engine bundles its own grammars and does NOT use tree-sitter-language-pack).
Treesit exits 0 and prints no error in that case, so zero symbols is the only
signal you get. There is no Errors: line: that one appears for parse
failures, and an absent parser never reaches parsing. The full signal is in
the tree-sitting skill's Setup section.
1. Get the repo (tarball, not per-file)
OWNER=...
REPO=...
REF=main # branch name, tag, or SHA. For a PR: pull/N/head
curl -sL -H "Authorization: Bearer $GH_TOKEN" \
"https://api.github.com/repos/$OWNER/$REPO/tarball/$REF" -o /tmp/$REPO.tar.gz
mkdir -p /tmp/$REPO && tar -xzf /tmp/$REPO.tar.gz -C /tmp/$REPO --strip-components=1
ls /tmp/$REPO | head # sanity check — did extraction land?
One HTTP call gets the whole repo. Do NOT curl README, cat files, or
fetch via contents/PATH first — they're in the tarball. The
Authorization header is only needed for private repos; public repos
work without it.
Ref selection matters. If exploring a feature branch, PR, or tag,
set REF accordingly. The default main will silently give you stale
code if the question is about an unmerged branch.
2. Structural scan
$PYTHON $TREESIT /tmp/$REPO --stats
Read the output. It gives file counts, symbol counts, languages, and per-directory symbol density. This IS the orienting artifact — treat it as the product of this step, not warm-up.
Drill only if you have a specific question. For pure "what is this repo" exploration, skip drilling and go to step 3 — featuring surfaces the interesting paths for you. Drill when a user asked about a specific subsystem, or when step 3's output raises a question that needs source.
When you do drill, batch queries in one invocation. Every treesit call pays the full scan cost. Multiple queries added to the same command share that scan and each additional query adds ~0ms. If you're about to make a second treesit call on the same path, fold it into the first.
# GOOD — one scan, three answers
$PYTHON $TREESIT /tmp/$REPO --path=SUBDIR --detail=full \
'find:*Handler*:function' 'source:main' 'refs:Config'
# BAD — three scans, three answers (3× the cost for the same information)
$PYTHON $TREESIT /tmp/$REPO --path=SUBDIR --detail=full
$PYTHON $TREESIT /tmp/$REPO 'find:*Handler*:function'
$PYTHON $TREESIT /tmp/$REPO 'refs:Config'
3. Feature synthesis
Pick the mode from your DELIVERABLE, before you run it.
| Your deliverable | Command | Size |
|---|---|---|
| Your own understanding — a review, an orientation read, answering a question | --orient |
~115 lines |
A written _FEATURES.md that must cite every symbol |
full output | thousands of lines |
# Default. Complexity assessment, decomposition ranking, directory tree, entry points.
$PYTHON $GATHER /tmp/$REPO --skip tests,.github,node_modules --orient
# Only when you are about to WRITE the inventory into a file:
$PYTHON $GATHER /tmp/$REPO --skip tests,.github,node_modules --source-budget 8000
Output includes a "Candidate areas for sub-files (by symbol density)" list near the top — that's your drill-target picker, ranked.
Never pipe the full output through head. If you are about to truncate it,
--orient was the correct mode and you have paid for thousands of lines you
will not read. One review's full gather ran to 5,697 lines and was cut at line
120; every finding in it came from treesit drilling and targeted reads
instead. --orient returns the ~115 lines that get used. The full mode's
symbol inventory exists to be CITED, not read.
4. Reason about the combined output
Synthesize 2+3: capabilities, feature groups, architecture, entry
points, anomalies. Produce _FEATURES.md when warranted. This is the
LLM step; everything before was mechanical.
When to Use This vs Other Skills
| Situation | Use |
|---|---|
| "I just cloned this, what is it?" | exploring-codebases (this skill) |
| "Where is the retry logic?" | searching-codebases |
"Find all files matching class.*Error" |
searching-codebases |
| "Show me the symbols in auth.py" | tree-sitting directly |
| "Which files are most about CSRF / sessions / queryset filtering?" | bm25 |
| "Rank these docs by relevance to a multi-word concept" | bm25 |
| "Document what this codebase does" | featuring directly |
| "Teach me this codebase" (a human is learning) | orienting-codebases |
| "Get me this repo" — fetch, no analysis | accessing-github-repos, cloning-project |
Exploring is the divergent skill — you don't know what you're looking for yet. Searching is the convergent skill — you know what you want.
orienting-codebases runs the same tree-sitting + featuring pipeline and is
the nearest thing in the catalogue to this skill. The split is the audience:
this one builds Claude's understanding so work can proceed; that one builds
the user's understanding through guided exercises and HTML artifacts. If
nobody is being taught, this is the right skill.
Pairing bm25 with this workflow
Once steps 2–3 have surfaced the rough shape of the repo, bm25 is the
natural complement when you want ranked content search beyond grep
and beyond exact-symbol lookup. It ranks files by lexical relevance to a
multi-word query, which is useful for "what's this codebase actually
about when I search for X?" — particularly when you don't yet know the
symbol name to feed to tree-sitting.
BM25=/mnt/skills/user/bm25/scripts/bm25.py
# Pass multiple queries — index builds once, all queries reuse it
python3 $BM25 /tmp/$REPO 'auth flow' 'session backend' 'middleware pipeline' \
--exclude 'tests/*' --exclude '*/tests/*' --top-k 5
Two patterns that pair especially well:
- bm25 → tree-sitting. Use bm25 to find the top-ranked files for a
concept; then
tree-sitting source:Symbol:path/to/file.pyto read the actual implementation. - bm25 with
--exclude 'tests/*'. Test directories tend to dominate keyword queries because test names redundantly mention domain terms. Excluding them up front lands you on implementation files.
bm25 is corpus-agnostic — it'll also work on project knowledge stores
or uploads/ if your exploration spans docs, transcripts, or PDFs.
Delegating to subagents
Only when the repo is large (>1000 files or several distinct subsystems) and this environment exposes a subagent tool (Agent/Task in Claude Code and CCotw). Claude.ai chat and bare-skill runs have none: run steps 2-4 inline and skip this entirely. Never simulate fan-out by other means when the tool is absent.
Steps 2-3 stay inline either way. Only step 4's judgment work fans out, one agent per subsystem, and a subagent inherits nothing -- not the conversation, not this file, not the knowledge that scan artifacts are already on disk. Read references/subagent-delegation.md before writing the first agent prompt; it carries the four things every prompt must include and what happens when they are missing.
Notes
- Large repos (>100 files): use
--skip tests,vendored,docs,...in step 2 to focus the scan. - Monorepos: treat each package/service as a separate exploration.
Generate per-subsystem
_FEATURES.mdfiles linked from a root index. - Drill heuristics (if step 2 drilling is warranted): directories
with high symbol-to-file ratio (dense logic), entry-point names
(
main,cli,app,server,routes), files with many imports (integration points).
Files (claude-skills)
-
references
-
subagent-delegation.md 2.2 KB
# Delegating exploration to subagents Loaded from the `exploring-codebases` workflow. Applies only when the repo is large (>1000 files or several distinct subsystems) AND this environment exposes a subagent tool. Steps 2-3 of the main workflow stay inline either way -- they are mechanical and cheap; only step 4's judgment work fans out. **Gate first: does this environment expose a subagent tool** (Agent/Task in Claude Code and CCotw)? Claude.ai chat and bare-skill runs have none — run steps 2–4 inline and skip this section entirely. Never simulate fan-out by other means when the tool is absent. When the tool exists and the repo is large (>1000 files or several distinct subsystems), keep steps 2–3 inline — they're mechanical and cheap — and fan out only step 4's judgment work, one agent per subsystem. **Subagents inherit nothing.** Not your conversation, not this SKILL.md, not the knowledge that scan artifacts exist on disk. An agent prompted only with "explore `crates/foo`" will re-derive structure by `ls`/glob crawling at full tier cost. (Observed 2026-07-16: four Sonnet agents launched onto a 2,300-file repo without the handoff spent their opening turns running `ls`, with the full symbol index already on disk.) Every subagent prompt must therefore carry: 1. **Its structure slice, pre-computed.** Partition the gather output's `## Public API` section by subsystem path prefix, write each slice to a file, and point the agent at its file: "grep/Read this instead of listing directories." Small slices (<50KB) can be pasted inline instead. 2. **The treesit recipe verbatim** — the full command including the venv python path, plus the batch rule (one invocation, many queries; each invocation pays the scan, extra queries are free). 3. **Anti-crawl instructions** — no `ls`/Glob for discovery; `Read` only to confirm or expand a line range the slice or treesit already located. 4. **An output spec** — what to report, a line budget, and "file paths + line refs" so results are verifiable. Routing (see the `agent-routing` skill): multi-turn exploration is outside Haiku's calibrated zone — use `sonnet` for subsystem agents and keep the final cross-cluster synthesis in the orchestrator.
-
-
CHANGELOG.md 3.1 KB
# exploring-codebases - Changelog All notable changes to the `exploring-codebases` skill are documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/). ## [2.5.2] - 2026-09-09 ### Other - prompt-audit: dated prompting patterns across the skill catalogue (#791) ## [2.5.1] - 2026-08-25 ### Other - creating-skill: use Anthropic's quick_validate.py instead of a hand-rolled check (#775) ## [2.5.0] - 2026-08-25 ### Fixed - repair broken frontmatter, mark obsolete skills, close registry gaps (#746) ### Other - Applicability boundaries, real failure signals, findable descriptions (#774) - featuring: fix source=None crash on cache hit, add --orient (#770) ## [2.4.0] - 2026-07-16 ### Added - subagent context-handoff guidance (exploring-codebases 2.4.0, agent-routing 1.2.0) (#733) ## [2.3.1] - 2026-06-21 ### Fixed - install tree-sitter core, not language-pack (step 0) (#708) ## [2.3.0] - 2026-05-20 ### Other - exploring-codebases v2.3.0: add bm25 pairing ## [2.3.0] - 2026-05-18 ### Added - bm25 skill pairing — routing table entries for "which files are most about X?" / "rank by concept" queries, plus a "Pairing bm25 with this workflow" subsection covering the bm25→tree-sitting follow-through pattern and the `--exclude tests/*` lesson from fly #656. ## [2.2.1] - 2026-04-21 ### Other - exploring-codebases v2.2.1: good/bad batching examples (salvaged from #559) (#565) ## [2.2.0] - 2026-04-20 ### Other - exploring-codebases v2.2.0: step-0 setup, ref variants, sanity check, drill-optional (#561) ## [2.1.0] - 2026-04-20 ### Other - exploring-codebases: TODO-style workflow (tarball-first, batched treesit) (#560) - Remove _MAP.md files, direct agents to tree-sitting for code navigation (#545) ## [2.0.0] - 2026-04-08 ### Added - add treesit.py CLI, fix cross-process cache loss, fix Symbol dict bug (#536) ### Other - exploring-codebases: remove dead search.py (replaced by tree-sitting) ## [1.0.0] - 2026-03-31 ### Added - add mapping-features skill for behavioral web app documentation (#432) - implement issues #229, #231, #281, #282, #283 (v4.3.0) ### Other - exploring v1.0.0, searching v2.0.0: tree-sitting replaces mapping-codebases - Regenerate _MAP.md files after @lat: backlink insertion (#504) - Lattice v2: bidirectional source-anchored knowledge graph (#503) ## [0.3.1] - 2026-02-13 ### Fixed - parse multi-line signatures in --use-maps mode ## [0.3.1] - 2026-02-13 ### Fixed - Fix `--use-maps` mode failing to parse multi-line signatures in `_MAP.md` files - Functions with long parameter lists (spanning multiple lines in map files) are now correctly parsed - Full multi-line signatures are accumulated and displayed properly ### Added - Document scope and limitations: which code elements are returned vs not ## [0.3.0] - 2026-02-13 ### Added - Address top 3 priority GitHub issues (#253, #254, #276) - fix three issues - aliases, expansion threshold, return format (v3.7.0) ## [0.2.0] - 2026-02-03 ### Added - Add/Update skill: exploring-codebases - Add/Update skill: exploring-codebases ### Other - Update README with project inspiration details -
README.md 340 B
# exploring-codebases Semantic search for codebases. Locates matches with ripgrep and expands them into full AST nodes (functions/classes) using tree-sitter. Returns complete, syntactically valid code blocks rather than fragmented lines. Use when looking for specific implementations, examples, or references where full context is needed. -
SKILL.md 9.1 KB
--- name: exploring-codebases description: >- First-encounter orientation on a repository nobody here has worked in yet. Runs a fixed five-step workflow — venv setup, tarball fetch, tree-sitting structural scan, featuring synthesis, then reasoning over the two — and yields an account of what the repo contains and how it is arranged, optionally written out as _FEATURES.md. Use for "I just cloned this", "what is this repo", "what does this do", "explore this repo", "give me an orientation", "what are the main features", "review what's new in this repo", or before starting work in a codebase you have not seen. This is the divergent what's-here skill. Route elsewhere for: a named symbol, a file's structure or a line range (tree-sitting); all callers of a Python symbol (searching-codebases); teaching a human the codebase through exercises (orienting-codebases); fetching or cloning a repo without analysing it (accessing-github-repos, cloning-project). metadata: version: 2.5.2 --- # Exploring Codebases Exploratory code analysis for unfamiliar repositories. Orchestrates tree-sitting (structural) and featuring (semantic) over a local copy. ## Workflow Five numbered steps, in order. Do not skip step 0. ### 0. Setup (once per session) ```bash uv venv /home/claude/.venv 2>/dev/null uv pip install tree-sitter --python /home/claude/.venv/bin/python export PYTHON=/home/claude/.venv/bin/python export TREESIT=/mnt/skills/user/tree-sitting/scripts/treesit.py export GATHER=/mnt/skills/user/featuring/scripts/gather.py ``` If step 2's `--stats` reports `Symbols: 0` on a repo you know contains code, the `tree-sitter` core package isn't installed — come back here and install it (the engine bundles its own grammars and does NOT use tree-sitter-language-pack). Treesit exits 0 and prints no error in that case, so zero symbols is the only signal you get. There is no `Errors:` line: that one appears for parse failures, and an absent parser never reaches parsing. The full signal is in the tree-sitting skill's Setup section. ### 1. Get the repo (tarball, not per-file) ```bash OWNER=... REPO=... REF=main # branch name, tag, or SHA. For a PR: pull/N/head curl -sL -H "Authorization: Bearer $GH_TOKEN" \ "https://api.github.com/repos/$OWNER/$REPO/tarball/$REF" -o /tmp/$REPO.tar.gz mkdir -p /tmp/$REPO && tar -xzf /tmp/$REPO.tar.gz -C /tmp/$REPO --strip-components=1 ls /tmp/$REPO | head # sanity check — did extraction land? ``` One HTTP call gets the whole repo. Do NOT curl README, cat files, or fetch via `contents/PATH` first — they're in the tarball. The Authorization header is only needed for private repos; public repos work without it. **Ref selection matters.** If exploring a feature branch, PR, or tag, set `REF` accordingly. The default `main` will silently give you stale code if the question is about an unmerged branch. ### 2. Structural scan ```bash $PYTHON $TREESIT /tmp/$REPO --stats ``` Read the output. It gives file counts, symbol counts, languages, and per-directory symbol density. This IS the orienting artifact — treat it as the product of this step, not warm-up. **Drill only if you have a specific question.** For pure "what is this repo" exploration, skip drilling and go to step 3 — featuring surfaces the interesting paths for you. Drill when a user asked about a specific subsystem, or when step 3's output raises a question that needs source. **When you do drill, batch queries in one invocation.** Every treesit call pays the full scan cost. Multiple queries added to the same command share that scan and each additional query adds ~0ms. If you're about to make a second treesit call on the same path, fold it into the first. ```bash # GOOD — one scan, three answers $PYTHON $TREESIT /tmp/$REPO --path=SUBDIR --detail=full \ 'find:*Handler*:function' 'source:main' 'refs:Config' # BAD — three scans, three answers (3× the cost for the same information) $PYTHON $TREESIT /tmp/$REPO --path=SUBDIR --detail=full $PYTHON $TREESIT /tmp/$REPO 'find:*Handler*:function' $PYTHON $TREESIT /tmp/$REPO 'refs:Config' ``` ### 3. Feature synthesis **Pick the mode from your DELIVERABLE, before you run it.** | Your deliverable | Command | Size | |---|---|---| | Your own understanding — a review, an orientation read, answering a question | `--orient` | ~115 lines | | A written `_FEATURES.md` that must cite every symbol | full output | thousands of lines | ```bash # Default. Complexity assessment, decomposition ranking, directory tree, entry points. $PYTHON $GATHER /tmp/$REPO --skip tests,.github,node_modules --orient # Only when you are about to WRITE the inventory into a file: $PYTHON $GATHER /tmp/$REPO --skip tests,.github,node_modules --source-budget 8000 ``` Output includes a "Candidate areas for sub-files (by symbol density)" list near the top — that's your drill-target picker, ranked. **Never pipe the full output through `head`.** If you are about to truncate it, `--orient` was the correct mode and you have paid for thousands of lines you will not read. One review's full gather ran to 5,697 lines and was cut at line 120; every finding in it came from `treesit` drilling and targeted reads instead. `--orient` returns the ~115 lines that get used. The full mode's symbol inventory exists to be CITED, not read. ### 4. Reason about the combined output Synthesize 2+3: capabilities, feature groups, architecture, entry points, anomalies. Produce `_FEATURES.md` when warranted. This is the LLM step; everything before was mechanical. ## When to Use This vs Other Skills | Situation | Use | |-----------|-----| | "I just cloned this, what is it?" | **exploring-codebases** (this skill) | | "Where is the retry logic?" | searching-codebases | | "Find all files matching `class.*Error`" | searching-codebases | | "Show me the symbols in auth.py" | tree-sitting directly | | "Which files are most about CSRF / sessions / queryset filtering?" | bm25 | | "Rank these docs by relevance to a multi-word concept" | bm25 | | "Document what this codebase does" | featuring directly | | "Teach me this codebase" (a human is learning) | orienting-codebases | | "Get me this repo" — fetch, no analysis | accessing-github-repos, cloning-project | Exploring is the **divergent** skill — you don't know what you're looking for yet. Searching is the **convergent** skill — you know what you want. `orienting-codebases` runs the same tree-sitting + featuring pipeline and is the nearest thing in the catalogue to this skill. The split is the audience: this one builds Claude's understanding so work can proceed; that one builds the *user's* understanding through guided exercises and HTML artifacts. If nobody is being taught, this is the right skill. ### Pairing bm25 with this workflow Once steps 2–3 have surfaced the rough shape of the repo, `bm25` is the natural complement when you want **ranked content search** beyond grep and beyond exact-symbol lookup. It ranks files by lexical relevance to a multi-word query, which is useful for "what's this codebase actually *about* when I search for X?" — particularly when you don't yet know the symbol name to feed to `tree-sitting`. ```bash BM25=/mnt/skills/user/bm25/scripts/bm25.py # Pass multiple queries — index builds once, all queries reuse it python3 $BM25 /tmp/$REPO 'auth flow' 'session backend' 'middleware pipeline' \ --exclude 'tests/*' --exclude '*/tests/*' --top-k 5 ``` Two patterns that pair especially well: 1. **bm25 → tree-sitting.** Use bm25 to find the top-ranked files for a concept; then `tree-sitting source:Symbol:path/to/file.py` to read the actual implementation. 2. **bm25 with `--exclude 'tests/*'`.** Test directories tend to dominate keyword queries because test names redundantly mention domain terms. Excluding them up front lands you on implementation files. bm25 is corpus-agnostic — it'll also work on `project` knowledge stores or `uploads/` if your exploration spans docs, transcripts, or PDFs. ## Delegating to subagents Only when the repo is large (>1000 files or several distinct subsystems) **and** this environment exposes a subagent tool (Agent/Task in Claude Code and CCotw). Claude.ai chat and bare-skill runs have none: run steps 2-4 inline and skip this entirely. Never simulate fan-out by other means when the tool is absent. Steps 2-3 stay inline either way. Only step 4's judgment work fans out, one agent per subsystem, and a subagent inherits nothing -- not the conversation, not this file, not the knowledge that scan artifacts are already on disk. Read [references/subagent-delegation.md](references/subagent-delegation.md) before writing the first agent prompt; it carries the four things every prompt must include and what happens when they are missing. ## Notes - **Large repos (>100 files)**: use `--skip tests,vendored,docs,...` in step 2 to focus the scan. - **Monorepos**: treat each package/service as a separate exploration. Generate per-subsystem `_FEATURES.md` files linked from a root index. - **Drill heuristics** (if step 2 drilling is warranted): directories with high symbol-to-file ratio (dense logic), entry-point names (`main`, `cli`, `app`, `server`, `routes`), files with many imports (integration points).
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.