Claude Skill

codebase-onboarding

Use when you land in an unfamiliar or inherited codebase and must get productive fast: a breadth-first map of entry points, request flow, module ownership, hidden side effects (cron, webhooks, workers) and churn hotspots, committed as CODEBASE-MAP.md. NOT a deep audit of one modu

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies

#architecture

Virus-scanned Reviewed automatically before listing.

Full trust report

Download ericrisco-rsc-harness-skills_codebase-onboarding-953fef5.zip · 10 KB
Part of ericrisco/rsc-harness — 46 skills

Install

skills CLI npx skills add https://github.com/ericrisco/rsc-harness/tree/main/skills/codebase-onboarding
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ericrisco-rsc-harness@llmmart
Git git clone https://github.com/ericrisco/rsc-harness.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ericrisco/rsc-harness collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Codebase onboarding — get oriented fast, leave a map

You have just landed in a codebase you did not write: a fresh clone, an inherited project, an acquired repo, an abandoned side project someone handed you. The instinct is to start reading files top-to-bottom. Resist it. That is how a week disappears and you still cannot answer "where does X happen". This skill runs a disciplined breadth-first reconnaissance pass and produces one durable artifact: a map a teammate can trust and you can re-read tomorrow.

The payoff is measured. Engineers using AI to onboard reach the same milestones roughly 2x faster — productive in 1–2 weeks instead of 4–6 — and the biggest gains are exactly in searching for code, decoding undocumented patterns, and tracing data flows (super-productivity.com, accessed 2026-06-02). That is what the recon pass below targets, in order.

Lead with the deliverable

Before you grep a single line, know the target: a single living file, CODEBASE-MAP.md, committed at the repo root. You work backward from its sections — every recon step fills one. Minimal schema:

# CODEBASE-MAP.md — <repo name>

## Stack          # languages, framework + versions, package manager, run scripts
## Entry points   # main / server bootstrap / route registration / CLI commands
## Request flow   # one real path traced transport -> business logic -> persistence
## Module ownership   # who owns transport / business logic / persistence / UI
## Hidden behavior    # cron, webhooks, queue workers, event listeners, env branches
## Hotspots       # most-churned + most-complex files = highest risk
## How to run     # the exact commands to boot it and hit one path locally

Why a file and not a chat answer: a map that lives only in the conversation dies when the session ends, and the next agent re-does the work. The artifact is the point. verify.sh checks these sections exist (structure, not content).

Two operating rules

  1. Breadth before depth. First pass maps where things are, not how they work. You are drawing the subway map, not reading every passenger's diary. Depth is analyze/debug work, on demand, later. — Reading everything is the failure mode onboarding exists to replace.
  2. Hypothesis before answer. Spend ~5 minutes forming your own guess ("auth probably lives in src/middleware"), then grep to confirm or kill it. — Verifying a hypothesis builds the mental model that makes you fast; a handed-to-you answer does not stick (martinfowler.com, Böckeler, accessed 2026-06-02).

The recon pass — ordered

Run these in order. Each step writes one map section. Stop escalating the moment the section is answerable.

a. Orient — read the manifest, size the repo. Read the manifest(s) and lockfile, not the README first: package.json / pyproject.toml / go.mod / Gemfile / pom.xml tell you the real stack, framework version and run scripts; the lockfile tells you what is actually installed. Then size it:

scc --by-file --sort lines .   # LOC, complexity, COCOMO estimate per file

scc (Sloc Cloc and Code, pure-Go, v3.7.0 Apr 2026) is the fast structural counter of record — materially faster than cloc/tokei and it reports per-file complexity, which you reuse for hotspots. Why manifest-first: the README describes intent (often stale); the manifest describes reality.

b. Find the entry points. Where does execution start? Look for main, the server bootstrap, route registration, CLI command definitions. Let the framework's convention guide you (Next.js app//pages/, Express app.use/router mounts, Django urls.py, FastAPI @app/APIRouter, Rails routes.rb, Spring @RestController). See references/recon-playbook.md for per-ecosystem patterns.

c. Trace one real request end-to-end. Pick a single meaningful path (a login, a checkout, the main CLI command) and follow it: transport (route/handler) → business logic → persistence → response. One path traced beats ten skimmed. This is the spine of the map.

d. Map module ownership. For the directories scc flagged as large, label each: transport, business logic, persistence, UI, shared/util. You are answering "if I need to change pricing, which folder do I open" — the question teammates actually ask.

e. Hunt hidden behavior. The bugs live in what runs without a request. Grep for cron schedules, webhook receivers, queue/background workers, event listeners and env-driven branches (see the appendix). Why this step is non-negotiable: side effects are invisible in a top-down read and they are where inherited codebases bite.

f. Rank hotspots. Git churn is the cheapest risk signal — no extra tooling, and high-churn files are a proxy for "lacks tests/abstraction":

git log --format=format: --name-only --since=12.month \
  | grep -v '^$' | sort | uniq -c | sort -nr | head -50

The richer move is churn × complexity: the top-right quadrant (changes constantly and is hard to read) is your real danger zone (understandlegacycode.com Hotspots, accessed 2026-06-02). Escalate to that — or to a tree-sitter dependency graph — only when grep + churn is not enough: references/recon-playbook.md has the churn×complexity recipe (code-maat) and when codegraph PageRank / FileScopeMCP earn their setup cost.

g. Confirm hands-on. Run the app, walk one end-user journey, send a real request, watch the logs. Reading alone leaves the map unverified; a single real request validates the whole trace in step c.

Scope the effort

Match the pass to the repo. Do not stand up heavy tooling on a small project.

Repo shape Map fully Skip / defer Escalate to a graph tool?
Tiny (<20 files) a, b, c, g churn, ownership table No — grep is faster than setup
Single-service app all a–g — Only if ownership is unclear after grep
Large monorepo a, b, then per-package c–f mapping every package at once Yes — codegraph PageRank to find the load-bearing packages
Polyglot a, b, c per language boundary one unified flow diagram Yes, if cross-language calls obscure the flow

Command appendix

Language-tagged, copy-ready. Per-ecosystem depth lives in references/recon-playbook.md.

# Size + complexity (reuse the complexity column for hotspots)
scc --by-file --sort complexity .

# Route registration (adjust per framework)
rg -n "app\.(get|post|put|delete|use)\(|@app\.(get|post)|APIRouter|router\.(get|post)" --type-add 'web:*.{js,ts,py}' -tweb

# Cron / scheduled jobs
rg -n "cron|schedule|@scheduled|setInterval|celery\.beat|node-cron" -i

# Webhook receivers
rg -n "webhook|/hooks/|stripe.*signature|x-hub-signature" -i

# Queue / background workers
rg -n "queue|worker|bull|sidekiq|celery|sqs|rabbitmq|kafka|@task" -i

# Env-driven branches (hidden config-conditional behavior)
rg -n "process\.env\.|os\.environ|ENV\[|getenv" 

Writing & maintaining the map

  • Commit it. CODEBASE-MAP.md at the repo root, in version control. A map outside the repo rots silently.
  • Keep it living. When the recon reveals you guessed wrong, fix the line — the map is the record of what is true now, not your first impression.
  • Link out, do not duplicate. The map says what is. For why a choice was made, write an ADR (../decision-records/SKILL.md). To scaffold project tooling and a wiki, that is ../harness/SKILL.md. To write the agent-memory CLAUDE.md, that is ../init/SKILL.md — onboarding is the broader recon that feeds it. For generic note/wiki capture, ../knowledge-ops/SKILL.md.
  • Hand off to depth tools. Once the map exists, a deep correctness/security read of one module is ../analyze/SKILL.md; chasing a specific failure through the system is ../debug/SKILL.md. Onboarding builds the map you debug with.

Anti-patterns

Bad Why it bites Good
Read every file top-to-bottom Burns the week; you finish exhausted and still can't trace one request Breadth-first: map locations first, depth on demand
Trust the README over the code READMEs drift; the manifest and the routes are the truth Read manifest + lockfile first, confirm by grep
Map everything at full depth Analysis paralysis on a monorepo; you map dead modules Trace one real flow end-to-end; expand only where needed
Skip the run step An unverified map is a hypothesis, not a map Boot it, hit one path, watch logs before you trust the trace
Map lives only in chat Dies with the session; next agent redoes it Write & commit CODEBASE-MAP.md
Guess instead of grep Confident-wrong is worse than slow-right Form the hypothesis, then rg to confirm or kill it
Stand up a tree-sitter MCP graph on a 5-file repo Setup costs more than the whole recon Reserve graph tools for large monorepos; grep + churn first
Files (rsc-harness)
  • evals
    • cases.yaml 4.4 KB
      skill: codebase-onboarding
      
      # Prompts that MUST load `codebase-onboarding`. This is the breadth-first
      # reconnaissance skill: it maps an unfamiliar/inherited codebase (entry points,
      # request/data flow, module ownership, hidden side effects, hotspots) into a
      # committed CODEBASE-MAP.md. Several prompts deliberately avoid "onboarding" to
      # test that the intent — orientation in code you didn't write — is the trigger.
      should_trigger:
        - prompt: "I just cloned this repo and have no idea where anything is — help me get oriented."
          why: "Canonical cold-start: fresh clone, no mental model, needs breadth-first orientation — the core job."
      
        - prompt: "I inherited this project from a dev who left. Map it out before I change anything."
          why: "Inherited/legacy framing with the explicit 'map before touching' instinct this skill enforces."
      
        - prompt: "Where does the business logic actually live in this codebase, versus the transport layer?"
          why: "Module-ownership question (recon step d) phrased without naming the skill."
      
        - prompt: "What happens behind the scenes when a request hits /checkout — any cron jobs or webhooks involved?"
          why: "Non-obvious hidden-side-effects framing (step e) plus a request trace (step c); side effects are invisible in a top-down read."
      
        - prompt: "acabo de heredar este repo y no sé por dónde empezar, mápamelo."
          why: "Spanish, inherited codebase, explicit 'map it' — same cold-start intent in another language."
      
        - prompt: "Which files are the dangerous ones here — the stuff that changes constantly and nobody understands?"
          why: "Hotspot/churn×complexity framing (step f); 'changes constantly' is the churn signal, 'nobody understands' is complexity."
      
        - prompt: "Give me a high-level architecture map of this monorepo so I can find the load-bearing packages."
          why: "Large-monorepo recon that escalates to a structural-importance graph — the decision-table large-repo row."
      
      # NEAR-MISS prompts that must NOT load `codebase-onboarding`. Each routes to the
      # genuinely correct sibling. The recurring trap: anything that audits depth,
      # scaffolds tooling, records a decision, or chases a specific failure rather than
      # building breadth-first orientation.
      should_not_trigger:
        - prompt: "Review this PR diff for correctness bugs before we merge."
          route_to: "analyze"
          why: "Depth audit of a specific change, not breadth orientation of an unfamiliar repo."
      
        - prompt: "Set up 01-TOOLS and 02-DOCS and generate the root CLAUDE.md for this workspace."
          route_to: "harness"
          why: "Operational tooling/wiki scaffolding, not understanding — onboarding produces a map, not a tooling layer."
      
        - prompt: "Write an ADR for why we picked Postgres over Mongo."
          route_to: "decision-records"
          why: "Capturing the 'why' of a decision; the map captures 'what is', not the rationale."
      
        - prompt: "The login flow throws a 500 intermittently — find the bug."
          route_to: "debug"
          why: "Chasing a specific failure through the system; onboarding builds the map you debug with, it does not root-cause."
      
        - prompt: "Create the initial CLAUDE.md memory file for this repo."
          route_to: "init"
          why: "Narrow single-file agent memory; onboarding is the broader recon that feeds init, not the file write itself."
      
      capability:
        - scenario: >
            Given an unfamiliar Node/Express + Postgres repo (no prior context), produce
            the onboarding map without reading every file line-by-line.
          must_include:
            - "Identifies stack + framework version from package.json and the lockfile, not the README first"
            - "Sizes the repo with scc (per-file LOC/complexity)"
            - "Locates the entry point and route registration (app.get/post/use or router mounts)"
            - "Traces one real request end-to-end: transport handler -> business logic -> persistence"
            - "Labels module ownership (transport / business logic / persistence / UI)"
            - "Surfaces at least one hidden behavior (cron / webhook / queue worker / listener) via ripgrep"
            - "Ranks hotspots using the git churn one-liner (and notes churn×complexity as the richer move)"
            - "States the exact commands to run it locally and hit one path"
            - "Writes results into CODEBASE-MAP.md with the required sections (Stack, Entry points, Request flow, Module ownership, Hidden behavior, Hotspots, How to run)"
            - "Does NOT attempt a line-by-line full read; works breadth-first"
      
    • README.md 1 KB
      # Evals — codebase-onboarding
      
      These cases are split into routing checks and a capability rubric, run through the repo's standard skill-eval harness. The `should_trigger` and `should_not_trigger` prompts validate the description's routing: each `should_trigger` prompt must select this skill (including the non-obvious hotspot/hidden-behavior and Spanish phrasings), and each `should_not_trigger` prompt must route to the named sibling (`analyze`, `harness`, `decision-records`, `debug`, `init`) instead — that is how we confirm the boundary lines hold. The `capability` block is graded by a model against the `must_include` rubric, not asserted by a script: given an unfamiliar Node/Express + Postgres repo, the run is judged on whether it produced a breadth-first map (stack, entry points, one traced flow, ownership, a hidden side effect, churn hotspots, run instructions) written to `CODEBASE-MAP.md` — and crucially did *not* read the repo line-by-line. Run the whole file with the harness; there is no separate setup.
      
  • references
    • recon-playbook.md 4.4 KB
      # Recon playbook — per-ecosystem patterns & escalation
      
      Offloaded depth for the recon pass in `../SKILL.md`. Pull the rows you need; do not run them all.
      
      ## Entry points & hidden side effects by ecosystem
      
      ### Node / Express / Next.js
      - **Entry**: `package.json` `scripts.start`/`dev` → the bootstrap file (`server.js`, `src/index.ts`). Next.js: routes are files under `app/` (route handlers `route.ts`) or `pages/api/`.
      - **Routes**: `rg -n "app\.(get|post|put|delete|use)\(|router\.(get|post|use)\(" src`
      - **Side effects**: `setInterval`, `node-cron`, BullMQ (`new Queue`/`Worker`), `EventEmitter.on(`, Next.js `middleware.ts`, `instrumentation.ts`, and `next.config` rewrites/redirects.
      
      ### Python / Django / FastAPI
      - **Entry (Django)**: `manage.py` → `settings.ROOT_URLCONF` → `urls.py` trees. **FastAPI**: the `FastAPI()` instance + `APIRouter` includes; `uvicorn` target in the run command.
      - **Routes**: `rg -n "@app\.(get|post|put|delete)|APIRouter|path\(|re_path\(|router\.register" .`
      - **Side effects**: Celery (`@shared_task`, `celery.beat` schedule), Django signals (`@receiver`, `post_save`), management commands (`management/commands/`), `apps.py` `ready()`, middleware list in settings.
      
      ### Ruby / Rails
      - **Entry**: `config/routes.rb`, `config/application.rb`, initializers in `config/initializers/`.
      - **Side effects**: Sidekiq/ActiveJob workers (`app/jobs`, `*_worker.rb`), `whenever` cron schedule, ActiveRecord callbacks (`after_save`, `before_create`), `config/schedule.rb`.
      
      ### Go
      - **Entry**: `func main()` (find with `rg -n "func main\(\)"`), the router setup (`chi`, `gin`, `mux`, stdlib `http.HandleFunc`).
      - **Side effects**: goroutines launched at boot (`go func()`), `time.Ticker`/`cron`, `init()` functions (run before main, easy to miss).
      
      ### JVM / Spring Boot
      - **Entry**: the `@SpringBootApplication` main class; `application.yml`/`.properties` for active profiles.
      - **Routes**: `rg -n "@RestController|@RequestMapping|@GetMapping|@PostMapping" src`
      - **Side effects**: `@Scheduled`, `@EventListener`, `@PostConstruct`, `@KafkaListener`/`@RabbitListener`, `@Async`.
      
      ### PHP / Laravel
      - **Entry**: `routes/web.php`, `routes/api.php`, `public/index.php`.
      - **Side effects**: `app/Console/Kernel.php` schedule, jobs in `app/Jobs`, model events/observers (`app/Observers`), event listeners in `EventServiceProvider`.
      
      ## Churn × complexity (the richer hotspot signal)
      
      Git churn alone ranks files by change frequency. Multiply by complexity to find the true danger zone — files that change constantly *and* are hard to read.
      
      ```bash
      # 1. Churn (revisions per file) over a year, via code-maat
      git log --pretty=format:'[%h] %an %ad %s' --date=short --numstat --since=12.month > /tmp/log.txt
      java -jar code-maat.jar -l /tmp/log.txt -c git2 -a revisions > /tmp/churn.csv
      
      # 2. Complexity per file (reuse scc, or indentation as a cheap proxy)
      scc --by-file --format csv . > /tmp/complexity.csv
      ```
      
      Join the two CSVs on filename; the **top-right quadrant** (high revisions, high complexity) is where bugs concentrate and where tests are missing. Start refactoring/test-writing there, not in the file that merely looks ugly. (adamtornhill/code-maat; understandlegacycode.com Hotspots Analysis; accessed 2026-06-02.)
      
      ## When to bring in a tree-sitter dependency graph
      
      Grep + churn answers most first-pass questions. Reach for a structural graph only when the codebase is large enough that "who depends on this" is no longer obvious by eye.
      
      - **codegraph** (tree-sitter + PageRank, no embeddings/GPU): one-shot ranking of files by structural importance, biased toward query keywords. Prefer this one-shot rank for onboarding — it is fast and stateless. (github.com/tarunms7/codegraph)
      - **FileScopeMCP**: scores each file 0–10 by how many things depend on it — a quick "what is load-bearing" lens. (github.com/admica/FileScopeMCP)
      - **Heavier MCP-native graphs** (codegraph-ai/CodeGraph: ~42 MCP tools across 38 languages; tree-sitter-analyzer): full function/class/import + call/inheritance graphs. Powerful, but standing up a persistent MCP graph server during *first-pass* onboarding is usually overkill — defer until the map exists and you have a concrete deep-dive question. (accessed 2026-06-02.)
      
      **Rule of thumb**: if the repo is under a few hundred files, a tree-sitter graph costs more setup than it saves. Use it for monorepos and for finding the load-bearing packages, not for a service you can grep through in ten minutes.
      
  • scripts
    • verify.sh 1.8 KB
      #!/usr/bin/env bash
      # verify.sh — structural check for the codebase-onboarding artifact.
      #
      # Validates that a CODEBASE-MAP.md contains the required sections. This is a
      # STRUCTURE check, not a content-correctness check (this is a process skill;
      # its real rigor is the capability eval). Read-only: it never writes or edits.
      #
      # Usage:   verify.sh [path-to-map]      (default: ./CODEBASE-MAP.md)
      # Exit 0:  map present with all required sections, OR no map yet (nothing to
      #          check — an empty/clean target must not produce a false failure).
      # Exit 1:  map present but missing one or more required sections.
      
      set -euo pipefail
      
      MAP="${1:-CODEBASE-MAP.md}"
      
      # No artifact yet => nothing to validate. Clean target, exit 0 (no false fail).
      if [[ ! -f "$MAP" ]]; then
        echo "OK: no $MAP yet — nothing to verify (run the recon pass to produce one)."
        exit 0
      fi
      
      # Required section headers. Each entry is an extended-regex alternation so that
      # reasonable wording variants (e.g. "Request flow" / "Data flow") still pass.
      declare -a SECTIONS=(
        "Stack"
        "Entry points?"
        "Request flow|Data flow"
        "Module ownership"
        "Hidden behavior|Hidden behaviour|Side effects?"
        "Hotspots?"
        "How to run|Running locally"
      )
      
      missing=()
      for pat in "${SECTIONS[@]}"; do
        # Match a markdown heading line (## ...) containing the pattern, case-insensitive.
        if ! grep -Eiq "^#{1,6}[[:space:]].*(${pat})" "$MAP"; then
          # Use the first alternative as the human-readable label.
          missing+=("${pat%%|*}")
        fi
      done
      
      if (( ${#missing[@]} > 0 )); then
        echo "FAIL: $MAP is missing required section(s):" >&2
        for m in "${missing[@]}"; do
          echo "  - $m" >&2
        done
        echo "Required: Stack, Entry points, Request flow, Module ownership, Hidden behavior, Hotspots, How to run." >&2
        exit 1
      fi
      
      echo "OK: $MAP has all required sections."
      exit 0
      
  • SKILL.md 9.3 KB
    ---
    name: codebase-onboarding
    description: "Use when you land in an unfamiliar or inherited codebase and must get productive fast: a breadth-first map of entry points, request flow, module ownership, hidden side effects (cron, webhooks, workers) and churn hotspots, committed as CODEBASE-MAP.md. NOT a deep audit of one module (that is `analyze`) or chasing one failure (that is `debug`)."
    tags: [onboarding, codebase, code-mapping, legacy-code, architecture, reverse-engineering, hotspots]
    recommends: [analyze, debug, decision-records, harness, init, knowledge-ops]
    origin: risco
    ---
    
    # Codebase onboarding — get oriented fast, leave a map
    
    You have just landed in a codebase you did not write: a fresh clone, an inherited project, an acquired repo, an abandoned side project someone handed you. The instinct is to start reading files top-to-bottom. Resist it. That is how a week disappears and you still cannot answer "where does X happen". This skill runs a disciplined **breadth-first reconnaissance pass** and produces one durable artifact: a map a teammate can trust and you can re-read tomorrow.
    
    The payoff is measured. Engineers using AI to onboard reach the same milestones roughly **2x faster** — productive in 1–2 weeks instead of 4–6 — and the biggest gains are exactly in *searching for code, decoding undocumented patterns, and tracing data flows* (super-productivity.com, accessed 2026-06-02). That is what the recon pass below targets, in order.
    
    ## Lead with the deliverable
    
    Before you grep a single line, know the target: a single living file, `CODEBASE-MAP.md`, committed at the repo root. You work *backward* from its sections — every recon step fills one. Minimal schema:
    
    ```markdown
    # CODEBASE-MAP.md — <repo name>
    
    ## Stack          # languages, framework + versions, package manager, run scripts
    ## Entry points   # main / server bootstrap / route registration / CLI commands
    ## Request flow   # one real path traced transport -> business logic -> persistence
    ## Module ownership   # who owns transport / business logic / persistence / UI
    ## Hidden behavior    # cron, webhooks, queue workers, event listeners, env branches
    ## Hotspots       # most-churned + most-complex files = highest risk
    ## How to run     # the exact commands to boot it and hit one path locally
    ```
    
    Why a file and not a chat answer: a map that lives only in the conversation dies when the session ends, and the next agent re-does the work. The artifact is the point. `verify.sh` checks these sections exist (structure, not content).
    
    ## Two operating rules
    
    1. **Breadth before depth.** First pass maps *where things are*, not *how they work*. You are drawing the subway map, not reading every passenger's diary. Depth is `analyze`/`debug` work, on demand, later. — Reading everything is the failure mode onboarding exists to replace.
    2. **Hypothesis before answer.** Spend ~5 minutes forming your own guess ("auth probably lives in `src/middleware`"), then grep to confirm or kill it. — Verifying a hypothesis builds the mental model that makes you fast; a handed-to-you answer does not stick (martinfowler.com, Böckeler, accessed 2026-06-02).
    
    ## The recon pass — ordered
    
    Run these in order. Each step writes one map section. Stop escalating the moment the section is answerable.
    
    **a. Orient — read the manifest, size the repo.**
    Read the manifest(s) and lockfile, not the README first: `package.json` / `pyproject.toml` / `go.mod` / `Gemfile` / `pom.xml` tell you the real stack, framework version and run scripts; the lockfile tells you what is actually installed. Then size it:
    
    ```bash
    scc --by-file --sort lines .   # LOC, complexity, COCOMO estimate per file
    ```
    
    `scc` (Sloc Cloc and Code, pure-Go, v3.7.0 Apr 2026) is the fast structural counter of record — materially faster than cloc/tokei and it reports per-file complexity, which you reuse for hotspots. Why manifest-first: the README describes intent (often stale); the manifest describes reality.
    
    **b. Find the entry points.**
    Where does execution start? Look for `main`, the server bootstrap, route registration, CLI command definitions. Let the framework's convention guide you (Next.js `app/`/`pages/`, Express `app.use`/router mounts, Django `urls.py`, FastAPI `@app`/`APIRouter`, Rails `routes.rb`, Spring `@RestController`). See `references/recon-playbook.md` for per-ecosystem patterns.
    
    **c. Trace one real request end-to-end.**
    Pick a single meaningful path (a login, a checkout, the main CLI command) and follow it: transport (route/handler) → business logic → persistence → response. One path traced beats ten skimmed. This is the spine of the map.
    
    **d. Map module ownership.**
    For the directories `scc` flagged as large, label each: transport, business logic, persistence, UI, shared/util. You are answering "if I need to change pricing, which folder do I open" — the question teammates actually ask.
    
    **e. Hunt hidden behavior.**
    The bugs live in what runs *without* a request. Grep for cron schedules, webhook receivers, queue/background workers, event listeners and env-driven branches (see the appendix). Why this step is non-negotiable: side effects are invisible in a top-down read and they are where inherited codebases bite.
    
    **f. Rank hotspots.**
    Git churn is the cheapest risk signal — no extra tooling, and high-churn files are a proxy for "lacks tests/abstraction":
    
    ```bash
    git log --format=format: --name-only --since=12.month \
      | grep -v '^$' | sort | uniq -c | sort -nr | head -50
    ```
    
    The richer move is **churn × complexity**: the top-right quadrant (changes constantly *and* is hard to read) is your real danger zone (understandlegacycode.com Hotspots, accessed 2026-06-02). Escalate to that — or to a tree-sitter dependency graph — only when grep + churn is not enough: `references/recon-playbook.md` has the churn×complexity recipe (code-maat) and when codegraph PageRank / FileScopeMCP earn their setup cost.
    
    **g. Confirm hands-on.**
    Run the app, walk one end-user journey, send a real request, watch the logs. Reading alone leaves the map unverified; a single real request validates the whole trace in step c.
    
    ## Scope the effort
    
    Match the pass to the repo. Do not stand up heavy tooling on a small project.
    
    | Repo shape | Map fully | Skip / defer | Escalate to a graph tool? |
    | --- | --- | --- | --- |
    | Tiny (<20 files) | a, b, c, g | churn, ownership table | No — grep is faster than setup |
    | Single-service app | all a–g | — | Only if ownership is unclear after grep |
    | Large monorepo | a, b, then per-package c–f | mapping every package at once | Yes — codegraph PageRank to find the load-bearing packages |
    | Polyglot | a, b, c per language boundary | one unified flow diagram | Yes, if cross-language calls obscure the flow |
    
    ## Command appendix
    
    Language-tagged, copy-ready. Per-ecosystem depth lives in `references/recon-playbook.md`.
    
    ```bash
    # Size + complexity (reuse the complexity column for hotspots)
    scc --by-file --sort complexity .
    
    # Route registration (adjust per framework)
    rg -n "app\.(get|post|put|delete|use)\(|@app\.(get|post)|APIRouter|router\.(get|post)" --type-add 'web:*.{js,ts,py}' -tweb
    
    # Cron / scheduled jobs
    rg -n "cron|schedule|@scheduled|setInterval|celery\.beat|node-cron" -i
    
    # Webhook receivers
    rg -n "webhook|/hooks/|stripe.*signature|x-hub-signature" -i
    
    # Queue / background workers
    rg -n "queue|worker|bull|sidekiq|celery|sqs|rabbitmq|kafka|@task" -i
    
    # Env-driven branches (hidden config-conditional behavior)
    rg -n "process\.env\.|os\.environ|ENV\[|getenv" 
    ```
    
    ## Writing & maintaining the map
    
    - **Commit it.** `CODEBASE-MAP.md` at the repo root, in version control. A map outside the repo rots silently.
    - **Keep it living.** When the recon reveals you guessed wrong, fix the line — the map is the record of what is *true now*, not your first impression.
    - **Link out, do not duplicate.** The map says *what is*. For *why a choice was made*, write an ADR (`../decision-records/SKILL.md`). To scaffold project tooling and a wiki, that is `../harness/SKILL.md`. To write the agent-memory `CLAUDE.md`, that is `../init/SKILL.md` — onboarding is the broader recon that feeds it. For generic note/wiki capture, `../knowledge-ops/SKILL.md`.
    - **Hand off to depth tools.** Once the map exists, a deep correctness/security read of one module is `../analyze/SKILL.md`; chasing a specific failure through the system is `../debug/SKILL.md`. Onboarding builds the map you debug *with*.
    
    ## Anti-patterns
    
    | Bad | Why it bites | Good |
    | --- | --- | --- |
    | Read every file top-to-bottom | Burns the week; you finish exhausted and still can't trace one request | Breadth-first: map locations first, depth on demand |
    | Trust the README over the code | READMEs drift; the manifest and the routes are the truth | Read manifest + lockfile first, confirm by grep |
    | Map everything at full depth | Analysis paralysis on a monorepo; you map dead modules | Trace one real flow end-to-end; expand only where needed |
    | Skip the run step | An unverified map is a hypothesis, not a map | Boot it, hit one path, watch logs before you trust the trace |
    | Map lives only in chat | Dies with the session; next agent redoes it | Write & commit `CODEBASE-MAP.md` |
    | Guess instead of grep | Confident-wrong is worse than slow-right | Form the hypothesis, then `rg` to confirm or kill it |
    | Stand up a tree-sitter MCP graph on a 5-file repo | Setup costs more than the whole recon | Reserve graph tools for large monorepos; grep + churn first |
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related