Claude Skill

searching-codebases

Binding-resolved Python symbol queries via pyright — every true caller (--refs), the real definition (--def), or an inferred signature (--hover) of a .py symbol, excluding the same-named false positives text search cannot tell apart. Use when a task needs ALL callers or users of

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download oaustegard-claude-skills-plugins_code-intelligence_skills_searching-codebases-e39c726.zip · 36 KB
Part of oaustegard/claude-skills — 39 skills

Install

skills CLI npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/code-intelligence/skills/searching-codebases
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
Git git clone https://github.com/oaustegard/claude-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.

README

searching-codebases

Find code in any codebase by regex pattern or natural language concept. Auto-routes between n-gram indexed regex search (2-20x faster than ripgrep) and TF-IDF semantic search. Expands results to full functions via tree-sitting AST data. For Python sources, a binding-resolved reference/definition tier (pyright, via python-lsp) finds real callers and definitions without same-name false positives.

Features

  • Dual search modes — regex (pattern/identifier) and semantic (natural language concepts), auto-routed per query
  • N-gram indexed regex — sparse inverted index narrows candidate files by 90-99% before ripgrep verification
  • TF-IDF semantic search — cosine similarity ranking over code chunks (functions, classes)
  • Binding-resolved Python tier — --refs/--def/--hover via pyright: true find-all-callers and go-to-definition that exclude same-named, unrelated symbols and follow imports. Engaged lazily; degrades to the regex text path when pyright/node is unavailable
  • AST context expansion — --expand returns complete function/class bodies instead of line fragments
  • Flexible sources — accepts GitHub URLs, local directories, uploaded files/archives, or project knowledge
  • Mixed queries — multiple queries with different modes in a single invocation; indexes built once per mode

Dependencies

  • ripgrep — required for regex verification
  • tree-sitting — auto-installs the bare tree-sitter package when needed: for --expand context and for the symbol→position resolution that seeds the binding-resolved tier (grammars ship bundled). Regex and semantic search work without it
  • scikit-learn — required for semantic mode (auto-installs)
  • python-lsp — provides the binding-resolved tier (--refs/--def/--hover); self-bootstraps pyright on first use and needs system node (v18+). Without it those flags degrade to the regex text path

Skill manifest

Searching Codebases

Find code in any codebase by pattern or concept. One entry point, two search strategies, automatic routing.

When NOT to use this skill

Python callers and definitions only. Every other code question is cheaper elsewhere, and the measurement below is why this skill is edge-case-only.

Situation Use
What symbols does this file contain? tree-sitting
Where is X defined, so I can read it? tree-sitting
Any structure question, any language tree-sitting
A literal string or regex plain ripgrep
Which files are most about a concept? bm25
First look at an unfamiliar repo exploring-codebases

Non-Python code has no pyright binding resolution to offer, so there is no version of this skill that applies to it.

Prerequisites

uv tool install ripgrep

tree-sitting installs automatically when needed — for --expand context expansion and for the binding-resolved --refs/--def/--hover tier, which uses it to resolve symbol positions. Only the bare tree-sitter package is fetched; the language grammars ship bundled.

Primary Command

SKILL_DIR=/mnt/skills/user/searching-codebases

python3 $SKILL_DIR/scripts/search.py SOURCE "query1" ["query2" ...] [OPTIONS]

SOURCE is any of:

  • Local directory path
  • GitHub URL (downloads tarball automatically)
  • uploads (uses /mnt/user-data/uploads/)
  • project (uses /mnt/project/)
  • Path to a .zip or .tar.gz archive

Search Modes

Regex mode (patterns, identifiers, literal text):

python3 $SKILL_DIR/scripts/search.py ./repo "def handle_error"
python3 $SKILL_DIR/scripts/search.py ./repo "class.*Exception" --regex
python3 $SKILL_DIR/scripts/search.py ./repo "TODO|FIXME|HACK"

Semantic mode (concepts, natural language):

python3 $SKILL_DIR/scripts/search.py ./repo "retry logic with backoff" --semantic
python3 $SKILL_DIR/scripts/search.py ./repo "authentication flow"
python3 $SKILL_DIR/scripts/search.py ./repo "error handling strategy"

Auto-detection: short queries and code-like tokens → regex. Multi-word natural language → semantic. Override with --regex or --semantic.

Binding-resolved mode (Python only — pyright via the python-lsp skill):

python3 $SKILL_DIR/scripts/search.py ./repo --refs SYMBOL    # find all real uses
python3 $SKILL_DIR/scripts/search.py ./repo --def SYMBOL     # go-to-definition
python3 $SKILL_DIR/scripts/search.py ./repo --hover SYMBOL   # inferred type/signature

Regex mode matches text, so a cross-reference for a function false-positives on shadowed and same-named-but-unrelated symbols. --refs is binding-resolved: pyright excludes the unrelated same-named symbol and follows imports. Use it when you need a true "find all callers/users" for a .py symbol, not a text grep.

The tier is engaged lazily — pyright's index cost is paid only when you ask for --refs/--def/--hover, never on ordinary searches. It is Python-only; for non-.py sources, or when pyright/node is unavailable, it prints a one-line degradation note and falls back to the regex text path. Each takes a single bare symbol name and is mutually exclusive with the other two and with text queries.

Options

  • --regex / --semantic: Force search mode
  • --refs SYMBOL / --def SYMBOL / --hover SYMBOL: Binding-resolved Python queries via pyright (see Binding-resolved mode above)
  • --expand: Return full function bodies via tree-sitting AST context
  • --benchmark: Compare indexed regex vs brute-force ripgrep
  • --branch NAME: Git branch for GitHub URLs (default: main)
  • --skip DIRS: Comma-separated directories to skip
  • --json: Machine-readable output
  • -v: Show index stats and query routing decisions

How It Works

Regex search builds a sparse n-gram inverted index over all files. Queries are decomposed into literal fragments, looked up in the index to identify candidate files (typically 90-99% reduction), then verified with ripgrep. Frequency-weighted n-grams make rare character sequences more selective.

Semantic search builds a TF-IDF index over code chunks (functions, classes, structural entries). Queries are ranked by cosine similarity.

Context expansion (--expand) uses tree-sitting's AST cache to identify function/class boundaries, returning complete structural units rather than line fragments. On first use, tree-sitting scans the repo (~700ms for 250 files); subsequent expansions are sub-millisecond.

Small codebases (< 20 files) skip indexing entirely — direct ripgrep is faster when there's nothing to narrow.

Mixed Queries

Multiple queries can use different modes in a single invocation. Each query is auto-routed independently, and indexes are built once per mode:

python3 $SKILL_DIR/scripts/search.py ./repo \
  "class.*Error" \
  "error recovery strategy" \
  "def retry"

Dependencies

  • tree-sitting: Provides AST context expansion for --expand and the symbol→position resolution that seeds the binding-resolved tier (--refs/--def/--hover). Auto-installs the bare tree-sitter package when either is used (grammars are bundled). Regex and semantic search work without it.
  • ripgrep: Required for regex verification. Install via uv tool install ripgrep.
  • scikit-learn: Required for semantic mode. Installs automatically.
  • python-lsp: Provides the binding-resolved tier (--refs/--def/--hover). Self-bootstraps pyright on first use and requires system node (v18+). Not required — without it those flags degrade to the regex text path.

When to Use — narrow, by design

The ONE recommended use: binding-resolved Python symbol queries.

  • "find all callers of X" / "where is X really defined" for a .py symbol, when same-named-but-unrelated symbols would pollute a text grep. Empirical basis: rg get on psf/requests returned 232 hits, 224 of them false; --refs get excluded all 224 (2026-06-15).

When NOT to Use — which is most of the time

Everything else. Measured head-to-head on real issue-localization tasks (7 scikit-learn issues with merged fix-PRs, gold = PR diff files, 2026-07-04, replicating the file-discovery metric of arXiv:2602.11988):

  • Literal tokens / identifiers: naive rg -l tied or beat the indexed tier on recall@10 in every instance, at 0.4s vs 25s.
  • Concept / natural-language search: the TF-IDF semantic tier never beat identifier grep — not even on identifier-poor issues, which are themselves rare (~0.3% of merged-PR traffic in the sample).
  • First encounter / "what is this repo": use exploring-codebases.
  • Repos under ~20 files: read them.

The self-test before invoking: would plain rg return the same answer? If yes, use rg. The indexed-regex and semantic tiers are retained for completeness and for corpora where they may yet earn their cost (very large repos, non-code document collections), but they carry the burden of proof.

Files

  • scripts/search.py — Entry point, query routing, output formatting
  • scripts/resolve.py — Input source resolution (GitHub, uploads, archives)
  • scripts/context.py — tree-sitting-based AST context expansion
  • scripts/ngram_index.py — Sparse n-gram inverted index, regex decomposition
  • scripts/sparse_ngrams.py — Core n-gram algorithms, frequency weights
  • scripts/code_rag.py — TF-IDF semantic search over code chunks
  • scripts/lsp_refs.py — Binding-resolved Python tier: symbol→position resolution (tree-sitting), pyright queries (python-lsp), soft fallback
Files (claude-skills)
  • scripts
    • code_rag.py 20.4 KB
      """
      code_rag.py — TF-IDF semantic search over codebases
      ====================================================
      
      Semantic search layer that bridges natural language intent to actual
      codebase identifiers. Indexes docstrings, comments, function signatures,
      and markdown sections.
      
      Zero dependencies beyond scikit-learn + numpy (pre-installed).
      
      Usage:
          python3 code_rag.py index /path/to/repo
          python3 code_rag.py search /path/to/repo "retry logic with backoff"
          python3 code_rag.py search /path/to/repo "authentication flow" --top 10
          python3 code_rag.py search /path/to/repo "error handling" --grouped
          python3 code_rag.py search /path/to/repo "middleware" --rg
      """
      
      import json
      import os
      import re
      import sys
      import time
      from dataclasses import dataclass, field
      from pathlib import Path
      
      from sklearn.feature_extraction.text import TfidfVectorizer
      from sklearn.metrics.pairwise import cosine_similarity
      
      # ── Chunk extraction ─────────────────────────────────────────────
      
      @dataclass
      class Chunk:
          """A searchable unit extracted from a source file."""
          file: str          # relative path from repo root
          line: int          # starting line number (1-indexed)
          kind: str          # function, class, module_doc, section, map_entry
          name: str          # identifier or heading
          text: str          # searchable content (docstring + signature + comments)
          end_line: int = 0
      
          @property
          def loc(self) -> str:
              return f"{self.file}:{self.line}"
      
      
      def _extract_python(filepath: str, rel_path: str) -> list[Chunk]:
          """Extract functions, classes, and module docstrings from Python files.
          
          Uses regex rather than AST intentionally — we want docstrings, comments,
          and identifiers for TF-IDF vocabulary, not syntactic correctness.
          """
          try:
              with open(filepath, "r", errors="replace") as f:
                  content = f.read()
                  lines = content.split('\n')
          except OSError:
              return []
      
          chunks = []
      
          # Module-level docstring
          mod_match = re.match(
              r'\s*(?:#[^\n]*\n)*\s*("""[\s\S]*?"""|\'\'\'[\s\S]*?\'\'\')', content
          )
          if mod_match:
              doc = mod_match.group(1).strip('"\' \n')
              if len(doc) > 20:
                  chunks.append(Chunk(
                      file=rel_path, line=1, kind="module_doc",
                      name=Path(rel_path).stem, text=doc,
                  ))
      
          # Functions and classes with decorators, signatures, docstrings, nearby comments
          pattern = re.compile(
              r'^((?:[ \t]*@\w+[^\n]*\n)*)'   # decorators
              r'^([ \t]*)(def|class)\s+'        # indent + keyword
              r'(\w+)'                          # name
              r'([^\n]*)\n'                     # rest of signature line
              r'((?:[ \t]*"""[\s\S]*?"""'       # optional docstring
              r"|[ \t]*'''[\s\S]*?''')?)",
              re.MULTILINE
          )
      
          for match in pattern.finditer(content):
              decorators = match.group(1).strip()
              kind = "class" if match.group(3) == "class" else "function"
              name = match.group(4)
              signature = f"{match.group(3)} {name}{match.group(5).strip()}"
              docstring = match.group(6).strip('"\' \n\t') if match.group(6) else ""
              line_num = content[:match.start()].count('\n') + 1
      
              text_parts = [name, signature]
              if decorators:
                  text_parts.append(decorators)
              if docstring:
                  text_parts.append(docstring)
      
              # Grab inline comments in the next ~20 lines
              body_start = match.end()
              body_lines = content[body_start:body_start + 2000].split('\n')[:20]
              comments = [
                  l.strip().lstrip('#').strip()
                  for l in body_lines if l.strip().startswith('#')
              ]
              if comments:
                  text_parts.extend(comments)
      
              chunks.append(Chunk(
                  file=rel_path, line=line_num, kind=kind,
                  name=name, text=" ".join(text_parts),
                  end_line=line_num + len((match.group(6) or "").split('\n')),
              ))
      
          return chunks
      
      
      def _extract_markdown(filepath: str, rel_path: str) -> list[Chunk]:
          """Extract heading-based sections from Markdown files."""
          try:
              with open(filepath, "r", errors="replace") as f:
                  lines = f.readlines()
          except OSError:
              return []
      
          chunks = []
          heading = None
          section_lines = []
          start = 1
      
          for i, line in enumerate(lines, 1):
              m = re.match(r'^(#{1,3})\s+(.+)', line)
              if m:
                  if heading and section_lines:
                      text = " ".join(section_lines)
                      if len(text) > 30:
                          chunks.append(Chunk(
                              file=rel_path, line=start, kind="section",
                              name=heading, text=text, end_line=i - 1,
                          ))
                  heading = m.group(2).strip()
                  section_lines = [heading]
                  start = i
              elif heading:
                  stripped = line.strip()
                  if stripped and not stripped.startswith('```'):
                      section_lines.append(stripped)
      
          if heading and section_lines:
              text = " ".join(section_lines)
              if len(text) > 30:
                  chunks.append(Chunk(
                      file=rel_path, line=start, kind="section",
                      name=heading, text=text, end_line=len(lines),
                  ))
      
          return chunks
      
      
      def _extract_js_ts(filepath: str, rel_path: str) -> list[Chunk]:
          """Extract functions, classes, and interfaces from JS/TS/TSX/JSX files.
      
          Covers: function declarations, arrow functions assigned to const/let/var,
          class declarations, interface/type declarations, exported members,
          and JSDoc comments preceding any of the above.
          """
          try:
              with open(filepath, "r", errors="replace") as f:
                  content = f.read()
          except OSError:
              return []
      
          chunks = []
      
          # Module-level JSDoc or leading block comment
          mod_match = re.match(r'\s*(/\*\*[\s\S]*?\*/)', content)
          if mod_match:
              doc = mod_match.group(1).strip('/* \n')
              if len(doc) > 20:
                  chunks.append(Chunk(
                      file=rel_path, line=1, kind="module_doc",
                      name=Path(rel_path).stem, text=doc,
                  ))
      
          # --- Pattern 1: function/class/interface declarations ---
          # Captures: export? async? function name(...), class Name, interface Name, type Name
          decl_pattern = re.compile(
              r'(/\*\*[\s\S]*?\*/\s*)?'                        # optional JSDoc (group 1)
              r'^[ \t]*(export\s+(?:default\s+)?)?'             # optional export (group 2)
              r'(async\s+)?'                                     # optional async (group 3)
              r'(function\*?|class|interface|type|enum)\s+'      # keyword (group 4)
              r'(\w+)'                                           # name (group 5)
              r'([^\n]*)',                                        # rest of line (group 6)
              re.MULTILINE
          )
      
          for match in decl_pattern.finditer(content):
              jsdoc = match.group(1) or ""
              jsdoc_clean = re.sub(r'[/*]', ' ', jsdoc).strip()
              keyword = match.group(4)
              name = match.group(5)
              rest = match.group(6).strip()
              line_num = content[:match.start()].count('\n') + 1
      
              kind_map = {
                  'function': 'function', 'class': 'class',
                  'interface': 'class', 'type': 'class', 'enum': 'class',
              }
              # Strip trailing * from function*
              kind = kind_map.get(keyword.rstrip('*'), 'function')
      
              text_parts = [name, f"{keyword} {name}{rest}"]
              if jsdoc_clean:
                  text_parts.append(jsdoc_clean)
      
              chunks.append(Chunk(
                  file=rel_path, line=line_num, kind=kind,
                  name=name, text=" ".join(text_parts),
              ))
      
          # --- Pattern 2: arrow functions assigned to variables ---
          # const myFunc = (...) => { ... }
          # const myFunc: Type = (...) => ...
          arrow_pattern = re.compile(
              r'(/\*\*[\s\S]*?\*/\s*)?'                         # optional JSDoc
              r'^[ \t]*(export\s+(?:default\s+)?)?'              # optional export
              r'(const|let|var)\s+'                              # binding keyword
              r'(\w+)'                                           # name (group 4)
              r'([^=]*?)\s*='                                    # type annotation etc
              r'\s*(?:async\s+)?'                                # optional async
              r'\([^)]*\)\s*(?::\s*[^=]+?)?\s*=>',              # arrow function signature
              re.MULTILINE
          )
      
          for match in arrow_pattern.finditer(content):
              jsdoc = match.group(1) or ""
              jsdoc_clean = re.sub(r'[/*]', ' ', jsdoc).strip()
              name = match.group(4)
              line_num = content[:match.start()].count('\n') + 1
      
              # Skip if already captured by declaration pattern (unlikely but safe)
              text_parts = [name, match.group(0).split('=>')[0].strip()]
              if jsdoc_clean:
                  text_parts.append(jsdoc_clean)
      
              chunks.append(Chunk(
                  file=rel_path, line=line_num, kind="function",
                  name=name, text=" ".join(text_parts),
              ))
      
          # --- Pattern 3: class methods (inside class bodies) ---
          method_pattern = re.compile(
              r'(/\*\*[\s\S]*?\*/\s*)?'                         # optional JSDoc
              r'^[ \t]+((?:static|async|get|set|private|protected|public|readonly)\s+)*'
              r'(\w+)\s*\(',                                     # method name + open paren
              re.MULTILINE
          )
      
          for match in method_pattern.finditer(content):
              jsdoc = match.group(1) or ""
              jsdoc_clean = re.sub(r'[/*]', ' ', jsdoc).strip()
              name = match.group(3)
              line_num = content[:match.start()].count('\n') + 1
      
              # Skip common false positives
              if name in ('if', 'for', 'while', 'switch', 'catch', 'return',
                           'require', 'import', 'console', 'throw', 'new',
                           'typeof', 'instanceof', 'delete', 'void', 'yield',
                           'await', 'super', 'this'):
                  continue
      
              text_parts = [name]
              if jsdoc_clean:
                  text_parts.append(jsdoc_clean)
      
              chunks.append(Chunk(
                  file=rel_path, line=line_num, kind="function",
                  name=name, text=" ".join(text_parts),
              ))
      
          return chunks
      
      
      def _extract_yaml(filepath: str, rel_path: str) -> list[Chunk]:
          """Extract top-level keys and commented sections from YAML files.
      
          YAML files often contain infrastructure config, CI/CD pipelines,
          Kubernetes manifests, docker-compose services, etc. The top-level
          keys and their comments are the most searchable units.
          """
          try:
              with open(filepath, "r", errors="replace") as f:
                  lines = f.readlines()
          except OSError:
              return []
      
          # Skip very large YAML files (likely generated, e.g. lockfiles)
          if len(lines) > 2000:
              return []
      
          chunks = []
          current_key = None
          current_lines = []
          current_start = 1
          comment_buffer = []
      
          for i, line in enumerate(lines, 1):
              stripped = line.strip()
      
              # Accumulate comment lines preceding a key
              if stripped.startswith('#'):
                  comment_buffer.append(stripped.lstrip('#').strip())
                  continue
      
              # Top-level key: no leading whitespace, ends with ':'
              top_key_match = re.match(r'^([a-zA-Z_][\w.-]*)\s*:', line)
              if top_key_match:
                  # Flush previous section
                  if current_key and current_lines:
                      text = " ".join(current_lines)
                      if len(text) > 15:
                          chunks.append(Chunk(
                              file=rel_path, line=current_start, kind="section",
                              name=current_key, text=text,
                          ))
      
                  current_key = top_key_match.group(1)
                  current_lines = [current_key]
                  current_start = i
      
                  # Include preceding comments as searchable text
                  if comment_buffer:
                      current_lines.extend(comment_buffer)
                  comment_buffer = []
      
              elif current_key and stripped:
                  # Nested content: include values and inline comments
                  # Strip YAML syntax noise but keep identifiers and values
                  clean = re.sub(r'^\s*-\s*', '', stripped)
                  clean = clean.split('#')[0].strip()  # strip inline comments into separate add
                  inline_comment = stripped.split('#')[1].strip() if '#' in stripped else ""
                  if clean:
                      current_lines.append(clean)
                  if inline_comment:
                      current_lines.append(inline_comment)
              else:
                  comment_buffer = []
      
          # Flush last section
          if current_key and current_lines:
              text = " ".join(current_lines)
              if len(text) > 15:
                  chunks.append(Chunk(
                      file=rel_path, line=current_start, kind="section",
                      name=current_key, text=text,
                  ))
      
          # Also create a file-level chunk with the whole file as context
          # (useful for small config files like docker-compose.yml)
          if len(lines) < 100:
              all_text = " ".join(
                  l.strip() for l in lines
                  if l.strip() and not l.strip().startswith('---')
              )
              if len(all_text) > 30:
                  chunks.append(Chunk(
                      file=rel_path, line=1, kind="module_doc",
                      name=Path(rel_path).stem, text=all_text,
                  ))
      
          return chunks
      
      
      # ── Index ────────────────────────────────────────────────────────
      
      SKIP_DIRS = {
          '.git', 'node_modules', '__pycache__', '.venv', 'venv', 'dist',
          'build', '.next', '.mypy_cache', '.pytest_cache', '.tox', '.eggs',
          '.ruff_cache', 'target', 'coverage', '.coverage',
      }
      
      EXTRACTORS = {
          '.py': _extract_python,
          '.md': _extract_markdown,
          '.js': _extract_js_ts,
          '.jsx': _extract_js_ts,
          '.ts': _extract_js_ts,
          '.tsx': _extract_js_ts,
          '.mjs': _extract_js_ts,
          '.mts': _extract_js_ts,
          '.yaml': _extract_yaml,
          '.yml': _extract_yaml,
      }
      
      
      @dataclass
      class Index:
          """TF-IDF index over code chunks."""
          chunks: list[Chunk] = field(default_factory=list)
          vectorizer: TfidfVectorizer | None = None
          matrix: object | None = None
          build_time_ms: float = 0
          repo_path: str = ""
      
          def build(self, repo_path: str, skip_dirs: set[str] = None):
              """Walk repo, extract chunks, build TF-IDF matrix."""
              t0 = time.monotonic()
              self.repo_path = str(repo_path)
              skip = skip_dirs or SKIP_DIRS
              self.chunks = []
      
              for root, dirs, files in os.walk(repo_path):
                  dirs[:] = [d for d in dirs if d not in skip and not d.startswith('.')]
      
                  for fname in files:
                      fpath = os.path.join(root, fname)
                      rel = os.path.relpath(fpath, repo_path)
                      ext = Path(fname).suffix.lower()
      
                      if ext in EXTRACTORS:
                          self.chunks.extend(EXTRACTORS[ext](fpath, rel))
      
              if not self.chunks:
                  print("WARNING: No chunks extracted", file=sys.stderr)
                  return self
      
              self.vectorizer = TfidfVectorizer(
                  ngram_range=(1, 2),
                  sublinear_tf=True,
                  max_df=0.80,
                  min_df=2,
                  stop_words="english",
                  max_features=50000,
                  token_pattern=r'(?u)\b[a-zA-Z_]\w{1,}\b',
              )
              self.matrix = self.vectorizer.fit_transform([c.text for c in self.chunks])
              self.build_time_ms = (time.monotonic() - t0) * 1000
              return self
      
          def search(self, query: str, top_k: int = 5, min_score: float = 0.01
                     ) -> list[tuple[Chunk, float]]:
              """Semantic search: rank chunks by cosine similarity to query."""
              if self.vectorizer is None or self.matrix is None:
                  return []
      
              q_vec = self.vectorizer.transform([query])
              scores = cosine_similarity(q_vec, self.matrix).flatten()
              indices = scores.argsort()[::-1]
      
              results = []
              for idx in indices:
                  if scores[idx] < min_score:
                      break
                  results.append((self.chunks[idx], float(scores[idx])))
                  if len(results) >= top_k:
                      break
              return results
      
          def search_grouped(self, query: str, top_k: int = 10, min_score: float = 0.01
                             ) -> dict[str, list[tuple[Chunk, float]]]:
              """Search and group by file — useful for feeding targets to grep."""
              results = self.search(query, top_k=top_k * 3, min_score=min_score)
              grouped = {}
              for chunk, score in results:
                  grouped.setdefault(chunk.file, []).append((chunk, score))
              sorted_files = sorted(
                  grouped.items(), key=lambda x: x[1][0][1], reverse=True
              )
              return dict(sorted_files[:top_k])
      
          def stats(self) -> dict:
              if not self.chunks:
                  return {"chunks": 0}
              kinds = {}
              for c in self.chunks:
                  kinds[c.kind] = kinds.get(c.kind, 0) + 1
              return {
                  "chunks": len(self.chunks),
                  "files": len(set(c.file for c in self.chunks)),
                  "vocabulary": len(self.vectorizer.get_feature_names_out()) if self.vectorizer else 0,
                  "build_ms": round(self.build_time_ms),
                  "kinds": kinds,
              }
      
      
      # ── Output formatting ────────────────────────────────────────────
      
      def _sanitize(text: str) -> str:
          """Strip HTML entities and normalize whitespace."""
          import html
          return html.unescape(text).replace('\n', ' ').strip()
      
      
      def _format_results(results: list[tuple[Chunk, float]]) -> str:
          lines = []
          for chunk, score in results:
              preview = _sanitize(chunk.text[:140])
              if len(chunk.text) > 140:
                  preview += "..."
              lines.append(f"  {score:.3f}  [{chunk.kind}] {chunk.name}  {chunk.loc}")
              lines.append(f"         {preview}")
              lines.append("")
          return "\n".join(lines)
      
      
      def _format_grouped(grouped: dict[str, list[tuple[Chunk, float]]]) -> str:
          lines = []
          for filepath, hits in grouped.items():
              lines.append(f"\n  {filepath}")
              for chunk, score in hits:
                  name = _sanitize(chunk.name)
                  lines.append(
                      f"    {score:.3f}  {chunk.kind:12s}  {name} :{chunk.line}"
                  )
          return "\n".join(lines)
      
      
      def _format_for_grep(grouped: dict[str, list[tuple[Chunk, float]]],
                           repo_path: str) -> str:
          lines = ["# Files ranked by relevance (use with grep -n -A25):"]
          for filepath, hits in grouped.items():
              best_score = hits[0][1]
              names = [_sanitize(h[0].name) for h in hits[:3]]
              lines.append(f"# score={best_score:.3f} matches: {', '.join(names)}")
              lines.append(os.path.join(repo_path, filepath))
          return "\n".join(lines)
      
      
      # ── CLI ──────────────────────────────────────────────────────────
      
      def main():
          if len(sys.argv) < 3:
              print(__doc__)
              sys.exit(1)
      
          cmd = sys.argv[1]
          repo_path = sys.argv[2]
      
          if cmd == "index":
              idx = Index()
              idx.build(repo_path)
              print(json.dumps(idx.stats(), indent=2))
      
          elif cmd == "search":
              if len(sys.argv) < 4:
                  print("Usage: code_rag.py search /path/to/repo \"query\" [options]",
                        file=sys.stderr)
                  sys.exit(1)
      
              query = sys.argv[3]
              top_k = 5
              grouped = "--grouped" in sys.argv
              rg = "--rg" in sys.argv
      
              for i, arg in enumerate(sys.argv):
                  if arg == "--top" and i + 1 < len(sys.argv):
                      top_k = int(sys.argv[i + 1])
      
              idx = Index()
              idx.build(repo_path)
              stats = idx.stats()
              print(f"Indexed {stats['chunks']} chunks from {stats['files']} files "
                    f"({stats['vocabulary']} features, {stats['build_ms']}ms)",
                    file=sys.stderr)
      
              if grouped or rg:
                  results = idx.search_grouped(query, top_k=top_k)
                  if rg:
                      print(_format_for_grep(results, repo_path))
                  else:
                      print(_format_grouped(results))
              else:
                  results = idx.search(query, top_k=top_k)
                  if results:
                      print(_format_results(results))
                  else:
                      print("  No results above threshold.", file=sys.stderr)
      
          else:
              print(f"Unknown command: {cmd}", file=sys.stderr)
              sys.exit(1)
      
      
      if __name__ == "__main__":
          main()
      
    • context.py 6.3 KB
      """
      Expand search match lines into full structural context (functions/classes).
      
      Uses tree-sitting's AST cache for symbol boundaries. Scans the repo once
      on first expand call (~700ms), then all expansions are sub-millisecond.
      Falls back to a fixed-size context window if tree-sitting is unavailable.
      """
      
      import os
      import sys
      from dataclasses import dataclass
      
      
      @dataclass
      class CodeContext:
          """A structural code unit containing a match."""
          file_path: str
          start_line: int
          end_line: int
          match_line: int
          node_type: str  # "function", "class", "method"
          name: str
          source: str
          language: str | None = None
          signature: str | None = None
      
      
      # Lazily initialized tree-sitting cache
      _cache = None
      _cache_root = None
      
      
      def _ensure_cache(search_root: str):
          """Scan the repo with tree-sitting on first call. No-op on subsequent calls."""
          global _cache, _cache_root
          if _cache is not None and _cache_root == search_root:
              return _cache
      
          try:
              ts_scripts = "/mnt/skills/user/tree-sitting/scripts"
              if ts_scripts not in sys.path:
                  sys.path.insert(0, ts_scripts)
              from engine import CodeCache
      
              _cache = CodeCache()
              _cache.scan(search_root)
              _cache_root = search_root
              return _cache
          except Exception as e:
              print(f"tree-sitting unavailable ({e}), using window fallback",
                    file=sys.stderr)
              return None
      
      
      def expand_match(file_path: str, line_number: int, search_root: str,
                       signatures_only: bool = True) -> CodeContext | None:
          """
          Expand a match at file:line into its containing function/class.
      
          Uses tree-sitting AST data for structural boundaries. Falls back to
          a context window around the match if tree-sitting is unavailable.
      
          Args:
              file_path: Absolute path to matched file
              line_number: 1-indexed line number of the match
              search_root: Root directory of the codebase
              signatures_only: Return only signature, not full body
          """
          cache = _ensure_cache(search_root)
          if cache is not None:
              return _expand_from_ast(file_path, line_number, search_root,
                                      cache, signatures_only)
          return _expand_window(file_path, line_number)
      
      
      def _expand_from_ast(file_path: str, line_number: int, search_root: str,
                           cache, signatures_only: bool) -> CodeContext | None:
          """Expand using tree-sitting's parsed AST symbols."""
          relpath = os.path.relpath(file_path, search_root)
      
          # Get symbols for this file from the cache
          entry = cache.files.get(relpath)
          if not entry or not entry.symbols:
              return _expand_window(file_path, line_number)
      
          # Find the innermost symbol containing this line
          containing = None
          for sym in entry.symbols:
              if sym.line <= line_number <= sym.end_line:
                  # Prefer the most specific (innermost) match
                  if containing is None:
                      containing = sym
                  else:
                      if sym.line >= containing.line:
                          containing = sym
              # Also check children (methods within classes)
              for child in getattr(sym, 'children', []):
                  if child.line <= line_number <= child.end_line:
                      if containing is None or child.line >= containing.line:
                          containing = child
      
          if not containing:
              # No containing symbol — try nearest preceding symbol
              best = None
              for sym in entry.symbols:
                  if sym.line <= line_number:
                      if best is None or sym.line > best.line:
                          best = sym
              containing = best
      
          if not containing:
              return _expand_window(file_path, line_number)
      
          start_line = containing.line
          end_line = containing.end_line or start_line
      
          try:
              with open(file_path, "r") as f:
                  lines = f.readlines()
          except (FileNotFoundError, PermissionError):
              return None
      
          # Ensure end_line doesn't exceed file
          end_line = min(end_line, len(lines))
      
          # Trim trailing blanks
          while end_line > start_line and not lines[end_line - 1].strip():
              end_line -= 1
      
          source = "".join(lines[start_line - 1:end_line])
      
          name = containing.name or ""
          kind = containing.kind or "function"
          # Normalize kind to node_type
          kind_map = {
              "class": "class", "struct": "class", "interface": "class",
              "enum": "class", "trait": "class",
              "method": "method", "impl_method": "method",
              "function": "function", "func": "function",
          }
          node_type = kind_map.get(kind, kind)
      
          signature = None
          if signatures_only:
              sig = containing.signature or ""
              if sig:
                  signature = sig
              else:
                  # Use first line as signature fallback
                  first_line = lines[start_line - 1].rstrip() if start_line <= len(lines) else name
                  signature = first_line
      
          return CodeContext(
              file_path=file_path, start_line=start_line, end_line=end_line,
              match_line=line_number, node_type=node_type, name=name,
              source=source, language=entry.lang, signature=signature,
          )
      
      
      def _expand_window(file_path: str, line_number: int,
                         context: int = 10) -> CodeContext | None:
          """Fallback: return a fixed window around the match."""
          try:
              with open(file_path, "r") as f:
                  lines = f.readlines()
          except (FileNotFoundError, PermissionError):
              return None
      
          start = max(1, line_number - context)
          end = min(len(lines), line_number + context)
          source = "".join(lines[start - 1:end])
      
          ext = os.path.splitext(file_path)[1].lower()
          ext_to_lang = {
              ".py": "python", ".js": "javascript", ".ts": "typescript",
              ".go": "go", ".rs": "rust", ".rb": "ruby", ".java": "java",
              ".c": "c", ".cpp": "cpp", ".cs": "csharp",
          }
      
          return CodeContext(
              file_path=file_path, start_line=start, end_line=end,
              match_line=line_number, node_type="context", name="",
              source=source, language=ext_to_lang.get(ext),
          )
      
      
      def deduplicate_contexts(contexts: list[CodeContext]) -> list[CodeContext]:
          """Remove duplicate expansions (same function from multiple match lines)."""
          seen = set()
          unique = []
          for ctx in contexts:
              key = (ctx.file_path, ctx.start_line, ctx.end_line)
              if key not in seen:
                  seen.add(key)
                  unique.append(ctx)
          return unique
      
    • lsp_refs.py 18.4 KB
      """Binding-resolved reference/definition tier for Python sources.
      
      Bridges searching-codebases to the ``python-lsp`` skill so that
      cross-reference queries on ``.py`` files are resolved by pyright's binder
      instead of by text matching. The win over the regex tier: a same-named but
      unrelated symbol (a shadowed name, an unrelated ``helper`` in another module)
      is *excluded*, and ``definition`` follows imports across files — neither of
      which ripgrep or tree-sitter can do.
      
      This tier is engaged lazily: only when a true cross-reference / definition /
      hover is requested, so pyright's index cost is never paid on ordinary text
      searches.
      
      Soft-fallback contract: every path that cannot produce a binding-resolved
      answer raises :class:`LspUnavailable` with a one-line reason. The caller
      (``search.py``) catches it, emits a one-line degradation note, and falls back
      to the existing regex/tree-sitting text path. Causes that trigger fallback:
      non-Python target (the symbol has no Python definition), pyright/node absent,
      or the python-lsp client being unimportable.
      
      Coordinate seam: tree-sitting / ripgrep positions are 1-based; LSP is 0-based.
      Conversion happens at the boundary via ``Position.from_one_based(line, col)``.
      """
      
      from __future__ import annotations
      
      import os
      import re
      import subprocess
      import sys
      from dataclasses import dataclass
      from pathlib import Path
      
      # Default directories to skip when collecting Python files / occurrences.
      _SKIP_DIRS = {
          ".git", ".hg", ".svn", "node_modules", "__pycache__", ".venv", "venv",
          "env", ".tox", "build", "dist", ".mypy_cache", ".pytest_cache", ".idea",
          ".eggs", "site-packages",
      }
      
      # Bound the number of files opened in the language server and the number of
      # occurrence positions queried, so a huge repo doesn't make the lazy path hang.
      _MAX_OPEN_FILES = 600
      _MAX_OCCURRENCES = 40
      
      
      class LspUnavailable(Exception):
          """Signals the binding-resolved path can't run — caller should fall back.
      
          The message is a single line suitable for a user-facing degradation note.
          """
      
      
      @dataclass
      class Site:
          """A 1-based position on a symbol token, plus context for display."""
          file: str          # path relative to root
          line: int          # 1-based line
          col: int           # 1-based column of the symbol token
          kind: str = ""     # tree-sitting symbol kind (function/class/method/...)
          line_text: str = ""
      
      
      # ── locating sibling skills ──────────────────────────────────────────────────
      
      def _skill_scripts(name: str, env_var: str | None = None) -> str | None:
          """Locate a sibling skill's ``scripts/`` dir across deploy and dev layouts.
      
          Order: explicit env override, the installed ``/mnt/skills/user`` location,
          then a sibling of this repo checkout (``<repo>/<name>/scripts``).
          """
          candidates: list[str] = []
          if env_var and os.environ.get(env_var):
              candidates.append(os.environ[env_var])
          candidates.append(f"/mnt/skills/user/{name}/scripts")
          # parents[0]=scripts, [1]=searching-codebases, [2]=repo root
          repo_root = Path(__file__).resolve().parents[2]
          candidates.append(str(repo_root / name / "scripts"))
          for c in candidates:
              if c and os.path.isdir(c):
                  return c
          return None
      
      
      def _import_lsp_client():
          """Import the python-lsp client, or raise LspUnavailable on a soft miss."""
          scripts = _skill_scripts("python-lsp", env_var="PYTHON_LSP_SCRIPTS")
          if not scripts:
              raise LspUnavailable("python-lsp skill not found (binding-resolved tier unavailable)")
          if scripts not in sys.path:
              sys.path.insert(0, scripts)
          try:
              import lsp_client
              return lsp_client
          except ImportError as e:  # pragma: no cover - import-path dependent
              raise LspUnavailable(f"cannot import python-lsp client: {e}")
      
      
      def _code_cache(root: str):
          """Build a tree-sitting CodeCache over ``root``, or raise LspUnavailable."""
          scripts = _skill_scripts("tree-sitting", env_var="TREE_SITTING_SCRIPTS")
          if not scripts:
              raise LspUnavailable("tree-sitting skill not found (cannot resolve symbol position)")
          if scripts not in sys.path:
              sys.path.insert(0, scripts)
          # The engine imports fine without tree_sitter, but its grammar loader
          # swallows the missing-dependency ImportError and silently returns no
          # parsers, so scan() yields zero symbols and the whole tier degrades to
          # regex with no real error. tree_sitter is not installed at boot and only
          # the semantic path installs its own dep, so ensure it here — same
          # `uv pip install --system` pattern search_semantic uses for scikit-learn.
          _ensure_tree_sitter()
          try:
              from engine import CodeCache
          except ImportError as e:  # pragma: no cover - import-path dependent
              raise LspUnavailable(f"cannot import tree-sitting engine: {e}")
          cache = CodeCache()
          cache.scan(root)
          return cache
      
      
      def _ensure_tree_sitter() -> None:
          """Install the bare ``tree-sitter`` package if absent (no-op when present).
      
          The tree-sitting engine loads bundled grammar ``.so`` files via ctypes
          against ``tree_sitter.Language``; without the package the loader returns
          None and scans produce nothing. Tries install strategies in order and
          verifies the import after each — ``uv pip --system`` silently no-ops under
          PEP 668 in the locked-down container, so ``pip --break-system-packages``
          leads.
          """
          try:
              import tree_sitter
              return
          except ImportError:
              pass
          for cmd in (
              [sys.executable, "-m", "pip", "install", "tree-sitter",
               "--break-system-packages", "-q"],
              ["uv", "pip", "install", "tree-sitter", "--system"],
          ):
              try:
                  subprocess.run(cmd, capture_output=True, timeout=120)
              except Exception:
                  continue
              try:
                  import tree_sitter  # noqa: F401
                  return
              except ImportError:
                  continue
      
      
      # ── symbol → position resolution ─────────────────────────────────────────────
      
      def _col_of(line_text: str, name: str) -> int:
          """1-based column of the first whole-word occurrence of ``name`` in a line.
      
          Falls back to a substring search, then to column 1, so a position is always
          produced even if the token is adjacent to non-word characters.
          """
          m = re.search(rf"\b{re.escape(name)}\b", line_text)
          if m:
              return m.start() + 1
          idx = line_text.find(name)
          return (idx + 1) if idx >= 0 else 1
      
      
      def _read_line(root: str, rel: str, line1: int) -> str:
          try:
              with open(os.path.join(root, rel), "r", errors="replace") as f:
                  lines = f.readlines()
          except OSError:
              return ""
          return lines[line1 - 1].rstrip("\n") if 0 < line1 <= len(lines) else ""
      
      
      def definition_sites(root: str, symbol: str) -> list[Site]:
          """Definition sites of ``symbol`` in Python files, via tree-sitting.
      
          Returns one :class:`Site` per ``def`` / ``class`` / method named exactly
          ``symbol`` (children included). Empty when the symbol has no Python
          definition — which is the non-Python / wrong-target fallback signal.
          """
          cache = _code_cache(root)
          sites: list[Site] = []
          for sym in cache.find_symbol(symbol, limit=200):
              if sym.name != symbol:
                  continue  # find_symbol does substring matching; we want exact
              if not sym.file.endswith(".py"):
                  continue
              text = _read_line(root, sym.file, sym.line)
              sites.append(Site(file=sym.file, line=sym.line, col=_col_of(text, symbol),
                                kind=sym.kind, line_text=text))
          return sites
      
      
      def _find_ripgrep() -> str | None:
          from shutil import which
          return which("rg")
      
      
      def occurrence_sites(root: str, symbol: str, limit: int = _MAX_OCCURRENCES) -> list[Site]:
          """Every textual ``\\bsymbol\\b`` occurrence in ``.py`` files under root.
      
          Used as anchors for ``definition`` (so it can follow an import from any use
          site) and to bound the set of files opened in the language server. Text
          matching here is intentional and safe — it is a *superset* of the real
          references; pyright narrows it to the binding-resolved set.
          """
          rg = _find_ripgrep()
          sites: list[Site] = []
          pattern = rf"\b{re.escape(symbol)}\b"
          if rg:
              cmd = [rg, "--no-heading", "--line-number", "--column", "--color=never",
                     "-t", "py"]
              for d in _SKIP_DIRS:
                  cmd.extend(["--glob", f"!{d}"])
              cmd.extend(["-e", pattern, root])
              try:
                  out = subprocess.run(cmd, capture_output=True, text=True, timeout=60).stdout
              except (subprocess.TimeoutExpired, subprocess.SubprocessError):
                  out = ""
              for line in out.splitlines():
                  parts = line.split(":", 3)
                  if len(parts) >= 4 and parts[1].isdigit() and parts[2].isdigit():
                      rel = os.path.relpath(parts[0], root)
                      sites.append(Site(file=rel, line=int(parts[1]), col=int(parts[2]),
                                        line_text=parts[3]))
                      if len(sites) >= limit:
                          break
              return sites
          # No ripgrep: fall back to a Python walk.
          rx = re.compile(pattern)
          for dirpath, dirnames, filenames in os.walk(root):
              dirnames[:] = [d for d in dirnames if d not in _SKIP_DIRS and not d.startswith(".")]
              for fn in filenames:
                  if not fn.endswith(".py"):
                      continue
                  fp = os.path.join(dirpath, fn)
                  try:
                      with open(fp, "r", errors="replace") as f:
                          for i, text in enumerate(f, start=1):
                              m = rx.search(text)
                              if m:
                                  sites.append(Site(file=os.path.relpath(fp, root), line=i,
                                                    col=m.start() + 1, line_text=text.rstrip("\n")))
                                  if len(sites) >= limit:
                                      return sites
                  except OSError:
                      continue
          return sites
      
      
      # ── the query ────────────────────────────────────────────────────────────────
      
      def _files_with_symbol(root: str, symbol: str) -> set[str]:
          """All ``.py`` files containing ``symbol`` (word-boundary), relative to ``root``.
      
          Unlike :func:`occurrence_sites`, this is **uncapped**. It bounds the set of
          files pyright opens; capping it — the old behaviour reused the
          40-occurrence-capped occurrence list — silently starved recall for frequent
          symbols, because call sites in files past the cap were never opened and so
          were invisible to pyright (e.g. requests ``get``: 3 refs reported vs 84
          real, the missing 81 living in unopened test files). A files-with-matches
          scan stays cheap even when occurrences number in the hundreds.
          """
          pattern = rf"\b{re.escape(symbol)}\b"
          found: set[str] = set()
          rg = _find_ripgrep()
          if rg:
              cmd = [rg, "-l", "-t", "py"]
              for d in _SKIP_DIRS:
                  cmd.extend(["--glob", f"!{d}"])
              cmd.extend(["-e", pattern, root])
              try:
                  out = subprocess.run(cmd, capture_output=True, text=True, timeout=60).stdout
              except (subprocess.TimeoutExpired, subprocess.SubprocessError):
                  out = ""
              for line in out.splitlines():
                  if line.strip():
                      found.add(os.path.relpath(line.strip(), root))
              return found
          # No ripgrep: Python walk.
          rx = re.compile(pattern)
          for dirpath, dirnames, filenames in os.walk(root):
              dirnames[:] = [d for d in dirnames if d not in _SKIP_DIRS and not d.startswith(".")]
              for fn in filenames:
                  if not fn.endswith(".py"):
                      continue
                  fp = os.path.join(dirpath, fn)
                  try:
                      with open(fp, "r", errors="replace") as f:
                          if any(rx.search(t) for t in f):
                              found.add(os.path.relpath(fp, root))
                  except OSError:
                      continue
          return found
      
      
      def lsp_query(root: str, symbol: str, op: str = "references",
                    max_open: int = _MAX_OPEN_FILES, verbose: bool = False) -> dict:
          """Run a binding-resolved ``op`` for ``symbol`` over the Python sources.
      
          ``op`` is one of ``references`` | ``definition`` | ``hover``.
      
          - ``references``: anchored at each definition site; pyright returns the
            binding-resolved uses, excluding same-named symbols of other bindings.
            Results are grouped per definition (per binding).
          - ``definition``: anchored at each occurrence; follows imports to the real
            definition(s); results are unioned and de-duplicated.
          - ``hover``: anchored at each definition site; returns the inferred
            type/signature string.
      
          Raises :class:`LspUnavailable` (soft fallback) when the binding-resolved
          answer cannot be produced.
          """
          if op not in ("references", "definition", "hover"):
              raise ValueError(f"unknown op: {op!r}")
      
          lsp = _import_lsp_client()
          Position = lsp.Position
      
          defs = definition_sites(root, symbol)
          occurrences = occurrence_sites(root, symbol)
          if not defs and not occurrences:
              raise LspUnavailable(f"no Python occurrence of {symbol!r} found")
          if op in ("references", "hover") and not defs:
              raise LspUnavailable(f"no Python definition of {symbol!r} found")
      
          # Ensure pyright is available; fail soft (one-line reason) if not.
          try:
              lsp.ensure_pyright(install=True)
          except lsp.BootstrapError as e:
              raise LspUnavailable(str(e).splitlines()[0])
      
          # Relevant files to open = every file that mentions the symbol, plus the
          # definition files. This builds pyright's model for exactly the files that
          # could contain a reference, without indexing the whole repo.
          open_files: set[str] = _files_with_symbol(root, symbol) | {s.file for s in defs}
          files = sorted(f for f in open_files if f.endswith(".py"))
          truncated = len(files) > max_open
          if truncated:
              files = files[:max_open]
      
          note_parts: list[str] = []
          results: list[dict] = []
          try:
              with lsp.LSPClient(root) as client:
                  if files:
                      client.open_all(*files)
                  if not client.wait_for_index(timeout=30):
                      note_parts.append("pyright index wait timed out; results may be incomplete")
      
                  if op == "references":
                      for d in defs:
                          pos = Position.from_one_based(d.line, d.col)
                          locs = client.references(d.file, pos.line, pos.character)
                          results.append({"anchor": d, "locations": locs})
                  elif op == "hover":
                      for d in defs:
                          pos = Position.from_one_based(d.line, d.col)
                          results.append({"anchor": d, "hover": client.hover(d.file, pos.line, pos.character)})
                  else:  # definition
                      anchors = occurrences or defs
                      seen = set()
                      union = []
                      for a in anchors:
                          pos = Position.from_one_based(a.line, a.col)
                          for loc in client.definition(a.file, pos.line, pos.character):
                              key = (loc.path, loc.start_line, loc.start_char)
                              if key not in seen:
                                  seen.add(key)
                                  union.append(loc)
                      results.append({"anchor": None, "locations": union})
          except LspUnavailable:
              raise
          except Exception as e:  # any client/runtime failure → soft fallback
              raise LspUnavailable(f"binding-resolved query failed: {e}")
      
          if truncated:
              note_parts.append(f"opened first {max_open} of {len(open_files)} candidate files")
      
          return {
              "op": op,
              "symbol": symbol,
              "results": results,
              "opened": len(files),
              "note": "; ".join(note_parts) or None,
          }
      
      
      # ── formatting ───────────────────────────────────────────────────────────────
      
      def _rel(path: str, root: str) -> str:
          try:
              return os.path.relpath(path, root)
          except ValueError:
              return path
      
      
      def format_lsp(result: dict, root: str, json_out: bool = False) -> str:
          """Render an :func:`lsp_query` result. LSP 0-based positions become 1-based."""
          if json_out:
              import json
              payload = {"op": result["op"], "symbol": result["symbol"],
                         "opened": result["opened"], "note": result["note"], "results": []}
              for grp in result["results"]:
                  entry: dict = {}
                  anchor = grp.get("anchor")
                  if anchor is not None:
                      entry["anchor"] = {"file": anchor.file, "line": anchor.line,
                                         "col": anchor.col, "kind": anchor.kind}
                  if "hover" in grp:
                      entry["hover"] = grp["hover"]
                  if "locations" in grp:
                      entry["locations"] = [
                          {"file": _rel(loc.path, root), "line": loc.start_line + 1,
                           "col": loc.start_char + 1} for loc in grp["locations"]
                      ]
                  payload["results"].append(entry)
              return json.dumps(payload, indent=2)
      
          op = result["op"]
          sym = result["symbol"]
          lines = [f"{sym} — binding-resolved {op} (python-lsp)"]
          total = 0
          for grp in result["results"]:
              anchor = grp.get("anchor")
              if "hover" in grp:
                  where = f"{anchor.file}:{anchor.line}" if anchor else ""
                  hover = grp["hover"] or "(no type information)"
                  lines.append(f"  {where}  {hover}")
                  total += 1
                  continue
              locs = grp.get("locations", [])
              if anchor is not None:
                  lines.append(f"  definition {anchor.file}:{anchor.line}"
                               f"{f' ({anchor.kind})' if anchor.kind else ''}")
              for loc in locs:
                  lines.append(f"    {_rel(loc.path, root)}:{loc.start_line + 1}:{loc.start_char + 1}")
                  total += 1
              if anchor is not None and not locs:
                  lines.append("    (no references)")
          if op == "references":
              lines.append(f"  {total} reference{'s' if total != 1 else ''} "
                           f"(binding-resolved; same-named symbols of other bindings excluded)")
          if result.get("note"):
              lines.append(f"  note: {result['note']}")
          return "\n".join(lines)
      
    • ngram_index.py 20.7 KB
      """
      Inverted index for fast regex search using sparse n-grams.
      
      Architecture:
      - Index maps n-gram hashes → sets of file IDs
      - File IDs map back to file paths
      - Query decomposes regex into literals → covering n-grams → posting list intersection
      - Candidates verified by actual regex match (ripgrep or Python re)
      """
      
      import os
      import re
      import sre_parse
      import subprocess
      import time
      from pathlib import Path
      
      from sparse_ngrams import (
          FrequencyWeights,
          build_all,
          build_covering,
          compute_weights,
          ngram_hash,
          weight_crc32,
      )
      
      # Default extensions to index (source code)
      DEFAULT_EXTENSIONS = {
          ".py", ".js", ".jsx", ".ts", ".tsx", ".mjs", ".mts",
          ".go", ".rs", ".rb", ".java", ".c", ".h", ".cpp", ".hpp", ".cc",
          ".cs", ".php", ".swift", ".kt", ".scala", ".lua", ".zig",
          ".sh", ".bash", ".zsh", ".fish",
          ".html", ".css", ".scss", ".less",
          ".json", ".yaml", ".yml", ".toml", ".xml", ".md", ".txt", ".rst",
          ".sql", ".graphql", ".proto",
          ".dockerfile", ".env", ".ini", ".cfg", ".conf",
          ".r", ".R", ".jl", ".ex", ".exs", ".erl", ".hrl",
          ".vim", ".el", ".clj", ".cljs", ".ml", ".mli", ".hs",
      }
      
      # Default directories to skip
      DEFAULT_SKIP_DIRS = {
          ".git", "node_modules", "__pycache__", ".venv", "venv",
          "dist", "build", ".next", ".cache", "target", "vendor",
          ".tox", ".mypy_cache", ".pytest_cache", "coverage",
          ".idea", ".vscode", ".eclipse",
      }
      
      # Max file size to index (skip giant generated files)
      MAX_FILE_SIZE = 1_000_000  # 1MB
      
      
      # @lat: [[code-intelligence#N-gram Indexing]]
      class NgramIndex:
          """
          Sparse n-gram inverted index for a directory of source files.
      
          Build once, query many times. Candidate files from index lookup
          are verified with actual regex matching for correctness.
          """
      
          def __init__(self, weight_fn=None):
              # n-gram hash → set of file IDs
              self.postings: dict[int, set[int]] = {}
              # file ID → file path
              self.files: dict[int, str] = {}
              # file path → file ID (reverse lookup)
              self._path_to_id: dict[str, int] = {}
              # next file ID
              self._next_id: int = 0
              # weight function
              self._weight_fn = weight_fn or weight_crc32
              # frequency weights (trained from corpus)
              self._freq_weights: FrequencyWeights | None = None
              # stats
              self.stats = {
                  "files_indexed": 0,
                  "files_skipped": 0,
                  "total_ngrams": 0,
                  "unique_ngrams": 0,
                  "index_time_ms": 0,
                  "total_bytes": 0,
              }
      
          def _assign_id(self, path: str) -> int:
              """Assign a numeric ID to a file path."""
              if path in self._path_to_id:
                  return self._path_to_id[path]
              fid = self._next_id
              self._next_id += 1
              self.files[fid] = path
              self._path_to_id[path] = fid
              return fid
      
          def _should_index(self, path: str, skip_dirs: set[str]) -> bool:
              """Check if a file should be indexed."""
              p = Path(path)
      
              # Check directory exclusions
              for part in p.parts:
                  if part in skip_dirs:
                      return False
      
              # Check extension
              suffix = p.suffix.lower()
              if suffix and suffix not in DEFAULT_EXTENSIONS:
                  # Also accept extensionless files (Makefile, Dockerfile, etc.)
                  if suffix:
                      return False
      
              # Check if it's a Makefile/Dockerfile/etc (no extension)
              name = p.name.lower()
              known_names = {
                  "makefile", "dockerfile", "vagrantfile", "gemfile",
                  "rakefile", "procfile", "brewfile", "justfile",
              }
              if not suffix and name not in known_names:
                  return False
      
              return True
      
          def build(
              self,
              root: str,
              skip_dirs: set[str] | None = None,
              use_frequency_weights: bool = True,
              verbose: bool = False,
          ):
              """
              Build the index from all source files under root.
      
              If use_frequency_weights is True, makes two passes:
              1. Train frequency table on all file contents
              2. Index using frequency-based weights (rare pairs = high weight)
              """
              skip = skip_dirs or DEFAULT_SKIP_DIRS
              t_start = time.monotonic()
      
              # Collect files
              file_paths = []
              for dirpath, dirnames, filenames in os.walk(root):
                  # Prune skip dirs
                  dirnames[:] = [d for d in dirnames if d not in skip]
                  for fname in filenames:
                      fpath = os.path.join(dirpath, fname)
                      if self._should_index(fpath, skip):
                          try:
                              size = os.path.getsize(fpath)
                              if size <= MAX_FILE_SIZE and size > 0:
                                  file_paths.append(fpath)
                              else:
                                  self.stats["files_skipped"] += 1
                          except OSError:
                              self.stats["files_skipped"] += 1
      
              if verbose:
                  print(f"Found {len(file_paths)} files to index")
      
              # Pass 1: train frequency weights (optional)
              if use_frequency_weights:
                  if verbose:
                      print("Pass 1: Training frequency weights...")
                  fw = FrequencyWeights()
                  for fpath in file_paths:
                      try:
                          with open(fpath, "rb") as f:
                              data = f.read()
                          fw.train(data)
                          self.stats["total_bytes"] += len(data)
                      except (OSError, UnicodeDecodeError):
                          pass
                  fw.freeze()
                  self._freq_weights = fw
                  self._weight_fn = fw.weight
                  if verbose:
                      print(f"  Trained on {self.stats['total_bytes']:,} bytes")
              else:
                  # Single pass, count bytes as we go
                  pass
      
              # Pass 2 (or only pass): build index
              if verbose:
                  print("Building sparse n-gram index...")
              total_ngrams = 0
              for i, fpath in enumerate(file_paths):
                  try:
                      with open(fpath, "rb") as f:
                          data = f.read()
      
                      if not use_frequency_weights:
                          self.stats["total_bytes"] += len(data)
      
                      fid = self._assign_id(fpath)
                      weights = compute_weights(data, self._weight_fn)
                      ngrams = build_all(weights)
      
                      for start, end in ngrams:
                          h = ngram_hash(data, start, end)
                          if h not in self.postings:
                              self.postings[h] = set()
                          self.postings[h].add(fid)
                          total_ngrams += 1
      
                      self.stats["files_indexed"] += 1
      
                      if verbose and (i + 1) % 500 == 0:
                          print(f"  Indexed {i + 1}/{len(file_paths)} files...")
      
                  except (OSError, UnicodeDecodeError):
                      self.stats["files_skipped"] += 1
      
              elapsed_ms = (time.monotonic() - t_start) * 1000
              self.stats["total_ngrams"] = total_ngrams
              self.stats["unique_ngrams"] = len(self.postings)
              self.stats["index_time_ms"] = round(elapsed_ms, 1)
      
              if verbose:
                  print(f"Index built: {self.stats['files_indexed']} files, "
                        f"{self.stats['unique_ngrams']:,} unique n-grams, "
                        f"{elapsed_ms:.0f}ms")
      
          def _query_literal(self, literal: bytes) -> set[int] | None:
              """
              Query the index with a literal byte string.
              Returns set of candidate file IDs, or None if no n-grams extracted.
              """
              if len(literal) < 3:
                  # Too short for meaningful n-gram lookup
                  return None
      
              weights = compute_weights(literal, self._weight_fn)
              covering = build_covering(weights)
      
              if not covering:
                  return None
      
              # Intersect posting lists for all covering n-grams
              candidates = None
              for start, end in covering:
                  h = ngram_hash(literal, start, end)
                  posting = self.postings.get(h, set())
                  if candidates is None:
                      candidates = set(posting)
                  else:
                      candidates &= posting
                  # Early termination: empty intersection
                  if candidates is not None and not candidates:
                      return set()
      
              return candidates
      
          def _eval_plan(self, plan: "QueryPlan") -> set[int] | None:
              """
              Evaluate a query plan tree against the index.
      
              AND → intersect posting lists
              OR  → union posting lists
              LITERAL → look up covering n-grams
              """
              if plan.op == "literal":
                  return self._query_literal(plan.literal)
              elif plan.op == "and":
                  result = None
                  for child in plan.children:
                      child_ids = self._eval_plan(child)
                      if child_ids is not None:
                          if result is None:
                              result = set(child_ids)
                          else:
                              result &= child_ids
                          if not result:
                              return set()
                  return result
              elif plan.op == "or":
                  result = set()
                  for child in plan.children:
                      child_ids = self._eval_plan(child)
                      if child_ids is None:
                          # This branch can't be indexed — any file could match
                          # So the whole OR is unbounded
                          return None
                      result |= child_ids
                  return result
              return None
      
          def search(
              self,
              pattern: str,
              root: str,
              max_results: int = 100,
              verbose: bool = False,
          ) -> list[dict]:
              """
              Search for a regex pattern. Returns list of matches.
      
              Each match: {"file": path, "line": line_number, "text": line_text}
      
              Pipeline:
              1. Parse regex into query plan tree (preserving AND/OR structure)
              2. Evaluate plan against index (intersect/union posting lists)
              3. Verify candidates with ripgrep (or Python re fallback)
              """
              start = time.monotonic()
      
              # Build query plan from regex
              plan = extract_query_plan(pattern)
      
              if verbose:
                  print(f"Pattern: {pattern}")
                  print(f"Query plan: {plan}")
      
              # Evaluate plan against index
              candidate_ids: set[int] | None = None
              if plan is not None:
                  candidate_ids = self._eval_plan(plan)
      
              if candidate_ids is None:
                  # No usable plan — fall back to scanning all files
                  candidate_files = list(self.files.values())
                  if verbose:
                      print(f"No usable literals — scanning all {len(candidate_files)} files")
              else:
                  candidate_files = [self.files[fid] for fid in candidate_ids if fid in self.files]
                  if verbose:
                      reduction = (1 - len(candidate_files) / max(len(self.files), 1)) * 100
                      print(f"Index narrowed to {len(candidate_files)}/{len(self.files)} "
                            f"files ({reduction:.0f}% reduction)")
      
              # Verify with actual regex matching
              matches = verify_candidates(pattern, candidate_files, root, max_results, verbose)
      
              elapsed_ms = (time.monotonic() - start) * 1000
              if verbose:
                  print(f"Search completed: {len(matches)} matches in {elapsed_ms:.0f}ms")
      
              return matches
      
      
      class QueryPlan:
          """
          Tree structure representing how to combine index lookups.
          
          AND nodes: all children must match (intersect posting lists)
          OR nodes: any child can match (union posting lists)
          LITERAL nodes: leaf — look up in index
          """
          __slots__ = ("children", "literal", "op")
      
          def __init__(self, op: str, children=None, literal: bytes = None):
              self.op = op  # "and", "or", "literal"
              self.children = children or []
              self.literal = literal
      
          def __repr__(self):
              if self.op == "literal":
                  return f"LIT({self.literal!r})"
              kids = ", ".join(repr(c) for c in self.children)
              return f"{self.op.upper()}({kids})"
      
      
      def extract_query_plan(pattern: str) -> QueryPlan | None:
          """
          Parse a regex and produce a QueryPlan tree preserving AND/OR structure.
      
          Sequential literals → AND (intersect)
          Alternations (a|b) → OR (union)
          """
          try:
              parsed = sre_parse.parse(pattern)
          except Exception:
              encoded = pattern.encode("utf-8")
              if len(encoded) >= 3:
                  return QueryPlan("literal", literal=encoded)
              return None
      
          plan = _build_plan(parsed)
          return _simplify(plan)
      
      
      def _build_plan(parsed) -> QueryPlan | None:
          """Recursively build a query plan from parsed regex."""
          parts = []  # AND children from sequential processing
          current = bytearray()
      
          def flush():
              nonlocal current
              if current:
                  parts.append(QueryPlan("literal", literal=bytes(current)))
              current = bytearray()
      
          for op, av in parsed:
              if op == sre_parse.LITERAL:
                  current.append(av)
              elif op == sre_parse.AT:
                  pass  # anchors
              elif op == sre_parse.SUBPATTERN:
                  if av[3] is not None:
                      sub = _build_plan(av[3])
                      if sub:
                          flush()
                          parts.append(sub)
              elif op == sre_parse.BRANCH:
                  flush()
                  branches = []
                  for branch in av[1]:
                      bp = _build_plan(branch)
                      if bp:
                          branches.append(bp)
                  if branches:
                      if len(branches) == 1:
                          parts.append(branches[0])
                      else:
                          parts.append(QueryPlan("or", children=branches))
              else:
                  flush()
      
          flush()
      
          if not parts:
              return None
          if len(parts) == 1:
              return parts[0]
          return QueryPlan("and", children=parts)
      
      
      def _simplify(plan: QueryPlan | None) -> QueryPlan | None:
          """Flatten nested AND(AND(...)) and OR(OR(...)) nodes."""
          if plan is None:
              return None
          if plan.op == "literal":
              return plan
      
          new_children = []
          for child in plan.children:
              s = _simplify(child)
              if s is None:
                  continue
              # Flatten: AND(AND(a,b), c) → AND(a,b,c)
              if s.op == plan.op:
                  new_children.extend(s.children)
              else:
                  new_children.append(s)
      
          if not new_children:
              return None
          if len(new_children) == 1:
              return new_children[0]
          return QueryPlan(plan.op, children=new_children)
      
      
      def extract_literals(pattern: str) -> list[bytes]:
          """
          Legacy interface: extract flat list of AND-literals from a pattern.
          For simple patterns without alternation — all literals must be present.
          """
          plan = extract_query_plan(pattern)
          if plan is None:
              return []
          lits = []
          _collect_and_literals(plan, lits)
          return lits
      
      
      def _collect_and_literals(plan: QueryPlan, out: list):
          """Collect literals from AND nodes only (conservative: no OR)."""
          if plan.op == "literal":
              out.append(plan.literal)
          elif plan.op == "and":
              for child in plan.children:
                  _collect_and_literals(child, out)
          # OR nodes are skipped — can't AND their children
      
      
      def verify_candidates(
          pattern: str,
          candidate_files: list[str],
          root: str,
          max_results: int = 100,
          verbose: bool = False,
      ) -> list[dict]:
          """
          Verify candidate files by actually running the regex.
          Uses ripgrep if available, falls back to Python re.
          """
          if not candidate_files:
              return []
      
          # Try ripgrep first
          rg = _find_ripgrep()
          if rg:
              return _verify_ripgrep(rg, pattern, candidate_files, max_results, verbose)
          else:
              return _verify_python(pattern, candidate_files, max_results, verbose)
      
      
      def _find_ripgrep() -> str | None:
          """Find ripgrep binary."""
          for name in ["rg", "ripgrep"]:
              try:
                  result = subprocess.run(
                      ["which", name], capture_output=True, text=True
                  )
                  if result.returncode == 0:
                      return result.stdout.strip()
              except FileNotFoundError:
                  pass
          return None
      
      
      def _verify_ripgrep(
          rg_path: str,
          pattern: str,
          files: list[str],
          max_results: int,
          verbose: bool,
      ) -> list[dict]:
          """Verify candidates using ripgrep for speed."""
          matches = []
      
          # ripgrep can accept file list via stdin with --files-from
          # But simpler: pass files as arguments (may hit ARG_MAX for huge lists)
          # For large lists, batch them
          BATCH_SIZE = 500
          for i in range(0, len(files), BATCH_SIZE):
              batch = files[i : i + BATCH_SIZE]
              cmd = [
                  rg_path,
                  "--no-heading",
                  "--with-filename",
                  "--line-number",
                  "--color=never",
                  "--max-count=50",  # limit per file
                  "-e", pattern,
              ] + batch
      
              try:
                  result = subprocess.run(
                      cmd, capture_output=True, text=True, timeout=30
                  )
                  for line in result.stdout.splitlines():
                      # Format: file:line:text — but file paths could contain colons
                      # Use --line-number to guarantee line field is numeric
                      parts = line.split(":", 2)
                      if len(parts) >= 3 and parts[1].isdigit():
                          matches.append({
                              "file": parts[0],
                              "line": int(parts[1]),
                              "text": parts[2],
                          })
                          if len(matches) >= max_results:
                              return matches
              except (subprocess.TimeoutExpired, subprocess.SubprocessError) as e:
                  if verbose:
                      print(f"ripgrep error: {e}")
      
          return matches
      
      
      def _verify_python(
          pattern: str,
          files: list[str],
          max_results: int,
          verbose: bool,
      ) -> list[dict]:
          """Verify candidates using Python re (fallback)."""
          matches = []
          try:
              compiled = re.compile(pattern)
          except re.error as e:
              if verbose:
                  print(f"Invalid regex: {e}")
              return []
      
          for fpath in files:
              try:
                  with open(fpath, "r", errors="replace") as f:
                      for line_num, line in enumerate(f, 1):
                          if compiled.search(line):
                              matches.append({
                                  "file": fpath,
                                  "line": line_num,
                                  "text": line.rstrip(),
                              })
                              if len(matches) >= max_results:
                                  return matches
              except OSError:
                  pass
      
          return matches
      
      
      def _brute_force_search(
          pattern: str,
          root: str,
          skip_dirs: set[str],
          max_results: int = 100,
      ) -> list[dict]:
          """
          Brute-force search (no index) for benchmarking comparison.
          Walks all files and matches with ripgrep or Python re.
          """
          rg = _find_ripgrep()
          if rg:
              cmd = [
                  rg,
                  "--no-heading",
                  "--line-number",
                  "--color=never",
              ]
              for d in skip_dirs:
                  cmd.extend(["--glob", f"!{d}"])
              cmd.extend(["-e", pattern, root])
      
              try:
                  result = subprocess.run(cmd, capture_output=True, text=True, timeout=60)
                  matches = []
                  for line in result.stdout.splitlines():
                      parts = line.split(":", 2)
                      if len(parts) >= 3 and parts[1].isdigit():
                          matches.append({
                              "file": parts[0],
                              "line": int(parts[1]),
                              "text": parts[2],
                          })
                          if len(matches) >= max_results:
                              break
                  return matches
              except (subprocess.TimeoutExpired, subprocess.SubprocessError):
                  pass
      
          # Fallback: Python re over all files
          try:
              compiled = re.compile(pattern)
          except re.error:
              return []
      
          matches = []
          for dirpath, dirnames, filenames in os.walk(root):
              dirnames[:] = [d for d in dirnames if d not in skip_dirs]
              for fname in filenames:
                  fpath = os.path.join(dirpath, fname)
                  try:
                      size = os.path.getsize(fpath)
                      if size > MAX_FILE_SIZE or size == 0:
                          continue
                      with open(fpath, "r", errors="replace") as f:
                          for line_num, line in enumerate(f, 1):
                              if compiled.search(line):
                                  matches.append({
                                      "file": fpath,
                                      "line": line_num,
                                      "text": line.rstrip(),
                                  })
                                  if len(matches) >= max_results:
                                      return matches
                  except (OSError, UnicodeDecodeError):
                      pass
          return matches
      
    • resolve.py 4.5 KB
      """
      Resolve code sources to a local directory path.
      
      Handles: GitHub URLs, local directories, uploaded archives,
      uploaded files, project knowledge files.
      """
      
      import os
      import shutil
      import tarfile
      import urllib.request
      import zipfile
      
      WORK_DIR = "/home/claude/code-search-workspace"
      
      
      def resolve(source: str, branch: str = "main") -> str:
          """
          Resolve a source to a local directory path.
      
          Args:
              source: GitHub URL, local path, or "uploads" / "project"
              branch: Git branch for GitHub URLs
      
          Returns:
              Absolute path to a directory containing the code
          """
          # GitHub URL
          if source.startswith(("http://", "https://")):
              return _resolve_github(source, branch)
      
          # Explicit "uploads" keyword
          if source.lower() in ("uploads", "uploaded"):
              return _resolve_uploads()
      
          # Explicit "project" keyword
          if source.lower() in ("project", "project-knowledge"):
              return _resolve_project()
      
          # Archive file path
          p = os.path.expanduser(source)
          if os.path.isfile(p):
              return _resolve_archive(p)
      
          # Local directory
          if os.path.isdir(p):
              return os.path.abspath(p)
      
          raise FileNotFoundError(f"Cannot resolve source: {source}")
      
      
      def _resolve_github(url: str, branch: str) -> str:
          """Download a GitHub repo tarball and extract it."""
          url = url.rstrip("/")
          url = url.removesuffix(".git")
      
          parts = url.replace("https://github.com/", "").split("/")
          if len(parts) < 2:
              raise ValueError(f"Cannot parse GitHub URL: {url}")
      
          owner, repo = parts[0], parts[1]
          dest = os.path.join(WORK_DIR, f"{owner}-{repo}")
      
          # Clean previous download
          if os.path.exists(dest):
              shutil.rmtree(dest)
          os.makedirs(dest, exist_ok=True)
      
          tarball_url = f"https://codeload.github.com/{owner}/{repo}/tar.gz/{branch}"
          tar_path = os.path.join(dest, "repo.tar.gz")
      
          for attempt in range(3):
              try:
                  urllib.request.urlretrieve(tarball_url, tar_path)
                  with tarfile.open(tar_path) as tf:
                      tf.extractall(dest)
                  os.remove(tar_path)
                  # Find extracted directory
                  for entry in os.listdir(dest):
                      ep = os.path.join(dest, entry)
                      if os.path.isdir(ep):
                          return ep
              except Exception as e:
                  if attempt == 2:
                      raise RuntimeError(f"Failed to download {owner}/{repo}@{branch}: {e}")
                  continue
      
          return dest
      
      
      def _resolve_uploads() -> str:
          """Use uploaded files from /mnt/user-data/uploads/."""
          uploads = "/mnt/user-data/uploads"
          if not os.path.isdir(uploads):
              raise FileNotFoundError("No uploads directory found")
      
          entries = os.listdir(uploads)
          if not entries:
              raise FileNotFoundError("No uploaded files found")
      
          # If there's a single archive, extract it
          archives = [e for e in entries if e.endswith((".zip", ".tar.gz", ".tgz", ".tar"))]
          if len(archives) == 1 and len(entries) == 1:
              return _resolve_archive(os.path.join(uploads, archives[0]))
      
          # Otherwise use the uploads directory as-is
          return uploads
      
      
      def _resolve_project() -> str:
          """Use project knowledge files from /mnt/project/."""
          project = "/mnt/project"
          if not os.path.isdir(project):
              raise FileNotFoundError("No project directory found")
          return project
      
      
      def _resolve_archive(path: str) -> str:
          """Extract an archive to a temp directory."""
          dest = os.path.join(WORK_DIR, "extracted")
          if os.path.exists(dest):
              shutil.rmtree(dest)
          os.makedirs(dest)
      
          if path.endswith(".zip"):
              with zipfile.ZipFile(path) as zf:
                  zf.extractall(dest)
          elif path.endswith((".tar.gz", ".tgz")) or path.endswith(".tar"):
              with tarfile.open(path) as tf:
                  tf.extractall(dest)
          else:
              raise ValueError(f"Unknown archive format: {path}")
      
          # If archive contained a single directory, return that
          entries = os.listdir(dest)
          if len(entries) == 1 and os.path.isdir(os.path.join(dest, entries[0])):
              return os.path.join(dest, entries[0])
          return dest
      
      
      def count_files(root: str, skip_dirs: set = None) -> int:
          """Quick file count for deciding whether indexing is worthwhile."""
          skip = skip_dirs or {".git", "node_modules", "__pycache__", ".venv", "venv",
                               "dist", "build", ".next", "target", "vendor"}
          count = 0
          for _, dirs, files in os.walk(root):
              dirs[:] = [d for d in dirs if d not in skip]
              count += len(files)
          return count
      
    • search.py 12.8 KB
      #!/usr/bin/env python3
      """
      Unified code search: regex (n-gram indexed) and semantic (TF-IDF).
      
      Usage:
          # Auto-detect query type
          python search.py /path/to/repo "def handle_error"
          python search.py /path/to/repo "retry logic with backoff"
      
          # Explicit mode
          python search.py /path/to/repo "class.*Error" --regex
          python search.py /path/to/repo "error handling" --semantic
      
          # Multiple queries
          python search.py /path/to/repo "def test_" "import os" "TODO|FIXME"
      
          # GitHub repo
          python search.py https://github.com/org/repo "authentication flow"
      
          # Expand to full function bodies via tree-sitting AST
          python search.py /path/to/repo "query" --expand
      
          # Benchmark regex search: indexed vs brute-force
          python search.py /path/to/repo "pattern" --benchmark
      
          # JSON output
          python search.py /path/to/repo "query" --json
      """
      
      import argparse
      import json
      import os
      import subprocess
      import sys
      import time
      
      # Add script directory to path
      sys.path.insert(0, os.path.dirname(__file__))
      
      from resolve import count_files, resolve
      
      # Regex metacharacters that signal "this is a regex, not natural language"
      _REGEX_META = {'*', '+', '?', '[', ']', '(', ')', '{', '}', '|', '^', '$', '\\', '.'}
      # Only flag as regex if the "exotic" ones appear (not just . or parens)
      _STRONG_REGEX_META = {'*', '+', '?', '[', ']', '{', '}', '^', '$', '\\', '|'}
      
      
      # @lat: [[code-intelligence#Multi-Modal Search]]
      def detect_mode(query: str) -> str:
          """
          Heuristic: is this query a regex/literal or a conceptual search?
      
          Returns "regex" or "semantic".
          """
          # Explicit regex markers
          if any(c in query for c in _STRONG_REGEX_META):
              return "regex"
      
          # Short queries with code-like tokens → regex (literal search)
          words = query.split()
          if len(words) <= 3:
              # Looks like code: contains underscores, dots, camelCase, parens
              if any(c in query for c in "_.()"):
                  return "regex"
              # Single identifier
              if len(words) == 1:
                  return "regex"
      
          # Multi-word queries without code markers → semantic
          if len(words) >= 3:
              return "semantic"
      
          return "regex"
      
      
      def search_regex(root: str, queries: list, expand: bool = False,
                       benchmark: bool = False, verbose: bool = False,
                       skip_dirs: set = None) -> dict:
          """
          Regex/literal search using sparse n-gram index.
      
          Returns {query: [matches]} where each match has file, line, text,
          and optionally context (expanded function).
          """
          from ngram_index import NgramIndex, _brute_force_search
      
          # Build index
          index = NgramIndex()
          file_count = count_files(root, skip_dirs)
      
          # Skip indexing for tiny codebases — just use ripgrep directly
          if file_count < 20 and not benchmark:
              if verbose:
                  print(f"Small codebase ({file_count} files), using direct search", file=sys.stderr)
              results = {}
              for q in queries:
                  from ngram_index import DEFAULT_SKIP_DIRS, _brute_force_search
                  matches = _brute_force_search(q, root, skip_dirs or DEFAULT_SKIP_DIRS)
                  results[q] = _maybe_expand(matches, root, expand)
              return results
      
          index.build(root, skip_dirs=skip_dirs,
                      use_frequency_weights=True, verbose=verbose)
      
          if verbose:
              s = index.stats
              print(f"Index: {s['files_indexed']} files, {s['unique_ngrams']:,} n-grams, "
                    f"{s['index_time_ms']:.0f}ms", file=sys.stderr)
      
          results = {}
          for q in queries:
              if benchmark:
                  _run_benchmark(index, q, root, skip_dirs or set())
              else:
                  matches = index.search(q, root, max_results=500, verbose=verbose)
                  results[q] = _maybe_expand(matches, root, expand)
      
          return results
      
      
      def search_semantic(root: str, queries: list, expand: bool = False,
                          verbose: bool = False, skip_dirs: set = None) -> dict:
          """
          Semantic search using TF-IDF over code chunks.
      
          Returns {query: [matches]} with file, line, text, score.
          """
          # Ensure sklearn is available
          try:
              from code_rag import Index
          except ImportError:
              subprocess.run(
                  ["uv", "pip", "install", "scikit-learn", "--system"],
                  capture_output=True,
              )
              from code_rag import Index
      
          index = Index()
          index.build(root, skip_dirs=skip_dirs)
      
          if verbose:
              s = index.stats()
              print(f"TF-IDF index: {s.get('chunks', 0)} chunks, {s.get('vocabulary', 0)} terms, "
                    f"{s.get('build_ms', 0)}ms", file=sys.stderr)
      
          results = {}
          for q in queries:
              hits = index.search(q, top_k=20)
              matches = []
              for chunk, score in hits:
                  matches.append({
                      "file": os.path.join(root, chunk.file) if not os.path.isabs(chunk.file) else chunk.file,
                      "line": chunk.line,
                      "text": chunk.text[:200],
                      "score": round(score, 4),
                      "kind": chunk.kind,
                      "name": chunk.name,
                  })
              results[q] = matches
      
          return results
      
      
      def _maybe_expand(matches: list, root: str, expand: bool) -> list:
          """Optionally expand matches to full function context."""
          if not expand:
              return matches
      
          from context import expand_match
      
          contexts = []
          for m in matches:
              ctx = expand_match(m["file"], m["line"], root, signatures_only=False)
              if ctx:
                  m["context"] = {
                      "name": ctx.name,
                      "type": ctx.node_type,
                      "start_line": ctx.start_line,
                      "end_line": ctx.end_line,
                      "source": ctx.source,
                  }
          return matches
      
      
      def _run_benchmark(index, pattern, root, skip_dirs):
          """Compare indexed search vs brute-force ripgrep."""
          from ngram_index import _brute_force_search
      
          t0 = time.monotonic()
          indexed = index.search(pattern, root, max_results=5000, verbose=False)
          t_idx = (time.monotonic() - t0) * 1000
      
          t0 = time.monotonic()
          brute = _brute_force_search(pattern, root, skip_dirs, max_results=5000)
          t_brute = (time.monotonic() - t0) * 1000
      
          idx_files = {m["file"] for m in indexed}
          brute_files = {m["file"] for m in brute}
          missed = brute_files - idx_files
      
          print(f"\n{'='*60}")
          print(f"BENCHMARK: '{pattern}'")
          print(f"{'='*60}")
          print(f"  Indexed:  {t_idx:8.1f}ms  ({len(indexed)} matches, {len(idx_files)} files)")
          print(f"  Brute rg: {t_brute:8.1f}ms  ({len(brute)} matches, {len(brute_files)} files)")
          if t_brute > 0:
              print(f"  Speedup:  {t_brute / max(t_idx, 0.1):.1f}x")
          if missed:
              print(f"  ⚠ Missed: {len(missed)} files")
          elif not (idx_files - brute_files):
              print("  ✓ Results match")
      
      
      def format_results(results: dict, root: str, output_json: bool = False) -> str:
          """Format search results for display."""
          if output_json:
              # Make paths relative
              for q, matches in results.items():
                  for m in matches:
                      try:
                          m["file"] = os.path.relpath(m["file"], root)
                      except ValueError:
                          pass
              return json.dumps(results, indent=2)
      
          lines = []
          for query, matches in results.items():
              if len(results) > 1:
                  lines.append(f"\n--- {query} ---")
      
              if not matches:
                  lines.append("No matches found.")
                  continue
      
              lines.append(f"{len(matches)} match{'es' if len(matches) != 1 else ''}")
      
              for m in matches[:30]:
                  try:
                      rel = os.path.relpath(m["file"], root)
                  except ValueError:
                      rel = m["file"]
      
                  if "score" in m:
                      lines.append(f"  {rel}:{m['line']}  [{m['score']:.3f}]  {m.get('name', '')}")
                  else:
                      text = m["text"][:150].rstrip()
                      lines.append(f"  {rel}:{m['line']}: {text}")
      
                  if "context" in m:
                      ctx = m["context"]
                      lines.append(f"    → {ctx['type']} {ctx['name']} "
                                   f"(lines {ctx['start_line']}-{ctx['end_line']})")
      
              if len(matches) > 30:
                  lines.append(f"  ... and {len(matches) - 30} more")
      
          return "\n".join(lines)
      
      
      def run_lsp_query(root: str, symbol: str, op: str, json_out: bool,
                        verbose: bool) -> None:
          """Binding-resolved (pyright) tier for Python symbols, with soft fallback.
      
          On any condition where the binding-resolved answer can't be produced
          (non-Python target, pyright/node absent, client unimportable), emit a
          one-line degradation note and fall back to the regex text path so the user
          still gets results.
          """
          from lsp_refs import LspUnavailable, format_lsp, lsp_query
      
          try:
              result = lsp_query(root, symbol, op=op, verbose=verbose)
          except LspUnavailable as e:
              print(f"[python-lsp unavailable: {e}] falling back to regex text search",
                    file=sys.stderr)
              results = search_regex(root, [symbol], verbose=verbose)
              print(format_results(results, root, json_out))
              return
          print(format_lsp(result, root, json_out))
      
      
      def main():
          parser = argparse.ArgumentParser(description="Unified code search")
          parser.add_argument("source", help="Path, GitHub URL, 'uploads', or 'project'")
          parser.add_argument("queries", nargs="*", help="Search queries")
          parser.add_argument("--regex", action="store_true", help="Force regex mode")
          parser.add_argument("--semantic", action="store_true", help="Force semantic mode")
          parser.add_argument("--expand", action="store_true", help="Expand to full function bodies")
          parser.add_argument("--benchmark", action="store_true", help="Benchmark indexed vs brute-force")
          parser.add_argument("--refs", metavar="SYMBOL", default=None,
                              help="Binding-resolved references for a Python SYMBOL (pyright; "
                                   "excludes same-named unrelated symbols)")
          parser.add_argument("--def", dest="defn", metavar="SYMBOL", default=None,
                              help="Binding-resolved go-to-definition for a Python SYMBOL "
                                   "(follows imports across files)")
          parser.add_argument("--hover", metavar="SYMBOL", default=None,
                              help="Inferred type/signature for a Python SYMBOL (pyright)")
          parser.add_argument("--branch", default="main", help="Git branch for GitHub URLs")
          parser.add_argument("--skip", default=None, help="Comma-separated directories to skip")
          parser.add_argument("--json", action="store_true", help="JSON output")
          parser.add_argument("-v", "--verbose", action="store_true")
      
          args = parser.parse_args()
      
          # Resolve source
          root = resolve(args.source, args.branch)
          if args.verbose:
              print(f"Resolved: {root} ({count_files(root)} files)", file=sys.stderr)
      
          # Binding-resolved tier (Python only, engaged lazily). Mutually exclusive
          # with the text-search queries — these answer a single symbol query.
          lsp_ops = [("references", args.refs), ("definition", args.defn), ("hover", args.hover)]
          active = [(op, sym) for op, sym in lsp_ops if sym]
          if active:
              if len(active) > 1:
                  parser.error("--refs / --def / --hover are mutually exclusive")
              op, symbol = active[0]
              run_lsp_query(root, symbol, op, args.json, args.verbose)
              return
      
          if not args.queries:
              parser.error("no queries given (provide search terms, or use --refs/--def/--hover)")
      
          skip_dirs = None
          if args.skip:
              skip_dirs = set(args.skip.split(","))
      
          # Route queries
          all_results = {}
          for query in args.queries:
              if args.regex:
                  mode = "regex"
              elif args.semantic:
                  mode = "semantic"
              else:
                  mode = detect_mode(query)
      
              if args.verbose:
                  print(f"Query: '{query}' → {mode} mode", file=sys.stderr)
      
          # Batch by mode for efficiency (one index build per mode)
          regex_queries = []
          semantic_queries = []
          for query in args.queries:
              if args.regex:
                  regex_queries.append(query)
              elif args.semantic:
                  semantic_queries.append(query)
              else:
                  mode = detect_mode(query)
                  if mode == "regex":
                      regex_queries.append(query)
                  else:
                      semantic_queries.append(query)
      
          if regex_queries:
              results = search_regex(root, regex_queries, expand=args.expand,
                                     benchmark=args.benchmark, verbose=args.verbose,
                                     skip_dirs=skip_dirs)
              all_results.update(results)
      
          if semantic_queries:
              results = search_semantic(root, semantic_queries, expand=args.expand,
                                        verbose=args.verbose, skip_dirs=skip_dirs)
              all_results.update(results)
      
          if not args.benchmark:
              print(format_results(all_results, root, args.json))
      
      
      if __name__ == "__main__":
          main()
      
    • sparse_ngrams.py 6.1 KB
      """
      Sparse N-gram extraction for fast regex search indexing.
      
      Based on the approach described by Cursor (2026): variable-length n-grams
      selected deterministically via a weight function over character pairs.
      
      Two modes:
      - build_all: Extract ALL valid sparse n-grams (used at index time)
      - build_covering: Extract MINIMAL covering set (used at query time)
      
      A sparse n-gram is a substring where the character-pair weights at both
      boundary positions are strictly greater than all interior weights.
      """
      
      import zlib
      
      
      def weight_crc32(a: int, b: int) -> int:
          """CRC32-based weight for a character pair. Deterministic, uniform."""
          return zlib.crc32(bytes([a, b])) & 0xFFFFFFFF
      
      
      class FrequencyWeights:
          """
          Frequency-based weight function: rare character pairs get HIGH weights,
          common pairs get LOW weights. This produces longer n-grams at rare
          boundaries (more selective posting lists) and shorter n-grams at common
          boundaries (acceptable since they appear everywhere anyway).
          """
      
          def __init__(self):
              self._freq: dict[tuple[int, int], int] = {}
              self._max_freq: int = 1
              self._frozen = False
      
          def train(self, data: bytes):
              """Accumulate character pair frequencies from training data."""
              if self._frozen:
                  raise RuntimeError("Cannot train after freezing")
              for i in range(len(data) - 1):
                  pair = (data[i], data[i + 1])
                  self._freq[pair] = self._freq.get(pair, 0) + 1
      
          def freeze(self):
              """Finalize the frequency table. Converts frequencies to weights."""
              if self._freq:
                  self._max_freq = max(self._freq.values())
              self._frozen = True
      
          def weight(self, a: int, b: int) -> int:
              """
              Weight for a character pair. Higher = rarer.
              Uses inverted frequency: rare pairs get high weights.
              Falls back to CRC32 for unseen pairs (treated as very rare).
              """
              if not self._frozen:
                  raise RuntimeError("Must freeze() before computing weights")
              freq = self._freq.get((a, b), 0)
              if freq == 0:
                  # Unseen pair = very rare = high weight
                  return self._max_freq + weight_crc32(a, b) % (self._max_freq // 2 + 1)
              # Invert: rare = high weight
              return self._max_freq - freq + 1
      
          def save(self) -> bytes:
              """Serialize frequency table."""
              import json
              data = {
                  "freq": {f"{a},{b}": c for (a, b), c in self._freq.items()},
                  "max_freq": self._max_freq,
              }
              return json.dumps(data).encode()
      
          @classmethod
          def load(cls, raw: bytes) -> "FrequencyWeights":
              """Deserialize frequency table."""
              import json
              data = json.loads(raw)
              w = cls()
              w._freq = {
                  (int(k.split(",")[0]), int(k.split(",")[1])): v
                  for k, v in data["freq"].items()
              }
              w._max_freq = data["max_freq"]
              w._frozen = True
              return w
      
      
      def compute_weights(
          text: bytes, weight_fn=weight_crc32
      ) -> list[int]:
          """Compute weights for all consecutive character pairs in text."""
          if len(text) < 2:
              return []
          return [weight_fn(text[i], text[i + 1]) for i in range(len(text) - 1)]
      
      
      def build_all(weights: list[int]) -> list[tuple[int, int]]:
          """
          Extract ALL valid sparse n-grams from a weight sequence.
      
          Uses a monotone stack algorithm (O(n) amortized).
      
          Returns list of (start_pair_pos, end_pair_pos) where each n-gram
          spans characters [start_pair_pos, end_pair_pos + 2) in the original text.
      
          A sparse n-gram from pair position a to pair position b is valid iff
          w[a] > w[k] and w[b] > w[k] for all a < k < b.
          """
          n = len(weights)
          if n == 0:
              return []
          if n == 1:
              return [(0, 0)]
      
          ngrams = []
          # Monotone decreasing stack of pair positions
          stack: list[int] = []
      
          for i in range(n):
              # Pop positions dominated by current weight
              while stack and weights[i] >= weights[stack[-1]]:
                  j = stack.pop()
                  # (j, i) is valid: j and i are both >= weights[j],
                  # and everything between j and i on the stack was already
                  # popped (so had weight < w[j] < w[i])
                  ngrams.append((j, i))
      
              # Adjacent stack entry to current position forms valid n-gram
              if stack:
                  ngrams.append((stack[-1], i))
      
              stack.append(i)
      
          return ngrams
      
      
      def build_covering(weights: list[int]) -> list[tuple[int, int]]:
          """
          Extract the MINIMAL covering set of sparse n-grams.
      
          Used at query time: produces the fewest, longest n-grams needed
          to look up in the index. Any document containing the query text
          must contain all of these n-grams.
      
          Greedy: from each position, jump to the farthest valid endpoint
          (the first position with weight >= current).
          """
          n = len(weights)
          if n == 0:
              return []
          if n == 1:
              return [(0, 0)]
      
          ngrams = []
          i = 0
      
          while i < n:
              # Find the first position j > i where w[j] >= w[i]
              j = i + 1
              while j < n and weights[j] < weights[i]:
                  j += 1
      
              if j >= n:
                  # No higher weight found — take the highest remaining position
                  # as the endpoint (best available boundary)
                  if i < n - 1:
                      best = i + 1
                      for k in range(i + 2, n):
                          if weights[k] > weights[best]:
                              best = k
                      ngrams.append((i, best))
                      i = best
                  else:
                      # At the last position, nothing more to cover
                      break
              else:
                  ngrams.append((i, j))
                  i = j
      
          return ngrams
      
      
      def ngram_text(text: bytes, start: int, end: int) -> bytes:
          """
          Extract the n-gram substring from text given pair positions.
          Pair position p corresponds to characters text[p:p+2].
          N-gram from pair a to pair b spans text[a:b+2].
          """
          return text[start : end + 2]
      
      
      def ngram_hash(text: bytes, start: int, end: int) -> int:
          """Hash an n-gram for use as index key. Uses CRC32 for speed."""
          return zlib.crc32(text[start : end + 2]) & 0xFFFFFFFF
      
  • tests
    • fixture
      • pkg
        • models.py 188 B
          class User:
              def __init__(self, name: str) -> None:
                  self.name = name
          
              def greet(self) -> str:
                  return f"hi {self.name}"
          
          
          def helper(x: int) -> int:
              return x + 1
          
        • other.py 62 B
          def helper() -> str:
              return "unrelated"
          
          
          print(helper())
          
        • service.py 125 B
          from pkg.models import User, helper
          
          
          def make_user(name: str) -> User:
              u = User(name)
              print(helper(3))
              return u
          
        • __init__.py 0 B
    • test_lsp_refs.py 5.5 KB
      """Tests for the binding-resolved (python-lsp) tier of searching-codebases.
      
      Round-trips against a small multi-file fixture driving a real
      ``pyright-langserver``, and exercises the mandatory soft-fallback contract.
      
      Run: python -m pytest searching-codebases/tests/test_lsp_refs.py -v
      Or:  python searching-codebases/tests/test_lsp_refs.py   (standalone)
      
      The semantic tests require pyright (and system node); they are skipped if the
      bootstrap can't make pyright available. The fallback tests have no such
      dependency — they assert the degradation path.
      """
      
      import subprocess
      import sys
      from pathlib import Path
      
      import pytest
      
      SCRIPTS = Path(__file__).resolve().parent.parent / "scripts"
      sys.path.insert(0, str(SCRIPTS))
      
      import lsp_refs
      from lsp_refs import (
          LspUnavailable,
          definition_sites,
          lsp_query,
          occurrence_sites,
      )
      
      FIXTURE = str(Path(__file__).parent / "fixture")
      
      
      def _pyright_available() -> bool:
          try:
              lsp = lsp_refs._import_lsp_client()
              lsp.ensure_pyright(install=True)
              return True
          except Exception:
              return False
      
      
      pyright = pytest.mark.skipif(not _pyright_available(),
                                   reason="pyright/node unavailable")
      
      
      def _pyright_proc_count() -> int:
          out = subprocess.run(["pgrep", "-f", "pyright-langserver"],
                               capture_output=True, text=True).stdout
          return len([ln for ln in out.splitlines() if ln.strip()])
      
      
      # ── symbol → position resolution (no server) ─────────────────────────────────
      
      def test_definition_sites_finds_both_helpers():
          sites = definition_sites(FIXTURE, "helper")
          by_file = {s.file: s for s in sites}
          assert "pkg/models.py" in by_file and "pkg/other.py" in by_file
          # Column lands on the symbol token, not the `def` keyword (1-based).
          assert by_file["pkg/models.py"].col == 5  # "def helper" -> h at col 5
          assert by_file["pkg/models.py"].line == 9
      
      
      def test_occurrence_sites_are_a_superset():
          occ = occurrence_sites(FIXTURE, "helper")
          files = {s.file for s in occ}
          # Text matching catches all three files (including the unrelated other.py).
          assert {"pkg/models.py", "pkg/other.py", "pkg/service.py"} <= files
      
      
      # ── binding-resolved queries (the win over regex) ────────────────────────────
      
      @pyright
      def test_references_exclude_unrelated_same_named_symbol():
          result = lsp_query(FIXTURE, "helper", op="references")
          groups = {g["anchor"].file: g for g in result["results"]}
          # The models.py binding's references include its def and the use in
          # service.py, but EXCLUDE the unrelated helper in other.py.
          models = {Path(loc.path).name for loc in groups["pkg/models.py"]["locations"]}
          models_files = {str(Path(loc.path).relative_to(Path(FIXTURE)))
                          for loc in groups["pkg/models.py"]["locations"]}
          assert "pkg/models.py" in models_files
          assert "pkg/service.py" in models_files
          assert "pkg/other.py" not in models_files, (
              "binding-resolved references must exclude the same-named unrelated symbol"
          )
          assert "models.py" in models  # sanity on the name extraction
      
      
      @pyright
      def test_definition_follows_import_across_files():
          result = lsp_query(FIXTURE, "User", op="definition")
          locs = result["results"][0]["locations"]
          files = {str(Path(loc.path).relative_to(Path(FIXTURE))) for loc in locs}
          # Anchored at every occurrence (incl. the import + use in service.py),
          # definition resolves across the import to the class in models.py.
          assert "pkg/models.py" in files
      
      
      @pyright
      def test_hover_returns_inferred_signature():
          result = lsp_query(FIXTURE, "helper", op="hover")
          hovers = " ".join(g["hover"] or "" for g in result["results"])
          assert "helper" in hovers and "int" in hovers
      
      
      @pyright
      def test_indexing_wait_makes_queries_deterministic():
          # Without wait_for_index this would flake to empty. Repeat must be stable.
          for _ in range(2):
              result = lsp_query(FIXTURE, "User", op="references")
              total = sum(len(g["locations"]) for g in result["results"])
              assert total >= 3
      
      
      @pyright
      def test_no_orphaned_subprocess():
          before = _pyright_proc_count()
          lsp_query(FIXTURE, "User", op="references")
          # Context manager in lsp_query must reap the server.
          assert _pyright_proc_count() == before, "pyright-langserver leaked"
      
      
      # ── soft-fallback contract (no pyright dependency) ───────────────────────────
      
      def test_missing_symbol_raises_lsp_unavailable():
          with pytest.raises(LspUnavailable):
              lsp_query(FIXTURE, "NoSuchSymbolAnywhere", op="references")
      
      
      def test_bootstrap_failure_raises_lsp_unavailable(monkeypatch):
          # Simulate node/pyright absence: ensure_pyright fails loud, the tier must
          # degrade (raise LspUnavailable) rather than hang.
          lsp = lsp_refs._import_lsp_client()
      
          def boom(*a, **k):
              raise lsp.BootstrapError("pyright requires system 'node' but none found")
      
          monkeypatch.setattr(lsp, "ensure_pyright", boom)
          with pytest.raises(LspUnavailable) as exc:
              lsp_query(FIXTURE, "helper", op="references")
          assert "node" in str(exc.value).lower()
      
      
      # ── standalone runner ────────────────────────────────────────────────────────
      
      if __name__ == "__main__":
          raise SystemExit(pytest.main([__file__, "-v"]))
      
  • CHANGELOG.md 3.6 KB
    # searching-codebases - Changelog
    
    All notable changes to the `searching-codebases` skill are documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
    
    ## [2.5.0] - 2026-08-25
    
    ### Other
    
    - top skills: separate by omission, and correct the guidance that said otherwise (#777)
    - Deprecate mapping-codebases; adopt ruff 0.16.0 baseline (#747)
    
    ## [2.2.0] - 2026-07-04
    
    ### Other
    
    - searching-codebases 2.2.0: narrow to binding-resolved edge case (#722)
    
    ## [2.2.0] - 2026-07-04
    
    ### Changed
    
    - Repositioned the skill as edge-case: the binding-resolved Python tier
      (`--refs`/`--def`/`--hover`) is the one recommended use. Frontmatter
      description and When-to-Use rewritten accordingly.
    - Empirical basis: head-to-head localization eval on 7 scikit-learn issues
      with merged fix-PRs (replicating the file-discovery metric of
      arXiv:2602.11988). Indexed-regex and TF-IDF semantic tiers tied or lost
      vs naive rg on every instance at 4-60x wall-clock; semantic never won,
      including on identifier-poor issues (~0.3% of merged-PR traffic).
    - Tiers remain functional; they carry the burden of proof.
    
    ## [2.1.1] - 2026-06-15
    
    ### Fixed
    
    - open all symbol-bearing files for the binding-resolved tier (#698)
    - auto-install tree-sitter for the binding-resolved tier
    
    ## [2.1.0] - 2026-06-15
    
    ### Added
    
    - add treesit.py CLI, fix cross-process cache loss, fix Symbol dict bug (#536)
    
    ### Other
    
    - searching-codebases: binding-resolved Python tier via python-lsp (#695) (#697)
    - Remove _MAP.md files, direct agents to tree-sitting for code navigation (#545)
    - Add missing READMEs for searching-codebases, featuring, tree-sitting (#521)
    
    ## [2.1.0] - 2026-06-15
    
    ### Added
    
    - Binding-resolved reference/definition tier for Python via the `python-lsp` skill (pyright): `--refs`, `--def`, `--hover` (#695)
      - `--refs SYMBOL` excludes same-named-but-unrelated symbols and follows imports — the precision the regex cross-reference tier lacks
      - Engaged lazily (pyright index cost paid only on these flags); Python-only with a mandatory soft fallback to the regex text path (non-`.py`, or pyright/node absent) that emits a one-line degradation note
      - `scripts/lsp_refs.py`: tree-sitting symbol→position resolution, 1-based↔0-based conversion at the LSP boundary, lifecycle-safe queries (index-wait + subprocess reaping)
      - `tests/test_lsp_refs.py`: binding-resolution, import-following, indexing-wait determinism, no orphaned subprocess, and soft-fallback coverage
    
    ## [2.0.0] - 2026-03-31
    
    ### Other
    
    - exploring v1.0.0, searching v2.0.0: tree-sitting replaces mapping-codebases
    - Regenerate _MAP.md files after @lat: backlink insertion (#504)
    - Lattice v2: bidirectional source-anchored knowledge graph (#503)
    
    ## [1.0.0] - 2026-03-25
    
    ### Added
    
    - add mapping-features skill for behavioral web app documentation (#432)
    
    ### Other
    
    - Remove searching-codebases/scripts/_MAP.md (absorbed into search.py)
    - Remove searching-codebases/scripts/flowing.py (absorbed into search.py)
    - Remove searching-codebases/scripts/pipeline.py (absorbed into search.py)
    - searching-codebases/scripts/sparse_ngrams.py
    - searching-codebases/scripts/ngram_index.py
    - searching-codebases/scripts/context.py
    - searching-codebases/scripts/resolve.py
    - searching-codebases/scripts/search.py
    - searching-codebases/SKILL.md
    - Implement #414, #415, #416: perch findings digests, boot flight awareness, inline links
    - Add searching-codebases skill
    ## 2.3.0 — 2026-07-26
    
    - Dropped `_MAP.md` as a corpus source (`_extract_map_entries` removed). Code and
      markdown extractors are unaffected; `_MAP.md` files no longer exist to index.
  • README.md 1.9 KB
    # searching-codebases
    
    Find code in any codebase by regex pattern or natural language concept. Auto-routes between n-gram indexed regex search (2-20x faster than ripgrep) and TF-IDF semantic search. Expands results to full functions via tree-sitting AST data. For Python sources, a binding-resolved reference/definition tier (pyright, via python-lsp) finds real callers and definitions without same-name false positives.
    
    ## Features
    
    - **Dual search modes** — regex (pattern/identifier) and semantic (natural language concepts), auto-routed per query
    - **N-gram indexed regex** — sparse inverted index narrows candidate files by 90-99% before ripgrep verification
    - **TF-IDF semantic search** — cosine similarity ranking over code chunks (functions, classes)
    - **Binding-resolved Python tier** — `--refs`/`--def`/`--hover` via pyright: true find-all-callers and go-to-definition that exclude same-named, unrelated symbols and follow imports. Engaged lazily; degrades to the regex text path when pyright/node is unavailable
    - **AST context expansion** — `--expand` returns complete function/class bodies instead of line fragments
    - **Flexible sources** — accepts GitHub URLs, local directories, uploaded files/archives, or project knowledge
    - **Mixed queries** — multiple queries with different modes in a single invocation; indexes built once per mode
    
    ## Dependencies
    
    - **ripgrep** — required for regex verification
    - **tree-sitting** — auto-installs the bare `tree-sitter` package when needed: for `--expand` context and for the symbol→position resolution that seeds the binding-resolved tier (grammars ship bundled). Regex and semantic search work without it
    - **scikit-learn** — required for semantic mode (auto-installs)
    - **python-lsp** — provides the binding-resolved tier (`--refs`/`--def`/`--hover`); self-bootstraps pyright on first use and needs system `node` (v18+). Without it those flags degrade to the regex text path
    
  • SKILL.md 8.1 KB
    ---
    name: searching-codebases
    description: >-
      Binding-resolved Python symbol queries via pyright — every true caller
      (--refs), the real definition (--def), or an inferred signature (--hover)
      of a .py symbol, excluding the same-named false positives text search
      cannot tell apart. Use when a task needs ALL callers or users of a named
      Python symbol and grep would over-match. Python only, and only for the
      caller/definition question: measured 2026-07-04 on real issue-localization
      tasks, plain ripgrep tied or beat the semantic and indexed-regex tiers at
      4-60x less wall-clock, so those tiers survive below without being a
      default.
    metadata:
      version: 2.5.0
    ---
    
    # Searching Codebases
    
    Find code in any codebase by pattern or concept. One entry point, two
    search strategies, automatic routing.
    
    ## When NOT to use this skill
    
    Python callers and definitions only. Every other code question is cheaper
    elsewhere, and the measurement below is why this skill is edge-case-only.
    
    | Situation | Use |
    |---|---|
    | What symbols does this file contain? | tree-sitting |
    | Where is X defined, so I can read it? | tree-sitting |
    | Any structure question, any language | tree-sitting |
    | A literal string or regex | plain ripgrep |
    | Which files are most about a concept? | bm25 |
    | First look at an unfamiliar repo | exploring-codebases |
    
    Non-Python code has no pyright binding resolution to offer, so there is no
    version of this skill that applies to it.
    
    ## Prerequisites
    
    ```bash
    uv tool install ripgrep
    ```
    
    tree-sitting installs automatically when needed — for `--expand` context
    expansion and for the binding-resolved `--refs`/`--def`/`--hover` tier, which
    uses it to resolve symbol positions. Only the bare `tree-sitter` package is
    fetched; the language grammars ship bundled.
    
    ## Primary Command
    
    ```bash
    SKILL_DIR=/mnt/skills/user/searching-codebases
    
    python3 $SKILL_DIR/scripts/search.py SOURCE "query1" ["query2" ...] [OPTIONS]
    ```
    
    SOURCE is any of:
    - Local directory path
    - GitHub URL (downloads tarball automatically)
    - `uploads` (uses `/mnt/user-data/uploads/`)
    - `project` (uses `/mnt/project/`)
    - Path to a `.zip` or `.tar.gz` archive
    
    ## Search Modes
    
    **Regex mode** (patterns, identifiers, literal text):
    ```bash
    python3 $SKILL_DIR/scripts/search.py ./repo "def handle_error"
    python3 $SKILL_DIR/scripts/search.py ./repo "class.*Exception" --regex
    python3 $SKILL_DIR/scripts/search.py ./repo "TODO|FIXME|HACK"
    ```
    
    **Semantic mode** (concepts, natural language):
    ```bash
    python3 $SKILL_DIR/scripts/search.py ./repo "retry logic with backoff" --semantic
    python3 $SKILL_DIR/scripts/search.py ./repo "authentication flow"
    python3 $SKILL_DIR/scripts/search.py ./repo "error handling strategy"
    ```
    
    Auto-detection: short queries and code-like tokens → regex. Multi-word
    natural language → semantic. Override with `--regex` or `--semantic`.
    
    **Binding-resolved mode** (Python only — pyright via the `python-lsp` skill):
    ```bash
    python3 $SKILL_DIR/scripts/search.py ./repo --refs SYMBOL    # find all real uses
    python3 $SKILL_DIR/scripts/search.py ./repo --def SYMBOL     # go-to-definition
    python3 $SKILL_DIR/scripts/search.py ./repo --hover SYMBOL   # inferred type/signature
    ```
    
    Regex mode matches *text*, so a cross-reference for a function false-positives
    on shadowed and same-named-but-unrelated symbols. `--refs` is **binding-resolved**:
    pyright excludes the unrelated same-named symbol and follows imports. Use it when
    you need a true "find all callers/users" for a `.py` symbol, not a text grep.
    
    The tier is engaged **lazily** — pyright's index cost is paid only when you ask
    for `--refs`/`--def`/`--hover`, never on ordinary searches. It is **Python-only**;
    for non-`.py` sources, or when pyright/node is unavailable, it prints a one-line
    degradation note and falls back to the regex text path. Each takes a single bare
    symbol name and is mutually exclusive with the other two and with text queries.
    
    ## Options
    
    - `--regex` / `--semantic`: Force search mode
    - `--refs SYMBOL` / `--def SYMBOL` / `--hover SYMBOL`: Binding-resolved Python
      queries via pyright (see Binding-resolved mode above)
    - `--expand`: Return full function bodies via tree-sitting AST context
    - `--benchmark`: Compare indexed regex vs brute-force ripgrep
    - `--branch NAME`: Git branch for GitHub URLs (default: main)
    - `--skip DIRS`: Comma-separated directories to skip
    - `--json`: Machine-readable output
    - `-v`: Show index stats and query routing decisions
    
    ## How It Works
    
    **Regex search** builds a sparse n-gram inverted index over all files.
    Queries are decomposed into literal fragments, looked up in the index
    to identify candidate files (typically 90-99% reduction), then verified
    with ripgrep. Frequency-weighted n-grams make rare character sequences
    more selective.
    
    **Semantic search** builds a TF-IDF index over code chunks (functions,
    classes, structural entries). Queries are ranked by cosine similarity.
    
    **Context expansion** (`--expand`) uses tree-sitting's AST cache to
    identify function/class boundaries, returning complete structural units
    rather than line fragments. On first use, tree-sitting scans the repo
    (~700ms for 250 files); subsequent expansions are sub-millisecond.
    
    **Small codebases** (< 20 files) skip indexing entirely — direct ripgrep is
    faster when there's nothing to narrow.
    
    ## Mixed Queries
    
    Multiple queries can use different modes in a single invocation. Each query
    is auto-routed independently, and indexes are built once per mode:
    
    ```bash
    python3 $SKILL_DIR/scripts/search.py ./repo \
      "class.*Error" \
      "error recovery strategy" \
      "def retry"
    ```
    
    ## Dependencies
    
    - **tree-sitting**: Provides AST context expansion for `--expand` *and* the
      symbol→position resolution that seeds the binding-resolved tier
      (`--refs`/`--def`/`--hover`). Auto-installs the bare `tree-sitter` package
      when either is used (grammars are bundled). Regex and semantic search work
      without it.
    - **ripgrep**: Required for regex verification. Install via `uv tool install ripgrep`.
    - **scikit-learn**: Required for semantic mode. Installs automatically.
    - **python-lsp**: Provides the binding-resolved tier (`--refs`/`--def`/`--hover`).
      Self-bootstraps pyright on first use and requires system `node` (v18+). Not
      required — without it those flags degrade to the regex text path.
    
    ## When to Use — narrow, by design
    
    The ONE recommended use: **binding-resolved Python symbol queries**.
    
    - "find all callers of `X`" / "where is `X` really defined" for a `.py`
      symbol, when same-named-but-unrelated symbols would pollute a text grep.
      Empirical basis: `rg get` on psf/requests returned 232 hits, 224 of them
      false; `--refs get` excluded all 224 (2026-06-15).
    
    ## When NOT to Use — which is most of the time
    
    Everything else. Measured head-to-head on real issue-localization tasks
    (7 scikit-learn issues with merged fix-PRs, gold = PR diff files,
    2026-07-04, replicating the file-discovery metric of arXiv:2602.11988):
    
    - **Literal tokens / identifiers**: naive `rg -l` tied or beat the indexed
      tier on recall@10 in every instance, at 0.4s vs 25s.
    - **Concept / natural-language search**: the TF-IDF semantic tier never
      beat identifier grep — not even on identifier-poor issues, which are
      themselves rare (~0.3% of merged-PR traffic in the sample).
    - **First encounter / "what is this repo"**: use exploring-codebases.
    - **Repos under ~20 files**: read them.
    
    The self-test before invoking: would plain `rg` return the same answer?
    If yes, use rg. The indexed-regex and semantic tiers are retained for
    completeness and for corpora where they may yet earn their cost (very
    large repos, non-code document collections), but they carry the burden
    of proof.
    
    ## Files
    
    - `scripts/search.py` — Entry point, query routing, output formatting
    - `scripts/resolve.py` — Input source resolution (GitHub, uploads, archives)
    - `scripts/context.py` — tree-sitting-based AST context expansion
    - `scripts/ngram_index.py` — Sparse n-gram inverted index, regex decomposition
    - `scripts/sparse_ngrams.py` — Core n-gram algorithms, frequency weights
    - `scripts/code_rag.py` — TF-IDF semantic search over code chunks
    - `scripts/lsp_refs.py` — Binding-resolved Python tier: symbol→position
      resolution (tree-sitting), pyright queries (python-lsp), soft fallback
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related