Claude Skill

semantic-grep

In-process semantic search over text files or in-memory strings, using Gemini embeddings via the CF AI Gateway. Use when user wants fuzzy/conceptual search where exact-keyword grep would miss — "sessions discussing regulatory constraints", "code about retry logic", "notes mention

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download oaustegard-claude-skills-plugins_code-intelligence_skills_semantic-grep-e39c726.zip · 10 KB
Part of oaustegard/claude-skills — 39 skills

Install

skills CLI npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/code-intelligence/skills/semantic-grep
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
Git git clone https://github.com/oaustegard/claude-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.

README

semantic-grep

Semantic search over text files using Gemini embeddings. In-process Python, not a subprocess CLI. See SKILL.md for usage.

Skill manifest

Semantic Grep

jina-grep-style semantic search, done in-process via Python rather than as an external CLI. Embeds query + corpus chunks with gemini-embedding-2, ranks by cosine similarity, returns grep-format output.

When Semantic Search Helps

The core trade-off (lifted from jina-grep-cli's own docs and validated in testing):

Task Tool
Known exact string, filename, or regex grep / rg / searching-codebases
"What files discuss concept X" when X may not appear verbatim semantic-grep
Hybrid: prefilter with grep, rerank by concept grep → rerank_candidates()

Regression test result (workshop session corpus, 135 docs):

  • "handling regulatory constraints" → top hit "Engineering AI Systems Under Sovereignty Constraints" (0.67). ✓
  • "sessions about GEPA" → top hit "Gemma, DeepMind's Family of Open Models" (0.69). ✗ — false positive on phonetic neighbor. GEPA is mentioned verbatim in one session description; grep would find it correctly.

Rule: when the user query reads like a named entity or keyword, try grep first. Only reach for semantic-grep when paraphrase/concept matching is actually needed.

Setup

Credentials via proxy.env (Cloudflare AI Gateway w/ BYOK — same pattern as invoking-gemini):

CF_ACCOUNT_ID=...
CF_GATEWAY_ID=...
CF_API_TOKEN=...

Direct-API fallback: GOOGLE_API_KEY or GEMINI_API_KEY env var. No dependencies beyond requests + numpy.

Quick Start

import sys
sys.path.insert(0, '/mnt/skills/user/semantic-grep/scripts')
from semantic_grep import semantic_grep, format_grep

# Directory of .txt files
results = semantic_grep("error handling under load", "/path/to/notes",
                        top_k=5, granularity="paragraph")
print(format_grep(results))
# notes/incidents.txt:42:  When the queue depth exceeds... [0.71]
# notes/postmortem.txt:8:  Under sustained traffic we saw... [0.68]

Core API

semantic_grep(query, corpus, *, top_k=10, threshold=None, ...)

Main search function.

  • query (str) — the search query (embedded with RETRIEVAL_QUERY task type)
  • corpus (str | Path | list[Chunk]) — a file, directory, or pre-chunked list
  • top_k (int | None) — max results; None = all above threshold
  • threshold (float | None) — cosine similarity cutoff; None = no filter (top_k only)
  • granularity ("paragraph" | "line") — how to chunk files (default paragraph)
  • include (str) — filename-glob filter when corpus is a directory (default "*.txt"). Matches against Path.name only, not the full path — "*.md" works, "docs/*.md" does not.
  • model (str) — default "gemini-embedding-2". gemini-embedding-001 is retired (text-only) and warns if passed explicitly.
  • dim (int) — 128 / 768 / 1536 / 3072 (default 768; MRL-truncated + renormalized)
  • task ("text" | "code") — selects text vs code task types

Returns list[Match] where Match has path, line, text, score.

load_corpus(path, *, include="*.txt", granularity="paragraph") -> list[Chunk]

Load and chunk a file or directory without embedding. Useful for inspecting what gets embedded before paying for the API call.

embed_batch(texts, task_type, *, model, dim, group_size=100) -> np.ndarray

Lower-level: embed a list of strings directly via :batchEmbedContents. Returns (N, dim) float32 array, rows normalized when dim < 3072.

format_grep(matches, *, max_text_chars=200, show_score=True) -> str

Format matches as grep output: path:line: snippet [score].

Pipe-mode Rerank Pattern

The highest-leverage use isn't naive full-corpus semantic search — it's hybrid retrieval: fast coarse filter → semantic rerank.

import subprocess
from semantic_grep import Chunk, semantic_grep, format_grep

# Stage 1: fast exact/regex prefilter with rg
result = subprocess.run(
    ["rg", "-n", "--no-heading", "error|fail|timeout", "logs/"],
    capture_output=True, text=True,
)

# Parse `path:line:text` into Chunks
chunks = []
for raw in result.stdout.splitlines():
    path, line, text = raw.split(":", 2)
    chunks.append(Chunk(path=path, line=int(line), text=text))

# Stage 2: semantic rerank on the prefiltered subset
ranked = semantic_grep("intermittent queue saturation during peak traffic",
                       chunks, top_k=10)
print(format_grep(ranked))

This is how you scale past the "embed the whole corpus every call" limit without needing a vector DB. The exact-match stage cheaply cuts millions of lines to thousands; semantic reranks those.

Task Types (Gemini)

  • text mode (default): query → RETRIEVAL_QUERY, docs → RETRIEVAL_DOCUMENT. Asymmetric — documented to outperform symmetric encoding for retrieval.
  • code mode: query → CODE_RETRIEVAL_QUERY, docs → RETRIEVAL_DOCUMENT. Use when searching code with natural-language queries.

Use SEMANTIC_SIMILARITY (symmetric) only if you're doing pairwise sim, not retrieval. This module doesn't expose that path yet.

Model Notes

gemini-embedding-2 (GA since 2026-04-22) — general-purpose and multimodal. Verified 2026-07-21 via the CF gateway: text, image and audio all embed to the same space at the requested dim, L2-normalized. The retired gemini-embedding-001 was text-only and rejected non-text input with HTTP 400:

  • 2,048 input token limit per text. Longer texts are truncated at ~8K chars (approximation).
  • Matryoshka (MRL) — 3072 native dims, safely truncatable to 1536/768/256/128.
  • 3072 is auto-normalized; lower dims need client-side renorm (handled here).
  • Pricing: $0.15 / 1M input tokens. 135 medium paragraphs ≈ 15K tokens ≈ $0.002 per query.

gemini-embedding-2-preview (March 2026) is multimodal and currently top of MTEB. Set model="gemini-embedding-2-preview" to opt in once the preview stabilizes.

Limitations

  • No persistent index. Every call re-embeds the corpus. Fine for <~1K chunks; prohibitive for real knowledge bases. Phase 2: cache embeddings by content hash.
  • Token budget is approximated by char count (×1.5). Conservative for mixed-script text; over-truncates English slightly. Real tokenizer would use the Gemini tokenizer endpoint but costs an extra call per embed.
  • Batch bulk-failure diagnostic. If one text in a group of 100 overflows or is rejected by safety filters, the whole batch fails and the 99 good ones are lost. No per-index fallback yet.
  • No memory ceiling on corpus size. semantic_grep pre-allocates (N, dim) float32; 1M chunks at dim=768 ≈ 3GB. Caller is responsible for sane chunk counts. load_corpus also follows symlinks via rglob — fine in a trusted single-user container, not for untrusted paths.
  • Sequential batch groups. group_size=100 per HTTP call; groups run serially. For >1K chunks, add asyncio — not needed yet.
  • No CLI shim. Called as a Python module, not a subprocess. Per design: "within an LLM rather than calling out to one."
  • Embedding function lives here, not in invoking-gemini. Should be factored up when invoking-gemini adds embedding support. Tracked as followup.

Related Skills

  • invoking-gemini — sibling; handles Gemini text + image generation through the same CF gateway. Shares credential pattern.
  • searching-codebases — regex/AST search. Use first when the query is a known pattern.
  • extracting-keywords — YAKE keyword extraction; orthogonal, but pairs well for building query terms from a long prompt.
  • exploring-codebases — for understanding repo structure. Semantic-grep doesn't replace AST-based navigation.

Attribution

Conceptually inspired by jina-grep-cli — we kept the retrieval shape (grep-compatible output, asymmetric query/doc embeddings, threshold + top-k) but swapped the MLX/Apple-Silicon backend for a portable Gemini API call. The original's pipe-mode rerank pattern is the most generalizable idea it contributes and is preserved here.

Files (claude-skills)
  • scripts
    • semantic_grep.py 14 KB
      """
      semantic_grep: a jina-grep reproduction using Gemini embeddings.
      
      In-process semantic search over text files or in-memory strings.
      Routes through Cloudflare AI Gateway (proxy.env) when available.
      
      Model: gemini-embedding-2 (GA 2026-04-22; general-purpose AND multimodal, MRL
             truncation supported). Supersedes gemini-embedding-001, which is text-only.
             Verified 2026-07-21 through the CF gateway: text, image (image/png) and
             audio (audio/wav) all embed to the same space at the requested dim,
             L2-normalized. 001 rejects non-text with HTTP 400.
      Task types used: RETRIEVAL_QUERY (query) + RETRIEVAL_DOCUMENT (corpus) — asymmetric.
      
      Design notes:
      - Serverless-mode equivalent: every call re-embeds the corpus. No persistent index.
      - Caller-facing API returns structured dicts. Grep-format is a formatter, not the core.
      - Flag surface mirrors jina-grep where it makes sense for a Python API (top_k, threshold,
        granularity, include). -r recursion is implicit when path is a directory.
      """
      
      from __future__ import annotations
      
      import fnmatch
      import os
      import time
      from dataclasses import dataclass
      from pathlib import Path
      from typing import Literal
      
      import numpy as np
      import requests
      
      # ---------------------------------------------------------------------------
      # Credentials (reuse invoking-gemini's proxy pattern)
      # ---------------------------------------------------------------------------
      
      _CF_GATEWAY_BASE = "https://gateway.ai.cloudflare.com/v1"
      _DIRECT_BASE = "https://generativelanguage.googleapis.com/v1beta"
      _DEFAULT_MODEL = "gemini-embedding-2"
      
      # Retired 2026-07-21. gemini-embedding-001 is text-only and superseded by
      # gemini-embedding-2 (general-purpose AND multimodal, same request shape and
      # dims). Still callable if explicitly requested, but warns — nothing in this
      # skill persists vectors, so there is no index to migrate.
      _RETIRED_MODELS = {
          "gemini-embedding-001": "gemini-embedding-2",
      }
      
      
      def _check_retired(model: str) -> None:
          """Warn once if a caller explicitly asks for a retired embedding model."""
          if model in _RETIRED_MODELS:
              import warnings
              warnings.warn(
                  f"{model} is retired; use {_RETIRED_MODELS[model]} "
                  f"(general-purpose and multimodal, same dims). Vectors from "
                  f"different encoders are not comparable — do not mix them.",
                  DeprecationWarning, stacklevel=3,
              )
      _MAX_INPUT_TOKENS = 2048  # conservative; documented for embedding-001, unverified for -2
      # Conservative char-per-token ratio. English prose is ~4 chars/token,
      # but CJK is closer to 1 char/token and emoji can be <1. Picking 1.5 as a safe
      # lower bound means we under-fill for English (some wasted context) in exchange
      # for not silently overflowing the API limit for multi-script text.
      _APPROX_CHARS_PER_TOKEN = 1.5
      
      
      def _load_env(path: Path) -> dict:
          if not path.exists():
              return {}
          out = {}
          for line in path.read_text().splitlines():
              line = line.strip()
              if not line or line.startswith("#") or "=" not in line:
                  continue
              if line.startswith("export "):
                  line = line[7:].lstrip()
              k, v = line.split("=", 1)
              out[k.strip()] = v.strip().strip("'\"")
          return out
      
      
      def _embed_url(model: str, *, batch: bool = False) -> tuple[str, dict]:
          """Return (url, extra_headers) for embedContent / batchEmbedContents call."""
          endpoint = "batchEmbedContents" if batch else "embedContent"
          proxy = _load_env(Path("/mnt/project/proxy.env"))
          cf_account = proxy.get("CF_ACCOUNT_ID") or os.environ.get("CF_ACCOUNT_ID")
          cf_gateway = proxy.get("CF_GATEWAY_ID") or os.environ.get("CF_GATEWAY_ID")
          cf_token = proxy.get("CF_API_TOKEN") or os.environ.get("CF_API_TOKEN")
          if cf_account and cf_gateway and cf_token:
              url = (
                  f"{_CF_GATEWAY_BASE}/{cf_account}/{cf_gateway}"
                  f"/google-ai-studio/v1beta/models/{model}:{endpoint}"
              )
              headers = {"cf-aig-authorization": f"Bearer {cf_token}"}
              return url, headers
          key = os.environ.get("GOOGLE_API_KEY") or os.environ.get("GEMINI_API_KEY")
          if not key:
              raise RuntimeError("No credentials: missing proxy.env / CF_* env vars / GOOGLE_API_KEY")
          url = f"{_DIRECT_BASE}/models/{model}:{endpoint}"
          return url, {"x-goog-api-key": key}
      
      
      # ---------------------------------------------------------------------------
      # Embedding
      # ---------------------------------------------------------------------------
      
      TaskType = Literal[
          "RETRIEVAL_QUERY", "RETRIEVAL_DOCUMENT", "SEMANTIC_SIMILARITY",
          "CLASSIFICATION", "CLUSTERING", "QUESTION_ANSWERING", "FACT_VERIFICATION",
          "CODE_RETRIEVAL_QUERY",
      ]
      
      
      def _embed_one(text: str, task_type: TaskType, *, model: str, dim: int,
                     timeout: float = 30.0, retries: int = 2) -> np.ndarray:
          """Embed a single string. Normalizes output if dim < 3072."""
          _check_retired(model)
          url, headers = _embed_url(model)
          headers = {**headers, "Content-Type": "application/json"}
          body = {
              "content": {"parts": [{"text": text}]},
              "taskType": task_type,
              "outputDimensionality": dim,
          }
          last_err = None
          for attempt in range(retries + 1):
              try:
                  r = requests.post(url, headers=headers, json=body, timeout=timeout)
                  if r.status_code == 200:
                      data = r.json()
                      vals = data["embedding"]["values"]
                      arr = np.asarray(vals, dtype=np.float32)
                      if dim < 3072:
                          n = np.linalg.norm(arr)
                          if n > 0:
                              arr = arr / n
                      return arr
                  # Retry on 429/5xx
                  if r.status_code in (429, 500, 502, 503, 504) and attempt < retries:
                      time.sleep(1.5 ** attempt)
                      continue
                  raise RuntimeError(f"Embed failed {r.status_code}: {r.text[:300]}")
              except requests.exceptions.RequestException as e:
                  last_err = e
                  if attempt < retries:
                      time.sleep(1.5 ** attempt)
                      continue
                  raise
          raise RuntimeError(f"Embed failed after retries: {last_err}")
      
      
      def _truncate_to_token_budget(text: str) -> str:
          """Approximate 2048-token cap by char count. Gemini will reject overflow otherwise.
      
          Uses a conservative chars-per-token ratio (1.5) so mixed-script / CJK text
          is not silently truncated-too-long.
          """
          limit = int(_MAX_INPUT_TOKENS * _APPROX_CHARS_PER_TOKEN)  # ~3K chars
          return text[:limit] if len(text) > limit else text
      
      
      def embed_batch(texts: list[str], task_type: TaskType, *,
                      model: str = _DEFAULT_MODEL, dim: int = 768,
                      group_size: int = 100, timeout: float = 90.0,
                      retries: int = 3) -> np.ndarray:
          """Embed N strings via :batchEmbedContents (single HTTP call per group of `group_size`).
      
          Returns (N, dim) array. Normalizes output rows when dim < 3072.
          """
          if not texts:
              return np.zeros((0, dim), dtype=np.float32)
      
          url, base_headers = _embed_url(model, batch=True)
          headers = {**base_headers, "Content-Type": "application/json"}
          out = np.zeros((len(texts), dim), dtype=np.float32)
      
          for start in range(0, len(texts), group_size):
              group = texts[start:start + group_size]
              body = {
                  "requests": [
                      {
                          "model": f"models/{model}",
                          "content": {"parts": [{"text": _truncate_to_token_budget(t)}]},
                          "taskType": task_type,
                          "outputDimensionality": dim,
                      }
                      for t in group
                  ]
              }
      
              for attempt in range(retries + 1):
                  try:
                      r = requests.post(url, headers=headers, json=body, timeout=timeout)
                      if r.status_code == 200:
                          data = r.json()
                          embs = data.get("embeddings", [])
                          if len(embs) != len(group):
                              raise RuntimeError(
                                  f"Batch size mismatch: sent {len(group)}, got {len(embs)}"
                              )
                          for i, e in enumerate(embs):
                              arr = np.asarray(e["values"], dtype=np.float32)
                              if dim < 3072:
                                  n = np.linalg.norm(arr)
                                  if n > 0:
                                      arr = arr / n
                              out[start + i] = arr
                          break  # success, next group
                      if r.status_code in (429, 500, 502, 503, 504) and attempt < retries:
                          time.sleep(1.5 ** attempt)
                          continue
                      raise RuntimeError(f"Batch embed failed {r.status_code}: {r.text[:400]}")
                  except requests.exceptions.RequestException as e:
                      if attempt < retries:
                          time.sleep(1.5 ** attempt)
                          continue
                      raise RuntimeError(f"Batch embed network error: {e}") from e
      
          return out
      
      
      # ---------------------------------------------------------------------------
      # Chunking
      # ---------------------------------------------------------------------------
      
      Granularity = Literal["line", "paragraph"]
      
      
      @dataclass
      class Chunk:
          path: str       # source identifier (filepath or logical id)
          line: int       # 1-indexed starting line
          text: str
      
      
      def chunk_text(text: str, *, path: str = "<stdin>",
                     granularity: Granularity = "paragraph") -> list[Chunk]:
          """Split text into chunks with source line numbers preserved."""
          chunks: list[Chunk] = []
          if granularity == "line":
              for i, line in enumerate(text.splitlines(), start=1):
                  s = line.strip()
                  if s:
                      chunks.append(Chunk(path=path, line=i, text=s))
              return chunks
      
          # paragraph: split on blank lines, track the starting line of each paragraph
          cur_lines: list[str] = []
          cur_start: int | None = None
          for i, line in enumerate(text.splitlines(), start=1):
              if line.strip() == "":
                  if cur_lines:
                      chunks.append(Chunk(path=path, line=cur_start,
                                          text="\n".join(cur_lines).strip()))
                      cur_lines = []
                      cur_start = None
              else:
                  if cur_start is None:
                      cur_start = i
                  cur_lines.append(line)
          if cur_lines:
              chunks.append(Chunk(path=path, line=cur_start or 1,
                                  text="\n".join(cur_lines).strip()))
          return chunks
      
      
      def load_corpus(path: str | Path, *, include: str = "*.txt",
                      granularity: Granularity = "paragraph") -> list[Chunk]:
          """Load and chunk a file, or recursively a directory matching `include`."""
          p = Path(path)
          chunks: list[Chunk] = []
          if p.is_file():
              chunks.extend(chunk_text(p.read_text(errors="replace"),
                                       path=str(p), granularity=granularity))
              return chunks
          if p.is_dir():
              for fp in sorted(p.rglob("*")):
                  if fp.is_file() and fnmatch.fnmatch(fp.name, include):
                      chunks.extend(chunk_text(fp.read_text(errors="replace"),
                                               path=str(fp), granularity=granularity))
              return chunks
          raise FileNotFoundError(path)
      
      
      # ---------------------------------------------------------------------------
      # Search
      # ---------------------------------------------------------------------------
      
      @dataclass
      class Match:
          path: str
          line: int
          text: str
          score: float
      
      
      def semantic_grep(query: str, corpus: str | Path | list[Chunk], *,
                        top_k: int = 10, threshold: float | None = None,
                        granularity: Granularity = "paragraph",
                        include: str = "*.txt",
                        model: str = _DEFAULT_MODEL, dim: int = 768,
                        task: Literal["text", "code"] = "text") -> list[Match]:
          """Semantic search over a file, directory, or pre-chunked list.
      
          - top_k: max results (set to None for all above threshold)
          - threshold: cosine similarity cutoff (None = no filter, use top_k only)
          - granularity: paragraph (default) or line
          - task: 'text' → RETRIEVAL_QUERY/DOCUMENT; 'code' → CODE_RETRIEVAL_QUERY/DOCUMENT
      
          Raises ValueError on empty query. Returns [] for empty corpus without
          hitting the API.
          """
          if not query or not query.strip():
              raise ValueError("query must be non-empty")
      
          # Resolve corpus
          if isinstance(corpus, list):
              chunks = corpus
          else:
              chunks = load_corpus(corpus, include=include, granularity=granularity)
          if not chunks:
              return []
      
          q_task: TaskType = "CODE_RETRIEVAL_QUERY" if task == "code" else "RETRIEVAL_QUERY"
          d_task: TaskType = "RETRIEVAL_DOCUMENT"
      
          q_vec = _embed_one(_truncate_to_token_budget(query), q_task, model=model, dim=dim)
          d_vecs = embed_batch([c.text for c in chunks], d_task, model=model, dim=dim)
      
          # Cosine sim — vectors are normalized when dim < 3072 (handled in _embed_one)
          scores = d_vecs @ q_vec  # (N,)
      
          # Rank
          order = np.argsort(-scores)
          matches: list[Match] = []
          for idx in order:
              s = float(scores[idx])
              if threshold is not None and s < threshold:
                  break
              matches.append(Match(path=chunks[idx].path, line=chunks[idx].line,
                                   text=chunks[idx].text, score=s))
              if top_k is not None and len(matches) >= top_k:
                  break
          return matches
      
      
      # ---------------------------------------------------------------------------
      # Formatting
      # ---------------------------------------------------------------------------
      
      def format_grep(matches: list[Match], *, max_text_chars: int = 200,
                      show_score: bool = True) -> str:
          """Format matches in grep-compatible `path:line: text` form."""
          lines = []
          for m in matches:
              snippet = m.text.replace("\n", " ")
              if len(snippet) > max_text_chars:
                  snippet = snippet[:max_text_chars - 1] + "…"
              prefix = f"{m.path}:{m.line}:"
              tail = f"  [{m.score:.3f}]" if show_score else ""
              lines.append(f"{prefix} {snippet}{tail}")
          return "\n".join(lines)
      
  • CHANGELOG.md 3.1 KB
    # semantic-grep - Changelog
    
    ## 2026-07-21
    
    ### Changed — default encoder is now `gemini-embedding-2`
    - `_DEFAULT_MODEL`: `gemini-embedding-001` → `gemini-embedding-2` (GA 2026-04-22).
    - Embedding-2 is a strict superset: same request shape and dims for text, plus
      image and audio into the same vector space. Verified through the CF gateway
      2026-07-21 — 001 rejects non-text input with HTTP 400.
    
    ### Retired — `gemini-embedding-001`
    - Listed in `_RETIRED_MODELS`; passing it explicitly still works but raises a
      `DeprecationWarning` pointing at the replacement.
    - **No migration needed.** This skill is serverless and re-embeds on every call
      (no persistent index), and the memory system stopped generating embeddings in
      v0.13.0 — so no stored 001 vectors exist to invalidate.
    - Vectors from different encoders are not comparable; never mix them in one index.
    
    All notable changes to the `semantic-grep` skill are documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
    
    ## [0.2.1] - 2026-09-09
    
    ### Other
    
    - prompt-audit: dated prompting patterns across the skill catalogue (#791)
    - Deprecate mapping-codebases; adopt ruff 0.16.0 baseline (#747)
    
    ## [0.2.0] - 2026-07-23
    
    ### Fixed
    
    - fall back to env vars for CF gateway credentials (#707)
    
    ### Other
    
    - semantic-grep: default to gemini-embedding-2, retire gemini-embedding-001 (#743)
    
    ## [0.1.1] - 2026-04-19
    
    ### Other
    
    - Add semantic-grep skill (v0.1.0) (#558)
    
    ## [0.1.1] - 2026-04-19
    
    ### Fixed
    
    - tighten char-per-token ratio (4 → 1.5) so non-ASCII text does not silently overflow Gemini's 2048-token input limit
    - `_load_env` now strips `export ` prefix, tolerating shell-sourceable .env files
    - `semantic_grep()` raises `ValueError` on empty query instead of hitting the API with empty payload
    
    ### Documentation
    
    - SKILL.md clarifies `include` glob matches filename only (not path)
    - Limitations section expanded with memory ceiling, symlink-following, and batch bulk-failure notes surfaced by adversarial review
    
    ## [0.1.0] - 2026-04-19
    
    ### Added
    
    - initial skill: semantic search over text files via `gemini-embedding-001`
    - `semantic_grep()` — main search function with `top_k`/`threshold`/`granularity` flags
    - `embed_batch()` using Gemini's `:batchEmbedContents` endpoint (1 HTTP call per 100 chunks)
    - `load_corpus()` / `chunk_text()` — paragraph or line granularity
    - `format_grep()` — grep-compatible `path:line: text [score]` output
    - asymmetric `RETRIEVAL_QUERY` / `RETRIEVAL_DOCUMENT` task types; `code` task mode via `CODE_RETRIEVAL_QUERY`
    - MRL truncation (128/768/1536/3072 dims) with client-side renormalization for dims < 3072
    - credential loading via `/mnt/project/proxy.env` (CF AI Gateway BYOK) with direct-API fallback
    
    ### Notes
    
    - conceptually inspired by [`jina-grep-cli`](https://github.com/jina-ai/jina-grep-cli); swaps MLX backend for Gemini API
    - no persistent index yet — every call re-embeds the corpus (serverless mode equivalent)
    - embedding function is duplicated here vs `invoking-gemini/scripts/gemini_client.py`; should be factored up when invoking-gemini adds embedding support
  • README.md 153 B
    # semantic-grep
    
    Semantic search over text files using Gemini embeddings. In-process Python, not a subprocess CLI. See [SKILL.md](./SKILL.md) for usage.
    
  • SKILL.md 8.4 KB
    ---
    name: semantic-grep
    description: In-process semantic search over text files or in-memory strings, using Gemini embeddings via the CF AI Gateway. Use when user wants fuzzy/conceptual search where exact-keyword grep would miss — "sessions discussing regulatory constraints", "code about retry logic", "notes mentioning burnout even if the word isn't there". Complements searching-codebases (regex/AST) and extracting-keywords (YAKE). Do NOT use when an exact string/regex match is what's wanted — grep/rg wins on speed and precision there.
    metadata:
      version: 0.2.1
    ---
    
    # Semantic Grep
    
    jina-grep-style semantic search, done in-process via Python rather than as an external CLI. Embeds query + corpus chunks with `gemini-embedding-2`, ranks by cosine similarity, returns grep-format output.
    
    ## When Semantic Search Helps
    
    The core trade-off (lifted from `jina-grep-cli`'s own docs and validated in testing):
    
    | Task | Tool |
    |------|------|
    | Known exact string, filename, or regex | `grep` / `rg` / `searching-codebases` |
    | "What files discuss concept X" when X may not appear verbatim | **semantic-grep** |
    | Hybrid: prefilter with grep, rerank by concept | grep → `rerank_candidates()` |
    
    **Regression test result (workshop session corpus, 135 docs):**
    - *"handling regulatory constraints"* → top hit *"Engineering AI Systems Under Sovereignty Constraints"* (0.67). ✓
    - *"sessions about GEPA"* → top hit *"Gemma, DeepMind's Family of Open Models"* (0.69). ✗ — false positive on phonetic neighbor. GEPA is mentioned verbatim in one session description; grep would find it correctly.
    
    **Rule: when the user query reads like a named entity or keyword, try grep first. Only reach for semantic-grep when paraphrase/concept matching is actually needed.**
    
    ## Setup
    
    Credentials via `proxy.env` (Cloudflare AI Gateway w/ BYOK — same pattern as `invoking-gemini`):
    
    ```
    CF_ACCOUNT_ID=...
    CF_GATEWAY_ID=...
    CF_API_TOKEN=...
    ```
    
    Direct-API fallback: `GOOGLE_API_KEY` or `GEMINI_API_KEY` env var. No dependencies beyond `requests` + `numpy`.
    
    ## Quick Start
    
    ```python
    import sys
    sys.path.insert(0, '/mnt/skills/user/semantic-grep/scripts')
    from semantic_grep import semantic_grep, format_grep
    
    # Directory of .txt files
    results = semantic_grep("error handling under load", "/path/to/notes",
                            top_k=5, granularity="paragraph")
    print(format_grep(results))
    # notes/incidents.txt:42:  When the queue depth exceeds... [0.71]
    # notes/postmortem.txt:8:  Under sustained traffic we saw... [0.68]
    ```
    
    ## Core API
    
    ### `semantic_grep(query, corpus, *, top_k=10, threshold=None, ...)`
    
    Main search function.
    
    - `query` *(str)* — the search query (embedded with `RETRIEVAL_QUERY` task type)
    - `corpus` *(str | Path | list[Chunk])* — a file, directory, or pre-chunked list
    - `top_k` *(int | None)* — max results; `None` = all above threshold
    - `threshold` *(float | None)* — cosine similarity cutoff; `None` = no filter (top_k only)
    - `granularity` *("paragraph" | "line")* — how to chunk files (default paragraph)
    - `include` *(str)* — filename-glob filter when `corpus` is a directory (default `"*.txt"`). Matches against `Path.name` only, not the full path — `"*.md"` works, `"docs/*.md"` does not.
    - `model` *(str)* — default `"gemini-embedding-2"`. `gemini-embedding-001` is **retired** (text-only) and warns if passed explicitly.
    - `dim` *(int)* — 128 / 768 / 1536 / 3072 (default 768; MRL-truncated + renormalized)
    - `task` *("text" | "code")* — selects text vs code task types
    
    Returns `list[Match]` where `Match` has `path`, `line`, `text`, `score`.
    
    ### `load_corpus(path, *, include="*.txt", granularity="paragraph") -> list[Chunk]`
    
    Load and chunk a file or directory without embedding. Useful for inspecting what gets embedded before paying for the API call.
    
    ### `embed_batch(texts, task_type, *, model, dim, group_size=100) -> np.ndarray`
    
    Lower-level: embed a list of strings directly via `:batchEmbedContents`. Returns `(N, dim)` float32 array, rows normalized when `dim < 3072`.
    
    ### `format_grep(matches, *, max_text_chars=200, show_score=True) -> str`
    
    Format matches as grep output: `path:line: snippet  [score]`.
    
    ## Pipe-mode Rerank Pattern
    
    The highest-leverage use isn't naive full-corpus semantic search — it's hybrid retrieval: **fast coarse filter → semantic rerank**.
    
    ```python
    import subprocess
    from semantic_grep import Chunk, semantic_grep, format_grep
    
    # Stage 1: fast exact/regex prefilter with rg
    result = subprocess.run(
        ["rg", "-n", "--no-heading", "error|fail|timeout", "logs/"],
        capture_output=True, text=True,
    )
    
    # Parse `path:line:text` into Chunks
    chunks = []
    for raw in result.stdout.splitlines():
        path, line, text = raw.split(":", 2)
        chunks.append(Chunk(path=path, line=int(line), text=text))
    
    # Stage 2: semantic rerank on the prefiltered subset
    ranked = semantic_grep("intermittent queue saturation during peak traffic",
                           chunks, top_k=10)
    print(format_grep(ranked))
    ```
    
    This is how you scale past the "embed the whole corpus every call" limit without needing a vector DB. The exact-match stage cheaply cuts millions of lines to thousands; semantic reranks those.
    
    ## Task Types (Gemini)
    
    - **text mode** (default): query → `RETRIEVAL_QUERY`, docs → `RETRIEVAL_DOCUMENT`. Asymmetric — documented to outperform symmetric encoding for retrieval.
    - **code mode**: query → `CODE_RETRIEVAL_QUERY`, docs → `RETRIEVAL_DOCUMENT`. Use when searching code with natural-language queries.
    
    Use `SEMANTIC_SIMILARITY` (symmetric) only if you're doing pairwise sim, not retrieval. This module doesn't expose that path yet.
    
    ## Model Notes
    
    `gemini-embedding-2` (GA since 2026-04-22) — general-purpose **and** multimodal.
    Verified 2026-07-21 via the CF gateway: text, image and audio all embed to the
    same space at the requested dim, L2-normalized. The retired `gemini-embedding-001`
    was text-only and rejected non-text input with HTTP 400:
    - 2,048 input token limit per text. Longer texts are truncated at ~8K chars (approximation).
    - Matryoshka (MRL) — 3072 native dims, safely truncatable to 1536/768/256/128.
    - 3072 is auto-normalized; lower dims need client-side renorm (handled here).
    - Pricing: $0.15 / 1M input tokens. 135 medium paragraphs ≈ 15K tokens ≈ $0.002 per query.
    
    `gemini-embedding-2-preview` (March 2026) is multimodal and currently top of MTEB. Set `model="gemini-embedding-2-preview"` to opt in once the preview stabilizes.
    
    ## Limitations
    
    - **No persistent index.** Every call re-embeds the corpus. Fine for <~1K chunks; prohibitive for real knowledge bases. Phase 2: cache embeddings by content hash.
    - **Token budget is approximated by char count (×1.5).** Conservative for mixed-script text; over-truncates English slightly. Real tokenizer would use the Gemini tokenizer endpoint but costs an extra call per embed.
    - **Batch bulk-failure diagnostic.** If one text in a group of 100 overflows or is rejected by safety filters, the whole batch fails and the 99 good ones are lost. No per-index fallback yet.
    - **No memory ceiling on corpus size.** `semantic_grep` pre-allocates `(N, dim)` float32; 1M chunks at dim=768 ≈ 3GB. Caller is responsible for sane chunk counts. `load_corpus` also follows symlinks via `rglob` — fine in a trusted single-user container, not for untrusted paths.
    - **Sequential batch groups.** `group_size=100` per HTTP call; groups run serially. For >1K chunks, add asyncio — not needed yet.
    - **No CLI shim.** Called as a Python module, not a subprocess. Per design: "within an LLM rather than calling out to one."
    - **Embedding function lives here, not in `invoking-gemini`.** Should be factored up when invoking-gemini adds embedding support. Tracked as followup.
    
    ## Related Skills
    
    - `invoking-gemini` — sibling; handles Gemini text + image generation through the same CF gateway. Shares credential pattern.
    - `searching-codebases` — regex/AST search. Use first when the query is a known pattern.
    - `extracting-keywords` — YAKE keyword extraction; orthogonal, but pairs well for building query terms from a long prompt.
    - `exploring-codebases` — for understanding repo structure. Semantic-grep doesn't replace AST-based navigation.
    
    ## Attribution
    
    Conceptually inspired by [`jina-grep-cli`](https://github.com/jina-ai/jina-grep-cli) — we kept the retrieval shape (grep-compatible output, asymmetric query/doc embeddings, threshold + top-k) but swapped the MLX/Apple-Silicon backend for a portable Gemini API call. The original's pipe-mode rerank pattern is the most generalizable idea it contributes and is preserved here.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related