semantic-grep
In-process semantic search over text files or in-memory strings, using Gemini embeddings via the CF AI Gateway. Use when user wants fuzzy/conceptual search where exact-keyword grep would miss — "sessions discussing regulatory constraints", "code about retry logic", "notes mention
Install
npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/code-intelligence/skills/semantic-grep
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
git clone https://github.com/oaustegard/claude-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.
README
semantic-grep
Semantic search over text files using Gemini embeddings. In-process Python, not a subprocess CLI. See SKILL.md for usage.
Skill manifest
Semantic Grep
jina-grep-style semantic search, done in-process via Python rather than as an external CLI. Embeds query + corpus chunks with gemini-embedding-2, ranks by cosine similarity, returns grep-format output.
When Semantic Search Helps
The core trade-off (lifted from jina-grep-cli's own docs and validated in testing):
| Task | Tool |
|---|---|
| Known exact string, filename, or regex | grep / rg / searching-codebases |
| "What files discuss concept X" when X may not appear verbatim | semantic-grep |
| Hybrid: prefilter with grep, rerank by concept | grep → rerank_candidates() |
Regression test result (workshop session corpus, 135 docs):
- "handling regulatory constraints" → top hit "Engineering AI Systems Under Sovereignty Constraints" (0.67). ✓
- "sessions about GEPA" → top hit "Gemma, DeepMind's Family of Open Models" (0.69). ✗ — false positive on phonetic neighbor. GEPA is mentioned verbatim in one session description; grep would find it correctly.
Rule: when the user query reads like a named entity or keyword, try grep first. Only reach for semantic-grep when paraphrase/concept matching is actually needed.
Setup
Credentials via proxy.env (Cloudflare AI Gateway w/ BYOK — same pattern as invoking-gemini):
CF_ACCOUNT_ID=...
CF_GATEWAY_ID=...
CF_API_TOKEN=...
Direct-API fallback: GOOGLE_API_KEY or GEMINI_API_KEY env var. No dependencies beyond requests + numpy.
Quick Start
import sys
sys.path.insert(0, '/mnt/skills/user/semantic-grep/scripts')
from semantic_grep import semantic_grep, format_grep
# Directory of .txt files
results = semantic_grep("error handling under load", "/path/to/notes",
top_k=5, granularity="paragraph")
print(format_grep(results))
# notes/incidents.txt:42: When the queue depth exceeds... [0.71]
# notes/postmortem.txt:8: Under sustained traffic we saw... [0.68]
Core API
semantic_grep(query, corpus, *, top_k=10, threshold=None, ...)
Main search function.
query(str) — the search query (embedded withRETRIEVAL_QUERYtask type)corpus(str | Path | list[Chunk]) — a file, directory, or pre-chunked listtop_k(int | None) — max results;None= all above thresholdthreshold(float | None) — cosine similarity cutoff;None= no filter (top_k only)granularity("paragraph" | "line") — how to chunk files (default paragraph)include(str) — filename-glob filter whencorpusis a directory (default"*.txt"). Matches againstPath.nameonly, not the full path —"*.md"works,"docs/*.md"does not.model(str) — default"gemini-embedding-2".gemini-embedding-001is retired (text-only) and warns if passed explicitly.dim(int) — 128 / 768 / 1536 / 3072 (default 768; MRL-truncated + renormalized)task("text" | "code") — selects text vs code task types
Returns list[Match] where Match has path, line, text, score.
load_corpus(path, *, include="*.txt", granularity="paragraph") -> list[Chunk]
Load and chunk a file or directory without embedding. Useful for inspecting what gets embedded before paying for the API call.
embed_batch(texts, task_type, *, model, dim, group_size=100) -> np.ndarray
Lower-level: embed a list of strings directly via :batchEmbedContents. Returns (N, dim) float32 array, rows normalized when dim < 3072.
format_grep(matches, *, max_text_chars=200, show_score=True) -> str
Format matches as grep output: path:line: snippet [score].
Pipe-mode Rerank Pattern
The highest-leverage use isn't naive full-corpus semantic search — it's hybrid retrieval: fast coarse filter → semantic rerank.
import subprocess
from semantic_grep import Chunk, semantic_grep, format_grep
# Stage 1: fast exact/regex prefilter with rg
result = subprocess.run(
["rg", "-n", "--no-heading", "error|fail|timeout", "logs/"],
capture_output=True, text=True,
)
# Parse `path:line:text` into Chunks
chunks = []
for raw in result.stdout.splitlines():
path, line, text = raw.split(":", 2)
chunks.append(Chunk(path=path, line=int(line), text=text))
# Stage 2: semantic rerank on the prefiltered subset
ranked = semantic_grep("intermittent queue saturation during peak traffic",
chunks, top_k=10)
print(format_grep(ranked))
This is how you scale past the "embed the whole corpus every call" limit without needing a vector DB. The exact-match stage cheaply cuts millions of lines to thousands; semantic reranks those.
Task Types (Gemini)
- text mode (default): query →
RETRIEVAL_QUERY, docs →RETRIEVAL_DOCUMENT. Asymmetric — documented to outperform symmetric encoding for retrieval. - code mode: query →
CODE_RETRIEVAL_QUERY, docs →RETRIEVAL_DOCUMENT. Use when searching code with natural-language queries.
Use SEMANTIC_SIMILARITY (symmetric) only if you're doing pairwise sim, not retrieval. This module doesn't expose that path yet.
Model Notes
gemini-embedding-2 (GA since 2026-04-22) — general-purpose and multimodal.
Verified 2026-07-21 via the CF gateway: text, image and audio all embed to the
same space at the requested dim, L2-normalized. The retired gemini-embedding-001
was text-only and rejected non-text input with HTTP 400:
- 2,048 input token limit per text. Longer texts are truncated at ~8K chars (approximation).
- Matryoshka (MRL) — 3072 native dims, safely truncatable to 1536/768/256/128.
- 3072 is auto-normalized; lower dims need client-side renorm (handled here).
- Pricing: $0.15 / 1M input tokens. 135 medium paragraphs ≈ 15K tokens ≈ $0.002 per query.
gemini-embedding-2-preview (March 2026) is multimodal and currently top of MTEB. Set model="gemini-embedding-2-preview" to opt in once the preview stabilizes.
Limitations
- No persistent index. Every call re-embeds the corpus. Fine for <~1K chunks; prohibitive for real knowledge bases. Phase 2: cache embeddings by content hash.
- Token budget is approximated by char count (×1.5). Conservative for mixed-script text; over-truncates English slightly. Real tokenizer would use the Gemini tokenizer endpoint but costs an extra call per embed.
- Batch bulk-failure diagnostic. If one text in a group of 100 overflows or is rejected by safety filters, the whole batch fails and the 99 good ones are lost. No per-index fallback yet.
- No memory ceiling on corpus size.
semantic_greppre-allocates(N, dim)float32; 1M chunks at dim=768 ≈ 3GB. Caller is responsible for sane chunk counts.load_corpusalso follows symlinks viarglob— fine in a trusted single-user container, not for untrusted paths. - Sequential batch groups.
group_size=100per HTTP call; groups run serially. For >1K chunks, add asyncio — not needed yet. - No CLI shim. Called as a Python module, not a subprocess. Per design: "within an LLM rather than calling out to one."
- Embedding function lives here, not in
invoking-gemini. Should be factored up when invoking-gemini adds embedding support. Tracked as followup.
Related Skills
invoking-gemini— sibling; handles Gemini text + image generation through the same CF gateway. Shares credential pattern.searching-codebases— regex/AST search. Use first when the query is a known pattern.extracting-keywords— YAKE keyword extraction; orthogonal, but pairs well for building query terms from a long prompt.exploring-codebases— for understanding repo structure. Semantic-grep doesn't replace AST-based navigation.
Attribution
Conceptually inspired by jina-grep-cli — we kept the retrieval shape (grep-compatible output, asymmetric query/doc embeddings, threshold + top-k) but swapped the MLX/Apple-Silicon backend for a portable Gemini API call. The original's pipe-mode rerank pattern is the most generalizable idea it contributes and is preserved here.
Files (claude-skills)
-
scripts
-
semantic_grep.py 14 KB
""" semantic_grep: a jina-grep reproduction using Gemini embeddings. In-process semantic search over text files or in-memory strings. Routes through Cloudflare AI Gateway (proxy.env) when available. Model: gemini-embedding-2 (GA 2026-04-22; general-purpose AND multimodal, MRL truncation supported). Supersedes gemini-embedding-001, which is text-only. Verified 2026-07-21 through the CF gateway: text, image (image/png) and audio (audio/wav) all embed to the same space at the requested dim, L2-normalized. 001 rejects non-text with HTTP 400. Task types used: RETRIEVAL_QUERY (query) + RETRIEVAL_DOCUMENT (corpus) — asymmetric. Design notes: - Serverless-mode equivalent: every call re-embeds the corpus. No persistent index. - Caller-facing API returns structured dicts. Grep-format is a formatter, not the core. - Flag surface mirrors jina-grep where it makes sense for a Python API (top_k, threshold, granularity, include). -r recursion is implicit when path is a directory. """ from __future__ import annotations import fnmatch import os import time from dataclasses import dataclass from pathlib import Path from typing import Literal import numpy as np import requests # --------------------------------------------------------------------------- # Credentials (reuse invoking-gemini's proxy pattern) # --------------------------------------------------------------------------- _CF_GATEWAY_BASE = "https://gateway.ai.cloudflare.com/v1" _DIRECT_BASE = "https://generativelanguage.googleapis.com/v1beta" _DEFAULT_MODEL = "gemini-embedding-2" # Retired 2026-07-21. gemini-embedding-001 is text-only and superseded by # gemini-embedding-2 (general-purpose AND multimodal, same request shape and # dims). Still callable if explicitly requested, but warns — nothing in this # skill persists vectors, so there is no index to migrate. _RETIRED_MODELS = { "gemini-embedding-001": "gemini-embedding-2", } def _check_retired(model: str) -> None: """Warn once if a caller explicitly asks for a retired embedding model.""" if model in _RETIRED_MODELS: import warnings warnings.warn( f"{model} is retired; use {_RETIRED_MODELS[model]} " f"(general-purpose and multimodal, same dims). Vectors from " f"different encoders are not comparable — do not mix them.", DeprecationWarning, stacklevel=3, ) _MAX_INPUT_TOKENS = 2048 # conservative; documented for embedding-001, unverified for -2 # Conservative char-per-token ratio. English prose is ~4 chars/token, # but CJK is closer to 1 char/token and emoji can be <1. Picking 1.5 as a safe # lower bound means we under-fill for English (some wasted context) in exchange # for not silently overflowing the API limit for multi-script text. _APPROX_CHARS_PER_TOKEN = 1.5 def _load_env(path: Path) -> dict: if not path.exists(): return {} out = {} for line in path.read_text().splitlines(): line = line.strip() if not line or line.startswith("#") or "=" not in line: continue if line.startswith("export "): line = line[7:].lstrip() k, v = line.split("=", 1) out[k.strip()] = v.strip().strip("'\"") return out def _embed_url(model: str, *, batch: bool = False) -> tuple[str, dict]: """Return (url, extra_headers) for embedContent / batchEmbedContents call.""" endpoint = "batchEmbedContents" if batch else "embedContent" proxy = _load_env(Path("/mnt/project/proxy.env")) cf_account = proxy.get("CF_ACCOUNT_ID") or os.environ.get("CF_ACCOUNT_ID") cf_gateway = proxy.get("CF_GATEWAY_ID") or os.environ.get("CF_GATEWAY_ID") cf_token = proxy.get("CF_API_TOKEN") or os.environ.get("CF_API_TOKEN") if cf_account and cf_gateway and cf_token: url = ( f"{_CF_GATEWAY_BASE}/{cf_account}/{cf_gateway}" f"/google-ai-studio/v1beta/models/{model}:{endpoint}" ) headers = {"cf-aig-authorization": f"Bearer {cf_token}"} return url, headers key = os.environ.get("GOOGLE_API_KEY") or os.environ.get("GEMINI_API_KEY") if not key: raise RuntimeError("No credentials: missing proxy.env / CF_* env vars / GOOGLE_API_KEY") url = f"{_DIRECT_BASE}/models/{model}:{endpoint}" return url, {"x-goog-api-key": key} # --------------------------------------------------------------------------- # Embedding # --------------------------------------------------------------------------- TaskType = Literal[ "RETRIEVAL_QUERY", "RETRIEVAL_DOCUMENT", "SEMANTIC_SIMILARITY", "CLASSIFICATION", "CLUSTERING", "QUESTION_ANSWERING", "FACT_VERIFICATION", "CODE_RETRIEVAL_QUERY", ] def _embed_one(text: str, task_type: TaskType, *, model: str, dim: int, timeout: float = 30.0, retries: int = 2) -> np.ndarray: """Embed a single string. Normalizes output if dim < 3072.""" _check_retired(model) url, headers = _embed_url(model) headers = {**headers, "Content-Type": "application/json"} body = { "content": {"parts": [{"text": text}]}, "taskType": task_type, "outputDimensionality": dim, } last_err = None for attempt in range(retries + 1): try: r = requests.post(url, headers=headers, json=body, timeout=timeout) if r.status_code == 200: data = r.json() vals = data["embedding"]["values"] arr = np.asarray(vals, dtype=np.float32) if dim < 3072: n = np.linalg.norm(arr) if n > 0: arr = arr / n return arr # Retry on 429/5xx if r.status_code in (429, 500, 502, 503, 504) and attempt < retries: time.sleep(1.5 ** attempt) continue raise RuntimeError(f"Embed failed {r.status_code}: {r.text[:300]}") except requests.exceptions.RequestException as e: last_err = e if attempt < retries: time.sleep(1.5 ** attempt) continue raise raise RuntimeError(f"Embed failed after retries: {last_err}") def _truncate_to_token_budget(text: str) -> str: """Approximate 2048-token cap by char count. Gemini will reject overflow otherwise. Uses a conservative chars-per-token ratio (1.5) so mixed-script / CJK text is not silently truncated-too-long. """ limit = int(_MAX_INPUT_TOKENS * _APPROX_CHARS_PER_TOKEN) # ~3K chars return text[:limit] if len(text) > limit else text def embed_batch(texts: list[str], task_type: TaskType, *, model: str = _DEFAULT_MODEL, dim: int = 768, group_size: int = 100, timeout: float = 90.0, retries: int = 3) -> np.ndarray: """Embed N strings via :batchEmbedContents (single HTTP call per group of `group_size`). Returns (N, dim) array. Normalizes output rows when dim < 3072. """ if not texts: return np.zeros((0, dim), dtype=np.float32) url, base_headers = _embed_url(model, batch=True) headers = {**base_headers, "Content-Type": "application/json"} out = np.zeros((len(texts), dim), dtype=np.float32) for start in range(0, len(texts), group_size): group = texts[start:start + group_size] body = { "requests": [ { "model": f"models/{model}", "content": {"parts": [{"text": _truncate_to_token_budget(t)}]}, "taskType": task_type, "outputDimensionality": dim, } for t in group ] } for attempt in range(retries + 1): try: r = requests.post(url, headers=headers, json=body, timeout=timeout) if r.status_code == 200: data = r.json() embs = data.get("embeddings", []) if len(embs) != len(group): raise RuntimeError( f"Batch size mismatch: sent {len(group)}, got {len(embs)}" ) for i, e in enumerate(embs): arr = np.asarray(e["values"], dtype=np.float32) if dim < 3072: n = np.linalg.norm(arr) if n > 0: arr = arr / n out[start + i] = arr break # success, next group if r.status_code in (429, 500, 502, 503, 504) and attempt < retries: time.sleep(1.5 ** attempt) continue raise RuntimeError(f"Batch embed failed {r.status_code}: {r.text[:400]}") except requests.exceptions.RequestException as e: if attempt < retries: time.sleep(1.5 ** attempt) continue raise RuntimeError(f"Batch embed network error: {e}") from e return out # --------------------------------------------------------------------------- # Chunking # --------------------------------------------------------------------------- Granularity = Literal["line", "paragraph"] @dataclass class Chunk: path: str # source identifier (filepath or logical id) line: int # 1-indexed starting line text: str def chunk_text(text: str, *, path: str = "<stdin>", granularity: Granularity = "paragraph") -> list[Chunk]: """Split text into chunks with source line numbers preserved.""" chunks: list[Chunk] = [] if granularity == "line": for i, line in enumerate(text.splitlines(), start=1): s = line.strip() if s: chunks.append(Chunk(path=path, line=i, text=s)) return chunks # paragraph: split on blank lines, track the starting line of each paragraph cur_lines: list[str] = [] cur_start: int | None = None for i, line in enumerate(text.splitlines(), start=1): if line.strip() == "": if cur_lines: chunks.append(Chunk(path=path, line=cur_start, text="\n".join(cur_lines).strip())) cur_lines = [] cur_start = None else: if cur_start is None: cur_start = i cur_lines.append(line) if cur_lines: chunks.append(Chunk(path=path, line=cur_start or 1, text="\n".join(cur_lines).strip())) return chunks def load_corpus(path: str | Path, *, include: str = "*.txt", granularity: Granularity = "paragraph") -> list[Chunk]: """Load and chunk a file, or recursively a directory matching `include`.""" p = Path(path) chunks: list[Chunk] = [] if p.is_file(): chunks.extend(chunk_text(p.read_text(errors="replace"), path=str(p), granularity=granularity)) return chunks if p.is_dir(): for fp in sorted(p.rglob("*")): if fp.is_file() and fnmatch.fnmatch(fp.name, include): chunks.extend(chunk_text(fp.read_text(errors="replace"), path=str(fp), granularity=granularity)) return chunks raise FileNotFoundError(path) # --------------------------------------------------------------------------- # Search # --------------------------------------------------------------------------- @dataclass class Match: path: str line: int text: str score: float def semantic_grep(query: str, corpus: str | Path | list[Chunk], *, top_k: int = 10, threshold: float | None = None, granularity: Granularity = "paragraph", include: str = "*.txt", model: str = _DEFAULT_MODEL, dim: int = 768, task: Literal["text", "code"] = "text") -> list[Match]: """Semantic search over a file, directory, or pre-chunked list. - top_k: max results (set to None for all above threshold) - threshold: cosine similarity cutoff (None = no filter, use top_k only) - granularity: paragraph (default) or line - task: 'text' → RETRIEVAL_QUERY/DOCUMENT; 'code' → CODE_RETRIEVAL_QUERY/DOCUMENT Raises ValueError on empty query. Returns [] for empty corpus without hitting the API. """ if not query or not query.strip(): raise ValueError("query must be non-empty") # Resolve corpus if isinstance(corpus, list): chunks = corpus else: chunks = load_corpus(corpus, include=include, granularity=granularity) if not chunks: return [] q_task: TaskType = "CODE_RETRIEVAL_QUERY" if task == "code" else "RETRIEVAL_QUERY" d_task: TaskType = "RETRIEVAL_DOCUMENT" q_vec = _embed_one(_truncate_to_token_budget(query), q_task, model=model, dim=dim) d_vecs = embed_batch([c.text for c in chunks], d_task, model=model, dim=dim) # Cosine sim — vectors are normalized when dim < 3072 (handled in _embed_one) scores = d_vecs @ q_vec # (N,) # Rank order = np.argsort(-scores) matches: list[Match] = [] for idx in order: s = float(scores[idx]) if threshold is not None and s < threshold: break matches.append(Match(path=chunks[idx].path, line=chunks[idx].line, text=chunks[idx].text, score=s)) if top_k is not None and len(matches) >= top_k: break return matches # --------------------------------------------------------------------------- # Formatting # --------------------------------------------------------------------------- def format_grep(matches: list[Match], *, max_text_chars: int = 200, show_score: bool = True) -> str: """Format matches in grep-compatible `path:line: text` form.""" lines = [] for m in matches: snippet = m.text.replace("\n", " ") if len(snippet) > max_text_chars: snippet = snippet[:max_text_chars - 1] + "…" prefix = f"{m.path}:{m.line}:" tail = f" [{m.score:.3f}]" if show_score else "" lines.append(f"{prefix} {snippet}{tail}") return "\n".join(lines)
-
-
CHANGELOG.md 3.1 KB
# semantic-grep - Changelog ## 2026-07-21 ### Changed — default encoder is now `gemini-embedding-2` - `_DEFAULT_MODEL`: `gemini-embedding-001` → `gemini-embedding-2` (GA 2026-04-22). - Embedding-2 is a strict superset: same request shape and dims for text, plus image and audio into the same vector space. Verified through the CF gateway 2026-07-21 — 001 rejects non-text input with HTTP 400. ### Retired — `gemini-embedding-001` - Listed in `_RETIRED_MODELS`; passing it explicitly still works but raises a `DeprecationWarning` pointing at the replacement. - **No migration needed.** This skill is serverless and re-embeds on every call (no persistent index), and the memory system stopped generating embeddings in v0.13.0 — so no stored 001 vectors exist to invalidate. - Vectors from different encoders are not comparable; never mix them in one index. All notable changes to the `semantic-grep` skill are documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/). ## [0.2.1] - 2026-09-09 ### Other - prompt-audit: dated prompting patterns across the skill catalogue (#791) - Deprecate mapping-codebases; adopt ruff 0.16.0 baseline (#747) ## [0.2.0] - 2026-07-23 ### Fixed - fall back to env vars for CF gateway credentials (#707) ### Other - semantic-grep: default to gemini-embedding-2, retire gemini-embedding-001 (#743) ## [0.1.1] - 2026-04-19 ### Other - Add semantic-grep skill (v0.1.0) (#558) ## [0.1.1] - 2026-04-19 ### Fixed - tighten char-per-token ratio (4 → 1.5) so non-ASCII text does not silently overflow Gemini's 2048-token input limit - `_load_env` now strips `export ` prefix, tolerating shell-sourceable .env files - `semantic_grep()` raises `ValueError` on empty query instead of hitting the API with empty payload ### Documentation - SKILL.md clarifies `include` glob matches filename only (not path) - Limitations section expanded with memory ceiling, symlink-following, and batch bulk-failure notes surfaced by adversarial review ## [0.1.0] - 2026-04-19 ### Added - initial skill: semantic search over text files via `gemini-embedding-001` - `semantic_grep()` — main search function with `top_k`/`threshold`/`granularity` flags - `embed_batch()` using Gemini's `:batchEmbedContents` endpoint (1 HTTP call per 100 chunks) - `load_corpus()` / `chunk_text()` — paragraph or line granularity - `format_grep()` — grep-compatible `path:line: text [score]` output - asymmetric `RETRIEVAL_QUERY` / `RETRIEVAL_DOCUMENT` task types; `code` task mode via `CODE_RETRIEVAL_QUERY` - MRL truncation (128/768/1536/3072 dims) with client-side renormalization for dims < 3072 - credential loading via `/mnt/project/proxy.env` (CF AI Gateway BYOK) with direct-API fallback ### Notes - conceptually inspired by [`jina-grep-cli`](https://github.com/jina-ai/jina-grep-cli); swaps MLX backend for Gemini API - no persistent index yet — every call re-embeds the corpus (serverless mode equivalent) - embedding function is duplicated here vs `invoking-gemini/scripts/gemini_client.py`; should be factored up when invoking-gemini adds embedding support -
README.md 153 B
# semantic-grep Semantic search over text files using Gemini embeddings. In-process Python, not a subprocess CLI. See [SKILL.md](./SKILL.md) for usage. -
SKILL.md 8.4 KB
--- name: semantic-grep description: In-process semantic search over text files or in-memory strings, using Gemini embeddings via the CF AI Gateway. Use when user wants fuzzy/conceptual search where exact-keyword grep would miss — "sessions discussing regulatory constraints", "code about retry logic", "notes mentioning burnout even if the word isn't there". Complements searching-codebases (regex/AST) and extracting-keywords (YAKE). Do NOT use when an exact string/regex match is what's wanted — grep/rg wins on speed and precision there. metadata: version: 0.2.1 --- # Semantic Grep jina-grep-style semantic search, done in-process via Python rather than as an external CLI. Embeds query + corpus chunks with `gemini-embedding-2`, ranks by cosine similarity, returns grep-format output. ## When Semantic Search Helps The core trade-off (lifted from `jina-grep-cli`'s own docs and validated in testing): | Task | Tool | |------|------| | Known exact string, filename, or regex | `grep` / `rg` / `searching-codebases` | | "What files discuss concept X" when X may not appear verbatim | **semantic-grep** | | Hybrid: prefilter with grep, rerank by concept | grep → `rerank_candidates()` | **Regression test result (workshop session corpus, 135 docs):** - *"handling regulatory constraints"* → top hit *"Engineering AI Systems Under Sovereignty Constraints"* (0.67). ✓ - *"sessions about GEPA"* → top hit *"Gemma, DeepMind's Family of Open Models"* (0.69). ✗ — false positive on phonetic neighbor. GEPA is mentioned verbatim in one session description; grep would find it correctly. **Rule: when the user query reads like a named entity or keyword, try grep first. Only reach for semantic-grep when paraphrase/concept matching is actually needed.** ## Setup Credentials via `proxy.env` (Cloudflare AI Gateway w/ BYOK — same pattern as `invoking-gemini`): ``` CF_ACCOUNT_ID=... CF_GATEWAY_ID=... CF_API_TOKEN=... ``` Direct-API fallback: `GOOGLE_API_KEY` or `GEMINI_API_KEY` env var. No dependencies beyond `requests` + `numpy`. ## Quick Start ```python import sys sys.path.insert(0, '/mnt/skills/user/semantic-grep/scripts') from semantic_grep import semantic_grep, format_grep # Directory of .txt files results = semantic_grep("error handling under load", "/path/to/notes", top_k=5, granularity="paragraph") print(format_grep(results)) # notes/incidents.txt:42: When the queue depth exceeds... [0.71] # notes/postmortem.txt:8: Under sustained traffic we saw... [0.68] ``` ## Core API ### `semantic_grep(query, corpus, *, top_k=10, threshold=None, ...)` Main search function. - `query` *(str)* — the search query (embedded with `RETRIEVAL_QUERY` task type) - `corpus` *(str | Path | list[Chunk])* — a file, directory, or pre-chunked list - `top_k` *(int | None)* — max results; `None` = all above threshold - `threshold` *(float | None)* — cosine similarity cutoff; `None` = no filter (top_k only) - `granularity` *("paragraph" | "line")* — how to chunk files (default paragraph) - `include` *(str)* — filename-glob filter when `corpus` is a directory (default `"*.txt"`). Matches against `Path.name` only, not the full path — `"*.md"` works, `"docs/*.md"` does not. - `model` *(str)* — default `"gemini-embedding-2"`. `gemini-embedding-001` is **retired** (text-only) and warns if passed explicitly. - `dim` *(int)* — 128 / 768 / 1536 / 3072 (default 768; MRL-truncated + renormalized) - `task` *("text" | "code")* — selects text vs code task types Returns `list[Match]` where `Match` has `path`, `line`, `text`, `score`. ### `load_corpus(path, *, include="*.txt", granularity="paragraph") -> list[Chunk]` Load and chunk a file or directory without embedding. Useful for inspecting what gets embedded before paying for the API call. ### `embed_batch(texts, task_type, *, model, dim, group_size=100) -> np.ndarray` Lower-level: embed a list of strings directly via `:batchEmbedContents`. Returns `(N, dim)` float32 array, rows normalized when `dim < 3072`. ### `format_grep(matches, *, max_text_chars=200, show_score=True) -> str` Format matches as grep output: `path:line: snippet [score]`. ## Pipe-mode Rerank Pattern The highest-leverage use isn't naive full-corpus semantic search — it's hybrid retrieval: **fast coarse filter → semantic rerank**. ```python import subprocess from semantic_grep import Chunk, semantic_grep, format_grep # Stage 1: fast exact/regex prefilter with rg result = subprocess.run( ["rg", "-n", "--no-heading", "error|fail|timeout", "logs/"], capture_output=True, text=True, ) # Parse `path:line:text` into Chunks chunks = [] for raw in result.stdout.splitlines(): path, line, text = raw.split(":", 2) chunks.append(Chunk(path=path, line=int(line), text=text)) # Stage 2: semantic rerank on the prefiltered subset ranked = semantic_grep("intermittent queue saturation during peak traffic", chunks, top_k=10) print(format_grep(ranked)) ``` This is how you scale past the "embed the whole corpus every call" limit without needing a vector DB. The exact-match stage cheaply cuts millions of lines to thousands; semantic reranks those. ## Task Types (Gemini) - **text mode** (default): query → `RETRIEVAL_QUERY`, docs → `RETRIEVAL_DOCUMENT`. Asymmetric — documented to outperform symmetric encoding for retrieval. - **code mode**: query → `CODE_RETRIEVAL_QUERY`, docs → `RETRIEVAL_DOCUMENT`. Use when searching code with natural-language queries. Use `SEMANTIC_SIMILARITY` (symmetric) only if you're doing pairwise sim, not retrieval. This module doesn't expose that path yet. ## Model Notes `gemini-embedding-2` (GA since 2026-04-22) — general-purpose **and** multimodal. Verified 2026-07-21 via the CF gateway: text, image and audio all embed to the same space at the requested dim, L2-normalized. The retired `gemini-embedding-001` was text-only and rejected non-text input with HTTP 400: - 2,048 input token limit per text. Longer texts are truncated at ~8K chars (approximation). - Matryoshka (MRL) — 3072 native dims, safely truncatable to 1536/768/256/128. - 3072 is auto-normalized; lower dims need client-side renorm (handled here). - Pricing: $0.15 / 1M input tokens. 135 medium paragraphs ≈ 15K tokens ≈ $0.002 per query. `gemini-embedding-2-preview` (March 2026) is multimodal and currently top of MTEB. Set `model="gemini-embedding-2-preview"` to opt in once the preview stabilizes. ## Limitations - **No persistent index.** Every call re-embeds the corpus. Fine for <~1K chunks; prohibitive for real knowledge bases. Phase 2: cache embeddings by content hash. - **Token budget is approximated by char count (×1.5).** Conservative for mixed-script text; over-truncates English slightly. Real tokenizer would use the Gemini tokenizer endpoint but costs an extra call per embed. - **Batch bulk-failure diagnostic.** If one text in a group of 100 overflows or is rejected by safety filters, the whole batch fails and the 99 good ones are lost. No per-index fallback yet. - **No memory ceiling on corpus size.** `semantic_grep` pre-allocates `(N, dim)` float32; 1M chunks at dim=768 ≈ 3GB. Caller is responsible for sane chunk counts. `load_corpus` also follows symlinks via `rglob` — fine in a trusted single-user container, not for untrusted paths. - **Sequential batch groups.** `group_size=100` per HTTP call; groups run serially. For >1K chunks, add asyncio — not needed yet. - **No CLI shim.** Called as a Python module, not a subprocess. Per design: "within an LLM rather than calling out to one." - **Embedding function lives here, not in `invoking-gemini`.** Should be factored up when invoking-gemini adds embedding support. Tracked as followup. ## Related Skills - `invoking-gemini` — sibling; handles Gemini text + image generation through the same CF gateway. Shares credential pattern. - `searching-codebases` — regex/AST search. Use first when the query is a known pattern. - `extracting-keywords` — YAKE keyword extraction; orthogonal, but pairs well for building query terms from a long prompt. - `exploring-codebases` — for understanding repo structure. Semantic-grep doesn't replace AST-based navigation. ## Attribution Conceptually inspired by [`jina-grep-cli`](https://github.com/jina-ai/jina-grep-cli) — we kept the retrieval shape (grep-compatible output, asymmetric query/doc embeddings, threshold + top-k) but swapped the MLX/Apple-Silicon backend for a portable Gemini API call. The original's pipe-mode rerank pattern is the most generalizable idea it contributes and is preserved here.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.