api-vector-db-chroma
Chroma vector database -- collection management, automatic embedding, metadata filtering, document storage, query patterns
Install
npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/api-vector-db-chroma/skills/api-vector-db-chroma
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
git clone https://github.com/agents-inc/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Chroma Patterns
Quick Guide: Use
chromadb(v3.x) with@chroma-core/default-embedfor automatic embedding. Chroma auto-embeds documents if no embeddings are provided -- just passdocumentsandidstocollection.add(). Usewherefor metadata filtering andwhereDocumentfor document content filtering ($contains,$regex). Default distance metric isl2(Euclidean); usecosinefor most embedding models viaconfiguration: { hnsw: { space: "cosine" } }. Query results return nested arrays (ids: string[][]) because queries are batched -- always accessresults.ids[0]for a single query. Include only the fields you need via theincludeparameter to reduce payload size.
<critical_requirements>
CRITICAL: Before Using This Skill
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST install @chroma-core/default-embed alongside chromadb -- the default embedding function ships as a separate package since v3)
(You MUST access query results as nested arrays -- results.ids[0], results.documents[0] -- because Chroma batches queries and returns string[][] not string[])
(You MUST use the configuration parameter for HNSW settings -- the legacy metadata: { "hnsw:space": "cosine" } approach is deprecated)
(You MUST use flat metadata values only (string, number, boolean, typed arrays) -- nested objects are not supported and will be rejected)
</critical_requirements>
Examples
- Core Patterns -- Client setup, collection management, add, query, get, update, upsert, delete
- Metadata Filtering -- Filter operators, compound filters, document content filters, whereDocument
- Embedding Functions -- Default, OpenAI, custom embedding functions, provider packages
Additional resources:
- reference.md -- API quick reference, filter operators, include options, limits, production checklist
Auto-detection: Chroma, chromadb, ChromaClient, CloudClient, createCollection, getOrCreateCollection, collection.add, collection.query, collection.get, collection.upsert, queryTexts, queryEmbeddings, nResults, whereDocument, $contains, @chroma-core/default-embed, @chroma-core/openai, EmbeddingFunction, vector database, semantic search, embedding, RAG retrieval, hnsw:space
When to use:
- Semantic search over document embeddings (RAG retrieval)
- Rapid prototyping with automatic embedding generation (no external embedding pipeline needed)
- Metadata-filtered vector search with compound logical operators
- Document content filtering with
$containsand$regex - Local development with in-process or Docker-based Chroma server
Key patterns covered:
- Client setup (HTTP, Cloud, Docker)
- Collection management (create, get, delete, configure HNSW)
- Document CRUD with automatic embedding (add, query, get, update, upsert, delete)
- Metadata filtering (
where) with comparison, set, array, and logical operators - Document content filtering (
whereDocument) with$containsand$regex - Embedding function configuration (default, OpenAI, custom)
- Query result handling (nested array structure, include options)
When NOT to use:
- Full-text search with complex boolean ranking (use a dedicated search engine)
- Relational data with joins and transactions (use a relational database)
- Multi-modal image+text embeddings in TypeScript (currently Python-only in Chroma)
- High-scale production with millions of vectors and strict SLAs (evaluate managed vector databases)
<decision_framework>
Decision Framework
Which Distance Metric?
Which distance metric should I use?
|-- Using embeddings from a language model? -> cosine (normalized, most common)
|-- Need dot product similarity? -> ip (inner product)
|-- Comparing raw feature vectors? -> l2 (Euclidean, Chroma default)
'-- Unsure? -> cosine (safe default for most embedding models)
Which Embedding Function?
Which embedding function should I use?
|-- Quick prototyping, English text? -> @chroma-core/default-embed (all-MiniLM-L6-v2, runs locally)
|-- Need high-quality embeddings? -> @chroma-core/openai (text-embedding-3-small)
|-- Have your own embedding pipeline? -> Pass embeddings directly (skip embedding function)
|-- Need custom model? -> Implement EmbeddingFunction interface
'-- Want all providers? -> npm install @chroma-core/all
Where vs WhereDocument?
How should I filter results?
|-- Structured attributes (category, year, status)? -> where (metadata filter)
|-- Full-text content search? -> whereDocument ($contains, $regex)
|-- Both? -> Combine where + whereDocument in same query
'-- Need exact match on specific IDs? -> get({ ids: [...] })
ChromaClient vs CloudClient?
Which client should I use?
|-- Local development or self-hosted? -> ChromaClient({ path: "http://localhost:8000" })
|-- Chroma Cloud (managed)? -> CloudClient({ apiKey, tenant, database })
|-- Docker deployment? -> ChromaClient with Docker host URL
'-- Testing? -> ChromaClient against local Docker container
</decision_framework>
<red_flags>
RED FLAGS
High Priority Issues:
- Accessing query results as flat arrays instead of nested --
results.idsisstring[][], notstring[]; always useresults.ids[0]for single-query results - Missing
@chroma-core/default-embedpackage -- since v3, the default embedding function ships separately;npm install chromadb @chroma-core/default-embed - Using deprecated
metadata: { "hnsw:space": "cosine" }for HNSW config -- useconfiguration: { hnsw: { space: "cosine" } }instead - Nested objects in metadata -- Chroma only supports flat key-value metadata; nested objects are rejected
Medium Priority Issues:
- Not specifying
includein queries -- default includes vary (queryreturns documents, metadatas, distances;getreturns documents, metadatas); explicitly setincludefor clarity and to control payload size - Using
l2(default) whencosineis appropriate -- most embedding models are normalized for cosine similarity;l2may produce worse results - Calling
add()withoutdocumentsorembeddings-- at least one must be provided; metadata alone is insufficient - Not handling empty results --
results.ids[0]may be an empty array; check length before processing
Common Mistakes:
- Passing
queryEmbeddingsANDqueryTextstogether -- use one or the other, not both - Expecting
update()to create missing records --update()silently ignores non-existent IDs; useupsert()for create-or-update semantics - Calling
delete()with no arguments -- deletes nothing (not everything); passidsorwhereto target specific records - Using
$gt/$lton string metadata -- comparison operators only work on numeric values (int or float)
Gotchas & Edge Cases:
- HNSW configuration (
space,ef_construction,max_neighbors) cannot be changed after collection creation -- you must delete and recreate the collection collection.count()returns total records in the collection, not filtered counts -- there is no filtered count APIpeek()returns the firstlimititems (default 10) in insertion order, not by relevance -- useful for debugging, not querying$containsinwhereDocumentis case-sensitive -- searching for "Neural" will not match "neural"$regexinwhereDocumentuses full regex syntax but can be slow on large collections- Array metadata values (
string[],number[]) must be homogeneous -- mixing types within an array is rejected - Metadata keys are case-sensitive --
Categoryandcategoryare different fields - The
nResultsdefault is 10 if not specified inquery() - Multimodal embedding (images + text) is currently Python-only -- TypeScript support is not yet available
</red_flags>
<critical_reminders>
CRITICAL REMINDERS
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST install @chroma-core/default-embed alongside chromadb -- the default embedding function ships as a separate package since v3)
(You MUST access query results as nested arrays -- results.ids[0], results.documents[0] -- because Chroma batches queries and returns string[][] not string[])
(You MUST use the configuration parameter for HNSW settings -- the legacy metadata: { "hnsw:space": "cosine" } approach is deprecated)
(You MUST use flat metadata values only (string, number, boolean, typed arrays) -- nested objects are not supported and will be rejected)
Failure to follow these rules will cause embedding failures, incorrect result access, deprecated configuration warnings, and rejected metadata.
</critical_reminders>
Files (skills)
-
examples
-
core.md 14.7 KB
# Chroma -- Core Pattern Examples > Client setup, collection management, and fundamental operations (add, query, get, update, upsert, delete). Reference from [SKILL.md](../SKILL.md). **Related examples:** - [metadata-filtering.md](metadata-filtering.md) -- Filter operators and compound filters - [embedding-functions.md](embedding-functions.md) -- Default, OpenAI, custom embedding functions --- ## Client Initialization (HTTP) ```typescript import { ChromaClient } from "chromadb"; function createChromaClient(): ChromaClient { const chromaUrl = process.env.CHROMA_URL; if (!chromaUrl) { throw new Error("CHROMA_URL environment variable is required"); } return new ChromaClient({ path: chromaUrl }); } export { createChromaClient }; ``` **Why good:** Server URL from environment variable (never hardcoded), validation before construction, named export By default, `ChromaClient()` connects to `http://localhost:8000`. Explicit URL is preferred to avoid deployment-specific bugs. --- ## Cloud Client Initialization ```typescript import { CloudClient } from "chromadb"; function createCloudClient(): CloudClient { const apiKey = process.env.CHROMA_API_KEY; if (!apiKey) { throw new Error("CHROMA_API_KEY environment variable is required"); } return new CloudClient({ apiKey, tenant: process.env.CHROMA_TENANT ?? "default_tenant", database: process.env.CHROMA_DATABASE ?? "default_database", }); } export { createCloudClient }; ``` **Why good:** API key from environment, explicit tenant and database, named export --- ## Client with Token Authentication ```typescript import { ChromaClient } from "chromadb"; function createAuthenticatedClient(): ChromaClient { const chromaUrl = process.env.CHROMA_URL; const chromaToken = process.env.CHROMA_TOKEN; if (!chromaUrl || !chromaToken) { throw new Error( "CHROMA_URL and CHROMA_TOKEN environment variables required", ); } return new ChromaClient({ path: chromaUrl, auth: { provider: "token", credentials: chromaToken, tokenHeaderType: "X_CHROMA_TOKEN", }, }); } export { createAuthenticatedClient }; ``` **Why good:** Token from environment, explicit auth provider and header type --- ## Create a Collection with Cosine Distance ```typescript import { ChromaClient } from "chromadb"; const COLLECTION_NAME = "documents"; async function createCosineCollection(client: ChromaClient) { const collection = await client.createCollection({ name: COLLECTION_NAME, configuration: { hnsw: { space: "cosine" }, }, }); return collection; } export { createCosineCollection }; ``` **Why good:** Named constant for collection name, `cosine` metric via `configuration` parameter (not deprecated metadata approach) ```typescript // Bad Example -- deprecated HNSW config approach const collection = await client.createCollection({ name: "docs", metadata: { "hnsw:space": "cosine" }, // DEPRECATED in v3 }); ``` **Why bad:** `metadata` prefix for HNSW settings is deprecated; use `configuration: { hnsw: { space } }` instead --- ## Get or Create Collection (Idempotent) ```typescript import { ChromaClient } from "chromadb"; const COLLECTION_NAME = "articles"; async function getArticlesCollection(client: ChromaClient) { const collection = await client.getOrCreateCollection({ name: COLLECTION_NAME, configuration: { hnsw: { space: "cosine" }, }, }); return collection; } export { getArticlesCollection }; ``` **Why good:** Idempotent -- safe to call repeatedly, creates on first call, returns existing on subsequent calls **Gotcha:** If the collection already exists, `getOrCreateCollection` ignores the `configuration` parameter. It does NOT update the existing collection's configuration. --- ## Add Documents with Automatic Embedding ```typescript import type { ChromaClient } from "chromadb"; interface ArticleMetadata { title: string; category: string; year: number; } const COLLECTION_NAME = "articles"; async function addArticles( client: ChromaClient, articles: Array<{ id: string; text: string; metadata: ArticleMetadata }>, ): Promise<void> { const collection = await client.getOrCreateCollection({ name: COLLECTION_NAME, }); await collection.add({ ids: articles.map((a) => a.id), documents: articles.map((a) => a.text), metadatas: articles.map((a) => a.metadata), }); } export { addArticles }; export type { ArticleMetadata }; ``` **Why good:** Typed metadata interface, documents auto-embedded by collection's embedding function, columnar format (parallel arrays) ```typescript // Bad Example -- invalid metadata and missing content await collection.add({ ids: ["doc-1"], metadatas: [ { title: "Guide", author: { name: "Alice", org: "Acme" }, // INVALID: nested object }, ], // Missing documents AND embeddings -- at least one required }); ``` **Why bad:** Nested metadata objects are rejected, either `documents` or `embeddings` must be provided --- ## Add Documents with Pre-Computed Embeddings ```typescript import type { ChromaClient } from "chromadb"; const COLLECTION_NAME = "articles"; async function addWithEmbeddings( client: ChromaClient, records: Array<{ id: string; embedding: number[]; text: string; metadata: Record<string, string | number>; }>, ): Promise<void> { const collection = await client.getOrCreateCollection({ name: COLLECTION_NAME, configuration: { hnsw: { space: "cosine" } }, }); await collection.add({ ids: records.map((r) => r.id), embeddings: records.map((r) => r.embedding), documents: records.map((r) => r.text), metadatas: records.map((r) => r.metadata), }); } export { addWithEmbeddings }; ``` **Why good:** Pre-computed embeddings bypass the collection's embedding function, documents stored for retrieval, metadata for filtering **Note:** When both `embeddings` and `documents` are provided, Chroma stores the documents but uses the provided embeddings (does not re-embed). --- ## Query by Text Similarity ```typescript import type { ChromaClient } from "chromadb"; const N_RESULTS = 10; const COLLECTION_NAME = "articles"; interface SearchResult { id: string; document: string | null; distance: number | null; metadata: Record<string, string | number | boolean> | null; } async function searchArticles( client: ChromaClient, queryText: string, ): Promise<SearchResult[]> { const collection = await client.getCollection({ name: COLLECTION_NAME }); const results = await collection.query({ queryTexts: [queryText], nResults: N_RESULTS, include: ["documents", "metadatas", "distances"], }); // Results are nested arrays -- [0] for the first query return results.ids[0].map((id, i) => ({ id, document: results.documents?.[0]?.[i] ?? null, distance: results.distances?.[0]?.[i] ?? null, metadata: results.metadatas?.[0]?.[i] ?? null, })); } export { searchArticles }; ``` **Why good:** Named constant for nResults, explicit include, correct nested array access `[0]`, handles nullable fields, typed return value ```typescript // Bad Example -- incorrect result access const results = await collection.query({ queryTexts: ["query"], nResults: 5, }); // BUG: results.ids is string[][], not string[] console.log(results.ids[0]); // This is correct console.log(results.ids.length); // This is the number of QUERIES, not results! ``` **Why bad:** Treating `results.ids` as a flat array leads to bugs; must use `results.ids[0]` for single-query results --- ## Query with Multiple Queries (Batched) ```typescript const N_RESULTS = 5; async function batchQuery( collection: Awaited<ReturnType<ChromaClient["getCollection"]>>, queries: string[], ) { const results = await collection.query({ queryTexts: queries, // Multiple queries in one call nResults: N_RESULTS, }); // Each query gets its own result set return queries.map((query, queryIndex) => ({ query, results: results.ids[queryIndex].map((id, i) => ({ id, distance: results.distances?.[queryIndex]?.[i] ?? null, })), })); } export { batchQuery }; ``` **Why good:** Multiple queries in single API call, correct nested array indexing per query, efficient batching --- ## Query with Pre-Computed Embeddings ```typescript const N_RESULTS = 10; const results = await collection.query({ queryEmbeddings: [queryVector], // Pre-computed embedding nResults: N_RESULTS, include: ["documents", "metadatas", "distances"], }); ``` **Why good:** Bypasses collection's embedding function, useful when you manage your own embedding pipeline **Note:** Do not pass both `queryTexts` and `queryEmbeddings` -- use one or the other. --- ## Get Documents by ID ```typescript async function getDocuments( collection: Awaited<ReturnType<ChromaClient["getCollection"]>>, ids: string[], ) { const results = await collection.get({ ids, include: ["documents", "metadatas"], }); // get() returns flat arrays (not nested like query()) return results.ids.map((id, i) => ({ id, document: results.documents?.[i] ?? null, metadata: results.metadatas?.[i] ?? null, })); } export { getDocuments }; ``` **Why good:** `get()` returns flat arrays (unlike `query()`), explicit include, handles nullable fields **Gotcha:** `get()` silently omits IDs that don't exist. If you request 5 IDs and 2 don't exist, you get 3 results with no error. --- ## Get with Pagination ```typescript const PAGE_SIZE = 50; async function getAllDocuments( collection: Awaited<ReturnType<ChromaClient["getCollection"]>>, ): Promise<Array<{ id: string; document: string | null }>> { const allResults: Array<{ id: string; document: string | null }> = []; let offset = 0; while (true) { const page = await collection.get({ limit: PAGE_SIZE, offset, include: ["documents"], }); if (page.ids.length === 0) { break; } allResults.push( ...page.ids.map((id, i) => ({ id, document: page.documents?.[i] ?? null, })), ); offset += page.ids.length; } return allResults; } export { getAllDocuments }; ``` **Why good:** Paginated retrieval for large collections, named constant for page size, terminates when no more results --- ## Update Existing Records ```typescript async function updateDocumentMetadata( collection: Awaited<ReturnType<ChromaClient["getCollection"]>>, id: string, newCategory: string, ): Promise<void> { await collection.update({ ids: [id], metadatas: [{ category: newCategory }], }); } export { updateDocumentMetadata }; ``` **Why good:** Updates specific fields without re-embedding (if only metadata changes) **Gotcha:** `update()` silently ignores non-existent IDs. If the ID doesn't exist, nothing happens and no error is thrown. Use `upsert()` for create-or-update semantics. **Note:** If `documents` are provided in `update()`, Chroma re-embeds them using the collection's embedding function. --- ## Upsert (Create or Update) ```typescript async function upsertDocuments( collection: Awaited<ReturnType<ChromaClient["getCollection"]>>, docs: Array<{ id: string; text: string; metadata: Record<string, string | number | boolean>; }>, ): Promise<void> { await collection.upsert({ ids: docs.map((d) => d.id), documents: docs.map((d) => d.text), metadatas: docs.map((d) => d.metadata), }); } export { upsertDocuments }; ``` **Why good:** Idempotent -- creates new records or updates existing ones, safer than `add()` for pipelines that may run multiple times --- ## Delete Records ```typescript // Delete by IDs async function deleteByIds( collection: Awaited<ReturnType<ChromaClient["getCollection"]>>, ids: string[], ): Promise<void> { await collection.delete({ ids }); } // Delete by metadata filter async function deleteByCategory( collection: Awaited<ReturnType<ChromaClient["getCollection"]>>, category: string, ): Promise<void> { await collection.delete({ where: { category: { $eq: category } }, }); } export { deleteByIds, deleteByCategory }; ``` **Why good:** Two deletion patterns (by ID, by filter), named exports **Gotcha:** `delete()` with no arguments is a no-op -- it does not delete everything. Pass `ids` or `where` to target specific records. --- ## Delete Collection ```typescript async function removeCollection( client: ChromaClient, collectionName: string, ): Promise<void> { await client.deleteCollection({ name: collectionName }); } export { removeCollection }; ``` **Why good:** Clean deletion of entire collection including all vectors, metadata, and configuration **Note:** This is irreversible. To recreate with different HNSW settings, delete and recreate the collection. --- ## Collection Health Check ```typescript import { ChromaClient } from "chromadb"; async function healthCheck(client: ChromaClient): Promise<{ serverAlive: boolean; version: string | null; collectionCount: number; }> { try { await client.heartbeat(); const version = await client.version(); const count = await client.countCollections(); return { serverAlive: true, version, collectionCount: count }; } catch { return { serverAlive: false, version: null, collectionCount: 0 }; } } export { healthCheck }; ``` **Why good:** Uses `heartbeat()` for connectivity check, `version()` for compatibility verification, handles connection failures gracefully --- ## Peek at Collection Contents ```typescript const PEEK_LIMIT = 5; async function peekCollection( collection: Awaited<ReturnType<ChromaClient["getCollection"]>>, ) { const preview = await collection.peek({ limit: PEEK_LIMIT }); console.log(`Total records: ${await collection.count()}`); for (let i = 0; i < preview.ids.length; i++) { console.log(preview.ids[i], preview.documents?.[i]?.slice(0, 100)); } } export { peekCollection }; ``` **Why good:** Quick debugging tool, `peek()` returns first N items in insertion order, `count()` for total size **Note:** `peek()` is for debugging, not querying. Results are in insertion order, not relevance order. --- ## Result Iteration with `.rows()` ```typescript // Get results -- flat iteration const getResults = await collection.get({ include: ["documents", "metadatas"], }); for (const row of getResults.rows()) { console.log(row.id, row.document, row.metadata); } // Query results -- nested iteration (one batch per query) const queryResults = await collection.query({ queryTexts: ["search term"], nResults: 5, }); for (const batch of queryResults.rows()) { for (const row of batch) { console.log(row.id, row.document, row.metadata, row.distance); } } ``` **Why good:** `.rows()` provides a cleaner iteration API than manual index access, handles the nested/flat difference between query and get results --- _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_ -
embedding-functions.md 6.5 KB
# Chroma -- Embedding Function Examples > Default, OpenAI, and custom embedding function configuration. See [core.md](core.md) for basic operations and [reference.md](../reference.md) for the full provider list. **Related examples:** - [core.md](core.md) -- Client setup, add, query - [metadata-filtering.md](metadata-filtering.md) -- Filter operators --- ## Default Embedding Function The default embedding function uses `all-MiniLM-L6-v2` (Sentence Transformers), running locally. Requires a separate package since v3. ```bash npm install chromadb @chroma-core/default-embed ``` ```typescript import { ChromaClient } from "chromadb"; const COLLECTION_NAME = "articles"; async function createWithDefaultEmbedding(client: ChromaClient) { // No embeddingFunction parameter needed -- uses default automatically const collection = await client.createCollection({ name: COLLECTION_NAME, configuration: { hnsw: { space: "cosine" } }, }); // Documents are embedded automatically using all-MiniLM-L6-v2 await collection.add({ ids: ["doc-1", "doc-2"], documents: [ "Introduction to machine learning", "Advanced neural network architectures", ], }); return collection; } export { createWithDefaultEmbedding }; ``` **Why good:** Simplest setup -- just pass documents, embedding happens automatically. No API keys needed, runs locally. **Note:** `@chroma-core/default-embed` must be installed even though it's not explicitly imported. Chroma resolves it at runtime. --- ## OpenAI Embedding Function Use OpenAI's `text-embedding-3-small` for higher-quality embeddings. ```bash npm install @chroma-core/openai ``` ```typescript import { ChromaClient } from "chromadb"; import { OpenAIEmbeddingFunction } from "@chroma-core/openai"; const COLLECTION_NAME = "knowledge-base"; async function createWithOpenAI(client: ChromaClient) { const openAiKey = process.env.OPENAI_API_KEY; if (!openAiKey) { throw new Error("OPENAI_API_KEY environment variable is required"); } const embeddingFunction = new OpenAIEmbeddingFunction({ modelName: "text-embedding-3-small", }); const collection = await client.createCollection({ name: COLLECTION_NAME, embeddingFunction, configuration: { hnsw: { space: "cosine" } }, }); return collection; } export { createWithOpenAI }; ``` **Why good:** API key from environment (reads `OPENAI_API_KEY` automatically), explicit model name, cosine metric matches OpenAI's normalized embeddings **Note:** When retrieving an existing collection with `getCollection`, you must pass the same `embeddingFunction` so that queries are embedded correctly: ```typescript const collection = await client.getCollection({ name: COLLECTION_NAME, embeddingFunction: new OpenAIEmbeddingFunction({ modelName: "text-embedding-3-small", }), }); ``` --- ## Custom Embedding Function Implement the `EmbeddingFunction` interface for custom models. ```typescript import type { EmbeddingFunction } from "chromadb"; class CustomEmbeddingFunction implements EmbeddingFunction { public readonly name = "custom-embedder"; private readonly model: string; constructor(args: { model: string }) { this.model = args.model; } async generate(texts: string[]): Promise<number[][]> { // Replace with your embedding logic // Must return one embedding vector per input text const embeddings = await Promise.all( texts.map(async (text) => { // Call your embedding API or model return new Array(384).fill(0).map(() => Math.random()); }), ); return embeddings; } getConfig(): Record<string, unknown> { return { model: this.model }; } validateConfigUpdate(config: Record<string, unknown>): void { if ("model" in config) { throw new Error("Model cannot be updated after creation"); } } static buildFromConfig( config: Record<string, unknown>, ): CustomEmbeddingFunction { return new CustomEmbeddingFunction({ model: config.model as string }); } } export { CustomEmbeddingFunction }; ``` **Why good:** Implements required interface methods, immutable model config, `generate()` returns one vector per input --- ## Using Embedding Functions Directly You can invoke embedding functions independently for debugging or pre-computation. ```typescript import { DefaultEmbeddingFunction } from "@chroma-core/default-embed"; async function debugEmbeddings(): Promise<void> { const embeddingFn = new DefaultEmbeddingFunction(); // Generate embeddings outside of a collection const embeddings = await embeddingFn.generate([ "test document one", "test document two", ]); console.log("Dimension:", embeddings[0].length); console.log("First embedding (first 5 values):", embeddings[0].slice(0, 5)); } export { debugEmbeddings }; ``` **Why good:** Useful for verifying embedding dimensions, debugging embedding quality, or pre-computing embeddings for external use --- ## Collection with Ollama Embeddings Use locally-hosted models via Ollama. ```bash npm install @chroma-core/ollama ``` ```typescript import { ChromaClient } from "chromadb"; import { OllamaEmbeddingFunction } from "@chroma-core/ollama"; const COLLECTION_NAME = "local-knowledge"; async function createWithOllama(client: ChromaClient) { const embeddingFunction = new OllamaEmbeddingFunction({ url: process.env.OLLAMA_URL ?? "http://localhost:11434", model: "nomic-embed-text", }); const collection = await client.createCollection({ name: COLLECTION_NAME, embeddingFunction, configuration: { hnsw: { space: "cosine" } }, }); return collection; } export { createWithOllama }; ``` **Why good:** Fully local embedding pipeline (no API keys), configurable Ollama URL, explicit model name --- ## Embedding Function Gotchas **Dimension mismatch:** If you change embedding functions on an existing collection (e.g., switch from default 384-dim to OpenAI 1536-dim), queries will fail with dimension mismatch errors. Delete and recreate the collection with the new function. **Consistency requirement:** Always pass the same embedding function to `getCollection()` / `getOrCreateCollection()` as was used at creation. If you omit it, Chroma uses the default, which may produce different-dimension embeddings. **Query vs add:** The same embedding function is used for both `add()` (document embedding) and `query()` (query embedding). Do not mix different embedding models between add and query operations on the same collection. --- _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_ -
metadata-filtering.md 8.8 KB
# Chroma -- Metadata Filtering Examples > Filter operators, compound filters, document content filters, and where/whereDocument patterns. See [core.md](core.md) for basic query patterns and [reference.md](../reference.md) for the operator table. **Related examples:** - [core.md](core.md) -- Client setup, basic query with filters - [embedding-functions.md](embedding-functions.md) -- Embedding function configuration --- ## Comparison Operators ```typescript const N_RESULTS = 10; // Exact match (explicit) const drama = await collection.query({ queryTexts: ["dramatic story"], nResults: N_RESULTS, where: { genre: { $eq: "drama" } }, }); // Exact match (shorthand -- equivalent to $eq) const dramaShorthand = await collection.query({ queryTexts: ["dramatic story"], nResults: N_RESULTS, where: { genre: "drama" }, }); // Not equal const notComedy = await collection.query({ queryTexts: ["funny story"], nResults: N_RESULTS, where: { genre: { $ne: "comedy" } }, }); // Numeric range (greater than or equal) const recent = await collection.query({ queryTexts: ["recent articles"], nResults: N_RESULTS, where: { year: { $gte: 2023 } }, }); // Numeric range (less than) const older = await collection.query({ queryTexts: ["historical articles"], nResults: N_RESULTS, where: { year: { $lt: 2020 } }, }); ``` **Why good:** Each operator targets specific types, shorthand `{ genre: "drama" }` is equivalent to `{ genre: { $eq: "drama" } }` --- ## Set Membership Operators ```typescript const N_RESULTS = 10; // $in -- value in set const selectedGenres = await collection.query({ queryTexts: ["search query"], nResults: N_RESULTS, where: { genre: { $in: ["drama", "action", "thriller"] } }, }); // $nin -- value not in set const excludedGenres = await collection.query({ queryTexts: ["search query"], nResults: N_RESULTS, where: { genre: { $nin: ["horror", "comedy"] } }, }); ``` **Why good:** `$in` replaces multiple `$or` / `$eq` clauses, cleaner and more efficient --- ## Array Metadata Operators For metadata fields that are arrays (e.g., tags, authors). Requires Chroma >= 1.5.0. ```typescript const N_RESULTS = 10; // $contains -- array includes specific value const chenPapers = await collection.query({ queryTexts: ["machine learning"], nResults: N_RESULTS, where: { authors: { $contains: "Chen" } }, }); // $not_contains -- array excludes specific value const withoutDraft = await collection.query({ queryTexts: ["published papers"], nResults: N_RESULTS, where: { tags: { $not_contains: "draft" } }, }); ``` **Why good:** `$contains` checks if an array metadata field includes a scalar value, works with string/int/float/boolean arrays **Gotcha:** `$contains` and `$not_contains` only work on array metadata fields, not on scalar values. For scalar equality, use `$eq`. --- ## Compound Filters with $and / $or ```typescript const N_RESULTS = 10; // AND: all conditions must match const filteredResults = await collection.query({ queryTexts: ["search query"], nResults: N_RESULTS, where: { $and: [ { genre: { $eq: "drama" } }, { year: { $gte: 2020 } }, { rating: { $gt: 7.5 } }, ], }, }); // OR: any condition matches const broadResults = await collection.query({ queryTexts: ["search query"], nResults: N_RESULTS, where: { $or: [{ genre: { $eq: "drama" } }, { genre: { $eq: "documentary" } }], }, }); // Nested AND + OR const complexResults = await collection.query({ queryTexts: ["search query"], nResults: N_RESULTS, where: { $and: [ { year: { $gte: 2020 } }, { $or: [{ genre: { $eq: "drama" } }, { genre: { $eq: "thriller" } }], }, ], }, }); ``` **Why good:** Nested logical operators for complex queries, `$or` inside `$and` for flexible filtering --- ## Document Content Filtering (whereDocument) Filter by document text content using `$contains`, `$not_contains`, `$regex`, and `$not_regex`. ```typescript const N_RESULTS = 10; // $contains -- case-sensitive substring match const containsNeural = await collection.query({ queryTexts: ["deep learning"], nResults: N_RESULTS, whereDocument: { $contains: "neural network" }, }); // $not_contains -- exclude documents with substring const excludeDeprecated = await collection.query({ queryTexts: ["current practices"], nResults: N_RESULTS, whereDocument: { $not_contains: "deprecated" }, }); // $regex -- pattern matching const regexMatch = await collection.query({ queryTexts: ["technical papers"], nResults: N_RESULTS, whereDocument: { $regex: "deep\\s+learning" }, }); // $not_regex -- exclude by pattern const excludePattern = await collection.query({ queryTexts: ["final papers"], nResults: N_RESULTS, whereDocument: { $not_regex: "draft.*v[0-9]" }, }); ``` **Why good:** `$contains` for simple substring search, `$regex` for pattern matching, `$not_*` variants for exclusion **Gotcha:** `$contains` is case-sensitive -- "Neural" will not match "neural". Use `$regex` with case-insensitive patterns if needed. --- ## Compound Document Filters ```typescript const N_RESULTS = 10; // AND -- both substrings must appear in document const bothTerms = await collection.query({ queryTexts: ["AI research"], nResults: N_RESULTS, whereDocument: { $and: [{ $contains: "transformer" }, { $contains: "attention" }], }, }); // OR -- either substring matches const eitherTerm = await collection.query({ queryTexts: ["AI research"], nResults: N_RESULTS, whereDocument: { $or: [{ $contains: "transformer" }, { $contains: "recurrent" }], }, }); ``` **Why good:** Compound document filters for multi-term matching --- ## Combined Metadata + Document Filters Use `where` and `whereDocument` together for maximum filtering precision. ```typescript const N_RESULTS = 10; const results = await collection.query({ queryTexts: ["machine learning tutorial"], nResults: N_RESULTS, where: { $and: [{ category: { $eq: "tutorial" } }, { year: { $gte: 2023 } }], }, whereDocument: { $contains: "neural network" }, include: ["documents", "metadatas", "distances"], }); // Access results (nested arrays for query) for (let i = 0; i < results.ids[0].length; i++) { console.log( results.ids[0][i], results.distances?.[0]?.[i], results.metadatas?.[0]?.[i], ); } ``` **Why good:** Combined metadata + document filters for precise retrieval, correct nested array access --- ## Filtering with get() (No Vector Query) Use `get()` with filters to retrieve documents without similarity search. ```typescript const PAGE_SIZE = 50; // Get by metadata filter const tutorials = await collection.get({ where: { category: { $eq: "tutorial" } }, limit: PAGE_SIZE, include: ["documents", "metadatas"], }); // Get by document content filter const withCode = await collection.get({ whereDocument: { $contains: "function" }, limit: PAGE_SIZE, include: ["documents"], }); // Combined metadata + document filter const recentTutorialsWithCode = await collection.get({ where: { $and: [{ category: { $eq: "tutorial" } }, { year: { $gte: 2023 } }], }, whereDocument: { $contains: "function" }, limit: PAGE_SIZE, include: ["documents", "metadatas"], }); ``` **Why good:** `get()` does not require a query vector, useful for data management and non-similarity-based retrieval --- ## Delete by Metadata Filter ```typescript // Delete all records matching a filter await collection.delete({ where: { $and: [{ status: { $eq: "archived" } }, { year: { $lt: 2020 } }], }, }); ``` **Why good:** Bulk deletion without knowing record IDs, compound filter for precise targeting --- ## Common Filter Mistakes ```typescript // Bad: $gt on string value (only works on numbers) where: { title: { $gt: "A"; } } // Fix: use $eq, $ne, $in, or $nin for strings // Bad: Nested object in metadata metadatas: [{ author: { name: "Alice", org: "Acme" } }]; // Fix: flatten to top-level keys metadatas: [{ authorName: "Alice", authorOrg: "Acme" }]; // Bad: Mixed types in array metadata metadatas: [{ tags: ["draft", 2024, true] }]; // INVALID: mixed types // Fix: arrays must be homogeneous metadatas: [{ tags: ["draft", "2024"] }]; // All strings // Bad: $contains on scalar metadata field where: { category: { $contains: "tut"; } } // INVALID: $contains is for arrays // Fix: use $eq for scalar fields, $contains for array fields where: { category: { $eq: "tutorial"; } } // Bad: Case-sensitive surprise with whereDocument whereDocument: { $contains: "Neural"; } // Won't match "neural" // Fix: use $regex for case-insensitive matching whereDocument: { $regex: "(?i)neural"; } ``` **Why bad:** Each mistake causes either a rejection or silently empty results; type-restricted operators and case sensitivity are common sources of "no results" bugs --- _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_
-
-
reference.md 15 KB
# Chroma Quick Reference > API reference, filter operators, include options, limits, and production checklist. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples. --- ## ChromaClient Methods | Method | Description | Returns | | ----------------------------------------------- | ----------------------------------------------------- | ----------------------- | | `new ChromaClient({ path })` | Create HTTP client (default: `http://localhost:8000`) | `ChromaClient` | | `new CloudClient({ apiKey, tenant, database })` | Create Chroma Cloud client | `CloudClient` | | `client.heartbeat()` | Check server connectivity | `Promise<number>` | | `client.version()` | Get server version | `Promise<string>` | | `client.createCollection(options)` | Create a new collection | `Promise<Collection>` | | `client.getCollection(options)` | Get an existing collection | `Promise<Collection>` | | `client.getOrCreateCollection(options)` | Get or create a collection | `Promise<Collection>` | | `client.deleteCollection({ name })` | Delete a collection | `Promise<void>` | | `client.listCollections({ limit?, offset? })` | List collections (paginated) | `Promise<Collection[]>` | | `client.countCollections()` | Count total collections | `Promise<number>` | | `client.reset()` | Reset entire database (dev only) | `Promise<void>` | ## Collection Methods | Method | Description | Returns | | ---------------------------- | ------------------------------- | ---------------------- | | `collection.add(options)` | Add documents/embeddings | `Promise<void>` | | `collection.query(options)` | Find similar documents | `Promise<QueryResult>` | | `collection.get(options?)` | Get by IDs or filter | `Promise<GetResult>` | | `collection.update(options)` | Update existing records | `Promise<void>` | | `collection.upsert(options)` | Insert or update records | `Promise<void>` | | `collection.delete(options)` | Delete by IDs or filter | `Promise<void>` | | `collection.peek(options?)` | Preview first N items | `Promise<GetResult>` | | `collection.count()` | Count records in collection | `Promise<number>` | | `collection.modify(options)` | Update collection name/metadata | `Promise<void>` | --- ## Method Parameters ### add / upsert | Parameter | Type | Required | Description | | ------------ | --------------------------------- | --------------------------- | ---------------------------------- | | `ids` | `string[]` | Yes | Unique identifiers for each record | | `documents` | `string[]` | One of documents/embeddings | Text documents (auto-embedded) | | `embeddings` | `number[][]` | One of documents/embeddings | Pre-computed embedding vectors | | `metadatas` | `Record<string, MetadataValue>[]` | No | Metadata for filtering | | `uris` | `string[]` | No | URIs for data loaders | ### query | Parameter | Type | Required | Description | | ----------------- | --------------- | ----------------------- | ------------------------------- | | `queryTexts` | `string[]` | One of texts/embeddings | Text queries (auto-embedded) | | `queryEmbeddings` | `number[][]` | One of texts/embeddings | Pre-computed query embeddings | | `nResults` | `number` | No | Number of results (default: 10) | | `where` | `Where` | No | Metadata filter | | `whereDocument` | `WhereDocument` | No | Document content filter | | `include` | `Include[]` | No | Fields to include in response | ### get | Parameter | Type | Required | Description | | --------------- | --------------- | -------- | ----------------------------- | | `ids` | `string[]` | No | Specific IDs to retrieve | | `where` | `Where` | No | Metadata filter | | `whereDocument` | `WhereDocument` | No | Document content filter | | `limit` | `number` | No | Max results to return | | `offset` | `number` | No | Skip first N results | | `include` | `Include[]` | No | Fields to include in response | ### delete | Parameter | Type | Required | Description | | --------------- | --------------- | -------- | ------------------------------------ | | `ids` | `string[]` | No | Specific IDs to delete | | `where` | `Where` | No | Metadata filter for deletion | | `whereDocument` | `WhereDocument` | No | Document content filter for deletion | --- ## Metadata Filter Operators (`where`) | Operator | Description | Supported Types | Example | | --------------- | --------------------- | ----------------------- | ----------------------------------------- | | `$eq` | Equal to (default) | string, number, boolean | `{ genre: { $eq: "drama" } }` | | `$ne` | Not equal to | string, number, boolean | `{ genre: { $ne: "comedy" } }` | | `$gt` | Greater than | number only | `{ year: { $gt: 2020 } }` | | `$gte` | Greater than or equal | number only | `{ year: { $gte: 2020 } }` | | `$lt` | Less than | number only | `{ year: { $lt: 2020 } }` | | `$lte` | Less than or equal | number only | `{ year: { $lte: 2020 } }` | | `$in` | In set | string, number, boolean | `{ genre: { $in: ["drama", "action"] } }` | | `$nin` | Not in set | string, number, boolean | `{ genre: { $nin: ["horror"] } }` | | `$contains` | Array contains value | typed arrays | `{ authors: { $contains: "Chen" } }` | | `$not_contains` | Array excludes value | typed arrays | `{ tags: { $not_contains: "draft" } }` | | `$and` | Logical AND | filter[] | `{ $and: [filter1, filter2] }` | | `$or` | Logical OR | filter[] | `{ $or: [filter1, filter2] }` | **Rules:** - Direct equality shorthand: `{ field: "value" }` equals `{ field: { $eq: "value" } }` - `$gt`, `$gte`, `$lt`, `$lte` work on numeric values only - `$contains` / `$not_contains` work on array metadata fields only (Chroma >= 1.5.0) - Array metadata must be homogeneous (all same type: string, int, float, or boolean) - Nested objects are not supported in metadata -- flat key-value pairs only - Metadata keys are case-sensitive --- ## Document Content Filter Operators (`whereDocument`) | Operator | Description | Example | | --------------- | --------------------------------- | ------------------------------------------------ | | `$contains` | Case-sensitive substring match | `{ $contains: "neural network" }` | | `$not_contains` | Excludes documents with substring | `{ $not_contains: "deprecated" }` | | `$regex` | Regex pattern match | `{ $regex: "deep\\s+learning" }` | | `$not_regex` | Excludes documents matching regex | `{ $not_regex: "draft.*v[0-9]" }` | | `$and` | Logical AND for document filters | `{ $and: [{$contains: "A"}, {$contains: "B"}] }` | | `$or` | Logical OR for document filters | `{ $or: [{$contains: "A"}, {$contains: "B"}] }` | --- ## Include Options | Value | In `query()` default? | In `get()` default? | Description | | -------------- | --------------------- | ------------------- | --------------------------------- | | `"documents"` | Yes | Yes | Document text content | | `"metadatas"` | Yes | Yes | Metadata key-value pairs | | `"distances"` | Yes | No | Similarity distances (query only) | | `"embeddings"` | No | No | Raw embedding vectors | | `"uris"` | No | No | Data loader URIs | --- ## Return Type Structures ### QueryResult (from `query()`) ```typescript interface QueryResult { ids: string[][]; // Nested: [query1_ids, query2_ids, ...] documents: (string | null)[][]; metadatas: (Record<string, MetadataValue> | null)[][]; embeddings: (number[] | null)[][] | null; distances: (number | null)[][] | null; uris: (string | null)[][] | null; include: Include[]; } ``` ### GetResult (from `get()`, `peek()`) ```typescript interface GetResult { ids: string[]; // Flat array documents: (string | null)[]; metadatas: (Record<string, MetadataValue> | null)[]; embeddings: number[][] | null; uris: (string | null)[] | null; include: Include[]; } ``` **Critical difference:** `query()` returns nested arrays (batched queries), `get()` returns flat arrays. --- ## HNSW Configuration Parameters | Parameter | Default | Modifiable after creation? | Description | | ----------------- | --------- | -------------------------- | --------------------------------------- | | `space` | `"l2"` | No | Distance function: `l2`, `cosine`, `ip` | | `ef_construction` | `100` | No | Candidate list size during index build | | `ef_search` | `100` | Yes | Candidate list during queries | | `max_neighbors` | `16` | No | Max connections per graph node | | `num_threads` | CPU cores | Yes | Threads for operations | | `batch_size` | `100` | Yes | Vectors per batch | | `sync_threshold` | `1000` | Yes | Index sync trigger | | `resize_factor` | `1.2` | Yes | Growth multiplier on resize | --- ## Distance Metrics | Metric | Description | When to Use | | -------- | ------------------------------------ | --------------------------------------------------- | | `l2` | Euclidean distance squared (default) | Raw feature vectors where absolute distance matters | | `cosine` | Cosine similarity (normalized) | Most embedding models (recommended for text) | | `ip` | Inner product | Pre-normalized embeddings, dot product similarity | --- ## Supported Metadata Types | Type | Example Value | Filter Support | | -------------- | --------------- | ---------------------------- | | String | `"drama"` | All operators | | Number (int) | `2024` | All operators | | Number (float) | `0.95` | All operators | | Boolean | `true` | `$eq`, `$ne`, `$in`, `$nin` | | String array | `["a", "b"]` | `$contains`, `$not_contains` | | Int array | `[1, 2, 3]` | `$contains`, `$not_contains` | | Float array | `[0.1, 0.2]` | `$contains`, `$not_contains` | | Boolean array | `[true, false]` | `$contains`, `$not_contains` | **Constraints:** Arrays must be homogeneous (all elements same type). Empty arrays and nested arrays are not allowed. Nested objects are not supported. --- ## Embedding Function Packages | Package | Model | Use Case | | --------------------------------- | ---------------------- | -------------------------------------- | | `@chroma-core/default-embed` | all-MiniLM-L6-v2 | Local, English text, quick prototyping | | `@chroma-core/openai` | text-embedding-3-small | High-quality, production use | | `@chroma-core/cohere` | Cohere embed models | Multilingual support | | `@chroma-core/google-gemini` | Gemini embeddings | Google ecosystem | | `@chroma-core/ollama` | Ollama models | Self-hosted, local LLMs | | `@chroma-core/huggingface-server` | HF models | Custom HF models | | `@chroma-core/voyageai` | Voyage AI models | Specialized embeddings | | `@chroma-core/jina` | Jina AI models | Long-context embeddings | | `@chroma-core/all` | All providers | Install everything | --- ## Production Checklist ### Security - [ ] Chroma server behind authentication (token or basic auth) - [ ] API tokens stored in environment variables, not in code - [ ] `client.reset()` disabled in production (dangerous -- deletes everything) ### Collection Configuration - [ ] Distance metric set explicitly (`cosine` for text embeddings) - [ ] Embedding function configured per collection (not relying on defaults) - [ ] `ef_construction` increased for large collections (better recall, slower build) ### Data Management - [ ] Metadata values are flat (no nested objects) - [ ] IDs are unique and deterministic - [ ] `upsert` used instead of `add` for idempotent operations - [ ] Large datasets added in batches (avoid single massive `add` call) ### Query Optimization - [ ] `nResults` set to minimum needed (lower = faster) - [ ] `include` parameter used to exclude unnecessary fields - [ ] `where` filters use indexed metadata fields - [ ] `whereDocument` used sparingly (can be slow on large collections) ### Monitoring - [ ] `collection.count()` tracked for growth monitoring - [ ] `client.heartbeat()` used for health checks - [ ] Error handling for connection failures and timeouts --- _Full skill documentation: [SKILL.md](SKILL.md) | Examples: [examples/](examples/)_ -
SKILL.md 12.9 KB
--- name: api-vector-db-chroma description: Chroma vector database -- collection management, automatic embedding, metadata filtering, document storage, query patterns --- # Chroma Patterns > **Quick Guide:** Use `chromadb` (v3.x) with `@chroma-core/default-embed` for automatic embedding. Chroma auto-embeds documents if no embeddings are provided -- just pass `documents` and `ids` to `collection.add()`. Use `where` for metadata filtering and `whereDocument` for document content filtering (`$contains`, `$regex`). Default distance metric is `l2` (Euclidean); use `cosine` for most embedding models via `configuration: { hnsw: { space: "cosine" } }`. Query results return nested arrays (`ids: string[][]`) because queries are batched -- always access `results.ids[0]` for a single query. Include only the fields you need via the `include` parameter to reduce payload size. --- <critical_requirements> ## CRITICAL: Before Using This Skill > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST install `@chroma-core/default-embed` alongside `chromadb` -- the default embedding function ships as a separate package since v3)** **(You MUST access query results as nested arrays -- `results.ids[0]`, `results.documents[0]` -- because Chroma batches queries and returns `string[][]` not `string[]`)** **(You MUST use the `configuration` parameter for HNSW settings -- the legacy `metadata: { "hnsw:space": "cosine" }` approach is deprecated)** **(You MUST use flat metadata values only (string, number, boolean, typed arrays) -- nested objects are not supported and will be rejected)** </critical_requirements> --- ## Examples - [Core Patterns](examples/core.md) -- Client setup, collection management, add, query, get, update, upsert, delete - [Metadata Filtering](examples/metadata-filtering.md) -- Filter operators, compound filters, document content filters, whereDocument - [Embedding Functions](examples/embedding-functions.md) -- Default, OpenAI, custom embedding functions, provider packages **Additional resources:** - [reference.md](reference.md) -- API quick reference, filter operators, include options, limits, production checklist --- **Auto-detection:** Chroma, chromadb, ChromaClient, CloudClient, createCollection, getOrCreateCollection, collection.add, collection.query, collection.get, collection.upsert, queryTexts, queryEmbeddings, nResults, whereDocument, $contains, @chroma-core/default-embed, @chroma-core/openai, EmbeddingFunction, vector database, semantic search, embedding, RAG retrieval, hnsw:space **When to use:** - Semantic search over document embeddings (RAG retrieval) - Rapid prototyping with automatic embedding generation (no external embedding pipeline needed) - Metadata-filtered vector search with compound logical operators - Document content filtering with `$contains` and `$regex` - Local development with in-process or Docker-based Chroma server **Key patterns covered:** - Client setup (HTTP, Cloud, Docker) - Collection management (create, get, delete, configure HNSW) - Document CRUD with automatic embedding (add, query, get, update, upsert, delete) - Metadata filtering (`where`) with comparison, set, array, and logical operators - Document content filtering (`whereDocument`) with `$contains` and `$regex` - Embedding function configuration (default, OpenAI, custom) - Query result handling (nested array structure, include options) **When NOT to use:** - Full-text search with complex boolean ranking (use a dedicated search engine) - Relational data with joins and transactions (use a relational database) - Multi-modal image+text embeddings in TypeScript (currently Python-only in Chroma) - High-scale production with millions of vectors and strict SLAs (evaluate managed vector databases) --- <philosophy> ## Philosophy Chroma is a **lightweight, developer-friendly embedding database** designed for rapid prototyping and production RAG applications. The core principle: **pass documents in, get relevant results out -- Chroma handles embedding automatically.** **Core principles:** 1. **Documents first, vectors optional** -- Unlike most vector databases, Chroma can embed documents automatically using a configured embedding function. You never need to manage embeddings directly unless you want to. 2. **Collections are self-contained** -- Each collection has its own embedding function, distance metric, and HNSW configuration. No global index management needed. 3. **Metadata is for filtering, documents are for content** -- Use `where` for structured metadata filters and `whereDocument` for full-text content filters. Both can be combined in a single query. 4. **Batteries included** -- The default embedding function (`all-MiniLM-L6-v2` via `@chroma-core/default-embed`) works out of the box for English text. Swap to OpenAI, Cohere, or any provider with a single package change. 5. **Query results are batched** -- Chroma supports multiple queries in a single call. Results are always nested arrays (`string[][]`), even for single queries. Always access `[0]` for the first query's results. </philosophy> --- <patterns> ## Core Patterns ### Pattern 1: Client Initialization Always pass the server URL explicitly from an environment variable -- never rely on the implicit `http://localhost:8000` default. See [examples/core.md](examples/core.md) for HTTP, Cloud, and token-authenticated client examples. ```typescript const chromaUrl = process.env.CHROMA_URL; if (!chromaUrl) throw new Error("CHROMA_URL environment variable is required"); return new ChromaClient({ path: chromaUrl }); ``` --- ### Pattern 2: Collection with Distance Metric Use the `configuration` parameter for HNSW settings -- never the deprecated `metadata: { "hnsw:space": "cosine" }` approach. See [examples/core.md](examples/core.md). ```typescript const collection = await client.createCollection({ name: COLLECTION_NAME, configuration: { hnsw: { space: "cosine" } }, }); ``` --- ### Pattern 3: Add Documents with Automatic Embedding Pass `documents` and `ids` -- Chroma embeds automatically. Either `documents` or `embeddings` must be provided; metadata alone is insufficient. See [examples/core.md](examples/core.md). ```typescript await collection.add({ ids: articles.map((a) => a.id), documents: articles.map((a) => a.text), metadatas: articles.map((a) => ({ category: a.category })), }); ``` --- ### Pattern 4: Query with Metadata and Document Filters Combine `where` (metadata) and `whereDocument` (content) filters. Results are nested arrays -- always access `[0]` for single-query results. See [examples/metadata-filtering.md](examples/metadata-filtering.md) for all operators. ```typescript const results = await collection.query({ queryTexts: ["machine learning fundamentals"], nResults: N_RESULTS, where: { $and: [{ category: { $eq: "tutorial" } }, { year: { $gte: 2023 } }], }, whereDocument: { $contains: "neural network" }, include: ["documents", "metadatas", "distances"], }); // results.ids[0] -- nested array, access [0] for first query ``` --- ### Pattern 5: Get with Pagination Use `get()` with `limit`/`offset` for non-similarity retrieval. Unlike `query()`, `get()` returns flat arrays. See [examples/core.md](examples/core.md). ```typescript const results = await collection.get({ where: { category: { $eq: category } }, limit: PAGE_SIZE, offset, include: ["documents", "metadatas"], }); ``` --- ### Pattern 6: Upsert for Idempotent Operations Use `upsert()` instead of `add()` for create-or-update semantics -- safer for pipelines that run multiple times. See [examples/core.md](examples/core.md). ```typescript await collection.upsert({ ids: docs.map((d) => d.id), documents: docs.map((d) => d.text), metadatas: docs.map((d) => d.metadata), }); ``` </patterns> --- <decision_framework> ## Decision Framework ### Which Distance Metric? ``` Which distance metric should I use? |-- Using embeddings from a language model? -> cosine (normalized, most common) |-- Need dot product similarity? -> ip (inner product) |-- Comparing raw feature vectors? -> l2 (Euclidean, Chroma default) '-- Unsure? -> cosine (safe default for most embedding models) ``` ### Which Embedding Function? ``` Which embedding function should I use? |-- Quick prototyping, English text? -> @chroma-core/default-embed (all-MiniLM-L6-v2, runs locally) |-- Need high-quality embeddings? -> @chroma-core/openai (text-embedding-3-small) |-- Have your own embedding pipeline? -> Pass embeddings directly (skip embedding function) |-- Need custom model? -> Implement EmbeddingFunction interface '-- Want all providers? -> npm install @chroma-core/all ``` ### Where vs WhereDocument? ``` How should I filter results? |-- Structured attributes (category, year, status)? -> where (metadata filter) |-- Full-text content search? -> whereDocument ($contains, $regex) |-- Both? -> Combine where + whereDocument in same query '-- Need exact match on specific IDs? -> get({ ids: [...] }) ``` ### ChromaClient vs CloudClient? ``` Which client should I use? |-- Local development or self-hosted? -> ChromaClient({ path: "http://localhost:8000" }) |-- Chroma Cloud (managed)? -> CloudClient({ apiKey, tenant, database }) |-- Docker deployment? -> ChromaClient with Docker host URL '-- Testing? -> ChromaClient against local Docker container ``` </decision_framework> --- <red_flags> ## RED FLAGS **High Priority Issues:** - Accessing query results as flat arrays instead of nested -- `results.ids` is `string[][]`, not `string[]`; always use `results.ids[0]` for single-query results - Missing `@chroma-core/default-embed` package -- since v3, the default embedding function ships separately; `npm install chromadb @chroma-core/default-embed` - Using deprecated `metadata: { "hnsw:space": "cosine" }` for HNSW config -- use `configuration: { hnsw: { space: "cosine" } }` instead - Nested objects in metadata -- Chroma only supports flat key-value metadata; nested objects are rejected **Medium Priority Issues:** - Not specifying `include` in queries -- default includes vary (`query` returns documents, metadatas, distances; `get` returns documents, metadatas); explicitly set `include` for clarity and to control payload size - Using `l2` (default) when `cosine` is appropriate -- most embedding models are normalized for cosine similarity; `l2` may produce worse results - Calling `add()` without `documents` or `embeddings` -- at least one must be provided; metadata alone is insufficient - Not handling empty results -- `results.ids[0]` may be an empty array; check length before processing **Common Mistakes:** - Passing `queryEmbeddings` AND `queryTexts` together -- use one or the other, not both - Expecting `update()` to create missing records -- `update()` silently ignores non-existent IDs; use `upsert()` for create-or-update semantics - Calling `delete()` with no arguments -- deletes nothing (not everything); pass `ids` or `where` to target specific records - Using `$gt`/`$lt` on string metadata -- comparison operators only work on numeric values (int or float) **Gotchas & Edge Cases:** - HNSW configuration (`space`, `ef_construction`, `max_neighbors`) cannot be changed after collection creation -- you must delete and recreate the collection - `collection.count()` returns total records in the collection, not filtered counts -- there is no filtered count API - `peek()` returns the first `limit` items (default 10) in insertion order, not by relevance -- useful for debugging, not querying - `$contains` in `whereDocument` is case-sensitive -- searching for "Neural" will not match "neural" - `$regex` in `whereDocument` uses full regex syntax but can be slow on large collections - Array metadata values (`string[]`, `number[]`) must be homogeneous -- mixing types within an array is rejected - Metadata keys are case-sensitive -- `Category` and `category` are different fields - The `nResults` default is 10 if not specified in `query()` - Multimodal embedding (images + text) is currently Python-only -- TypeScript support is not yet available </red_flags> --- <critical_reminders> ## CRITICAL REMINDERS > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST install `@chroma-core/default-embed` alongside `chromadb` -- the default embedding function ships as a separate package since v3)** **(You MUST access query results as nested arrays -- `results.ids[0]`, `results.documents[0]` -- because Chroma batches queries and returns `string[][]` not `string[]`)** **(You MUST use the `configuration` parameter for HNSW settings -- the legacy `metadata: { "hnsw:space": "cosine" }` approach is deprecated)** **(You MUST use flat metadata values only (string, number, boolean, typed arrays) -- nested objects are not supported and will be rejected)** **Failure to follow these rules will cause embedding failures, incorrect result access, deprecated configuration warnings, and rejected metadata.** </critical_reminders>
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.