Claude Skill

api-vector-db-chroma

Chroma vector database -- collection management, automatic embedding, metadata filtering, document storage, query patterns

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download agents-inc-skills-dist_plugins_api-vector-db-chroma_skills_api-vector-db-chroma-3a51ef5.zip · 18 KB
Part of agents-inc/skills — 130 skills

Install

skills CLI npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/api-vector-db-chroma/skills/api-vector-db-chroma
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
Git git clone https://github.com/agents-inc/skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Chroma Patterns

Quick Guide: Use chromadb (v3.x) with @chroma-core/default-embed for automatic embedding. Chroma auto-embeds documents if no embeddings are provided -- just pass documents and ids to collection.add(). Use where for metadata filtering and whereDocument for document content filtering ($contains, $regex). Default distance metric is l2 (Euclidean); use cosine for most embedding models via configuration: { hnsw: { space: "cosine" } }. Query results return nested arrays (ids: string[][]) because queries are batched -- always access results.ids[0] for a single query. Include only the fields you need via the include parameter to reduce payload size.


<critical_requirements>

CRITICAL: Before Using This Skill

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST install @chroma-core/default-embed alongside chromadb -- the default embedding function ships as a separate package since v3)

(You MUST access query results as nested arrays -- results.ids[0], results.documents[0] -- because Chroma batches queries and returns string[][] not string[])

(You MUST use the configuration parameter for HNSW settings -- the legacy metadata: { "hnsw:space": "cosine" } approach is deprecated)

(You MUST use flat metadata values only (string, number, boolean, typed arrays) -- nested objects are not supported and will be rejected)

</critical_requirements>


Examples

  • Core Patterns -- Client setup, collection management, add, query, get, update, upsert, delete
  • Metadata Filtering -- Filter operators, compound filters, document content filters, whereDocument
  • Embedding Functions -- Default, OpenAI, custom embedding functions, provider packages

Additional resources:

  • reference.md -- API quick reference, filter operators, include options, limits, production checklist

Auto-detection: Chroma, chromadb, ChromaClient, CloudClient, createCollection, getOrCreateCollection, collection.add, collection.query, collection.get, collection.upsert, queryTexts, queryEmbeddings, nResults, whereDocument, $contains, @chroma-core/default-embed, @chroma-core/openai, EmbeddingFunction, vector database, semantic search, embedding, RAG retrieval, hnsw:space

When to use:

  • Semantic search over document embeddings (RAG retrieval)
  • Rapid prototyping with automatic embedding generation (no external embedding pipeline needed)
  • Metadata-filtered vector search with compound logical operators
  • Document content filtering with $contains and $regex
  • Local development with in-process or Docker-based Chroma server

Key patterns covered:

  • Client setup (HTTP, Cloud, Docker)
  • Collection management (create, get, delete, configure HNSW)
  • Document CRUD with automatic embedding (add, query, get, update, upsert, delete)
  • Metadata filtering (where) with comparison, set, array, and logical operators
  • Document content filtering (whereDocument) with $contains and $regex
  • Embedding function configuration (default, OpenAI, custom)
  • Query result handling (nested array structure, include options)

When NOT to use:

  • Full-text search with complex boolean ranking (use a dedicated search engine)
  • Relational data with joins and transactions (use a relational database)
  • Multi-modal image+text embeddings in TypeScript (currently Python-only in Chroma)
  • High-scale production with millions of vectors and strict SLAs (evaluate managed vector databases)



<decision_framework>

Decision Framework

Which Distance Metric?

Which distance metric should I use?
|-- Using embeddings from a language model? -> cosine (normalized, most common)
|-- Need dot product similarity? -> ip (inner product)
|-- Comparing raw feature vectors? -> l2 (Euclidean, Chroma default)
'-- Unsure? -> cosine (safe default for most embedding models)

Which Embedding Function?

Which embedding function should I use?
|-- Quick prototyping, English text? -> @chroma-core/default-embed (all-MiniLM-L6-v2, runs locally)
|-- Need high-quality embeddings? -> @chroma-core/openai (text-embedding-3-small)
|-- Have your own embedding pipeline? -> Pass embeddings directly (skip embedding function)
|-- Need custom model? -> Implement EmbeddingFunction interface
'-- Want all providers? -> npm install @chroma-core/all

Where vs WhereDocument?

How should I filter results?
|-- Structured attributes (category, year, status)? -> where (metadata filter)
|-- Full-text content search? -> whereDocument ($contains, $regex)
|-- Both? -> Combine where + whereDocument in same query
'-- Need exact match on specific IDs? -> get({ ids: [...] })

ChromaClient vs CloudClient?

Which client should I use?
|-- Local development or self-hosted? -> ChromaClient({ path: "http://localhost:8000" })
|-- Chroma Cloud (managed)? -> CloudClient({ apiKey, tenant, database })
|-- Docker deployment? -> ChromaClient with Docker host URL
'-- Testing? -> ChromaClient against local Docker container

</decision_framework>


<red_flags>

RED FLAGS

High Priority Issues:

  • Accessing query results as flat arrays instead of nested -- results.ids is string[][], not string[]; always use results.ids[0] for single-query results
  • Missing @chroma-core/default-embed package -- since v3, the default embedding function ships separately; npm install chromadb @chroma-core/default-embed
  • Using deprecated metadata: { "hnsw:space": "cosine" } for HNSW config -- use configuration: { hnsw: { space: "cosine" } } instead
  • Nested objects in metadata -- Chroma only supports flat key-value metadata; nested objects are rejected

Medium Priority Issues:

  • Not specifying include in queries -- default includes vary (query returns documents, metadatas, distances; get returns documents, metadatas); explicitly set include for clarity and to control payload size
  • Using l2 (default) when cosine is appropriate -- most embedding models are normalized for cosine similarity; l2 may produce worse results
  • Calling add() without documents or embeddings -- at least one must be provided; metadata alone is insufficient
  • Not handling empty results -- results.ids[0] may be an empty array; check length before processing

Common Mistakes:

  • Passing queryEmbeddings AND queryTexts together -- use one or the other, not both
  • Expecting update() to create missing records -- update() silently ignores non-existent IDs; use upsert() for create-or-update semantics
  • Calling delete() with no arguments -- deletes nothing (not everything); pass ids or where to target specific records
  • Using $gt/$lt on string metadata -- comparison operators only work on numeric values (int or float)

Gotchas & Edge Cases:

  • HNSW configuration (space, ef_construction, max_neighbors) cannot be changed after collection creation -- you must delete and recreate the collection
  • collection.count() returns total records in the collection, not filtered counts -- there is no filtered count API
  • peek() returns the first limit items (default 10) in insertion order, not by relevance -- useful for debugging, not querying
  • $contains in whereDocument is case-sensitive -- searching for "Neural" will not match "neural"
  • $regex in whereDocument uses full regex syntax but can be slow on large collections
  • Array metadata values (string[], number[]) must be homogeneous -- mixing types within an array is rejected
  • Metadata keys are case-sensitive -- Category and category are different fields
  • The nResults default is 10 if not specified in query()
  • Multimodal embedding (images + text) is currently Python-only -- TypeScript support is not yet available

</red_flags>


<critical_reminders>

CRITICAL REMINDERS

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST install @chroma-core/default-embed alongside chromadb -- the default embedding function ships as a separate package since v3)

(You MUST access query results as nested arrays -- results.ids[0], results.documents[0] -- because Chroma batches queries and returns string[][] not string[])

(You MUST use the configuration parameter for HNSW settings -- the legacy metadata: { "hnsw:space": "cosine" } approach is deprecated)

(You MUST use flat metadata values only (string, number, boolean, typed arrays) -- nested objects are not supported and will be rejected)

Failure to follow these rules will cause embedding failures, incorrect result access, deprecated configuration warnings, and rejected metadata.

</critical_reminders>

Files (skills)
  • examples
    • core.md 14.7 KB
      # Chroma -- Core Pattern Examples
      
      > Client setup, collection management, and fundamental operations (add, query, get, update, upsert, delete). Reference from [SKILL.md](../SKILL.md).
      
      **Related examples:**
      
      - [metadata-filtering.md](metadata-filtering.md) -- Filter operators and compound filters
      - [embedding-functions.md](embedding-functions.md) -- Default, OpenAI, custom embedding functions
      
      ---
      
      ## Client Initialization (HTTP)
      
      ```typescript
      import { ChromaClient } from "chromadb";
      
      function createChromaClient(): ChromaClient {
        const chromaUrl = process.env.CHROMA_URL;
        if (!chromaUrl) {
          throw new Error("CHROMA_URL environment variable is required");
        }
        return new ChromaClient({ path: chromaUrl });
      }
      
      export { createChromaClient };
      ```
      
      **Why good:** Server URL from environment variable (never hardcoded), validation before construction, named export
      
      By default, `ChromaClient()` connects to `http://localhost:8000`. Explicit URL is preferred to avoid deployment-specific bugs.
      
      ---
      
      ## Cloud Client Initialization
      
      ```typescript
      import { CloudClient } from "chromadb";
      
      function createCloudClient(): CloudClient {
        const apiKey = process.env.CHROMA_API_KEY;
        if (!apiKey) {
          throw new Error("CHROMA_API_KEY environment variable is required");
        }
        return new CloudClient({
          apiKey,
          tenant: process.env.CHROMA_TENANT ?? "default_tenant",
          database: process.env.CHROMA_DATABASE ?? "default_database",
        });
      }
      
      export { createCloudClient };
      ```
      
      **Why good:** API key from environment, explicit tenant and database, named export
      
      ---
      
      ## Client with Token Authentication
      
      ```typescript
      import { ChromaClient } from "chromadb";
      
      function createAuthenticatedClient(): ChromaClient {
        const chromaUrl = process.env.CHROMA_URL;
        const chromaToken = process.env.CHROMA_TOKEN;
        if (!chromaUrl || !chromaToken) {
          throw new Error(
            "CHROMA_URL and CHROMA_TOKEN environment variables required",
          );
        }
      
        return new ChromaClient({
          path: chromaUrl,
          auth: {
            provider: "token",
            credentials: chromaToken,
            tokenHeaderType: "X_CHROMA_TOKEN",
          },
        });
      }
      
      export { createAuthenticatedClient };
      ```
      
      **Why good:** Token from environment, explicit auth provider and header type
      
      ---
      
      ## Create a Collection with Cosine Distance
      
      ```typescript
      import { ChromaClient } from "chromadb";
      
      const COLLECTION_NAME = "documents";
      
      async function createCosineCollection(client: ChromaClient) {
        const collection = await client.createCollection({
          name: COLLECTION_NAME,
          configuration: {
            hnsw: { space: "cosine" },
          },
        });
        return collection;
      }
      
      export { createCosineCollection };
      ```
      
      **Why good:** Named constant for collection name, `cosine` metric via `configuration` parameter (not deprecated metadata approach)
      
      ```typescript
      // Bad Example -- deprecated HNSW config approach
      const collection = await client.createCollection({
        name: "docs",
        metadata: { "hnsw:space": "cosine" }, // DEPRECATED in v3
      });
      ```
      
      **Why bad:** `metadata` prefix for HNSW settings is deprecated; use `configuration: { hnsw: { space } }` instead
      
      ---
      
      ## Get or Create Collection (Idempotent)
      
      ```typescript
      import { ChromaClient } from "chromadb";
      
      const COLLECTION_NAME = "articles";
      
      async function getArticlesCollection(client: ChromaClient) {
        const collection = await client.getOrCreateCollection({
          name: COLLECTION_NAME,
          configuration: {
            hnsw: { space: "cosine" },
          },
        });
        return collection;
      }
      
      export { getArticlesCollection };
      ```
      
      **Why good:** Idempotent -- safe to call repeatedly, creates on first call, returns existing on subsequent calls
      
      **Gotcha:** If the collection already exists, `getOrCreateCollection` ignores the `configuration` parameter. It does NOT update the existing collection's configuration.
      
      ---
      
      ## Add Documents with Automatic Embedding
      
      ```typescript
      import type { ChromaClient } from "chromadb";
      
      interface ArticleMetadata {
        title: string;
        category: string;
        year: number;
      }
      
      const COLLECTION_NAME = "articles";
      
      async function addArticles(
        client: ChromaClient,
        articles: Array<{ id: string; text: string; metadata: ArticleMetadata }>,
      ): Promise<void> {
        const collection = await client.getOrCreateCollection({
          name: COLLECTION_NAME,
        });
      
        await collection.add({
          ids: articles.map((a) => a.id),
          documents: articles.map((a) => a.text),
          metadatas: articles.map((a) => a.metadata),
        });
      }
      
      export { addArticles };
      export type { ArticleMetadata };
      ```
      
      **Why good:** Typed metadata interface, documents auto-embedded by collection's embedding function, columnar format (parallel arrays)
      
      ```typescript
      // Bad Example -- invalid metadata and missing content
      await collection.add({
        ids: ["doc-1"],
        metadatas: [
          {
            title: "Guide",
            author: { name: "Alice", org: "Acme" }, // INVALID: nested object
          },
        ],
        // Missing documents AND embeddings -- at least one required
      });
      ```
      
      **Why bad:** Nested metadata objects are rejected, either `documents` or `embeddings` must be provided
      
      ---
      
      ## Add Documents with Pre-Computed Embeddings
      
      ```typescript
      import type { ChromaClient } from "chromadb";
      
      const COLLECTION_NAME = "articles";
      
      async function addWithEmbeddings(
        client: ChromaClient,
        records: Array<{
          id: string;
          embedding: number[];
          text: string;
          metadata: Record<string, string | number>;
        }>,
      ): Promise<void> {
        const collection = await client.getOrCreateCollection({
          name: COLLECTION_NAME,
          configuration: { hnsw: { space: "cosine" } },
        });
      
        await collection.add({
          ids: records.map((r) => r.id),
          embeddings: records.map((r) => r.embedding),
          documents: records.map((r) => r.text),
          metadatas: records.map((r) => r.metadata),
        });
      }
      
      export { addWithEmbeddings };
      ```
      
      **Why good:** Pre-computed embeddings bypass the collection's embedding function, documents stored for retrieval, metadata for filtering
      
      **Note:** When both `embeddings` and `documents` are provided, Chroma stores the documents but uses the provided embeddings (does not re-embed).
      
      ---
      
      ## Query by Text Similarity
      
      ```typescript
      import type { ChromaClient } from "chromadb";
      
      const N_RESULTS = 10;
      const COLLECTION_NAME = "articles";
      
      interface SearchResult {
        id: string;
        document: string | null;
        distance: number | null;
        metadata: Record<string, string | number | boolean> | null;
      }
      
      async function searchArticles(
        client: ChromaClient,
        queryText: string,
      ): Promise<SearchResult[]> {
        const collection = await client.getCollection({ name: COLLECTION_NAME });
      
        const results = await collection.query({
          queryTexts: [queryText],
          nResults: N_RESULTS,
          include: ["documents", "metadatas", "distances"],
        });
      
        // Results are nested arrays -- [0] for the first query
        return results.ids[0].map((id, i) => ({
          id,
          document: results.documents?.[0]?.[i] ?? null,
          distance: results.distances?.[0]?.[i] ?? null,
          metadata: results.metadatas?.[0]?.[i] ?? null,
        }));
      }
      
      export { searchArticles };
      ```
      
      **Why good:** Named constant for nResults, explicit include, correct nested array access `[0]`, handles nullable fields, typed return value
      
      ```typescript
      // Bad Example -- incorrect result access
      const results = await collection.query({
        queryTexts: ["query"],
        nResults: 5,
      });
      
      // BUG: results.ids is string[][], not string[]
      console.log(results.ids[0]); // This is correct
      console.log(results.ids.length); // This is the number of QUERIES, not results!
      ```
      
      **Why bad:** Treating `results.ids` as a flat array leads to bugs; must use `results.ids[0]` for single-query results
      
      ---
      
      ## Query with Multiple Queries (Batched)
      
      ```typescript
      const N_RESULTS = 5;
      
      async function batchQuery(
        collection: Awaited<ReturnType<ChromaClient["getCollection"]>>,
        queries: string[],
      ) {
        const results = await collection.query({
          queryTexts: queries, // Multiple queries in one call
          nResults: N_RESULTS,
        });
      
        // Each query gets its own result set
        return queries.map((query, queryIndex) => ({
          query,
          results: results.ids[queryIndex].map((id, i) => ({
            id,
            distance: results.distances?.[queryIndex]?.[i] ?? null,
          })),
        }));
      }
      
      export { batchQuery };
      ```
      
      **Why good:** Multiple queries in single API call, correct nested array indexing per query, efficient batching
      
      ---
      
      ## Query with Pre-Computed Embeddings
      
      ```typescript
      const N_RESULTS = 10;
      
      const results = await collection.query({
        queryEmbeddings: [queryVector], // Pre-computed embedding
        nResults: N_RESULTS,
        include: ["documents", "metadatas", "distances"],
      });
      ```
      
      **Why good:** Bypasses collection's embedding function, useful when you manage your own embedding pipeline
      
      **Note:** Do not pass both `queryTexts` and `queryEmbeddings` -- use one or the other.
      
      ---
      
      ## Get Documents by ID
      
      ```typescript
      async function getDocuments(
        collection: Awaited<ReturnType<ChromaClient["getCollection"]>>,
        ids: string[],
      ) {
        const results = await collection.get({
          ids,
          include: ["documents", "metadatas"],
        });
      
        // get() returns flat arrays (not nested like query())
        return results.ids.map((id, i) => ({
          id,
          document: results.documents?.[i] ?? null,
          metadata: results.metadatas?.[i] ?? null,
        }));
      }
      
      export { getDocuments };
      ```
      
      **Why good:** `get()` returns flat arrays (unlike `query()`), explicit include, handles nullable fields
      
      **Gotcha:** `get()` silently omits IDs that don't exist. If you request 5 IDs and 2 don't exist, you get 3 results with no error.
      
      ---
      
      ## Get with Pagination
      
      ```typescript
      const PAGE_SIZE = 50;
      
      async function getAllDocuments(
        collection: Awaited<ReturnType<ChromaClient["getCollection"]>>,
      ): Promise<Array<{ id: string; document: string | null }>> {
        const allResults: Array<{ id: string; document: string | null }> = [];
        let offset = 0;
      
        while (true) {
          const page = await collection.get({
            limit: PAGE_SIZE,
            offset,
            include: ["documents"],
          });
      
          if (page.ids.length === 0) {
            break;
          }
      
          allResults.push(
            ...page.ids.map((id, i) => ({
              id,
              document: page.documents?.[i] ?? null,
            })),
          );
      
          offset += page.ids.length;
        }
      
        return allResults;
      }
      
      export { getAllDocuments };
      ```
      
      **Why good:** Paginated retrieval for large collections, named constant for page size, terminates when no more results
      
      ---
      
      ## Update Existing Records
      
      ```typescript
      async function updateDocumentMetadata(
        collection: Awaited<ReturnType<ChromaClient["getCollection"]>>,
        id: string,
        newCategory: string,
      ): Promise<void> {
        await collection.update({
          ids: [id],
          metadatas: [{ category: newCategory }],
        });
      }
      
      export { updateDocumentMetadata };
      ```
      
      **Why good:** Updates specific fields without re-embedding (if only metadata changes)
      
      **Gotcha:** `update()` silently ignores non-existent IDs. If the ID doesn't exist, nothing happens and no error is thrown. Use `upsert()` for create-or-update semantics.
      
      **Note:** If `documents` are provided in `update()`, Chroma re-embeds them using the collection's embedding function.
      
      ---
      
      ## Upsert (Create or Update)
      
      ```typescript
      async function upsertDocuments(
        collection: Awaited<ReturnType<ChromaClient["getCollection"]>>,
        docs: Array<{
          id: string;
          text: string;
          metadata: Record<string, string | number | boolean>;
        }>,
      ): Promise<void> {
        await collection.upsert({
          ids: docs.map((d) => d.id),
          documents: docs.map((d) => d.text),
          metadatas: docs.map((d) => d.metadata),
        });
      }
      
      export { upsertDocuments };
      ```
      
      **Why good:** Idempotent -- creates new records or updates existing ones, safer than `add()` for pipelines that may run multiple times
      
      ---
      
      ## Delete Records
      
      ```typescript
      // Delete by IDs
      async function deleteByIds(
        collection: Awaited<ReturnType<ChromaClient["getCollection"]>>,
        ids: string[],
      ): Promise<void> {
        await collection.delete({ ids });
      }
      
      // Delete by metadata filter
      async function deleteByCategory(
        collection: Awaited<ReturnType<ChromaClient["getCollection"]>>,
        category: string,
      ): Promise<void> {
        await collection.delete({
          where: { category: { $eq: category } },
        });
      }
      
      export { deleteByIds, deleteByCategory };
      ```
      
      **Why good:** Two deletion patterns (by ID, by filter), named exports
      
      **Gotcha:** `delete()` with no arguments is a no-op -- it does not delete everything. Pass `ids` or `where` to target specific records.
      
      ---
      
      ## Delete Collection
      
      ```typescript
      async function removeCollection(
        client: ChromaClient,
        collectionName: string,
      ): Promise<void> {
        await client.deleteCollection({ name: collectionName });
      }
      
      export { removeCollection };
      ```
      
      **Why good:** Clean deletion of entire collection including all vectors, metadata, and configuration
      
      **Note:** This is irreversible. To recreate with different HNSW settings, delete and recreate the collection.
      
      ---
      
      ## Collection Health Check
      
      ```typescript
      import { ChromaClient } from "chromadb";
      
      async function healthCheck(client: ChromaClient): Promise<{
        serverAlive: boolean;
        version: string | null;
        collectionCount: number;
      }> {
        try {
          await client.heartbeat();
          const version = await client.version();
          const count = await client.countCollections();
      
          return { serverAlive: true, version, collectionCount: count };
        } catch {
          return { serverAlive: false, version: null, collectionCount: 0 };
        }
      }
      
      export { healthCheck };
      ```
      
      **Why good:** Uses `heartbeat()` for connectivity check, `version()` for compatibility verification, handles connection failures gracefully
      
      ---
      
      ## Peek at Collection Contents
      
      ```typescript
      const PEEK_LIMIT = 5;
      
      async function peekCollection(
        collection: Awaited<ReturnType<ChromaClient["getCollection"]>>,
      ) {
        const preview = await collection.peek({ limit: PEEK_LIMIT });
      
        console.log(`Total records: ${await collection.count()}`);
        for (let i = 0; i < preview.ids.length; i++) {
          console.log(preview.ids[i], preview.documents?.[i]?.slice(0, 100));
        }
      }
      
      export { peekCollection };
      ```
      
      **Why good:** Quick debugging tool, `peek()` returns first N items in insertion order, `count()` for total size
      
      **Note:** `peek()` is for debugging, not querying. Results are in insertion order, not relevance order.
      
      ---
      
      ## Result Iteration with `.rows()`
      
      ```typescript
      // Get results -- flat iteration
      const getResults = await collection.get({
        include: ["documents", "metadatas"],
      });
      
      for (const row of getResults.rows()) {
        console.log(row.id, row.document, row.metadata);
      }
      
      // Query results -- nested iteration (one batch per query)
      const queryResults = await collection.query({
        queryTexts: ["search term"],
        nResults: 5,
      });
      
      for (const batch of queryResults.rows()) {
        for (const row of batch) {
          console.log(row.id, row.document, row.metadata, row.distance);
        }
      }
      ```
      
      **Why good:** `.rows()` provides a cleaner iteration API than manual index access, handles the nested/flat difference between query and get results
      
      ---
      
      _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_
      
    • embedding-functions.md 6.5 KB
      # Chroma -- Embedding Function Examples
      
      > Default, OpenAI, and custom embedding function configuration. See [core.md](core.md) for basic operations and [reference.md](../reference.md) for the full provider list.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, add, query
      - [metadata-filtering.md](metadata-filtering.md) -- Filter operators
      
      ---
      
      ## Default Embedding Function
      
      The default embedding function uses `all-MiniLM-L6-v2` (Sentence Transformers), running locally. Requires a separate package since v3.
      
      ```bash
      npm install chromadb @chroma-core/default-embed
      ```
      
      ```typescript
      import { ChromaClient } from "chromadb";
      
      const COLLECTION_NAME = "articles";
      
      async function createWithDefaultEmbedding(client: ChromaClient) {
        // No embeddingFunction parameter needed -- uses default automatically
        const collection = await client.createCollection({
          name: COLLECTION_NAME,
          configuration: { hnsw: { space: "cosine" } },
        });
      
        // Documents are embedded automatically using all-MiniLM-L6-v2
        await collection.add({
          ids: ["doc-1", "doc-2"],
          documents: [
            "Introduction to machine learning",
            "Advanced neural network architectures",
          ],
        });
      
        return collection;
      }
      
      export { createWithDefaultEmbedding };
      ```
      
      **Why good:** Simplest setup -- just pass documents, embedding happens automatically. No API keys needed, runs locally.
      
      **Note:** `@chroma-core/default-embed` must be installed even though it's not explicitly imported. Chroma resolves it at runtime.
      
      ---
      
      ## OpenAI Embedding Function
      
      Use OpenAI's `text-embedding-3-small` for higher-quality embeddings.
      
      ```bash
      npm install @chroma-core/openai
      ```
      
      ```typescript
      import { ChromaClient } from "chromadb";
      import { OpenAIEmbeddingFunction } from "@chroma-core/openai";
      
      const COLLECTION_NAME = "knowledge-base";
      
      async function createWithOpenAI(client: ChromaClient) {
        const openAiKey = process.env.OPENAI_API_KEY;
        if (!openAiKey) {
          throw new Error("OPENAI_API_KEY environment variable is required");
        }
      
        const embeddingFunction = new OpenAIEmbeddingFunction({
          modelName: "text-embedding-3-small",
        });
      
        const collection = await client.createCollection({
          name: COLLECTION_NAME,
          embeddingFunction,
          configuration: { hnsw: { space: "cosine" } },
        });
      
        return collection;
      }
      
      export { createWithOpenAI };
      ```
      
      **Why good:** API key from environment (reads `OPENAI_API_KEY` automatically), explicit model name, cosine metric matches OpenAI's normalized embeddings
      
      **Note:** When retrieving an existing collection with `getCollection`, you must pass the same `embeddingFunction` so that queries are embedded correctly:
      
      ```typescript
      const collection = await client.getCollection({
        name: COLLECTION_NAME,
        embeddingFunction: new OpenAIEmbeddingFunction({
          modelName: "text-embedding-3-small",
        }),
      });
      ```
      
      ---
      
      ## Custom Embedding Function
      
      Implement the `EmbeddingFunction` interface for custom models.
      
      ```typescript
      import type { EmbeddingFunction } from "chromadb";
      
      class CustomEmbeddingFunction implements EmbeddingFunction {
        public readonly name = "custom-embedder";
        private readonly model: string;
      
        constructor(args: { model: string }) {
          this.model = args.model;
        }
      
        async generate(texts: string[]): Promise<number[][]> {
          // Replace with your embedding logic
          // Must return one embedding vector per input text
          const embeddings = await Promise.all(
            texts.map(async (text) => {
              // Call your embedding API or model
              return new Array(384).fill(0).map(() => Math.random());
            }),
          );
          return embeddings;
        }
      
        getConfig(): Record<string, unknown> {
          return { model: this.model };
        }
      
        validateConfigUpdate(config: Record<string, unknown>): void {
          if ("model" in config) {
            throw new Error("Model cannot be updated after creation");
          }
        }
      
        static buildFromConfig(
          config: Record<string, unknown>,
        ): CustomEmbeddingFunction {
          return new CustomEmbeddingFunction({ model: config.model as string });
        }
      }
      
      export { CustomEmbeddingFunction };
      ```
      
      **Why good:** Implements required interface methods, immutable model config, `generate()` returns one vector per input
      
      ---
      
      ## Using Embedding Functions Directly
      
      You can invoke embedding functions independently for debugging or pre-computation.
      
      ```typescript
      import { DefaultEmbeddingFunction } from "@chroma-core/default-embed";
      
      async function debugEmbeddings(): Promise<void> {
        const embeddingFn = new DefaultEmbeddingFunction();
      
        // Generate embeddings outside of a collection
        const embeddings = await embeddingFn.generate([
          "test document one",
          "test document two",
        ]);
      
        console.log("Dimension:", embeddings[0].length);
        console.log("First embedding (first 5 values):", embeddings[0].slice(0, 5));
      }
      
      export { debugEmbeddings };
      ```
      
      **Why good:** Useful for verifying embedding dimensions, debugging embedding quality, or pre-computing embeddings for external use
      
      ---
      
      ## Collection with Ollama Embeddings
      
      Use locally-hosted models via Ollama.
      
      ```bash
      npm install @chroma-core/ollama
      ```
      
      ```typescript
      import { ChromaClient } from "chromadb";
      import { OllamaEmbeddingFunction } from "@chroma-core/ollama";
      
      const COLLECTION_NAME = "local-knowledge";
      
      async function createWithOllama(client: ChromaClient) {
        const embeddingFunction = new OllamaEmbeddingFunction({
          url: process.env.OLLAMA_URL ?? "http://localhost:11434",
          model: "nomic-embed-text",
        });
      
        const collection = await client.createCollection({
          name: COLLECTION_NAME,
          embeddingFunction,
          configuration: { hnsw: { space: "cosine" } },
        });
      
        return collection;
      }
      
      export { createWithOllama };
      ```
      
      **Why good:** Fully local embedding pipeline (no API keys), configurable Ollama URL, explicit model name
      
      ---
      
      ## Embedding Function Gotchas
      
      **Dimension mismatch:** If you change embedding functions on an existing collection (e.g., switch from default 384-dim to OpenAI 1536-dim), queries will fail with dimension mismatch errors. Delete and recreate the collection with the new function.
      
      **Consistency requirement:** Always pass the same embedding function to `getCollection()` / `getOrCreateCollection()` as was used at creation. If you omit it, Chroma uses the default, which may produce different-dimension embeddings.
      
      **Query vs add:** The same embedding function is used for both `add()` (document embedding) and `query()` (query embedding). Do not mix different embedding models between add and query operations on the same collection.
      
      ---
      
      _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_
      
    • metadata-filtering.md 8.8 KB
      # Chroma -- Metadata Filtering Examples
      
      > Filter operators, compound filters, document content filters, and where/whereDocument patterns. See [core.md](core.md) for basic query patterns and [reference.md](../reference.md) for the operator table.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, basic query with filters
      - [embedding-functions.md](embedding-functions.md) -- Embedding function configuration
      
      ---
      
      ## Comparison Operators
      
      ```typescript
      const N_RESULTS = 10;
      
      // Exact match (explicit)
      const drama = await collection.query({
        queryTexts: ["dramatic story"],
        nResults: N_RESULTS,
        where: { genre: { $eq: "drama" } },
      });
      
      // Exact match (shorthand -- equivalent to $eq)
      const dramaShorthand = await collection.query({
        queryTexts: ["dramatic story"],
        nResults: N_RESULTS,
        where: { genre: "drama" },
      });
      
      // Not equal
      const notComedy = await collection.query({
        queryTexts: ["funny story"],
        nResults: N_RESULTS,
        where: { genre: { $ne: "comedy" } },
      });
      
      // Numeric range (greater than or equal)
      const recent = await collection.query({
        queryTexts: ["recent articles"],
        nResults: N_RESULTS,
        where: { year: { $gte: 2023 } },
      });
      
      // Numeric range (less than)
      const older = await collection.query({
        queryTexts: ["historical articles"],
        nResults: N_RESULTS,
        where: { year: { $lt: 2020 } },
      });
      ```
      
      **Why good:** Each operator targets specific types, shorthand `{ genre: "drama" }` is equivalent to `{ genre: { $eq: "drama" } }`
      
      ---
      
      ## Set Membership Operators
      
      ```typescript
      const N_RESULTS = 10;
      
      // $in -- value in set
      const selectedGenres = await collection.query({
        queryTexts: ["search query"],
        nResults: N_RESULTS,
        where: { genre: { $in: ["drama", "action", "thriller"] } },
      });
      
      // $nin -- value not in set
      const excludedGenres = await collection.query({
        queryTexts: ["search query"],
        nResults: N_RESULTS,
        where: { genre: { $nin: ["horror", "comedy"] } },
      });
      ```
      
      **Why good:** `$in` replaces multiple `$or` / `$eq` clauses, cleaner and more efficient
      
      ---
      
      ## Array Metadata Operators
      
      For metadata fields that are arrays (e.g., tags, authors). Requires Chroma >= 1.5.0.
      
      ```typescript
      const N_RESULTS = 10;
      
      // $contains -- array includes specific value
      const chenPapers = await collection.query({
        queryTexts: ["machine learning"],
        nResults: N_RESULTS,
        where: { authors: { $contains: "Chen" } },
      });
      
      // $not_contains -- array excludes specific value
      const withoutDraft = await collection.query({
        queryTexts: ["published papers"],
        nResults: N_RESULTS,
        where: { tags: { $not_contains: "draft" } },
      });
      ```
      
      **Why good:** `$contains` checks if an array metadata field includes a scalar value, works with string/int/float/boolean arrays
      
      **Gotcha:** `$contains` and `$not_contains` only work on array metadata fields, not on scalar values. For scalar equality, use `$eq`.
      
      ---
      
      ## Compound Filters with $and / $or
      
      ```typescript
      const N_RESULTS = 10;
      
      // AND: all conditions must match
      const filteredResults = await collection.query({
        queryTexts: ["search query"],
        nResults: N_RESULTS,
        where: {
          $and: [
            { genre: { $eq: "drama" } },
            { year: { $gte: 2020 } },
            { rating: { $gt: 7.5 } },
          ],
        },
      });
      
      // OR: any condition matches
      const broadResults = await collection.query({
        queryTexts: ["search query"],
        nResults: N_RESULTS,
        where: {
          $or: [{ genre: { $eq: "drama" } }, { genre: { $eq: "documentary" } }],
        },
      });
      
      // Nested AND + OR
      const complexResults = await collection.query({
        queryTexts: ["search query"],
        nResults: N_RESULTS,
        where: {
          $and: [
            { year: { $gte: 2020 } },
            {
              $or: [{ genre: { $eq: "drama" } }, { genre: { $eq: "thriller" } }],
            },
          ],
        },
      });
      ```
      
      **Why good:** Nested logical operators for complex queries, `$or` inside `$and` for flexible filtering
      
      ---
      
      ## Document Content Filtering (whereDocument)
      
      Filter by document text content using `$contains`, `$not_contains`, `$regex`, and `$not_regex`.
      
      ```typescript
      const N_RESULTS = 10;
      
      // $contains -- case-sensitive substring match
      const containsNeural = await collection.query({
        queryTexts: ["deep learning"],
        nResults: N_RESULTS,
        whereDocument: { $contains: "neural network" },
      });
      
      // $not_contains -- exclude documents with substring
      const excludeDeprecated = await collection.query({
        queryTexts: ["current practices"],
        nResults: N_RESULTS,
        whereDocument: { $not_contains: "deprecated" },
      });
      
      // $regex -- pattern matching
      const regexMatch = await collection.query({
        queryTexts: ["technical papers"],
        nResults: N_RESULTS,
        whereDocument: { $regex: "deep\\s+learning" },
      });
      
      // $not_regex -- exclude by pattern
      const excludePattern = await collection.query({
        queryTexts: ["final papers"],
        nResults: N_RESULTS,
        whereDocument: { $not_regex: "draft.*v[0-9]" },
      });
      ```
      
      **Why good:** `$contains` for simple substring search, `$regex` for pattern matching, `$not_*` variants for exclusion
      
      **Gotcha:** `$contains` is case-sensitive -- "Neural" will not match "neural". Use `$regex` with case-insensitive patterns if needed.
      
      ---
      
      ## Compound Document Filters
      
      ```typescript
      const N_RESULTS = 10;
      
      // AND -- both substrings must appear in document
      const bothTerms = await collection.query({
        queryTexts: ["AI research"],
        nResults: N_RESULTS,
        whereDocument: {
          $and: [{ $contains: "transformer" }, { $contains: "attention" }],
        },
      });
      
      // OR -- either substring matches
      const eitherTerm = await collection.query({
        queryTexts: ["AI research"],
        nResults: N_RESULTS,
        whereDocument: {
          $or: [{ $contains: "transformer" }, { $contains: "recurrent" }],
        },
      });
      ```
      
      **Why good:** Compound document filters for multi-term matching
      
      ---
      
      ## Combined Metadata + Document Filters
      
      Use `where` and `whereDocument` together for maximum filtering precision.
      
      ```typescript
      const N_RESULTS = 10;
      
      const results = await collection.query({
        queryTexts: ["machine learning tutorial"],
        nResults: N_RESULTS,
        where: {
          $and: [{ category: { $eq: "tutorial" } }, { year: { $gte: 2023 } }],
        },
        whereDocument: { $contains: "neural network" },
        include: ["documents", "metadatas", "distances"],
      });
      
      // Access results (nested arrays for query)
      for (let i = 0; i < results.ids[0].length; i++) {
        console.log(
          results.ids[0][i],
          results.distances?.[0]?.[i],
          results.metadatas?.[0]?.[i],
        );
      }
      ```
      
      **Why good:** Combined metadata + document filters for precise retrieval, correct nested array access
      
      ---
      
      ## Filtering with get() (No Vector Query)
      
      Use `get()` with filters to retrieve documents without similarity search.
      
      ```typescript
      const PAGE_SIZE = 50;
      
      // Get by metadata filter
      const tutorials = await collection.get({
        where: { category: { $eq: "tutorial" } },
        limit: PAGE_SIZE,
        include: ["documents", "metadatas"],
      });
      
      // Get by document content filter
      const withCode = await collection.get({
        whereDocument: { $contains: "function" },
        limit: PAGE_SIZE,
        include: ["documents"],
      });
      
      // Combined metadata + document filter
      const recentTutorialsWithCode = await collection.get({
        where: {
          $and: [{ category: { $eq: "tutorial" } }, { year: { $gte: 2023 } }],
        },
        whereDocument: { $contains: "function" },
        limit: PAGE_SIZE,
        include: ["documents", "metadatas"],
      });
      ```
      
      **Why good:** `get()` does not require a query vector, useful for data management and non-similarity-based retrieval
      
      ---
      
      ## Delete by Metadata Filter
      
      ```typescript
      // Delete all records matching a filter
      await collection.delete({
        where: {
          $and: [{ status: { $eq: "archived" } }, { year: { $lt: 2020 } }],
        },
      });
      ```
      
      **Why good:** Bulk deletion without knowing record IDs, compound filter for precise targeting
      
      ---
      
      ## Common Filter Mistakes
      
      ```typescript
      // Bad: $gt on string value (only works on numbers)
      where: {
        title: {
          $gt: "A";
        }
      }
      // Fix: use $eq, $ne, $in, or $nin for strings
      
      // Bad: Nested object in metadata
      metadatas: [{ author: { name: "Alice", org: "Acme" } }];
      // Fix: flatten to top-level keys
      metadatas: [{ authorName: "Alice", authorOrg: "Acme" }];
      
      // Bad: Mixed types in array metadata
      metadatas: [{ tags: ["draft", 2024, true] }]; // INVALID: mixed types
      // Fix: arrays must be homogeneous
      metadatas: [{ tags: ["draft", "2024"] }]; // All strings
      
      // Bad: $contains on scalar metadata field
      where: {
        category: {
          $contains: "tut";
        }
      } // INVALID: $contains is for arrays
      // Fix: use $eq for scalar fields, $contains for array fields
      where: {
        category: {
          $eq: "tutorial";
        }
      }
      
      // Bad: Case-sensitive surprise with whereDocument
      whereDocument: {
        $contains: "Neural";
      } // Won't match "neural"
      // Fix: use $regex for case-insensitive matching
      whereDocument: {
        $regex: "(?i)neural";
      }
      ```
      
      **Why bad:** Each mistake causes either a rejection or silently empty results; type-restricted operators and case sensitivity are common sources of "no results" bugs
      
      ---
      
      _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_
      
  • reference.md 15 KB
    # Chroma Quick Reference
    
    > API reference, filter operators, include options, limits, and production checklist. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples.
    
    ---
    
    ## ChromaClient Methods
    
    | Method                                          | Description                                           | Returns                 |
    | ----------------------------------------------- | ----------------------------------------------------- | ----------------------- |
    | `new ChromaClient({ path })`                    | Create HTTP client (default: `http://localhost:8000`) | `ChromaClient`          |
    | `new CloudClient({ apiKey, tenant, database })` | Create Chroma Cloud client                            | `CloudClient`           |
    | `client.heartbeat()`                            | Check server connectivity                             | `Promise<number>`       |
    | `client.version()`                              | Get server version                                    | `Promise<string>`       |
    | `client.createCollection(options)`              | Create a new collection                               | `Promise<Collection>`   |
    | `client.getCollection(options)`                 | Get an existing collection                            | `Promise<Collection>`   |
    | `client.getOrCreateCollection(options)`         | Get or create a collection                            | `Promise<Collection>`   |
    | `client.deleteCollection({ name })`             | Delete a collection                                   | `Promise<void>`         |
    | `client.listCollections({ limit?, offset? })`   | List collections (paginated)                          | `Promise<Collection[]>` |
    | `client.countCollections()`                     | Count total collections                               | `Promise<number>`       |
    | `client.reset()`                                | Reset entire database (dev only)                      | `Promise<void>`         |
    
    ## Collection Methods
    
    | Method                       | Description                     | Returns                |
    | ---------------------------- | ------------------------------- | ---------------------- |
    | `collection.add(options)`    | Add documents/embeddings        | `Promise<void>`        |
    | `collection.query(options)`  | Find similar documents          | `Promise<QueryResult>` |
    | `collection.get(options?)`   | Get by IDs or filter            | `Promise<GetResult>`   |
    | `collection.update(options)` | Update existing records         | `Promise<void>`        |
    | `collection.upsert(options)` | Insert or update records        | `Promise<void>`        |
    | `collection.delete(options)` | Delete by IDs or filter         | `Promise<void>`        |
    | `collection.peek(options?)`  | Preview first N items           | `Promise<GetResult>`   |
    | `collection.count()`         | Count records in collection     | `Promise<number>`      |
    | `collection.modify(options)` | Update collection name/metadata | `Promise<void>`        |
    
    ---
    
    ## Method Parameters
    
    ### add / upsert
    
    | Parameter    | Type                              | Required                    | Description                        |
    | ------------ | --------------------------------- | --------------------------- | ---------------------------------- |
    | `ids`        | `string[]`                        | Yes                         | Unique identifiers for each record |
    | `documents`  | `string[]`                        | One of documents/embeddings | Text documents (auto-embedded)     |
    | `embeddings` | `number[][]`                      | One of documents/embeddings | Pre-computed embedding vectors     |
    | `metadatas`  | `Record<string, MetadataValue>[]` | No                          | Metadata for filtering             |
    | `uris`       | `string[]`                        | No                          | URIs for data loaders              |
    
    ### query
    
    | Parameter         | Type            | Required                | Description                     |
    | ----------------- | --------------- | ----------------------- | ------------------------------- |
    | `queryTexts`      | `string[]`      | One of texts/embeddings | Text queries (auto-embedded)    |
    | `queryEmbeddings` | `number[][]`    | One of texts/embeddings | Pre-computed query embeddings   |
    | `nResults`        | `number`        | No                      | Number of results (default: 10) |
    | `where`           | `Where`         | No                      | Metadata filter                 |
    | `whereDocument`   | `WhereDocument` | No                      | Document content filter         |
    | `include`         | `Include[]`     | No                      | Fields to include in response   |
    
    ### get
    
    | Parameter       | Type            | Required | Description                   |
    | --------------- | --------------- | -------- | ----------------------------- |
    | `ids`           | `string[]`      | No       | Specific IDs to retrieve      |
    | `where`         | `Where`         | No       | Metadata filter               |
    | `whereDocument` | `WhereDocument` | No       | Document content filter       |
    | `limit`         | `number`        | No       | Max results to return         |
    | `offset`        | `number`        | No       | Skip first N results          |
    | `include`       | `Include[]`     | No       | Fields to include in response |
    
    ### delete
    
    | Parameter       | Type            | Required | Description                          |
    | --------------- | --------------- | -------- | ------------------------------------ |
    | `ids`           | `string[]`      | No       | Specific IDs to delete               |
    | `where`         | `Where`         | No       | Metadata filter for deletion         |
    | `whereDocument` | `WhereDocument` | No       | Document content filter for deletion |
    
    ---
    
    ## Metadata Filter Operators (`where`)
    
    | Operator        | Description           | Supported Types         | Example                                   |
    | --------------- | --------------------- | ----------------------- | ----------------------------------------- |
    | `$eq`           | Equal to (default)    | string, number, boolean | `{ genre: { $eq: "drama" } }`             |
    | `$ne`           | Not equal to          | string, number, boolean | `{ genre: { $ne: "comedy" } }`            |
    | `$gt`           | Greater than          | number only             | `{ year: { $gt: 2020 } }`                 |
    | `$gte`          | Greater than or equal | number only             | `{ year: { $gte: 2020 } }`                |
    | `$lt`           | Less than             | number only             | `{ year: { $lt: 2020 } }`                 |
    | `$lte`          | Less than or equal    | number only             | `{ year: { $lte: 2020 } }`                |
    | `$in`           | In set                | string, number, boolean | `{ genre: { $in: ["drama", "action"] } }` |
    | `$nin`          | Not in set            | string, number, boolean | `{ genre: { $nin: ["horror"] } }`         |
    | `$contains`     | Array contains value  | typed arrays            | `{ authors: { $contains: "Chen" } }`      |
    | `$not_contains` | Array excludes value  | typed arrays            | `{ tags: { $not_contains: "draft" } }`    |
    | `$and`          | Logical AND           | filter[]                | `{ $and: [filter1, filter2] }`            |
    | `$or`           | Logical OR            | filter[]                | `{ $or: [filter1, filter2] }`             |
    
    **Rules:**
    
    - Direct equality shorthand: `{ field: "value" }` equals `{ field: { $eq: "value" } }`
    - `$gt`, `$gte`, `$lt`, `$lte` work on numeric values only
    - `$contains` / `$not_contains` work on array metadata fields only (Chroma >= 1.5.0)
    - Array metadata must be homogeneous (all same type: string, int, float, or boolean)
    - Nested objects are not supported in metadata -- flat key-value pairs only
    - Metadata keys are case-sensitive
    
    ---
    
    ## Document Content Filter Operators (`whereDocument`)
    
    | Operator        | Description                       | Example                                          |
    | --------------- | --------------------------------- | ------------------------------------------------ |
    | `$contains`     | Case-sensitive substring match    | `{ $contains: "neural network" }`                |
    | `$not_contains` | Excludes documents with substring | `{ $not_contains: "deprecated" }`                |
    | `$regex`        | Regex pattern match               | `{ $regex: "deep\\s+learning" }`                 |
    | `$not_regex`    | Excludes documents matching regex | `{ $not_regex: "draft.*v[0-9]" }`                |
    | `$and`          | Logical AND for document filters  | `{ $and: [{$contains: "A"}, {$contains: "B"}] }` |
    | `$or`           | Logical OR for document filters   | `{ $or: [{$contains: "A"}, {$contains: "B"}] }`  |
    
    ---
    
    ## Include Options
    
    | Value          | In `query()` default? | In `get()` default? | Description                       |
    | -------------- | --------------------- | ------------------- | --------------------------------- |
    | `"documents"`  | Yes                   | Yes                 | Document text content             |
    | `"metadatas"`  | Yes                   | Yes                 | Metadata key-value pairs          |
    | `"distances"`  | Yes                   | No                  | Similarity distances (query only) |
    | `"embeddings"` | No                    | No                  | Raw embedding vectors             |
    | `"uris"`       | No                    | No                  | Data loader URIs                  |
    
    ---
    
    ## Return Type Structures
    
    ### QueryResult (from `query()`)
    
    ```typescript
    interface QueryResult {
      ids: string[][]; // Nested: [query1_ids, query2_ids, ...]
      documents: (string | null)[][];
      metadatas: (Record<string, MetadataValue> | null)[][];
      embeddings: (number[] | null)[][] | null;
      distances: (number | null)[][] | null;
      uris: (string | null)[][] | null;
      include: Include[];
    }
    ```
    
    ### GetResult (from `get()`, `peek()`)
    
    ```typescript
    interface GetResult {
      ids: string[]; // Flat array
      documents: (string | null)[];
      metadatas: (Record<string, MetadataValue> | null)[];
      embeddings: number[][] | null;
      uris: (string | null)[] | null;
      include: Include[];
    }
    ```
    
    **Critical difference:** `query()` returns nested arrays (batched queries), `get()` returns flat arrays.
    
    ---
    
    ## HNSW Configuration Parameters
    
    | Parameter         | Default   | Modifiable after creation? | Description                             |
    | ----------------- | --------- | -------------------------- | --------------------------------------- |
    | `space`           | `"l2"`    | No                         | Distance function: `l2`, `cosine`, `ip` |
    | `ef_construction` | `100`     | No                         | Candidate list size during index build  |
    | `ef_search`       | `100`     | Yes                        | Candidate list during queries           |
    | `max_neighbors`   | `16`      | No                         | Max connections per graph node          |
    | `num_threads`     | CPU cores | Yes                        | Threads for operations                  |
    | `batch_size`      | `100`     | Yes                        | Vectors per batch                       |
    | `sync_threshold`  | `1000`    | Yes                        | Index sync trigger                      |
    | `resize_factor`   | `1.2`     | Yes                        | Growth multiplier on resize             |
    
    ---
    
    ## Distance Metrics
    
    | Metric   | Description                          | When to Use                                         |
    | -------- | ------------------------------------ | --------------------------------------------------- |
    | `l2`     | Euclidean distance squared (default) | Raw feature vectors where absolute distance matters |
    | `cosine` | Cosine similarity (normalized)       | Most embedding models (recommended for text)        |
    | `ip`     | Inner product                        | Pre-normalized embeddings, dot product similarity   |
    
    ---
    
    ## Supported Metadata Types
    
    | Type           | Example Value   | Filter Support               |
    | -------------- | --------------- | ---------------------------- |
    | String         | `"drama"`       | All operators                |
    | Number (int)   | `2024`          | All operators                |
    | Number (float) | `0.95`          | All operators                |
    | Boolean        | `true`          | `$eq`, `$ne`, `$in`, `$nin`  |
    | String array   | `["a", "b"]`    | `$contains`, `$not_contains` |
    | Int array      | `[1, 2, 3]`     | `$contains`, `$not_contains` |
    | Float array    | `[0.1, 0.2]`    | `$contains`, `$not_contains` |
    | Boolean array  | `[true, false]` | `$contains`, `$not_contains` |
    
    **Constraints:** Arrays must be homogeneous (all elements same type). Empty arrays and nested arrays are not allowed. Nested objects are not supported.
    
    ---
    
    ## Embedding Function Packages
    
    | Package                           | Model                  | Use Case                               |
    | --------------------------------- | ---------------------- | -------------------------------------- |
    | `@chroma-core/default-embed`      | all-MiniLM-L6-v2       | Local, English text, quick prototyping |
    | `@chroma-core/openai`             | text-embedding-3-small | High-quality, production use           |
    | `@chroma-core/cohere`             | Cohere embed models    | Multilingual support                   |
    | `@chroma-core/google-gemini`      | Gemini embeddings      | Google ecosystem                       |
    | `@chroma-core/ollama`             | Ollama models          | Self-hosted, local LLMs                |
    | `@chroma-core/huggingface-server` | HF models              | Custom HF models                       |
    | `@chroma-core/voyageai`           | Voyage AI models       | Specialized embeddings                 |
    | `@chroma-core/jina`               | Jina AI models         | Long-context embeddings                |
    | `@chroma-core/all`                | All providers          | Install everything                     |
    
    ---
    
    ## Production Checklist
    
    ### Security
    
    - [ ] Chroma server behind authentication (token or basic auth)
    - [ ] API tokens stored in environment variables, not in code
    - [ ] `client.reset()` disabled in production (dangerous -- deletes everything)
    
    ### Collection Configuration
    
    - [ ] Distance metric set explicitly (`cosine` for text embeddings)
    - [ ] Embedding function configured per collection (not relying on defaults)
    - [ ] `ef_construction` increased for large collections (better recall, slower build)
    
    ### Data Management
    
    - [ ] Metadata values are flat (no nested objects)
    - [ ] IDs are unique and deterministic
    - [ ] `upsert` used instead of `add` for idempotent operations
    - [ ] Large datasets added in batches (avoid single massive `add` call)
    
    ### Query Optimization
    
    - [ ] `nResults` set to minimum needed (lower = faster)
    - [ ] `include` parameter used to exclude unnecessary fields
    - [ ] `where` filters use indexed metadata fields
    - [ ] `whereDocument` used sparingly (can be slow on large collections)
    
    ### Monitoring
    
    - [ ] `collection.count()` tracked for growth monitoring
    - [ ] `client.heartbeat()` used for health checks
    - [ ] Error handling for connection failures and timeouts
    
    ---
    
    _Full skill documentation: [SKILL.md](SKILL.md) | Examples: [examples/](examples/)_
    
  • SKILL.md 12.9 KB
    ---
    name: api-vector-db-chroma
    description: Chroma vector database -- collection management, automatic embedding, metadata filtering, document storage, query patterns
    ---
    
    # Chroma Patterns
    
    > **Quick Guide:** Use `chromadb` (v3.x) with `@chroma-core/default-embed` for automatic embedding. Chroma auto-embeds documents if no embeddings are provided -- just pass `documents` and `ids` to `collection.add()`. Use `where` for metadata filtering and `whereDocument` for document content filtering (`$contains`, `$regex`). Default distance metric is `l2` (Euclidean); use `cosine` for most embedding models via `configuration: { hnsw: { space: "cosine" } }`. Query results return nested arrays (`ids: string[][]`) because queries are batched -- always access `results.ids[0]` for a single query. Include only the fields you need via the `include` parameter to reduce payload size.
    
    ---
    
    <critical_requirements>
    
    ## CRITICAL: Before Using This Skill
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST install `@chroma-core/default-embed` alongside `chromadb` -- the default embedding function ships as a separate package since v3)**
    
    **(You MUST access query results as nested arrays -- `results.ids[0]`, `results.documents[0]` -- because Chroma batches queries and returns `string[][]` not `string[]`)**
    
    **(You MUST use the `configuration` parameter for HNSW settings -- the legacy `metadata: { "hnsw:space": "cosine" }` approach is deprecated)**
    
    **(You MUST use flat metadata values only (string, number, boolean, typed arrays) -- nested objects are not supported and will be rejected)**
    
    </critical_requirements>
    
    ---
    
    ## Examples
    
    - [Core Patterns](examples/core.md) -- Client setup, collection management, add, query, get, update, upsert, delete
    - [Metadata Filtering](examples/metadata-filtering.md) -- Filter operators, compound filters, document content filters, whereDocument
    - [Embedding Functions](examples/embedding-functions.md) -- Default, OpenAI, custom embedding functions, provider packages
    
    **Additional resources:**
    
    - [reference.md](reference.md) -- API quick reference, filter operators, include options, limits, production checklist
    
    ---
    
    **Auto-detection:** Chroma, chromadb, ChromaClient, CloudClient, createCollection, getOrCreateCollection, collection.add, collection.query, collection.get, collection.upsert, queryTexts, queryEmbeddings, nResults, whereDocument, $contains, @chroma-core/default-embed, @chroma-core/openai, EmbeddingFunction, vector database, semantic search, embedding, RAG retrieval, hnsw:space
    
    **When to use:**
    
    - Semantic search over document embeddings (RAG retrieval)
    - Rapid prototyping with automatic embedding generation (no external embedding pipeline needed)
    - Metadata-filtered vector search with compound logical operators
    - Document content filtering with `$contains` and `$regex`
    - Local development with in-process or Docker-based Chroma server
    
    **Key patterns covered:**
    
    - Client setup (HTTP, Cloud, Docker)
    - Collection management (create, get, delete, configure HNSW)
    - Document CRUD with automatic embedding (add, query, get, update, upsert, delete)
    - Metadata filtering (`where`) with comparison, set, array, and logical operators
    - Document content filtering (`whereDocument`) with `$contains` and `$regex`
    - Embedding function configuration (default, OpenAI, custom)
    - Query result handling (nested array structure, include options)
    
    **When NOT to use:**
    
    - Full-text search with complex boolean ranking (use a dedicated search engine)
    - Relational data with joins and transactions (use a relational database)
    - Multi-modal image+text embeddings in TypeScript (currently Python-only in Chroma)
    - High-scale production with millions of vectors and strict SLAs (evaluate managed vector databases)
    
    ---
    
    <philosophy>
    
    ## Philosophy
    
    Chroma is a **lightweight, developer-friendly embedding database** designed for rapid prototyping and production RAG applications. The core principle: **pass documents in, get relevant results out -- Chroma handles embedding automatically.**
    
    **Core principles:**
    
    1. **Documents first, vectors optional** -- Unlike most vector databases, Chroma can embed documents automatically using a configured embedding function. You never need to manage embeddings directly unless you want to.
    2. **Collections are self-contained** -- Each collection has its own embedding function, distance metric, and HNSW configuration. No global index management needed.
    3. **Metadata is for filtering, documents are for content** -- Use `where` for structured metadata filters and `whereDocument` for full-text content filters. Both can be combined in a single query.
    4. **Batteries included** -- The default embedding function (`all-MiniLM-L6-v2` via `@chroma-core/default-embed`) works out of the box for English text. Swap to OpenAI, Cohere, or any provider with a single package change.
    5. **Query results are batched** -- Chroma supports multiple queries in a single call. Results are always nested arrays (`string[][]`), even for single queries. Always access `[0]` for the first query's results.
    
    </philosophy>
    
    ---
    
    <patterns>
    
    ## Core Patterns
    
    ### Pattern 1: Client Initialization
    
    Always pass the server URL explicitly from an environment variable -- never rely on the implicit `http://localhost:8000` default. See [examples/core.md](examples/core.md) for HTTP, Cloud, and token-authenticated client examples.
    
    ```typescript
    const chromaUrl = process.env.CHROMA_URL;
    if (!chromaUrl) throw new Error("CHROMA_URL environment variable is required");
    return new ChromaClient({ path: chromaUrl });
    ```
    
    ---
    
    ### Pattern 2: Collection with Distance Metric
    
    Use the `configuration` parameter for HNSW settings -- never the deprecated `metadata: { "hnsw:space": "cosine" }` approach. See [examples/core.md](examples/core.md).
    
    ```typescript
    const collection = await client.createCollection({
      name: COLLECTION_NAME,
      configuration: { hnsw: { space: "cosine" } },
    });
    ```
    
    ---
    
    ### Pattern 3: Add Documents with Automatic Embedding
    
    Pass `documents` and `ids` -- Chroma embeds automatically. Either `documents` or `embeddings` must be provided; metadata alone is insufficient. See [examples/core.md](examples/core.md).
    
    ```typescript
    await collection.add({
      ids: articles.map((a) => a.id),
      documents: articles.map((a) => a.text),
      metadatas: articles.map((a) => ({ category: a.category })),
    });
    ```
    
    ---
    
    ### Pattern 4: Query with Metadata and Document Filters
    
    Combine `where` (metadata) and `whereDocument` (content) filters. Results are nested arrays -- always access `[0]` for single-query results. See [examples/metadata-filtering.md](examples/metadata-filtering.md) for all operators.
    
    ```typescript
    const results = await collection.query({
      queryTexts: ["machine learning fundamentals"],
      nResults: N_RESULTS,
      where: {
        $and: [{ category: { $eq: "tutorial" } }, { year: { $gte: 2023 } }],
      },
      whereDocument: { $contains: "neural network" },
      include: ["documents", "metadatas", "distances"],
    });
    // results.ids[0] -- nested array, access [0] for first query
    ```
    
    ---
    
    ### Pattern 5: Get with Pagination
    
    Use `get()` with `limit`/`offset` for non-similarity retrieval. Unlike `query()`, `get()` returns flat arrays. See [examples/core.md](examples/core.md).
    
    ```typescript
    const results = await collection.get({
      where: { category: { $eq: category } },
      limit: PAGE_SIZE,
      offset,
      include: ["documents", "metadatas"],
    });
    ```
    
    ---
    
    ### Pattern 6: Upsert for Idempotent Operations
    
    Use `upsert()` instead of `add()` for create-or-update semantics -- safer for pipelines that run multiple times. See [examples/core.md](examples/core.md).
    
    ```typescript
    await collection.upsert({
      ids: docs.map((d) => d.id),
      documents: docs.map((d) => d.text),
      metadatas: docs.map((d) => d.metadata),
    });
    ```
    
    </patterns>
    
    ---
    
    <decision_framework>
    
    ## Decision Framework
    
    ### Which Distance Metric?
    
    ```
    Which distance metric should I use?
    |-- Using embeddings from a language model? -> cosine (normalized, most common)
    |-- Need dot product similarity? -> ip (inner product)
    |-- Comparing raw feature vectors? -> l2 (Euclidean, Chroma default)
    '-- Unsure? -> cosine (safe default for most embedding models)
    ```
    
    ### Which Embedding Function?
    
    ```
    Which embedding function should I use?
    |-- Quick prototyping, English text? -> @chroma-core/default-embed (all-MiniLM-L6-v2, runs locally)
    |-- Need high-quality embeddings? -> @chroma-core/openai (text-embedding-3-small)
    |-- Have your own embedding pipeline? -> Pass embeddings directly (skip embedding function)
    |-- Need custom model? -> Implement EmbeddingFunction interface
    '-- Want all providers? -> npm install @chroma-core/all
    ```
    
    ### Where vs WhereDocument?
    
    ```
    How should I filter results?
    |-- Structured attributes (category, year, status)? -> where (metadata filter)
    |-- Full-text content search? -> whereDocument ($contains, $regex)
    |-- Both? -> Combine where + whereDocument in same query
    '-- Need exact match on specific IDs? -> get({ ids: [...] })
    ```
    
    ### ChromaClient vs CloudClient?
    
    ```
    Which client should I use?
    |-- Local development or self-hosted? -> ChromaClient({ path: "http://localhost:8000" })
    |-- Chroma Cloud (managed)? -> CloudClient({ apiKey, tenant, database })
    |-- Docker deployment? -> ChromaClient with Docker host URL
    '-- Testing? -> ChromaClient against local Docker container
    ```
    
    </decision_framework>
    
    ---
    
    <red_flags>
    
    ## RED FLAGS
    
    **High Priority Issues:**
    
    - Accessing query results as flat arrays instead of nested -- `results.ids` is `string[][]`, not `string[]`; always use `results.ids[0]` for single-query results
    - Missing `@chroma-core/default-embed` package -- since v3, the default embedding function ships separately; `npm install chromadb @chroma-core/default-embed`
    - Using deprecated `metadata: { "hnsw:space": "cosine" }` for HNSW config -- use `configuration: { hnsw: { space: "cosine" } }` instead
    - Nested objects in metadata -- Chroma only supports flat key-value metadata; nested objects are rejected
    
    **Medium Priority Issues:**
    
    - Not specifying `include` in queries -- default includes vary (`query` returns documents, metadatas, distances; `get` returns documents, metadatas); explicitly set `include` for clarity and to control payload size
    - Using `l2` (default) when `cosine` is appropriate -- most embedding models are normalized for cosine similarity; `l2` may produce worse results
    - Calling `add()` without `documents` or `embeddings` -- at least one must be provided; metadata alone is insufficient
    - Not handling empty results -- `results.ids[0]` may be an empty array; check length before processing
    
    **Common Mistakes:**
    
    - Passing `queryEmbeddings` AND `queryTexts` together -- use one or the other, not both
    - Expecting `update()` to create missing records -- `update()` silently ignores non-existent IDs; use `upsert()` for create-or-update semantics
    - Calling `delete()` with no arguments -- deletes nothing (not everything); pass `ids` or `where` to target specific records
    - Using `$gt`/`$lt` on string metadata -- comparison operators only work on numeric values (int or float)
    
    **Gotchas & Edge Cases:**
    
    - HNSW configuration (`space`, `ef_construction`, `max_neighbors`) cannot be changed after collection creation -- you must delete and recreate the collection
    - `collection.count()` returns total records in the collection, not filtered counts -- there is no filtered count API
    - `peek()` returns the first `limit` items (default 10) in insertion order, not by relevance -- useful for debugging, not querying
    - `$contains` in `whereDocument` is case-sensitive -- searching for "Neural" will not match "neural"
    - `$regex` in `whereDocument` uses full regex syntax but can be slow on large collections
    - Array metadata values (`string[]`, `number[]`) must be homogeneous -- mixing types within an array is rejected
    - Metadata keys are case-sensitive -- `Category` and `category` are different fields
    - The `nResults` default is 10 if not specified in `query()`
    - Multimodal embedding (images + text) is currently Python-only -- TypeScript support is not yet available
    
    </red_flags>
    
    ---
    
    <critical_reminders>
    
    ## CRITICAL REMINDERS
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST install `@chroma-core/default-embed` alongside `chromadb` -- the default embedding function ships as a separate package since v3)**
    
    **(You MUST access query results as nested arrays -- `results.ids[0]`, `results.documents[0]` -- because Chroma batches queries and returns `string[][]` not `string[]`)**
    
    **(You MUST use the `configuration` parameter for HNSW settings -- the legacy `metadata: { "hnsw:space": "cosine" }` approach is deprecated)**
    
    **(You MUST use flat metadata values only (string, number, boolean, typed arrays) -- nested objects are not supported and will be rejected)**
    
    **Failure to follow these rules will cause embedding failures, incorrect result access, deprecated configuration warnings, and rejected metadata.**
    
    </critical_reminders>
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related