Claude Skill

api-vector-db-weaviate

Weaviate vector database patterns with weaviate-client v3 -- collection management, vectorizer modules, hybrid search, filtering, generative search (RAG), multi-tenancy, batch imports

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download agents-inc-skills-dist_plugins_api-vector-db-weaviate_skills_api-vector-db-weaviate-3a51ef5.zip · 20 KB
Part of agents-inc/skills — 130 skills

Install

skills CLI npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/api-vector-db-weaviate/skills/api-vector-db-weaviate
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
Git git clone https://github.com/agents-inc/skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Weaviate Patterns

Quick Guide: Use Weaviate for semantic search and RAG applications. Use weaviate-client (v3.x) as the TypeScript client -- it uses gRPC for performance and provides full type safety with generics. Connect via connectToWeaviateCloud() for managed instances or connectToLocal() for Docker. Collections are the central abstraction -- configure vectorizers at collection level, not per-query. Use collection.query.* for search, collection.generate.* for RAG, and collection.data.* for CRUD. Always call client.close() when done. Increase query timeout to 60s+ when using generative search. The v3 client does NOT support browsers or Embedded Weaviate.


<critical_requirements>

CRITICAL: Before Using This Skill

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST call client.close() when done with the Weaviate client -- it maintains persistent gRPC connections that will leak if not closed)

(You MUST configure vectorizers at the COLLECTION level during client.collections.create() -- you cannot add a vectorizer after creation, only add new named vectors)

(You MUST use a SEPARATE client.collections.use() call with .withTenant() for multi-tenant queries -- queries without tenant context on multi-tenant collections will fail)

(You MUST increase query timeout to 60+ seconds when using generate.* (RAG) submodule -- generative model calls are slow and the default timeout causes failures)

</critical_requirements>


Examples

Additional resources:

  • reference.md -- API cheat sheet, vectorizer comparison, data types, decision frameworks

Auto-detection: Weaviate, weaviate-client, connectToWeaviateCloud, connectToLocal, nearText, nearVector, hybrid search, bm25, vector database, semantic search, RAG, generative search, generate.nearText, insertMany, vectorizer, text2vec, multi-tenancy, withTenant, collection.query, collection.generate, collection.data

When to use:

  • Semantic search over text, images, or multimodal data
  • Retrieval Augmented Generation (RAG) with built-in generative search
  • Hybrid search combining vector similarity and keyword (BM25) ranking
  • Multi-tenant applications needing isolated vector stores per customer
  • Applications requiring built-in vectorization (no external embedding pipeline)
  • Real-time similarity search with filtering on structured properties

Key patterns covered:

  • weaviate-client v3 connection setup and configuration
  • Collection management with vectorizer modules (text2vec-openai, text2vec-cohere, etc.)
  • Object CRUD (insert, insertMany, update, replace, deleteById, deleteMany)
  • Search types (nearText, nearVector, hybrid, bm25, fetchObjects)
  • Filtering with operators (equal, greaterThan, like, containsAny, and/or/not)
  • Generative search (RAG) with singlePrompt and groupedTask
  • Multi-tenancy with tenant lifecycle management
  • Batch imports with insertMany and error handling
  • Cross-references between collections
  • Named vectors for multi-vector collections

When NOT to use:

  • Relational data with complex joins (use a relational database)
  • Full-text search without vector component (use a dedicated search engine)
  • Key-value caching (use a key-value store)
  • Time-series data (use a time-series database)
  • Graph traversal queries (use a graph database)
  • Browser-side applications (v3 client is Node.js only)



<decision_framework>

Decision Framework

Which Search Type?

What kind of search do I need?
├─ Natural language query, semantic meaning? -> nearText (requires vectorizer module)
├─ Have pre-computed vector embedding? -> nearVector
├─ Exact keyword matching? -> bm25
├─ Both semantic and keyword relevance? -> hybrid (alpha controls blend)
├─ Just list/filter objects without search? -> fetchObjects
└─ Search + LLM generation? -> generate.nearText / generate.hybrid

Which Vectorizer?

Which vectorizer module should I use?
├─ OpenAI models (text-embedding-3-small/large)? -> text2VecOpenAI
├─ Cohere models (embed-v3)? -> text2VecCohere
├─ Self-hosted models? -> text2VecOllama or text2VecTransformers
├─ Bring your own embeddings? -> none (use selfProvided for named vectors)
├─ Multimodal (images + text)? -> multi2VecClip or multi2VecBind
└─ Multiple embedding strategies? -> Named vectors (array of vectorizers)

Single vs Named Vectors?

How many vector representations do I need?
├─ One embedding per object (most common)? -> Single default vectorizer
├─ Different embeddings for different properties? -> Named vectors
├─ Mix of auto-vectorized and self-provided? -> Named vectors with selfProvided
└─ Different models for different search use cases? -> Named vectors

When to Use Multi-Tenancy?

Do I need data isolation?
├─ Each customer/user needs isolated data? -> Enable multi-tenancy
├─ Shared dataset, filter by user? -> Single tenant with filters
├─ Need to offload inactive tenants? -> Multi-tenancy with tenant states
└─ Small number of distinct datasets? -> Separate collections may be simpler

</decision_framework>


<red_flags>

RED FLAGS

High Priority Issues:

  • Missing client.close() -- gRPC connections persist and leak memory/file descriptors
  • Trying to add a default vectorizer after collection creation -- vectorizer must be configured in create(). Only named vectors can be added later with config.addVector()
  • Querying a multi-tenant collection without .withTenant() -- all operations fail with an error
  • Using default query timeout with generate.* -- generative calls need 60+ seconds; default is often too short

Medium Priority Issues:

  • Using replace() when update() is intended -- replace deletes all properties not included in the call; update merges
  • Not checking insertMany response for errors -- partial failures are silent; check response.hasErrors and response.errors
  • Passing alpha: 1.0 to hybrid search -- equivalent to pure vector search; use nearText instead for clarity
  • Not specifying targetVector with named vectors -- queries default to the first vector, which may not be the intended one

Common Mistakes:

  • Using v2 class-based API (client.schema.classCreator()) with v3 client -- the API is completely different; v3 uses client.collections.create()
  • Forgetting to pass API key headers for vectorizer modules -- X-OpenAI-Api-Key, X-Cohere-Api-Key etc. must be in connection headers
  • Using connectToWCS() (deprecated) instead of connectToWeaviateCloud()
  • Adding a property after data import without reindexing -- pre-existing objects won't have that property indexed

Gotchas & Edge Cases:

  • insertMany uses server-side batching but the TS client does NOT have a streaming batch API -- for very large imports (100K+), chunk into batches of 100-1000 objects
  • Filters.and() and Filters.or() take a flat list of filter conditions, NOT nested arrays -- Filters.and(a, b, c) not Filters.and([a, b, c])
  • fetchObjects() without limit returns 25 objects by default (server-side default), not all objects
  • Property names in Weaviate must start with a lowercase letter -- the client silently lowercases the first character
  • distance metadata varies by vector distance metric -- cosine distance range [0, 2], not [0, 1]
  • deleteMany has a server-side maximum of 10,000 objects per call (configurable via QUERY_MAXIMUM_RESULTS)
  • Weaviate auto-detects property types on first insert if not defined in the schema -- this can cause type mismatches if first object has atypical data
  • fetchObjectById returns null for non-existent IDs, not an empty object -- always check for null before accessing properties
  • Cross-references in multi-tenant collections can only reference objects in the same tenant or in non-multi-tenant collections

</red_flags>


<critical_reminders>

CRITICAL REMINDERS

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST call client.close() when done with the Weaviate client -- it maintains persistent gRPC connections that will leak if not closed)

(You MUST configure vectorizers at the COLLECTION level during client.collections.create() -- you cannot add a vectorizer after creation, only add new named vectors)

(You MUST use a SEPARATE client.collections.use() call with .withTenant() for multi-tenant queries -- queries without tenant context on multi-tenant collections will fail)

(You MUST increase query timeout to 60+ seconds when using generate.* (RAG) submodule -- generative model calls are slow and the default timeout causes failures)

Failure to follow these rules will cause connection leaks, missing vectorization, multi-tenant query failures, and RAG timeouts.

</critical_reminders>

Files (skills)
  • examples
    • core.md 12 KB
      # Weaviate -- Core Patterns
      
      > Connection setup, collection management, object CRUD, and basic queries. Reference from [SKILL.md](../SKILL.md).
      
      **Related examples:**
      
      - [search.md](search.md) -- nearText, nearVector, hybrid, bm25, filters, generative search (RAG)
      - [multi-tenancy.md](multi-tenancy.md) -- Tenant management, batch imports, cross-references
      
      ---
      
      ## Connection: Weaviate Cloud
      
      ```typescript
      import weaviate from "weaviate-client";
      
      const QUERY_TIMEOUT_SECONDS = 30;
      const INSERT_TIMEOUT_SECONDS = 120;
      const INIT_TIMEOUT_SECONDS = 5;
      
      async function createCloudClient() {
        const url = process.env.WEAVIATE_URL;
        const apiKey = process.env.WEAVIATE_API_KEY;
        if (!url || !apiKey) {
          throw new Error(
            "WEAVIATE_URL and WEAVIATE_API_KEY environment variables are required",
          );
        }
      
        const client = await weaviate.connectToWeaviateCloud(url, {
          authCredentials: new weaviate.ApiKey(apiKey),
          headers: {
            "X-OpenAI-Api-Key": process.env.OPENAI_API_KEY ?? "",
          },
          timeout: {
            query: QUERY_TIMEOUT_SECONDS,
            insert: INSERT_TIMEOUT_SECONDS,
            init: INIT_TIMEOUT_SECONDS,
          },
        });
      
        return client;
      }
      
      export { createCloudClient };
      ```
      
      **Why good:** Environment variable validation, named timeout constants, API key header for vectorizer module, `authCredentials` for Weaviate Cloud authentication
      
      ```typescript
      // Bad Example -- Leaked connections, missing headers
      import weaviate from "weaviate-client";
      
      const client = await weaviate.connectToWeaviateCloud(
        "https://my-instance.weaviate.network",
        {
          authCredentials: new weaviate.ApiKey("hardcoded-key"),
        },
      );
      // No X-OpenAI-Api-Key header -- nearText queries will fail silently
      // No client.close() -- gRPC connection leaks
      // Hardcoded credentials -- leak in version control
      ```
      
      **Why bad:** Hardcoded credentials, missing vectorizer API key header causes silent search failures, no `client.close()` leaks gRPC connections
      
      ---
      
      ## Connection: Local Docker
      
      ```typescript
      import weaviate from "weaviate-client";
      
      const LOCAL_HTTP_PORT = 8080;
      const LOCAL_GRPC_PORT = 50051;
      
      async function createLocalClient() {
        const client = await weaviate.connectToLocal({
          port: LOCAL_HTTP_PORT,
          grpcPort: LOCAL_GRPC_PORT,
          headers: {
            "X-OpenAI-Api-Key": process.env.OPENAI_API_KEY ?? "",
          },
        });
      
        return client;
      }
      
      export { createLocalClient };
      ```
      
      **Why good:** Explicit ports with named constants, headers still provided for vectorizer modules even locally
      
      ---
      
      ## Connection: Cleanup Pattern
      
      Always close the client when done. Use try/finally in scripts or shutdown hooks in servers.
      
      ```typescript
      // Script pattern -- try/finally
      const client = await createCloudClient();
      try {
        // ... perform operations
      } finally {
        client.close();
      }
      ```
      
      ```typescript
      // Server pattern -- shutdown hook
      const client = await createCloudClient();
      
      process.on("SIGTERM", () => {
        client.close();
        process.exit(0);
      });
      
      process.on("SIGINT", () => {
        client.close();
        process.exit(0);
      });
      ```
      
      ---
      
      ## Collection: Basic Creation
      
      ```typescript
      import { vectors, dataType, generative } from "weaviate-client";
      
      async function createArticleCollection(client: WeaviateClient) {
        await client.collections.create({
          name: "Article",
          vectorizers: vectors.text2VecOpenAI({
            model: "text-embedding-3-small",
          }),
          generative: generative.openAI({
            model: "gpt-4o",
          }),
          properties: [
            { name: "title", dataType: dataType.TEXT },
            { name: "body", dataType: dataType.TEXT },
            { name: "category", dataType: dataType.TEXT },
            { name: "author", dataType: dataType.TEXT },
            { name: "publishedAt", dataType: dataType.DATE },
            { name: "wordCount", dataType: dataType.INT },
          ],
        });
      }
      
      export { createArticleCollection };
      ```
      
      **Why good:** Explicit property data types prevent auto-detection surprises, vectorizer and generative model configured at creation
      
      ---
      
      ## Collection: Named Vectors
      
      Use named vectors when objects need multiple embedding representations (e.g., title vs body, or different models).
      
      ```typescript
      import { vectors, dataType, configure } from "weaviate-client";
      
      await client.collections.create({
        name: "Product",
        vectorizers: [
          vectors.text2VecOpenAI({
            name: "title_vector",
            sourceProperties: ["title", "brand"],
            vectorIndexConfig: configure.vectorIndex.hnsw(),
          }),
          vectors.text2VecOpenAI({
            name: "description_vector",
            sourceProperties: ["description"],
            vectorIndexConfig: configure.vectorIndex.hnsw(),
          }),
          vectors.selfProvided({
            name: "image_vector",
            vectorIndexConfig: configure.vectorIndex.hnsw(),
          }),
        ],
        properties: [
          { name: "title", dataType: dataType.TEXT },
          { name: "brand", dataType: dataType.TEXT },
          { name: "description", dataType: dataType.TEXT },
          { name: "price", dataType: dataType.NUMBER },
        ],
      });
      ```
      
      **Why good:** Separate vectors for different semantic fields, `selfProvided` for externally computed image embeddings, `sourceProperties` controls which fields each vector covers
      
      ```typescript
      // Bad Example -- Named vectors without targetVector in query
      const products = client.collections.use("Product");
      const result = await products.query.nearText("leather jacket", { limit: 5 });
      // Defaults to first named vector -- may search title_vector when you wanted description_vector
      ```
      
      **Why bad:** Without `targetVector`, Weaviate uses the first named vector, which may not be the intended search field
      
      ---
      
      ## Collection: Vectorizer Property Controls
      
      Skip vectorization for properties that shouldn't influence search, or include the property name in the embedding.
      
      ```typescript
      import { vectors, dataType, tokenization } from "weaviate-client";
      
      await client.collections.create({
        name: "Document",
        vectorizers: vectors.text2VecOpenAI(),
        properties: [
          {
            name: "title",
            dataType: dataType.TEXT,
            vectorizePropertyName: true, // "title: My Article" vectorized together
            tokenization: tokenization.LOWERCASE,
          },
          {
            name: "body",
            dataType: dataType.TEXT,
            tokenization: tokenization.WHITESPACE,
          },
          {
            name: "internalId",
            dataType: dataType.TEXT,
            skipVectorization: true, // Don't include in embedding
          },
          {
            name: "createdAt",
            dataType: dataType.DATE,
            skipVectorization: true,
          },
        ],
      });
      ```
      
      **Why good:** `skipVectorization` on non-semantic fields prevents noise in embeddings, `vectorizePropertyName` adds context for short fields, tokenization controls keyword search behavior
      
      ---
      
      ## Collection: Check and Delete
      
      ```typescript
      // Check if collection exists before creating
      const exists = await client.collections.exists("Article");
      if (!exists) {
        await client.collections.create({ name: "Article" /* ... */ });
      }
      
      // Get collection configuration
      const articles = client.collections.use("Article");
      const config = await articles.config.get();
      console.log(config);
      
      // List all collections
      const allCollections = await client.collections.listAll();
      
      // Delete collection (permanent -- deletes all data)
      await client.collections.delete("Article");
      ```
      
      ---
      
      ## Object: Insert Single
      
      ```typescript
      const articles = client.collections.use("Article");
      
      const uuid = await articles.data.insert({
        title: "Introduction to Vector Databases",
        body: "Vector databases store data alongside embeddings...",
        category: "technology",
        author: "Jane Smith",
        publishedAt: new Date("2024-06-15").toISOString(),
        wordCount: 1500,
      });
      
      console.log("Inserted:", uuid);
      ```
      
      ---
      
      ## Object: Insert with Explicit ID
      
      Use `generateUuid5` for deterministic, idempotent IDs based on content.
      
      ```typescript
      import { generateUuid5 } from "weaviate-client";
      
      const articles = client.collections.use("Article");
      const data = {
        title: "Deterministic ID Example",
        body: "Content here...",
        category: "tutorial",
      };
      
      const deterministicId = generateUuid5("Article", JSON.stringify(data));
      
      const uuid = await articles.data.insert({
        properties: data,
        id: deterministicId,
      });
      ```
      
      **Why good:** `generateUuid5` produces the same UUID for the same input -- safe to retry without creating duplicates
      
      ---
      
      ## Object: Insert with Pre-Computed Vector
      
      ```typescript
      const articles = client.collections.use("Article");
      const EMBEDDING_DIM = 1536;
      
      // Single default vector
      await articles.data.insert({
        properties: { title: "Custom Vector Example", body: "..." },
        vectors: myEmbeddingArray, // number[] matching collection vector dimension
      });
      
      // Named vectors
      const products = client.collections.use("Product");
      await products.data.insert({
        properties: { title: "Jacket", description: "Warm winter jacket" },
        vectors: {
          title_vector: titleEmbedding,
          description_vector: descEmbedding,
          image_vector: imageEmbedding,
        },
      });
      ```
      
      ---
      
      ## Object: Update vs Replace
      
      ```typescript
      const articles = client.collections.use("Article");
      const objectId = "ed89d9e7-4c9d-4a6a-8d20-095cb0026f54";
      
      // Update (merge) -- preserves properties NOT included
      await articles.data.update({
        id: objectId,
        properties: {
          wordCount: 2000, // Only this property changes
        },
      });
      
      // Replace (overwrite) -- DELETES properties NOT included
      await articles.data.replace({
        id: objectId,
        properties: {
          title: "Replaced Title",
          body: "Replaced body",
          // category, author, publishedAt, wordCount are DELETED
        },
      });
      ```
      
      **Why good:** Clear distinction between merge and overwrite semantics
      
      ```typescript
      // Bad Example -- Using replace when update was intended
      await articles.data.replace({
        id: objectId,
        properties: { wordCount: 2000 },
      });
      // All other properties (title, body, category, etc.) are now DELETED
      ```
      
      **Why bad:** `replace` deletes every property not explicitly provided -- use `update` for partial changes
      
      ---
      
      ## Object: Delete
      
      ```typescript
      const articles = client.collections.use("Article");
      
      // Delete single object by ID
      await articles.data.deleteById("ed89d9e7-4c9d-4a6a-8d20-095cb0026f54");
      
      // Delete multiple objects matching a filter
      const deleteResult = await articles.data.deleteMany(
        articles.filter.byProperty("category").equal("draft"),
      );
      console.log("Deleted:", deleteResult);
      
      // Delete by ID list
      const idsToDelete = ["id-1", "id-2", "id-3"];
      await articles.data.deleteMany(articles.filter.byId().containsAny(idsToDelete));
      
      // Dry run -- check what would be deleted without deleting
      const dryResult = await articles.data.deleteMany(
        articles.filter.byProperty("category").equal("old"),
        { dryRun: true, verbose: true },
      );
      console.log("Would delete:", dryResult);
      ```
      
      ---
      
      ## Object: Iterate Over Entire Collection
      
      Use `iterator()` to process all objects without loading everything into memory.
      
      ```typescript
      const articles = client.collections.use("Article");
      
      for await (const item of articles.iterator()) {
        console.log(item.uuid, item.properties.title);
      }
      
      // With specific return properties
      for await (const item of articles.iterator({
        returnProperties: ["title", "category"],
      })) {
        processArticle(item);
      }
      ```
      
      **Why good:** Memory-efficient -- streams objects in batches internally, unlike `fetchObjects` which loads a fixed page
      
      ---
      
      ## TypeScript Generics
      
      Use generics for compile-time type safety on collection objects.
      
      ```typescript
      interface Article {
        title: string;
        body: string;
        category: string;
        author: string;
        publishedAt: string;
        wordCount: number;
      }
      
      const articles = client.collections.use<Article>("Article");
      
      // Now insert and query methods are typed
      const uuid = await articles.data.insert({
        title: "Typed Insert",
        body: "This is type-checked at compile time",
        category: "tutorial",
        author: "Developer",
        publishedAt: new Date().toISOString(),
        wordCount: 500,
        // misspelledField: "error" // TypeScript error!
      });
      
      const result = await articles.query.fetchObjects({ limit: 5 });
      for (const obj of result.objects) {
        // obj.properties is typed as Article
        console.log(obj.properties.title); // string, not unknown
      }
      ```
      
      **Why good:** Compile-time type checking catches property name typos, wrong types, and missing fields before runtime
      
      ---
      
      _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_
      
    • multi-tenancy.md 11 KB
      # Weaviate -- Multi-Tenancy & Batch Examples
      
      > Tenant management, batch imports, and cross-references. See [core.md](core.md) for connection and collection setup.
      
      **Related examples:**
      
      - [core.md](core.md) -- Connection, collection setup, object CRUD
      - [search.md](search.md) -- nearText, nearVector, hybrid, bm25, filters, generative search
      
      ---
      
      ## Pattern 1: Enable Multi-Tenancy
      
      Multi-tenancy must be enabled at collection creation time.
      
      ```typescript
      import weaviate from "weaviate-client";
      import { vectors, dataType } from "weaviate-client";
      
      await client.collections.create({
        name: "CustomerDocument",
        multiTenancy: weaviate.configure.multiTenancy({
          enabled: true,
          autoTenantCreation: true, // Create tenants on first insert
        }),
        vectorizers: vectors.text2VecOpenAI(),
        properties: [
          { name: "title", dataType: dataType.TEXT },
          { name: "content", dataType: dataType.TEXT },
          { name: "docType", dataType: dataType.TEXT },
        ],
      });
      ```
      
      **Why good:** `autoTenantCreation: true` avoids manual tenant creation before inserts, useful for dynamic SaaS applications
      
      ```typescript
      // Bad Example -- Querying multi-tenant collection without tenant context
      const docs = client.collections.use("CustomerDocument");
      const result = await docs.query.fetchObjects({ limit: 10 });
      // FAILS: multi-tenant collection requires .withTenant()
      ```
      
      **Why bad:** All operations on multi-tenant collections require `.withTenant()` -- queries without tenant context throw an error
      
      ---
      
      ## Pattern 2: Tenant Lifecycle Management
      
      ```typescript
      const docs = client.collections.use("CustomerDocument");
      
      // Create tenants manually
      await docs.tenants.create([
        { name: "tenant-acme" },
        { name: "tenant-globex" },
        { name: "tenant-initech" },
      ]);
      
      // List all tenants
      const allTenants = await docs.tenants.get();
      console.log("Tenants:", Object.keys(allTenants));
      
      // Get specific tenants
      const specific = await docs.tenants.getByNames([
        "tenant-acme",
        "tenant-globex",
      ]);
      console.log(specific);
      
      // Get single tenant
      const acme = await docs.tenants.getByName("tenant-acme");
      console.log(acme);
      
      // Remove tenants (non-existent names are silently ignored)
      await docs.tenants.remove([
        { name: "tenant-initech" },
        { name: "tenant-nonexistent" }, // Ignored
      ]);
      ```
      
      ---
      
      ## Pattern 3: Tenant State Management
      
      Deactivate tenants to free memory, reactivate when needed.
      
      ```typescript
      const docs = client.collections.use("CustomerDocument");
      
      // Deactivate tenant (data stays on disk, freed from memory)
      await docs.tenants.update({
        name: "tenant-acme",
        activityStatus: "INACTIVE",
      });
      
      // Offload tenant to cold storage (cloud deployments only)
      await docs.tenants.update({
        name: "tenant-globex",
        activityStatus: "OFFLOADED",
      });
      
      // Reactivate tenant before querying
      await docs.tenants.update({
        name: "tenant-acme",
        activityStatus: "ACTIVE",
      });
      
      // Now queries work again
      const acmeDocs = docs.withTenant("tenant-acme");
      const result = await acmeDocs.query.fetchObjects({ limit: 10 });
      ```
      
      **Why good:** Tenant states manage memory for large multi-tenant deployments -- inactive tenants consume no memory
      
      ### Auto-Tenant Activation
      
      ```typescript
      import weaviate from "weaviate-client";
      
      // Enable auto-activation so queries automatically activate inactive tenants
      const docs = client.collections.use("CustomerDocument");
      await docs.config.update({
        multiTenancy: weaviate.reconfigure.multiTenancy({
          autoTenantActivation: true,
        }),
      });
      ```
      
      ---
      
      ## Pattern 4: Multi-Tenant CRUD Operations
      
      All data operations require tenant context via `.withTenant()`.
      
      ```typescript
      const docs = client.collections.use("CustomerDocument");
      const acmeDocs = docs.withTenant("tenant-acme");
      
      // Insert
      const uuid = await acmeDocs.data.insert({
        title: "Q3 Report",
        content: "Revenue increased by 15%...",
        docType: "report",
      });
      
      // Query
      const result = await acmeDocs.query.nearText("quarterly revenue", {
        limit: 5,
        returnMetadata: ["distance"],
      });
      
      // Update
      await acmeDocs.data.update({
        id: uuid,
        properties: { docType: "financial-report" },
      });
      
      // Delete
      await acmeDocs.data.deleteById(uuid);
      
      // Delete many
      await acmeDocs.data.deleteMany(
        acmeDocs.filter.byProperty("docType").equal("draft"),
      );
      ```
      
      ---
      
      ## Pattern 5: Batch Import with insertMany
      
      Use `insertMany` for bulk data loading. Always check for errors -- partial failures are silent.
      
      ```typescript
      const articles = client.collections.use("Article");
      const BATCH_SIZE = 100;
      
      interface ArticleData {
        title: string;
        body: string;
        category: string;
      }
      
      async function batchImport(
        collection: ReturnType<typeof client.collections.use>,
        data: ArticleData[],
      ) {
        // Process in chunks to avoid memory pressure
        for (let i = 0; i < data.length; i += BATCH_SIZE) {
          const chunk = data.slice(i, i + BATCH_SIZE);
      
          const response = await collection.data.insertMany(chunk);
      
          // CRITICAL: Check for partial failures
          if (response.hasErrors) {
            for (const err of Object.values(response.errors)) {
              console.error(`Insert error at index ${err.index}:`, err.message);
            }
            throw new Error(
              `Batch insert failed with ${Object.keys(response.errors).length} errors`,
            );
          }
      
          console.log(
            `Imported ${Math.min(i + BATCH_SIZE, data.length)}/${data.length}`,
          );
        }
      }
      
      export { batchImport };
      ```
      
      **Why good:** Chunked processing avoids memory pressure for large datasets, `hasErrors` check catches partial failures that would otherwise be silent
      
      ```typescript
      // Bad Example -- Ignoring insertMany errors
      const response = await articles.data.insertMany(largeDataset);
      // response.hasErrors might be true, but we never check
      // Some objects silently failed to insert
      ```
      
      **Why bad:** `insertMany` can partially fail -- some objects insert, some don't. Without checking `hasErrors`, you have incomplete data
      
      ---
      
      ## Pattern 6: Batch Import with Deterministic IDs
      
      Use `generateUuid5` for idempotent imports -- safe to retry without duplicates.
      
      ```typescript
      import { generateUuid5 } from "weaviate-client";
      
      const COLLECTION_NAME = "Article";
      
      function prepareObjects(data: ArticleData[]) {
        return data.map((item) => ({
          properties: item,
          id: generateUuid5(COLLECTION_NAME, item.title), // Deterministic UUID from title
        }));
      }
      
      const articles = client.collections.use(COLLECTION_NAME);
      const objects = prepareObjects(rawData);
      const response = await articles.data.insertMany(objects);
      
      if (response.hasErrors) {
        // Objects with duplicate IDs will fail -- expected on retry
        const realErrors = Object.values(response.errors).filter(
          (err) => !err.message.includes("already exists"),
        );
        if (realErrors.length > 0) {
          throw new Error(`Import failed: ${realErrors.length} non-duplicate errors`);
        }
      }
      ```
      
      **Why good:** Deterministic IDs make imports idempotent, filtering "already exists" errors enables safe retries
      
      ---
      
      ## Pattern 7: Batch Import with Custom Vectors
      
      ```typescript
      const products = client.collections.use("Product");
      const BATCH_SIZE = 100;
      
      interface ProductWithEmbedding {
        properties: { title: string; description: string; price: number };
        vectors: number[];
      }
      
      async function importWithVectors(data: ProductWithEmbedding[]) {
        for (let i = 0; i < data.length; i += BATCH_SIZE) {
          const chunk = data.slice(i, i + BATCH_SIZE);
          const response = await products.data.insertMany(chunk);
      
          if (response.hasErrors) {
            console.error("Batch errors:", response.errors);
            throw new Error("Import with vectors failed");
          }
        }
      }
      ```
      
      ### Named Vectors in Batch
      
      ```typescript
      const multiVectorObjects = data.map((item) => ({
        properties: { title: item.title, description: item.description },
        vectors: {
          title_vector: item.titleEmbedding,
          description_vector: item.descEmbedding,
        },
      }));
      
      await products.data.insertMany(multiVectorObjects);
      ```
      
      ---
      
      ## Pattern 8: Cross-References
      
      Link objects across collections. Requires adding a reference property to the collection.
      
      ```typescript
      import { dataType } from "weaviate-client";
      
      // Create collections
      await client.collections.create({
        name: "Author",
        vectorizers: vectors.text2VecOpenAI(),
        properties: [
          { name: "name", dataType: dataType.TEXT },
          { name: "bio", dataType: dataType.TEXT },
        ],
      });
      
      await client.collections.create({
        name: "Article",
        vectorizers: vectors.text2VecOpenAI(),
        properties: [
          { name: "title", dataType: dataType.TEXT },
          { name: "body", dataType: dataType.TEXT },
        ],
      });
      
      // Add cross-reference property
      const articles = client.collections.use("Article");
      await articles.config.addReference({
        name: "writtenBy",
        targetCollection: "Author",
      });
      ```
      
      ### Create and Query References
      
      ```typescript
      const authors = client.collections.use("Author");
      const articles = client.collections.use("Article");
      
      // Insert an author
      const authorId = await authors.data.insert({
        name: "Jane Smith",
        bio: "AI researcher and writer",
      });
      
      // Insert an article with reference
      const articleId = await articles.data.insert({
        properties: { title: "AI Ethics", body: "..." },
        references: { writtenBy: authorId },
      });
      
      // OR add reference after creation
      await articles.data.referenceAdd({
        fromUuid: articleId,
        fromProperty: "writtenBy",
        to: authorId,
      });
      
      // Query with resolved references
      const result = await articles.query.fetchObjects({
        limit: 5,
        returnReferences: [
          {
            linkOn: "writtenBy",
            returnProperties: ["name", "bio"],
          },
        ],
      });
      
      for (const obj of result.objects) {
        console.log("Article:", obj.properties.title);
        console.log(
          "Author:",
          obj.references?.writtenBy?.objects[0]?.properties.name,
        );
      }
      ```
      
      ### Replace and Delete References
      
      ```typescript
      // Replace all references on a property
      await articles.data.referenceReplace({
        fromUuid: articleId,
        fromProperty: "writtenBy",
        to: [newAuthorId], // Replaces all existing references
      });
      
      // Delete specific reference
      await articles.data.referenceDelete({
        fromUuid: articleId,
        fromProperty: "writtenBy",
        to: authorId,
      });
      ```
      
      ---
      
      ## Pattern 9: Multi-Tenant Cross-References
      
      Cross-references in multi-tenant collections can only point to objects in the same tenant or in non-multi-tenant collections.
      
      ```typescript
      const tenantDocs = docs.withTenant("tenant-acme");
      
      // Add reference property
      await docs.config.addReference({
        name: "hasCategory",
        targetCollection: "Category", // Non-multi-tenant collection
      });
      
      // Create cross-reference
      await tenantDocs.data.referenceAdd({
        fromUuid: documentId,
        fromProperty: "hasCategory",
        to: categoryId, // Category object in non-MT collection
      });
      ```
      
      **Why good:** Non-multi-tenant reference targets work across all tenants (shared lookup data)
      
      ```typescript
      // Bad Example -- Cross-tenant reference
      const acmeDocs = docs.withTenant("tenant-acme");
      await acmeDocs.data.referenceAdd({
        fromUuid: acmeDocId,
        fromProperty: "relatedDoc",
        to: globexDocId, // Object in tenant-globex -- FAILS
      });
      ```
      
      **Why bad:** Multi-tenant cross-references cannot span tenants -- only same-tenant or non-multi-tenant targets
      
      ---
      
      _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_
      
    • search.md 12.6 KB
      # Weaviate -- Search & Filtering Examples
      
      > Search types, filtering, and generative search (RAG). See [core.md](core.md) for connection and collection setup.
      
      **Related examples:**
      
      - [core.md](core.md) -- Connection, collection setup, object CRUD
      - [multi-tenancy.md](multi-tenancy.md) -- Tenant management, batch imports, cross-references
      
      ---
      
      ## Pattern 1: nearText (Semantic Search)
      
      Search by natural language query. Requires a vectorizer module configured on the collection.
      
      ```typescript
      const articles = client.collections.use("Article");
      const SEARCH_LIMIT = 10;
      const MAX_DISTANCE = 0.3;
      
      // Basic nearText
      const result = await articles.query.nearText("climate change policy", {
        limit: SEARCH_LIMIT,
        returnMetadata: ["distance"],
      });
      
      for (const obj of result.objects) {
        console.log(obj.properties.title, "distance:", obj.metadata?.distance);
      }
      ```
      
      **Why good:** `returnMetadata: ['distance']` enables relevance debugging, named constant for limit
      
      ### With Distance Threshold
      
      ```typescript
      // Only return results within distance threshold
      const result = await articles.query.nearText("renewable energy", {
        distance: MAX_DISTANCE,
        returnMetadata: ["distance"],
      });
      ```
      
      **When to use:** When you want to ensure a minimum relevance level rather than a fixed count.
      
      ### With Named Vector Target
      
      ```typescript
      const products = client.collections.use("Product");
      
      const result = await products.query.nearText("comfortable running shoes", {
        limit: SEARCH_LIMIT,
        targetVector: "description_vector", // Explicitly target the description embedding
        returnMetadata: ["distance"],
      });
      ```
      
      **When to use:** Collections with named vectors -- always specify `targetVector` to avoid defaulting to the first vector.
      
      ---
      
      ## Pattern 2: nearVector (Vector Search)
      
      Search with a pre-computed embedding vector. No vectorizer module needed.
      
      ```typescript
      const articles = client.collections.use("Article");
      const SEARCH_LIMIT = 5;
      
      // Use your own embedding
      const queryVector = await myEmbeddingModel.encode("search query");
      
      const result = await articles.query.nearVector(queryVector, {
        limit: SEARCH_LIMIT,
        returnMetadata: ["distance"],
      });
      
      for (const obj of result.objects) {
        console.log(obj.properties.title, obj.metadata?.distance);
      }
      ```
      
      **When to use:** When you compute embeddings externally (e.g., using your own embedding service or a model Weaviate doesn't support).
      
      ---
      
      ## Pattern 3: BM25 (Keyword Search)
      
      Traditional keyword search based on term frequency. No vector component.
      
      ```typescript
      const articles = client.collections.use("Article");
      const SEARCH_LIMIT = 10;
      
      // Basic BM25
      const result = await articles.query.bm25("vector database comparison", {
        limit: SEARCH_LIMIT,
        returnMetadata: ["score"],
      });
      
      for (const obj of result.objects) {
        console.log(obj.properties.title, "score:", obj.metadata?.score);
      }
      ```
      
      ### BM25 with Property Targeting and Boosting
      
      ```typescript
      // Search specific properties with weight boosting
      const result = await articles.query.bm25("machine learning", {
        limit: SEARCH_LIMIT,
        queryProperties: ["title^3", "body"], // title weighted 3x
        returnMetadata: ["score"],
      });
      ```
      
      **Why good:** Property boosting (`title^3`) prioritizes matches in important fields
      
      ---
      
      ## Pattern 4: Hybrid Search
      
      Combines vector search and keyword search with configurable weighting.
      
      ```typescript
      import { Filters } from "weaviate-client";
      
      const articles = client.collections.use("Article");
      const SEARCH_LIMIT = 10;
      const HYBRID_ALPHA = 0.75; // 0.0 = pure keyword, 1.0 = pure vector
      
      const result = await articles.query.hybrid("neural network architecture", {
        alpha: HYBRID_ALPHA,
        limit: SEARCH_LIMIT,
        returnMetadata: ["score", "explainScore"],
      });
      
      for (const obj of result.objects) {
        console.log(obj.properties.title);
        console.log("Score:", obj.metadata?.score);
        console.log("Explanation:", obj.metadata?.explainScore);
      }
      ```
      
      **Why good:** `explainScore` shows the vector and keyword contributions, helping tune alpha
      
      ### Hybrid with Fusion Type
      
      ```typescript
      // RelativeScore uses actual similarity scores (default since v1.24)
      const result = await articles.query.hybrid("search query", {
        fusionType: "RelativeScore", // or "Ranked"
        alpha: HYBRID_ALPHA,
        limit: SEARCH_LIMIT,
      });
      ```
      
      **When to use:** `RelativeScore` (default) for most cases. `Ranked` when you want rank-based fusion regardless of actual similarity distances.
      
      ### Hybrid with Keyword Operator Control
      
      ```typescript
      import { Bm25Operator } from "weaviate-client";
      
      // Require at least 2 of 3 query tokens to match
      const result = await articles.query.hybrid("Australian mammal cute", {
        bm25Operator: Bm25Operator.or({ minimumMatch: 2 }),
        alpha: HYBRID_ALPHA,
        limit: SEARCH_LIMIT,
      });
      
      // Require ALL query tokens to match
      const strictResult = await articles.query.hybrid("neural network training", {
        bm25Operator: Bm25Operator.and(),
        alpha: HYBRID_ALPHA,
        limit: SEARCH_LIMIT,
      });
      ```
      
      ---
      
      ## Pattern 5: Filtering
      
      Filters work with all search methods. They narrow results AFTER vector/keyword retrieval.
      
      ### Single Filter
      
      ```typescript
      const articles = client.collections.use("Article");
      
      const result = await articles.query.nearText("technology trends", {
        limit: SEARCH_LIMIT,
        filters: articles.filter.byProperty("category").equal("technology"),
        returnMetadata: ["distance"],
      });
      ```
      
      ### Combined Filters
      
      ```typescript
      import { Filters } from "weaviate-client";
      
      const result = await articles.query.hybrid("AI research", {
        alpha: HYBRID_ALPHA,
        limit: SEARCH_LIMIT,
        filters: Filters.and(
          articles.filter.byProperty("category").equal("technology"),
          articles.filter.byProperty("wordCount").greaterThan(500),
          Filters.not(articles.filter.byProperty("author").equal("Bot")),
        ),
      });
      ```
      
      **Why good:** Flat argument list to `Filters.and()`, not nested arrays
      
      ```typescript
      // Bad Example -- Array argument to Filters.and
      const result = await articles.query.hybrid("search", {
        filters: Filters.and([filterA, filterB]), // WRONG: expects flat args, not array
      });
      ```
      
      **Why bad:** `Filters.and()` takes variadic arguments, not an array -- `Filters.and(a, b, c)` not `Filters.and([a, b, c])`
      
      ### Nested Filters
      
      ```typescript
      const result = await articles.query.fetchObjects({
        filters: Filters.and(
          articles.filter.byProperty("category").equal("technology"),
          Filters.or(
            articles.filter.byProperty("wordCount").greaterThan(2000),
            articles.filter.byProperty("wordCount").lessThan(500),
          ),
        ),
        limit: SEARCH_LIMIT,
      });
      ```
      
      ### Filter with Like (Wildcard)
      
      ```typescript
      const result = await articles.query.fetchObjects({
        filters: articles.filter.byProperty("title").like("*machine learning*"),
        limit: SEARCH_LIMIT,
      });
      ```
      
      ### Filter by Date
      
      ```typescript
      const cutoffDate = new Date("2024-01-01");
      const result = await articles.query.fetchObjects({
        filters: articles.filter.byProperty("publishedAt").greaterThan(cutoffDate),
        limit: SEARCH_LIMIT,
      });
      ```
      
      ### Filter by Cross-Reference
      
      ```typescript
      const result = await articles.query.fetchObjects({
        filters: articles.filter
          .byRef("hasCategory")
          .byProperty("title")
          .equal("Science"),
        returnReferences: [
          {
            linkOn: "hasCategory",
            returnProperties: ["title"],
          },
        ],
        limit: SEARCH_LIMIT,
      });
      ```
      
      ### Filter by Object ID
      
      ```typescript
      const result = await articles.query.fetchObjects({
        filters: articles.filter.byId().equal(targetId),
      });
      ```
      
      ### Filter by Property Length
      
      ```typescript
      // Second arg `true` enables length-based filtering
      const result = await articles.query.fetchObjects({
        filters: articles.filter.byProperty("title", true).greaterThan(20),
        limit: SEARCH_LIMIT,
      });
      ```
      
      ---
      
      ## Pattern 6: Pagination
      
      ### Offset-Based Pagination
      
      ```typescript
      const PAGE_SIZE = 20;
      
      async function fetchPage(collection: Collection, page: number) {
        return collection.query.fetchObjects({
          limit: PAGE_SIZE,
          offset: page * PAGE_SIZE,
        });
      }
      
      // Page 0 (first 20), Page 1 (next 20), etc.
      const page0 = await fetchPage(articles, 0);
      const page1 = await fetchPage(articles, 1);
      ```
      
      **When to use:** Simple pagination for user-facing pages. Offset has a 10,000 limit by default.
      
      ### Cursor-Based Pagination
      
      ```typescript
      // Use after cursor for deep pagination beyond offset limits
      const firstPage = await articles.query.fetchObjects({
        limit: PAGE_SIZE,
        returnMetadata: ["creationTime"],
      });
      
      // Get the last object's ID for cursor
      const lastId = firstPage.objects[firstPage.objects.length - 1]?.uuid;
      if (lastId) {
        const nextPage = await articles.query.fetchObjects({
          limit: PAGE_SIZE,
          after: lastId,
        });
      }
      ```
      
      **When to use:** Deep pagination or iterating large result sets where offset would be slow.
      
      ---
      
      ## Pattern 7: Return Options
      
      ```typescript
      const articles = client.collections.use("Article");
      
      // Return specific properties only
      const result = await articles.query.fetchObjects({
        limit: SEARCH_LIMIT,
        returnProperties: ["title", "category"],
      });
      
      // Include vector in response
      const withVector = await articles.query.fetchObjects({
        limit: SEARCH_LIMIT,
        includeVector: true,
      });
      
      // Return metadata (distance, score, creation time, etc.)
      const withMeta = await articles.query.nearText("search term", {
        limit: SEARCH_LIMIT,
        returnMetadata: ["distance", "creationTime"],
      });
      
      // Return cross-references
      const withRefs = await articles.query.fetchObjects({
        limit: SEARCH_LIMIT,
        returnReferences: [
          {
            linkOn: "hasCategory",
            returnProperties: ["title"],
          },
        ],
      });
      ```
      
      ---
      
      ## Pattern 8: Generative Search (RAG) -- Single Prompt
      
      Each search result is individually processed by the LLM. Requires a generative model configured on the collection.
      
      ```typescript
      const articles = client.collections.use("Article");
      const RAG_LIMIT = 5;
      
      const result = await articles.generate.nearText(
        "climate change solutions",
        {
          singlePrompt: "Summarize this article in one tweet: {title} - {body}",
        },
        {
          limit: RAG_LIMIT,
          returnMetadata: ["distance"],
        },
      );
      
      for (const obj of result.objects) {
        console.log("Source:", obj.properties.title);
        console.log("Generated:", obj.generative?.text);
      }
      ```
      
      **Why good:** `{title}` and `{body}` are interpolated from each object's properties -- no manual string building needed
      
      ### With Metadata and Debug
      
      ```typescript
      import { generativeParameters } from "weaviate-client";
      
      const result = await articles.generate.nearText(
        "renewable energy",
        {
          singlePrompt: {
            prompt: "Extract 3 key points from: {title} - {body}",
            metadata: true,
            debug: true,
          },
          config: generativeParameters.openAI({ model: "gpt-4o" }),
        },
        { limit: RAG_LIMIT },
      );
      
      for (const obj of result.objects) {
        console.log("Generated:", obj.generative?.text);
        console.log("Debug:", obj.generative?.debug);
        console.log("Metadata:", obj.generative?.metadata);
      }
      ```
      
      ---
      
      ## Pattern 9: Generative Search (RAG) -- Grouped Task
      
      All search results are sent to the LLM as a single context. One output for the entire group.
      
      ```typescript
      const articles = client.collections.use("Article");
      const RAG_LIMIT = 5;
      
      const result = await articles.generate.nearText(
        "artificial intelligence ethics",
        {
          groupedTask:
            "Based on these articles, write a summary of the main ethical concerns in AI.",
        },
        { limit: RAG_LIMIT },
      );
      
      // Single generated response for all results
      console.log("Combined summary:", result.generative?.text);
      
      // Individual source objects are still available
      for (const obj of result.objects) {
        console.log("Source:", obj.properties.title);
      }
      ```
      
      **Why good:** `result.generative?.text` gives one combined output, individual `obj.properties` still accessible for citations
      
      ### Grouped Task with Selected Properties
      
      ```typescript
      const result = await articles.generate.nearText(
        "machine learning applications",
        {
          groupedTask: "Compare and contrast these articles.",
          groupedProperties: ["title", "body"], // Only send title and body to LLM
        },
        { limit: RAG_LIMIT },
      );
      ```
      
      **When to use:** When you want to limit what properties the LLM sees, reducing token usage and improving focus.
      
      ---
      
      ## Pattern 10: Generative Search with Override Model
      
      Override the collection's default generative model at query time.
      
      ```typescript
      import { generativeParameters } from "weaviate-client";
      
      const result = await articles.generate.nearText(
        "quantum computing",
        {
          singlePrompt: "Explain this to a 5-year-old: {title}",
          config: generativeParameters.anthropic({
            model: "claude-haiku-4-5",
            maxTokens: 200,
          }),
        },
        { limit: RAG_LIMIT },
      );
      ```
      
      **When to use:** When you need a different model for specific queries (e.g., faster model for summaries, more capable model for analysis).
      
      ---
      
      _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_
      
  • reference.md 14.4 KB
    # Weaviate Quick Reference
    
    > API cheat sheet, vectorizer comparison, data types, and decision frameworks. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples.
    
    ---
    
    ## Connection Methods
    
    | Method                              | Use Case                                         | Example                                                                                             |
    | ----------------------------------- | ------------------------------------------------ | --------------------------------------------------------------------------------------------------- |
    | `connectToWeaviateCloud(url, opts)` | Weaviate Cloud managed instances                 | `weaviate.connectToWeaviateCloud(url, { authCredentials: new weaviate.ApiKey(key) })`               |
    | `connectToLocal(opts?)`             | Local Docker instances (default: localhost:8080) | `weaviate.connectToLocal()`                                                                         |
    | `connectToCustom(opts)`             | Custom host/port/protocol                        | `weaviate.connectToCustom({ httpHost: 'host', httpPort: 8080, grpcHost: 'host', grpcPort: 50051 })` |
    
    ### Connection Options
    
    | Option            | Default | Description                                                   |
    | ----------------- | ------- | ------------------------------------------------------------- |
    | `authCredentials` | none    | `new weaviate.ApiKey(key)` for API key auth                   |
    | `headers`         | `{}`    | API keys for vectorizer modules (`X-OpenAI-Api-Key`, etc.)    |
    | `timeout.query`   | 30      | Query timeout in seconds                                      |
    | `timeout.insert`  | 120     | Insert timeout in seconds                                     |
    | `timeout.init`    | 2       | Init check timeout in seconds                                 |
    | `skipInitChecks`  | false   | Skip version and port checks (temporary troubleshooting only) |
    
    ---
    
    ## Collection Operations
    
    | Operation    | Method                                                          | Notes                                              |
    | ------------ | --------------------------------------------------------------- | -------------------------------------------------- |
    | Create       | `client.collections.create({ name, vectorizers, properties })`  | Vectorizer must be set here                        |
    | Get          | `client.collections.use('Name')`                                | Returns collection object for queries              |
    | Exists       | `client.collections.exists('Name')`                             | Returns boolean                                    |
    | List         | `client.collections.listAll()`                                  | Returns all collection configs                     |
    | Config       | `collection.config.get()`                                       | Returns full collection configuration              |
    | Update       | `collection.config.update({ ... })`                             | Limited -- can update index params, not vectorizer |
    | Delete       | `client.collections.delete('Name')`                             | Permanent -- deletes all data                      |
    | Add property | `collection.config.addProperty({ name, dataType })`             | Existing objects not reindexed                     |
    | Add vector   | `collection.config.addVector(vectors.text2VecOpenAI({ name }))` | Named vectors only                                 |
    
    ---
    
    ## Data Operations
    
    | Operation      | Method                                        | Returns                             |
    | -------------- | --------------------------------------------- | ----------------------------------- |
    | Insert one     | `collection.data.insert({ properties })`      | UUID string                         |
    | Insert many    | `collection.data.insertMany(objects)`         | Response with `hasErrors`, `errors` |
    | Update (merge) | `collection.data.update({ id, properties })`  | Preserves unspecified properties    |
    | Replace (full) | `collection.data.replace({ id, properties })` | Deletes unspecified properties      |
    | Delete one     | `collection.data.deleteById(id)`              | boolean                             |
    | Delete many    | `collection.data.deleteMany(filter)`          | Count of deleted objects            |
    | Exists         | `collection.data.exists(id)`                  | boolean                             |
    | Fetch by ID    | `collection.query.fetchObjectById(id)`        | Object or null                      |
    
    ---
    
    ## Search Methods
    
    | Method                               | Description                           | Key Options                                            |
    | ------------------------------------ | ------------------------------------- | ------------------------------------------------------ |
    | `query.nearText(text, opts)`         | Semantic search via vectorizer module | `limit`, `distance`, `filters`, `returnMetadata`       |
    | `query.nearVector(vec, opts)`        | Search by raw vector                  | `limit`, `distance`, `filters`                         |
    | `query.hybrid(text, opts)`           | Vector + keyword blend                | `alpha` (0=keyword, 1=vector), `fusionType`, `filters` |
    | `query.bm25(text, opts)`             | Keyword search (BM25)                 | `queryProperties`, `filters`                           |
    | `query.fetchObjects(opts)`           | List/filter without search            | `limit`, `offset`, `filters`, `sort`                   |
    | `query.fetchObjectById(id)`          | Get single object                     | `includeVector`, `returnReferences`                    |
    | `generate.nearText(text, gen, opts)` | RAG with semantic search              | `singlePrompt`, `groupedTask`                          |
    | `generate.hybrid(text, gen, opts)`   | RAG with hybrid search                | Same as hybrid + generate options                      |
    | `generate.fetchObjects(gen, opts)`   | RAG without search ranking            | `singlePrompt`, `groupedTask`                          |
    
    ---
    
    ## Filter Operators
    
    | Operator                | Example                                       | Notes             |
    | ----------------------- | --------------------------------------------- | ----------------- |
    | `equal(value)`          | `.byProperty('status').equal('active')`       | Exact match       |
    | `notEqual(value)`       | `.byProperty('status').notEqual('draft')`     | Negation          |
    | `greaterThan(value)`    | `.byProperty('price').greaterThan(100)`       | Exclusive         |
    | `greaterOrEqual(value)` | `.byProperty('price').greaterOrEqual(100)`    | Inclusive         |
    | `lessThan(value)`       | `.byProperty('price').lessThan(50)`           | Exclusive         |
    | `lessOrEqual(value)`    | `.byProperty('price').lessOrEqual(50)`        | Inclusive         |
    | `like(pattern)`         | `.byProperty('name').like('*smith*')`         | Wildcard match    |
    | `containsAny(arr)`      | `.byProperty('tags').containsAny(['a', 'b'])` | Any token matches |
    | `containsAll(arr)`      | `.byProperty('tags').containsAll(['a', 'b'])` | All tokens match  |
    | `containsNone(arr)`     | `.byProperty('tags').containsNone(['x'])`     | No tokens match   |
    | `isNull(bool)`          | `.byProperty('field').isNull(true)`           | Null check        |
    | `withinGeoRange(opts)`  | `.byProperty('loc').withinGeoRange({...})`    | Geo proximity     |
    
    ### Combining Filters
    
    ```typescript
    import { Filters } from "weaviate-client";
    
    // AND
    Filters.and(filterA, filterB, filterC);
    
    // OR
    Filters.or(filterA, filterB);
    
    // NOT
    Filters.not(filterA);
    
    // Nested
    Filters.and(filterA, Filters.or(filterB, filterC));
    ```
    
    ### Metadata Filters
    
    ```typescript
    // By object ID
    collection.filter.byId().equal(targetId);
    
    // By creation time
    collection.filter.byCreationTime().greaterOrEqual("2024-01-01T00:00:00Z");
    
    // By property length (second arg = true)
    collection.filter.byProperty("title", true).greaterThan(10);
    
    // By cross-reference property
    collection.filter.byRef("hasCategory").byProperty("title").equal("Science");
    ```
    
    ---
    
    ## Vectorizer Comparison
    
    | Vectorizer             | Provider       | Use Case                | Requires API Key Header |
    | ---------------------- | -------------- | ----------------------- | ----------------------- |
    | `text2VecOpenAI`       | OpenAI         | General text embedding  | `X-OpenAI-Api-Key`      |
    | `text2VecCohere`       | Cohere         | Multilingual, general   | `X-Cohere-Api-Key`      |
    | `text2VecHuggingFace`  | HuggingFace    | Open-source models      | `X-HuggingFace-Api-Key` |
    | `text2VecOllama`       | Ollama (local) | Self-hosted models      | None (local)            |
    | `text2VecTransformers` | Custom         | Self-hosted transformer | None (local)            |
    | `multi2VecClip`        | CLIP           | Image + text multimodal | Depends on provider     |
    | `selfProvided`         | You            | Bring your own vectors  | None                    |
    
    ---
    
    ## Data Types
    
    | `dataType.*`      | TypeScript Type | Description                 |
    | ----------------- | --------------- | --------------------------- |
    | `TEXT`            | string          | Tokenized text (searchable) |
    | `TEXT_ARRAY`      | string[]        | Array of text values        |
    | `INT`             | number          | Integer                     |
    | `INT_ARRAY`       | number[]        | Array of integers           |
    | `NUMBER`          | number          | Float                       |
    | `NUMBER_ARRAY`    | number[]        | Array of floats             |
    | `BOOLEAN`         | boolean         | True/false                  |
    | `DATE`            | string/Date     | ISO 8601 date               |
    | `UUID`            | string          | UUID reference              |
    | `GEO_COORDINATES` | object          | `{ latitude, longitude }`   |
    | `BLOB`            | string          | Base64 encoded binary       |
    | `OBJECT`          | object          | Nested object               |
    | `OBJECT_ARRAY`    | object[]        | Array of nested objects     |
    
    ---
    
    ## Generative Model Configuration
    
    | Provider  | Import                   | Example                                               |
    | --------- | ------------------------ | ----------------------------------------------------- |
    | OpenAI    | `generative.openAI()`    | `generative.openAI({ model: "gpt-4o" })`              |
    | Cohere    | `generative.cohere()`    | `generative.cohere({ model: "command-r-plus" })`      |
    | Anthropic | `generative.anthropic()` | `generative.anthropic({ model: "claude-haiku-4-5" })` |
    | Ollama    | `generative.ollama()`    | `generative.ollama({ model: "llama3" })`              |
    
    ### Reranker Configuration
    
    | Provider | Import                | Example               |
    | -------- | --------------------- | --------------------- |
    | Cohere   | `reranker.cohere()`   | `reranker.cohere()`   |
    | VoyageAI | `reranker.voyageAI()` | `reranker.voyageAI()` |
    
    ---
    
    ## Vector Index Types
    
    | Type      | Use Case                                       | Config                            |
    | --------- | ---------------------------------------------- | --------------------------------- |
    | `hnsw`    | Default, good for most use cases               | `configure.vectorIndex.hnsw()`    |
    | `flat`    | Small collections (< 10K objects)              | `configure.vectorIndex.flat()`    |
    | `dynamic` | Auto-switches flat -> hnsw as collection grows | `configure.vectorIndex.dynamic()` |
    
    ### Quantization (Compression)
    
    ```typescript
    configure.vectorIndex.hnsw({
      quantizer: configure.vectorIndex.quantizer.pq(), // Product quantization
    });
    
    configure.vectorIndex.flat({
      quantizer: configure.vectorIndex.quantizer.bq(), // Binary quantization
    });
    ```
    
    ---
    
    ## Multi-Tenancy Quick Reference
    
    | Operation          | Method                                                                       |
    | ------------------ | ---------------------------------------------------------------------------- |
    | Enable             | `multiTenancy: weaviate.configure.multiTenancy({ enabled: true })` in create |
    | Auto-create        | `autoTenantCreation: true` in multiTenancy config                            |
    | Add tenants        | `collection.tenants.create([{ name: 'tenantA' }])`                           |
    | List tenants       | `collection.tenants.get()`                                                   |
    | Get by name        | `collection.tenants.getByName('tenantA')`                                    |
    | Delete tenants     | `collection.tenants.remove([{ name: 'tenantB' }])`                           |
    | Set state          | `collection.tenants.update({ name: 'tenantA', activityStatus: 'ACTIVE' })`   |
    | Query with tenant  | `collection.withTenant('tenantA').query.fetchObjects()`                      |
    | Insert with tenant | `collection.withTenant('tenantA').data.insert({ ... })`                      |
    
    ### Tenant States
    
    | State       | Description                                    |
    | ----------- | ---------------------------------------------- |
    | `ACTIVE`    | Tenant is loaded and queryable (default)       |
    | `INACTIVE`  | Tenant data on disk, not loaded in memory      |
    | `OFFLOADED` | Tenant data moved to cold storage (cloud only) |
    
    ---
    
    ## Production Checklist
    
    ### Connection
    
    - [ ] `client.close()` called in cleanup/shutdown handlers
    - [ ] API key headers for vectorizer modules (`X-OpenAI-Api-Key`, etc.)
    - [ ] Query timeout increased for RAG operations (`query: 60`)
    - [ ] `skipInitChecks: false` (only true for temporary debugging)
    
    ### Collections
    
    - [ ] Vectorizer configured at creation time
    - [ ] Properties defined with explicit data types (not auto-detected)
    - [ ] Generative model configured if using RAG
    - [ ] Vector index type appropriate for collection size
    
    ### Data
    
    - [ ] `insertMany` response checked for `hasErrors`
    - [ ] Deterministic UUIDs via `generateUuid5` for idempotent imports
    - [ ] `update` (merge) vs `replace` (overwrite) chosen correctly
    
    ### Multi-Tenancy
    
    - [ ] Tenant name validated (alphanumeric, underscore, hyphen; 4-64 chars)
    - [ ] `.withTenant()` used on all operations for multi-tenant collections
    - [ ] Inactive tenant activation before queries
    - [ ] Backups only include ACTIVE tenants
    
    ### Search
    
    - [ ] `targetVector` specified for named vector collections
    - [ ] `returnMetadata: ['distance']` or `['score']` for relevance debugging
    - [ ] Filters combined with `Filters.and()` / `Filters.or()` (not arrays)
    - [ ] `alpha` parameter documented for hybrid search tuning
    
    ---
    
    _Full skill documentation: [SKILL.md](SKILL.md) | Examples: [examples/](examples/)_
    
  • SKILL.md 15.3 KB
    ---
    name: api-vector-db-weaviate
    description: Weaviate vector database patterns with weaviate-client v3 -- collection management, vectorizer modules, hybrid search, filtering, generative search (RAG), multi-tenancy, batch imports
    ---
    
    # Weaviate Patterns
    
    > **Quick Guide:** Use Weaviate for semantic search and RAG applications. Use **weaviate-client** (v3.x) as the TypeScript client -- it uses gRPC for performance and provides full type safety with generics. Connect via `connectToWeaviateCloud()` for managed instances or `connectToLocal()` for Docker. Collections are the central abstraction -- configure vectorizers at collection level, not per-query. Use `collection.query.*` for search, `collection.generate.*` for RAG, and `collection.data.*` for CRUD. Always call `client.close()` when done. Increase query timeout to 60s+ when using generative search. The v3 client does NOT support browsers or Embedded Weaviate.
    
    ---
    
    <critical_requirements>
    
    ## CRITICAL: Before Using This Skill
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST call `client.close()` when done with the Weaviate client -- it maintains persistent gRPC connections that will leak if not closed)**
    
    **(You MUST configure vectorizers at the COLLECTION level during `client.collections.create()` -- you cannot add a vectorizer after creation, only add new named vectors)**
    
    **(You MUST use a SEPARATE `client.collections.use()` call with `.withTenant()` for multi-tenant queries -- queries without tenant context on multi-tenant collections will fail)**
    
    **(You MUST increase query timeout to 60+ seconds when using `generate.*` (RAG) submodule -- generative model calls are slow and the default timeout causes failures)**
    
    </critical_requirements>
    
    ---
    
    ## Examples
    
    - [Core Patterns](examples/core.md) -- Connection, collection setup, object CRUD, basic search
    - [Search & Filtering](examples/search.md) -- nearText, nearVector, hybrid, bm25, filters, generative search (RAG)
    - [Multi-Tenancy & Batch](examples/multi-tenancy.md) -- Tenant management, batch imports, cross-references
    
    **Additional resources:**
    
    - [reference.md](reference.md) -- API cheat sheet, vectorizer comparison, data types, decision frameworks
    
    ---
    
    **Auto-detection:** Weaviate, weaviate-client, connectToWeaviateCloud, connectToLocal, nearText, nearVector, hybrid search, bm25, vector database, semantic search, RAG, generative search, generate.nearText, insertMany, vectorizer, text2vec, multi-tenancy, withTenant, collection.query, collection.generate, collection.data
    
    **When to use:**
    
    - Semantic search over text, images, or multimodal data
    - Retrieval Augmented Generation (RAG) with built-in generative search
    - Hybrid search combining vector similarity and keyword (BM25) ranking
    - Multi-tenant applications needing isolated vector stores per customer
    - Applications requiring built-in vectorization (no external embedding pipeline)
    - Real-time similarity search with filtering on structured properties
    
    **Key patterns covered:**
    
    - weaviate-client v3 connection setup and configuration
    - Collection management with vectorizer modules (text2vec-openai, text2vec-cohere, etc.)
    - Object CRUD (insert, insertMany, update, replace, deleteById, deleteMany)
    - Search types (nearText, nearVector, hybrid, bm25, fetchObjects)
    - Filtering with operators (equal, greaterThan, like, containsAny, and/or/not)
    - Generative search (RAG) with singlePrompt and groupedTask
    - Multi-tenancy with tenant lifecycle management
    - Batch imports with insertMany and error handling
    - Cross-references between collections
    - Named vectors for multi-vector collections
    
    **When NOT to use:**
    
    - Relational data with complex joins (use a relational database)
    - Full-text search without vector component (use a dedicated search engine)
    - Key-value caching (use a key-value store)
    - Time-series data (use a time-series database)
    - Graph traversal queries (use a graph database)
    - Browser-side applications (v3 client is Node.js only)
    
    ---
    
    <philosophy>
    
    ## Philosophy
    
    Weaviate is a **vector database** that stores data objects alongside their vector embeddings. The core principle: **configure once at the collection level, then query with simple method calls.**
    
    **Core principles:**
    
    1. **Collection-centric design** -- All configuration (vectorizer, generative model, reranker, properties) is set at collection creation. Queries operate on collection objects obtained via `client.collections.use()`.
    2. **Built-in vectorization** -- Weaviate can vectorize data automatically using configured modules (text2vec-openai, text2vec-cohere, etc.). You don't need an external embedding pipeline unless you want one.
    3. **Search is a spectrum** -- Use `nearText` for semantic similarity, `bm25` for keyword matching, `hybrid` for a weighted combination. The `alpha` parameter controls the vector-vs-keyword balance in hybrid search.
    4. **RAG is a search mode, not a separate system** -- Switch from `collection.query.nearText()` to `collection.generate.nearText()` to add LLM generation on top of search results.
    5. **Filters are additive** -- Filters narrow results after vector/keyword retrieval. Combine with `Filters.and()` and `Filters.or()` for complex conditions.
    
    </philosophy>
    
    ---
    
    <patterns>
    
    ## Core Patterns
    
    ### Pattern 1: Connection Setup
    
    Connect to Weaviate Cloud or local Docker instance. Always close the client when done. See [examples/core.md](examples/core.md) for full examples.
    
    ```typescript
    // Good Example -- Cloud connection with API key headers
    import weaviate from "weaviate-client";
    
    const QUERY_TIMEOUT_SECONDS = 30;
    const INSERT_TIMEOUT_SECONDS = 120;
    
    async function createWeaviateClient() {
      const client = await weaviate.connectToWeaviateCloud(
        process.env.WEAVIATE_URL!,
        {
          authCredentials: new weaviate.ApiKey(process.env.WEAVIATE_API_KEY!),
          headers: {
            "X-OpenAI-Api-Key": process.env.OPENAI_API_KEY!,
          },
          timeout: {
            query: QUERY_TIMEOUT_SECONDS,
            insert: INSERT_TIMEOUT_SECONDS,
          },
        },
      );
      return client;
    }
    
    export { createWeaviateClient };
    ```
    
    **Why good:** Environment variables for credentials, explicit timeouts, API key headers for vectorizer modules
    
    ```typescript
    // Bad Example -- Missing cleanup, no timeout config
    import weaviate from "weaviate-client";
    const client = await weaviate.connectToLocal();
    // No client.close() -- gRPC connections leak
    // No timeout config -- generative queries will timeout
    ```
    
    **Why bad:** Missing `client.close()` leaks gRPC connections, default timeout too short for RAG queries
    
    ---
    
    ### Pattern 2: Collection with Vectorizer
    
    Configure vectorizer and properties at creation time. See [examples/core.md](examples/core.md) for named vectors and advanced configuration.
    
    ```typescript
    // Good Example -- Collection with vectorizer and generative model
    import { vectors, dataType, generative } from "weaviate-client";
    
    await client.collections.create({
      name: "Article",
      vectorizers: vectors.text2VecOpenAI({
        model: "text-embedding-3-small",
      }),
      generative: generative.openAI({
        model: "gpt-4o",
      }),
      properties: [
        { name: "title", dataType: dataType.TEXT },
        { name: "body", dataType: dataType.TEXT },
        { name: "category", dataType: dataType.TEXT },
        { name: "publishedAt", dataType: dataType.DATE },
      ],
    });
    ```
    
    **Why good:** Vectorizer and generative model configured at collection level, typed properties with explicit data types
    
    ```typescript
    // Bad Example -- Trying to add vectorizer after creation
    await client.collections.create({ name: "Article" });
    // No way to add a vectorizer to an existing collection without named vectors
    // Must delete and recreate, or use addVector() for named vectors only
    ```
    
    **Why bad:** Vectorizer must be set at creation time; cannot be added to an existing default vector after the fact
    
    ---
    
    ### Pattern 3: Hybrid Search with Filters
    
    Combine vector and keyword search with property filters. See [examples/search.md](examples/search.md) for all search types.
    
    ```typescript
    // Good Example -- Hybrid search with filter
    import { Filters } from "weaviate-client";
    
    const articles = client.collections.use("Article");
    const SEARCH_LIMIT = 10;
    const HYBRID_ALPHA = 0.75; // Favor vector search
    
    const result = await articles.query.hybrid("machine learning trends", {
      alpha: HYBRID_ALPHA,
      limit: SEARCH_LIMIT,
      filters: Filters.and(
        articles.filter.byProperty("category").equal("technology"),
        articles.filter
          .byProperty("publishedAt")
          .greaterThan(new Date("2024-01-01")),
      ),
      returnMetadata: ["score", "explainScore"],
    });
    
    for (const obj of result.objects) {
      console.log(obj.properties.title, obj.metadata?.score);
    }
    ```
    
    **Why good:** Named constants for limits and alpha, combined filter with `Filters.and()`, metadata for debugging relevance
    
    ---
    
    ### Pattern 4: Generative Search (RAG)
    
    Switch from `query.*` to `generate.*` for RAG. See [examples/search.md](examples/search.md) for singlePrompt and groupedTask patterns.
    
    ```typescript
    // Good Example -- RAG with single prompt per result
    const articles = client.collections.use("Article");
    const RAG_RESULT_LIMIT = 5;
    
    const result = await articles.generate.nearText(
      "climate change policy",
      {
        singlePrompt: "Summarize this article in one sentence: {title} - {body}",
      },
      {
        limit: RAG_RESULT_LIMIT,
        returnMetadata: ["distance"],
      },
    );
    
    for (const obj of result.objects) {
      console.log("Source:", obj.properties.title);
      console.log("Generated:", obj.generative?.text);
    }
    ```
    
    **Why good:** Uses property interpolation `{title}` in prompt, accesses generated text via `obj.generative?.text`
    
    ```typescript
    // Bad Example -- Using query instead of generate for RAG
    const result = await articles.query.nearText("climate change", { limit: 5 });
    // Then manually calling OpenAI API with results
    // Weaviate does this natively with generate.*
    ```
    
    **Why bad:** Misses Weaviate's built-in RAG -- extra network hops, no automatic prompt interpolation
    
    </patterns>
    
    ---
    
    <decision_framework>
    
    ## Decision Framework
    
    ### Which Search Type?
    
    ```
    What kind of search do I need?
    ├─ Natural language query, semantic meaning? -> nearText (requires vectorizer module)
    ├─ Have pre-computed vector embedding? -> nearVector
    ├─ Exact keyword matching? -> bm25
    ├─ Both semantic and keyword relevance? -> hybrid (alpha controls blend)
    ├─ Just list/filter objects without search? -> fetchObjects
    └─ Search + LLM generation? -> generate.nearText / generate.hybrid
    ```
    
    ### Which Vectorizer?
    
    ```
    Which vectorizer module should I use?
    ├─ OpenAI models (text-embedding-3-small/large)? -> text2VecOpenAI
    ├─ Cohere models (embed-v3)? -> text2VecCohere
    ├─ Self-hosted models? -> text2VecOllama or text2VecTransformers
    ├─ Bring your own embeddings? -> none (use selfProvided for named vectors)
    ├─ Multimodal (images + text)? -> multi2VecClip or multi2VecBind
    └─ Multiple embedding strategies? -> Named vectors (array of vectorizers)
    ```
    
    ### Single vs Named Vectors?
    
    ```
    How many vector representations do I need?
    ├─ One embedding per object (most common)? -> Single default vectorizer
    ├─ Different embeddings for different properties? -> Named vectors
    ├─ Mix of auto-vectorized and self-provided? -> Named vectors with selfProvided
    └─ Different models for different search use cases? -> Named vectors
    ```
    
    ### When to Use Multi-Tenancy?
    
    ```
    Do I need data isolation?
    ├─ Each customer/user needs isolated data? -> Enable multi-tenancy
    ├─ Shared dataset, filter by user? -> Single tenant with filters
    ├─ Need to offload inactive tenants? -> Multi-tenancy with tenant states
    └─ Small number of distinct datasets? -> Separate collections may be simpler
    ```
    
    </decision_framework>
    
    ---
    
    <red_flags>
    
    ## RED FLAGS
    
    **High Priority Issues:**
    
    - Missing `client.close()` -- gRPC connections persist and leak memory/file descriptors
    - Trying to add a default vectorizer after collection creation -- vectorizer must be configured in `create()`. Only named vectors can be added later with `config.addVector()`
    - Querying a multi-tenant collection without `.withTenant()` -- all operations fail with an error
    - Using default query timeout with `generate.*` -- generative calls need 60+ seconds; default is often too short
    
    **Medium Priority Issues:**
    
    - Using `replace()` when `update()` is intended -- `replace` deletes all properties not included in the call; `update` merges
    - Not checking `insertMany` response for errors -- partial failures are silent; check `response.hasErrors` and `response.errors`
    - Passing `alpha: 1.0` to hybrid search -- equivalent to pure vector search; use `nearText` instead for clarity
    - Not specifying `targetVector` with named vectors -- queries default to the first vector, which may not be the intended one
    
    **Common Mistakes:**
    
    - Using v2 class-based API (`client.schema.classCreator()`) with v3 client -- the API is completely different; v3 uses `client.collections.create()`
    - Forgetting to pass API key headers for vectorizer modules -- `X-OpenAI-Api-Key`, `X-Cohere-Api-Key` etc. must be in connection headers
    - Using `connectToWCS()` (deprecated) instead of `connectToWeaviateCloud()`
    - Adding a property after data import without reindexing -- pre-existing objects won't have that property indexed
    
    **Gotchas & Edge Cases:**
    
    - `insertMany` uses server-side batching but the TS client does NOT have a streaming batch API -- for very large imports (100K+), chunk into batches of 100-1000 objects
    - `Filters.and()` and `Filters.or()` take a flat list of filter conditions, NOT nested arrays -- `Filters.and(a, b, c)` not `Filters.and([a, b, c])`
    - `fetchObjects()` without `limit` returns 25 objects by default (server-side default), not all objects
    - Property names in Weaviate must start with a lowercase letter -- the client silently lowercases the first character
    - `distance` metadata varies by vector distance metric -- cosine distance range [0, 2], not [0, 1]
    - `deleteMany` has a server-side maximum of 10,000 objects per call (configurable via `QUERY_MAXIMUM_RESULTS`)
    - Weaviate auto-detects property types on first insert if not defined in the schema -- this can cause type mismatches if first object has atypical data
    - `fetchObjectById` returns `null` for non-existent IDs, not an empty object -- always check for null before accessing properties
    - Cross-references in multi-tenant collections can only reference objects in the same tenant or in non-multi-tenant collections
    
    </red_flags>
    
    ---
    
    <critical_reminders>
    
    ## CRITICAL REMINDERS
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST call `client.close()` when done with the Weaviate client -- it maintains persistent gRPC connections that will leak if not closed)**
    
    **(You MUST configure vectorizers at the COLLECTION level during `client.collections.create()` -- you cannot add a vectorizer after creation, only add new named vectors)**
    
    **(You MUST use a SEPARATE `client.collections.use()` call with `.withTenant()` for multi-tenant queries -- queries without tenant context on multi-tenant collections will fail)**
    
    **(You MUST increase query timeout to 60+ seconds when using `generate.*` (RAG) submodule -- generative model calls are slow and the default timeout causes failures)**
    
    **Failure to follow these rules will cause connection leaks, missing vectorization, multi-tenant query failures, and RAG timeouts.**
    
    </critical_reminders>
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related