api-vector-db-qdrant
Qdrant vector database -- collection management, point operations, payload filtering, named vectors, quantization, recommendations, snapshots
Install
npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/api-vector-db-qdrant/skills/api-vector-db-qdrant
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
git clone https://github.com/agents-inc/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Qdrant Patterns
Quick Guide: Use
@qdrant/js-client-rest(v1.17.x) for high-performance vector search. Collections define vector dimensions and distance metrics upfront -- mismatches cause silent failures. Usemust/should/must_notfilter clauses with payload conditions (not Pinecone-style$eq/$and). Payload indexes are optional but critical for filter performance at scale -- create them explicitly withcreatePayloadIndex(). Named vectors let you store multiple embeddings per point (e.g., title + content). Quantization (scalar/binary/product) trades accuracy for memory and speed. Thequery()method is the universal search endpoint -- prefer it over the oldersearch()method.
<critical_requirements>
CRITICAL: Before Using This Skill
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST create payload indexes with createPayloadIndex() for any field used in filters -- unindexed fields cause full scans that degrade linearly with collection size)
(You MUST use must/should/must_not filter syntax -- Qdrant does NOT use $eq/$and/$or operators like Pinecone)
(You MUST match vector dimensions exactly between embedding model output and collection config -- dimension mismatches cause silent upsert failures or corrupt search results)
(You MUST set wait: true on writes when subsequent reads depend on the data -- Qdrant writes are asynchronous by default and may not be immediately visible)
</critical_requirements>
Examples
- Core Patterns -- Client setup, collection creation, upsert, query, scroll, delete
- Filtering -- must/should/must_not conditions, match/range operators, payload indexes
- Named Vectors & Quantization -- Multiple vectors per point, scalar/binary/product quantization
- Recommendations & Batch -- Recommend API, batch operations, snapshots
Additional resources:
- reference.md -- API quick reference, filter operators, limits, decision frameworks, production checklist
Auto-detection: Qdrant, QdrantClient, @qdrant/js-client-rest, createCollection, upsert, query, scroll, recommend, setPayload, createPayloadIndex, must, should, must_not, payload, named vectors, quantization, vector database, similarity search, semantic search, RAG retrieval, embedding search
When to use:
- Semantic search over document embeddings (RAG retrieval pipelines)
- Similarity search for recommendations, deduplication, or classification
- Multi-vector search with named vectors (e.g., title embedding + content embedding per document)
- Filtered vector search with complex payload conditions (must/should/must_not)
- Memory-optimized deployments using scalar, binary, or product quantization
Key patterns covered:
- Client setup and collection management (distance metrics, HNSW config)
- Point CRUD operations (upsert, query, scroll, retrieve, delete, count)
- Payload filtering with must/should/must_not and match/range conditions
- Named vectors for multiple embeddings per point
- Quantization configuration (scalar, binary, product)
- Recommendation API with positive/negative examples
- Batch operations and snapshot management
- Payload indexing for filter performance
When NOT to use:
- Full-text search with BM25 ranking (use a dedicated search engine)
- Relational data with joins and transactions (use a relational database)
- Key-value lookups without vector similarity (use a KV store)
- Storing large documents or binary blobs (store embeddings + metadata references only)
<decision_framework>
Decision Framework
Which Distance Metric?
Which distance metric should I use?
|-- Using normalized embeddings (OpenAI, Cohere)? -> Cosine (most common, safe default)
|-- Pre-normalized embeddings and need speed? -> Dot (faster, same results as Cosine for unit vectors)
|-- Raw feature vectors where magnitude matters? -> Euclid (L2 distance)
|-- City-block distance needed? -> Manhattan
'-- Unsure? -> Cosine (works with any embedding model)
Single Vector vs Named Vectors?
How many embeddings per point?
|-- One embedding model? -> Single vector (simpler config)
|-- Multiple embedding models (title + content)? -> Named vectors
|-- Same model, different text segments? -> Named vectors
|-- Multi-modal (text + image)? -> Named vectors with different dimensions
'-- Want to avoid duplicating payloads across collections? -> Named vectors
Which Quantization Method?
How should I optimize memory?
|-- Good default, balanced accuracy/speed? -> Scalar (int8, 4x compression)
|-- Maximum speed, can tolerate accuracy loss? -> Binary (32x compression)
| '-- Best with high-dimensional models (>= 1024 dims)
|-- Maximum compression, speed not critical? -> Product (up to 64x compression)
| '-- Slowest quantization, most accuracy loss
'-- No memory pressure? -> Skip quantization (full float32 precision)
Payload Index Strategy?
Should I create a payload index?
|-- Field used in filter conditions? -> YES, always index
|-- Field used in order_by for scroll? -> YES, index for sort performance
|-- Field only read after search (display only)? -> NO, skip index
|-- High-cardinality field (UUIDs, timestamps)? -> YES, but evaluate index type
'-- Low-cardinality field (enum-like)? -> YES, keyword index is very efficient
</decision_framework>
<red_flags>
RED FLAGS
High Priority Issues:
- Using Pinecone-style filter syntax (
$eq,$and,$or) -- Qdrant usesmust/should/must_notwithmatch/rangeconditions - Vector dimension mismatch between embedding model and collection config -- causes silent failures or garbage results
- Missing payload indexes on filtered fields -- causes full collection scans that degrade linearly with size
- Not setting
wait: truewhen read-after-write consistency is needed -- writes are async by default
Medium Priority Issues:
- Using deprecated
search()method instead ofquery()--query()is the universal endpoint with prefetch and fusion support - Forgetting
with_payload: truein queries -- payload is NOT included by default - Creating payload indexes after bulk upsert instead of before -- retroactive indexing is slower than indexing during upsert
- Using
offsetfor deep pagination in scroll -- performance degrades; useoffsetas cursor (point ID), not page number
Common Mistakes:
- Passing
filterat the wrong nesting level -- filter goes at the top level of the query args, not nested inside another object - Using
id: 0as a point ID -- Qdrant requires positive integers or UUID strings; 0 is invalid - Confusing
setPayload(merge) withoverwritePayload(replace) --setPayloadmerges fields,overwritePayloadreplaces the entire payload - Calling
deletePayloadwith field names but no point selector -- you must specify which points to update viapointsarray orfilter
Gotchas & Edge Cases:
- Point IDs must be positive integers or UUID strings -- negative numbers, zero, and non-UUID strings are rejected
scroll()withorder_byrequires a payload index on the sort field -- without it, the request failscount()withexact: trueis slow on large collections -- useexact: false(default) for approximate counts- Snapshot recovery requires matching Qdrant minor versions -- a v1.14.x snapshot cannot be restored to a v1.15.x cluster
- Binary quantization works best with high-dimensional vectors (>= 1024 dims) -- for smaller vectors, scalar quantization is more accurate
query()withprefetchenables multi-stage retrieval (retrieve 1000, then re-rank to top 10) -- but requires understanding the prefetch pipeline- Named vector search requires the
usingparameter -- omitting it searches the default (unnamed) vector, which may not exist deletePayloadremoves specific keys,clearPayloadremoves ALL keys -- they are different operations
</red_flags>
<critical_reminders>
CRITICAL REMINDERS
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST create payload indexes with createPayloadIndex() for any field used in filters -- unindexed fields cause full scans that degrade linearly with collection size)
(You MUST use must/should/must_not filter syntax -- Qdrant does NOT use $eq/$and/$or operators like Pinecone)
(You MUST match vector dimensions exactly between embedding model output and collection config -- dimension mismatches cause silent upsert failures or corrupt search results)
(You MUST set wait: true on writes when subsequent reads depend on the data -- Qdrant writes are asynchronous by default and may not be immediately visible)
Failure to follow these rules will cause empty search results, degraded filter performance, data consistency issues, and hard-to-debug dimension mismatch errors.
</critical_reminders>
Files (skills)
-
examples
-
core.md 10.4 KB
# Qdrant -- Core Pattern Examples > Client setup, collection management, and fundamental point operations (upsert, query, scroll, retrieve, delete, count). Reference from [SKILL.md](../SKILL.md). **Related examples:** - [filtering.md](filtering.md) -- must/should/must_not conditions, match/range operators - [named-vectors-quantization.md](named-vectors-quantization.md) -- Multiple vectors per point, quantization - [recommendations-batch.md](recommendations-batch.md) -- Recommend API, batch operations, snapshots --- ## Client Initialization ```typescript import { QdrantClient } from "@qdrant/js-client-rest"; function createQdrantClient(): QdrantClient { const url = process.env.QDRANT_URL; const apiKey = process.env.QDRANT_API_KEY; if (!url) { throw new Error("QDRANT_URL environment variable is required"); } // apiKey is optional for local instances return new QdrantClient({ url, apiKey }); } export { createQdrantClient }; ``` **Why good:** URL from environment variable (never hardcoded), optional API key for local vs cloud, explicit validation, named export For local development without authentication: ```typescript const client = new QdrantClient({ url: "http://localhost:6333" }); ``` For Qdrant Cloud: ```typescript const client = new QdrantClient({ url: "https://your-cluster.cloud.qdrant.io", apiKey: process.env.QDRANT_API_KEY, }); ``` **Note:** You can also use `host` + `port` instead of `url`, but `url` is preferred because it includes the protocol and avoids ambiguity with HTTPS. --- ## Create a Collection ```typescript import { QdrantClient } from "@qdrant/js-client-rest"; const EMBEDDING_DIMENSION = 1536; // Must match your embedding model const COLLECTION_NAME = "documents"; async function createDocumentsCollection(client: QdrantClient): Promise<void> { const exists = await client.collectionExists(COLLECTION_NAME); if (exists.exists) { return; // Collection already exists } await client.createCollection(COLLECTION_NAME, { vectors: { size: EMBEDDING_DIMENSION, distance: "Cosine", }, }); } export { createDocumentsCollection, COLLECTION_NAME, EMBEDDING_DIMENSION }; ``` **Why good:** Named constants for dimension and collection name, idempotent (checks existence first), explicit distance metric ```typescript // Bad Example -- missing existence check, hardcoded values await client.createCollection("docs", { vectors: { size: 768, distance: "Cosine" }, }); // Throws if collection already exists; dimension may not match model ``` **Why bad:** No existence check causes errors on re-run, hardcoded dimension risks mismatch --- ## Create Collection with HNSW Tuning ```typescript const EMBEDDING_DIMENSION = 1536; const HNSW_M = 16; // Connections per node (higher = better recall, more memory) const HNSW_EF_CONSTRUCT = 200; // Build-time search width (higher = better index, slower build) await client.createCollection("high-recall-docs", { vectors: { size: EMBEDDING_DIMENSION, distance: "Cosine", }, hnsw_config: { m: HNSW_M, ef_construct: HNSW_EF_CONSTRUCT, }, }); ``` **Why good:** Named constants for HNSW parameters, explicit tuning for recall-critical workloads **Gotcha:** Higher `m` and `ef_construct` values improve recall but increase memory usage and index build time. The defaults (`m: 16`, `ef_construct: 100`) work well for most use cases. --- ## Upsert Points with Typed Payload ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; interface ArticlePayload { title: string; category: string; publishedAt: number; // Unix timestamp tags: string[]; } async function upsertArticle( client: QdrantClient, id: string, embedding: number[], payload: ArticlePayload, ): Promise<void> { await client.upsert("articles", { wait: true, points: [{ id, vector: embedding, payload }], }); } export { upsertArticle }; export type { ArticlePayload }; ``` **Why good:** Type-safe payload interface, `wait: true` for read-after-write consistency, string ID (UUID-format), named export ```typescript // Bad Example -- no wait, numeric ID issues await client.upsert("articles", { points: [ { id: 0, // INVALID: point IDs must be positive integers or UUID strings vector: embedding, payload: { title: "Guide" }, }, ], }); // No wait: true -- data may not be queryable immediately ``` **Why bad:** `id: 0` is invalid (must be positive integer or UUID string), missing `wait: true` means subsequent queries may miss this point --- ## Query by Vector Similarity ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; const TOP_K = 10; const COLLECTION_NAME = "articles"; interface SearchResult { id: string | number; score: number; payload: Record<string, unknown> | null; } async function searchArticles( client: QdrantClient, queryEmbedding: number[], category?: string, ): Promise<SearchResult[]> { const filter = category ? { must: [{ key: "category", match: { value: category } }] } : undefined; const response = await client.query(COLLECTION_NAME, { query: queryEmbedding, filter, with_payload: true, limit: TOP_K, }); return response.points.map((point) => ({ id: point.id, score: point.score ?? 0, payload: point.payload ?? null, })); } export { searchArticles }; ``` **Why good:** Uses `query()` (universal endpoint), optional filter, `with_payload: true`, named constants, typed response --- ## Retrieve Points by ID ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; async function getArticles( client: QdrantClient, ids: (string | number)[], ): Promise<void> { const records = await client.retrieve("articles", { ids, with_payload: true, with_vector: false, // Skip vectors to reduce response size }); for (const record of records) { console.log(record.id, record.payload); } } export { getArticles }; ``` **Why good:** Retrieves by ID without vector search, `with_vector: false` reduces response size, explicit payload inclusion **Gotcha:** `retrieve()` returns an array, not an object keyed by ID. Missing IDs are silently omitted (no error thrown) -- always check the returned array length if you need to confirm all IDs were found. --- ## Scroll Through All Points ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; const PAGE_SIZE = 100; async function scrollAllPoints( client: QdrantClient, collectionName: string, ): Promise< Array<{ id: string | number; payload: Record<string, unknown> | null }> > { const allPoints: Array<{ id: string | number; payload: Record<string, unknown> | null; }> = []; let offset: string | number | undefined; do { const response = await client.scroll(collectionName, { limit: PAGE_SIZE, offset, with_payload: true, with_vector: false, }); for (const point of response.points) { allPoints.push({ id: point.id, payload: point.payload ?? null }); } offset = response.next_page_offset ?? undefined; } while (offset !== undefined); return allPoints; } export { scrollAllPoints }; ``` **Why good:** Cursor-based pagination using `next_page_offset`, skips vectors for efficiency, named constant for page size **Gotcha:** `scroll()` returns `next_page_offset` as the cursor for the next page. This is NOT a numeric page offset -- it is a point ID. Pass it directly as `offset` for the next request. --- ## Delete Points ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; // Delete by IDs async function deleteByIds( client: QdrantClient, collectionName: string, ids: (string | number)[], ): Promise<void> { await client.delete(collectionName, { wait: true, points: ids, }); } // Delete by filter async function deleteByCategory( client: QdrantClient, collectionName: string, category: string, ): Promise<void> { await client.delete(collectionName, { wait: true, filter: { must: [{ key: "category", match: { value: category } }], }, }); } export { deleteByIds, deleteByCategory }; ``` **Why good:** Two deletion patterns (by ID, by filter), `wait: true` for consistency, Qdrant filter syntax --- ## Count Points ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; async function countByCategory( client: QdrantClient, collectionName: string, category: string, ): Promise<number> { const result = await client.count(collectionName, { filter: { must: [{ key: "category", match: { value: category } }], }, exact: true, // Exact count (slower) vs approximate (faster, default) }); return result.count; } export { countByCategory }; ``` **Why good:** Filtered count, explicit `exact: true` when precision matters **Gotcha:** `exact: true` performs a full scan and is slow on large collections. Use `exact: false` (default) for approximate counts in dashboards or monitoring. --- ## Payload Operations (Set, Overwrite, Delete, Clear) ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; // Merge new fields into existing payload (preserves other fields) async function addPayloadFields( client: QdrantClient, collectionName: string, pointIds: (string | number)[], fields: Record<string, unknown>, ): Promise<void> { await client.setPayload(collectionName, { payload: fields, points: pointIds, wait: true, }); } // Replace entire payload (removes all existing fields) async function replacePayload( client: QdrantClient, collectionName: string, pointIds: (string | number)[], payload: Record<string, unknown>, ): Promise<void> { await client.overwritePayload(collectionName, { payload, points: pointIds, wait: true, }); } // Remove specific payload keys async function removePayloadKeys( client: QdrantClient, collectionName: string, pointIds: (string | number)[], keys: string[], ): Promise<void> { await client.deletePayload(collectionName, { keys, points: pointIds, wait: true, }); } export { addPayloadFields, replacePayload, removePayloadKeys }; ``` **Why good:** Three distinct operations for payload management, `wait: true` for consistency **Critical distinction:** `setPayload` MERGES fields (like `Object.assign`), while `overwritePayload` REPLACES the entire payload. Use `setPayload` to add a field without losing existing data. --- _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_ -
filtering.md 8.7 KB
# Qdrant -- Filtering Examples > Payload filtering with must/should/must_not conditions, match/range operators, and payload indexing. See [core.md](core.md) for basic query patterns and [reference.md](../reference.md) for the operator table. **Related examples:** - [core.md](core.md) -- Client setup, upsert, query, scroll, delete - [named-vectors-quantization.md](named-vectors-quantization.md) -- Named vectors, quantization config --- ## Create Payload Indexes **Always create indexes on fields used in filter conditions.** Without indexes, filters cause full collection scans. ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; async function setupPayloadIndexes( client: QdrantClient, collectionName: string, ): Promise<void> { // Keyword index for exact match / enum-like fields await client.createPayloadIndex(collectionName, { field_name: "category", field_schema: "keyword", wait: true, }); // Integer index for range queries await client.createPayloadIndex(collectionName, { field_name: "publishedAt", field_schema: "integer", wait: true, }); // Text index for full-text search await client.createPayloadIndex(collectionName, { field_name: "description", field_schema: "text", wait: true, }); // Float index for numeric range await client.createPayloadIndex(collectionName, { field_name: "score", field_schema: "float", wait: true, }); } export { setupPayloadIndexes }; ``` **Why good:** Indexes created per field type, `wait: true` ensures index is ready before queries, covers common field types **Gotcha:** Create indexes BEFORE bulk data ingestion when possible. Retroactive indexing after millions of points is slower than indexing during upsert. --- ## Must (AND) Conditions All conditions must match. Equivalent to logical AND. ```typescript const TOP_K = 10; const results = await client.query("documents", { query: embedding, filter: { must: [ { key: "category", match: { value: "tutorial" } }, { key: "publishedAt", range: { gte: 1700000000 } }, ], }, with_payload: true, limit: TOP_K, }); ``` **Why good:** Both conditions enforced, explicit match and range operators with correct nesting --- ## Should (OR) Conditions At least one condition must match. Equivalent to logical OR. ```typescript const TOP_K = 10; const results = await client.query("documents", { query: embedding, filter: { should: [ { key: "category", match: { value: "tutorial" } }, { key: "category", match: { value: "guide" } }, ], }, with_payload: true, limit: TOP_K, }); ``` **Why good:** OR logic via `should`, each condition independent --- ## Must Not (NOT) Conditions No listed conditions may match. Equivalent to logical NOT. ```typescript const TOP_K = 10; const results = await client.query("documents", { query: embedding, filter: { must_not: [ { key: "category", match: { value: "archived" } }, { key: "status", match: { value: "draft" } }, ], }, with_payload: true, limit: TOP_K, }); ``` **Why good:** Excludes both archived and draft documents, clean negation syntax --- ## Combined must + should + must_not Mix filter logic at the top level. ```typescript const TOP_K = 10; const results = await client.query("documents", { query: embedding, filter: { must: [{ key: "publishedAt", range: { gte: 1700000000 } }], should: [ { key: "category", match: { value: "tutorial" } }, { key: "category", match: { value: "guide" } }, ], must_not: [{ key: "status", match: { value: "draft" } }], }, with_payload: true, limit: TOP_K, }); ``` **Why good:** AND + OR + NOT in a single filter, readable structure **Semantics:** `must AND (at least one should) AND (none of must_not)` -- all three clauses are combined with AND at the top level. --- ## Match Operators ```typescript const TOP_K = 10; // Exact value match const exact = await client.query("products", { query: embedding, filter: { must: [{ key: "brand", match: { value: "Acme" } }] }, limit: TOP_K, }); // Match any value from a list (like SQL IN) const anyOf = await client.query("products", { query: embedding, filter: { must: [{ key: "brand", match: { any: ["Acme", "Globex", "Initech"] } }], }, limit: TOP_K, }); // Exclude values from a list (like SQL NOT IN) const exceptThese = await client.query("products", { query: embedding, filter: { must: [{ key: "brand", match: { except: ["Acme", "Globex"] } }] }, limit: TOP_K, }); // Full-text substring match (requires text index) const textMatch = await client.query("products", { query: embedding, filter: { must: [{ key: "description", match: { text: "wireless" } }] }, limit: TOP_K, }); ``` **Why good:** Four match variants for different use cases, comments clarify SQL equivalents --- ## Range Operators ```typescript const TOP_K = 10; // Numeric range (price between 10 and 100) const priceRange = await client.query("products", { query: embedding, filter: { must: [{ key: "price", range: { gte: 10, lte: 100 } }], }, limit: TOP_K, }); // Open-ended range (published after a date) const recent = await client.query("documents", { query: embedding, filter: { must: [{ key: "publishedAt", range: { gt: 1700000000 } }], }, limit: TOP_K, }); ``` **Why good:** Range supports `gt`, `gte`, `lt`, `lte` -- any subset can be used, both bounded and open-ended ranges shown --- ## Existence and Null Checks ```typescript const TOP_K = 10; // Points where "summary" field has no value const empty = await client.query("documents", { query: embedding, filter: { must: [{ is_empty: { key: "summary" } }], }, limit: TOP_K, }); // Points where "deleted_at" is explicitly null const nullField = await client.query("documents", { query: embedding, filter: { must: [{ is_null: { key: "deleted_at" } }], }, limit: TOP_K, }); // Points by specific IDs const byId = await client.query("documents", { query: embedding, filter: { must: [{ has_id: [1, 42, 100] }], }, limit: TOP_K, }); ``` **Why good:** `is_empty` checks for missing/empty fields, `is_null` checks for explicit nulls, `has_id` filters by point IDs **Gotcha:** `is_empty` matches when the field does not exist OR has an empty value. `is_null` matches only when the field exists and is explicitly `null`. --- ## Nested Payload Filtering Filter on fields inside nested JSON objects. ```typescript const TOP_K = 10; // Payload structure: { items: [{ name: "Widget", price: 29.99 }] } const results = await client.query("orders", { query: embedding, filter: { must: [ { nested: { key: "items", filter: { must: [ { key: "name", match: { value: "Widget" } }, { key: "price", range: { lte: 50 } }, ], }, }, }, ], }, limit: TOP_K, }); ``` **Why good:** Nested filter reaches into array-of-objects payload structure, conditions apply within each nested object **Gotcha:** Nested filtering requires the parent field to be an array of objects. Conditions within a `nested` block apply to each object independently -- a point matches if ANY object in the array satisfies ALL conditions. --- ## Values Count Filter Filter by the number of values a multi-value payload field has. ```typescript const TOP_K = 10; // Points with at least 3 tags const wellTagged = await client.query("documents", { query: embedding, filter: { must: [{ key: "tags", values_count: { gte: 3 } }], }, limit: TOP_K, }); ``` **Why good:** Filters by cardinality of multi-value fields, useful for quality filtering --- ## Common Filter Mistakes ```typescript // Bad: Pinecone-style syntax (does NOT work in Qdrant) filter: { $and: [{ category: { $eq: "tutorial" } }]; } // Fix: Use Qdrant must/match syntax filter: { must: [{ key: "category", match: { value: "tutorial" } }]; } // Bad: Missing "key" field in condition filter: { must: [{ match: { value: "tutorial" } }]; } // Fix: Always include the payload field key filter: { must: [{ key: "category", match: { value: "tutorial" } }]; } // Bad: Range on unindexed field (full scan) filter: { must: [{ key: "price", range: { gte: 10 } }]; } // Fix: Create index first await client.createPayloadIndex("products", { field_name: "price", field_schema: "float", }); // Bad: Using "or" instead of "should" filter: { or: [{ key: "a", match: { value: 1 } }]; } // Fix: Use "should" for OR logic filter: { should: [{ key: "a", match: { value: 1 } }]; } ``` **Why bad:** Each mistake causes either query rejection, silent empty results, or severe performance degradation --- _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_ -
named-vectors-quantization.md 8.7 KB
# Qdrant -- Named Vectors & Quantization Examples > Multiple vectors per point and quantization configuration. See [core.md](core.md) for basic operations and [reference.md](../reference.md) for quantization method comparison. **Related examples:** - [core.md](core.md) -- Client setup, upsert, query, scroll, delete - [filtering.md](filtering.md) -- Payload filtering conditions - [recommendations-batch.md](recommendations-batch.md) -- Recommend API, batch operations --- ## Named Vectors -- Collection Setup Store multiple embeddings per point (e.g., title + content embeddings from different models). ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; const TITLE_DIMENSION = 384; // e.g., all-MiniLM-L6-v2 const CONTENT_DIMENSION = 1536; // e.g., text-embedding-3-small async function createMultiVectorCollection( client: QdrantClient, ): Promise<void> { await client.createCollection("articles", { vectors: { title: { size: TITLE_DIMENSION, distance: "Cosine" }, content: { size: CONTENT_DIMENSION, distance: "Cosine" }, }, }); } export { createMultiVectorCollection, TITLE_DIMENSION, CONTENT_DIMENSION }; ``` **Why good:** Named constants for dimensions, different models can have different dimensions, each named vector has its own distance metric --- ## Named Vectors -- Upsert ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; interface ArticleVectors { title: number[]; content: number[]; } async function upsertArticleWithVectors( client: QdrantClient, id: string, vectors: ArticleVectors, payload: Record<string, unknown>, ): Promise<void> { await client.upsert("articles", { wait: true, points: [ { id, vector: vectors, // Object with named vector keys payload, }, ], }); } export { upsertArticleWithVectors }; ``` **Why good:** Typed vector interface enforces both vectors are provided, vector is an object keyed by name (not an array) --- ## Named Vectors -- Search by Specific Vector ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; const TOP_K = 10; // Search by title similarity async function searchByTitle( client: QdrantClient, titleEmbedding: number[], ): Promise<void> { const results = await client.query("articles", { query: titleEmbedding, using: "title", // Specify which named vector to search with_payload: true, limit: TOP_K, }); for (const point of results.points) { console.log(point.id, point.score, point.payload); } } // Search by content similarity async function searchByContent( client: QdrantClient, contentEmbedding: number[], ): Promise<void> { const results = await client.query("articles", { query: contentEmbedding, using: "content", with_payload: true, limit: TOP_K, }); for (const point of results.points) { console.log(point.id, point.score, point.payload); } } export { searchByTitle, searchByContent }; ``` **Why good:** `using` specifies which named vector to compare against, avoids searching the wrong vector space **Gotcha:** If you omit `using`, Qdrant searches the default (unnamed) vector. If the collection only has named vectors (no default), the query will fail. --- ## Scalar Quantization 4x memory compression by converting float32 to int8. Best default choice for most workloads. ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; const EMBEDDING_DIMENSION = 1536; const SCALAR_QUANTILE = 0.99; // Exclude top 1% outliers for better range mapping async function createScalarQuantizedCollection( client: QdrantClient, ): Promise<void> { await client.createCollection("documents-sq", { vectors: { size: EMBEDDING_DIMENSION, distance: "Cosine", }, quantization_config: { scalar: { type: "int8", quantile: SCALAR_QUANTILE, always_ram: true, // Keep quantized vectors in RAM for speed }, }, }); } export { createScalarQuantizedCollection }; ``` **Why good:** Named constants, `quantile: 0.99` excludes outliers for better compression, `always_ram: true` keeps quantized data in memory **Tradeoff:** ~4x memory savings with minimal accuracy loss (< 1% for most embedding models). Recommended as the default quantization method. --- ## Binary Quantization 32x memory compression by converting each float dimension to a single bit. Best for high-dimensional vectors (>= 1024 dims). ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; const EMBEDDING_DIMENSION = 1536; async function createBinaryQuantizedCollection( client: QdrantClient, ): Promise<void> { await client.createCollection("documents-bq", { vectors: { size: EMBEDDING_DIMENSION, distance: "Cosine", }, quantization_config: { binary: { always_ram: true, }, }, }); } export { createBinaryQuantizedCollection }; ``` **Why good:** Maximum speed (CPU XNOR + popcount operations), 32x compression, suitable for high-dimensional embeddings **Tradeoff:** Higher accuracy loss than scalar, especially for low-dimensional vectors (< 1024 dims). Use `rescore: true` in search params to improve accuracy by re-scoring top candidates with original vectors. --- ## Product Quantization Up to 64x compression. Slowest quantization method but maximizes memory savings. ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; const EMBEDDING_DIMENSION = 1536; async function createProductQuantizedCollection( client: QdrantClient, ): Promise<void> { await client.createCollection("documents-pq", { vectors: { size: EMBEDDING_DIMENSION, distance: "Cosine", }, quantization_config: { product: { compression: "x16", // x4, x8, x16, x32, x64 always_ram: true, }, }, }); } export { createProductQuantizedCollection }; ``` **Why good:** Configurable compression ratio, `always_ram: true` for speed **Tradeoff:** Significant accuracy loss and slower search compared to scalar/binary. Use only when memory is the primary constraint and accuracy is secondary. --- ## Quantization with Search Params (Rescore) Enable rescoring to improve accuracy for quantized collections. ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; const TOP_K = 10; const OVERSAMPLING = 2.0; // Retrieve 2x candidates for rescoring const results = await client.query("documents-bq", { query: embedding, params: { quantization: { rescore: true, // Re-score top candidates with original float32 vectors oversampling: OVERSAMPLING, // Retrieve more candidates before rescoring }, }, with_payload: true, limit: TOP_K, }); ``` **Why good:** `rescore: true` re-evaluates top candidates using full-precision vectors, `oversampling` retrieves extra candidates to improve recall after rescoring **When to use:** Always enable rescoring for binary quantization. Optional but recommended for product quantization. Usually unnecessary for scalar quantization. --- ## Per-Vector Quantization (Named Vectors) Apply different quantization to different named vectors within the same collection. ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; const TITLE_DIMENSION = 384; const CONTENT_DIMENSION = 1536; await client.createCollection("articles-optimized", { vectors: { title: { size: TITLE_DIMENSION, distance: "Cosine", quantization_config: { scalar: { type: "int8", always_ram: true }, }, }, content: { size: CONTENT_DIMENSION, distance: "Cosine", quantization_config: { binary: { always_ram: true }, }, }, }, }); ``` **Why good:** Scalar quantization for low-dim title vectors (better accuracy), binary quantization for high-dim content vectors (better compression), per-vector optimization --- ## Add Quantization to Existing Collection ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; const SCALAR_QUANTILE = 0.99; async function enableQuantization( client: QdrantClient, collectionName: string, ): Promise<void> { await client.updateCollection(collectionName, { quantization_config: { scalar: { type: "int8", quantile: SCALAR_QUANTILE, always_ram: true, }, }, }); } export { enableQuantization }; ``` **Why good:** Quantization can be added retroactively via `updateCollection`, no re-ingestion needed **Gotcha:** Adding quantization to a large existing collection triggers a background rebuild of quantized vectors. Collection remains available during rebuild, but CPU usage increases temporarily. --- _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_ -
recommendations-batch.md 10.3 KB
# Qdrant -- Recommendations, Batch Operations & Snapshots > Recommend API, batch operations, and snapshot management. See [core.md](core.md) for basic operations. **Related examples:** - [core.md](core.md) -- Client setup, upsert, query, scroll, delete - [filtering.md](filtering.md) -- Payload filtering conditions - [named-vectors-quantization.md](named-vectors-quantization.md) -- Named vectors, quantization --- ## Recommend by Positive/Negative Examples Find points similar to positive examples and dissimilar to negative examples. ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; const TOP_K = 10; async function recommendSimilar( client: QdrantClient, collectionName: string, positiveIds: (string | number)[], negativeIds?: (string | number)[], ): Promise<void> { const results = await client.recommend(collectionName, { positive: positiveIds, negative: negativeIds ?? [], limit: TOP_K, with_payload: true, }); for (const point of results) { console.log(point.id, point.score, point.payload); } } export { recommendSimilar }; ``` **Why good:** Accepts point IDs as examples, optional negatives, `with_payload` for context --- ## Recommend with Strategy Selection ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; const TOP_K = 10; // Using query() with recommend -- preferred API async function recommendWithStrategy( client: QdrantClient, collectionName: string, ): Promise<void> { // "average_vector" (default): averages positive/negative vectors // Good when examples are from the same cluster const averaged = await client.query(collectionName, { query: { recommend: { positive: [1, 42], negative: [7], strategy: "average_vector", }, }, limit: TOP_K, with_payload: true, }); // "best_score": scores each candidate against each example individually // Better when positives span multiple clusters or negatives are important const bestScore = await client.query(collectionName, { query: { recommend: { positive: [1, 42], negative: [7], strategy: "best_score", }, }, limit: TOP_K, with_payload: true, }); } export { recommendWithStrategy }; ``` **Why good:** Shows both strategies with clear use-case guidance, uses `query()` universal endpoint **When to use each:** - `average_vector`: Fast, works well when positive examples are conceptually close - `best_score`: Better accuracy when positives span diverse topics or when negatives are critical --- ## Recommend with Raw Vectors Mix point IDs and raw vectors as positive/negative examples. ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; const TOP_K = 10; const results = await client.query("documents", { query: { recommend: { positive: [ 42, // Existing point ID [0.2, 0.3, 0.4, 0.5], // Raw vector (from external source) ], negative: [7], }, }, limit: TOP_K, with_payload: true, }); ``` **Why good:** Combines stored point IDs with external vectors, useful when some reference items are not in the collection --- ## Recommend with Named Vectors ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; const TOP_K = 10; // Recommend based on "content" named vector const results = await client.recommend("articles", { positive: [1, 42], using: "content", // Which named vector to use for similarity limit: TOP_K, with_payload: true, }); ``` **Why good:** `using` specifies which named vector drives the recommendation --- ## Recommend with Filter ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; const TOP_K = 10; const results = await client.recommend("documents", { positive: [1, 42], filter: { must: [{ key: "category", match: { value: "tutorial" } }], must_not: [{ key: "status", match: { value: "draft" } }], }, limit: TOP_K, with_payload: true, }); ``` **Why good:** Combines recommendation with payload filtering, ensures results match business constraints --- ## Batch Upsert with Chunking ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; const UPSERT_BATCH_SIZE = 200; interface VectorPoint { id: string; vector: number[]; payload: Record<string, unknown>; } function chunkArray<T>(array: T[], size: number): T[][] { const chunks: T[][] = []; for (let i = 0; i < array.length; i += size) { chunks.push(array.slice(i, i + size)); } return chunks; } async function batchUpsert( client: QdrantClient, collectionName: string, points: VectorPoint[], ): Promise<void> { const batches = chunkArray(points, UPSERT_BATCH_SIZE); for (const batch of batches) { await client.upsert(collectionName, { wait: true, points: batch, }); } } export { batchUpsert, chunkArray }; ``` **Why good:** Named constant for batch size, reusable chunking utility, `wait: true` per batch for consistency --- ## Parallel Batch Upsert ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; const UPSERT_BATCH_SIZE = 200; const MAX_CONCURRENT = 5; async function parallelBatchUpsert( client: QdrantClient, collectionName: string, points: Array<{ id: string; vector: number[]; payload: Record<string, unknown>; }>, ): Promise<void> { const batches = chunkArray(points, UPSERT_BATCH_SIZE); for (let i = 0; i < batches.length; i += MAX_CONCURRENT) { const window = batches.slice(i, i + MAX_CONCURRENT); await Promise.all( window.map((batch) => client.upsert(collectionName, { wait: true, points: batch }), ), ); } } function chunkArray<T>(array: T[], size: number): T[][] { const chunks: T[][] = []; for (let i = 0; i < array.length; i += size) { chunks.push(array.slice(i, i + size)); } return chunks; } export { parallelBatchUpsert }; ``` **Why good:** Bounded concurrency prevents overwhelming the server, `Promise.all` for parallel execution within each window --- ## Batch Upsert with Retry Logic ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; const UPSERT_BATCH_SIZE = 200; const MAX_RETRIES = 3; const INITIAL_BACKOFF_MS = 1000; async function upsertWithRetry( client: QdrantClient, collectionName: string, points: Array<{ id: string; vector: number[]; payload: Record<string, unknown>; }>, ): Promise<void> { for (let attempt = 0; attempt < MAX_RETRIES; attempt++) { try { await client.upsert(collectionName, { wait: true, points }); return; } catch (error) { const isRetryable = error instanceof Error && (error.message.includes("429") || error.message.includes("503") || error.message.includes("ECONNRESET")); if (!isRetryable || attempt === MAX_RETRIES - 1) { throw error; } const backoffMs = INITIAL_BACKOFF_MS * Math.pow(2, attempt); await new Promise((resolve) => setTimeout(resolve, backoffMs)); } } } export { upsertWithRetry }; ``` **Why good:** Exponential backoff for rate limits (429), server errors (503), and network issues, named constants, non-retryable errors propagate immediately --- ## Batch Update Operations Use `batchUpdate` for multiple different operations in a single request. ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; async function batchOperations(client: QdrantClient): Promise<void> { await client.batchUpdate("documents", { wait: true, operations: [ // Upsert new points { upsert: { points: [ { id: "new-1", vector: [0.1, 0.2, 0.3], payload: { title: "New Doc" }, }, ], }, }, // Update payload on existing points { set_payload: { payload: { reviewed: true }, points: ["existing-1", "existing-2"], }, }, // Delete old points { delete: { points: ["old-1", "old-2"], }, }, ], }); } export { batchOperations }; ``` **Why good:** Multiple operation types in one atomic request, reduces network round-trips, `wait: true` for consistency --- ## Create and Recover Snapshots ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; // Create a snapshot for backup async function backupCollection( client: QdrantClient, collectionName: string, ): Promise<string> { const snapshot = await client.createSnapshot(collectionName); console.log("Snapshot created:", snapshot.name, "size:", snapshot.size); return snapshot.name; } // List available snapshots async function listBackups( client: QdrantClient, collectionName: string, ): Promise<void> { const snapshots = await client.listSnapshots(collectionName); for (const snap of snapshots) { console.log(snap.name, snap.creation_time, snap.size); } } // Recover from a snapshot URL async function restoreCollection( client: QdrantClient, collectionName: string, snapshotUrl: string, ): Promise<void> { await client.recoverSnapshot(collectionName, { location: snapshotUrl, }); } // Delete a snapshot async function deleteBackup( client: QdrantClient, collectionName: string, snapshotName: string, ): Promise<void> { await client.deleteSnapshot(collectionName, snapshotName); } export { backupCollection, listBackups, restoreCollection, deleteBackup }; ``` **Why good:** Complete snapshot lifecycle (create, list, recover, delete), named exports **Gotcha:** Snapshot recovery requires matching Qdrant minor versions. A v1.14.x snapshot cannot be restored to a v1.15.x cluster. Always verify version compatibility before recovery. --- ## Full Instance Snapshot ```typescript import type { QdrantClient } from "@qdrant/js-client-rest"; async function fullBackup(client: QdrantClient): Promise<string> { const snapshot = await client.createFullSnapshot(); console.log("Full snapshot:", snapshot.name); return snapshot.name; } async function listFullBackups(client: QdrantClient): Promise<void> { const snapshots = await client.listFullSnapshots(); for (const snap of snapshots) { console.log(snap.name, snap.creation_time, snap.size); } } export { fullBackup, listFullBackups }; ``` **Why good:** Backs up ALL collections in one snapshot, useful for disaster recovery --- _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_
-
-
reference.md 12.8 KB
# Qdrant Quick Reference > API reference, filter operators, limits, decision frameworks, and production checklist. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples. --- ## QdrantClient Methods -- Collections | Method | Description | Returns | | ------------------------------------- | ------------------------------------ | ------------------------------ | | `new QdrantClient({ url, apiKey })` | Create client (local or cloud) | `QdrantClient` | | `client.createCollection(name, args)` | Create collection with vector config | `Promise<boolean>` | | `client.getCollection(name)` | Get collection info | `Promise<CollectionInfo>` | | `client.getCollections()` | List all collections | `Promise<CollectionsResponse>` | | `client.updateCollection(name, args)` | Update optimizer/quantization config | `Promise<boolean>` | | `client.deleteCollection(name)` | Delete a collection | `Promise<boolean>` | | `client.collectionExists(name)` | Check if collection exists | `Promise<CollectionExistence>` | ## QdrantClient Methods -- Points | Method | Description | Returns | | -------------------------------- | ----------------------------------- | ------------------------- | | `client.upsert(name, args)` | Insert or update points | `Promise<UpdateResult>` | | `client.retrieve(name, args)` | Fetch points by ID | `Promise<Record[]>` | | `client.delete(name, args)` | Delete by IDs or filter | `Promise<UpdateResult>` | | `client.scroll(name, args)` | Paginate through points | `Promise<ScrollResult>` | | `client.count(name, args)` | Count points (exact or approximate) | `Promise<CountResult>` | | `client.batchUpdate(name, args)` | Multiple operations in one call | `Promise<UpdateResult[]>` | ## QdrantClient Methods -- Search & Query | Method | Description | Returns | | ------------------------------------- | --------------------------------- | -------------------------- | | `client.query(name, args)` | Universal search (preferred) | `Promise<QueryResponse>` | | `client.queryBatch(name, args)` | Multiple queries in one call | `Promise<QueryResponse[]>` | | `client.queryPointGroups(name, args)` | Grouped search by payload field | `Promise<GroupsResult>` | | `client.search(name, args)` | Vector similarity search (legacy) | `Promise<ScoredPoint[]>` | | `client.searchBatch(name, args)` | Multiple searches in one call | `Promise<ScoredPoint[][]>` | ## QdrantClient Methods -- Recommendations | Method | Description | Returns | | ----------------------------------------- | --------------------------------------- | -------------------------- | | `client.recommend(name, args)` | Recommend by positive/negative examples | `Promise<ScoredPoint[]>` | | `client.recommendBatch(name, args)` | Multiple recommendations in one call | `Promise<ScoredPoint[][]>` | | `client.recommendPointGroups(name, args)` | Grouped recommendations | `Promise<GroupsResult>` | ## QdrantClient Methods -- Payload | Method | Description | Returns | | --------------------------------------- | ----------------------------------- | ----------------------- | | `client.setPayload(name, args)` | Merge payload fields into points | `Promise<UpdateResult>` | | `client.overwritePayload(name, args)` | Replace entire payload | `Promise<UpdateResult>` | | `client.deletePayload(name, args)` | Remove specific payload keys | `Promise<UpdateResult>` | | `client.clearPayload(name, args)` | Remove all payload from points | `Promise<UpdateResult>` | | `client.createPayloadIndex(name, args)` | Index a payload field for filtering | `Promise<UpdateResult>` | | `client.deletePayloadIndex(name, args)` | Remove a payload field index | `Promise<UpdateResult>` | ## QdrantClient Methods -- Vectors | Method | Description | Returns | | ---------------------------------- | ---------------------------------- | ----------------------- | | `client.updateVectors(name, args)` | Update vectors for existing points | `Promise<UpdateResult>` | | `client.deleteVectors(name, args)` | Remove named vectors from points | `Promise<UpdateResult>` | ## QdrantClient Methods -- Snapshots | Method | Description | Returns | | ------------------------------------------- | ------------------------------------ | -------------------------------- | | `client.createSnapshot(name)` | Create collection snapshot | `Promise<SnapshotDescription>` | | `client.listSnapshots(name)` | List collection snapshots | `Promise<SnapshotDescription[]>` | | `client.deleteSnapshot(name, snapshotName)` | Delete a snapshot | `Promise<boolean>` | | `client.recoverSnapshot(name, args)` | Recover collection from snapshot URL | `Promise<boolean>` | | `client.createFullSnapshot()` | Snapshot entire Qdrant instance | `Promise<SnapshotDescription>` | | `client.listFullSnapshots()` | List full instance snapshots | `Promise<SnapshotDescription[]>` | --- ## Filter Condition Types | Condition | Description | Example | | ------------------ | -------------------------------- | ------------------------------------------------------------------- | | `match.value` | Exact equality | `{ key: "city", match: { value: "London" } }` | | `match.text` | Full-text substring match | `{ key: "description", match: { text: "vector" } }` | | `match.any` | Match any value in list | `{ key: "city", match: { any: ["London", "Paris"] } }` | | `match.except` | Exclude values in list | `{ key: "city", match: { except: ["Berlin"] } }` | | `range` | Numeric range (gt, gte, lt, lte) | `{ key: "price", range: { gte: 10, lte: 100 } }` | | `values_count` | Count of payload values | `{ key: "tags", values_count: { gte: 2 } }` | | `geo_bounding_box` | Within rectangular area | `{ key: "location", geo_bounding_box: { top_left, bottom_right } }` | | `geo_radius` | Within circular area | `{ key: "location", geo_radius: { center, radius } }` | | `is_empty` | Field has no value | `{ is_empty: { key: "description" } }` | | `is_null` | Field is explicitly null | `{ is_null: { key: "deleted_at" } }` | | `has_id` | Match by point IDs | `{ has_id: [1, 2, 3] }` | | `nested` | Filter on nested payload objects | `{ nested: { key: "items", filter: { must: [...] } } }` | ## Filter Logic Operators | Operator | Description | Behavior | | ---------- | ------------------------------- | ------------------------------------------ | | `must` | All conditions must match (AND) | `filter: { must: [cond1, cond2] }` | | `should` | At least one must match (OR) | `filter: { should: [cond1, cond2] }` | | `must_not` | None may match (NOT) | `filter: { must_not: [cond1] }` | | Combined | Mix operators at top level | `filter: { must: [...], must_not: [...] }` | --- ## Distance Metrics | Metric | Description | When to Use | | ----------- | ------------------------------ | ---------------------------------------------- | | `Cosine` | Cosine similarity (normalized) | Most embedding models (safe default) | | `Dot` | Dot product | Pre-normalized embeddings (faster than Cosine) | | `Euclid` | L2 distance | Raw feature vectors where magnitude matters | | `Manhattan` | L1 / city-block distance | Specialized use cases | --- ## Payload Field Types & Index Types | Payload Type | Example Value | Index Type | Notes | | ------------ | -------------------------- | ---------- | --------------------------------- | | keyword | `"tutorial"` | `keyword` | Exact match, enum-like values | | integer | `2024` | `integer` | Range and exact match | | float | `0.95` | `float` | Range queries | | bool | `true` | `bool` | Exact match only | | text | `"long description"` | `text` | Full-text tokenized search | | geo | `{ lat: 51.5, lon: -0.1 }` | `geo` | Bounding box and radius queries | | datetime | `"2024-01-15T10:30:00Z"` | `datetime` | Range queries on ISO 8601 strings | --- ## Quantization Methods | Method | Compression | Speed | Accuracy | Best For | | -------------- | ----------- | ------- | -------- | ------------------------------------------ | | Scalar (int8) | 4x | Fast | High | Default choice, balanced tradeoff | | Binary (1-bit) | 32x | Fastest | Lower | High-dim vectors (>= 1024), speed-critical | | Product | Up to 64x | Slowest | Lowest | Memory-critical, accuracy secondary | --- ## Key Limits | Resource | Limit | | ----------------------------- | -------------------------------------------------------------- | | Point ID | Positive integer or UUID string (no zero, no negative) | | Vector dimensions | No hard limit (practical: up to 65,535) | | Payload per point | No hard limit (practical: keep under a few KB for performance) | | Points per upsert batch | No hard limit (practical: 100-500 for network reliability) | | `scroll` page size (`limit`) | Max 10,000 per page | | `query`/`search` result limit | Configurable (practical: keep under 1,000) | | Snapshot recovery | Must match Qdrant minor version (v1.14.x to v1.14.x) | --- ## Production Checklist ### Security - [ ] API key stored in environment variable, not in code - [ ] Client used server-side only (never expose API key in browser) - [ ] HTTPS enabled for cloud deployments (`url` includes `https://`) ### Collection Configuration - [ ] Vector dimension matches embedding model output exactly - [ ] Distance metric matches model recommendation (Cosine for most) - [ ] HNSW parameters tuned if needed (default `m: 16`, `ef_construct: 100` works for most) - [ ] Quantization configured if memory optimization needed ### Payload Indexes - [ ] `createPayloadIndex()` called for every field used in filter conditions - [ ] `createPayloadIndex()` called for fields used in `order_by` (scroll) - [ ] Index type matches field type (`keyword` for strings, `integer` for numbers) - [ ] Indexes created before bulk data ingestion (faster than retroactive indexing) ### Write Operations - [ ] `wait: true` set on upserts when read-after-write consistency is required - [ ] Batch size reasonable (100-500 points per upsert for network reliability) - [ ] Point IDs are positive integers or UUID strings (not zero, not negative) - [ ] Retry logic with exponential backoff for transient failures ### Query Optimization - [ ] `with_payload: true` included when payload data is needed (not included by default) - [ ] `limit` set to minimum needed (lower = faster) - [ ] Filters use indexed payload fields - [ ] `score_threshold` set to skip low-relevance results when applicable ### Monitoring - [ ] Collection info tracked via `getCollection()` (vector count, segment info) - [ ] Query latency monitored - [ ] Snapshot schedule configured for backup --- _Full skill documentation: [SKILL.md](SKILL.md) | Examples: [examples/](examples/)_ -
SKILL.md 15.4 KB
--- name: api-vector-db-qdrant description: Qdrant vector database -- collection management, point operations, payload filtering, named vectors, quantization, recommendations, snapshots --- # Qdrant Patterns > **Quick Guide:** Use `@qdrant/js-client-rest` (v1.17.x) for high-performance vector search. Collections define vector dimensions and distance metrics upfront -- mismatches cause silent failures. Use `must`/`should`/`must_not` filter clauses with payload conditions (not Pinecone-style `$eq`/`$and`). Payload indexes are optional but critical for filter performance at scale -- create them explicitly with `createPayloadIndex()`. Named vectors let you store multiple embeddings per point (e.g., title + content). Quantization (scalar/binary/product) trades accuracy for memory and speed. The `query()` method is the universal search endpoint -- prefer it over the older `search()` method. --- <critical_requirements> ## CRITICAL: Before Using This Skill > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST create payload indexes with `createPayloadIndex()` for any field used in filters -- unindexed fields cause full scans that degrade linearly with collection size)** **(You MUST use `must`/`should`/`must_not` filter syntax -- Qdrant does NOT use `$eq`/`$and`/`$or` operators like Pinecone)** **(You MUST match vector dimensions exactly between embedding model output and collection config -- dimension mismatches cause silent upsert failures or corrupt search results)** **(You MUST set `wait: true` on writes when subsequent reads depend on the data -- Qdrant writes are asynchronous by default and may not be immediately visible)** </critical_requirements> --- ## Examples - [Core Patterns](examples/core.md) -- Client setup, collection creation, upsert, query, scroll, delete - [Filtering](examples/filtering.md) -- must/should/must_not conditions, match/range operators, payload indexes - [Named Vectors & Quantization](examples/named-vectors-quantization.md) -- Multiple vectors per point, scalar/binary/product quantization - [Recommendations & Batch](examples/recommendations-batch.md) -- Recommend API, batch operations, snapshots **Additional resources:** - [reference.md](reference.md) -- API quick reference, filter operators, limits, decision frameworks, production checklist --- **Auto-detection:** Qdrant, QdrantClient, @qdrant/js-client-rest, createCollection, upsert, query, scroll, recommend, setPayload, createPayloadIndex, must, should, must_not, payload, named vectors, quantization, vector database, similarity search, semantic search, RAG retrieval, embedding search **When to use:** - Semantic search over document embeddings (RAG retrieval pipelines) - Similarity search for recommendations, deduplication, or classification - Multi-vector search with named vectors (e.g., title embedding + content embedding per document) - Filtered vector search with complex payload conditions (must/should/must_not) - Memory-optimized deployments using scalar, binary, or product quantization **Key patterns covered:** - Client setup and collection management (distance metrics, HNSW config) - Point CRUD operations (upsert, query, scroll, retrieve, delete, count) - Payload filtering with must/should/must_not and match/range conditions - Named vectors for multiple embeddings per point - Quantization configuration (scalar, binary, product) - Recommendation API with positive/negative examples - Batch operations and snapshot management - Payload indexing for filter performance **When NOT to use:** - Full-text search with BM25 ranking (use a dedicated search engine) - Relational data with joins and transactions (use a relational database) - Key-value lookups without vector similarity (use a KV store) - Storing large documents or binary blobs (store embeddings + metadata references only) --- <philosophy> ## Philosophy Qdrant is a **high-performance open-source vector database** built in Rust, designed for filtered similarity search at scale. The core principle: **store vectors with rich payloads, search by similarity, filter by payload conditions.** **Core principles:** 1. **Payload is first-class** -- Unlike databases that treat metadata as secondary, Qdrant's payload system supports complex nested JSON, multiple data types, and granular indexing. Use payloads for filtering, not just annotation. 2. **Index what you filter** -- Payload indexes are not automatic. Create explicit indexes on fields used in filters via `createPayloadIndex()`. Without indexes, filters cause full collection scans. 3. **Named vectors for multi-modal** -- A single point can hold multiple named vectors (e.g., title embedding + content embedding). Search targets a specific named vector. This avoids duplicating payloads across collections. 4. **Quantization for scale** -- Scalar (4x compression), binary (32x), and product quantization trade accuracy for memory savings. Configure at collection or per-vector level. Use `always_ram: true` to keep quantized vectors in memory for speed. 5. **Writes are async by default** -- Upserts return before data is persisted to all replicas. Set `wait: true` when immediate consistency matters (e.g., read-after-write flows). </philosophy> --- <patterns> ## Core Patterns ### Pattern 1: Client Initialization Create a QdrantClient connected to a local instance or Qdrant Cloud. See [examples/core.md](examples/core.md) for full examples. ```typescript // Good Example import { QdrantClient } from "@qdrant/js-client-rest"; function createQdrantClient(): QdrantClient { const url = process.env.QDRANT_URL; const apiKey = process.env.QDRANT_API_KEY; if (!url) { throw new Error("QDRANT_URL environment variable is required"); } return new QdrantClient({ url, apiKey }); } export { createQdrantClient }; ``` **Why good:** URL and API key from environment, validation before construction, named export ```typescript // Bad Example import { QdrantClient } from "@qdrant/js-client-rest"; const client = new QdrantClient({ host: "my-cluster.cloud.qdrant.io", apiKey: "sk-abc123...", }); // Hardcoded credentials leak in version control ``` **Why bad:** Hardcoded API key, host without HTTPS (use `url` with full protocol for cloud) --- ### Pattern 2: Collection Creation Define vector dimensions and distance metric. Dimension must exactly match your embedding model output. See [examples/core.md](examples/core.md). ```typescript // Good Example const EMBEDDING_DIMENSION = 1536; await client.createCollection("documents", { vectors: { size: EMBEDDING_DIMENSION, distance: "Cosine", }, }); export { EMBEDDING_DIMENSION }; ``` **Why good:** Named constant for dimension, explicit distance metric, clean config ```typescript // Bad Example await client.createCollection("documents", { vectors: { size: 768, distance: "Cosine" }, // Dimension mismatch if using a 1536-dim model -- upserts may silently fail or produce garbage search results }); ``` **Why bad:** Hardcoded dimension that may not match embedding model, no named constant --- ### Pattern 3: Upsert Points with Payload Upsert vectors with payload (Qdrant's term for metadata). See [examples/core.md](examples/core.md). ```typescript // Good Example interface DocumentPayload { title: string; category: string; createdAt: number; tags: string[]; } await client.upsert("documents", { wait: true, points: [ { id: "doc-1", vector: embedding, payload: { title: "Guide", category: "tutorial", createdAt: 1710000000, tags: ["ai", "search"], }, }, ], }); ``` **Why good:** Typed payload interface, `wait: true` for immediate consistency, structured payload --- ### Pattern 4: Query with Payload Filter Use `must`/`should`/`must_not` filter clauses -- NOT Pinecone-style `$eq`/`$and`. See [examples/filtering.md](examples/filtering.md). ```typescript // Good Example const TOP_K = 10; const results = await client.query("documents", { query: queryEmbedding, filter: { must: [ { key: "category", match: { value: "tutorial" } }, { key: "createdAt", range: { gte: 1700000000 } }, ], }, with_payload: true, limit: TOP_K, }); for (const point of results.points) { console.log(point.id, point.score, point.payload); } ``` **Why good:** Named constant for limit, Qdrant filter syntax (must + match/range), `with_payload` included ```typescript // Bad Example -- Pinecone syntax does NOT work in Qdrant const results = await client.query("documents", { query: embedding, filter: { $and: [{ category: { $eq: "tutorial" } }], }, limit: 100, }); ``` **Why bad:** Pinecone-style `$and`/`$eq` operators are invalid in Qdrant, magic number for limit --- ### Pattern 5: Named Vectors Store multiple embeddings per point. See [examples/named-vectors-quantization.md](examples/named-vectors-quantization.md). ```typescript // Good Example const TITLE_DIM = 384; const CONTENT_DIM = 1536; await client.createCollection("articles", { vectors: { title: { size: TITLE_DIM, distance: "Cosine" }, content: { size: CONTENT_DIM, distance: "Cosine" }, }, }); // Upsert with named vectors await client.upsert("articles", { wait: true, points: [ { id: "article-1", vector: { title: titleEmbedding, content: contentEmbedding }, payload: { title: "Intro to Vectors" }, }, ], }); // Search by specific named vector const results = await client.query("articles", { query: queryEmbedding, using: "content", limit: TOP_K, }); ``` **Why good:** Different dimensions per named vector, `using` specifies which vector to search, avoids duplicating payloads across collections --- ### Pattern 6: Recommendation API Find similar points using positive/negative examples. See [examples/recommendations-batch.md](examples/recommendations-batch.md). ```typescript // Good Example const results = await client.query("documents", { query: { recommend: { positive: [1, 42], negative: [7], strategy: "best_score", }, }, limit: TOP_K, with_payload: true, }); ``` **Why good:** Uses point IDs as positive/negative examples, `best_score` strategy handles negatives better than default `average_vector` </patterns> --- <decision_framework> ## Decision Framework ### Which Distance Metric? ``` Which distance metric should I use? |-- Using normalized embeddings (OpenAI, Cohere)? -> Cosine (most common, safe default) |-- Pre-normalized embeddings and need speed? -> Dot (faster, same results as Cosine for unit vectors) |-- Raw feature vectors where magnitude matters? -> Euclid (L2 distance) |-- City-block distance needed? -> Manhattan '-- Unsure? -> Cosine (works with any embedding model) ``` ### Single Vector vs Named Vectors? ``` How many embeddings per point? |-- One embedding model? -> Single vector (simpler config) |-- Multiple embedding models (title + content)? -> Named vectors |-- Same model, different text segments? -> Named vectors |-- Multi-modal (text + image)? -> Named vectors with different dimensions '-- Want to avoid duplicating payloads across collections? -> Named vectors ``` ### Which Quantization Method? ``` How should I optimize memory? |-- Good default, balanced accuracy/speed? -> Scalar (int8, 4x compression) |-- Maximum speed, can tolerate accuracy loss? -> Binary (32x compression) | '-- Best with high-dimensional models (>= 1024 dims) |-- Maximum compression, speed not critical? -> Product (up to 64x compression) | '-- Slowest quantization, most accuracy loss '-- No memory pressure? -> Skip quantization (full float32 precision) ``` ### Payload Index Strategy? ``` Should I create a payload index? |-- Field used in filter conditions? -> YES, always index |-- Field used in order_by for scroll? -> YES, index for sort performance |-- Field only read after search (display only)? -> NO, skip index |-- High-cardinality field (UUIDs, timestamps)? -> YES, but evaluate index type '-- Low-cardinality field (enum-like)? -> YES, keyword index is very efficient ``` </decision_framework> --- <red_flags> ## RED FLAGS **High Priority Issues:** - Using Pinecone-style filter syntax (`$eq`, `$and`, `$or`) -- Qdrant uses `must`/`should`/`must_not` with `match`/`range` conditions - Vector dimension mismatch between embedding model and collection config -- causes silent failures or garbage results - Missing payload indexes on filtered fields -- causes full collection scans that degrade linearly with size - Not setting `wait: true` when read-after-write consistency is needed -- writes are async by default **Medium Priority Issues:** - Using deprecated `search()` method instead of `query()` -- `query()` is the universal endpoint with prefetch and fusion support - Forgetting `with_payload: true` in queries -- payload is NOT included by default - Creating payload indexes after bulk upsert instead of before -- retroactive indexing is slower than indexing during upsert - Using `offset` for deep pagination in scroll -- performance degrades; use `offset` as cursor (point ID), not page number **Common Mistakes:** - Passing `filter` at the wrong nesting level -- filter goes at the top level of the query args, not nested inside another object - Using `id: 0` as a point ID -- Qdrant requires positive integers or UUID strings; 0 is invalid - Confusing `setPayload` (merge) with `overwritePayload` (replace) -- `setPayload` merges fields, `overwritePayload` replaces the entire payload - Calling `deletePayload` with field names but no point selector -- you must specify which points to update via `points` array or `filter` **Gotchas & Edge Cases:** - Point IDs must be positive integers or UUID strings -- negative numbers, zero, and non-UUID strings are rejected - `scroll()` with `order_by` requires a payload index on the sort field -- without it, the request fails - `count()` with `exact: true` is slow on large collections -- use `exact: false` (default) for approximate counts - Snapshot recovery requires matching Qdrant minor versions -- a v1.14.x snapshot cannot be restored to a v1.15.x cluster - Binary quantization works best with high-dimensional vectors (>= 1024 dims) -- for smaller vectors, scalar quantization is more accurate - `query()` with `prefetch` enables multi-stage retrieval (retrieve 1000, then re-rank to top 10) -- but requires understanding the prefetch pipeline - Named vector search requires the `using` parameter -- omitting it searches the default (unnamed) vector, which may not exist - `deletePayload` removes specific keys, `clearPayload` removes ALL keys -- they are different operations </red_flags> --- <critical_reminders> ## CRITICAL REMINDERS > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST create payload indexes with `createPayloadIndex()` for any field used in filters -- unindexed fields cause full scans that degrade linearly with collection size)** **(You MUST use `must`/`should`/`must_not` filter syntax -- Qdrant does NOT use `$eq`/`$and`/`$or` operators like Pinecone)** **(You MUST match vector dimensions exactly between embedding model output and collection config -- dimension mismatches cause silent upsert failures or corrupt search results)** **(You MUST set `wait: true` on writes when subsequent reads depend on the data -- Qdrant writes are asynchronous by default and may not be immediately visible)** **Failure to follow these rules will cause empty search results, degraded filter performance, data consistency issues, and hard-to-debug dimension mismatch errors.** </critical_reminders>
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.