api-vector-db-pinecone
Pinecone serverless vector database -- index management, vector operations, metadata filtering, namespaces, hybrid search, inference API
Install
npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/api-vector-db-pinecone/skills/api-vector-db-pinecone
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
git clone https://github.com/agents-inc/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Pinecone Patterns
Quick Guide: Use
@pinecone-database/pinecone(v7.x) for serverless vector database operations. Target indexes by host (pc.index({ host })), not by name. Use namespaces for multi-tenant isolation (physically separate, cheaper queries). Batch upserts at 200 records (max 1,000 or 2 MB). Metadata is limited to 40 KB per record with flat key-value pairs only (no nested objects). Pinecone is eventually consistent -- vectors may not appear in queries immediately after upsert. UsedescribeIndexStats()to verify indexing progress. For hybrid search, usedotproductmetric with sparse+dense vectors in a single index.
<critical_requirements>
CRITICAL: Before Using This Skill
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST target indexes by host URL, not by name -- pc.index({ host }) is the v7 API; pc.index('name') is deprecated)
(You MUST batch upserts to max 1,000 records or 2 MB per request -- exceeding either limit causes a 400 error)
(You MUST use flat key-value metadata only -- nested objects, null values, and keys starting with $ are rejected by Pinecone)
(You MUST handle eventual consistency -- vectors are not queryable immediately after upsert; use describeIndexStats() or retry logic for freshness-critical flows)
</critical_requirements>
Examples
- Core Patterns -- Client setup, index creation, upsert, query, fetch, update, delete
- Namespaces & Multi-Tenancy -- Namespace isolation, multi-tenant patterns, namespace management API
- Metadata Filtering -- Filter operators, compound filters, best practices
- Hybrid Search -- Sparse-dense vectors, hybrid index setup, alpha weighting
- Inference API -- Embedding generation, reranking, integrated inference indexes
- Batch Operations -- Chunked upserts, parallel ingestion, bulk import
Additional resources:
- reference.md -- API quick reference, filter operators, limits, decision frameworks, production checklist
Auto-detection: Pinecone, @pinecone-database/pinecone, createIndex, createIndexForModel, upsert, query, topK, includeMetadata, sparseValues, namespace, describeIndexStats, vector database, similarity search, embedding, cosine, dotproduct, euclidean, RAG retrieval, semantic search, pinecone-sparse-english, rerank, searchRecords, upsertRecords, fetchByMetadata
When to use:
- Semantic search over document embeddings (RAG retrieval)
- Similarity search for recommendations, deduplication, or classification
- Multi-tenant vector isolation using namespaces
- Hybrid semantic + keyword search using sparse-dense vectors
- Embedding generation and result reranking via Pinecone Inference API
Key patterns covered:
- Client setup and index management (serverless vs pod-based)
- Vector CRUD operations (upsert, query, fetch, update, delete)
- Metadata filtering with compound operators
- Namespace-based multi-tenancy
- Sparse-dense hybrid search
- Pinecone Inference API (embed, rerank)
- Batch ingestion with chunking and parallelism
- Integrated inference indexes (automatic embedding)
When NOT to use:
- Full-text search with complex boolean queries (use a dedicated search engine)
- Relational data with joins and transactions (use a relational database)
- Real-time streaming or pub/sub messaging (use a message broker)
- Storing large binary blobs or documents (use object storage; store only embeddings + metadata references)
<decision_framework>
Decision Framework
Which Index Type?
Which Pinecone index type should I use?
|-- Serverless? (recommended for most use cases)
| |-- Variable or unpredictable traffic? -> Serverless (auto-scales, pay-per-use)
| |-- Starting a new project? -> Serverless (simpler, no capacity planning)
| '-- Need hybrid sparse-dense search? -> Serverless with dotproduct metric
|
'-- Pod-based? (legacy, specific needs)
|-- Need guaranteed low latency SLAs? -> Pod-based (dedicated compute)
'-- Using collections for snapshots? -> Pod-based (collections are pod-only)
Which Metric?
Which distance metric should I use?
|-- Using embeddings from a language model? -> cosine (normalized, most common)
|-- Need hybrid search (sparse + dense)? -> dotproduct (REQUIRED for hybrid)
|-- Comparing raw feature vectors? -> euclidean (absolute distance matters)
'-- Unsure? -> cosine (safe default for most embedding models)
Namespaces vs Metadata Filtering?
How should I isolate tenant data?
|-- Strict data isolation required? -> Namespaces (physical separation)
|-- Need to query across tenants? -> Metadata filtering (logical separation)
|-- Cost-sensitive at scale? -> Namespaces (query cost = tenant size, not total)
|-- Few tenants (< 10)? -> Either approach works
'-- Many tenants (100+)? -> Namespaces (metadata filtering scans everything)
Embedded Inference vs External Embeddings?
How should I generate embeddings?
|-- Want simplest architecture? -> Integrated inference (createIndexForModel)
| (Pinecone handles embedding automatically on upsert/query)
|
|-- Need a specific embedding model not hosted by Pinecone? -> External
| (Generate embeddings yourself, upsert raw vectors)
|
|-- Need hybrid search with sparse vectors? -> External sparse model
| (Use pinecone-sparse-english-v0 via inference API + your dense model)
|
'-- Need full control over embedding pipeline? -> External
(Custom preprocessing, chunking, model selection)
</decision_framework>
<red_flags>
RED FLAGS
High Priority Issues:
- Targeting index by name instead of host --
pc.index("name")is deprecated in v7; usepc.index({ host })to avoid an extra API call - Upserting vectors with wrong dimensions -- dimension mismatch causes a 400 error; verify your embedding model's output dimension matches the index
- Nested metadata objects -- Pinecone only supports flat key-value metadata; nested objects are silently ignored or rejected
- Using metadata filtering for multi-tenancy at scale -- scans the entire namespace regardless of filter selectivity; use namespaces instead
Medium Priority Issues:
- Missing
includeMetadata: truein queries -- metadata is NOT included by default; omitting this returns only IDs and scores - Upsert batches exceeding 1,000 records or 2 MB -- triggers a 400 error; chunk at 200 records for safety margin
- Using
Dateobjects in metadata -- Pinecone metadata supports strings, numbers, booleans, and string arrays only; convert dates to Unix timestamps - Not awaiting
createIndex()readiness -- index creation is async; the index is not ready for operations immediately aftercreateIndex()returns
Common Mistakes:
- Querying immediately after upsert and expecting results -- eventual consistency means freshly upserted vectors may not be queryable for seconds
- Using
topK> 1,000 withincludeMetadata: true-- the maxtopKis 1,000 when including metadata or values; without them, max is 10,000 - Passing array values as metadata filter values (
{ tags: ["a", "b"] }) -- use$inoperator instead:{ tags: { $in: ["a", "b"] } } - Forgetting that
deleteAll()without a namespace deletes from the default namespace only, not the entire index
Gotchas & Edge Cases:
describeIndexStats()returns approximate counts -- record counts are not exact in real-time, especially after recent upserts or deletes- Metadata values are always returned as their original types, but filter comparisons are type-strict --
$eq: "42"does not match numeric42 - Sparse vector indices must be positive 32-bit integers (uint32), and values must be non-zero floats -- zero values are silently dropped
listPaginated()returns vector IDs only (no values or metadata) -- usefetch()to get full vector data- The
$inand$ninoperators accept a maximum of 10,000 values each - Metadata keys cannot start with
$(reserved for operators) upsertis an upsert, not an insert -- upserting with an existing ID overwrites the previous vector and metadata entirely (no partial merge)update()merges metadata by default -- updating metadata replaces only the fields you specify, not the entire metadata object- The v7 SDK includes built-in automatic retry with exponential backoff for transient errors -- custom retry logic is only needed for fine-grained control or non-default retry policies
upsertRecords(integrated inference) accepts a direct array, not{ records: [...] }-- this differs from the regularupsertmethod which uses{ records: [...] }
</red_flags>
<critical_reminders>
CRITICAL REMINDERS
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST target indexes by host URL, not by name -- pc.index({ host }) is the v7 API; pc.index('name') is deprecated)
(You MUST batch upserts to max 1,000 records or 2 MB per request -- exceeding either limit causes a 400 error)
(You MUST use flat key-value metadata only -- nested objects, null values, and keys starting with $ are rejected by Pinecone)
(You MUST handle eventual consistency -- vectors are not queryable immediately after upsert; use describeIndexStats() or retry logic for freshness-critical flows)
Failure to follow these rules will cause index creation failures, rejected upserts, empty query results, and degraded multi-tenant performance.
</critical_reminders>
Files (skills)
-
examples
-
batch-operations.md 6.5 KB
# Pinecone -- Batch Operations Examples > Chunked upserts, parallel ingestion, and bulk data management. See [core.md](core.md) for single-record operations. **Related examples:** - [core.md](core.md) -- Basic upsert, query, fetch, delete - [inference.md](inference.md) -- Batch embedding generation --- ## Shared Utility: Chunk Array All batch patterns below use this chunking utility. ```typescript function chunkArray<T>(array: T[], size: number): T[][] { const chunks: T[][] = []; for (let i = 0; i < array.length; i += size) { chunks.push(array.slice(i, i + size)); } return chunks; } export { chunkArray }; ``` --- ## Chunked Sequential Upsert Batch upserts at 200 records to stay well within the 1,000-record / 2 MB limit. ```typescript import type { Pinecone, RecordMetadata } from "@pinecone-database/pinecone"; const UPSERT_BATCH_SIZE = 200; interface VectorRecord<T extends RecordMetadata = RecordMetadata> { id: string; values: number[]; metadata?: T; } async function batchUpsert<T extends RecordMetadata>( index: ReturnType<Pinecone["index"]>, namespace: string, records: VectorRecord<T>[], ): Promise<void> { const batches = chunkArray(records, UPSERT_BATCH_SIZE); const ns = index.namespace(namespace); for (const batch of batches) { await ns.upsert({ records: batch }); } } export { batchUpsert }; ``` **Why good:** Named constant for batch size, generic metadata type, sequential processing avoids overwhelming the API --- ## Parallel Upsert with Concurrency Control For large datasets, parallel upserts increase throughput while respecting rate limits. ```typescript import type { Pinecone, RecordMetadata } from "@pinecone-database/pinecone"; const UPSERT_BATCH_SIZE = 200; const MAX_CONCURRENT_UPSERTS = 5; async function parallelBatchUpsert<T extends RecordMetadata>( index: ReturnType<Pinecone["index"]>, namespace: string, records: Array<{ id: string; values: number[]; metadata?: T }>, ): Promise<void> { const batches = chunkArray(records, UPSERT_BATCH_SIZE); const ns = index.namespace(namespace); // Process batches with bounded concurrency for (let i = 0; i < batches.length; i += MAX_CONCURRENT_UPSERTS) { const concurrentBatches = batches.slice(i, i + MAX_CONCURRENT_UPSERTS); await Promise.all( concurrentBatches.map((batch) => ns.upsert({ records: batch })), ); } } export { parallelBatchUpsert }; ``` **Why good:** Bounded concurrency prevents rate limiting, `Promise.all` for parallel execution within each window, sequential windows prevent overload --- ## Batch Upsert with Progress Tracking ```typescript import type { Pinecone, RecordMetadata } from "@pinecone-database/pinecone"; const UPSERT_BATCH_SIZE = 200; async function batchUpsertWithProgress<T extends RecordMetadata>( index: ReturnType<Pinecone["index"]>, namespace: string, records: Array<{ id: string; values: number[]; metadata?: T }>, onProgress?: (completed: number, total: number) => void, ): Promise<void> { const batches = chunkArray(records, UPSERT_BATCH_SIZE); const ns = index.namespace(namespace); let completed = 0; for (const batch of batches) { await ns.upsert({ records: batch }); completed += batch.length; onProgress?.(completed, records.length); } } export { batchUpsertWithProgress }; ``` **Why good:** Progress callback for long-running ingestion jobs, accurate count tracking --- ## Batch Upsert with Retry Logic ```typescript import type { Pinecone, RecordMetadata } from "@pinecone-database/pinecone"; const UPSERT_BATCH_SIZE = 200; const MAX_RETRIES = 3; const INITIAL_BACKOFF_MS = 1000; async function upsertWithRetry<T extends RecordMetadata>( ns: ReturnType<Pinecone["index"]>, records: Array<{ id: string; values: number[]; metadata?: T }>, ): Promise<void> { for (let attempt = 0; attempt < MAX_RETRIES; attempt++) { try { await ns.upsert({ records }); return; } catch (error) { const isRetryable = error instanceof Error && (error.message.includes("429") || error.message.includes("503")); if (!isRetryable || attempt === MAX_RETRIES - 1) { throw error; } const backoffMs = INITIAL_BACKOFF_MS * Math.pow(2, attempt); await new Promise((resolve) => setTimeout(resolve, backoffMs)); } } } async function batchUpsertWithRetry<T extends RecordMetadata>( index: ReturnType<Pinecone["index"]>, namespace: string, records: Array<{ id: string; values: number[]; metadata?: T }>, ): Promise<void> { const batches = chunkArray(records, UPSERT_BATCH_SIZE); const ns = index.namespace(namespace); for (const batch of batches) { await upsertWithRetry(ns, batch); } } export { batchUpsertWithRetry }; ``` **Why good:** Exponential backoff for rate limits (429) and server errors (503), named constants, non-retryable errors propagate immediately --- ## Verify Upsert Completion ```typescript import type { Pinecone } from "@pinecone-database/pinecone"; const POLL_INTERVAL_MS = 2000; const MAX_POLL_ATTEMPTS = 30; async function waitForUpsertCompletion( index: ReturnType<Pinecone["index"]>, expectedCount: number, namespace?: string, ): Promise<boolean> { for (let attempt = 0; attempt < MAX_POLL_ATTEMPTS; attempt++) { const stats = await index.describeIndexStats(); const currentCount = namespace ? (stats.namespaces?.[namespace]?.recordCount ?? 0) : (stats.totalRecordCount ?? 0); if (currentCount >= expectedCount) { return true; } await new Promise((resolve) => setTimeout(resolve, POLL_INTERVAL_MS)); } return false; // Timed out waiting for vectors to be indexed } export { waitForUpsertCompletion }; ``` **Why good:** Handles eventual consistency by polling, namespace-aware count check, timeout prevents infinite waiting **Gotcha:** `describeIndexStats()` counts are approximate. For exact verification, query for a specific recently-upserted ID using `fetch`. --- ## Batch Delete by IDs ```typescript import type { Pinecone } from "@pinecone-database/pinecone"; const DELETE_BATCH_SIZE = 1000; // Max IDs per deleteMany call async function batchDelete( index: ReturnType<Pinecone["index"]>, namespace: string, ids: string[], ): Promise<void> { const ns = index.namespace(namespace); const batches = chunkArray(ids, DELETE_BATCH_SIZE); for (const batch of batches) { await ns.deleteMany({ ids: batch }); } } export { batchDelete }; ``` **Why good:** Respects 1,000 ID limit per delete request, sequential batching --- _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_ -
core.md 9.9 KB
# Pinecone -- Core Pattern Examples > Client setup, index management, and fundamental vector operations (upsert, query, fetch, update, delete). Reference from [SKILL.md](../SKILL.md). **Related examples:** - [namespaces.md](namespaces.md) -- Namespace isolation, multi-tenant patterns - [metadata-filtering.md](metadata-filtering.md) -- Filter operators and compound filters - [hybrid-search.md](hybrid-search.md) -- Sparse-dense hybrid search - [inference.md](inference.md) -- Embedding generation and reranking - [batch-operations.md](batch-operations.md) -- Chunked upserts, parallel ingestion --- ## Client Initialization ```typescript import { Pinecone } from "@pinecone-database/pinecone"; function createPineconeClient(): Pinecone { const apiKey = process.env.PINECONE_API_KEY; if (!apiKey) { throw new Error("PINECONE_API_KEY environment variable is required"); } return new Pinecone({ apiKey }); } export { createPineconeClient }; ``` **Why good:** API key from environment variable (never hardcoded), explicit validation, named export The `Pinecone` constructor also reads `PINECONE_API_KEY` from the environment automatically if no `apiKey` is passed. Explicit passing is preferred for clarity and to fail fast with a clear error. --- ## Create a Serverless Index ```typescript import { Pinecone } from "@pinecone-database/pinecone"; const EMBEDDING_DIMENSION = 1536; // Must match your embedding model const INDEX_NAME = "documents"; async function createServerlessIndex(pc: Pinecone): Promise<string> { const indexModel = await pc.createIndex({ name: INDEX_NAME, dimension: EMBEDDING_DIMENSION, metric: "cosine", spec: { serverless: { cloud: "aws", region: "us-east-1", }, }, waitUntilReady: true, // Block until index is ready for operations }); return indexModel.host; // Save this -- use it to target the index } export { createServerlessIndex }; ``` **Why good:** Named constants for dimension and name, `waitUntilReady: true` prevents premature operations, returns host URL for index targeting ```typescript // ❌ Bad Example -- wrong dimension, no wait const indexModel = await pc.createIndex({ name: "docs", dimension: 768, // Mismatch if using a 1536-dim model metric: "cosine", spec: { serverless: { cloud: "aws", region: "us-east-1" } }, }); // Index may not be ready yet -- operations will fail const index = pc.index({ host: indexModel.host }); await index.upsert({ records: [...] }); // May throw ``` **Why bad:** Dimension mismatch causes 400 errors on upsert, no `waitUntilReady` means index may not accept operations immediately --- ## Target an Index by Host ```typescript import { Pinecone } from "@pinecone-database/pinecone"; // v7 API -- target by host (preferred) const INDEX_HOST = process.env.PINECONE_INDEX_HOST; function getIndex(pc: Pinecone): ReturnType<Pinecone["index"]> { if (!INDEX_HOST) { throw new Error("PINECONE_INDEX_HOST environment variable is required"); } return pc.index({ host: INDEX_HOST }); } export { getIndex }; ``` **Why good:** Host URL from environment, avoids extra API call to resolve name to host If you don't have the host URL stored, use `describeIndex` to retrieve it: ```typescript const indexModel = await pc.describeIndex(INDEX_NAME); const index = pc.index({ host: indexModel.host }); ``` --- ## Upsert Vectors with Typed Metadata ```typescript import type { Pinecone, RecordMetadata } from "@pinecone-database/pinecone"; interface ArticleMetadata extends RecordMetadata { title: string; category: string; publishedAt: number; // Unix timestamp tags: string[]; // Only string arrays are supported } const NAMESPACE = "articles"; async function upsertArticle( index: ReturnType<Pinecone["index"]>, id: string, embedding: number[], metadata: ArticleMetadata, ): Promise<void> { await index.namespace(NAMESPACE).upsert({ records: [{ id, values: embedding, metadata }], }); } export { upsertArticle }; export type { ArticleMetadata }; ``` **Why good:** Type-safe metadata via `RecordMetadata` extension, namespace isolation, Unix timestamp for dates (not Date objects), string array for tags ```typescript // ❌ Bad Example -- invalid metadata await index.upsert({ records: [ { id: "doc-1", values: embedding, metadata: { title: "Guide", author: { name: "Alice", role: "admin" }, // INVALID: nested object createdAt: new Date(), // INVALID: Date object $priority: "high", // INVALID: key starts with $ score: null, // INVALID: null value }, }, ], }); ``` **Why bad:** Nested objects are not supported, Date objects are not serializable, `$`-prefixed keys are reserved for operators, null values are rejected --- ## Query by Vector Similarity ```typescript import type { Pinecone } from "@pinecone-database/pinecone"; import type { ArticleMetadata } from "./upsert-article"; const TOP_K = 10; const NAMESPACE = "articles"; interface SearchResult { id: string; score: number; metadata: ArticleMetadata | undefined; } async function searchArticles( index: ReturnType<Pinecone["index"]>, queryEmbedding: number[], category?: string, ): Promise<SearchResult[]> { const filter = category ? { category: { $eq: category } } : undefined; const response = await index.namespace(NAMESPACE).query({ vector: queryEmbedding, topK: TOP_K, includeMetadata: true, filter, }); return response.matches.map((match) => ({ id: match.id, score: match.score ?? 0, metadata: match.metadata as ArticleMetadata | undefined, })); } export { searchArticles }; ``` **Why good:** Optional filter, `includeMetadata: true` to get metadata back, typed response mapping, named constant for topK --- ## Fetch Vectors by ID ```typescript const NAMESPACE = "articles"; async function fetchArticles( index: ReturnType<Pinecone["index"]>, ids: string[], ): Promise<void> { // Max 1,000 IDs per fetch const response = await index.namespace(NAMESPACE).fetch({ ids }); for (const [id, record] of Object.entries(response.records)) { console.log(id, record.metadata); // record.values contains the vector (if stored) } } export { fetchArticles }; ``` **Why good:** Targets namespace, respects 1,000 ID limit, iterates response records correctly **Gotcha:** `fetch` returns an object keyed by ID, not an array. Missing IDs are silently omitted from the response (no error thrown). --- ## Update Vector Metadata ```typescript const NAMESPACE = "articles"; async function updateArticleCategory( index: ReturnType<Pinecone["index"]>, id: string, newCategory: string, ): Promise<void> { await index.namespace(NAMESPACE).update({ id, metadata: { category: newCategory }, // Only specified fields are updated -- other metadata fields are preserved }); } export { updateArticleCategory }; ``` **Why good:** Partial metadata update (only `category` changes, other fields preserved), targets namespace **Gotcha:** `update()` merges metadata fields, it does not replace the entire metadata object. To remove a metadata field, you must re-upsert the entire record. --- ## Delete Vectors ```typescript const NAMESPACE = "articles"; // Delete by IDs async function deleteArticles( index: ReturnType<Pinecone["index"]>, ids: string[], ): Promise<void> { await index.namespace(NAMESPACE).deleteMany({ ids }); } // Delete by metadata filter async function deleteByCategory( index: ReturnType<Pinecone["index"]>, category: string, ): Promise<void> { await index.namespace(NAMESPACE).deleteMany({ filter: { category: { $eq: category } }, }); } // Delete all vectors in a namespace async function clearNamespace( index: ReturnType<Pinecone["index"]>, ): Promise<void> { await index.namespace(NAMESPACE).deleteAll(); } export { deleteArticles, deleteByCategory, clearNamespace }; ``` **Why good:** Three deletion patterns (by ID, by filter, all), targets namespace, named exports **Gotcha:** `deleteAll()` without a namespace targets the default (empty) namespace. To delete everything across all namespaces, you must delete each namespace individually or delete and recreate the index. --- ## Check Index Statistics ```typescript async function getIndexStats( index: ReturnType<Pinecone["index"]>, ): Promise<void> { const stats = await index.describeIndexStats(); console.log("Total vectors:", stats.totalRecordCount); console.log("Index fullness:", stats.indexFullness); console.log("Dimension:", stats.dimension); // Per-namespace breakdown if (stats.namespaces) { for (const [ns, nsStats] of Object.entries(stats.namespaces)) { console.log(` Namespace "${ns}": ${nsStats.recordCount} vectors`); } } } export { getIndexStats }; ``` **Why good:** Shows total and per-namespace counts, useful for monitoring and verifying upsert completion **Gotcha:** Record counts from `describeIndexStats()` are approximate and may lag behind recent upserts by several seconds. --- ## List Vector IDs with Pagination ```typescript const PAGE_SIZE = 100; const NAMESPACE = "articles"; async function listAllVectorIds( index: ReturnType<Pinecone["index"]>, ): Promise<string[]> { const allIds: string[] = []; let paginationToken: string | undefined; do { const response = await index.namespace(NAMESPACE).listPaginated({ limit: PAGE_SIZE, paginationToken, prefix: "doc-", // Optional: filter by ID prefix }); if (response.vectors) { allIds.push(...response.vectors.map((v) => v.id)); } paginationToken = response.pagination?.next; } while (paginationToken); return allIds; } export { listAllVectorIds }; ``` **Why good:** Handles pagination correctly, optional prefix filter, collects all IDs across pages **Gotcha:** `listPaginated` returns only vector IDs, not values or metadata. Use `fetch` to get full vector data for specific IDs. --- _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_ -
hybrid-search.md 6 KB
# Pinecone -- Hybrid Search Examples > Sparse-dense hybrid search combining semantic and keyword retrieval. See [core.md](core.md) for basic query patterns and [inference.md](inference.md) for embedding generation. **Related examples:** - [core.md](core.md) -- Basic vector operations - [inference.md](inference.md) -- Sparse embedding generation with `pinecone-sparse-english-v0` --- ## Create a Hybrid Index Hybrid search requires `dotproduct` metric -- the only metric that supports combined sparse+dense queries. ```typescript import { Pinecone } from "@pinecone-database/pinecone"; const DENSE_DIMENSION = 1536; const HYBRID_INDEX_NAME = "hybrid-search"; async function createHybridIndex(pc: Pinecone): Promise<string> { const indexModel = await pc.createIndex({ name: HYBRID_INDEX_NAME, dimension: DENSE_DIMENSION, metric: "dotproduct", // REQUIRED for hybrid search spec: { serverless: { cloud: "aws", region: "us-east-1", }, }, waitUntilReady: true, }); return indexModel.host; } export { createHybridIndex }; ``` **Why good:** Uses `dotproduct` metric (required for sparse-dense), `waitUntilReady` prevents premature operations ```typescript // ❌ Bad Example -- wrong metric for hybrid await pc.createIndex({ name: "hybrid", dimension: 1536, metric: "cosine", // Does NOT support sparse-dense queries spec: { serverless: { cloud: "aws", region: "us-east-1" } }, }); ``` **Why bad:** `cosine` and `euclidean` metrics do not support sparse vector queries; only `dotproduct` works for hybrid search --- ## Upsert with Sparse + Dense Vectors Each record can have both `values` (dense) and `sparseValues` (sparse). ```typescript import type { Pinecone } from "@pinecone-database/pinecone"; const NAMESPACE = "documents"; interface SparseVector { indices: number[]; // uint32 positive integers values: number[]; // Non-zero floats } async function upsertHybridRecord( index: ReturnType<Pinecone["index"]>, id: string, denseEmbedding: number[], sparseEmbedding: SparseVector, metadata: Record<string, string | number>, ): Promise<void> { await index.namespace(NAMESPACE).upsert({ records: [ { id, values: denseEmbedding, sparseValues: sparseEmbedding, metadata, }, ], }); } export { upsertHybridRecord }; export type { SparseVector }; ``` **Why good:** Both dense and sparse vectors on the same record, typed sparse vector interface, metadata included **Gotcha:** Sparse vector `indices` must be positive 32-bit integers and `values` must be non-zero floats. Zero values are silently dropped. --- ## Hybrid Query with Alpha Weighting Control the balance between semantic (dense) and keyword (sparse) relevance using alpha weighting. ```typescript import type { Pinecone } from "@pinecone-database/pinecone"; import type { SparseVector } from "./upsert-hybrid"; const TOP_K = 10; const NAMESPACE = "documents"; function scaleVector(vector: number[], alpha: number): number[] { return vector.map((v) => v * alpha); } function scaleSparseVector(sparse: SparseVector, alpha: number): SparseVector { return { indices: sparse.indices, values: sparse.values.map((v) => v * alpha), }; } async function hybridSearch( index: ReturnType<Pinecone["index"]>, denseQuery: number[], sparseQuery: SparseVector, alpha: number = 0.7, // 0 = pure keyword, 1 = pure semantic ): Promise<void> { const results = await index.namespace(NAMESPACE).query({ vector: scaleVector(denseQuery, alpha), sparseVector: scaleSparseVector(sparseQuery, 1 - alpha), topK: TOP_K, includeMetadata: true, }); for (const match of results.matches) { console.log(match.id, match.score, match.metadata); } } export { hybridSearch }; ``` **Why good:** Alpha weighting controls semantic vs keyword balance, scaling applied before query, pure semantic (alpha=1) or pure keyword (alpha=0) modes available **Note:** Alpha tuning is domain-specific. Start at 0.7 (70% semantic, 30% keyword) and adjust based on relevance evaluation. --- ## Generate Sparse Embeddings via Inference API Use Pinecone's hosted sparse model for keyword-style embeddings. ```typescript import { Pinecone } from "@pinecone-database/pinecone"; import type { SparseVector } from "./upsert-hybrid"; async function generateSparseEmbedding( pc: Pinecone, text: string, inputType: "passage" | "query", ): Promise<SparseVector> { const result = await pc.inference.embed({ model: "pinecone-sparse-english-v0", inputs: [{ text }], parameters: { inputType, truncate: "END" }, }); const embedding = result.data[0]; if (!embedding.sparseValues) { throw new Error("Sparse model did not return sparse values"); } return { indices: embedding.sparseValues.indices, values: embedding.sparseValues.values, }; } export { generateSparseEmbedding }; ``` **Why good:** Uses `inputType` to distinguish passage indexing from query time, `truncate: "END"` handles long inputs **Important:** Use `inputType: "passage"` when generating embeddings for upsert and `inputType: "query"` when generating embeddings for search. Using the wrong type degrades relevance. --- ## Limitations of Single Hybrid Index ``` Hybrid index (dotproduct metric): ✅ Combined sparse + dense queries ✅ Alpha weighting for relevance tuning ❌ Cannot do sparse-only queries ❌ Cannot use integrated embedding (createIndexForModel) ❌ Cannot use searchRecords or upsertRecords convenience methods Separate dense index (cosine metric): ✅ Integrated inference (automatic embedding) ✅ searchRecords / upsertRecords ❌ No sparse vector support ❌ No hybrid queries ``` **When to choose hybrid:** You need combined semantic + keyword search and are willing to manage your own embedding pipeline. **When to skip hybrid:** Pure semantic search is sufficient, and you want the simplest architecture with integrated inference. --- _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_ -
inference.md 7.2 KB
# Pinecone -- Inference API Examples > Embedding generation, reranking, and integrated inference indexes. See [core.md](core.md) for basic operations and [hybrid-search.md](hybrid-search.md) for sparse embeddings. **Related examples:** - [core.md](core.md) -- Client setup, upsert, query - [hybrid-search.md](hybrid-search.md) -- Sparse embeddings for hybrid search - [batch-operations.md](batch-operations.md) -- Batch embedding generation --- ## Generate Dense Embeddings ```typescript import { Pinecone } from "@pinecone-database/pinecone"; async function embedTexts( pc: Pinecone, texts: string[], inputType: "passage" | "query", ): Promise<number[][]> { const result = await pc.inference.embed({ model: "multilingual-e5-large", inputs: texts.map((text) => ({ text })), parameters: { inputType, // "passage" for indexing, "query" for search truncate: "END", // Truncate if input exceeds model's context window }, }); return result.data.map((item) => { if (!item.values) { throw new Error("Embedding model did not return dense values"); } return item.values; }); } export { embedTexts }; ``` **Why good:** `inputType` distinguishes indexing from querying (critical for asymmetric models), `truncate` handles long inputs gracefully, validates return values **Important:** Always use `inputType: "passage"` when generating embeddings for documents being indexed, and `inputType: "query"` for search queries. Mixing these up degrades retrieval quality with asymmetric models. --- ## Rerank Query Results Reranking improves relevance by scoring query-document pairs with a cross-encoder model, which is more accurate than vector similarity alone. ```typescript import type { Pinecone } from "@pinecone-database/pinecone"; const RERANK_TOP_N = 5; interface RankedResult { id: string; text: string; relevanceScore: number; } async function rerankResults( pc: Pinecone, query: string, documents: Array<{ id: string; text: string }>, ): Promise<RankedResult[]> { const result = await pc.inference.rerank({ model: "pinecone-rerank-v0", query, documents, topN: RERANK_TOP_N, returnDocuments: true, rankFields: ["text"], }); return result.data.map((item) => ({ id: item.document?.id ?? "", text: item.document?.text ?? "", relevanceScore: item.score, })); } export { rerankResults }; ``` **Why good:** `topN` limits output, `returnDocuments: true` includes document content in response, `rankFields` specifies which fields to rank on, typed return value **Pattern:** First query Pinecone for top-K candidates (e.g., 50), then rerank to get the most relevant top-N (e.g., 5). This two-stage approach balances recall and precision. --- ## Two-Stage Retrieval: Query + Rerank ```typescript import type { Pinecone } from "@pinecone-database/pinecone"; const INITIAL_TOP_K = 50; // Broad retrieval const RERANK_TOP_N = 5; // Precise reranking async function searchAndRerank( pc: Pinecone, index: ReturnType<Pinecone["index"]>, queryEmbedding: number[], queryText: string, namespace: string, ) { // Stage 1: Broad vector search const candidates = await index.namespace(namespace).query({ vector: queryEmbedding, topK: INITIAL_TOP_K, includeMetadata: true, }); if (candidates.matches.length === 0) { return []; } // Stage 2: Rerank with cross-encoder const documents = candidates.matches.map((m) => ({ id: m.id, text: (m.metadata?.content as string) ?? "", })); const reranked = await pc.inference.rerank({ model: "pinecone-rerank-v0", query: queryText, documents, topN: RERANK_TOP_N, returnDocuments: true, rankFields: ["text"], }); return reranked.data.map((item) => ({ id: item.document?.id ?? "", text: item.document?.text ?? "", score: item.score, })); } export { searchAndRerank }; ``` **Why good:** Two-stage retrieval (broad recall then precise reranking), named constants for K and N, handles empty results --- ## Integrated Inference Index Create an index that automatically generates embeddings on upsert and query -- no external embedding pipeline needed. ```typescript import { Pinecone } from "@pinecone-database/pinecone"; const INDEX_NAME = "auto-embed"; const TEXT_FIELD = "chunk_text"; async function createIntegratedIndex(pc: Pinecone): Promise<string> { const indexModel = await pc.createIndexForModel({ name: INDEX_NAME, cloud: "aws", region: "us-east-1", embed: { model: "multilingual-e5-large", fieldMap: { text: TEXT_FIELD }, // Which metadata field to embed }, }); return indexModel.host; } export { createIntegratedIndex }; ``` **Why good:** Pinecone handles embedding automatically, `fieldMap` maps the text source field, no dimension/metric config needed (derived from model) --- ## Upsert and Search with Integrated Inference With an integrated inference index, you upsert text records and search with text queries -- no vectors involved. ```typescript import type { Pinecone } from "@pinecone-database/pinecone"; const NAMESPACE = "articles"; const TOP_K = 10; // Upsert raw text -- Pinecone generates embeddings automatically async function upsertTextRecords( index: ReturnType<Pinecone["index"]>, records: Array<{ id: string; chunkText: string; category: string }>, ): Promise<void> { // upsertRecords takes a direct array, not { records: [...] } // _id is the canonical field; id also works as an alias await index.namespace(NAMESPACE).upsertRecords( records.map((r) => ({ _id: r.id, chunk_text: r.chunkText, // Must match fieldMap.text from index creation category: r.category, })), ); } // Search with text query -- Pinecone embeds the query automatically async function searchByText( index: ReturnType<Pinecone["index"]>, queryText: string, ): Promise<void> { const results = await index.namespace(NAMESPACE).searchRecords({ query: { topK: TOP_K, inputs: { text: queryText } }, }); for (const hit of results.result.hits) { console.log(hit._id, hit._score, hit.fields); } } export { upsertTextRecords, searchByText }; ``` **Why good:** No vector generation code needed, `_id` identifies records (`id` also works as an alias), `upsertRecords` takes a direct array (not `{ records: [...] }`), search uses text input directly **Gotcha:** Integrated inference indexes use different method names (`upsertRecords`/`searchRecords` instead of `upsert`/`query`), different record shapes (`_id` with metadata fields as top-level properties), and `upsertRecords` accepts a max of 96 records per call (not 1,000 like `upsert`). --- ## List Available Models ```typescript import { Pinecone } from "@pinecone-database/pinecone"; async function listEmbeddingModels(pc: Pinecone): Promise<void> { const models = await pc.inference.listModels({ type: "embed", // Filter to embedding models only }); for (const model of models.models ?? []) { console.log( `${model.model}: ${model.vectorType}, dim=${model.defaultDimension}`, ); } } export { listEmbeddingModels }; ``` **Why good:** Filtered by model type, shows key properties (vector type, dimension) --- _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_ -
metadata-filtering.md 4.9 KB
# Pinecone -- Metadata Filtering Examples > Filter operators, compound filters, and best practices. See [core.md](core.md) for basic query patterns and [reference.md](../reference.md) for the operator table. **Related examples:** - [core.md](core.md) -- Client setup, basic query with filters - [namespaces.md](namespaces.md) -- Alternative to filtering for multi-tenancy --- ## Comparison Operators ```typescript const TOP_K = 10; // Exact match const drama = await index.namespace(ns).query({ vector: embedding, topK: TOP_K, includeMetadata: true, filter: { genre: { $eq: "drama" } }, }); // Not equal const notComedy = await index.namespace(ns).query({ vector: embedding, topK: TOP_K, includeMetadata: true, filter: { genre: { $ne: "comedy" } }, }); // Numeric range const recent = await index.namespace(ns).query({ vector: embedding, topK: TOP_K, includeMetadata: true, filter: { year: { $gte: 2020 } }, }); // Set membership const selected = await index.namespace(ns).query({ vector: embedding, topK: TOP_K, includeMetadata: true, filter: { genre: { $in: ["drama", "action", "thriller"] } }, }); ``` **Why good:** Each operator targets specific types (see reference.md for full operator table), `$in` for set membership avoids multiple `$or` clauses --- ## Compound Filters with $and / $or Only `$and` and `$or` are allowed at the top level of a filter expression. ```typescript const TOP_K = 10; // AND: all conditions must match const filteredResults = await index.namespace(ns).query({ vector: embedding, topK: TOP_K, includeMetadata: true, filter: { $and: [ { genre: { $eq: "drama" } }, { year: { $gte: 2020 } }, { rating: { $gt: 7.5 } }, ], }, }); // OR: any condition matches const broadResults = await index.namespace(ns).query({ vector: embedding, topK: TOP_K, includeMetadata: true, filter: { $or: [{ genre: { $eq: "drama" } }, { genre: { $eq: "documentary" } }], }, }); // Combined AND + OR const complexResults = await index.namespace(ns).query({ vector: embedding, topK: TOP_K, includeMetadata: true, filter: { $and: [ { year: { $gte: 2020 } }, { $or: [{ genre: { $eq: "drama" } }, { genre: { $eq: "thriller" } }], }, ], }, }); ``` **Why good:** Nested logical operators for complex queries, `$or` inside `$and` for flexible filtering --- ## Field Existence Check ```typescript const TOP_K = 10; // Only records that have a "summary" field const withSummary = await index.namespace(ns).query({ vector: embedding, topK: TOP_K, includeMetadata: true, filter: { summary: { $exists: true } }, }); // Records missing a "category" field const uncategorized = await index.namespace(ns).query({ vector: embedding, topK: TOP_K, includeMetadata: true, filter: { category: { $exists: false } }, }); ``` **Why good:** `$exists` checks for field presence without requiring a specific value, useful for incremental data enrichment workflows --- ## Delete by Metadata Filter ```typescript // Delete all records matching a filter await index.namespace(ns).deleteMany({ filter: { $and: [{ status: { $eq: "archived" } }, { createdAt: { $lt: 1700000000 } }], }, }); ``` **Why good:** Bulk deletion without knowing vector IDs, combines metadata filter with namespace targeting --- ## Fetch by Metadata Filter ```typescript // Fetch records by metadata (no vector required) const response = await index.namespace(ns).fetchByMetadata({ filter: { category: { $eq: "tutorial" } }, limit: 50, }); for (const record of response.records ?? []) { console.log(record.id, record.metadata); } ``` **Why good:** Retrieves records by metadata without needing vector IDs or a query vector, useful for data management tasks --- ## Common Filter Mistakes ```typescript // ❌ Bad: Array as filter value (not valid syntax) filter: { genre: ["drama", "action"] } // Fix: use $in operator filter: { genre: { $in: ["drama", "action"] } } // ❌ Bad: Nested object in metadata metadata: { author: { name: "Alice", org: "Acme" } } // Fix: flatten to top-level keys metadata: { authorName: "Alice", authorOrg: "Acme" } // ❌ Bad: $eq with array value filter: { genre: { $eq: ["drama", "action"] } } // Fix: use $in for set membership filter: { genre: { $in: ["drama", "action"] } } // ❌ Bad: Type mismatch in comparison // Metadata: { year: 2024 } (number) filter: { year: { $eq: "2024" } } // String "2024" does NOT match number 2024 // Fix: use matching type filter: { year: { $eq: 2024 } } // ❌ Bad: Null value in metadata metadata: { category: null } // Fix: omit the field entirely, or use a sentinel value metadata: { category: "uncategorized" } ``` **Why bad:** Each mistake causes either a rejection or silently empty results; type-strict comparisons are a common source of "no results" bugs --- _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_ -
namespaces.md 5.3 KB
# Pinecone -- Namespaces & Multi-Tenancy Examples > Namespace isolation patterns for multi-tenant applications. See [core.md](core.md) for basic vector operations. **Related examples:** - [core.md](core.md) -- Client setup, upsert, query, fetch, update, delete - [metadata-filtering.md](metadata-filtering.md) -- Filter operators (alternative to namespace isolation) --- ## Namespace-Based Tenant Isolation Namespaces physically separate data within an index. Queries scan only the target namespace, making them cheaper and faster than metadata filtering at scale. ```typescript import type { Pinecone } from "@pinecone-database/pinecone"; const TOP_K = 10; function getTenantIndex(pc: Pinecone, host: string, tenantId: string) { return pc.index({ host }).namespace(`tenant-${tenantId}`); } async function searchTenant( pc: Pinecone, host: string, tenantId: string, queryEmbedding: number[], ) { const ns = getTenantIndex(pc, host, tenantId); const results = await ns.query({ vector: queryEmbedding, topK: TOP_K, includeMetadata: true, }); return results.matches; } export { getTenantIndex, searchTenant }; ``` **Why good:** Physical isolation per tenant, query cost proportional to tenant data size (not total index size), simple namespace naming convention --- ## Namespace Management API The v7 SDK provides explicit namespace management methods. ```typescript import type { Pinecone } from "@pinecone-database/pinecone"; const INDEX_HOST = process.env.PINECONE_INDEX_HOST!; async function setupTenantNamespace( pc: Pinecone, tenantId: string, ): Promise<void> { const index = pc.index({ host: INDEX_HOST }); const namespaceName = `tenant-${tenantId}`; // Create namespace (v7 API -- explicit creation) await index.createNamespace({ name: namespaceName, // Optional: declare metadata schema for this namespace }); } async function removeTenantNamespace( pc: Pinecone, tenantId: string, ): Promise<void> { const index = pc.index({ host: INDEX_HOST }); // Deletes namespace and ALL vectors within it await index.deleteNamespace(`tenant-${tenantId}`); } async function listTenantNamespaces(pc: Pinecone): Promise<string[]> { const index = pc.index({ host: INDEX_HOST }); const response = await index.listNamespaces({ prefix: "tenant-", // Filter by prefix }); return response.namespaces?.map((ns) => ns.name) ?? []; } export { setupTenantNamespace, removeTenantNamespace, listTenantNamespaces }; ``` **Why good:** Explicit namespace lifecycle management, prefix-based listing for tenant discovery, `deleteNamespace` cleanly removes tenant data --- ## Tenant Onboarding and Offboarding ```typescript import type { Pinecone } from "@pinecone-database/pinecone"; interface TenantData { id: string; documents: Array<{ id: string; embedding: number[]; metadata: Record<string, string | number>; }>; } const INDEX_HOST = process.env.PINECONE_INDEX_HOST!; const UPSERT_BATCH_SIZE = 200; async function onboardTenant(pc: Pinecone, tenant: TenantData): Promise<void> { const ns = pc.index({ host: INDEX_HOST }).namespace(`tenant-${tenant.id}`); // Batch upsert tenant documents for (let i = 0; i < tenant.documents.length; i += UPSERT_BATCH_SIZE) { const batch = tenant.documents.slice(i, i + UPSERT_BATCH_SIZE); await ns.upsert({ records: batch.map((doc) => ({ id: doc.id, values: doc.embedding, metadata: doc.metadata, })), }); } } async function offboardTenant(pc: Pinecone, tenantId: string): Promise<void> { const index = pc.index({ host: INDEX_HOST }); // Single call removes all tenant data await index.deleteNamespace(`tenant-${tenantId}`); } export { onboardTenant, offboardTenant }; ``` **Why good:** Batched upserts during onboarding, single-call offboarding via `deleteNamespace`, clean data isolation --- ## Per-Namespace Statistics ```typescript import type { Pinecone } from "@pinecone-database/pinecone"; const INDEX_HOST = process.env.PINECONE_INDEX_HOST!; async function getTenantStats( pc: Pinecone, tenantId: string, ): Promise<{ recordCount: number }> { const index = pc.index({ host: INDEX_HOST }); const stats = await index.describeIndexStats(); const nsName = `tenant-${tenantId}`; const nsStats = stats.namespaces?.[nsName]; if (!nsStats) { return { recordCount: 0 }; } return { recordCount: nsStats.recordCount ?? 0 }; } export { getTenantStats }; ``` **Why good:** Per-namespace vector count, handles missing namespace gracefully **Gotcha:** `describeIndexStats()` returns approximate counts. For exact counts after bulk operations, allow a few seconds for indexing to complete. --- ## Namespace vs Metadata Filtering -- Cost Comparison ``` Scenario: 100 tenants, 1 GB data each, querying one tenant Namespace approach: Query scans: 1 GB (tenant's namespace only) Cost: ~1 read unit per query Metadata filtering approach (single namespace): Query scans: 100 GB (entire namespace, then filters) Cost: ~100 read units per query At 1,000 queries/day: namespace approach is 100x cheaper. ``` **Rule of thumb:** Use namespaces for multi-tenancy when you have more than a handful of tenants or when total data exceeds a few GB. Use metadata filtering only when you need cross-tenant queries. --- _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_
-
-
reference.md 11.7 KB
# Pinecone Quick Reference > API reference, filter operators, limits, decision frameworks, and production checklist. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples. --- ## Pinecone Client Methods | Method | Description | Returns | | --------------------------------- | ------------------------------------------------------- | --------------------- | | `new Pinecone({ apiKey })` | Create client (reads `PINECONE_API_KEY` env if omitted) | `Pinecone` | | `pc.createIndex(options)` | Create serverless or pod-based index | `Promise<IndexModel>` | | `pc.createIndexForModel(options)` | Create index with integrated inference | `Promise<IndexModel>` | | `pc.describeIndex(name)` | Get index metadata and host | `Promise<IndexModel>` | | `pc.listIndexes()` | List all indexes | `Promise<IndexList>` | | `pc.configureIndex(options)` | Update index config (replicas, pod type) | `Promise<IndexModel>` | | `pc.deleteIndex(name)` | Delete an index | `Promise<void>` | | `pc.index({ host })` | Target a specific index by host URL (preferred) | `Index<T>` | | `pc.index(name, host)` | Target by name + host (two-arg shorthand) | `Index<T>` | ## Index Methods (Vector Operations) | Method | Description | Returns | | -------------------------------- | ------------------------------------------------------- | ------------------------------------- | | `index.namespace(name)` | Target a namespace within the index | `Index<T>` | | `index.upsert({ records })` | Insert or update vectors | `Promise<void>` | | `index.upsertRecords(records[])` | Upsert with integrated inference (direct array, max 96) | `Promise<void>` | | `index.query(options)` | Find similar vectors | `Promise<QueryResponse<T>>` | | `index.searchRecords(options)` | Search with integrated inference | `Promise<SearchRecordsResponse>` | | `index.fetch({ ids })` | Get vectors by ID | `Promise<FetchResponse<T>>` | | `index.fetchByMetadata(options)` | Fetch vectors by metadata filter | `Promise<FetchByMetadataResponse<T>>` | | `index.update(options)` | Update vector values or metadata | `Promise<void>` | | `index.deleteOne({ id })` | Delete a single vector | `Promise<void>` | | `index.deleteMany(options)` | Delete by IDs or metadata filter | `Promise<void>` | | `index.deleteAll()` | Delete all vectors in namespace | `Promise<void>` | | `index.listPaginated(options)` | List vector IDs with pagination | `Promise<ListResponse>` | | `index.describeIndexStats()` | Get index statistics | `Promise<IndexStatsDescription>` | ## Namespace Management Methods | Method | Description | Returns | | -------------------------------- | ------------------------------------------------ | --------------------------------- | | `index.listNamespaces(options?)` | List all namespaces | `Promise<ListNamespacesResponse>` | | `index.createNamespace(options)` | Create a namespace with optional metadata schema | `Promise<NamespaceDescription>` | | `index.describeNamespace(name)` | Get namespace details | `Promise<NamespaceDescription>` | | `index.deleteNamespace(name)` | Delete a namespace and all its vectors | `Promise<void>` | ## Inference Methods | Method | Description | Returns | | ------------------------------ | ----------------------------- | ------------------------- | | `pc.inference.embed(options)` | Generate embeddings from text | `Promise<EmbeddingsList>` | | `pc.inference.rerank(options)` | Rerank documents by relevance | `Promise<RerankResult>` | | `pc.inference.getModel(name)` | Get model info | `Promise<ModelInfo>` | | `pc.inference.listModels()` | List available models | `Promise<ModelInfoList>` | ## Collections Methods (Pod-Based Only) | Method | Description | Returns | | ------------------------------ | ------------------------------------- | -------------------------- | | `pc.createCollection(options)` | Create collection snapshot from index | `Promise<CollectionModel>` | | `pc.listCollections()` | List all collections | `Promise<CollectionList>` | | `pc.describeCollection(name)` | Get collection details | `Promise<CollectionModel>` | | `pc.deleteCollection(name)` | Delete a collection | `Promise<void>` | --- ## Metadata Filter Operators | Operator | Description | Supported Types | Example | | --------- | -------------------------------- | ----------------------- | ----------------------------------------- | | `$eq` | Equal to | string, number, boolean | `{ genre: { $eq: "drama" } }` | | `$ne` | Not equal to | string, number, boolean | `{ genre: { $ne: "comedy" } }` | | `$gt` | Greater than | number | `{ year: { $gt: 2020 } }` | | `$gte` | Greater than or equal | number | `{ year: { $gte: 2020 } }` | | `$lt` | Less than | number | `{ year: { $lt: 2020 } }` | | `$lte` | Less than or equal | number | `{ year: { $lte: 2020 } }` | | `$in` | In array (max 10,000 values) | string, number | `{ genre: { $in: ["drama", "action"] } }` | | `$nin` | Not in array (max 10,000 values) | string, number | `{ genre: { $nin: ["horror"] } }` | | `$exists` | Field exists or not | boolean | `{ genre: { $exists: true } }` | | `$and` | Logical AND (top-level) | filter[] | `{ $and: [filter1, filter2] }` | | `$or` | Logical OR (top-level) | filter[] | `{ $or: [filter1, filter2] }` | **Rules:** - Only `$and` and `$or` allowed at the query's top level - `$in` and `$nin` accept max 10,000 values each - Metadata keys cannot start with `$` - No nested objects; flat key-value pairs only - Supported types: string, number (int/float), boolean, string[] - Null values are not supported --- ## Limits Quick Reference | Resource | Limit | | ----------------------------------------------------------- | --------------------- | | Max vectors per upsert | 1,000 records or 2 MB | | Max metadata per record | 40 KB | | Max record ID length | 512 characters | | Max dense vector dimensions | 20,000 | | Max sparse non-zero values | 2,048 per vector | | Max `topK` (without metadata) | 10,000 | | Max `topK` (with `includeMetadata` or `includeValues`) | 1,000 | | Max vectors per fetch/delete | 1,000 IDs | | Max `$in` / `$nin` values | 10,000 each | | Max text records per `upsertRecords` (integrated inference) | 96 | | Serverless indexes per project | 100 | | Pod-based indexes per project | 20 | --- ## Supported Distance Metrics | Metric | Description | When to Use | | ------------ | ------------------------------ | -------------------------------------------------------------------- | | `cosine` | Cosine similarity (normalized) | Most embedding models (default choice) | | `dotproduct` | Dot product | Hybrid search (required for sparse-dense), pre-normalized embeddings | | `euclidean` | L2 distance | Raw feature vectors where absolute distance matters | --- ## Supported Metadata Types | Type | Example Value | Filter Support | Notes | | -------------- | ------------- | ----------------------- | ----------------------------------------- | | String | `"drama"` | All operators | Keys cannot start with `$` | | Number (int) | `2024` | All operators | Comparisons are type-strict | | Number (float) | `0.95` | All operators | Stored as 64-bit floats | | Boolean | `true` | `$eq`, `$ne`, `$exists` | No `$gt`/`$lt` on booleans | | String array | `["a", "b"]` | `$in`, `$nin`, `$eq` | Arrays of strings ONLY (no number arrays) | --- ## Production Checklist ### Security - [ ] API key stored in environment variable, not in code - [ ] SDK used server-side only (never in browser -- exposes API key) - [ ] Index access restricted by project/API key scoping ### Index Configuration - [ ] Dimension matches your embedding model's output dimension exactly - [ ] Metric matches your embedding model's recommendation (cosine for most) - [ ] Serverless index used unless pod-based is specifically required - [ ] Index region chosen close to your application servers ### Data Management - [ ] Metadata is flat key-value pairs (no nested objects) - [ ] Metadata per record under 40 KB - [ ] Record IDs are unique, deterministic, and under 512 characters - [ ] Upserts batched at 200 records (well under 1,000/2 MB limit) - [ ] Namespaces used for multi-tenant data isolation ### Query Optimization - [ ] `topK` set to minimum needed (lower = faster + cheaper) - [ ] `includeMetadata` and `includeValues` only when needed (max topK drops to 1,000) - [ ] Metadata filters use indexed fields with reasonable cardinality - [ ] Namespaces used instead of metadata filtering for tenant isolation ### Consistency & Reliability - [ ] Application tolerates eventual consistency after upserts - [ ] Freshness-critical flows poll `describeIndexStats()` before querying - [ ] Error handling for 400 (bad request), 429 (rate limit), 500 (server error) - [ ] Retry logic with exponential backoff for transient failures ### Monitoring - [ ] Track vector count via `describeIndexStats()` - [ ] Monitor read/write unit consumption - [ ] Alert on dimension mismatch errors (misconfigured embedding pipeline) - [ ] Track query latency percentiles --- _Full skill documentation: [SKILL.md](SKILL.md) | Examples: [examples/](examples/)_ -
SKILL.md 16.2 KB
--- name: api-vector-db-pinecone description: Pinecone serverless vector database -- index management, vector operations, metadata filtering, namespaces, hybrid search, inference API --- # Pinecone Patterns > **Quick Guide:** Use `@pinecone-database/pinecone` (v7.x) for serverless vector database operations. Target indexes by host (`pc.index({ host })`), not by name. Use namespaces for multi-tenant isolation (physically separate, cheaper queries). Batch upserts at 200 records (max 1,000 or 2 MB). Metadata is limited to 40 KB per record with flat key-value pairs only (no nested objects). Pinecone is eventually consistent -- vectors may not appear in queries immediately after upsert. Use `describeIndexStats()` to verify indexing progress. For hybrid search, use `dotproduct` metric with sparse+dense vectors in a single index. --- <critical_requirements> ## CRITICAL: Before Using This Skill > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST target indexes by host URL, not by name -- `pc.index({ host })` is the v7 API; `pc.index('name')` is deprecated)** **(You MUST batch upserts to max 1,000 records or 2 MB per request -- exceeding either limit causes a 400 error)** **(You MUST use flat key-value metadata only -- nested objects, null values, and keys starting with `$` are rejected by Pinecone)** **(You MUST handle eventual consistency -- vectors are not queryable immediately after upsert; use `describeIndexStats()` or retry logic for freshness-critical flows)** </critical_requirements> --- ## Examples - [Core Patterns](examples/core.md) -- Client setup, index creation, upsert, query, fetch, update, delete - [Namespaces & Multi-Tenancy](examples/namespaces.md) -- Namespace isolation, multi-tenant patterns, namespace management API - [Metadata Filtering](examples/metadata-filtering.md) -- Filter operators, compound filters, best practices - [Hybrid Search](examples/hybrid-search.md) -- Sparse-dense vectors, hybrid index setup, alpha weighting - [Inference API](examples/inference.md) -- Embedding generation, reranking, integrated inference indexes - [Batch Operations](examples/batch-operations.md) -- Chunked upserts, parallel ingestion, bulk import **Additional resources:** - [reference.md](reference.md) -- API quick reference, filter operators, limits, decision frameworks, production checklist --- **Auto-detection:** Pinecone, @pinecone-database/pinecone, createIndex, createIndexForModel, upsert, query, topK, includeMetadata, sparseValues, namespace, describeIndexStats, vector database, similarity search, embedding, cosine, dotproduct, euclidean, RAG retrieval, semantic search, pinecone-sparse-english, rerank, searchRecords, upsertRecords, fetchByMetadata **When to use:** - Semantic search over document embeddings (RAG retrieval) - Similarity search for recommendations, deduplication, or classification - Multi-tenant vector isolation using namespaces - Hybrid semantic + keyword search using sparse-dense vectors - Embedding generation and result reranking via Pinecone Inference API **Key patterns covered:** - Client setup and index management (serverless vs pod-based) - Vector CRUD operations (upsert, query, fetch, update, delete) - Metadata filtering with compound operators - Namespace-based multi-tenancy - Sparse-dense hybrid search - Pinecone Inference API (embed, rerank) - Batch ingestion with chunking and parallelism - Integrated inference indexes (automatic embedding) **When NOT to use:** - Full-text search with complex boolean queries (use a dedicated search engine) - Relational data with joins and transactions (use a relational database) - Real-time streaming or pub/sub messaging (use a message broker) - Storing large binary blobs or documents (use object storage; store only embeddings + metadata references) --- <philosophy> ## Philosophy Pinecone is a **managed serverless vector database** purpose-built for similarity search at scale. The core principle: **store embeddings and metadata, query by vector similarity, filter by metadata.** **Core principles:** 1. **Vectors in, results out** -- Pinecone stores high-dimensional vectors and returns the most similar ones. It is not a general-purpose database. Structure your data as embeddings + metadata references. 2. **Namespaces for isolation** -- Use namespaces to physically separate tenant data. Queries scan only the target namespace, reducing cost and latency compared to metadata filtering across a shared namespace. 3. **Metadata is for filtering, not storage** -- Keep metadata small (40 KB limit) and flat. Store document content in your primary database; store only filterable attributes (category, date, tenant ID) as Pinecone metadata. 4. **Batch for throughput** -- Individual upserts are inefficient. Batch at 200 records for optimal throughput (max 1,000 or 2 MB per request). 5. **Eventual consistency is normal** -- Freshly upserted vectors may not appear in query results immediately. Design your application to tolerate brief staleness or poll `describeIndexStats()` before querying. </philosophy> --- <patterns> ## Core Patterns ### Pattern 1: Client Initialization Create a Pinecone client from an API key. See [examples/core.md](examples/core.md) for full examples. ```typescript // Good Example import { Pinecone } from "@pinecone-database/pinecone"; function createPineconeClient(): Pinecone { const apiKey = process.env.PINECONE_API_KEY; if (!apiKey) { throw new Error("PINECONE_API_KEY environment variable is required"); } return new Pinecone({ apiKey }); } export { createPineconeClient }; ``` **Why good:** API key from environment variable, validation before construction, named export ```typescript // Bad Example import { Pinecone } from "@pinecone-database/pinecone"; const pc = new Pinecone({ apiKey: "sk-abc123..." }); // Hardcoded key leaks in version control ``` **Why bad:** Hardcoded API key is a security risk, no validation --- ### Pattern 2: Index Targeting (v7 API) Always target an index by its host URL, not its name. See [examples/core.md](examples/core.md). ```typescript // Good Example -- target by host const indexModel = await pc.createIndex({ name: "products", dimension: EMBEDDING_DIMENSION, metric: "cosine", spec: { serverless: { cloud: "aws", region: "us-east-1" } }, }); const index = pc.index({ host: indexModel.host }); ``` **Why good:** `pc.index({ host })` is the v7 API, avoids an extra API call to resolve the name to a host ```typescript // Bad Example -- target by name (deprecated) const index = pc.index("products"); // Triggers an extra describeIndex call to resolve the host URL ``` **Why bad:** Targeting by name requires an extra network call and is deprecated in v7 --- ### Pattern 3: Upsert with Metadata Upsert vectors with flat metadata for filtering. See [examples/core.md](examples/core.md) for typed metadata. ```typescript // Good Example interface DocumentMetadata { title: string; category: string; createdAt: number; // Unix timestamp (numbers only, no Date objects) } const NAMESPACE = "articles"; await index.namespace(NAMESPACE).upsert({ records: [ { id: "doc-1", values: embedding, // number[] matching index dimension metadata: { title: "Guide", category: "tutorial", createdAt: 1710000000 }, }, ], }); ``` **Why good:** Typed metadata interface, flat key-value pairs, numeric timestamp (not Date), namespace isolation --- ### Pattern 4: Query with Metadata Filter Query for similar vectors with metadata filtering. See [examples/metadata-filtering.md](examples/metadata-filtering.md) for all operators. ```typescript // Good Example const TOP_K = 10; const results = await index.namespace(NAMESPACE).query({ vector: queryEmbedding, topK: TOP_K, includeMetadata: true, filter: { $and: [ { category: { $eq: "tutorial" } }, { createdAt: { $gte: 1700000000 } }, ], }, }); for (const match of results.matches) { console.log(match.id, match.score, match.metadata); } ``` **Why good:** Named constant for topK, structured filter with `$and`, includes metadata in response ```typescript // Bad Example const results = await index.query({ vector: queryEmbedding, topK: 100, includeMetadata: true, filter: { tags: ["a", "b"] }, // INVALID: arrays are not valid filter values }); ``` **Why bad:** Missing namespace (queries default namespace), array filter syntax is invalid (use `$in`), no named constant for topK --- ### Pattern 5: Namespace-Based Multi-Tenancy Use namespaces for tenant isolation. See [examples/namespaces.md](examples/namespaces.md). ```typescript // Good Example -- physically isolated tenant data function getTenantIndex(pc: Pinecone, host: string, tenantId: string) { return pc.index({ host }).namespace(`tenant-${tenantId}`); } // Each tenant's queries scan only their namespace const tenantIndex = getTenantIndex(pc, INDEX_HOST, "acme-corp"); const results = await tenantIndex.query({ vector: embedding, topK: TOP_K }); ``` **Why good:** Physical isolation per tenant, queries scan only the target namespace (lower cost and latency) ```typescript // Bad Example -- metadata filtering for multi-tenancy await index.query({ vector: embedding, topK: 10, filter: { tenantId: { $eq: "acme-corp" } }, // Scans ENTIRE index, filters after -- expensive at scale }); ``` **Why bad:** Metadata filtering scans the full namespace regardless of filter selectivity, cost scales with total data not tenant data --- ### Pattern 6: Pinecone Inference API Generate embeddings and rerank results. See [examples/inference.md](examples/inference.md). ```typescript // Good Example -- embed text const embedResult = await pc.inference.embed({ model: "multilingual-e5-large", inputs: [{ text: "What is machine learning?" }], parameters: { inputType: "query", truncate: "END" }, }); const queryVector = embedResult.data[0].values; ``` **Why good:** Specifies `inputType` (query vs passage), handles truncation for long inputs ```typescript // Good Example -- rerank results const rerankResult = await pc.inference.rerank({ model: "pinecone-rerank-v0", query: "machine learning basics", documents: results.matches.map((m) => ({ id: m.id, text: m.metadata?.content as string, })), topN: 5, returnDocuments: true, }); ``` **Why good:** Reranks query results for better relevance, limits output with `topN` </patterns> --- <decision_framework> ## Decision Framework ### Which Index Type? ``` Which Pinecone index type should I use? |-- Serverless? (recommended for most use cases) | |-- Variable or unpredictable traffic? -> Serverless (auto-scales, pay-per-use) | |-- Starting a new project? -> Serverless (simpler, no capacity planning) | '-- Need hybrid sparse-dense search? -> Serverless with dotproduct metric | '-- Pod-based? (legacy, specific needs) |-- Need guaranteed low latency SLAs? -> Pod-based (dedicated compute) '-- Using collections for snapshots? -> Pod-based (collections are pod-only) ``` ### Which Metric? ``` Which distance metric should I use? |-- Using embeddings from a language model? -> cosine (normalized, most common) |-- Need hybrid search (sparse + dense)? -> dotproduct (REQUIRED for hybrid) |-- Comparing raw feature vectors? -> euclidean (absolute distance matters) '-- Unsure? -> cosine (safe default for most embedding models) ``` ### Namespaces vs Metadata Filtering? ``` How should I isolate tenant data? |-- Strict data isolation required? -> Namespaces (physical separation) |-- Need to query across tenants? -> Metadata filtering (logical separation) |-- Cost-sensitive at scale? -> Namespaces (query cost = tenant size, not total) |-- Few tenants (< 10)? -> Either approach works '-- Many tenants (100+)? -> Namespaces (metadata filtering scans everything) ``` ### Embedded Inference vs External Embeddings? ``` How should I generate embeddings? |-- Want simplest architecture? -> Integrated inference (createIndexForModel) | (Pinecone handles embedding automatically on upsert/query) | |-- Need a specific embedding model not hosted by Pinecone? -> External | (Generate embeddings yourself, upsert raw vectors) | |-- Need hybrid search with sparse vectors? -> External sparse model | (Use pinecone-sparse-english-v0 via inference API + your dense model) | '-- Need full control over embedding pipeline? -> External (Custom preprocessing, chunking, model selection) ``` </decision_framework> --- <red_flags> ## RED FLAGS **High Priority Issues:** - Targeting index by name instead of host -- `pc.index("name")` is deprecated in v7; use `pc.index({ host })` to avoid an extra API call - Upserting vectors with wrong dimensions -- dimension mismatch causes a 400 error; verify your embedding model's output dimension matches the index - Nested metadata objects -- Pinecone only supports flat key-value metadata; nested objects are silently ignored or rejected - Using metadata filtering for multi-tenancy at scale -- scans the entire namespace regardless of filter selectivity; use namespaces instead **Medium Priority Issues:** - Missing `includeMetadata: true` in queries -- metadata is NOT included by default; omitting this returns only IDs and scores - Upsert batches exceeding 1,000 records or 2 MB -- triggers a 400 error; chunk at 200 records for safety margin - Using `Date` objects in metadata -- Pinecone metadata supports strings, numbers, booleans, and string arrays only; convert dates to Unix timestamps - Not awaiting `createIndex()` readiness -- index creation is async; the index is not ready for operations immediately after `createIndex()` returns **Common Mistakes:** - Querying immediately after upsert and expecting results -- eventual consistency means freshly upserted vectors may not be queryable for seconds - Using `topK` > 1,000 with `includeMetadata: true` -- the max `topK` is 1,000 when including metadata or values; without them, max is 10,000 - Passing array values as metadata filter values (`{ tags: ["a", "b"] }`) -- use `$in` operator instead: `{ tags: { $in: ["a", "b"] } }` - Forgetting that `deleteAll()` without a namespace deletes from the default namespace only, not the entire index **Gotchas & Edge Cases:** - `describeIndexStats()` returns approximate counts -- record counts are not exact in real-time, especially after recent upserts or deletes - Metadata values are always returned as their original types, but filter comparisons are type-strict -- `$eq: "42"` does not match numeric `42` - Sparse vector indices must be positive 32-bit integers (uint32), and values must be non-zero floats -- zero values are silently dropped - `listPaginated()` returns vector IDs only (no values or metadata) -- use `fetch()` to get full vector data - The `$in` and `$nin` operators accept a maximum of 10,000 values each - Metadata keys cannot start with `$` (reserved for operators) - `upsert` is an upsert, not an insert -- upserting with an existing ID overwrites the previous vector and metadata entirely (no partial merge) - `update()` merges metadata by default -- updating metadata replaces only the fields you specify, not the entire metadata object - The v7 SDK includes built-in automatic retry with exponential backoff for transient errors -- custom retry logic is only needed for fine-grained control or non-default retry policies - `upsertRecords` (integrated inference) accepts a direct array, not `{ records: [...] }` -- this differs from the regular `upsert` method which uses `{ records: [...] }` </red_flags> --- <critical_reminders> ## CRITICAL REMINDERS > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST target indexes by host URL, not by name -- `pc.index({ host })` is the v7 API; `pc.index('name')` is deprecated)** **(You MUST batch upserts to max 1,000 records or 2 MB per request -- exceeding either limit causes a 400 error)** **(You MUST use flat key-value metadata only -- nested objects, null values, and keys starting with `$` are rejected by Pinecone)** **(You MUST handle eventual consistency -- vectors are not queryable immediately after upsert; use `describeIndexStats()` or retry logic for freshness-critical flows)** **Failure to follow these rules will cause index creation failures, rejected upserts, empty query results, and degraded multi-tenant performance.** </critical_reminders>
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.