api-search-elasticsearch
Elasticsearch patterns -- client setup, index management, search DSL, aggregations, vector search, bulk operations, deep pagination
Install
npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/api-search-elasticsearch/skills/api-search-elasticsearch
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
git clone https://github.com/agents-inc/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Elasticsearch Patterns
Quick Guide: Use
@elastic/elasticsearch(v8.x/v9.x) as the TypeScript client. Elasticsearch is near real-time -- documents are NOT searchable immediately after indexing; they become visible after a refresh (default: every 1 second on active indices). You MUST define explicit mappings before indexing -- dynamic mapping infers types from the first document, and mismatched types in later documents cause hard failures you cannot fix without reindexing. Usesearch_after+ Point in Time (PIT) for deep pagination -- NOTfrom/sizebeyond 10,000 hits and NOT the scroll API (deprecated for search). Use thebulkAPI orclient.helpers.bulk()for any batch operation -- never loop individual index calls.
<critical_requirements>
CRITICAL: Before Using This Skill
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST define explicit index mappings BEFORE indexing documents -- dynamic mapping infers types from the first document, and if a later document sends a different type for the same field, indexing fails with a mapper_parsing_exception that CANNOT be fixed without reindexing into a new index)
(You MUST use the bulk API for batch operations -- looping individual client.index() calls is orders of magnitude slower and can overwhelm the cluster with HTTP connections)
(You MUST NOT use from/size pagination beyond 10,000 results -- Elasticsearch throws Result window is too large by default; use search_after + PIT instead)
(You MUST NOT use refresh: true or refresh: "wait_for" in production request handlers -- forcing a refresh on every write degrades cluster performance; let the default 1-second refresh interval handle it)
</critical_requirements>
Examples
- Core Patterns -- Client setup, index management, document CRUD, search basics, TypeScript integration
- Aggregations -- Terms, range, date_histogram, nested, pipeline aggregations
- Vector Search -- Dense vector fields, kNN queries, hybrid search, similarity metrics
- Pagination -- from/size, search_after, Point in Time, scroll helpers
- Bulk Operations -- Bulk API, bulk helper, reindexing patterns
Additional resources:
- reference.md -- Search DSL cheat sheet, mapping types, aggregation reference, decision frameworks
Auto-detection: Elasticsearch, elasticsearch, @elastic/elasticsearch, client.search, client.index, client.bulk, client.indices.create, client.indices.putMapping, dense_vector, knn, search_after, point in time, openPIT, aggregations, aggs, bool query, match query, term query, multi_match, nested query, range query, client.helpers.bulk, client.helpers.scrollSearch, BulkResponse, SearchResponse, MappingProperty
When to use:
- Full-text search with advanced relevance tuning (BM25, custom analyzers, boosting)
- Aggregations and analytics (terms, histograms, pipeline aggregations)
- Vector/semantic search with kNN on dense_vector fields
- Log and event data search with time-based queries
- Complex structured queries combining bool, nested, range, and geo filters
- Search across large datasets requiring deep pagination (search_after + PIT)
Key patterns covered:
- Client initialization and connection management
- Index management with explicit mappings and settings
- Document CRUD (index, get, update, delete)
- Search DSL (match, term, bool, range, nested, multi_match)
- Aggregations (terms, range, date_histogram, nested, pipeline)
- Full-text analysis (custom analyzers, tokenizers, filters)
- Vector search (dense_vector, kNN, hybrid text+vector)
- Bulk operations and reindexing
- Deep pagination (search_after + PIT)
When NOT to use:
- Simple keyword search on small datasets (client-side filtering or database LIKE queries are simpler)
- Primary data store (Elasticsearch is a search engine, not a database -- always have a source of truth elsewhere)
- Strong consistency requirements (Elasticsearch is eventually consistent by design)
- Simple autocomplete on a small list (a prefix trie or client-side filter is simpler)
<decision_framework>
Decision Framework
Which Query Type?
What kind of search do I need?
-- Full-text relevance search? -> match / multi_match in must
-- Exact value filtering? -> term / terms / range in filter
-- Combining text + filters? -> bool query (must for text, filter for exact)
-- Fuzzy matching? -> match with fuzziness: "AUTO"
-- Phrase matching? -> match_phrase
-- Complex nested objects? -> nested query with path
-- Vector similarity? -> knn with dense_vector field
-- Text + vector hybrid? -> query + knn in same request
Pagination Strategy?
How deep do results go?
-- Under 10,000 total? -> from/size (simplest)
-- Over 10,000 hits? -> search_after + PIT (recommended)
-- Bulk data export? -> client.helpers.scrollSearch() or scrollDocuments()
-- Real-time infinite scroll? -> search_after (no PIT needed for forward-only)
text vs keyword?
What will I do with this field?
-- Full-text search (tokenized, relevance)? -> text
-- Exact match, filtering, aggregations, sorting? -> keyword
-- Both? -> Multi-field: { type: "text", fields: { keyword: { type: "keyword" } } }
-- Neither (just stored, never queried)? -> { type: "keyword", index: false }
Filter vs Must?
Does relevance scoring matter for this clause?
-- YES (affects result order) -> must
-- NO (binary yes/no filter) -> filter (cached, no scoring overhead)
-- Exclude documents -> must_not (in filter context)
-- Boost if present (optional) -> should with minimum_should_match: 0
</decision_framework>
<red_flags>
RED FLAGS
High Priority Issues:
- Relying on dynamic mapping without explicit mappings -- wrong type inference causes
mapper_parsing_exceptionthat requires reindexing to fix - Using
from/sizebeyond 10,000 results -- Elasticsearch throwsResult window is too large; usesearch_after+ PIT - Looping individual
client.index()calls instead ofclient.bulk()orclient.helpers.bulk()-- orders of magnitude slower, can overwhelm the cluster - Using
refresh: trueorrefresh: "wait_for"in production request handlers -- forces a segment refresh on every write, degrades cluster performance under load
Medium Priority Issues:
- Putting exact-match conditions (term, range) in
mustinstead offilter-- wastes CPU on scoring, misses filter cache - Using
texttype for fields that need exact matching or aggregation -- text fields are analyzed (tokenized), making aggregations return individual tokens instead of full values - Not including a tiebreaker field in
sortwhen usingsearch_after-- documents with identical sort values may be skipped or duplicated across pages - Missing
_sourcecheck --hit._sourcecan beundefinedif_sourceis disabled or fields are excluded; always handle this
Gotchas & Edge Cases:
- Mapping types are immutable -- once a field is mapped as
text, you cannot change it tokeyword. The only fix is to create a new index with correct mappings and reindex all documents textvskeywordconfusion --textfields are tokenized ("New York" becomes ["new", "york"]). Aggregating on atextfield gives you individual tokens, not full values. Usekeywordor a.keywordsub-field for aggregationsmatchvstermon text fields --termon atextfield often returns no results becausetermdoes NOT analyze the query but the field value IS analyzed (e.g., term "New York" won't match the analyzed tokens "new" and "york")- Near real-time delay -- after indexing, documents are NOT searchable until the next refresh (default: 1 second). Tests that index then immediately search must use
refresh: "wait_for"or explicitclient.indices.refresh() - Default
index.max_result_windowis 10,000 -- increasing this is possible but NOT recommended; deep pagination withfrom/sizeholds all skipped results in memory - Nested objects require
nestedmapping type -- arrays of objects are flattened by default, losing the association between fields within each object. If you need to query "color: red AND size: large" on the same object in an array, usenested - Aggregation on
_idis disabled by default (8.x+) -- use a separateidfield if you need to aggregate by document ID _scoreis null in filter context -- clauses infilterdo not contribute to scoring; if you need scoring, usemust- Bulk API partial failures -- a bulk request can succeed overall but have individual failures. Always check
result.errorsand iterateresult.itemsto find failed operations - Scroll API is deprecated for search -- use
search_after+ PIT for deep pagination. Scroll is still valid for one-time data export but consumes cluster resources (open search contexts) - PIT must be closed -- failing to close Point in Time contexts leaks resources on the cluster; always close in a finally block
</red_flags>
<critical_reminders>
CRITICAL REMINDERS
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST define explicit index mappings BEFORE indexing documents -- dynamic mapping infers types from the first document, and if a later document sends a different type for the same field, indexing fails with a mapper_parsing_exception that CANNOT be fixed without reindexing into a new index)
(You MUST use the bulk API for batch operations -- looping individual client.index() calls is orders of magnitude slower and can overwhelm the cluster with HTTP connections)
(You MUST NOT use from/size pagination beyond 10,000 results -- Elasticsearch throws Result window is too large by default; use search_after + PIT instead)
(You MUST NOT use refresh: true or refresh: "wait_for" in production request handlers -- forcing a refresh on every write degrades cluster performance; let the default 1-second refresh interval handle it)
Failure to follow these rules will cause mapping conflicts, pagination failures, cluster performance degradation, and silent data loss.
</critical_reminders>
Files (skills)
-
examples
-
aggregations.md 9.9 KB
# Elasticsearch -- Aggregation Examples > Terms, range, date_histogram, nested, and pipeline aggregation patterns. Reference from [SKILL.md](../SKILL.md). **Prerequisites:** Understand client setup and index management from [core.md](core.md) first. **Related examples:** - [core.md](core.md) -- Client setup, document operations, search basics - [pagination.md](pagination.md) -- Paginating aggregation results --- ## Terms Aggregation Group by field values and get document counts. ```typescript import type { Client } from "@elastic/elasticsearch"; const INDEX_NAME = "products"; const AGGREGATION_SIZE = 50; async function getCategoryDistribution( client: Client, ): Promise<Array<{ key: string; count: number }>> { const result = await client.search({ index: INDEX_NAME, size: 0, // No hits needed -- only aggregations aggs: { categories: { terms: { field: "categories", size: AGGREGATION_SIZE }, }, }, }); const buckets = ( result.aggregations?.categories as { buckets: Array<{ key: string; doc_count: number }>; } )?.buckets ?? []; return buckets.map((b) => ({ key: b.key, count: b.doc_count })); } export { getCategoryDistribution }; ``` **Why good:** `size: 0` skips hits (faster when only aggregations matter), named constant for aggregation size, typed bucket extraction **Important:** `terms` aggregation returns an approximate count. The `size` parameter controls how many top buckets to return (default: 10), NOT how many documents to scan. --- ## Terms with Sub-Aggregations Nest metric aggregations inside bucket aggregations. ```typescript async function getCategoryStats( client: Client, ): Promise< Array<{ category: string; count: number; avgPrice: number; maxPrice: number }> > { const result = await client.search({ index: INDEX_NAME, size: 0, aggs: { categories: { terms: { field: "categories", size: AGGREGATION_SIZE }, aggs: { avgPrice: { avg: { field: "price" } }, maxPrice: { max: { field: "price" } }, }, }, }, }); const buckets = ( result.aggregations?.categories as { buckets: Array<{ key: string; doc_count: number; avgPrice: { value: number | null }; maxPrice: { value: number | null }; }>; } )?.buckets ?? []; return buckets.map((b) => ({ category: b.key, count: b.doc_count, avgPrice: b.avgPrice.value ?? 0, maxPrice: b.maxPrice.value ?? 0, })); } export { getCategoryStats }; ``` **Why good:** Nested sub-aggregations compute per-bucket metrics, `value` can be null (e.g., empty buckets), null handled with fallback --- ## Range Aggregation Create custom numeric ranges for bucketing. ```typescript async function getPriceRanges( client: Client, ): Promise<Array<{ range: string; count: number }>> { const result = await client.search({ index: INDEX_NAME, size: 0, aggs: { priceRanges: { range: { field: "price", ranges: [ { key: "budget", to: 50 }, { key: "mid-range", from: 50, to: 200 }, { key: "premium", from: 200, to: 500 }, { key: "luxury", from: 500 }, ], }, }, }, }); const buckets = ( result.aggregations?.priceRanges as { buckets: Array<{ key: string; doc_count: number }>; } )?.buckets ?? []; return buckets.map((b) => ({ range: b.key, count: b.doc_count })); } export { getPriceRanges }; ``` **Why good:** Named keys make bucket identification readable, ranges cover the full spectrum with no gaps **Gotcha:** `from` is inclusive, `to` is exclusive. A document with `price: 50` falls into "mid-range" (50-200), not "budget" (to: 50). --- ## Date Histogram Time-based bucketing for time series data. ```typescript async function getMonthlySales( client: Client, year: number, ): Promise<Array<{ month: string; count: number; revenue: number }>> { const result = await client.search({ index: "orders", size: 0, query: { range: { orderDate: { gte: `${year}-01-01`, lt: `${year + 1}-01-01`, }, }, }, aggs: { monthly: { date_histogram: { field: "orderDate", calendar_interval: "month", format: "yyyy-MM", min_doc_count: 0, // Include empty months }, aggs: { revenue: { sum: { field: "totalAmount" } }, }, }, }, }); const buckets = ( result.aggregations?.monthly as { buckets: Array<{ key_as_string: string; doc_count: number; revenue: { value: number }; }>; } )?.buckets ?? []; return buckets.map((b) => ({ month: b.key_as_string, count: b.doc_count, revenue: b.revenue.value, })); } export { getMonthlySales }; ``` **Why good:** `calendar_interval: "month"` handles varying month lengths correctly, `min_doc_count: 0` ensures empty months appear in results, `format` controls the `key_as_string` output **Gotcha:** Use `calendar_interval` for months/quarters/years (variable length). Use `fixed_interval` for exact durations like "30d", "1h", "5m". Using `fixed_interval: "1M"` is an error -- months are not a fixed duration. --- ## Nested Aggregation Aggregate inside nested objects. ```typescript // Requires "reviews" field mapped as "nested" type async function getRatingDistribution( client: Client, ): Promise<Array<{ rating: number; count: number }>> { const result = await client.search({ index: INDEX_NAME, size: 0, aggs: { reviewsNested: { nested: { path: "reviews" }, aggs: { ratingBuckets: { histogram: { field: "reviews.rating", interval: 1, min_doc_count: 0, }, }, }, }, }, }); const buckets = ( result.aggregations?.reviewsNested as { ratingBuckets: { buckets: Array<{ key: number; doc_count: number }>; }; } )?.ratingBuckets.buckets ?? []; return buckets.map((b) => ({ rating: b.key, count: b.doc_count })); } export { getRatingDistribution }; ``` **Why good:** `nested` aggregation scope enters the nested documents before aggregating, histogram with interval 1 creates one bucket per rating value **Important:** Without the `nested` aggregation wrapper, aggregating on `reviews.rating` would aggregate on the flattened array values, giving incorrect counts. --- ## Pipeline Aggregation Compute values from the output of other aggregations. ```typescript async function getMonthlyRevenueWithMovingAvg( client: Client, ): Promise< Array<{ month: string; revenue: number; movingAvg: number | null }> > { const MOVING_FN_WINDOW = 3; const result = await client.search({ index: "orders", size: 0, aggs: { monthly: { date_histogram: { field: "orderDate", calendar_interval: "month", }, aggs: { revenue: { sum: { field: "totalAmount" } }, revenueMovingAvg: { moving_fn: { buckets_path: "revenue", window: MOVING_FN_WINDOW, script: "MovingFunctions.unweightedAvg(values)", }, }, }, }, }, }); const buckets = ( result.aggregations?.monthly as { buckets: Array<{ key_as_string: string; revenue: { value: number }; revenueMovingAvg?: { value: number }; }>; } )?.buckets ?? []; return buckets.map((b) => ({ month: b.key_as_string, revenue: b.revenue.value, movingAvg: b.revenueMovingAvg?.value ?? null, })); } export { getMonthlyRevenueWithMovingAvg }; ``` **Why good:** `moving_fn` pipeline aggregation computes a rolling average from the `revenue` sub-aggregation using `MovingFunctions.unweightedAvg(values)`, `buckets_path` references the sibling aggregation by name, first N-1 buckets have null moving average (not enough data) **Important:** `moving_avg` was removed in Elasticsearch 8.0. Use `moving_fn` with a script instead. Available predefined functions: `MovingFunctions.unweightedAvg(values)`, `MovingFunctions.linearWeightedAvg(values)`, `MovingFunctions.ewma(values, alpha)`, `MovingFunctions.holt(values, alpha, beta)`, `MovingFunctions.holtWinters(values, alpha, beta, gamma, period, multiplicative)` --- ## Aggregation with Filtered Scope Apply a filter to an aggregation without affecting the main query. ```typescript async function getInStockVsOutOfStock( client: Client, ): Promise<{ inStock: number; outOfStock: number }> { const result = await client.search({ index: INDEX_NAME, size: 0, aggs: { inStockCount: { filter: { term: { inStock: true } }, }, outOfStockCount: { filter: { term: { inStock: false } }, }, }, }); return { inStock: (result.aggregations?.inStockCount as { doc_count: number })?.doc_count ?? 0, outOfStock: (result.aggregations?.outOfStockCount as { doc_count: number }) ?.doc_count ?? 0, }; } export { getInStockVsOutOfStock }; ``` **Why good:** `filter` aggregation applies a separate filter scope per aggregation, each counting only matching documents --- ## Cardinality (Approximate Distinct Count) ```typescript async function getUniqueBrandCount(client: Client): Promise<number> { const result = await client.search({ index: INDEX_NAME, size: 0, aggs: { uniqueBrands: { cardinality: { field: "brand" }, }, }, }); return (result.aggregations?.uniqueBrands as { value: number })?.value ?? 0; } export { getUniqueBrandCount }; ``` **Gotcha:** `cardinality` is an approximation using HyperLogLog++. For fields with < 1000 unique values, it's exact. For larger cardinalities, expect ~2-3% error margin. --- _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_ -
bulk-operations.md 9.6 KB
# Elasticsearch -- Bulk Operations Examples > Bulk API, bulk helper, error handling, and reindexing patterns. Reference from [SKILL.md](../SKILL.md). **Prerequisites:** Understand client setup and index management from [core.md](core.md) first. **Related examples:** - [core.md](core.md) -- Client setup, index management, document CRUD - [pagination.md](pagination.md) -- scrollDocuments for reading all documents during reindex --- ## Bulk API (Low-Level) The raw bulk API uses alternating action/document pairs. ```typescript import type { Client, BulkResponse } from "@elastic/elasticsearch"; const INDEX_NAME = "products"; interface Product { productId: string; name: string; price: number; categories: string[]; } async function bulkIndexProducts( client: Client, products: Product[], ): Promise<{ successful: number; failed: number }> { const operations = products.flatMap((doc) => [ { index: { _index: INDEX_NAME, _id: doc.productId } }, doc, ]); const result: BulkResponse = await client.bulk({ operations, refresh: "wait_for", // Only in scripts/seeds -- NOT in production handlers }); if (result.errors) { const failedItems = result.items.filter( (item) => item.index?.error !== undefined, ); for (const item of failedItems) { console.error( `Failed to index ${item.index?._id}: ${item.index?.error?.reason}`, ); } return { successful: products.length - failedItems.length, failed: failedItems.length, }; } return { successful: products.length, failed: 0 }; } export { bulkIndexProducts }; ``` **Why good:** `flatMap` creates the alternating action/document format, explicit `_id` for upsert behavior, error handling checks `result.errors` and iterates `result.items` for details **Gotcha:** A bulk request can return HTTP 200 but still have individual failures. Always check `result.errors` -- it's true if ANY item failed. Then iterate `result.items` to find which ones. **Gotcha:** A 429 status on individual items means "too many requests" -- these are retriable. Other error codes (400, 409) typically indicate data issues that require fixing the document. --- ## Bulk Helper (Recommended) The bulk helper handles batching, concurrency, retries, and back-pressure automatically. ```typescript async function bulkIndexWithHelper( client: Client, products: Product[], ): Promise<{ total: number; successful: number; failed: number }> { const result = await client.helpers.bulk<Product>({ datasource: products, onDocument(doc) { return { index: { _index: INDEX_NAME, _id: doc.productId } }; }, refreshOnCompletion: INDEX_NAME, }); return { total: result.total, successful: result.successful, failed: result.failed, }; } export { bulkIndexWithHelper }; ``` **Why good:** Automatic batching (default: 5MB per batch), automatic concurrency (default: 5 parallel requests), automatic retries on 429, `refreshOnCompletion` triggers one refresh at the end ### Bulk Helper with Streaming Data Source ```typescript import { createReadStream } from "node:fs"; import { createInterface } from "node:readline"; async function bulkIndexFromFile( client: Client, filePath: string, ): Promise<{ total: number; failed: number }> { const lineReader = createInterface({ input: createReadStream(filePath), }); async function* generateDocuments() { for await (const line of lineReader) { if (line.trim()) { yield JSON.parse(line) as Product; } } } const result = await client.helpers.bulk<Product>({ datasource: generateDocuments(), onDocument(doc) { return { index: { _index: INDEX_NAME, _id: doc.productId } }; }, refreshOnCompletion: INDEX_NAME, }); return { total: result.total, failed: result.failed }; } export { bulkIndexFromFile }; ``` **Why good:** Async generator streams documents from file without loading all into memory, supports NDJSON format (one JSON object per line) ### Bulk Update ```typescript async function bulkUpdatePrices( client: Client, updates: Array<{ productId: string; newPrice: number }>, ): Promise<{ total: number; failed: number }> { const result = await client.helpers.bulk({ datasource: updates, onDocument(item) { return [ { update: { _index: INDEX_NAME, _id: item.productId } }, { doc: { price: item.newPrice } }, ]; }, }); return { total: result.total, failed: result.failed }; } export { bulkUpdatePrices }; ``` **Why good:** `onDocument` returns a two-element array for update operations -- first element is the action, second is the update body with `doc` for partial update ### Bulk Delete ```typescript async function bulkDeleteProducts( client: Client, productIds: string[], ): Promise<{ total: number; failed: number }> { const result = await client.helpers.bulk({ datasource: productIds, onDocument(id) { return { delete: { _index: INDEX_NAME, _id: id } }; }, }); return { total: result.total, failed: result.failed }; } export { bulkDeleteProducts }; ``` **Why good:** Delete operations have no document body -- `onDocument` returns only the action --- ## Bulk Helper Configuration ```typescript const FLUSH_BYTES = 5_000_000; // 5MB -- max batch size before flushing const FLUSH_INTERVAL_MS = 30_000; // 30s -- max time before flushing const CONCURRENCY = 5; // Parallel bulk requests const RETRIES = 3; // Retry attempts per document on 429 const result = await client.helpers.bulk<Product>({ datasource: products, onDocument(doc) { return { index: { _index: INDEX_NAME, _id: doc.productId } }; }, flushBytes: FLUSH_BYTES, flushInterval: FLUSH_INTERVAL_MS, concurrency: CONCURRENCY, retries: RETRIES, refreshOnCompletion: INDEX_NAME, onDrop(doc) { // Called when a document fails after all retries console.error(`Dropped document: ${doc.document.productId}`); }, }); ``` **Why good:** Named constants for all tuning parameters, `onDrop` callback for monitoring permanently failed documents --- ## Reindexing with Alias Swap Zero-downtime reindex by creating a new index, copying data, and swapping aliases. ```typescript const READ_ALIAS = "products-read"; const WRITE_ALIAS = "products-write"; const BATCH_SIZE = 500; async function reindexWithAliasSwap( client: Client, newIndexName: string, ): Promise<void> { // 1. Create new index with updated mappings await client.indices.create({ index: newIndexName, mappings: { properties: { productId: { type: "keyword" }, name: { type: "text", fields: { keyword: { type: "keyword" } }, }, description: { type: "text" }, price: { type: "float" }, categories: { type: "keyword" }, brand: { type: "keyword" }, inStock: { type: "boolean" }, rating: { type: "float" }, // New field createdAt: { type: "date" }, }, }, }); // 2. Disable refresh during bulk copy (faster indexing) await client.indices.putSettings({ index: newIndexName, settings: { index: { refresh_interval: "-1" } }, }); // 3. Copy documents from old index using scroll const docs = client.helpers.scrollDocuments<Product>({ index: READ_ALIAS, query: { match_all: {} }, }); await client.helpers.bulk({ datasource: docs, onDocument(doc) { return { index: { _index: newIndexName, _id: doc.productId } }; }, refreshOnCompletion: newIndexName, }); // 4. Re-enable refresh await client.indices.putSettings({ index: newIndexName, settings: { index: { refresh_interval: "1s" } }, }); // 5. Atomic alias swap // Find the current index behind the alias const aliasInfo = await client.indices.getAlias({ name: READ_ALIAS }); const oldIndexNames = Object.keys(aliasInfo); await client.indices.updateAliases({ actions: [ ...oldIndexNames.map((oldIdx) => ({ remove: { index: oldIdx, alias: READ_ALIAS }, })), ...oldIndexNames.map((oldIdx) => ({ remove: { index: oldIdx, alias: WRITE_ALIAS }, })), { add: { index: newIndexName, alias: READ_ALIAS } }, { add: { index: newIndexName, alias: WRITE_ALIAS } }, ], }); } export { reindexWithAliasSwap }; ``` **Why good:** Zero-downtime reindex using alias swap, refresh disabled during bulk copy for speed, `scrollDocuments` streams data without loading all into memory, alias swap is atomic **Gotcha:** Disabling refresh during bulk copy is critical for performance -- without it, Elasticsearch creates a new segment every second during the copy. Always re-enable refresh after the copy completes. --- ## Server-Side Reindex API For simple reindexing without transformation, use the built-in reindex API. ```typescript async function reindexServerSide( client: Client, sourceIndex: string, destIndex: string, ): Promise<{ total: number }> { const result = await client.reindex({ source: { index: sourceIndex }, dest: { index: destIndex }, wait_for_completion: true, // Block until done -- only in scripts }); return { total: result.total ?? 0 }; } export { reindexServerSide }; ``` **Why good:** Server-side reindex avoids round-tripping documents through the client -- faster for large datasets. Use `wait_for_completion: false` for very large reindexes and poll the task API instead. **When to use:** When you need to copy data between indices without transforming the document structure. For transformations (adding fields, changing types), use the client-side scroll + bulk pattern above. --- _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_ -
core.md 14.4 KB
# Elasticsearch -- Core Pattern Examples > Client setup, index management, document CRUD, search basics, and TypeScript integration. Reference from [SKILL.md](../SKILL.md). **Related examples:** - [aggregations.md](aggregations.md) -- Terms, range, date_histogram, pipeline aggregations - [vector-search.md](vector-search.md) -- Dense vector fields, kNN, hybrid search - [pagination.md](pagination.md) -- search_after, PIT, scroll helpers - [bulk-operations.md](bulk-operations.md) -- Bulk API, bulk helper, reindexing --- ## Client Setup ### Basic Setup with API Key Auth ```typescript import { Client } from "@elastic/elasticsearch"; function createElasticsearchClient(): Client { const node = process.env.ELASTICSEARCH_URL; if (!node) { throw new Error("ELASTICSEARCH_URL environment variable is required"); } return new Client({ node, auth: { apiKey: process.env.ELASTICSEARCH_API_KEY ?? "", }, }); } export { createElasticsearchClient }; ``` **Why good:** Environment variable validation, API key auth (preferred over basic auth), named export ### Elastic Cloud Setup ```typescript import { Client } from "@elastic/elasticsearch"; function createCloudClient(): Client { const cloudId = process.env.ELASTIC_CLOUD_ID; if (!cloudId) { throw new Error("ELASTIC_CLOUD_ID environment variable is required"); } return new Client({ cloud: { id: cloudId }, auth: { apiKey: process.env.ELASTICSEARCH_API_KEY ?? "", }, }); } export { createCloudClient }; ``` **Why good:** `cloud.id` encodes the cluster URL -- no need to construct URLs manually, handles TLS automatically ### Health Check ```typescript import type { Client } from "@elastic/elasticsearch"; const HEALTH_TIMEOUT_MS = 5000; async function verifyConnection(client: Client): Promise<boolean> { try { const controller = new AbortController(); const timeout = setTimeout(() => controller.abort(), HEALTH_TIMEOUT_MS); await client.ping({ signal: controller.signal }); clearTimeout(timeout); return true; } catch { return false; } } export { verifyConnection }; ``` **Why good:** `client.ping()` is the lightest health check, AbortController prevents hanging on unresponsive cluster, named constant for timeout --- ## Index Management ### Creating an Index with Explicit Mappings ```typescript import type { Client } from "@elastic/elasticsearch"; const INDEX_NAME = "products"; async function createProductIndex(client: Client): Promise<void> { const exists = await client.indices.exists({ index: INDEX_NAME }); if (exists) return; await client.indices.create({ index: INDEX_NAME, settings: { number_of_replicas: 1, refresh_interval: "1s", analysis: { analyzer: { product_analyzer: { type: "custom", tokenizer: "standard", filter: ["lowercase", "asciifolding"], }, }, }, }, mappings: { dynamic: "strict", // Reject documents with unmapped fields properties: { productId: { type: "keyword" }, name: { type: "text", analyzer: "product_analyzer", fields: { keyword: { type: "keyword", ignore_above: 256 } }, }, description: { type: "text", analyzer: "product_analyzer" }, price: { type: "float" }, categories: { type: "keyword" }, brand: { type: "keyword" }, inStock: { type: "boolean" }, tags: { type: "keyword" }, createdAt: { type: "date" }, }, }, }); } export { createProductIndex }; ``` **Why good:** `dynamic: "strict"` rejects unmapped fields (prevents mapping explosion), custom analyzer with asciifolding for accent-insensitive search, `text` + `keyword` multi-field on `name` for both search and aggregation, existence check prevents errors on re-run ### Adding Fields to Existing Mapping ```typescript // You CAN add new fields to an existing mapping // You CANNOT change the type of an existing field async function addRatingField(client: Client): Promise<void> { await client.indices.putMapping({ index: INDEX_NAME, properties: { rating: { type: "float" }, reviewCount: { type: "integer" }, }, }); } export { addRatingField }; ``` **Important:** `putMapping` can only ADD new fields. Changing an existing field type (e.g., `text` to `keyword`) requires creating a new index with correct mappings and reindexing all documents. ### Index Aliases for Zero-Downtime Reindexing ```typescript const PRODUCTS_READ_ALIAS = "products-read"; const PRODUCTS_WRITE_ALIAS = "products-write"; async function swapIndex( client: Client, oldIndex: string, newIndex: string, ): Promise<void> { await client.indices.updateAliases({ actions: [ { remove: { index: oldIndex, alias: PRODUCTS_READ_ALIAS } }, { add: { index: newIndex, alias: PRODUCTS_READ_ALIAS } }, { remove: { index: oldIndex, alias: PRODUCTS_WRITE_ALIAS } }, { add: { index: newIndex, alias: PRODUCTS_WRITE_ALIAS } }, ], }); } export { swapIndex }; ``` **Why good:** `updateAliases` is atomic -- read and write aliases switch simultaneously, no downtime during reindex --- ## Document Operations ### Indexing a Document ```typescript import type { Client } from "@elastic/elasticsearch"; interface Product { productId: string; name: string; description: string; price: number; categories: string[]; brand: string; inStock: boolean; createdAt: string; } const INDEX_NAME = "products"; async function indexProduct(client: Client, product: Product): Promise<string> { const result = await client.index({ index: INDEX_NAME, id: product.productId, // Explicit ID for upsert behavior document: product, }); return result._id; } export { indexProduct }; export type { Product }; ``` **Why good:** Explicit `id` enables upsert (index or replace), typed document, returns the document ID ### Getting a Document by ID ```typescript async function getProduct( client: Client, productId: string, ): Promise<Product | null> { try { const result = await client.get<Product>({ index: INDEX_NAME, id: productId, }); return result._source ?? null; } catch (err) { if ( err instanceof Error && "statusCode" in err && (err as { statusCode: number }).statusCode === 404 ) { return null; } throw err; } } export { getProduct }; ``` **Why good:** `_source` can be undefined, 404 handled gracefully (document not found is not an error in most use cases), generic type flows through to `_source` ### Partial Update ```typescript async function updateProductPrice( client: Client, productId: string, newPrice: number, ): Promise<void> { await client.update({ index: INDEX_NAME, id: productId, doc: { price: newPrice }, }); } export { updateProductPrice }; ``` **Why good:** `doc` performs partial update -- only `price` changes, all other fields preserved. Compare with `client.index()` which replaces the entire document. ### Scripted Update (Atomic) ```typescript const PRICE_INCREASE_PERCENTAGE = 10; async function increasePriceByPercent( client: Client, productId: string, ): Promise<void> { await client.update({ index: INDEX_NAME, id: productId, script: { source: "ctx._source.price *= (1 + params.pct / 100.0)", params: { pct: PRICE_INCREASE_PERCENTAGE }, }, }); } export { increasePriceByPercent }; ``` **Why good:** Script runs on the shard -- atomic, no read-modify-write race condition. Named constant for the percentage value. ### Delete by ID and by Query ```typescript async function deleteProduct(client: Client, productId: string): Promise<void> { await client.delete({ index: INDEX_NAME, id: productId, }); } async function deleteOutOfStockProducts(client: Client): Promise<number> { const result = await client.deleteByQuery({ index: INDEX_NAME, query: { term: { inStock: false }, }, }); return result.deleted ?? 0; } export { deleteProduct, deleteOutOfStockProducts }; ``` **Why good:** `deleteByQuery` for batch deletion without knowing individual IDs, returns count of deleted documents --- ## Search Patterns ### Basic Search with Highlighting ```typescript import type { Client, SearchResponse } from "@elastic/elasticsearch"; const DEFAULT_SEARCH_LIMIT = 20; async function searchProducts( client: Client, query: string, options?: { limit?: number }, ): Promise<SearchResponse<Product>> { return client.search<Product>({ index: INDEX_NAME, query: { multi_match: { query, fields: ["name^3", "description", "brand^2"], type: "best_fields", fuzziness: "AUTO", }, }, highlight: { fields: { name: {}, description: { fragment_size: 150, number_of_fragments: 2 }, }, pre_tags: ["<mark>"], post_tags: ["</mark>"], }, size: options?.limit ?? DEFAULT_SEARCH_LIMIT, }); } // Usage: // const results = await searchProducts(client, "wireless headphones"); // results.hits.hits[0]._source?.name -- original // results.hits.hits[0].highlight?.name?.[0] -- "<mark>wireless</mark> <mark>headphones</mark>" export { searchProducts }; ``` **Why good:** `name^3` boosts name matches 3x, `fuzziness: "AUTO"` handles typos, highlighting configured with custom tags, typed search response ### Bool Query with Filter Context ```typescript interface SearchFilters { minPrice?: number; maxPrice?: number; categories?: string[]; inStock?: boolean; query?: string; } async function filteredSearch( client: Client, filters: SearchFilters, ): Promise<SearchResponse<Product>> { const must: object[] = []; const filterClauses: object[] = []; if (filters.query) { must.push({ multi_match: { query: filters.query, fields: ["name^3", "description", "brand"], }, }); } if (filters.minPrice !== undefined || filters.maxPrice !== undefined) { filterClauses.push({ range: { price: { ...(filters.minPrice !== undefined ? { gte: filters.minPrice } : {}), ...(filters.maxPrice !== undefined ? { lte: filters.maxPrice } : {}), }, }, }); } if (filters.categories?.length) { filterClauses.push({ terms: { categories: filters.categories } }); } if (filters.inStock !== undefined) { filterClauses.push({ term: { inStock: filters.inStock } }); } return client.search<Product>({ index: INDEX_NAME, query: { bool: { ...(must.length > 0 ? { must } : {}), ...(filterClauses.length > 0 ? { filter: filterClauses } : {}), }, }, size: DEFAULT_SEARCH_LIMIT, }); } export { filteredSearch }; ``` **Why good:** Exact-match conditions in `filter` (no scoring, cached), full-text in `must` (scoring), dynamic query construction based on provided filters ### Nested Query ```typescript // For arrays of objects where field association matters // Mapping must use "nested" type, not default "object" interface ProductWithReviews extends Product { reviews: Array<{ author: string; rating: number; text: string }>; } async function searchByReview( client: Client, minRating: number, reviewText: string, ): Promise<SearchResponse<ProductWithReviews>> { return client.search<ProductWithReviews>({ index: INDEX_NAME, query: { nested: { path: "reviews", query: { bool: { must: [{ match: { "reviews.text": reviewText } }], filter: [{ range: { "reviews.rating": { gte: minRating } } }], }, }, inner_hits: { size: 3 }, // Return matching nested docs }, }, }); } export { searchByReview }; ``` **Why good:** `nested` query preserves field association within each review object (without `nested`, searching for "great" with rating >= 4 could match "great" from one review and rating 5 from a different review), `inner_hits` returns the matching nested documents **Important:** The `reviews` field must be mapped as `"type": "nested"`. Default `object` mapping flattens the array, losing field-level association. --- ## TypeScript Integration ### Typed Search Results ```typescript import type { Client, SearchHit } from "@elastic/elasticsearch"; // Generic type flows through to hits async function searchAndTransform( client: Client, query: string, ): Promise<Array<{ id: string; name: string; score: number }>> { const result = await client.search<Product>({ index: INDEX_NAME, query: { match: { name: query } }, size: 10, }); return result.hits.hits .filter( (hit): hit is SearchHit<Product> & { _source: Product } => hit._source !== undefined, ) .map((hit) => ({ id: hit._id, name: hit._source.name, score: hit._score ?? 0, })); } export { searchAndTransform }; ``` **Why good:** Type guard filters out hits with missing `_source`, `_score` can be null (e.g., in filter context), generic type parameter provides type safety on `_source` ### Importing Elasticsearch Types ```typescript import type { SearchResponse, SearchHit, BulkResponse, MappingProperty, } from "@elastic/elasticsearch/lib/api/types"; // Or use the estypes namespace for full type coverage import type { estypes } from "@elastic/elasticsearch"; type MySearchResponse = estypes.SearchResponse<Product>; ``` **Why good:** `estypes` provides complete request/response types matching the Elasticsearch specification --- ## Error Handling ### Common Error Patterns ```typescript import { errors } from "@elastic/elasticsearch"; async function safeSearch( client: Client, query: string, ): Promise<SearchResponse<Product> | null> { try { return await client.search<Product>({ index: INDEX_NAME, query: { match: { name: query } }, }); } catch (err) { if (err instanceof errors.ResponseError) { if (err.statusCode === 404) { // Index doesn't exist return null; } if (err.statusCode === 400) { // Bad query (malformed DSL) throw new Error(`Invalid search query: ${err.message}`); } } if (err instanceof errors.ConnectionError) { // Cluster unreachable throw new Error("Elasticsearch cluster unavailable"); } throw err; } } export { safeSearch }; ``` **Why good:** `errors.ResponseError` for HTTP errors with status codes, `errors.ConnectionError` for network issues, specific handling per status code --- _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_ -
pagination.md 7.7 KB
# Elasticsearch -- Pagination Examples > from/size, search_after, Point in Time (PIT), and scroll helper patterns. Reference from [SKILL.md](../SKILL.md). **Prerequisites:** Understand client setup and search basics from [core.md](core.md) first. **Related examples:** - [core.md](core.md) -- Client setup, search basics - [bulk-operations.md](bulk-operations.md) -- Bulk data export with scrollDocuments --- ## Pagination Decision Framework ``` How deep do results go? -- Under 10,000 total and users jump between pages? -> from/size -- Over 10,000 hits, forward-only (infinite scroll)? -> search_after -- Over 10,000 hits, need consistent snapshot? -> search_after + PIT -- Bulk data export (all documents)? -> client.helpers.scrollDocuments() ``` --- ## from/size (Shallow Pagination) Simple offset-based pagination. Limited to 10,000 total hits by default. ```typescript import type { Client, SearchResponse } from "@elastic/elasticsearch"; const INDEX_NAME = "products"; const PAGE_SIZE = 20; interface Product { productId: string; name: string; price: number; } async function getPage( client: Client, query: string, page: number, ): Promise<SearchResponse<Product>> { const from = (page - 1) * PAGE_SIZE; return client.search<Product>({ index: INDEX_NAME, query: { match: { name: query } }, from, size: PAGE_SIZE, track_total_hits: true, // Get exact total count }); } export { getPage }; ``` **Why good:** Simplest pagination, supports random page access (jump to page 5), `track_total_hits: true` for exact total count **Limitations:** - Default `index.max_result_window` is 10,000 -- `from + size` cannot exceed this - Deep pages are expensive: Elasticsearch must fetch and discard `from` documents on every shard - Not suitable for infinite scroll or large result sets --- ## search_after (Deep Pagination) Stateless cursor-based pagination using sort values. No 10,000 hit limit. ```typescript const PAGE_SIZE = 20; async function searchAfterPage( client: Client, query: string, lastSort?: Array<string | number>, ): Promise<{ hits: Product[]; nextSort: Array<string | number> | null; total: number; }> { const result = await client.search<Product>({ index: INDEX_NAME, query: { match: { name: query } }, sort: [ { _score: "desc" }, { "name.keyword": "asc" }, // Tiebreaker -- REQUIRED ], size: PAGE_SIZE, ...(lastSort ? { search_after: lastSort } : {}), track_total_hits: true, }); const hits = result.hits.hits; const nextSort = hits.length > 0 ? (hits[hits.length - 1].sort as Array<string | number>) : null; return { hits: hits.filter((h) => h._source !== undefined).map((h) => h._source!), nextSort, total: typeof result.hits.total === "number" ? result.hits.total : (result.hits.total?.value ?? 0), }; } export { searchAfterPage }; ``` **Why good:** No 10,000 hit limit, tiebreaker field prevents missing/duplicate documents, stateless (no server-side cursor to maintain) **Gotcha:** `search_after` requires a `sort` parameter. You MUST include a tiebreaker field (a unique value like `_id` or a keyword field) -- without it, documents with identical sort values may be skipped or duplicated across pages. **Gotcha:** `search_after` is forward-only. You cannot jump to page 5 directly -- you must iterate through pages 1-4 first. For random page access on small result sets, use `from`/`size`. --- ## search_after + Point in Time (Consistent Deep Pagination) PIT creates a snapshot of the index state, ensuring consistent results even as documents are added/updated during pagination. ```typescript const PIT_KEEP_ALIVE = "1m"; const PAGE_SIZE = 100; async function paginateAllProducts( client: Client, onPage: (products: Product[]) => Promise<void>, ): Promise<void> { const pit = await client.openPointInTime({ index: INDEX_NAME, keep_alive: PIT_KEEP_ALIVE, }); let searchAfter: Array<string | number> | undefined; try { while (true) { const result = await client.search<Product>({ pit: { id: pit.id, keep_alive: PIT_KEEP_ALIVE }, sort: [{ createdAt: "desc" }, { _id: "asc" }], size: PAGE_SIZE, ...(searchAfter ? { search_after: searchAfter } : {}), }); const hits = result.hits.hits; if (hits.length === 0) break; const products = hits .filter((h) => h._source !== undefined) .map((h) => h._source!); await onPage(products); searchAfter = hits[hits.length - 1].sort as Array<string | number>; } } finally { await client.closePointInTime({ id: pit.id }); } } export { paginateAllProducts }; ``` **Why good:** PIT ensures consistent snapshot across all pages (no missing/duplicate documents from concurrent writes), `keep_alive` refreshed on each request, PIT closed in `finally` block (prevents resource leak) **Key differences from plain search_after:** - PIT provides a frozen view of the index -- new documents added during pagination are not visible - Without PIT, `search_after` sees the live index -- concurrent writes can cause documents to shift between pages - PIT requires no `index` parameter in the search (it's bound to the PIT) **Gotcha:** PIT search automatically adds `_shard_doc` as an implicit tiebreaker. You can still add your own tiebreaker for deterministic ordering. --- ## scrollSearch Helper (Page-Level Iteration) The built-in helper handles scroll management automatically. Use for batch processing. ```typescript async function processAllDocuments( client: Client, onBatch: (docs: Product[]) => Promise<void>, ): Promise<void> { const scrollSearch = client.helpers.scrollSearch<Product>({ index: INDEX_NAME, query: { match_all: {} }, size: 500, }); for await (const result of scrollSearch) { const docs = result.documents; await onBatch(docs); } } export { processAllDocuments }; ``` **Why good:** Async iterator handles scroll lifecycle automatically (open, fetch, clear), `documents` property provides typed `_source` values directly **When to use:** One-time batch processing or data export. NOT for user-facing pagination. --- ## scrollDocuments Helper (Document-Level Iteration) Yields individual documents instead of pages. Most memory-efficient for large exports. ```typescript async function exportAllProducts(client: Client): Promise<Product[]> { const allProducts: Product[] = []; const docs = client.helpers.scrollDocuments<Product>({ index: INDEX_NAME, query: { match_all: {} }, }); for await (const doc of docs) { allProducts.push(doc); } return allProducts; } export { exportAllProducts }; ``` **Why good:** Most memory-efficient -- processes one document at a time, automatic filter_path optimization **When to use:** Large data exports, ETL pipelines, index migrations. The scroll API is NOT deprecated for this use case -- it's only deprecated for user-facing search pagination. --- ## Pagination Comparison | Method | Max Depth | Random Access | Consistency | Use Case | | --------------------- | --------- | ------------- | ----------- | -------------------------------- | | `from`/`size` | 10,000 | Yes | Live index | UI with page numbers | | `search_after` | Unlimited | No (forward) | Live index | Infinite scroll | | `search_after` + PIT | Unlimited | No (forward) | Snapshot | Deep pagination with consistency | | `scrollSearch` helper | Unlimited | No (forward) | Snapshot | Batch processing (pages) | | `scrollDocuments` | Unlimited | No (forward) | Snapshot | Batch processing (documents) | --- _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_ -
vector-search.md 9.3 KB
# Elasticsearch -- Vector Search Examples > Dense vector fields, kNN queries, hybrid text+vector search, and similarity metrics. Reference from [SKILL.md](../SKILL.md). **Prerequisites:** Understand client setup and index management from [core.md](core.md) first. **Related examples:** - [core.md](core.md) -- Client setup, index management, search basics - [aggregations.md](aggregations.md) -- Combining aggregations with vector search --- ## Dense Vector Mapping Configure a `dense_vector` field for kNN search. ```typescript import type { Client } from "@elastic/elasticsearch"; const INDEX_NAME = "articles"; const EMBEDDING_DIMS = 768; // Must match your embedding model's output dimension async function createVectorIndex(client: Client): Promise<void> { await client.indices.create({ index: INDEX_NAME, mappings: { properties: { title: { type: "text", fields: { keyword: { type: "keyword" } }, }, content: { type: "text" }, embedding: { type: "dense_vector", dims: EMBEDDING_DIMS, index: true, // Required for kNN search (default: true in 8.11+) similarity: "cosine", // cosine | l2_norm | dot_product | max_inner_product }, category: { type: "keyword" }, publishedAt: { type: "date" }, }, }, }); } export { createVectorIndex }; ``` **Why good:** `dims` matches the embedding model's output dimension exactly (768 for most BERT-based models), `similarity: "cosine"` is the most common metric, `index: true` enables approximate kNN search **Gotcha:** If `dims` does not match the actual vector length when indexing documents, Elasticsearch rejects the document with a `mapper_parsing_exception`. There is no auto-truncation or padding. --- ## Indexing Documents with Vectors ```typescript interface Article { title: string; content: string; embedding: number[]; category: string; publishedAt: string; } async function indexArticleWithVector( client: Client, article: Article, id: string, ): Promise<void> { await client.index({ index: INDEX_NAME, id, document: article, }); } export { indexArticleWithVector }; ``` **Important:** The `embedding` array MUST have exactly `EMBEDDING_DIMS` elements. Generate embeddings using your embedding model before indexing -- Elasticsearch does not generate embeddings for you (unless you configure an ingest pipeline with a deployed model). --- ## Basic kNN Search Find documents with vectors most similar to a query vector. ```typescript const KNN_CANDIDATES = 100; const KNN_RESULTS = 10; async function vectorSearch( client: Client, queryVector: number[], ): Promise<Array<{ id: string; title: string; score: number }>> { const result = await client.search<Article>({ index: INDEX_NAME, knn: { field: "embedding", query_vector: queryVector, k: KNN_RESULTS, num_candidates: KNN_CANDIDATES, }, }); return result.hits.hits .filter((hit) => hit._source !== undefined) .map((hit) => ({ id: hit._id, title: hit._source!.title, score: hit._score ?? 0, })); } export { vectorSearch }; ``` **Why good:** `num_candidates` controls the accuracy/speed tradeoff (higher = more accurate but slower), `k` is the number of results to return, named constants for both **Key concept:** `num_candidates` determines how many candidates each shard considers before the final `k` results are selected. A ratio of `num_candidates / k >= 10` is a good starting point. --- ## kNN with Filters Apply filters during the kNN search to restrict the vector space. ```typescript async function filteredVectorSearch( client: Client, queryVector: number[], category: string, ): Promise<Array<{ id: string; title: string; score: number }>> { const result = await client.search<Article>({ index: INDEX_NAME, knn: { field: "embedding", query_vector: queryVector, k: KNN_RESULTS, num_candidates: KNN_CANDIDATES, filter: { term: { category }, }, }, }); return result.hits.hits .filter((hit) => hit._source !== undefined) .map((hit) => ({ id: hit._id, title: hit._source!.title, score: hit._score ?? 0, })); } export { filteredVectorSearch }; ``` **Why good:** Filter is applied DURING the kNN search, not after -- this ensures exactly `k` results from the filtered subset, not fewer **Gotcha:** Without the `filter` inside `knn`, you would need to use a top-level `query` filter which is applied AFTER kNN. Post-filtering can return fewer than `k` results because some of the `k` nearest neighbors may not match the filter. --- ## Hybrid Search (Text + Vector) Combine traditional full-text search with vector similarity. ```typescript const TEXT_BOOST = 0.7; const VECTOR_BOOST = 0.3; async function hybridSearch( client: Client, queryText: string, queryVector: number[], ): Promise<Array<{ id: string; title: string; score: number }>> { const result = await client.search<Article>({ index: INDEX_NAME, query: { match: { content: { query: queryText, boost: TEXT_BOOST, }, }, }, knn: { field: "embedding", query_vector: queryVector, k: KNN_RESULTS, num_candidates: KNN_CANDIDATES, boost: VECTOR_BOOST, }, size: KNN_RESULTS, }); return result.hits.hits .filter((hit) => hit._source !== undefined) .map((hit) => ({ id: hit._id, title: hit._source!.title, score: hit._score ?? 0, })); } export { hybridSearch }; ``` **Why good:** Named constants for boost values, text and vector scores are combined with weighted boosts, `size` limits total results from both sources **How scoring works:** The final `_score` = `TEXT_BOOST * text_score + VECTOR_BOOST * knn_score`. Adjust boosts to control the balance between keyword relevance and semantic similarity. **Gotcha:** Results are combined via disjunction (OR) -- a document can appear from text match only, vector match only, or both. Documents matching both get a combined score. --- ## Semantic Search with Model Inference Let Elasticsearch generate the query vector using a deployed model. ```typescript const MODEL_ID = "my-text-embedding-model"; async function semanticSearch( client: Client, queryText: string, ): Promise<Array<{ id: string; title: string; score: number }>> { const result = await client.search<Article>({ index: INDEX_NAME, knn: { field: "embedding", k: KNN_RESULTS, num_candidates: KNN_CANDIDATES, query_vector_builder: { text_embedding: { model_id: MODEL_ID, model_text: queryText, }, }, }, }); return result.hits.hits .filter((hit) => hit._source !== undefined) .map((hit) => ({ id: hit._id, title: hit._source!.title, score: hit._score ?? 0, })); } export { semanticSearch }; ``` **Why good:** `query_vector_builder` generates the vector server-side using a deployed ML model -- no need to call the embedding API separately in your application code **When to use:** When you have a model deployed in Elasticsearch (via Eland or the trained models API). When embedding is done externally (e.g., OpenAI API), pass `query_vector` directly instead. --- ## Similarity Metrics | Metric | Use Case | Score Interpretation | | ------------------- | --------------------------------------- | ------------------------------------------- | | `cosine` | General purpose, normalized embeddings | -1 to 1 (mapped to 0-1 for `_score`) | | `dot_product` | Optimized for unit-length vectors | Unbounded (requires pre-normalized vectors) | | `l2_norm` | When absolute distance matters | 0 = identical, higher = more different | | `max_inner_product` | When negative inner products are needed | Unbounded | **Recommendation:** Use `cosine` for most embedding models (sentence-transformers, OpenAI, Cohere). Use `dot_product` only if you know your vectors are pre-normalized to unit length. **Gotcha:** `dot_product` on non-normalized vectors gives meaningless scores. Always normalize vectors before indexing if using `dot_product`. --- ## Quantization for Memory Efficiency Dense vectors can be quantized to reduce memory usage. ```typescript async function createQuantizedVectorIndex(client: Client): Promise<void> { await client.indices.create({ index: "articles-quantized", mappings: { properties: { embedding: { type: "dense_vector", dims: EMBEDDING_DIMS, similarity: "cosine", index: true, index_options: { type: "int8_hnsw", // 4x less memory than float32 }, }, }, }, }); } export { createQuantizedVectorIndex }; ``` **Why good:** `int8_hnsw` quantizes float32 vectors to int8, using ~4x less memory with minimal accuracy loss **Available quantization types:** - `hnsw` -- Full float32 precision (default) - `int8_hnsw` -- 8-bit quantization (~4x memory reduction) - `int4_hnsw` -- 4-bit quantization (~8x memory reduction, more accuracy loss) - `bbq_hnsw` -- Better binary quantization (auto-selected for dims >= 384 in 8.15+) --- _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_
-
-
reference.md 16.7 KB
# Elasticsearch Quick Reference > Search DSL, mapping types, aggregation reference, client methods, and decision frameworks. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples. --- ## Search DSL Quick Reference ### Query Types | Query Type | Purpose | Context | Example | | -------------- | ----------------------------------------- | -------- | ----------------------------------------------------------------- | | `match` | Full-text search (analyzed) | `must` | `{ match: { title: "search engine" } }` | | `multi_match` | Full-text across multiple fields | `must` | `{ multi_match: { query: "search", fields: ["title", "body"] } }` | | `match_phrase` | Exact phrase in order | `must` | `{ match_phrase: { title: "search engine" } }` | | `term` | Exact value (NOT analyzed) | `filter` | `{ term: { status: "published" } }` | | `terms` | Match any of multiple exact values | `filter` | `{ terms: { status: ["published", "draft"] } }` | | `range` | Numeric/date range | `filter` | `{ range: { price: { gte: 10, lte: 100 } } }` | | `exists` | Field exists and is not null | `filter` | `{ exists: { field: "description" } }` | | `bool` | Combine queries with must/should/filter | any | `{ bool: { must: [...], filter: [...] } }` | | `nested` | Query nested objects preserving structure | `must` | `{ nested: { path: "comments", query: { ... } } }` | | `wildcard` | Wildcard pattern match | `filter` | `{ wildcard: { name: { value: "elast*" } } }` | | `fuzzy` | Approximate string match | `must` | `{ fuzzy: { name: { value: "serch", fuzziness: "AUTO" } } }` | | `prefix` | Prefix match | `filter` | `{ prefix: { name: { value: "elast" } } }` | | `ids` | Match by document IDs | `filter` | `{ ids: { values: ["1", "2", "3"] } }` | ### Bool Query Structure ``` bool: must: [ ... ] # Must match, contributes to score should: [ ... ] # Optional match, boosts score (OR logic unless minimum_should_match set) filter: [ ... ] # Must match, NO scoring (cached, fast) must_not: [ ... ] # Must NOT match, NO scoring ``` **Key rule:** Use `filter` for yes/no conditions (term, range, exists). Use `must` only when relevance scoring matters (match, multi_match). --- ## Mapping Types | Type | Purpose | Searchable? | Aggregatable? | Sortable? | | -------------- | --------------------------------- | ----------- | ------------- | --------- | | `text` | Full-text search (analyzed) | Yes | No\* | No\* | | `keyword` | Exact values, filtering, sorting | Yes (exact) | Yes | Yes | | `integer` | 32-bit integer | Yes | Yes | Yes | | `long` | 64-bit integer | Yes | Yes | Yes | | `float` | 32-bit floating point | Yes | Yes | Yes | | `double` | 64-bit floating point | Yes | Yes | Yes | | `boolean` | true/false | Yes | Yes | Yes | | `date` | ISO 8601 dates or epoch millis | Yes | Yes | Yes | | `object` | JSON object (flattened) | Yes | Depends | No | | `nested` | JSON object (preserves structure) | Yes | Yes (nested) | No | | `geo_point` | Latitude/longitude pair | Geo queries | Yes | Geo sort | | `dense_vector` | Float vector for kNN search | kNN queries | No | No | | `ip` | IPv4/IPv6 address | Yes | Yes | Yes | | `completion` | Autocomplete suggestions | Suggest API | No | No | \*`text` fields cannot be aggregated or sorted directly. Use a `.keyword` sub-field for aggregation/sort. ### Multi-Field Pattern The most common pattern: `text` for search, `keyword` sub-field for aggregation/filtering. ```json { "name": { "type": "text", "analyzer": "standard", "fields": { "keyword": { "type": "keyword", "ignore_above": 256 } } } } ``` - Search: `{ "match": { "name": "wireless headphones" } }` - Aggregate: `{ "terms": { "field": "name.keyword" } }` - Sort: `{ "sort": [{ "name.keyword": "asc" }] }` --- ## Aggregation Types ### Bucket Aggregations (grouping) | Aggregation | Purpose | Example | | ---------------- | ------------------------------- | ---------------------------------------------------------------------- | | `terms` | Group by exact field values | `{ terms: { field: "status", size: 100 } }` | | `range` | Custom numeric ranges | `{ range: { field: "price", ranges: [{ to: 50 }, ...] } }` | | `date_histogram` | Time-based buckets | `{ date_histogram: { field: "created", calendar_interval: "month" } }` | | `histogram` | Numeric interval buckets | `{ histogram: { field: "price", interval: 10 } }` | | `filter` | Single filter bucket | `{ filter: { term: { status: "active" } } }` | | `filters` | Named filter buckets | `{ filters: { filters: { active: ..., inactive: ... } } }` | | `nested` | Aggregate within nested objects | `{ nested: { path: "comments" } }` | ### Metric Aggregations (calculations) | Aggregation | Purpose | Example | | ------------- | -------------------------- | ---------------------------------------- | | `avg` | Average value | `{ avg: { field: "price" } }` | | `sum` | Total sum | `{ sum: { field: "quantity" } }` | | `min` / `max` | Minimum / Maximum | `{ min: { field: "price" } }` | | `value_count` | Count of values | `{ value_count: { field: "price" } }` | | `stats` | min/max/avg/sum/count | `{ stats: { field: "price" } }` | | `cardinality` | Approximate distinct count | `{ cardinality: { field: "userId" } }` | | `percentiles` | Distribution percentiles | `{ percentiles: { field: "latency" } }` | | `top_hits` | Top documents per bucket | `{ top_hits: { size: 3, sort: [...] } }` | ### Pipeline Aggregations (calculations on other aggs) | Aggregation | Purpose | buckets_path Example | | ---------------- | --------------------------------------------- | ----------------------------- | | `avg_bucket` | Average across sibling buckets | `"monthly_sales>total_sales"` | | `sum_bucket` | Sum across sibling buckets | `"monthly_sales>total_sales"` | | `max_bucket` | Max across sibling buckets | `"monthly_sales>total_sales"` | | `derivative` | Rate of change between buckets | `"total_sales"` | | `cumulative_sum` | Running total across buckets | `"total_sales"` | | `moving_fn` | Moving function across buckets (script-based) | `"total_sales"` + `script` | | `bucket_sort` | Sort parent buckets by sub-agg value | N/A (uses `sort` param) | --- ## Client API Quick Reference ### Client | Method | Returns | Description | | -------------------------- | -------------------------- | ---------------------------------- | | `search<T>(params)` | `SearchResponse<T>` | Search documents | | `index(params)` | `IndexResponse` | Index single document | | `get<T>(params)` | `GetResponse<T>` | Get document by ID | | `update(params)` | `UpdateResponse` | Partial update document | | `delete(params)` | `DeleteResponse` | Delete document by ID | | `bulk(params)` | `BulkResponse` | Batch index/update/delete | | `deleteByQuery(params)` | `DeleteByQueryResponse` | Delete matching documents | | `updateByQuery(params)` | `UpdateByQueryResponse` | Update matching documents | | `mget(params)` | `MgetResponse<T>` | Get multiple documents by IDs | | `msearch(params)` | `MsearchResponse` | Multi-search in single request | | `openPointInTime(params)` | `OpenPointInTimeResponse` | Open PIT for consistent pagination | | `closePointInTime(params)` | `ClosePointInTimeResponse` | Close PIT (release resources) | ### Indices | Method | Returns | Description | | ------------------------------- | ---------------------- | --------------------------------- | | `indices.create(params)` | `CreateIndexResponse` | Create index with mappings | | `indices.delete(params)` | `AcknowledgedResponse` | Delete index | | `indices.exists(params)` | `boolean` | Check if index exists | | `indices.putMapping(params)` | `AcknowledgedResponse` | Add new fields to mapping | | `indices.getMapping(params)` | `GetMappingResponse` | Get current mappings | | `indices.putSettings(params)` | `AcknowledgedResponse` | Update index settings | | `indices.refresh(params)` | `RefreshResponse` | Force refresh (make docs visible) | | `indices.putAlias(params)` | `AcknowledgedResponse` | Create index alias | | `indices.deleteAlias(params)` | `AcknowledgedResponse` | Delete index alias | | `indices.updateAliases(params)` | `AcknowledgedResponse` | Atomic alias swap | ### Helpers | Method | Returns | Description | | --------------------------------- | ------------------ | ----------------------------------- | | `helpers.bulk(params)` | `BulkStats` | Bulk with batching + retries | | `helpers.search(params)` | `T[]` | Returns only `_source` docs | | `helpers.scrollSearch(params)` | `AsyncIterable` | Scroll with async iteration (pages) | | `helpers.scrollDocuments(params)` | `AsyncIterable<T>` | Scroll with async iteration (docs) | | `helpers.msearch()` | `MsearchHelper` | Batched multi-search | --- ## Search Parameters | Parameter | Type | Default | Description | | ------------------ | --------------- | --------- | ------------------------------------------------------- | | `index` | `string` | required | Index name or pattern | | `query` | `object` | match_all | Query DSL | | `size` | `number` | `10` | Number of hits to return | | `from` | `number` | `0` | Offset for pagination (max 10,000 by default) | | `sort` | `array` | `_score` | Sort order | | `_source` | `boolean/array` | `true` | Fields to include in `_source` | | `aggs` | `object` | none | Aggregations | | `highlight` | `object` | none | Highlight matching terms | | `search_after` | `array` | none | Cursor for deep pagination | | `pit` | `object` | none | Point in Time for consistent pagination | | `knn` | `object` | none | kNN vector search | | `track_total_hits` | `bool/number` | `10000` | Track exact total hits (true = exact, number = up to N) | | `timeout` | `string` | none | Search timeout (e.g., "5s") | | `explain` | `boolean` | `false` | Include score explanation | --- ## Common Index Settings | Setting | Default | Description | | ---------------------------------- | ------- | -------------------------------------------------- | | `index.number_of_shards` | `1` | Primary shards (set at creation, immutable) | | `index.number_of_replicas` | `1` | Replica shards (can be changed dynamically) | | `index.refresh_interval` | `"1s"` | How often new docs become searchable | | `index.max_result_window` | `10000` | Max `from + size` for pagination | | `index.mapping.total_fields.limit` | `1000` | Max fields in mapping (prevents mapping explosion) | --- ## Anti-Patterns ### Dynamic Mapping Without Explicit Types ```typescript // ANTI-PATTERN: No mappings defined await client.index({ index: "events", document: { timestamp: "2024-01-15" }, // Mapped as "date" }); // Later: await client.index({ index: "events", document: { timestamp: "not-a-date" }, // FAILS: mapper_parsing_exception }); ``` **Why it's wrong:** Dynamic mapping inferred `timestamp` as `date` from the first document. The second document with a non-date string fails permanently. You cannot change the mapping -- you must reindex. **What to do instead:** Define explicit mappings before indexing any documents. --- ### Using term Query on text Fields ```typescript // ANTI-PATTERN: term query on analyzed field const result = await client.search({ index: "products", query: { term: { name: "Wireless Headphones" } }, // Returns 0 results! }); ``` **Why it's wrong:** `text` fields are analyzed -- "Wireless Headphones" is stored as tokens ["wireless", "headphones"]. The `term` query does NOT analyze the input, so it looks for the exact un-analyzed string "Wireless Headphones" which doesn't exist. **What to do instead:** Use `match` for text fields, or query the `.keyword` sub-field with `term`. --- ### Refresh on Every Write ```typescript // ANTI-PATTERN: Forcing refresh in request handler app.post("/products", async (req, res) => { await client.index({ index: "products", document: req.body, refresh: true, // Forces segment refresh on EVERY write }); res.json({ success: true }); }); ``` **Why it's wrong:** Each refresh creates a new Lucene segment. Under write-heavy load, this creates thousands of tiny segments that must be merged, consuming CPU and I/O. The default 1-second refresh interval batches writes into efficient segments. **What to do instead:** Let the default refresh interval handle it. Use `refresh: "wait_for"` only in tests and seed scripts. --- ## Production Checklist ### Mappings - [ ] Explicit mappings defined for all fields before first document - [ ] `text` + `keyword` multi-field on fields needing both search and aggregation - [ ] `nested` type used for arrays of objects that need field-level query association - [ ] `dynamic: "strict"` or `dynamic: false` to prevent unexpected field additions ### Performance - [ ] `filter` context used for all non-scoring queries (term, range, exists) - [ ] `size: 0` when only aggregations are needed - [ ] Bulk API used for all batch operations - [ ] No `refresh: true` in production request handlers - [ ] `_source` filtering used to return only needed fields ### Pagination - [ ] `search_after` + PIT used for results beyond 10,000 - [ ] Tiebreaker field included in sort for `search_after` - [ ] PIT closed in finally block after use ### Reliability - [ ] Bulk response `errors` flag checked for partial failures - [ ] `_source` checked for undefined (can be missing if disabled) - [ ] Connection error handling with retries - [ ] Health check endpoint using `client.ping()` or `client.info()` --- _Full skill documentation: [SKILL.md](SKILL.md) | Examples: [examples/](examples/)_ -
SKILL.md 19.4 KB
--- name: api-search-elasticsearch description: Elasticsearch patterns -- client setup, index management, search DSL, aggregations, vector search, bulk operations, deep pagination --- # Elasticsearch Patterns > **Quick Guide:** Use `@elastic/elasticsearch` (v8.x/v9.x) as the TypeScript client. Elasticsearch is **near real-time** -- documents are NOT searchable immediately after indexing; they become visible after a refresh (default: every 1 second on active indices). You MUST define explicit mappings before indexing -- dynamic mapping infers types from the first document, and mismatched types in later documents cause hard failures you cannot fix without reindexing. Use `search_after` + Point in Time (PIT) for deep pagination -- NOT `from`/`size` beyond 10,000 hits and NOT the scroll API (deprecated for search). Use the `bulk` API or `client.helpers.bulk()` for any batch operation -- never loop individual index calls. --- <critical_requirements> ## CRITICAL: Before Using This Skill > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST define explicit index mappings BEFORE indexing documents -- dynamic mapping infers types from the first document, and if a later document sends a different type for the same field, indexing fails with a `mapper_parsing_exception` that CANNOT be fixed without reindexing into a new index)** **(You MUST use the `bulk` API for batch operations -- looping individual `client.index()` calls is orders of magnitude slower and can overwhelm the cluster with HTTP connections)** **(You MUST NOT use `from`/`size` pagination beyond 10,000 results -- Elasticsearch throws `Result window is too large` by default; use `search_after` + PIT instead)** **(You MUST NOT use `refresh: true` or `refresh: "wait_for"` in production request handlers -- forcing a refresh on every write degrades cluster performance; let the default 1-second refresh interval handle it)** </critical_requirements> --- ## Examples - [Core Patterns](examples/core.md) -- Client setup, index management, document CRUD, search basics, TypeScript integration - [Aggregations](examples/aggregations.md) -- Terms, range, date_histogram, nested, pipeline aggregations - [Vector Search](examples/vector-search.md) -- Dense vector fields, kNN queries, hybrid search, similarity metrics - [Pagination](examples/pagination.md) -- from/size, search_after, Point in Time, scroll helpers - [Bulk Operations](examples/bulk-operations.md) -- Bulk API, bulk helper, reindexing patterns **Additional resources:** - [reference.md](reference.md) -- Search DSL cheat sheet, mapping types, aggregation reference, decision frameworks --- **Auto-detection:** Elasticsearch, elasticsearch, @elastic/elasticsearch, client.search, client.index, client.bulk, client.indices.create, client.indices.putMapping, dense_vector, knn, search_after, point in time, openPIT, aggregations, aggs, bool query, match query, term query, multi_match, nested query, range query, client.helpers.bulk, client.helpers.scrollSearch, BulkResponse, SearchResponse, MappingProperty **When to use:** - Full-text search with advanced relevance tuning (BM25, custom analyzers, boosting) - Aggregations and analytics (terms, histograms, pipeline aggregations) - Vector/semantic search with kNN on dense_vector fields - Log and event data search with time-based queries - Complex structured queries combining bool, nested, range, and geo filters - Search across large datasets requiring deep pagination (search_after + PIT) **Key patterns covered:** - Client initialization and connection management - Index management with explicit mappings and settings - Document CRUD (index, get, update, delete) - Search DSL (match, term, bool, range, nested, multi_match) - Aggregations (terms, range, date_histogram, nested, pipeline) - Full-text analysis (custom analyzers, tokenizers, filters) - Vector search (dense_vector, kNN, hybrid text+vector) - Bulk operations and reindexing - Deep pagination (search_after + PIT) **When NOT to use:** - Simple keyword search on small datasets (client-side filtering or database LIKE queries are simpler) - Primary data store (Elasticsearch is a search engine, not a database -- always have a source of truth elsewhere) - Strong consistency requirements (Elasticsearch is eventually consistent by design) - Simple autocomplete on a small list (a prefix trie or client-side filter is simpler) --- <philosophy> ## Philosophy Elasticsearch is a **distributed search and analytics engine** built on Apache Lucene. It excels at full-text search, structured queries, aggregations, and vector search at scale. Core principles: 1. **Near real-time, not real-time** -- Documents are indexed into segments. A refresh (default: every 1 second on active indices) makes new segments searchable. Do not expect immediate consistency after writes. 2. **Mappings are immutable** -- Once a field type is set (text, keyword, integer, etc.), it cannot be changed. Wrong types require reindexing into a new index. Always define mappings explicitly before first document. 3. **Search engine, not database** -- Elasticsearch should not be your source of truth. Always have a primary database and sync to Elasticsearch for search. 4. **Bulk everything** -- The bulk API amortizes HTTP overhead across thousands of operations. Never loop individual index/update/delete calls. 5. **Pagination has limits** -- `from`/`size` is capped at 10,000 hits by default (`index.max_result_window`). Deep pagination requires `search_after` + Point in Time (PIT). The scroll API is deprecated for search use cases. 6. **Text vs keyword matters** -- `text` fields are analyzed (tokenized, lowercased) for full-text search. `keyword` fields are exact-match only. Getting this wrong means either broken search or broken aggregations/filters. </philosophy> --- <patterns> ## Core Patterns ### Pattern 1: Client Setup Initialize the client with node URL and authentication. The client supports Elastic Cloud, API keys, basic auth, and bearer tokens. ```typescript // Good Example -- Typed client setup with environment validation import { Client } from "@elastic/elasticsearch"; function createElasticsearchClient(): Client { const node = process.env.ELASTICSEARCH_URL; if (!node) { throw new Error("ELASTICSEARCH_URL environment variable is required"); } return new Client({ node, auth: { apiKey: process.env.ELASTICSEARCH_API_KEY ?? "", }, }); } export { createElasticsearchClient }; ``` **Why good:** Environment variable validation, named export, API key auth (preferred over basic auth in production) ```typescript // Bad Example -- Hardcoded credentials import { Client } from "@elastic/elasticsearch"; const client = new Client({ node: "http://localhost:9200", auth: { username: "elastic", password: "changeme" }, }); ``` **Why bad:** Hardcoded node URL and credentials leak in version control, basic auth with default password See [examples/core.md](examples/core.md) for Elastic Cloud setup, health checks, and child clients. --- ### Pattern 2: Index with Explicit Mappings Always define mappings before indexing. Dynamic mapping infers types from the first document -- if wrong, you must reindex. ```typescript // Good Example -- Explicit mappings with text + keyword multi-field const INDEX_NAME = "products"; await client.indices.create({ index: INDEX_NAME, settings: { number_of_replicas: 1, refresh_interval: "1s", }, mappings: { properties: { name: { type: "text", fields: { keyword: { type: "keyword" } }, }, description: { type: "text", analyzer: "standard" }, price: { type: "float" }, categories: { type: "keyword" }, inStock: { type: "boolean" }, createdAt: { type: "date" }, }, }, }); ``` **Why good:** Explicit types prevent mapping conflicts, `text` + `keyword` multi-field allows both full-text search and exact-match filtering/aggregation on `name` ```typescript // Bad Example -- No mappings, relying on dynamic mapping await client.indices.create({ index: "products" }); await client.index({ index: "products", document: { price: "29.99" }, // Oops -- "29.99" is a string, mapped as text }); // All future numeric price documents will fail with mapper_parsing_exception ``` **Why bad:** Dynamic mapping infers `price` as `text` from the string "29.99", and this mapping is immutable -- all future documents with numeric `price` will fail See [examples/core.md](examples/core.md) for analysis settings, custom analyzers, and mapping migration. --- ### Pattern 3: Search with Bool Query The bool query is the workhorse of Elasticsearch. It combines `must`, `should`, `must_not`, and `filter` clauses. ```typescript // Good Example -- Bool query with filter context for exact matches const MIN_PRICE = 10; const MAX_PRICE = 100; const result = await client.search<Product>({ index: INDEX_NAME, query: { bool: { must: [{ match: { description: "wireless headphones" } }], filter: [ { range: { price: { gte: MIN_PRICE, lte: MAX_PRICE } } }, { term: { inStock: true } }, ], }, }, size: 20, }); // result.hits.hits[0]._source is typed as Product | undefined ``` **Why good:** `filter` context for exact matches (no scoring overhead, cacheable), `must` for full-text relevance scoring, named constants for range values, typed search with generic ```typescript // Bad Example -- Everything in must (no filter context) const result = await client.search({ index: "products", query: { bool: { must: [ { match: { description: "wireless headphones" } }, { range: { price: { gte: 10, lte: 100 } } }, // Wasteful scoring { term: { inStock: true } }, // Wasteful scoring ], }, }, }); ``` **Why bad:** Range and term queries in `must` waste CPU on relevance scoring for yes/no conditions; `filter` context skips scoring and enables Elasticsearch's filter cache See [examples/core.md](examples/core.md) for multi_match, nested queries, and function_score. --- ### Pattern 4: Aggregations Aggregations compute analytics over search results. `terms` for category counts, `range` for bucketing, `date_histogram` for time series. ```typescript // Good Example -- Terms aggregation with sub-aggregation const AGGREGATION_SIZE = 50; const result = await client.search({ index: INDEX_NAME, size: 0, // No hits needed, only aggregations aggs: { categories: { terms: { field: "categories", size: AGGREGATION_SIZE }, aggs: { avgPrice: { avg: { field: "price" } }, }, }, }, }); // result.aggregations?.categories.buckets -> [{ key: "electronics", doc_count: 42, avgPrice: { value: 89.5 } }] ``` **Why good:** `size: 0` skips hits when only aggregations are needed (faster), nested sub-aggregation for metrics per bucket, named constant for aggregation size See [examples/aggregations.md](examples/aggregations.md) for date_histogram, range, nested, and pipeline aggregations. --- ### Pattern 5: Bulk Operations The bulk API batches multiple index/update/delete operations in a single request. Use `client.helpers.bulk()` for the best developer experience. ```typescript // Good Example -- Bulk helper with async generator const result = await client.helpers.bulk<Product>({ datasource: products, onDocument(doc) { return { index: { _index: INDEX_NAME, _id: doc.productId } }; }, refreshOnCompletion: INDEX_NAME, }); // result.total, result.successful, result.failed ``` **Why good:** Bulk helper handles batching, concurrency, retries, and back-pressure automatically; `refreshOnCompletion` triggers one refresh at the end instead of per-document ```typescript // Bad Example -- Looping individual index calls for (const product of products) { await client.index({ index: "products", document: product, refresh: true, // Refresh after EVERY document! }); } ``` **Why bad:** N individual HTTP requests instead of 1 bulk request, `refresh: true` on every document causes N segment refreshes (devastating to cluster performance) See [examples/bulk-operations.md](examples/bulk-operations.md) for error handling, update operations, and reindexing. --- ### Pattern 6: Deep Pagination with search_after + PIT `from`/`size` is limited to 10,000 hits. For deep pagination, use `search_after` with a Point in Time (PIT) for consistent results. ```typescript // Good Example -- search_after with PIT const PIT_KEEP_ALIVE = "1m"; const pit = await client.openPointInTime({ index: INDEX_NAME, keep_alive: PIT_KEEP_ALIVE, }); let searchAfter: Array<string | number> | undefined; let allHits: Product[] = []; while (true) { const result = await client.search<Product>({ pit: { id: pit.id, keep_alive: PIT_KEEP_ALIVE }, sort: [{ createdAt: "desc" }, { _id: "asc" }], // Tiebreaker! size: 100, ...(searchAfter ? { search_after: searchAfter } : {}), }); const hits = result.hits.hits; if (hits.length === 0) break; allHits = allHits.concat( hits.filter((h) => h._source !== undefined).map((h) => h._source!), ); searchAfter = hits[hits.length - 1].sort as Array<string | number>; } await client.closePointInTime({ id: pit.id }); ``` **Why good:** PIT ensures consistent snapshot across pages, tiebreaker `_id` prevents missing/duplicate documents, `keep_alive` refreshed on each request See [examples/pagination.md](examples/pagination.md) for from/size limits, scroll helpers, and pagination decision framework. </patterns> --- <decision_framework> ## Decision Framework ### Which Query Type? ``` What kind of search do I need? -- Full-text relevance search? -> match / multi_match in must -- Exact value filtering? -> term / terms / range in filter -- Combining text + filters? -> bool query (must for text, filter for exact) -- Fuzzy matching? -> match with fuzziness: "AUTO" -- Phrase matching? -> match_phrase -- Complex nested objects? -> nested query with path -- Vector similarity? -> knn with dense_vector field -- Text + vector hybrid? -> query + knn in same request ``` ### Pagination Strategy? ``` How deep do results go? -- Under 10,000 total? -> from/size (simplest) -- Over 10,000 hits? -> search_after + PIT (recommended) -- Bulk data export? -> client.helpers.scrollSearch() or scrollDocuments() -- Real-time infinite scroll? -> search_after (no PIT needed for forward-only) ``` ### text vs keyword? ``` What will I do with this field? -- Full-text search (tokenized, relevance)? -> text -- Exact match, filtering, aggregations, sorting? -> keyword -- Both? -> Multi-field: { type: "text", fields: { keyword: { type: "keyword" } } } -- Neither (just stored, never queried)? -> { type: "keyword", index: false } ``` ### Filter vs Must? ``` Does relevance scoring matter for this clause? -- YES (affects result order) -> must -- NO (binary yes/no filter) -> filter (cached, no scoring overhead) -- Exclude documents -> must_not (in filter context) -- Boost if present (optional) -> should with minimum_should_match: 0 ``` </decision_framework> --- <red_flags> ## RED FLAGS **High Priority Issues:** - Relying on dynamic mapping without explicit mappings -- wrong type inference causes `mapper_parsing_exception` that requires reindexing to fix - Using `from`/`size` beyond 10,000 results -- Elasticsearch throws `Result window is too large`; use `search_after` + PIT - Looping individual `client.index()` calls instead of `client.bulk()` or `client.helpers.bulk()` -- orders of magnitude slower, can overwhelm the cluster - Using `refresh: true` or `refresh: "wait_for"` in production request handlers -- forces a segment refresh on every write, degrades cluster performance under load **Medium Priority Issues:** - Putting exact-match conditions (term, range) in `must` instead of `filter` -- wastes CPU on scoring, misses filter cache - Using `text` type for fields that need exact matching or aggregation -- text fields are analyzed (tokenized), making aggregations return individual tokens instead of full values - Not including a tiebreaker field in `sort` when using `search_after` -- documents with identical sort values may be skipped or duplicated across pages - Missing `_source` check -- `hit._source` can be `undefined` if `_source` is disabled or fields are excluded; always handle this **Gotchas & Edge Cases:** - **Mapping types are immutable** -- once a field is mapped as `text`, you cannot change it to `keyword`. The only fix is to create a new index with correct mappings and reindex all documents - **`text` vs `keyword` confusion** -- `text` fields are tokenized ("New York" becomes ["new", "york"]). Aggregating on a `text` field gives you individual tokens, not full values. Use `keyword` or a `.keyword` sub-field for aggregations - **`match` vs `term` on text fields** -- `term` on a `text` field often returns no results because `term` does NOT analyze the query but the field value IS analyzed (e.g., term "New York" won't match the analyzed tokens "new" and "york") - **Near real-time delay** -- after indexing, documents are NOT searchable until the next refresh (default: 1 second). Tests that index then immediately search must use `refresh: "wait_for"` or explicit `client.indices.refresh()` - **Default `index.max_result_window` is 10,000** -- increasing this is possible but NOT recommended; deep pagination with `from`/`size` holds all skipped results in memory - **Nested objects require `nested` mapping type** -- arrays of objects are flattened by default, losing the association between fields within each object. If you need to query "color: red AND size: large" on the same object in an array, use `nested` - **Aggregation on `_id` is disabled by default** (8.x+) -- use a separate `id` field if you need to aggregate by document ID - **`_score` is null in filter context** -- clauses in `filter` do not contribute to scoring; if you need scoring, use `must` - **Bulk API partial failures** -- a bulk request can succeed overall but have individual failures. Always check `result.errors` and iterate `result.items` to find failed operations - **Scroll API is deprecated for search** -- use `search_after` + PIT for deep pagination. Scroll is still valid for one-time data export but consumes cluster resources (open search contexts) - **PIT must be closed** -- failing to close Point in Time contexts leaks resources on the cluster; always close in a finally block </red_flags> --- <critical_reminders> ## CRITICAL REMINDERS > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST define explicit index mappings BEFORE indexing documents -- dynamic mapping infers types from the first document, and if a later document sends a different type for the same field, indexing fails with a `mapper_parsing_exception` that CANNOT be fixed without reindexing into a new index)** **(You MUST use the `bulk` API for batch operations -- looping individual `client.index()` calls is orders of magnitude slower and can overwhelm the cluster with HTTP connections)** **(You MUST NOT use `from`/`size` pagination beyond 10,000 results -- Elasticsearch throws `Result window is too large` by default; use `search_after` + PIT instead)** **(You MUST NOT use `refresh: true` or `refresh: "wait_for"` in production request handlers -- forcing a refresh on every write degrades cluster performance; let the default 1-second refresh interval handle it)** **Failure to follow these rules will cause mapping conflicts, pagination failures, cluster performance degradation, and silent data loss.** </critical_reminders>
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.