Claude Skill

api-search-elasticsearch

Elasticsearch patterns -- client setup, index management, search DSL, aggregations, vector search, bulk operations, deep pagination

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download agents-inc-skills-dist_plugins_api-search-elasticsearch_skills_api-search-elasticsearch-3a51ef5.zip · 28 KB
Part of agents-inc/skills — 130 skills

Install

skills CLI npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/api-search-elasticsearch/skills/api-search-elasticsearch
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
Git git clone https://github.com/agents-inc/skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Elasticsearch Patterns

Quick Guide: Use @elastic/elasticsearch (v8.x/v9.x) as the TypeScript client. Elasticsearch is near real-time -- documents are NOT searchable immediately after indexing; they become visible after a refresh (default: every 1 second on active indices). You MUST define explicit mappings before indexing -- dynamic mapping infers types from the first document, and mismatched types in later documents cause hard failures you cannot fix without reindexing. Use search_after + Point in Time (PIT) for deep pagination -- NOT from/size beyond 10,000 hits and NOT the scroll API (deprecated for search). Use the bulk API or client.helpers.bulk() for any batch operation -- never loop individual index calls.


<critical_requirements>

CRITICAL: Before Using This Skill

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST define explicit index mappings BEFORE indexing documents -- dynamic mapping infers types from the first document, and if a later document sends a different type for the same field, indexing fails with a mapper_parsing_exception that CANNOT be fixed without reindexing into a new index)

(You MUST use the bulk API for batch operations -- looping individual client.index() calls is orders of magnitude slower and can overwhelm the cluster with HTTP connections)

(You MUST NOT use from/size pagination beyond 10,000 results -- Elasticsearch throws Result window is too large by default; use search_after + PIT instead)

(You MUST NOT use refresh: true or refresh: "wait_for" in production request handlers -- forcing a refresh on every write degrades cluster performance; let the default 1-second refresh interval handle it)

</critical_requirements>


Examples

  • Core Patterns -- Client setup, index management, document CRUD, search basics, TypeScript integration
  • Aggregations -- Terms, range, date_histogram, nested, pipeline aggregations
  • Vector Search -- Dense vector fields, kNN queries, hybrid search, similarity metrics
  • Pagination -- from/size, search_after, Point in Time, scroll helpers
  • Bulk Operations -- Bulk API, bulk helper, reindexing patterns

Additional resources:

  • reference.md -- Search DSL cheat sheet, mapping types, aggregation reference, decision frameworks

Auto-detection: Elasticsearch, elasticsearch, @elastic/elasticsearch, client.search, client.index, client.bulk, client.indices.create, client.indices.putMapping, dense_vector, knn, search_after, point in time, openPIT, aggregations, aggs, bool query, match query, term query, multi_match, nested query, range query, client.helpers.bulk, client.helpers.scrollSearch, BulkResponse, SearchResponse, MappingProperty

When to use:

  • Full-text search with advanced relevance tuning (BM25, custom analyzers, boosting)
  • Aggregations and analytics (terms, histograms, pipeline aggregations)
  • Vector/semantic search with kNN on dense_vector fields
  • Log and event data search with time-based queries
  • Complex structured queries combining bool, nested, range, and geo filters
  • Search across large datasets requiring deep pagination (search_after + PIT)

Key patterns covered:

  • Client initialization and connection management
  • Index management with explicit mappings and settings
  • Document CRUD (index, get, update, delete)
  • Search DSL (match, term, bool, range, nested, multi_match)
  • Aggregations (terms, range, date_histogram, nested, pipeline)
  • Full-text analysis (custom analyzers, tokenizers, filters)
  • Vector search (dense_vector, kNN, hybrid text+vector)
  • Bulk operations and reindexing
  • Deep pagination (search_after + PIT)

When NOT to use:

  • Simple keyword search on small datasets (client-side filtering or database LIKE queries are simpler)
  • Primary data store (Elasticsearch is a search engine, not a database -- always have a source of truth elsewhere)
  • Strong consistency requirements (Elasticsearch is eventually consistent by design)
  • Simple autocomplete on a small list (a prefix trie or client-side filter is simpler)



<decision_framework>

Decision Framework

Which Query Type?

What kind of search do I need?
-- Full-text relevance search? -> match / multi_match in must
-- Exact value filtering? -> term / terms / range in filter
-- Combining text + filters? -> bool query (must for text, filter for exact)
-- Fuzzy matching? -> match with fuzziness: "AUTO"
-- Phrase matching? -> match_phrase
-- Complex nested objects? -> nested query with path
-- Vector similarity? -> knn with dense_vector field
-- Text + vector hybrid? -> query + knn in same request

Pagination Strategy?

How deep do results go?
-- Under 10,000 total? -> from/size (simplest)
-- Over 10,000 hits? -> search_after + PIT (recommended)
-- Bulk data export? -> client.helpers.scrollSearch() or scrollDocuments()
-- Real-time infinite scroll? -> search_after (no PIT needed for forward-only)

text vs keyword?

What will I do with this field?
-- Full-text search (tokenized, relevance)? -> text
-- Exact match, filtering, aggregations, sorting? -> keyword
-- Both? -> Multi-field: { type: "text", fields: { keyword: { type: "keyword" } } }
-- Neither (just stored, never queried)? -> { type: "keyword", index: false }

Filter vs Must?

Does relevance scoring matter for this clause?
-- YES (affects result order) -> must
-- NO (binary yes/no filter) -> filter (cached, no scoring overhead)
-- Exclude documents -> must_not (in filter context)
-- Boost if present (optional) -> should with minimum_should_match: 0

</decision_framework>


<red_flags>

RED FLAGS

High Priority Issues:

  • Relying on dynamic mapping without explicit mappings -- wrong type inference causes mapper_parsing_exception that requires reindexing to fix
  • Using from/size beyond 10,000 results -- Elasticsearch throws Result window is too large; use search_after + PIT
  • Looping individual client.index() calls instead of client.bulk() or client.helpers.bulk() -- orders of magnitude slower, can overwhelm the cluster
  • Using refresh: true or refresh: "wait_for" in production request handlers -- forces a segment refresh on every write, degrades cluster performance under load

Medium Priority Issues:

  • Putting exact-match conditions (term, range) in must instead of filter -- wastes CPU on scoring, misses filter cache
  • Using text type for fields that need exact matching or aggregation -- text fields are analyzed (tokenized), making aggregations return individual tokens instead of full values
  • Not including a tiebreaker field in sort when using search_after -- documents with identical sort values may be skipped or duplicated across pages
  • Missing _source check -- hit._source can be undefined if _source is disabled or fields are excluded; always handle this

Gotchas & Edge Cases:

  • Mapping types are immutable -- once a field is mapped as text, you cannot change it to keyword. The only fix is to create a new index with correct mappings and reindex all documents
  • text vs keyword confusion -- text fields are tokenized ("New York" becomes ["new", "york"]). Aggregating on a text field gives you individual tokens, not full values. Use keyword or a .keyword sub-field for aggregations
  • match vs term on text fields -- term on a text field often returns no results because term does NOT analyze the query but the field value IS analyzed (e.g., term "New York" won't match the analyzed tokens "new" and "york")
  • Near real-time delay -- after indexing, documents are NOT searchable until the next refresh (default: 1 second). Tests that index then immediately search must use refresh: "wait_for" or explicit client.indices.refresh()
  • Default index.max_result_window is 10,000 -- increasing this is possible but NOT recommended; deep pagination with from/size holds all skipped results in memory
  • Nested objects require nested mapping type -- arrays of objects are flattened by default, losing the association between fields within each object. If you need to query "color: red AND size: large" on the same object in an array, use nested
  • Aggregation on _id is disabled by default (8.x+) -- use a separate id field if you need to aggregate by document ID
  • _score is null in filter context -- clauses in filter do not contribute to scoring; if you need scoring, use must
  • Bulk API partial failures -- a bulk request can succeed overall but have individual failures. Always check result.errors and iterate result.items to find failed operations
  • Scroll API is deprecated for search -- use search_after + PIT for deep pagination. Scroll is still valid for one-time data export but consumes cluster resources (open search contexts)
  • PIT must be closed -- failing to close Point in Time contexts leaks resources on the cluster; always close in a finally block

</red_flags>


<critical_reminders>

CRITICAL REMINDERS

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST define explicit index mappings BEFORE indexing documents -- dynamic mapping infers types from the first document, and if a later document sends a different type for the same field, indexing fails with a mapper_parsing_exception that CANNOT be fixed without reindexing into a new index)

(You MUST use the bulk API for batch operations -- looping individual client.index() calls is orders of magnitude slower and can overwhelm the cluster with HTTP connections)

(You MUST NOT use from/size pagination beyond 10,000 results -- Elasticsearch throws Result window is too large by default; use search_after + PIT instead)

(You MUST NOT use refresh: true or refresh: "wait_for" in production request handlers -- forcing a refresh on every write degrades cluster performance; let the default 1-second refresh interval handle it)

Failure to follow these rules will cause mapping conflicts, pagination failures, cluster performance degradation, and silent data loss.

</critical_reminders>

Files (skills)
  • examples
    • aggregations.md 9.9 KB
      # Elasticsearch -- Aggregation Examples
      
      > Terms, range, date_histogram, nested, and pipeline aggregation patterns. Reference from [SKILL.md](../SKILL.md).
      
      **Prerequisites:** Understand client setup and index management from [core.md](core.md) first.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, document operations, search basics
      - [pagination.md](pagination.md) -- Paginating aggregation results
      
      ---
      
      ## Terms Aggregation
      
      Group by field values and get document counts.
      
      ```typescript
      import type { Client } from "@elastic/elasticsearch";
      
      const INDEX_NAME = "products";
      const AGGREGATION_SIZE = 50;
      
      async function getCategoryDistribution(
        client: Client,
      ): Promise<Array<{ key: string; count: number }>> {
        const result = await client.search({
          index: INDEX_NAME,
          size: 0, // No hits needed -- only aggregations
          aggs: {
            categories: {
              terms: { field: "categories", size: AGGREGATION_SIZE },
            },
          },
        });
      
        const buckets =
          (
            result.aggregations?.categories as {
              buckets: Array<{ key: string; doc_count: number }>;
            }
          )?.buckets ?? [];
      
        return buckets.map((b) => ({ key: b.key, count: b.doc_count }));
      }
      
      export { getCategoryDistribution };
      ```
      
      **Why good:** `size: 0` skips hits (faster when only aggregations matter), named constant for aggregation size, typed bucket extraction
      
      **Important:** `terms` aggregation returns an approximate count. The `size` parameter controls how many top buckets to return (default: 10), NOT how many documents to scan.
      
      ---
      
      ## Terms with Sub-Aggregations
      
      Nest metric aggregations inside bucket aggregations.
      
      ```typescript
      async function getCategoryStats(
        client: Client,
      ): Promise<
        Array<{ category: string; count: number; avgPrice: number; maxPrice: number }>
      > {
        const result = await client.search({
          index: INDEX_NAME,
          size: 0,
          aggs: {
            categories: {
              terms: { field: "categories", size: AGGREGATION_SIZE },
              aggs: {
                avgPrice: { avg: { field: "price" } },
                maxPrice: { max: { field: "price" } },
              },
            },
          },
        });
      
        const buckets =
          (
            result.aggregations?.categories as {
              buckets: Array<{
                key: string;
                doc_count: number;
                avgPrice: { value: number | null };
                maxPrice: { value: number | null };
              }>;
            }
          )?.buckets ?? [];
      
        return buckets.map((b) => ({
          category: b.key,
          count: b.doc_count,
          avgPrice: b.avgPrice.value ?? 0,
          maxPrice: b.maxPrice.value ?? 0,
        }));
      }
      
      export { getCategoryStats };
      ```
      
      **Why good:** Nested sub-aggregations compute per-bucket metrics, `value` can be null (e.g., empty buckets), null handled with fallback
      
      ---
      
      ## Range Aggregation
      
      Create custom numeric ranges for bucketing.
      
      ```typescript
      async function getPriceRanges(
        client: Client,
      ): Promise<Array<{ range: string; count: number }>> {
        const result = await client.search({
          index: INDEX_NAME,
          size: 0,
          aggs: {
            priceRanges: {
              range: {
                field: "price",
                ranges: [
                  { key: "budget", to: 50 },
                  { key: "mid-range", from: 50, to: 200 },
                  { key: "premium", from: 200, to: 500 },
                  { key: "luxury", from: 500 },
                ],
              },
            },
          },
        });
      
        const buckets =
          (
            result.aggregations?.priceRanges as {
              buckets: Array<{ key: string; doc_count: number }>;
            }
          )?.buckets ?? [];
      
        return buckets.map((b) => ({ range: b.key, count: b.doc_count }));
      }
      
      export { getPriceRanges };
      ```
      
      **Why good:** Named keys make bucket identification readable, ranges cover the full spectrum with no gaps
      
      **Gotcha:** `from` is inclusive, `to` is exclusive. A document with `price: 50` falls into "mid-range" (50-200), not "budget" (to: 50).
      
      ---
      
      ## Date Histogram
      
      Time-based bucketing for time series data.
      
      ```typescript
      async function getMonthlySales(
        client: Client,
        year: number,
      ): Promise<Array<{ month: string; count: number; revenue: number }>> {
        const result = await client.search({
          index: "orders",
          size: 0,
          query: {
            range: {
              orderDate: {
                gte: `${year}-01-01`,
                lt: `${year + 1}-01-01`,
              },
            },
          },
          aggs: {
            monthly: {
              date_histogram: {
                field: "orderDate",
                calendar_interval: "month",
                format: "yyyy-MM",
                min_doc_count: 0, // Include empty months
              },
              aggs: {
                revenue: { sum: { field: "totalAmount" } },
              },
            },
          },
        });
      
        const buckets =
          (
            result.aggregations?.monthly as {
              buckets: Array<{
                key_as_string: string;
                doc_count: number;
                revenue: { value: number };
              }>;
            }
          )?.buckets ?? [];
      
        return buckets.map((b) => ({
          month: b.key_as_string,
          count: b.doc_count,
          revenue: b.revenue.value,
        }));
      }
      
      export { getMonthlySales };
      ```
      
      **Why good:** `calendar_interval: "month"` handles varying month lengths correctly, `min_doc_count: 0` ensures empty months appear in results, `format` controls the `key_as_string` output
      
      **Gotcha:** Use `calendar_interval` for months/quarters/years (variable length). Use `fixed_interval` for exact durations like "30d", "1h", "5m". Using `fixed_interval: "1M"` is an error -- months are not a fixed duration.
      
      ---
      
      ## Nested Aggregation
      
      Aggregate inside nested objects.
      
      ```typescript
      // Requires "reviews" field mapped as "nested" type
      
      async function getRatingDistribution(
        client: Client,
      ): Promise<Array<{ rating: number; count: number }>> {
        const result = await client.search({
          index: INDEX_NAME,
          size: 0,
          aggs: {
            reviewsNested: {
              nested: { path: "reviews" },
              aggs: {
                ratingBuckets: {
                  histogram: {
                    field: "reviews.rating",
                    interval: 1,
                    min_doc_count: 0,
                  },
                },
              },
            },
          },
        });
      
        const buckets =
          (
            result.aggregations?.reviewsNested as {
              ratingBuckets: {
                buckets: Array<{ key: number; doc_count: number }>;
              };
            }
          )?.ratingBuckets.buckets ?? [];
      
        return buckets.map((b) => ({ rating: b.key, count: b.doc_count }));
      }
      
      export { getRatingDistribution };
      ```
      
      **Why good:** `nested` aggregation scope enters the nested documents before aggregating, histogram with interval 1 creates one bucket per rating value
      
      **Important:** Without the `nested` aggregation wrapper, aggregating on `reviews.rating` would aggregate on the flattened array values, giving incorrect counts.
      
      ---
      
      ## Pipeline Aggregation
      
      Compute values from the output of other aggregations.
      
      ```typescript
      async function getMonthlyRevenueWithMovingAvg(
        client: Client,
      ): Promise<
        Array<{ month: string; revenue: number; movingAvg: number | null }>
      > {
        const MOVING_FN_WINDOW = 3;
      
        const result = await client.search({
          index: "orders",
          size: 0,
          aggs: {
            monthly: {
              date_histogram: {
                field: "orderDate",
                calendar_interval: "month",
              },
              aggs: {
                revenue: { sum: { field: "totalAmount" } },
                revenueMovingAvg: {
                  moving_fn: {
                    buckets_path: "revenue",
                    window: MOVING_FN_WINDOW,
                    script: "MovingFunctions.unweightedAvg(values)",
                  },
                },
              },
            },
          },
        });
      
        const buckets =
          (
            result.aggregations?.monthly as {
              buckets: Array<{
                key_as_string: string;
                revenue: { value: number };
                revenueMovingAvg?: { value: number };
              }>;
            }
          )?.buckets ?? [];
      
        return buckets.map((b) => ({
          month: b.key_as_string,
          revenue: b.revenue.value,
          movingAvg: b.revenueMovingAvg?.value ?? null,
        }));
      }
      
      export { getMonthlyRevenueWithMovingAvg };
      ```
      
      **Why good:** `moving_fn` pipeline aggregation computes a rolling average from the `revenue` sub-aggregation using `MovingFunctions.unweightedAvg(values)`, `buckets_path` references the sibling aggregation by name, first N-1 buckets have null moving average (not enough data)
      
      **Important:** `moving_avg` was removed in Elasticsearch 8.0. Use `moving_fn` with a script instead. Available predefined functions: `MovingFunctions.unweightedAvg(values)`, `MovingFunctions.linearWeightedAvg(values)`, `MovingFunctions.ewma(values, alpha)`, `MovingFunctions.holt(values, alpha, beta)`, `MovingFunctions.holtWinters(values, alpha, beta, gamma, period, multiplicative)`
      
      ---
      
      ## Aggregation with Filtered Scope
      
      Apply a filter to an aggregation without affecting the main query.
      
      ```typescript
      async function getInStockVsOutOfStock(
        client: Client,
      ): Promise<{ inStock: number; outOfStock: number }> {
        const result = await client.search({
          index: INDEX_NAME,
          size: 0,
          aggs: {
            inStockCount: {
              filter: { term: { inStock: true } },
            },
            outOfStockCount: {
              filter: { term: { inStock: false } },
            },
          },
        });
      
        return {
          inStock:
            (result.aggregations?.inStockCount as { doc_count: number })?.doc_count ??
            0,
          outOfStock:
            (result.aggregations?.outOfStockCount as { doc_count: number })
              ?.doc_count ?? 0,
        };
      }
      
      export { getInStockVsOutOfStock };
      ```
      
      **Why good:** `filter` aggregation applies a separate filter scope per aggregation, each counting only matching documents
      
      ---
      
      ## Cardinality (Approximate Distinct Count)
      
      ```typescript
      async function getUniqueBrandCount(client: Client): Promise<number> {
        const result = await client.search({
          index: INDEX_NAME,
          size: 0,
          aggs: {
            uniqueBrands: {
              cardinality: { field: "brand" },
            },
          },
        });
      
        return (result.aggregations?.uniqueBrands as { value: number })?.value ?? 0;
      }
      
      export { getUniqueBrandCount };
      ```
      
      **Gotcha:** `cardinality` is an approximation using HyperLogLog++. For fields with < 1000 unique values, it's exact. For larger cardinalities, expect ~2-3% error margin.
      
      ---
      
      _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_
      
    • bulk-operations.md 9.6 KB
      # Elasticsearch -- Bulk Operations Examples
      
      > Bulk API, bulk helper, error handling, and reindexing patterns. Reference from [SKILL.md](../SKILL.md).
      
      **Prerequisites:** Understand client setup and index management from [core.md](core.md) first.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, index management, document CRUD
      - [pagination.md](pagination.md) -- scrollDocuments for reading all documents during reindex
      
      ---
      
      ## Bulk API (Low-Level)
      
      The raw bulk API uses alternating action/document pairs.
      
      ```typescript
      import type { Client, BulkResponse } from "@elastic/elasticsearch";
      
      const INDEX_NAME = "products";
      
      interface Product {
        productId: string;
        name: string;
        price: number;
        categories: string[];
      }
      
      async function bulkIndexProducts(
        client: Client,
        products: Product[],
      ): Promise<{ successful: number; failed: number }> {
        const operations = products.flatMap((doc) => [
          { index: { _index: INDEX_NAME, _id: doc.productId } },
          doc,
        ]);
      
        const result: BulkResponse = await client.bulk({
          operations,
          refresh: "wait_for", // Only in scripts/seeds -- NOT in production handlers
        });
      
        if (result.errors) {
          const failedItems = result.items.filter(
            (item) => item.index?.error !== undefined,
          );
          for (const item of failedItems) {
            console.error(
              `Failed to index ${item.index?._id}: ${item.index?.error?.reason}`,
            );
          }
          return {
            successful: products.length - failedItems.length,
            failed: failedItems.length,
          };
        }
      
        return { successful: products.length, failed: 0 };
      }
      
      export { bulkIndexProducts };
      ```
      
      **Why good:** `flatMap` creates the alternating action/document format, explicit `_id` for upsert behavior, error handling checks `result.errors` and iterates `result.items` for details
      
      **Gotcha:** A bulk request can return HTTP 200 but still have individual failures. Always check `result.errors` -- it's true if ANY item failed. Then iterate `result.items` to find which ones.
      
      **Gotcha:** A 429 status on individual items means "too many requests" -- these are retriable. Other error codes (400, 409) typically indicate data issues that require fixing the document.
      
      ---
      
      ## Bulk Helper (Recommended)
      
      The bulk helper handles batching, concurrency, retries, and back-pressure automatically.
      
      ```typescript
      async function bulkIndexWithHelper(
        client: Client,
        products: Product[],
      ): Promise<{ total: number; successful: number; failed: number }> {
        const result = await client.helpers.bulk<Product>({
          datasource: products,
          onDocument(doc) {
            return { index: { _index: INDEX_NAME, _id: doc.productId } };
          },
          refreshOnCompletion: INDEX_NAME,
        });
      
        return {
          total: result.total,
          successful: result.successful,
          failed: result.failed,
        };
      }
      
      export { bulkIndexWithHelper };
      ```
      
      **Why good:** Automatic batching (default: 5MB per batch), automatic concurrency (default: 5 parallel requests), automatic retries on 429, `refreshOnCompletion` triggers one refresh at the end
      
      ### Bulk Helper with Streaming Data Source
      
      ```typescript
      import { createReadStream } from "node:fs";
      import { createInterface } from "node:readline";
      
      async function bulkIndexFromFile(
        client: Client,
        filePath: string,
      ): Promise<{ total: number; failed: number }> {
        const lineReader = createInterface({
          input: createReadStream(filePath),
        });
      
        async function* generateDocuments() {
          for await (const line of lineReader) {
            if (line.trim()) {
              yield JSON.parse(line) as Product;
            }
          }
        }
      
        const result = await client.helpers.bulk<Product>({
          datasource: generateDocuments(),
          onDocument(doc) {
            return { index: { _index: INDEX_NAME, _id: doc.productId } };
          },
          refreshOnCompletion: INDEX_NAME,
        });
      
        return { total: result.total, failed: result.failed };
      }
      
      export { bulkIndexFromFile };
      ```
      
      **Why good:** Async generator streams documents from file without loading all into memory, supports NDJSON format (one JSON object per line)
      
      ### Bulk Update
      
      ```typescript
      async function bulkUpdatePrices(
        client: Client,
        updates: Array<{ productId: string; newPrice: number }>,
      ): Promise<{ total: number; failed: number }> {
        const result = await client.helpers.bulk({
          datasource: updates,
          onDocument(item) {
            return [
              { update: { _index: INDEX_NAME, _id: item.productId } },
              { doc: { price: item.newPrice } },
            ];
          },
        });
      
        return { total: result.total, failed: result.failed };
      }
      
      export { bulkUpdatePrices };
      ```
      
      **Why good:** `onDocument` returns a two-element array for update operations -- first element is the action, second is the update body with `doc` for partial update
      
      ### Bulk Delete
      
      ```typescript
      async function bulkDeleteProducts(
        client: Client,
        productIds: string[],
      ): Promise<{ total: number; failed: number }> {
        const result = await client.helpers.bulk({
          datasource: productIds,
          onDocument(id) {
            return { delete: { _index: INDEX_NAME, _id: id } };
          },
        });
      
        return { total: result.total, failed: result.failed };
      }
      
      export { bulkDeleteProducts };
      ```
      
      **Why good:** Delete operations have no document body -- `onDocument` returns only the action
      
      ---
      
      ## Bulk Helper Configuration
      
      ```typescript
      const FLUSH_BYTES = 5_000_000; // 5MB -- max batch size before flushing
      const FLUSH_INTERVAL_MS = 30_000; // 30s -- max time before flushing
      const CONCURRENCY = 5; // Parallel bulk requests
      const RETRIES = 3; // Retry attempts per document on 429
      
      const result = await client.helpers.bulk<Product>({
        datasource: products,
        onDocument(doc) {
          return { index: { _index: INDEX_NAME, _id: doc.productId } };
        },
        flushBytes: FLUSH_BYTES,
        flushInterval: FLUSH_INTERVAL_MS,
        concurrency: CONCURRENCY,
        retries: RETRIES,
        refreshOnCompletion: INDEX_NAME,
        onDrop(doc) {
          // Called when a document fails after all retries
          console.error(`Dropped document: ${doc.document.productId}`);
        },
      });
      ```
      
      **Why good:** Named constants for all tuning parameters, `onDrop` callback for monitoring permanently failed documents
      
      ---
      
      ## Reindexing with Alias Swap
      
      Zero-downtime reindex by creating a new index, copying data, and swapping aliases.
      
      ```typescript
      const READ_ALIAS = "products-read";
      const WRITE_ALIAS = "products-write";
      const BATCH_SIZE = 500;
      
      async function reindexWithAliasSwap(
        client: Client,
        newIndexName: string,
      ): Promise<void> {
        // 1. Create new index with updated mappings
        await client.indices.create({
          index: newIndexName,
          mappings: {
            properties: {
              productId: { type: "keyword" },
              name: {
                type: "text",
                fields: { keyword: { type: "keyword" } },
              },
              description: { type: "text" },
              price: { type: "float" },
              categories: { type: "keyword" },
              brand: { type: "keyword" },
              inStock: { type: "boolean" },
              rating: { type: "float" }, // New field
              createdAt: { type: "date" },
            },
          },
        });
      
        // 2. Disable refresh during bulk copy (faster indexing)
        await client.indices.putSettings({
          index: newIndexName,
          settings: { index: { refresh_interval: "-1" } },
        });
      
        // 3. Copy documents from old index using scroll
        const docs = client.helpers.scrollDocuments<Product>({
          index: READ_ALIAS,
          query: { match_all: {} },
        });
      
        await client.helpers.bulk({
          datasource: docs,
          onDocument(doc) {
            return { index: { _index: newIndexName, _id: doc.productId } };
          },
          refreshOnCompletion: newIndexName,
        });
      
        // 4. Re-enable refresh
        await client.indices.putSettings({
          index: newIndexName,
          settings: { index: { refresh_interval: "1s" } },
        });
      
        // 5. Atomic alias swap
        // Find the current index behind the alias
        const aliasInfo = await client.indices.getAlias({ name: READ_ALIAS });
        const oldIndexNames = Object.keys(aliasInfo);
      
        await client.indices.updateAliases({
          actions: [
            ...oldIndexNames.map((oldIdx) => ({
              remove: { index: oldIdx, alias: READ_ALIAS },
            })),
            ...oldIndexNames.map((oldIdx) => ({
              remove: { index: oldIdx, alias: WRITE_ALIAS },
            })),
            { add: { index: newIndexName, alias: READ_ALIAS } },
            { add: { index: newIndexName, alias: WRITE_ALIAS } },
          ],
        });
      }
      
      export { reindexWithAliasSwap };
      ```
      
      **Why good:** Zero-downtime reindex using alias swap, refresh disabled during bulk copy for speed, `scrollDocuments` streams data without loading all into memory, alias swap is atomic
      
      **Gotcha:** Disabling refresh during bulk copy is critical for performance -- without it, Elasticsearch creates a new segment every second during the copy. Always re-enable refresh after the copy completes.
      
      ---
      
      ## Server-Side Reindex API
      
      For simple reindexing without transformation, use the built-in reindex API.
      
      ```typescript
      async function reindexServerSide(
        client: Client,
        sourceIndex: string,
        destIndex: string,
      ): Promise<{ total: number }> {
        const result = await client.reindex({
          source: { index: sourceIndex },
          dest: { index: destIndex },
          wait_for_completion: true, // Block until done -- only in scripts
        });
      
        return { total: result.total ?? 0 };
      }
      
      export { reindexServerSide };
      ```
      
      **Why good:** Server-side reindex avoids round-tripping documents through the client -- faster for large datasets. Use `wait_for_completion: false` for very large reindexes and poll the task API instead.
      
      **When to use:** When you need to copy data between indices without transforming the document structure. For transformations (adding fields, changing types), use the client-side scroll + bulk pattern above.
      
      ---
      
      _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_
      
    • core.md 14.4 KB
      # Elasticsearch -- Core Pattern Examples
      
      > Client setup, index management, document CRUD, search basics, and TypeScript integration. Reference from [SKILL.md](../SKILL.md).
      
      **Related examples:**
      
      - [aggregations.md](aggregations.md) -- Terms, range, date_histogram, pipeline aggregations
      - [vector-search.md](vector-search.md) -- Dense vector fields, kNN, hybrid search
      - [pagination.md](pagination.md) -- search_after, PIT, scroll helpers
      - [bulk-operations.md](bulk-operations.md) -- Bulk API, bulk helper, reindexing
      
      ---
      
      ## Client Setup
      
      ### Basic Setup with API Key Auth
      
      ```typescript
      import { Client } from "@elastic/elasticsearch";
      
      function createElasticsearchClient(): Client {
        const node = process.env.ELASTICSEARCH_URL;
        if (!node) {
          throw new Error("ELASTICSEARCH_URL environment variable is required");
        }
      
        return new Client({
          node,
          auth: {
            apiKey: process.env.ELASTICSEARCH_API_KEY ?? "",
          },
        });
      }
      
      export { createElasticsearchClient };
      ```
      
      **Why good:** Environment variable validation, API key auth (preferred over basic auth), named export
      
      ### Elastic Cloud Setup
      
      ```typescript
      import { Client } from "@elastic/elasticsearch";
      
      function createCloudClient(): Client {
        const cloudId = process.env.ELASTIC_CLOUD_ID;
        if (!cloudId) {
          throw new Error("ELASTIC_CLOUD_ID environment variable is required");
        }
      
        return new Client({
          cloud: { id: cloudId },
          auth: {
            apiKey: process.env.ELASTICSEARCH_API_KEY ?? "",
          },
        });
      }
      
      export { createCloudClient };
      ```
      
      **Why good:** `cloud.id` encodes the cluster URL -- no need to construct URLs manually, handles TLS automatically
      
      ### Health Check
      
      ```typescript
      import type { Client } from "@elastic/elasticsearch";
      
      const HEALTH_TIMEOUT_MS = 5000;
      
      async function verifyConnection(client: Client): Promise<boolean> {
        try {
          const controller = new AbortController();
          const timeout = setTimeout(() => controller.abort(), HEALTH_TIMEOUT_MS);
      
          await client.ping({ signal: controller.signal });
          clearTimeout(timeout);
          return true;
        } catch {
          return false;
        }
      }
      
      export { verifyConnection };
      ```
      
      **Why good:** `client.ping()` is the lightest health check, AbortController prevents hanging on unresponsive cluster, named constant for timeout
      
      ---
      
      ## Index Management
      
      ### Creating an Index with Explicit Mappings
      
      ```typescript
      import type { Client } from "@elastic/elasticsearch";
      
      const INDEX_NAME = "products";
      
      async function createProductIndex(client: Client): Promise<void> {
        const exists = await client.indices.exists({ index: INDEX_NAME });
        if (exists) return;
      
        await client.indices.create({
          index: INDEX_NAME,
          settings: {
            number_of_replicas: 1,
            refresh_interval: "1s",
            analysis: {
              analyzer: {
                product_analyzer: {
                  type: "custom",
                  tokenizer: "standard",
                  filter: ["lowercase", "asciifolding"],
                },
              },
            },
          },
          mappings: {
            dynamic: "strict", // Reject documents with unmapped fields
            properties: {
              productId: { type: "keyword" },
              name: {
                type: "text",
                analyzer: "product_analyzer",
                fields: { keyword: { type: "keyword", ignore_above: 256 } },
              },
              description: { type: "text", analyzer: "product_analyzer" },
              price: { type: "float" },
              categories: { type: "keyword" },
              brand: { type: "keyword" },
              inStock: { type: "boolean" },
              tags: { type: "keyword" },
              createdAt: { type: "date" },
            },
          },
        });
      }
      
      export { createProductIndex };
      ```
      
      **Why good:** `dynamic: "strict"` rejects unmapped fields (prevents mapping explosion), custom analyzer with asciifolding for accent-insensitive search, `text` + `keyword` multi-field on `name` for both search and aggregation, existence check prevents errors on re-run
      
      ### Adding Fields to Existing Mapping
      
      ```typescript
      // You CAN add new fields to an existing mapping
      // You CANNOT change the type of an existing field
      async function addRatingField(client: Client): Promise<void> {
        await client.indices.putMapping({
          index: INDEX_NAME,
          properties: {
            rating: { type: "float" },
            reviewCount: { type: "integer" },
          },
        });
      }
      
      export { addRatingField };
      ```
      
      **Important:** `putMapping` can only ADD new fields. Changing an existing field type (e.g., `text` to `keyword`) requires creating a new index with correct mappings and reindexing all documents.
      
      ### Index Aliases for Zero-Downtime Reindexing
      
      ```typescript
      const PRODUCTS_READ_ALIAS = "products-read";
      const PRODUCTS_WRITE_ALIAS = "products-write";
      
      async function swapIndex(
        client: Client,
        oldIndex: string,
        newIndex: string,
      ): Promise<void> {
        await client.indices.updateAliases({
          actions: [
            { remove: { index: oldIndex, alias: PRODUCTS_READ_ALIAS } },
            { add: { index: newIndex, alias: PRODUCTS_READ_ALIAS } },
            { remove: { index: oldIndex, alias: PRODUCTS_WRITE_ALIAS } },
            { add: { index: newIndex, alias: PRODUCTS_WRITE_ALIAS } },
          ],
        });
      }
      
      export { swapIndex };
      ```
      
      **Why good:** `updateAliases` is atomic -- read and write aliases switch simultaneously, no downtime during reindex
      
      ---
      
      ## Document Operations
      
      ### Indexing a Document
      
      ```typescript
      import type { Client } from "@elastic/elasticsearch";
      
      interface Product {
        productId: string;
        name: string;
        description: string;
        price: number;
        categories: string[];
        brand: string;
        inStock: boolean;
        createdAt: string;
      }
      
      const INDEX_NAME = "products";
      
      async function indexProduct(client: Client, product: Product): Promise<string> {
        const result = await client.index({
          index: INDEX_NAME,
          id: product.productId, // Explicit ID for upsert behavior
          document: product,
        });
        return result._id;
      }
      
      export { indexProduct };
      export type { Product };
      ```
      
      **Why good:** Explicit `id` enables upsert (index or replace), typed document, returns the document ID
      
      ### Getting a Document by ID
      
      ```typescript
      async function getProduct(
        client: Client,
        productId: string,
      ): Promise<Product | null> {
        try {
          const result = await client.get<Product>({
            index: INDEX_NAME,
            id: productId,
          });
          return result._source ?? null;
        } catch (err) {
          if (
            err instanceof Error &&
            "statusCode" in err &&
            (err as { statusCode: number }).statusCode === 404
          ) {
            return null;
          }
          throw err;
        }
      }
      
      export { getProduct };
      ```
      
      **Why good:** `_source` can be undefined, 404 handled gracefully (document not found is not an error in most use cases), generic type flows through to `_source`
      
      ### Partial Update
      
      ```typescript
      async function updateProductPrice(
        client: Client,
        productId: string,
        newPrice: number,
      ): Promise<void> {
        await client.update({
          index: INDEX_NAME,
          id: productId,
          doc: { price: newPrice },
        });
      }
      
      export { updateProductPrice };
      ```
      
      **Why good:** `doc` performs partial update -- only `price` changes, all other fields preserved. Compare with `client.index()` which replaces the entire document.
      
      ### Scripted Update (Atomic)
      
      ```typescript
      const PRICE_INCREASE_PERCENTAGE = 10;
      
      async function increasePriceByPercent(
        client: Client,
        productId: string,
      ): Promise<void> {
        await client.update({
          index: INDEX_NAME,
          id: productId,
          script: {
            source: "ctx._source.price *= (1 + params.pct / 100.0)",
            params: { pct: PRICE_INCREASE_PERCENTAGE },
          },
        });
      }
      
      export { increasePriceByPercent };
      ```
      
      **Why good:** Script runs on the shard -- atomic, no read-modify-write race condition. Named constant for the percentage value.
      
      ### Delete by ID and by Query
      
      ```typescript
      async function deleteProduct(client: Client, productId: string): Promise<void> {
        await client.delete({
          index: INDEX_NAME,
          id: productId,
        });
      }
      
      async function deleteOutOfStockProducts(client: Client): Promise<number> {
        const result = await client.deleteByQuery({
          index: INDEX_NAME,
          query: {
            term: { inStock: false },
          },
        });
        return result.deleted ?? 0;
      }
      
      export { deleteProduct, deleteOutOfStockProducts };
      ```
      
      **Why good:** `deleteByQuery` for batch deletion without knowing individual IDs, returns count of deleted documents
      
      ---
      
      ## Search Patterns
      
      ### Basic Search with Highlighting
      
      ```typescript
      import type { Client, SearchResponse } from "@elastic/elasticsearch";
      
      const DEFAULT_SEARCH_LIMIT = 20;
      
      async function searchProducts(
        client: Client,
        query: string,
        options?: { limit?: number },
      ): Promise<SearchResponse<Product>> {
        return client.search<Product>({
          index: INDEX_NAME,
          query: {
            multi_match: {
              query,
              fields: ["name^3", "description", "brand^2"],
              type: "best_fields",
              fuzziness: "AUTO",
            },
          },
          highlight: {
            fields: {
              name: {},
              description: { fragment_size: 150, number_of_fragments: 2 },
            },
            pre_tags: ["<mark>"],
            post_tags: ["</mark>"],
          },
          size: options?.limit ?? DEFAULT_SEARCH_LIMIT,
        });
      }
      
      // Usage:
      // const results = await searchProducts(client, "wireless headphones");
      // results.hits.hits[0]._source?.name -- original
      // results.hits.hits[0].highlight?.name?.[0] -- "<mark>wireless</mark> <mark>headphones</mark>"
      
      export { searchProducts };
      ```
      
      **Why good:** `name^3` boosts name matches 3x, `fuzziness: "AUTO"` handles typos, highlighting configured with custom tags, typed search response
      
      ### Bool Query with Filter Context
      
      ```typescript
      interface SearchFilters {
        minPrice?: number;
        maxPrice?: number;
        categories?: string[];
        inStock?: boolean;
        query?: string;
      }
      
      async function filteredSearch(
        client: Client,
        filters: SearchFilters,
      ): Promise<SearchResponse<Product>> {
        const must: object[] = [];
        const filterClauses: object[] = [];
      
        if (filters.query) {
          must.push({
            multi_match: {
              query: filters.query,
              fields: ["name^3", "description", "brand"],
            },
          });
        }
      
        if (filters.minPrice !== undefined || filters.maxPrice !== undefined) {
          filterClauses.push({
            range: {
              price: {
                ...(filters.minPrice !== undefined ? { gte: filters.minPrice } : {}),
                ...(filters.maxPrice !== undefined ? { lte: filters.maxPrice } : {}),
              },
            },
          });
        }
      
        if (filters.categories?.length) {
          filterClauses.push({ terms: { categories: filters.categories } });
        }
      
        if (filters.inStock !== undefined) {
          filterClauses.push({ term: { inStock: filters.inStock } });
        }
      
        return client.search<Product>({
          index: INDEX_NAME,
          query: {
            bool: {
              ...(must.length > 0 ? { must } : {}),
              ...(filterClauses.length > 0 ? { filter: filterClauses } : {}),
            },
          },
          size: DEFAULT_SEARCH_LIMIT,
        });
      }
      
      export { filteredSearch };
      ```
      
      **Why good:** Exact-match conditions in `filter` (no scoring, cached), full-text in `must` (scoring), dynamic query construction based on provided filters
      
      ### Nested Query
      
      ```typescript
      // For arrays of objects where field association matters
      // Mapping must use "nested" type, not default "object"
      
      interface ProductWithReviews extends Product {
        reviews: Array<{ author: string; rating: number; text: string }>;
      }
      
      async function searchByReview(
        client: Client,
        minRating: number,
        reviewText: string,
      ): Promise<SearchResponse<ProductWithReviews>> {
        return client.search<ProductWithReviews>({
          index: INDEX_NAME,
          query: {
            nested: {
              path: "reviews",
              query: {
                bool: {
                  must: [{ match: { "reviews.text": reviewText } }],
                  filter: [{ range: { "reviews.rating": { gte: minRating } } }],
                },
              },
              inner_hits: { size: 3 }, // Return matching nested docs
            },
          },
        });
      }
      
      export { searchByReview };
      ```
      
      **Why good:** `nested` query preserves field association within each review object (without `nested`, searching for "great" with rating >= 4 could match "great" from one review and rating 5 from a different review), `inner_hits` returns the matching nested documents
      
      **Important:** The `reviews` field must be mapped as `"type": "nested"`. Default `object` mapping flattens the array, losing field-level association.
      
      ---
      
      ## TypeScript Integration
      
      ### Typed Search Results
      
      ```typescript
      import type { Client, SearchHit } from "@elastic/elasticsearch";
      
      // Generic type flows through to hits
      async function searchAndTransform(
        client: Client,
        query: string,
      ): Promise<Array<{ id: string; name: string; score: number }>> {
        const result = await client.search<Product>({
          index: INDEX_NAME,
          query: { match: { name: query } },
          size: 10,
        });
      
        return result.hits.hits
          .filter(
            (hit): hit is SearchHit<Product> & { _source: Product } =>
              hit._source !== undefined,
          )
          .map((hit) => ({
            id: hit._id,
            name: hit._source.name,
            score: hit._score ?? 0,
          }));
      }
      
      export { searchAndTransform };
      ```
      
      **Why good:** Type guard filters out hits with missing `_source`, `_score` can be null (e.g., in filter context), generic type parameter provides type safety on `_source`
      
      ### Importing Elasticsearch Types
      
      ```typescript
      import type {
        SearchResponse,
        SearchHit,
        BulkResponse,
        MappingProperty,
      } from "@elastic/elasticsearch/lib/api/types";
      
      // Or use the estypes namespace for full type coverage
      import type { estypes } from "@elastic/elasticsearch";
      
      type MySearchResponse = estypes.SearchResponse<Product>;
      ```
      
      **Why good:** `estypes` provides complete request/response types matching the Elasticsearch specification
      
      ---
      
      ## Error Handling
      
      ### Common Error Patterns
      
      ```typescript
      import { errors } from "@elastic/elasticsearch";
      
      async function safeSearch(
        client: Client,
        query: string,
      ): Promise<SearchResponse<Product> | null> {
        try {
          return await client.search<Product>({
            index: INDEX_NAME,
            query: { match: { name: query } },
          });
        } catch (err) {
          if (err instanceof errors.ResponseError) {
            if (err.statusCode === 404) {
              // Index doesn't exist
              return null;
            }
            if (err.statusCode === 400) {
              // Bad query (malformed DSL)
              throw new Error(`Invalid search query: ${err.message}`);
            }
          }
          if (err instanceof errors.ConnectionError) {
            // Cluster unreachable
            throw new Error("Elasticsearch cluster unavailable");
          }
          throw err;
        }
      }
      
      export { safeSearch };
      ```
      
      **Why good:** `errors.ResponseError` for HTTP errors with status codes, `errors.ConnectionError` for network issues, specific handling per status code
      
      ---
      
      _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_
      
    • pagination.md 7.7 KB
      # Elasticsearch -- Pagination Examples
      
      > from/size, search_after, Point in Time (PIT), and scroll helper patterns. Reference from [SKILL.md](../SKILL.md).
      
      **Prerequisites:** Understand client setup and search basics from [core.md](core.md) first.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, search basics
      - [bulk-operations.md](bulk-operations.md) -- Bulk data export with scrollDocuments
      
      ---
      
      ## Pagination Decision Framework
      
      ```
      How deep do results go?
      -- Under 10,000 total and users jump between pages? -> from/size
      -- Over 10,000 hits, forward-only (infinite scroll)? -> search_after
      -- Over 10,000 hits, need consistent snapshot? -> search_after + PIT
      -- Bulk data export (all documents)? -> client.helpers.scrollDocuments()
      ```
      
      ---
      
      ## from/size (Shallow Pagination)
      
      Simple offset-based pagination. Limited to 10,000 total hits by default.
      
      ```typescript
      import type { Client, SearchResponse } from "@elastic/elasticsearch";
      
      const INDEX_NAME = "products";
      const PAGE_SIZE = 20;
      
      interface Product {
        productId: string;
        name: string;
        price: number;
      }
      
      async function getPage(
        client: Client,
        query: string,
        page: number,
      ): Promise<SearchResponse<Product>> {
        const from = (page - 1) * PAGE_SIZE;
      
        return client.search<Product>({
          index: INDEX_NAME,
          query: { match: { name: query } },
          from,
          size: PAGE_SIZE,
          track_total_hits: true, // Get exact total count
        });
      }
      
      export { getPage };
      ```
      
      **Why good:** Simplest pagination, supports random page access (jump to page 5), `track_total_hits: true` for exact total count
      
      **Limitations:**
      
      - Default `index.max_result_window` is 10,000 -- `from + size` cannot exceed this
      - Deep pages are expensive: Elasticsearch must fetch and discard `from` documents on every shard
      - Not suitable for infinite scroll or large result sets
      
      ---
      
      ## search_after (Deep Pagination)
      
      Stateless cursor-based pagination using sort values. No 10,000 hit limit.
      
      ```typescript
      const PAGE_SIZE = 20;
      
      async function searchAfterPage(
        client: Client,
        query: string,
        lastSort?: Array<string | number>,
      ): Promise<{
        hits: Product[];
        nextSort: Array<string | number> | null;
        total: number;
      }> {
        const result = await client.search<Product>({
          index: INDEX_NAME,
          query: { match: { name: query } },
          sort: [
            { _score: "desc" },
            { "name.keyword": "asc" }, // Tiebreaker -- REQUIRED
          ],
          size: PAGE_SIZE,
          ...(lastSort ? { search_after: lastSort } : {}),
          track_total_hits: true,
        });
      
        const hits = result.hits.hits;
        const nextSort =
          hits.length > 0
            ? (hits[hits.length - 1].sort as Array<string | number>)
            : null;
      
        return {
          hits: hits.filter((h) => h._source !== undefined).map((h) => h._source!),
          nextSort,
          total:
            typeof result.hits.total === "number"
              ? result.hits.total
              : (result.hits.total?.value ?? 0),
        };
      }
      
      export { searchAfterPage };
      ```
      
      **Why good:** No 10,000 hit limit, tiebreaker field prevents missing/duplicate documents, stateless (no server-side cursor to maintain)
      
      **Gotcha:** `search_after` requires a `sort` parameter. You MUST include a tiebreaker field (a unique value like `_id` or a keyword field) -- without it, documents with identical sort values may be skipped or duplicated across pages.
      
      **Gotcha:** `search_after` is forward-only. You cannot jump to page 5 directly -- you must iterate through pages 1-4 first. For random page access on small result sets, use `from`/`size`.
      
      ---
      
      ## search_after + Point in Time (Consistent Deep Pagination)
      
      PIT creates a snapshot of the index state, ensuring consistent results even as documents are added/updated during pagination.
      
      ```typescript
      const PIT_KEEP_ALIVE = "1m";
      const PAGE_SIZE = 100;
      
      async function paginateAllProducts(
        client: Client,
        onPage: (products: Product[]) => Promise<void>,
      ): Promise<void> {
        const pit = await client.openPointInTime({
          index: INDEX_NAME,
          keep_alive: PIT_KEEP_ALIVE,
        });
      
        let searchAfter: Array<string | number> | undefined;
      
        try {
          while (true) {
            const result = await client.search<Product>({
              pit: { id: pit.id, keep_alive: PIT_KEEP_ALIVE },
              sort: [{ createdAt: "desc" }, { _id: "asc" }],
              size: PAGE_SIZE,
              ...(searchAfter ? { search_after: searchAfter } : {}),
            });
      
            const hits = result.hits.hits;
            if (hits.length === 0) break;
      
            const products = hits
              .filter((h) => h._source !== undefined)
              .map((h) => h._source!);
      
            await onPage(products);
      
            searchAfter = hits[hits.length - 1].sort as Array<string | number>;
          }
        } finally {
          await client.closePointInTime({ id: pit.id });
        }
      }
      
      export { paginateAllProducts };
      ```
      
      **Why good:** PIT ensures consistent snapshot across all pages (no missing/duplicate documents from concurrent writes), `keep_alive` refreshed on each request, PIT closed in `finally` block (prevents resource leak)
      
      **Key differences from plain search_after:**
      
      - PIT provides a frozen view of the index -- new documents added during pagination are not visible
      - Without PIT, `search_after` sees the live index -- concurrent writes can cause documents to shift between pages
      - PIT requires no `index` parameter in the search (it's bound to the PIT)
      
      **Gotcha:** PIT search automatically adds `_shard_doc` as an implicit tiebreaker. You can still add your own tiebreaker for deterministic ordering.
      
      ---
      
      ## scrollSearch Helper (Page-Level Iteration)
      
      The built-in helper handles scroll management automatically. Use for batch processing.
      
      ```typescript
      async function processAllDocuments(
        client: Client,
        onBatch: (docs: Product[]) => Promise<void>,
      ): Promise<void> {
        const scrollSearch = client.helpers.scrollSearch<Product>({
          index: INDEX_NAME,
          query: { match_all: {} },
          size: 500,
        });
      
        for await (const result of scrollSearch) {
          const docs = result.documents;
          await onBatch(docs);
        }
      }
      
      export { processAllDocuments };
      ```
      
      **Why good:** Async iterator handles scroll lifecycle automatically (open, fetch, clear), `documents` property provides typed `_source` values directly
      
      **When to use:** One-time batch processing or data export. NOT for user-facing pagination.
      
      ---
      
      ## scrollDocuments Helper (Document-Level Iteration)
      
      Yields individual documents instead of pages. Most memory-efficient for large exports.
      
      ```typescript
      async function exportAllProducts(client: Client): Promise<Product[]> {
        const allProducts: Product[] = [];
      
        const docs = client.helpers.scrollDocuments<Product>({
          index: INDEX_NAME,
          query: { match_all: {} },
        });
      
        for await (const doc of docs) {
          allProducts.push(doc);
        }
      
        return allProducts;
      }
      
      export { exportAllProducts };
      ```
      
      **Why good:** Most memory-efficient -- processes one document at a time, automatic filter_path optimization
      
      **When to use:** Large data exports, ETL pipelines, index migrations. The scroll API is NOT deprecated for this use case -- it's only deprecated for user-facing search pagination.
      
      ---
      
      ## Pagination Comparison
      
      | Method                | Max Depth | Random Access | Consistency | Use Case                         |
      | --------------------- | --------- | ------------- | ----------- | -------------------------------- |
      | `from`/`size`         | 10,000    | Yes           | Live index  | UI with page numbers             |
      | `search_after`        | Unlimited | No (forward)  | Live index  | Infinite scroll                  |
      | `search_after` + PIT  | Unlimited | No (forward)  | Snapshot    | Deep pagination with consistency |
      | `scrollSearch` helper | Unlimited | No (forward)  | Snapshot    | Batch processing (pages)         |
      | `scrollDocuments`     | Unlimited | No (forward)  | Snapshot    | Batch processing (documents)     |
      
      ---
      
      _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_
      
    • vector-search.md 9.3 KB
      # Elasticsearch -- Vector Search Examples
      
      > Dense vector fields, kNN queries, hybrid text+vector search, and similarity metrics. Reference from [SKILL.md](../SKILL.md).
      
      **Prerequisites:** Understand client setup and index management from [core.md](core.md) first.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, index management, search basics
      - [aggregations.md](aggregations.md) -- Combining aggregations with vector search
      
      ---
      
      ## Dense Vector Mapping
      
      Configure a `dense_vector` field for kNN search.
      
      ```typescript
      import type { Client } from "@elastic/elasticsearch";
      
      const INDEX_NAME = "articles";
      const EMBEDDING_DIMS = 768; // Must match your embedding model's output dimension
      
      async function createVectorIndex(client: Client): Promise<void> {
        await client.indices.create({
          index: INDEX_NAME,
          mappings: {
            properties: {
              title: {
                type: "text",
                fields: { keyword: { type: "keyword" } },
              },
              content: { type: "text" },
              embedding: {
                type: "dense_vector",
                dims: EMBEDDING_DIMS,
                index: true, // Required for kNN search (default: true in 8.11+)
                similarity: "cosine", // cosine | l2_norm | dot_product | max_inner_product
              },
              category: { type: "keyword" },
              publishedAt: { type: "date" },
            },
          },
        });
      }
      
      export { createVectorIndex };
      ```
      
      **Why good:** `dims` matches the embedding model's output dimension exactly (768 for most BERT-based models), `similarity: "cosine"` is the most common metric, `index: true` enables approximate kNN search
      
      **Gotcha:** If `dims` does not match the actual vector length when indexing documents, Elasticsearch rejects the document with a `mapper_parsing_exception`. There is no auto-truncation or padding.
      
      ---
      
      ## Indexing Documents with Vectors
      
      ```typescript
      interface Article {
        title: string;
        content: string;
        embedding: number[];
        category: string;
        publishedAt: string;
      }
      
      async function indexArticleWithVector(
        client: Client,
        article: Article,
        id: string,
      ): Promise<void> {
        await client.index({
          index: INDEX_NAME,
          id,
          document: article,
        });
      }
      
      export { indexArticleWithVector };
      ```
      
      **Important:** The `embedding` array MUST have exactly `EMBEDDING_DIMS` elements. Generate embeddings using your embedding model before indexing -- Elasticsearch does not generate embeddings for you (unless you configure an ingest pipeline with a deployed model).
      
      ---
      
      ## Basic kNN Search
      
      Find documents with vectors most similar to a query vector.
      
      ```typescript
      const KNN_CANDIDATES = 100;
      const KNN_RESULTS = 10;
      
      async function vectorSearch(
        client: Client,
        queryVector: number[],
      ): Promise<Array<{ id: string; title: string; score: number }>> {
        const result = await client.search<Article>({
          index: INDEX_NAME,
          knn: {
            field: "embedding",
            query_vector: queryVector,
            k: KNN_RESULTS,
            num_candidates: KNN_CANDIDATES,
          },
        });
      
        return result.hits.hits
          .filter((hit) => hit._source !== undefined)
          .map((hit) => ({
            id: hit._id,
            title: hit._source!.title,
            score: hit._score ?? 0,
          }));
      }
      
      export { vectorSearch };
      ```
      
      **Why good:** `num_candidates` controls the accuracy/speed tradeoff (higher = more accurate but slower), `k` is the number of results to return, named constants for both
      
      **Key concept:** `num_candidates` determines how many candidates each shard considers before the final `k` results are selected. A ratio of `num_candidates / k >= 10` is a good starting point.
      
      ---
      
      ## kNN with Filters
      
      Apply filters during the kNN search to restrict the vector space.
      
      ```typescript
      async function filteredVectorSearch(
        client: Client,
        queryVector: number[],
        category: string,
      ): Promise<Array<{ id: string; title: string; score: number }>> {
        const result = await client.search<Article>({
          index: INDEX_NAME,
          knn: {
            field: "embedding",
            query_vector: queryVector,
            k: KNN_RESULTS,
            num_candidates: KNN_CANDIDATES,
            filter: {
              term: { category },
            },
          },
        });
      
        return result.hits.hits
          .filter((hit) => hit._source !== undefined)
          .map((hit) => ({
            id: hit._id,
            title: hit._source!.title,
            score: hit._score ?? 0,
          }));
      }
      
      export { filteredVectorSearch };
      ```
      
      **Why good:** Filter is applied DURING the kNN search, not after -- this ensures exactly `k` results from the filtered subset, not fewer
      
      **Gotcha:** Without the `filter` inside `knn`, you would need to use a top-level `query` filter which is applied AFTER kNN. Post-filtering can return fewer than `k` results because some of the `k` nearest neighbors may not match the filter.
      
      ---
      
      ## Hybrid Search (Text + Vector)
      
      Combine traditional full-text search with vector similarity.
      
      ```typescript
      const TEXT_BOOST = 0.7;
      const VECTOR_BOOST = 0.3;
      
      async function hybridSearch(
        client: Client,
        queryText: string,
        queryVector: number[],
      ): Promise<Array<{ id: string; title: string; score: number }>> {
        const result = await client.search<Article>({
          index: INDEX_NAME,
          query: {
            match: {
              content: {
                query: queryText,
                boost: TEXT_BOOST,
              },
            },
          },
          knn: {
            field: "embedding",
            query_vector: queryVector,
            k: KNN_RESULTS,
            num_candidates: KNN_CANDIDATES,
            boost: VECTOR_BOOST,
          },
          size: KNN_RESULTS,
        });
      
        return result.hits.hits
          .filter((hit) => hit._source !== undefined)
          .map((hit) => ({
            id: hit._id,
            title: hit._source!.title,
            score: hit._score ?? 0,
          }));
      }
      
      export { hybridSearch };
      ```
      
      **Why good:** Named constants for boost values, text and vector scores are combined with weighted boosts, `size` limits total results from both sources
      
      **How scoring works:** The final `_score` = `TEXT_BOOST * text_score + VECTOR_BOOST * knn_score`. Adjust boosts to control the balance between keyword relevance and semantic similarity.
      
      **Gotcha:** Results are combined via disjunction (OR) -- a document can appear from text match only, vector match only, or both. Documents matching both get a combined score.
      
      ---
      
      ## Semantic Search with Model Inference
      
      Let Elasticsearch generate the query vector using a deployed model.
      
      ```typescript
      const MODEL_ID = "my-text-embedding-model";
      
      async function semanticSearch(
        client: Client,
        queryText: string,
      ): Promise<Array<{ id: string; title: string; score: number }>> {
        const result = await client.search<Article>({
          index: INDEX_NAME,
          knn: {
            field: "embedding",
            k: KNN_RESULTS,
            num_candidates: KNN_CANDIDATES,
            query_vector_builder: {
              text_embedding: {
                model_id: MODEL_ID,
                model_text: queryText,
              },
            },
          },
        });
      
        return result.hits.hits
          .filter((hit) => hit._source !== undefined)
          .map((hit) => ({
            id: hit._id,
            title: hit._source!.title,
            score: hit._score ?? 0,
          }));
      }
      
      export { semanticSearch };
      ```
      
      **Why good:** `query_vector_builder` generates the vector server-side using a deployed ML model -- no need to call the embedding API separately in your application code
      
      **When to use:** When you have a model deployed in Elasticsearch (via Eland or the trained models API). When embedding is done externally (e.g., OpenAI API), pass `query_vector` directly instead.
      
      ---
      
      ## Similarity Metrics
      
      | Metric              | Use Case                                | Score Interpretation                        |
      | ------------------- | --------------------------------------- | ------------------------------------------- |
      | `cosine`            | General purpose, normalized embeddings  | -1 to 1 (mapped to 0-1 for `_score`)        |
      | `dot_product`       | Optimized for unit-length vectors       | Unbounded (requires pre-normalized vectors) |
      | `l2_norm`           | When absolute distance matters          | 0 = identical, higher = more different      |
      | `max_inner_product` | When negative inner products are needed | Unbounded                                   |
      
      **Recommendation:** Use `cosine` for most embedding models (sentence-transformers, OpenAI, Cohere). Use `dot_product` only if you know your vectors are pre-normalized to unit length.
      
      **Gotcha:** `dot_product` on non-normalized vectors gives meaningless scores. Always normalize vectors before indexing if using `dot_product`.
      
      ---
      
      ## Quantization for Memory Efficiency
      
      Dense vectors can be quantized to reduce memory usage.
      
      ```typescript
      async function createQuantizedVectorIndex(client: Client): Promise<void> {
        await client.indices.create({
          index: "articles-quantized",
          mappings: {
            properties: {
              embedding: {
                type: "dense_vector",
                dims: EMBEDDING_DIMS,
                similarity: "cosine",
                index: true,
                index_options: {
                  type: "int8_hnsw", // 4x less memory than float32
                },
              },
            },
          },
        });
      }
      
      export { createQuantizedVectorIndex };
      ```
      
      **Why good:** `int8_hnsw` quantizes float32 vectors to int8, using ~4x less memory with minimal accuracy loss
      
      **Available quantization types:**
      
      - `hnsw` -- Full float32 precision (default)
      - `int8_hnsw` -- 8-bit quantization (~4x memory reduction)
      - `int4_hnsw` -- 4-bit quantization (~8x memory reduction, more accuracy loss)
      - `bbq_hnsw` -- Better binary quantization (auto-selected for dims >= 384 in 8.15+)
      
      ---
      
      _Full skill documentation: [SKILL.md](../SKILL.md) | Quick reference: [reference.md](../reference.md)_
      
  • reference.md 16.7 KB
    # Elasticsearch Quick Reference
    
    > Search DSL, mapping types, aggregation reference, client methods, and decision frameworks. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples.
    
    ---
    
    ## Search DSL Quick Reference
    
    ### Query Types
    
    | Query Type     | Purpose                                   | Context  | Example                                                           |
    | -------------- | ----------------------------------------- | -------- | ----------------------------------------------------------------- |
    | `match`        | Full-text search (analyzed)               | `must`   | `{ match: { title: "search engine" } }`                           |
    | `multi_match`  | Full-text across multiple fields          | `must`   | `{ multi_match: { query: "search", fields: ["title", "body"] } }` |
    | `match_phrase` | Exact phrase in order                     | `must`   | `{ match_phrase: { title: "search engine" } }`                    |
    | `term`         | Exact value (NOT analyzed)                | `filter` | `{ term: { status: "published" } }`                               |
    | `terms`        | Match any of multiple exact values        | `filter` | `{ terms: { status: ["published", "draft"] } }`                   |
    | `range`        | Numeric/date range                        | `filter` | `{ range: { price: { gte: 10, lte: 100 } } }`                     |
    | `exists`       | Field exists and is not null              | `filter` | `{ exists: { field: "description" } }`                            |
    | `bool`         | Combine queries with must/should/filter   | any      | `{ bool: { must: [...], filter: [...] } }`                        |
    | `nested`       | Query nested objects preserving structure | `must`   | `{ nested: { path: "comments", query: { ... } } }`                |
    | `wildcard`     | Wildcard pattern match                    | `filter` | `{ wildcard: { name: { value: "elast*" } } }`                     |
    | `fuzzy`        | Approximate string match                  | `must`   | `{ fuzzy: { name: { value: "serch", fuzziness: "AUTO" } } }`      |
    | `prefix`       | Prefix match                              | `filter` | `{ prefix: { name: { value: "elast" } } }`                        |
    | `ids`          | Match by document IDs                     | `filter` | `{ ids: { values: ["1", "2", "3"] } }`                            |
    
    ### Bool Query Structure
    
    ```
    bool:
      must:      [ ... ]   # Must match, contributes to score
      should:    [ ... ]   # Optional match, boosts score (OR logic unless minimum_should_match set)
      filter:    [ ... ]   # Must match, NO scoring (cached, fast)
      must_not:  [ ... ]   # Must NOT match, NO scoring
    ```
    
    **Key rule:** Use `filter` for yes/no conditions (term, range, exists). Use `must` only when relevance scoring matters (match, multi_match).
    
    ---
    
    ## Mapping Types
    
    | Type           | Purpose                           | Searchable? | Aggregatable? | Sortable? |
    | -------------- | --------------------------------- | ----------- | ------------- | --------- |
    | `text`         | Full-text search (analyzed)       | Yes         | No\*          | No\*      |
    | `keyword`      | Exact values, filtering, sorting  | Yes (exact) | Yes           | Yes       |
    | `integer`      | 32-bit integer                    | Yes         | Yes           | Yes       |
    | `long`         | 64-bit integer                    | Yes         | Yes           | Yes       |
    | `float`        | 32-bit floating point             | Yes         | Yes           | Yes       |
    | `double`       | 64-bit floating point             | Yes         | Yes           | Yes       |
    | `boolean`      | true/false                        | Yes         | Yes           | Yes       |
    | `date`         | ISO 8601 dates or epoch millis    | Yes         | Yes           | Yes       |
    | `object`       | JSON object (flattened)           | Yes         | Depends       | No        |
    | `nested`       | JSON object (preserves structure) | Yes         | Yes (nested)  | No        |
    | `geo_point`    | Latitude/longitude pair           | Geo queries | Yes           | Geo sort  |
    | `dense_vector` | Float vector for kNN search       | kNN queries | No            | No        |
    | `ip`           | IPv4/IPv6 address                 | Yes         | Yes           | Yes       |
    | `completion`   | Autocomplete suggestions          | Suggest API | No            | No        |
    
    \*`text` fields cannot be aggregated or sorted directly. Use a `.keyword` sub-field for aggregation/sort.
    
    ### Multi-Field Pattern
    
    The most common pattern: `text` for search, `keyword` sub-field for aggregation/filtering.
    
    ```json
    {
      "name": {
        "type": "text",
        "analyzer": "standard",
        "fields": {
          "keyword": {
            "type": "keyword",
            "ignore_above": 256
          }
        }
      }
    }
    ```
    
    - Search: `{ "match": { "name": "wireless headphones" } }`
    - Aggregate: `{ "terms": { "field": "name.keyword" } }`
    - Sort: `{ "sort": [{ "name.keyword": "asc" }] }`
    
    ---
    
    ## Aggregation Types
    
    ### Bucket Aggregations (grouping)
    
    | Aggregation      | Purpose                         | Example                                                                |
    | ---------------- | ------------------------------- | ---------------------------------------------------------------------- |
    | `terms`          | Group by exact field values     | `{ terms: { field: "status", size: 100 } }`                            |
    | `range`          | Custom numeric ranges           | `{ range: { field: "price", ranges: [{ to: 50 }, ...] } }`             |
    | `date_histogram` | Time-based buckets              | `{ date_histogram: { field: "created", calendar_interval: "month" } }` |
    | `histogram`      | Numeric interval buckets        | `{ histogram: { field: "price", interval: 10 } }`                      |
    | `filter`         | Single filter bucket            | `{ filter: { term: { status: "active" } } }`                           |
    | `filters`        | Named filter buckets            | `{ filters: { filters: { active: ..., inactive: ... } } }`             |
    | `nested`         | Aggregate within nested objects | `{ nested: { path: "comments" } }`                                     |
    
    ### Metric Aggregations (calculations)
    
    | Aggregation   | Purpose                    | Example                                  |
    | ------------- | -------------------------- | ---------------------------------------- |
    | `avg`         | Average value              | `{ avg: { field: "price" } }`            |
    | `sum`         | Total sum                  | `{ sum: { field: "quantity" } }`         |
    | `min` / `max` | Minimum / Maximum          | `{ min: { field: "price" } }`            |
    | `value_count` | Count of values            | `{ value_count: { field: "price" } }`    |
    | `stats`       | min/max/avg/sum/count      | `{ stats: { field: "price" } }`          |
    | `cardinality` | Approximate distinct count | `{ cardinality: { field: "userId" } }`   |
    | `percentiles` | Distribution percentiles   | `{ percentiles: { field: "latency" } }`  |
    | `top_hits`    | Top documents per bucket   | `{ top_hits: { size: 3, sort: [...] } }` |
    
    ### Pipeline Aggregations (calculations on other aggs)
    
    | Aggregation      | Purpose                                       | buckets_path Example          |
    | ---------------- | --------------------------------------------- | ----------------------------- |
    | `avg_bucket`     | Average across sibling buckets                | `"monthly_sales>total_sales"` |
    | `sum_bucket`     | Sum across sibling buckets                    | `"monthly_sales>total_sales"` |
    | `max_bucket`     | Max across sibling buckets                    | `"monthly_sales>total_sales"` |
    | `derivative`     | Rate of change between buckets                | `"total_sales"`               |
    | `cumulative_sum` | Running total across buckets                  | `"total_sales"`               |
    | `moving_fn`      | Moving function across buckets (script-based) | `"total_sales"` + `script`    |
    | `bucket_sort`    | Sort parent buckets by sub-agg value          | N/A (uses `sort` param)       |
    
    ---
    
    ## Client API Quick Reference
    
    ### Client
    
    | Method                     | Returns                    | Description                        |
    | -------------------------- | -------------------------- | ---------------------------------- |
    | `search<T>(params)`        | `SearchResponse<T>`        | Search documents                   |
    | `index(params)`            | `IndexResponse`            | Index single document              |
    | `get<T>(params)`           | `GetResponse<T>`           | Get document by ID                 |
    | `update(params)`           | `UpdateResponse`           | Partial update document            |
    | `delete(params)`           | `DeleteResponse`           | Delete document by ID              |
    | `bulk(params)`             | `BulkResponse`             | Batch index/update/delete          |
    | `deleteByQuery(params)`    | `DeleteByQueryResponse`    | Delete matching documents          |
    | `updateByQuery(params)`    | `UpdateByQueryResponse`    | Update matching documents          |
    | `mget(params)`             | `MgetResponse<T>`          | Get multiple documents by IDs      |
    | `msearch(params)`          | `MsearchResponse`          | Multi-search in single request     |
    | `openPointInTime(params)`  | `OpenPointInTimeResponse`  | Open PIT for consistent pagination |
    | `closePointInTime(params)` | `ClosePointInTimeResponse` | Close PIT (release resources)      |
    
    ### Indices
    
    | Method                          | Returns                | Description                       |
    | ------------------------------- | ---------------------- | --------------------------------- |
    | `indices.create(params)`        | `CreateIndexResponse`  | Create index with mappings        |
    | `indices.delete(params)`        | `AcknowledgedResponse` | Delete index                      |
    | `indices.exists(params)`        | `boolean`              | Check if index exists             |
    | `indices.putMapping(params)`    | `AcknowledgedResponse` | Add new fields to mapping         |
    | `indices.getMapping(params)`    | `GetMappingResponse`   | Get current mappings              |
    | `indices.putSettings(params)`   | `AcknowledgedResponse` | Update index settings             |
    | `indices.refresh(params)`       | `RefreshResponse`      | Force refresh (make docs visible) |
    | `indices.putAlias(params)`      | `AcknowledgedResponse` | Create index alias                |
    | `indices.deleteAlias(params)`   | `AcknowledgedResponse` | Delete index alias                |
    | `indices.updateAliases(params)` | `AcknowledgedResponse` | Atomic alias swap                 |
    
    ### Helpers
    
    | Method                            | Returns            | Description                         |
    | --------------------------------- | ------------------ | ----------------------------------- |
    | `helpers.bulk(params)`            | `BulkStats`        | Bulk with batching + retries        |
    | `helpers.search(params)`          | `T[]`              | Returns only `_source` docs         |
    | `helpers.scrollSearch(params)`    | `AsyncIterable`    | Scroll with async iteration (pages) |
    | `helpers.scrollDocuments(params)` | `AsyncIterable<T>` | Scroll with async iteration (docs)  |
    | `helpers.msearch()`               | `MsearchHelper`    | Batched multi-search                |
    
    ---
    
    ## Search Parameters
    
    | Parameter          | Type            | Default   | Description                                             |
    | ------------------ | --------------- | --------- | ------------------------------------------------------- |
    | `index`            | `string`        | required  | Index name or pattern                                   |
    | `query`            | `object`        | match_all | Query DSL                                               |
    | `size`             | `number`        | `10`      | Number of hits to return                                |
    | `from`             | `number`        | `0`       | Offset for pagination (max 10,000 by default)           |
    | `sort`             | `array`         | `_score`  | Sort order                                              |
    | `_source`          | `boolean/array` | `true`    | Fields to include in `_source`                          |
    | `aggs`             | `object`        | none      | Aggregations                                            |
    | `highlight`        | `object`        | none      | Highlight matching terms                                |
    | `search_after`     | `array`         | none      | Cursor for deep pagination                              |
    | `pit`              | `object`        | none      | Point in Time for consistent pagination                 |
    | `knn`              | `object`        | none      | kNN vector search                                       |
    | `track_total_hits` | `bool/number`   | `10000`   | Track exact total hits (true = exact, number = up to N) |
    | `timeout`          | `string`        | none      | Search timeout (e.g., "5s")                             |
    | `explain`          | `boolean`       | `false`   | Include score explanation                               |
    
    ---
    
    ## Common Index Settings
    
    | Setting                            | Default | Description                                        |
    | ---------------------------------- | ------- | -------------------------------------------------- |
    | `index.number_of_shards`           | `1`     | Primary shards (set at creation, immutable)        |
    | `index.number_of_replicas`         | `1`     | Replica shards (can be changed dynamically)        |
    | `index.refresh_interval`           | `"1s"`  | How often new docs become searchable               |
    | `index.max_result_window`          | `10000` | Max `from + size` for pagination                   |
    | `index.mapping.total_fields.limit` | `1000`  | Max fields in mapping (prevents mapping explosion) |
    
    ---
    
    ## Anti-Patterns
    
    ### Dynamic Mapping Without Explicit Types
    
    ```typescript
    // ANTI-PATTERN: No mappings defined
    await client.index({
      index: "events",
      document: { timestamp: "2024-01-15" }, // Mapped as "date"
    });
    // Later:
    await client.index({
      index: "events",
      document: { timestamp: "not-a-date" }, // FAILS: mapper_parsing_exception
    });
    ```
    
    **Why it's wrong:** Dynamic mapping inferred `timestamp` as `date` from the first document. The second document with a non-date string fails permanently. You cannot change the mapping -- you must reindex.
    
    **What to do instead:** Define explicit mappings before indexing any documents.
    
    ---
    
    ### Using term Query on text Fields
    
    ```typescript
    // ANTI-PATTERN: term query on analyzed field
    const result = await client.search({
      index: "products",
      query: { term: { name: "Wireless Headphones" } }, // Returns 0 results!
    });
    ```
    
    **Why it's wrong:** `text` fields are analyzed -- "Wireless Headphones" is stored as tokens ["wireless", "headphones"]. The `term` query does NOT analyze the input, so it looks for the exact un-analyzed string "Wireless Headphones" which doesn't exist.
    
    **What to do instead:** Use `match` for text fields, or query the `.keyword` sub-field with `term`.
    
    ---
    
    ### Refresh on Every Write
    
    ```typescript
    // ANTI-PATTERN: Forcing refresh in request handler
    app.post("/products", async (req, res) => {
      await client.index({
        index: "products",
        document: req.body,
        refresh: true, // Forces segment refresh on EVERY write
      });
      res.json({ success: true });
    });
    ```
    
    **Why it's wrong:** Each refresh creates a new Lucene segment. Under write-heavy load, this creates thousands of tiny segments that must be merged, consuming CPU and I/O. The default 1-second refresh interval batches writes into efficient segments.
    
    **What to do instead:** Let the default refresh interval handle it. Use `refresh: "wait_for"` only in tests and seed scripts.
    
    ---
    
    ## Production Checklist
    
    ### Mappings
    
    - [ ] Explicit mappings defined for all fields before first document
    - [ ] `text` + `keyword` multi-field on fields needing both search and aggregation
    - [ ] `nested` type used for arrays of objects that need field-level query association
    - [ ] `dynamic: "strict"` or `dynamic: false` to prevent unexpected field additions
    
    ### Performance
    
    - [ ] `filter` context used for all non-scoring queries (term, range, exists)
    - [ ] `size: 0` when only aggregations are needed
    - [ ] Bulk API used for all batch operations
    - [ ] No `refresh: true` in production request handlers
    - [ ] `_source` filtering used to return only needed fields
    
    ### Pagination
    
    - [ ] `search_after` + PIT used for results beyond 10,000
    - [ ] Tiebreaker field included in sort for `search_after`
    - [ ] PIT closed in finally block after use
    
    ### Reliability
    
    - [ ] Bulk response `errors` flag checked for partial failures
    - [ ] `_source` checked for undefined (can be missing if disabled)
    - [ ] Connection error handling with retries
    - [ ] Health check endpoint using `client.ping()` or `client.info()`
    
    ---
    
    _Full skill documentation: [SKILL.md](SKILL.md) | Examples: [examples/](examples/)_
    
  • SKILL.md 19.4 KB
    ---
    name: api-search-elasticsearch
    description: Elasticsearch patterns -- client setup, index management, search DSL, aggregations, vector search, bulk operations, deep pagination
    ---
    
    # Elasticsearch Patterns
    
    > **Quick Guide:** Use `@elastic/elasticsearch` (v8.x/v9.x) as the TypeScript client. Elasticsearch is **near real-time** -- documents are NOT searchable immediately after indexing; they become visible after a refresh (default: every 1 second on active indices). You MUST define explicit mappings before indexing -- dynamic mapping infers types from the first document, and mismatched types in later documents cause hard failures you cannot fix without reindexing. Use `search_after` + Point in Time (PIT) for deep pagination -- NOT `from`/`size` beyond 10,000 hits and NOT the scroll API (deprecated for search). Use the `bulk` API or `client.helpers.bulk()` for any batch operation -- never loop individual index calls.
    
    ---
    
    <critical_requirements>
    
    ## CRITICAL: Before Using This Skill
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST define explicit index mappings BEFORE indexing documents -- dynamic mapping infers types from the first document, and if a later document sends a different type for the same field, indexing fails with a `mapper_parsing_exception` that CANNOT be fixed without reindexing into a new index)**
    
    **(You MUST use the `bulk` API for batch operations -- looping individual `client.index()` calls is orders of magnitude slower and can overwhelm the cluster with HTTP connections)**
    
    **(You MUST NOT use `from`/`size` pagination beyond 10,000 results -- Elasticsearch throws `Result window is too large` by default; use `search_after` + PIT instead)**
    
    **(You MUST NOT use `refresh: true` or `refresh: "wait_for"` in production request handlers -- forcing a refresh on every write degrades cluster performance; let the default 1-second refresh interval handle it)**
    
    </critical_requirements>
    
    ---
    
    ## Examples
    
    - [Core Patterns](examples/core.md) -- Client setup, index management, document CRUD, search basics, TypeScript integration
    - [Aggregations](examples/aggregations.md) -- Terms, range, date_histogram, nested, pipeline aggregations
    - [Vector Search](examples/vector-search.md) -- Dense vector fields, kNN queries, hybrid search, similarity metrics
    - [Pagination](examples/pagination.md) -- from/size, search_after, Point in Time, scroll helpers
    - [Bulk Operations](examples/bulk-operations.md) -- Bulk API, bulk helper, reindexing patterns
    
    **Additional resources:**
    
    - [reference.md](reference.md) -- Search DSL cheat sheet, mapping types, aggregation reference, decision frameworks
    
    ---
    
    **Auto-detection:** Elasticsearch, elasticsearch, @elastic/elasticsearch, client.search, client.index, client.bulk, client.indices.create, client.indices.putMapping, dense_vector, knn, search_after, point in time, openPIT, aggregations, aggs, bool query, match query, term query, multi_match, nested query, range query, client.helpers.bulk, client.helpers.scrollSearch, BulkResponse, SearchResponse, MappingProperty
    
    **When to use:**
    
    - Full-text search with advanced relevance tuning (BM25, custom analyzers, boosting)
    - Aggregations and analytics (terms, histograms, pipeline aggregations)
    - Vector/semantic search with kNN on dense_vector fields
    - Log and event data search with time-based queries
    - Complex structured queries combining bool, nested, range, and geo filters
    - Search across large datasets requiring deep pagination (search_after + PIT)
    
    **Key patterns covered:**
    
    - Client initialization and connection management
    - Index management with explicit mappings and settings
    - Document CRUD (index, get, update, delete)
    - Search DSL (match, term, bool, range, nested, multi_match)
    - Aggregations (terms, range, date_histogram, nested, pipeline)
    - Full-text analysis (custom analyzers, tokenizers, filters)
    - Vector search (dense_vector, kNN, hybrid text+vector)
    - Bulk operations and reindexing
    - Deep pagination (search_after + PIT)
    
    **When NOT to use:**
    
    - Simple keyword search on small datasets (client-side filtering or database LIKE queries are simpler)
    - Primary data store (Elasticsearch is a search engine, not a database -- always have a source of truth elsewhere)
    - Strong consistency requirements (Elasticsearch is eventually consistent by design)
    - Simple autocomplete on a small list (a prefix trie or client-side filter is simpler)
    
    ---
    
    <philosophy>
    
    ## Philosophy
    
    Elasticsearch is a **distributed search and analytics engine** built on Apache Lucene. It excels at full-text search, structured queries, aggregations, and vector search at scale. Core principles:
    
    1. **Near real-time, not real-time** -- Documents are indexed into segments. A refresh (default: every 1 second on active indices) makes new segments searchable. Do not expect immediate consistency after writes.
    2. **Mappings are immutable** -- Once a field type is set (text, keyword, integer, etc.), it cannot be changed. Wrong types require reindexing into a new index. Always define mappings explicitly before first document.
    3. **Search engine, not database** -- Elasticsearch should not be your source of truth. Always have a primary database and sync to Elasticsearch for search.
    4. **Bulk everything** -- The bulk API amortizes HTTP overhead across thousands of operations. Never loop individual index/update/delete calls.
    5. **Pagination has limits** -- `from`/`size` is capped at 10,000 hits by default (`index.max_result_window`). Deep pagination requires `search_after` + Point in Time (PIT). The scroll API is deprecated for search use cases.
    6. **Text vs keyword matters** -- `text` fields are analyzed (tokenized, lowercased) for full-text search. `keyword` fields are exact-match only. Getting this wrong means either broken search or broken aggregations/filters.
    
    </philosophy>
    
    ---
    
    <patterns>
    
    ## Core Patterns
    
    ### Pattern 1: Client Setup
    
    Initialize the client with node URL and authentication. The client supports Elastic Cloud, API keys, basic auth, and bearer tokens.
    
    ```typescript
    // Good Example -- Typed client setup with environment validation
    import { Client } from "@elastic/elasticsearch";
    
    function createElasticsearchClient(): Client {
      const node = process.env.ELASTICSEARCH_URL;
      if (!node) {
        throw new Error("ELASTICSEARCH_URL environment variable is required");
      }
    
      return new Client({
        node,
        auth: {
          apiKey: process.env.ELASTICSEARCH_API_KEY ?? "",
        },
      });
    }
    
    export { createElasticsearchClient };
    ```
    
    **Why good:** Environment variable validation, named export, API key auth (preferred over basic auth in production)
    
    ```typescript
    // Bad Example -- Hardcoded credentials
    import { Client } from "@elastic/elasticsearch";
    const client = new Client({
      node: "http://localhost:9200",
      auth: { username: "elastic", password: "changeme" },
    });
    ```
    
    **Why bad:** Hardcoded node URL and credentials leak in version control, basic auth with default password
    
    See [examples/core.md](examples/core.md) for Elastic Cloud setup, health checks, and child clients.
    
    ---
    
    ### Pattern 2: Index with Explicit Mappings
    
    Always define mappings before indexing. Dynamic mapping infers types from the first document -- if wrong, you must reindex.
    
    ```typescript
    // Good Example -- Explicit mappings with text + keyword multi-field
    const INDEX_NAME = "products";
    
    await client.indices.create({
      index: INDEX_NAME,
      settings: {
        number_of_replicas: 1,
        refresh_interval: "1s",
      },
      mappings: {
        properties: {
          name: {
            type: "text",
            fields: { keyword: { type: "keyword" } },
          },
          description: { type: "text", analyzer: "standard" },
          price: { type: "float" },
          categories: { type: "keyword" },
          inStock: { type: "boolean" },
          createdAt: { type: "date" },
        },
      },
    });
    ```
    
    **Why good:** Explicit types prevent mapping conflicts, `text` + `keyword` multi-field allows both full-text search and exact-match filtering/aggregation on `name`
    
    ```typescript
    // Bad Example -- No mappings, relying on dynamic mapping
    await client.indices.create({ index: "products" });
    await client.index({
      index: "products",
      document: { price: "29.99" }, // Oops -- "29.99" is a string, mapped as text
    });
    // All future numeric price documents will fail with mapper_parsing_exception
    ```
    
    **Why bad:** Dynamic mapping infers `price` as `text` from the string "29.99", and this mapping is immutable -- all future documents with numeric `price` will fail
    
    See [examples/core.md](examples/core.md) for analysis settings, custom analyzers, and mapping migration.
    
    ---
    
    ### Pattern 3: Search with Bool Query
    
    The bool query is the workhorse of Elasticsearch. It combines `must`, `should`, `must_not`, and `filter` clauses.
    
    ```typescript
    // Good Example -- Bool query with filter context for exact matches
    const MIN_PRICE = 10;
    const MAX_PRICE = 100;
    
    const result = await client.search<Product>({
      index: INDEX_NAME,
      query: {
        bool: {
          must: [{ match: { description: "wireless headphones" } }],
          filter: [
            { range: { price: { gte: MIN_PRICE, lte: MAX_PRICE } } },
            { term: { inStock: true } },
          ],
        },
      },
      size: 20,
    });
    // result.hits.hits[0]._source is typed as Product | undefined
    ```
    
    **Why good:** `filter` context for exact matches (no scoring overhead, cacheable), `must` for full-text relevance scoring, named constants for range values, typed search with generic
    
    ```typescript
    // Bad Example -- Everything in must (no filter context)
    const result = await client.search({
      index: "products",
      query: {
        bool: {
          must: [
            { match: { description: "wireless headphones" } },
            { range: { price: { gte: 10, lte: 100 } } }, // Wasteful scoring
            { term: { inStock: true } }, // Wasteful scoring
          ],
        },
      },
    });
    ```
    
    **Why bad:** Range and term queries in `must` waste CPU on relevance scoring for yes/no conditions; `filter` context skips scoring and enables Elasticsearch's filter cache
    
    See [examples/core.md](examples/core.md) for multi_match, nested queries, and function_score.
    
    ---
    
    ### Pattern 4: Aggregations
    
    Aggregations compute analytics over search results. `terms` for category counts, `range` for bucketing, `date_histogram` for time series.
    
    ```typescript
    // Good Example -- Terms aggregation with sub-aggregation
    const AGGREGATION_SIZE = 50;
    
    const result = await client.search({
      index: INDEX_NAME,
      size: 0, // No hits needed, only aggregations
      aggs: {
        categories: {
          terms: { field: "categories", size: AGGREGATION_SIZE },
          aggs: {
            avgPrice: { avg: { field: "price" } },
          },
        },
      },
    });
    // result.aggregations?.categories.buckets -> [{ key: "electronics", doc_count: 42, avgPrice: { value: 89.5 } }]
    ```
    
    **Why good:** `size: 0` skips hits when only aggregations are needed (faster), nested sub-aggregation for metrics per bucket, named constant for aggregation size
    
    See [examples/aggregations.md](examples/aggregations.md) for date_histogram, range, nested, and pipeline aggregations.
    
    ---
    
    ### Pattern 5: Bulk Operations
    
    The bulk API batches multiple index/update/delete operations in a single request. Use `client.helpers.bulk()` for the best developer experience.
    
    ```typescript
    // Good Example -- Bulk helper with async generator
    const result = await client.helpers.bulk<Product>({
      datasource: products,
      onDocument(doc) {
        return { index: { _index: INDEX_NAME, _id: doc.productId } };
      },
      refreshOnCompletion: INDEX_NAME,
    });
    // result.total, result.successful, result.failed
    ```
    
    **Why good:** Bulk helper handles batching, concurrency, retries, and back-pressure automatically; `refreshOnCompletion` triggers one refresh at the end instead of per-document
    
    ```typescript
    // Bad Example -- Looping individual index calls
    for (const product of products) {
      await client.index({
        index: "products",
        document: product,
        refresh: true, // Refresh after EVERY document!
      });
    }
    ```
    
    **Why bad:** N individual HTTP requests instead of 1 bulk request, `refresh: true` on every document causes N segment refreshes (devastating to cluster performance)
    
    See [examples/bulk-operations.md](examples/bulk-operations.md) for error handling, update operations, and reindexing.
    
    ---
    
    ### Pattern 6: Deep Pagination with search_after + PIT
    
    `from`/`size` is limited to 10,000 hits. For deep pagination, use `search_after` with a Point in Time (PIT) for consistent results.
    
    ```typescript
    // Good Example -- search_after with PIT
    const PIT_KEEP_ALIVE = "1m";
    
    const pit = await client.openPointInTime({
      index: INDEX_NAME,
      keep_alive: PIT_KEEP_ALIVE,
    });
    
    let searchAfter: Array<string | number> | undefined;
    let allHits: Product[] = [];
    
    while (true) {
      const result = await client.search<Product>({
        pit: { id: pit.id, keep_alive: PIT_KEEP_ALIVE },
        sort: [{ createdAt: "desc" }, { _id: "asc" }], // Tiebreaker!
        size: 100,
        ...(searchAfter ? { search_after: searchAfter } : {}),
      });
    
      const hits = result.hits.hits;
      if (hits.length === 0) break;
    
      allHits = allHits.concat(
        hits.filter((h) => h._source !== undefined).map((h) => h._source!),
      );
      searchAfter = hits[hits.length - 1].sort as Array<string | number>;
    }
    
    await client.closePointInTime({ id: pit.id });
    ```
    
    **Why good:** PIT ensures consistent snapshot across pages, tiebreaker `_id` prevents missing/duplicate documents, `keep_alive` refreshed on each request
    
    See [examples/pagination.md](examples/pagination.md) for from/size limits, scroll helpers, and pagination decision framework.
    
    </patterns>
    
    ---
    
    <decision_framework>
    
    ## Decision Framework
    
    ### Which Query Type?
    
    ```
    What kind of search do I need?
    -- Full-text relevance search? -> match / multi_match in must
    -- Exact value filtering? -> term / terms / range in filter
    -- Combining text + filters? -> bool query (must for text, filter for exact)
    -- Fuzzy matching? -> match with fuzziness: "AUTO"
    -- Phrase matching? -> match_phrase
    -- Complex nested objects? -> nested query with path
    -- Vector similarity? -> knn with dense_vector field
    -- Text + vector hybrid? -> query + knn in same request
    ```
    
    ### Pagination Strategy?
    
    ```
    How deep do results go?
    -- Under 10,000 total? -> from/size (simplest)
    -- Over 10,000 hits? -> search_after + PIT (recommended)
    -- Bulk data export? -> client.helpers.scrollSearch() or scrollDocuments()
    -- Real-time infinite scroll? -> search_after (no PIT needed for forward-only)
    ```
    
    ### text vs keyword?
    
    ```
    What will I do with this field?
    -- Full-text search (tokenized, relevance)? -> text
    -- Exact match, filtering, aggregations, sorting? -> keyword
    -- Both? -> Multi-field: { type: "text", fields: { keyword: { type: "keyword" } } }
    -- Neither (just stored, never queried)? -> { type: "keyword", index: false }
    ```
    
    ### Filter vs Must?
    
    ```
    Does relevance scoring matter for this clause?
    -- YES (affects result order) -> must
    -- NO (binary yes/no filter) -> filter (cached, no scoring overhead)
    -- Exclude documents -> must_not (in filter context)
    -- Boost if present (optional) -> should with minimum_should_match: 0
    ```
    
    </decision_framework>
    
    ---
    
    <red_flags>
    
    ## RED FLAGS
    
    **High Priority Issues:**
    
    - Relying on dynamic mapping without explicit mappings -- wrong type inference causes `mapper_parsing_exception` that requires reindexing to fix
    - Using `from`/`size` beyond 10,000 results -- Elasticsearch throws `Result window is too large`; use `search_after` + PIT
    - Looping individual `client.index()` calls instead of `client.bulk()` or `client.helpers.bulk()` -- orders of magnitude slower, can overwhelm the cluster
    - Using `refresh: true` or `refresh: "wait_for"` in production request handlers -- forces a segment refresh on every write, degrades cluster performance under load
    
    **Medium Priority Issues:**
    
    - Putting exact-match conditions (term, range) in `must` instead of `filter` -- wastes CPU on scoring, misses filter cache
    - Using `text` type for fields that need exact matching or aggregation -- text fields are analyzed (tokenized), making aggregations return individual tokens instead of full values
    - Not including a tiebreaker field in `sort` when using `search_after` -- documents with identical sort values may be skipped or duplicated across pages
    - Missing `_source` check -- `hit._source` can be `undefined` if `_source` is disabled or fields are excluded; always handle this
    
    **Gotchas & Edge Cases:**
    
    - **Mapping types are immutable** -- once a field is mapped as `text`, you cannot change it to `keyword`. The only fix is to create a new index with correct mappings and reindex all documents
    - **`text` vs `keyword` confusion** -- `text` fields are tokenized ("New York" becomes ["new", "york"]). Aggregating on a `text` field gives you individual tokens, not full values. Use `keyword` or a `.keyword` sub-field for aggregations
    - **`match` vs `term` on text fields** -- `term` on a `text` field often returns no results because `term` does NOT analyze the query but the field value IS analyzed (e.g., term "New York" won't match the analyzed tokens "new" and "york")
    - **Near real-time delay** -- after indexing, documents are NOT searchable until the next refresh (default: 1 second). Tests that index then immediately search must use `refresh: "wait_for"` or explicit `client.indices.refresh()`
    - **Default `index.max_result_window` is 10,000** -- increasing this is possible but NOT recommended; deep pagination with `from`/`size` holds all skipped results in memory
    - **Nested objects require `nested` mapping type** -- arrays of objects are flattened by default, losing the association between fields within each object. If you need to query "color: red AND size: large" on the same object in an array, use `nested`
    - **Aggregation on `_id` is disabled by default** (8.x+) -- use a separate `id` field if you need to aggregate by document ID
    - **`_score` is null in filter context** -- clauses in `filter` do not contribute to scoring; if you need scoring, use `must`
    - **Bulk API partial failures** -- a bulk request can succeed overall but have individual failures. Always check `result.errors` and iterate `result.items` to find failed operations
    - **Scroll API is deprecated for search** -- use `search_after` + PIT for deep pagination. Scroll is still valid for one-time data export but consumes cluster resources (open search contexts)
    - **PIT must be closed** -- failing to close Point in Time contexts leaks resources on the cluster; always close in a finally block
    
    </red_flags>
    
    ---
    
    <critical_reminders>
    
    ## CRITICAL REMINDERS
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST define explicit index mappings BEFORE indexing documents -- dynamic mapping infers types from the first document, and if a later document sends a different type for the same field, indexing fails with a `mapper_parsing_exception` that CANNOT be fixed without reindexing into a new index)**
    
    **(You MUST use the `bulk` API for batch operations -- looping individual `client.index()` calls is orders of magnitude slower and can overwhelm the cluster with HTTP connections)**
    
    **(You MUST NOT use `from`/`size` pagination beyond 10,000 results -- Elasticsearch throws `Result window is too large` by default; use `search_after` + PIT instead)**
    
    **(You MUST NOT use `refresh: true` or `refresh: "wait_for"` in production request handlers -- forcing a refresh on every write degrades cluster performance; let the default 1-second refresh interval handle it)**
    
    **Failure to follow these rules will cause mapping conflicts, pagination failures, cluster performance degradation, and silent data loss.**
    
    </critical_reminders>
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related