Claude Skill

ai-orchestration-llamaindex

LlamaIndex.TS data framework for RAG, indexing, retrieval, query engines, chat engines, and agentic workflows in TypeScript

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download agents-inc-skills-dist_plugins_ai-orchestration-llamaindex_skills_ai-orchestration-llamaindex-3a51ef5.zip · 19 KB
Part of agents-inc/skills — 130 skills

Install

skills CLI npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/ai-orchestration-llamaindex/skills/ai-orchestration-llamaindex
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
Git git clone https://github.com/agents-inc/skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

LlamaIndex.TS Patterns

Quick Guide: LlamaIndex.TS is a data framework for building context-aware LLM applications in TypeScript. Use Settings singleton to configure LLM and embedding models globally. Load documents with SimpleDirectoryReader, chunk with SentenceSplitter, index with VectorStoreIndex.fromDocuments(), and query with index.asQueryEngine(). For agents, use agent() from @llamaindex/workflow with tool() definitions using Zod schemas. All core operations are async -- every function returns a Promise. The llamaindex package re-exports most things, but LLM providers require separate packages like @llamaindex/openai or @llamaindex/ollama.


<critical_requirements>

CRITICAL: Before Using This Skill

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST configure Settings.llm and Settings.embedModel before any indexing or querying -- the Settings singleton is lazily initialized and defaults to OpenAI, which will fail without an API key)

(You MUST await all LlamaIndex operations -- fromDocuments(), asQueryEngine(), query(), chat(), loadData() are ALL async)

(You MUST install provider packages separately -- @llamaindex/openai, @llamaindex/ollama, @llamaindex/anthropic are NOT included in the base llamaindex package)

(You MUST use storageContextFromDefaults({ persistDir }) to persist indexes -- without persistence, indexes are rebuilt from scratch on every restart)

(You MUST never hardcode API keys -- use environment variables and dotenv/config)

</critical_requirements>


Auto-detection: LlamaIndex, llamaindex, VectorStoreIndex, SimpleDirectoryReader, Settings.llm, Settings.embedModel, asQueryEngine, asChatEngine, ContextChatEngine, SentenceSplitter, storageContextFromDefaults, @llamaindex/openai, @llamaindex/ollama, @llamaindex/workflow, FunctionTool, QueryEngineTool, agentStreamEvent

When to use:

  • Building RAG (Retrieval-Augmented Generation) applications with custom documents
  • Loading, chunking, and indexing documents for LLM consumption
  • Creating query engines that answer questions from indexed data
  • Building chat interfaces with conversation memory over your data
  • Implementing agentic RAG with tool-calling agents that query indexes
  • Working with multiple data sources (files, PDFs, markdown, code)
  • Persisting vector indexes to avoid re-indexing on every restart

Key patterns covered:

  • Settings singleton for LLM and embedding model configuration
  • Document loading with SimpleDirectoryReader and custom readers
  • VectorStoreIndex creation, persistence, and querying
  • Query engines and chat engines
  • Agent creation with agent() and tool() using Zod schemas
  • Text splitting and chunking strategies
  • Streaming responses from query and chat engines
  • Storage context and index persistence

When NOT to use:

  • Simple one-shot LLM calls without document context -- use the LLM provider SDK directly
  • Applications that only need embeddings without indexing -- use the embedding API directly
  • Client-side / browser applications -- LlamaIndex.TS is server-side focused (Node.js >= 20)

Examples Index




<decision_framework>

Decision Framework

Which Index Type to Use

What is your use case?
+-- Semantic search over documents -> VectorStoreIndex (most common)
+-- Summarization of all documents -> SummaryIndex
+-- Both search AND summarization -> Create both, use as separate tools in an agent
+-- Hierarchical document structure -> Use MarkdownNodeParser + VectorStoreIndex

Query Engine vs Chat Engine vs Agent

How should users interact with your data?
+-- Single question, single answer -> Query Engine (index.asQueryEngine())
+-- Multi-turn conversation -> Chat Engine (ContextChatEngine)
+-- Multiple tools/indexes + reasoning -> Agent (agent() from @llamaindex/workflow)
+-- Complex multi-step workflow -> Multi-agent with handoffs

Which LLM Provider

Which LLM provider are you using?
+-- OpenAI -> npm install @llamaindex/openai
+-- Anthropic -> npm install @llamaindex/anthropic
+-- Local (Ollama) -> npm install @llamaindex/ollama
+-- Groq -> npm install @llamaindex/groq
+-- Google Gemini -> npm install @llamaindex/gemini

Chunk Size Selection

What kind of documents are you indexing?
+-- Short Q&A pairs -> chunkSize: 256-512
+-- Technical documentation -> chunkSize: 512-1024
+-- Long narratives/reports -> chunkSize: 1024-2048
+-- Code files -> Use CodeSplitter (AST-aware)
+-- Markdown -> Use MarkdownNodeParser (structure-aware)

</decision_framework>


<red_flags>

RED FLAGS

High Priority Issues:

  • Not configuring Settings.llm before indexing/querying -- defaults to OpenAI, fails silently without API key
  • Forgetting to await async operations -- fromDocuments(), query(), chat() all return Promises
  • Rebuilding indexes on every request instead of persisting with storageContextFromDefaults
  • Hardcoding API keys instead of using environment variables
  • Installing only llamaindex without provider packages (@llamaindex/openai, etc.)

Medium Priority Issues:

  • Using default chunk size (1024) without considering document characteristics -- causes poor retrieval
  • Not setting similarityTopK on retrievers -- default may return too few or too many results
  • Ignoring the response sourceNodes -- they contain the retrieved context for debugging and citations
  • Creating a new SimpleDirectoryReader per request instead of caching the loaded documents
  • Not handling the case where response.message.content might be empty on retrieval failure

Common Mistakes:

  • Confusing asQueryEngine() (single question) with ContextChatEngine (multi-turn conversation)
  • Using VectorStoreIndex.fromDocuments() when you should use VectorStoreIndex.init() to load from storage
  • Importing openai from llamaindex instead of @llamaindex/openai -- the llamaindex package may re-export some things but provider-specific imports are more reliable
  • Passing messages array to query() -- query engines take { query: string }, not a messages array
  • Using index.asQueryEngine() multiple times instead of storing the engine reference

Gotchas & Edge Cases:

  • Settings is a global singleton -- setting it in one module affects all others. Override locally by passing llm directly to constructors when you need different models for different operations.
  • SimpleDirectoryReader only works on Node.js -- it uses fs internally. For edge/serverless, load documents differently or use LlamaParse.
  • storageContextFromDefaults creates four JSON files in the persist directory (docstore.json, graph_store.json, index_store.json, vector_store.json). If any are corrupted, delete the directory and re-index.
  • Node.js >= 20 is required. Some modules use Web Stream API (ReadableStream, WritableStream), so add "DOM.AsyncIterable" to tsconfig.json lib if you get type errors.
  • tsconfig.json must use "moduleResolution": "bundler" or "nodenext" -- the classic "node" resolution will fail to resolve LlamaIndex sub-packages.
  • Default tokenizer is slow -- install gpt-tokenizer for 60x faster tokenization.
  • SentenceSplitter chunk size is in tokens, not characters. A 512-token chunk is roughly 2000 characters.
  • The llamaindex package is large (~2MB+). For production, consider importing specific sub-packages to reduce bundle size.
  • VectorStoreIndex.fromDocuments() makes embedding API calls for every chunk. For large document sets, this can be expensive. Monitor costs.
  • Chat engine conversation history grows unbounded -- implement history pruning for long-running sessions.

</red_flags>


<critical_reminders>

CRITICAL REMINDERS

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST configure Settings.llm and Settings.embedModel before any indexing or querying -- the Settings singleton is lazily initialized and defaults to OpenAI, which will fail without an API key)

(You MUST await all LlamaIndex operations -- fromDocuments(), asQueryEngine(), query(), chat(), loadData() are ALL async)

(You MUST install provider packages separately -- @llamaindex/openai, @llamaindex/ollama, @llamaindex/anthropic are NOT included in the base llamaindex package)

(You MUST use storageContextFromDefaults({ persistDir }) to persist indexes -- without persistence, indexes are rebuilt from scratch on every restart)

(You MUST never hardcode API keys -- use environment variables and dotenv/config)

Failure to follow these rules will produce broken RAG pipelines, wasted embedding API credits, or cryptic runtime errors.

</critical_reminders>

Files (skills)
  • examples
    • agents.md 7.3 KB
      # LlamaIndex.TS -- Agents & Tools Examples
      
      > Agent creation, tool definitions, multi-agent workflows, and QueryEngineTool patterns. See [core.md](core.md) for basic setup.
      
      **Prerequisites:** Understand Settings configuration and VectorStoreIndex from [core.md](core.md).
      
      **Related examples:**
      
      - [core.md](core.md) -- Setup, indexing, query engines
      - [chat-streaming.md](chat-streaming.md) -- Chat engines, streaming
      
      ---
      
      ## Basic Agent with Custom Tools
      
      ```typescript
      // agents/weather-agent.ts
      import "dotenv/config";
      import { tool, Settings } from "llamaindex";
      import { agent } from "@llamaindex/workflow";
      import { openai } from "@llamaindex/openai";
      import { z } from "zod";
      
      Settings.llm = openai({ model: "gpt-4o" });
      
      const getWeather = tool({
        name: "getWeather",
        description: "Get current weather for a city",
        parameters: z.object({
          city: z.string({ description: "City name (e.g. 'Paris')" }),
          unit: z
            .enum(["celsius", "fahrenheit"])
            .default("celsius")
            .describe("Temperature unit"),
        }),
        execute: async ({ city, unit }) => {
          // Replace with actual API call
          return {
            city,
            temperature: unit === "celsius" ? 22 : 72,
            condition: "sunny",
          };
        },
      });
      
      const weatherAgent = agent({
        tools: [getWeather],
      });
      
      const result = await weatherAgent.run("What's the weather in Paris?");
      console.log(result.data);
      ```
      
      **Why good:** Zod schema with descriptions guides the model, default values reduce friction, typed execute function
      
      ---
      
      ## Agent with Streaming Output
      
      ```typescript
      import { agentStreamEvent } from "@llamaindex/workflow";
      
      const events = weatherAgent.runStream(
        "Compare weather in Paris and Tokyo, recommend which to visit",
      );
      
      for await (const event of events) {
        if (agentStreamEvent.include(event)) {
          process.stdout.write(event.data.delta);
        }
      }
      console.log(); // Final newline
      ```
      
      **Why good:** Type-safe event filtering, progressive output, simple for-await pattern
      
      ---
      
      ## Agent with Structured Output
      
      ```typescript
      import { z } from "zod";
      
      const travelRecommendation = z.object({
        city: z.string(),
        reason: z.string(),
        bestSeason: z.string(),
        averageCost: z.number(),
      });
      
      const result = await weatherAgent.run("Recommend a city to visit in summer", {
        responseFormat: travelRecommendation,
      });
      
      // result.data.object is typed as { city, reason, bestSeason, averageCost }
      console.log(result.data.object);
      ```
      
      **Why good:** Zod schema validates and types the response, structured data extraction
      
      ---
      
      ## RAG Agent with QueryEngineTool
      
      ```typescript
      // agents/rag-agent.ts
      import {
        VectorStoreIndex,
        SimpleDirectoryReader,
        QueryEngineTool,
        Settings,
      } from "llamaindex";
      import { agent } from "@llamaindex/workflow";
      import { openai, OpenAIEmbedding } from "@llamaindex/openai";
      
      Settings.llm = openai({ model: "gpt-4o" });
      Settings.embedModel = new OpenAIEmbedding({ model: "text-embedding-3-small" });
      
      // Create index from documents
      const documents = await new SimpleDirectoryReader().loadData({
        directoryPath: "./data/product-docs",
      });
      const index = await VectorStoreIndex.fromDocuments(documents);
      
      // Wrap query engine as a tool
      const docSearchTool = new QueryEngineTool({
        queryEngine: index.asQueryEngine(),
        metadata: {
          name: "product_docs",
          description:
            "Search product documentation for features, pricing, and technical specs",
        },
      });
      
      const ragAgent = agent({
        tools: [docSearchTool],
        systemPrompt:
          "You are a product support agent. Use the product_docs tool to answer questions accurately. Cite sources.",
      });
      
      const response = await ragAgent.run("What are the pricing tiers?");
      console.log(response.data);
      ```
      
      **Why good:** QueryEngineTool bridges RAG and agents, descriptive metadata for accurate tool routing, system prompt sets behavior
      
      ---
      
      ## Multi-Agent Workflow
      
      ```typescript
      // agents/multi-agent.ts
      import { tool, Settings, QueryEngineTool, VectorStoreIndex } from "llamaindex";
      import { agent, multiAgent } from "@llamaindex/workflow";
      import { openai } from "@llamaindex/openai";
      import { z } from "zod";
      
      Settings.llm = openai({ model: "gpt-4o" });
      
      // Technical support agent
      const techAgent = agent({
        name: "TechSupport",
        description: "Handles technical questions about APIs and integrations",
        tools: [techDocsTool], // QueryEngineTool over technical docs
      });
      
      // Billing agent
      const billingAgent = agent({
        name: "Billing",
        description: "Handles billing, pricing, and subscription questions",
        tools: [billingDocsTool], // QueryEngineTool over billing docs
      });
      
      // Router agent -- delegates to specialists
      const routerAgent = agent({
        name: "Router",
        description: "Routes customer questions to the right specialist",
        tools: [],
        canHandoffTo: [techAgent, billingAgent],
      });
      
      // Create multi-agent system
      const system = multiAgent({
        agents: [routerAgent, techAgent, billingAgent],
        rootAgent: routerAgent,
      });
      
      const result = await system.run("How do I update my payment method?");
      console.log(result.data);
      // Router identifies this as billing -> hands off to Billing agent
      ```
      
      **Why good:** Named agents with descriptions for routing, `canHandoffTo` for delegation, root agent as entry point
      
      ---
      
      ## Agent with Multiple Tool Types
      
      ```typescript
      import { tool, FunctionTool, QueryEngineTool } from "llamaindex";
      import { agent } from "@llamaindex/workflow";
      import { z } from "zod";
      
      // Tool from Zod schema (recommended for new tools)
      const calculateTool = tool({
        name: "calculate",
        description: "Perform arithmetic calculations",
        parameters: z.object({
          expression: z.string({ description: "Math expression like '2 + 2'" }),
        }),
        execute: async ({ expression }) => {
          // Use a safe math parser, never eval()
          return { result: "4" };
        },
      });
      
      // FunctionTool wrapping an existing function
      function lookupUser(params: { userId: string }): {
        name: string;
        email: string;
      } {
        return { name: "Alice", email: "alice@example.com" };
      }
      
      const userLookupTool = new FunctionTool(lookupUser, {
        name: "lookupUser",
        description: "Look up user details by ID",
        parameters: {
          type: "object",
          properties: {
            userId: { type: "string", description: "User ID" },
          },
          required: ["userId"],
        },
      });
      
      // QueryEngineTool wrapping an index
      const knowledgeTool = new QueryEngineTool({
        queryEngine: index.asQueryEngine(),
        metadata: {
          name: "knowledge_base",
          description: "Search internal knowledge base for company policies",
        },
      });
      
      const supportAgent = agent({
        tools: [calculateTool, userLookupTool, knowledgeTool],
      });
      ```
      
      **Why good:** Shows all three tool types, Zod-based `tool()` preferred for new tools, FunctionTool for legacy functions, QueryEngineTool for RAG
      
      ---
      
      ## Agent Event Monitoring
      
      ```typescript
      import { agentToolCallEvent, agentStreamEvent } from "@llamaindex/workflow";
      
      const events = myAgent.runStream("Research the best practices for RAG");
      
      for await (const event of events) {
        // Track tool calls
        if (agentToolCallEvent.include(event)) {
          console.log(`Tool called: ${event.data.toolName}`);
          console.log(`  Args: ${JSON.stringify(event.data.toolKwargs)}`);
        }
      
        // Stream text output
        if (agentStreamEvent.include(event)) {
          process.stdout.write(event.data.delta);
        }
      }
      ```
      
      **Why good:** Typed event discrimination, tool call visibility for debugging, concurrent streaming output
      
      ---
      
      _For core setup, see [core.md](core.md). For chat and streaming details, see [chat-streaming.md](chat-streaming.md)._
      
    • chat-streaming.md 5.5 KB
      # LlamaIndex.TS -- Chat Engines & Streaming Examples
      
      > Conversational interfaces over indexed data, streaming responses, and system prompt configuration. See [core.md](core.md) for basic setup.
      
      **Prerequisites:** Understand Settings configuration and VectorStoreIndex from [core.md](core.md).
      
      **Related examples:**
      
      - [core.md](core.md) -- Setup, indexing, query engines
      - [agents.md](agents.md) -- Agent creation, tools
      - [ingestion.md](ingestion.md) -- Text splitters, custom readers
      
      ---
      
      ## ContextChatEngine -- Basic Usage
      
      ```typescript
      // chat/context-chat.ts
      import {
        VectorStoreIndex,
        ContextChatEngine,
        SimpleDirectoryReader,
        Settings,
      } from "llamaindex";
      import { openai, OpenAIEmbedding } from "@llamaindex/openai";
      
      Settings.llm = openai({ model: "gpt-4o" });
      Settings.embedModel = new OpenAIEmbedding({ model: "text-embedding-3-small" });
      
      const SIMILARITY_TOP_K = 3;
      
      const documents = await new SimpleDirectoryReader().loadData({
        directoryPath: "./data",
      });
      const index = await VectorStoreIndex.fromDocuments(documents);
      const retriever = index.asRetriever({ similarityTopK: SIMILARITY_TOP_K });
      
      const chatEngine = new ContextChatEngine({
        retriever,
        systemPrompt:
          "You are a helpful assistant. Answer questions using the provided context. " +
          "If the context does not contain the answer, say so explicitly.",
      });
      
      // First turn
      const response1 = await chatEngine.chat({
        message: "What is the return policy?",
      });
      console.log("Bot:", response1.message.content);
      
      // Follow-up (chat engine remembers history)
      const response2 = await chatEngine.chat({
        message: "Does that apply to digital products too?",
      });
      console.log("Bot:", response2.message.content);
      ```
      
      **Why good:** Custom system prompt sets behavior, retriever injects relevant context per turn, follow-up questions use conversation history
      
      ---
      
      ## Streaming Chat Engine
      
      ```typescript
      const chatEngine = new ContextChatEngine({
        retriever,
        systemPrompt: "You are a concise assistant.",
      });
      
      // Stream the response
      const stream = await chatEngine.chat({
        message: "Explain the benefits of TypeScript",
        stream: true,
      });
      
      for await (const chunk of stream) {
        process.stdout.write(chunk.message.content);
      }
      console.log(); // Final newline
      ```
      
      **Why good:** Progressive output for user-facing chat, same engine supports both streaming and non-streaming
      
      ---
      
      ## Query Engine Streaming
      
      ```typescript
      const queryEngine = index.asQueryEngine();
      
      // Stream a query response
      const stream = await queryEngine.query({
        query: "Summarize the key findings",
        stream: true,
      });
      
      for await (const chunk of stream) {
        process.stdout.write(chunk.message.content);
      }
      ```
      
      **Why good:** Same streaming pattern as chat engine, `stream: true` option on query
      
      ---
      
      ## Chat with Custom Response Synthesizer
      
      ```typescript
      import {
        ContextChatEngine,
        getResponseSynthesizer,
        responseModeSchema,
      } from "llamaindex";
      
      // Use tree_summarize for summarization-heavy conversations
      const synthesizer = getResponseSynthesizer(
        responseModeSchema.Enum.tree_summarize,
      );
      
      const chatEngine = new ContextChatEngine({
        retriever,
        responseSynthesizer: synthesizer,
        systemPrompt: "Summarize information concisely.",
      });
      ```
      
      **Why good:** Explicit response mode selection, tree_summarize is optimal for summarization tasks
      
      ---
      
      ## Chat History Management
      
      ```typescript
      const chatEngine = new ContextChatEngine({ retriever });
      
      // Chat maintains internal history
      await chatEngine.chat({ message: "What products do you offer?" });
      await chatEngine.chat({ message: "Tell me more about the premium plan." });
      
      // Reset history when starting a new conversation
      chatEngine.reset();
      await chatEngine.chat({ message: "I have a billing question." });
      ```
      
      **Why good:** Explicit reset for new conversations, prevents context bleed between sessions
      
      ---
      
      ## Response Synthesizer Streaming (Standalone)
      
      ```typescript
      import { getResponseSynthesizer, responseModeSchema } from "llamaindex";
      
      const synthesizer = getResponseSynthesizer(responseModeSchema.Enum.compact);
      
      // Synthesize from nodes directly
      const response = await synthesizer.synthesize({
        query: "What are the main themes?",
        nodesWithScore: retrievedNodes,
        stream: true,
      });
      
      for await (const chunk of response) {
        process.stdout.write(chunk.message.content);
      }
      ```
      
      **Why good:** Direct synthesizer usage for custom pipelines, bypasses query engine when you have your own retrieval
      
      ---
      
      ## Server-Sent Events (SSE) Streaming Pattern
      
      ```typescript
      // Integrate with your HTTP framework's request/response API
      import { ContextChatEngine } from "llamaindex";
      
      // Pre-initialize chat engine at startup (not per-request)
      let chatEngine: ContextChatEngine;
      
      async function handleChatStream(message: string): Promise<ReadableStream> {
        return new ReadableStream({
          async start(controller) {
            const encoder = new TextEncoder();
            const stream = await chatEngine.chat({ message, stream: true });
      
            for await (const chunk of stream) {
              const content = chunk.message.content;
              if (content) {
                controller.enqueue(
                  encoder.encode(`data: ${JSON.stringify({ content })}\n\n`),
                );
              }
            }
            controller.enqueue(encoder.encode("data: [DONE]\n\n"));
            controller.close();
          },
        });
      }
      
      // Return ReadableStream with Content-Type: text/event-stream from your route handler
      ```
      
      **Why good:** Framework-agnostic ReadableStream, SSE-compatible format, pre-initialized engine avoids per-request overhead
      
      ---
      
      _For core setup and indexing, see [core.md](core.md). For agent patterns, see [agents.md](agents.md)._
      
    • core.md 6.3 KB
      # LlamaIndex.TS -- Setup, Indexing & Querying Examples
      
      > Core patterns for Settings configuration, document loading, VectorStoreIndex, query engines, and persistence. See [SKILL.md](../SKILL.md) for concepts and decisions.
      
      **Related examples:**
      
      - [agents.md](agents.md) -- Agent creation, tools, multi-agent workflows
      - [chat-streaming.md](chat-streaming.md) -- Chat engines, streaming responses
      - [ingestion.md](ingestion.md) -- Text splitters, node parsers, custom readers
      
      ---
      
      ## Settings Configuration -- OpenAI
      
      ```typescript
      // lib/settings.ts
      import "dotenv/config";
      import { Settings } from "llamaindex";
      import { openai, OpenAIEmbedding } from "@llamaindex/openai";
      
      Settings.llm = openai({ model: "gpt-4o" });
      Settings.embedModel = new OpenAIEmbedding({ model: "text-embedding-3-small" });
      
      export { Settings };
      ```
      
      **Why good:** Centralized config in a single module, import triggers setup, explicit provider
      
      ---
      
      ## Settings Configuration -- Ollama (Local)
      
      ```typescript
      // lib/settings.ts
      import { Settings, SentenceSplitter } from "llamaindex";
      import { ollama } from "@llamaindex/ollama";
      import { HuggingFaceEmbedding } from "@llamaindex/huggingface";
      
      const CHUNK_SIZE = 512;
      const CHUNK_OVERLAP = 50;
      
      // Local LLM -- no API key needed
      Settings.llm = ollama({ model: "mixtral:8x7b" });
      
      // Local embeddings -- no API key needed
      Settings.embedModel = new HuggingFaceEmbedding({
        modelType: "BAAI/bge-small-en-v1.5",
      });
      
      Settings.nodeParser = new SentenceSplitter({
        chunkSize: CHUNK_SIZE,
        chunkOverlap: CHUNK_OVERLAP,
      });
      ```
      
      **Why good:** Fully local setup, no API keys, named constants for chunk params, explicit embedding model
      
      ---
      
      ## Basic RAG Pipeline
      
      ```typescript
      // rag.ts
      import "dotenv/config";
      import { Settings, SimpleDirectoryReader, VectorStoreIndex } from "llamaindex";
      import { openai, OpenAIEmbedding } from "@llamaindex/openai";
      
      Settings.llm = openai({ model: "gpt-4o" });
      Settings.embedModel = new OpenAIEmbedding({ model: "text-embedding-3-small" });
      
      async function main() {
        // Step 1: Load documents
        const documents = await new SimpleDirectoryReader().loadData({
          directoryPath: "./data",
        });
      
        // Step 2: Create vector index (embeds all chunks)
        const index = await VectorStoreIndex.fromDocuments(documents);
      
        // Step 3: Create query engine
        const queryEngine = index.asQueryEngine();
      
        // Step 4: Query
        const response = await queryEngine.query({
          query: "What are the main topics in these documents?",
        });
      
        console.log("Answer:", response.message.content);
      
        // Step 5: Inspect source nodes for debugging/citations
        for (const node of response.sourceNodes ?? []) {
          console.log(
            `Score: ${node.score}, Text: ${node.node.getText().slice(0, 100)}...`,
          );
        }
      }
      
      main().catch(console.error);
      ```
      
      **Why good:** Complete pipeline, error handling via catch, source node inspection for debugging
      
      ---
      
      ## Persisted Index -- Load or Create Pattern
      
      ```typescript
      // index-manager.ts
      import {
        VectorStoreIndex,
        storageContextFromDefaults,
        SimpleDirectoryReader,
      } from "llamaindex";
      import { existsSync } from "node:fs";
      
      const PERSIST_DIR = "./storage";
      const DATA_DIR = "./data";
      
      async function getOrCreateIndex(): Promise<VectorStoreIndex> {
        const storageContext = await storageContextFromDefaults({
          persistDir: PERSIST_DIR,
        });
      
        if (existsSync(PERSIST_DIR)) {
          // Load existing index
          console.log("Loading index from storage...");
          return await VectorStoreIndex.init({ storageContext });
        }
      
        // Create new index
        console.log("Creating new index...");
        const documents = await new SimpleDirectoryReader().loadData({
          directoryPath: DATA_DIR,
        });
      
        return await VectorStoreIndex.fromDocuments(documents, { storageContext });
      }
      
      const index = await getOrCreateIndex();
      const queryEngine = index.asQueryEngine();
      ```
      
      **Why good:** Checks for existing index first, avoids re-embedding, named constants for paths
      
      ---
      
      ## Custom Retriever Configuration
      
      ```typescript
      const SIMILARITY_TOP_K = 5;
      
      // Configure retriever with custom top-k
      const retriever = index.asRetriever({
        similarityTopK: SIMILARITY_TOP_K,
      });
      
      // Use retriever directly for retrieval without synthesis
      const nodes = await retriever.retrieve({ query: "TypeScript generics" });
      for (const node of nodes) {
        console.log(
          `[${node.score?.toFixed(3)}] ${node.node.getText().slice(0, 200)}`,
        );
      }
      
      // Or use with query engine
      const queryEngine = index.asQueryEngine({
        similarityTopK: SIMILARITY_TOP_K,
      });
      ```
      
      **Why good:** Named constant for top-k, direct retriever access for debugging, score inspection
      
      ---
      
      ## Multiple Indexes for Different Data Sources
      
      ```typescript
      import {
        VectorStoreIndex,
        SimpleDirectoryReader,
        QueryEngineTool,
      } from "llamaindex";
      
      // Separate indexes for different document types
      const techDocs = await new SimpleDirectoryReader().loadData({
        directoryPath: "./data/technical",
      });
      const techIndex = await VectorStoreIndex.fromDocuments(techDocs);
      
      const hrDocs = await new SimpleDirectoryReader().loadData({
        directoryPath: "./data/hr-policies",
      });
      const hrIndex = await VectorStoreIndex.fromDocuments(hrDocs);
      
      // Wrap as tools for agent use
      const techTool = new QueryEngineTool({
        queryEngine: techIndex.asQueryEngine(),
        metadata: {
          name: "technical_docs",
          description: "Search technical documentation for API references and guides",
        },
      });
      
      const hrTool = new QueryEngineTool({
        queryEngine: hrIndex.asQueryEngine(),
        metadata: {
          name: "hr_policies",
          description: "Search HR policies for leave, benefits, and company rules",
        },
      });
      ```
      
      **Why good:** Separate indexes for domain separation, descriptive tool metadata guides the LLM
      
      ---
      
      ## Response with Source Citations
      
      ```typescript
      const response = await queryEngine.query({
        query: "What is the refund policy?",
      });
      
      console.log("Answer:", response.message.content);
      console.log("\nSources:");
      
      for (const node of response.sourceNodes ?? []) {
        const metadata = node.node.metadata;
        const score = node.score?.toFixed(3) ?? "N/A";
        const preview = node.node.getText().slice(0, 150);
      
        console.log(
          `  - [${score}] ${metadata.file_path ?? "unknown"}: ${preview}...`,
        );
      }
      ```
      
      **Why good:** Extracts file path from metadata, shows relevance score, truncates preview text
      
      ---
      
      _For agent patterns, see [agents.md](agents.md). For chat and streaming, see [chat-streaming.md](chat-streaming.md). For API reference, see [reference.md](../reference.md)._
      
    • ingestion.md 6.6 KB
      # LlamaIndex.TS -- Ingestion & Text Splitting Examples
      
      > Text splitters, node parsers, custom readers, and ingestion patterns. See [core.md](core.md) for basic setup.
      
      **Prerequisites:** Understand Settings configuration and VectorStoreIndex from [core.md](core.md).
      
      **Related examples:**
      
      - [core.md](core.md) -- Setup, indexing, query engines
      - [agents.md](agents.md) -- Agent creation, tools
      - [chat-streaming.md](chat-streaming.md) -- Chat engines, streaming
      
      ---
      
      ## SentenceSplitter Configuration
      
      ```typescript
      import { SentenceSplitter, Settings } from "llamaindex";
      
      const CHUNK_SIZE = 512;
      const CHUNK_OVERLAP = 50;
      
      // Set globally -- affects all indexing operations
      Settings.nodeParser = new SentenceSplitter({
        chunkSize: CHUNK_SIZE,
        chunkOverlap: CHUNK_OVERLAP,
      });
      ```
      
      **Why good:** Named constants, sentence-aware splitting preserves meaning, overlap ensures context continuity at boundaries
      
      ```typescript
      // BAD: Magic numbers, no overlap
      Settings.nodeParser = new SentenceSplitter({ chunkSize: 100 });
      ```
      
      **Why bad:** Tiny chunks lose context, no overlap means important sentences get split across chunks, magic number
      
      ---
      
      ## Standalone Text Splitting
      
      ```typescript
      import { SentenceSplitter } from "llamaindex";
      
      const CHUNK_SIZE = 256;
      
      const splitter = new SentenceSplitter({ chunkSize: CHUNK_SIZE });
      
      // Split a single text
      const chunks = splitter.splitText(
        "LlamaIndex is a data framework for building LLM applications. " +
          "It provides tools for document loading, text splitting, indexing, " +
          "and querying. The framework supports multiple LLM providers.",
      );
      
      console.log(`Split into ${chunks.length} chunks`);
      chunks.forEach((chunk, i) => console.log(`Chunk ${i}: ${chunk}`));
      ```
      
      **Why good:** Useful for testing chunk sizes before indexing, named constant, standalone usage without Settings
      
      ---
      
      ## MarkdownNodeParser for Structured Documents
      
      ```typescript
      import { MarkdownNodeParser } from "llamaindex";
      import { MarkdownReader } from "@llamaindex/readers/markdown";
      
      const reader = new MarkdownReader();
      const documents = await reader.loadData("./docs/api-reference.md");
      
      const parser = new MarkdownNodeParser();
      const nodes = parser(documents);
      
      // Nodes preserve markdown structure in metadata
      for (const node of nodes) {
        console.log("Headers:", node.metadata);
        // e.g. { 'Header 1': 'API Reference', 'Header 2': 'Authentication' }
        console.log("Text:", node.getText().slice(0, 100));
      }
      ```
      
      **Why good:** Structure-aware parsing, header hierarchy preserved in metadata, enables hierarchical retrieval
      
      ---
      
      ## CodeSplitter for Source Code
      
      ```typescript
      import { CodeSplitter } from "@llamaindex/node-parser/code";
      import Parser from "tree-sitter";
      import TS from "tree-sitter-typescript";
      
      const MAX_CHARS = 1500;
      
      const parser = new Parser();
      parser.setLanguage(TS.typescript);
      
      const codeSplitter = new CodeSplitter({
        getParser: () => parser,
        maxChars: MAX_CHARS,
      });
      
      // Splits code at function/class boundaries, not arbitrary character positions
      const chunks = codeSplitter.splitText(`
      export function calculateTotal(items: Item[]): number {
        return items.reduce((sum, item) => sum + item.price * item.quantity, 0);
      }
      
      export function formatCurrency(amount: number): string {
        return new Intl.NumberFormat("en-US", {
          style: "currency",
          currency: "USD",
        }).format(amount);
      }
      `);
      ```
      
      **Why good:** AST-aware splitting preserves function boundaries, named constant for max chars, language-specific parsing
      
      ---
      
      ## Custom Reader for API Data
      
      ```typescript
      import { Document } from "llamaindex";
      import type { BaseReader } from "llamaindex";
      
      class ApiDocumentReader implements BaseReader {
        private readonly baseUrl: string;
      
        constructor(baseUrl: string) {
          this.baseUrl = baseUrl;
        }
      
        async loadData(): Promise<Document[]> {
          const response = await fetch(`${this.baseUrl}/documents`);
          const items = (await response.json()) as Array<{
            id: string;
            title: string;
            content: string;
          }>;
      
          return items.map(
            (item) =>
              new Document({
                text: item.content,
                metadata: {
                  source: this.baseUrl,
                  title: item.title,
                  id: item.id,
                },
              }),
          );
        }
      }
      
      // Usage
      const reader = new ApiDocumentReader("https://api.example.com");
      const documents = await reader.loadData();
      ```
      
      **Why good:** Implements BaseReader interface, metadata for source tracking, typed API response
      
      ---
      
      ## SimpleDirectoryReader with Custom File Handlers
      
      ```typescript
      import { SimpleDirectoryReader } from "llamaindex";
      import { PDFReader } from "@llamaindex/readers/pdf";
      
      const NUM_WORKERS = 4;
      
      const reader = new SimpleDirectoryReader();
      const documents = await reader.loadData({
        directoryPath: "./data/mixed-files",
        numWorkers: NUM_WORKERS,
        fileExtToReader: {
          // Override default PDF reader with a specific one
          ".pdf": new PDFReader(),
        },
        defaultReader: undefined, // Skip unsupported file types instead of using TextFileReader
      });
      
      console.log(`Loaded ${documents.length} documents`);
      ```
      
      **Why good:** Parallel loading with workers, custom reader mapping, explicit skip for unsupported types
      
      ---
      
      ## Chunk Size Selection Guide
      
      ```typescript
      import { SentenceSplitter, Settings } from "llamaindex";
      
      // Short Q&A pairs, FAQs
      const FAQ_CHUNK_SIZE = 256;
      const FAQ_CHUNK_OVERLAP = 25;
      
      // Technical documentation
      const DOCS_CHUNK_SIZE = 512;
      const DOCS_CHUNK_OVERLAP = 50;
      
      // Long narratives, reports
      const REPORT_CHUNK_SIZE = 1024;
      const REPORT_CHUNK_OVERLAP = 100;
      
      // Choose based on your document type
      Settings.nodeParser = new SentenceSplitter({
        chunkSize: DOCS_CHUNK_SIZE,
        chunkOverlap: DOCS_CHUNK_OVERLAP,
      });
      ```
      
      **Why good:** Named constants with descriptive names, chunk overlap proportional to chunk size (~10%)
      
      ---
      
      ## Document Metadata for Filtering
      
      ```typescript
      import { Document, VectorStoreIndex, MetadataFilter } from "llamaindex";
      
      const documents = [
        new Document({
          text: "Product A costs $99 per month...",
          metadata: { category: "pricing", product: "A" },
        }),
        new Document({
          text: "Product B includes enterprise features...",
          metadata: { category: "features", product: "B" },
        }),
      ];
      
      const index = await VectorStoreIndex.fromDocuments(documents);
      
      // Query with metadata filtering (when using compatible vector stores)
      const queryEngine = index.asQueryEngine({
        preFilters: {
          filters: [{ key: "category", value: "pricing", operator: "==" }],
        },
      });
      ```
      
      **Why good:** Metadata enables filtered retrieval, reduces noise in results, category-based routing
      
      ---
      
      _For core patterns, see [core.md](core.md). For agent patterns, see [agents.md](agents.md). For API reference, see [reference.md](../reference.md)._
      
  • reference.md 10.9 KB
    # LlamaIndex.TS Quick Reference
    
    > Package map, method signatures, response modes, and model providers. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples.
    
    ---
    
    ## Package Installation
    
    ```bash
    # Core framework (always required)
    npm install llamaindex
    
    # LLM providers (install one or more)
    npm install @llamaindex/openai       # OpenAI (GPT-4o, GPT-5, embeddings)
    npm install @llamaindex/anthropic    # Anthropic (Claude)
    npm install @llamaindex/ollama       # Local models via Ollama
    npm install @llamaindex/groq         # Groq
    npm install @llamaindex/gemini       # Google Gemini
    
    # Agent workflows
    npm install @llamaindex/workflow     # agent(), multiAgent(), agentStreamEvent
    
    # Readers (optional -- SimpleDirectoryReader is in core)
    npm install @llamaindex/readers      # Additional file readers
    
    # Embedding providers
    npm install @llamaindex/huggingface  # Local embeddings (BAAI/bge-small-en-v1.5)
    
    # Tools
    npm install @llamaindex/tools        # Built-in tools (wiki, MCP)
    
    # Validation
    npm install zod                      # For tool parameter schemas
    ```
    
    ---
    
    ## Import Map
    
    | Import                       | Package                | Purpose                                          |
    | ---------------------------- | ---------------------- | ------------------------------------------------ |
    | `Settings`                   | `llamaindex`           | Global singleton for LLM, embedding, node parser |
    | `VectorStoreIndex`           | `llamaindex`           | Vector-based document index                      |
    | `SummaryIndex`               | `llamaindex`           | Full-document summary index                      |
    | `SimpleDirectoryReader`      | `llamaindex`           | Load files from a directory                      |
    | `SentenceSplitter`           | `llamaindex`           | Sentence-aware text chunking                     |
    | `ContextChatEngine`          | `llamaindex`           | Chat engine with retrieval context               |
    | `storageContextFromDefaults` | `llamaindex`           | Storage context for persistence                  |
    | `tool`                       | `llamaindex`           | Define tools with Zod schemas                    |
    | `FunctionTool`               | `llamaindex`           | Wrap functions as agent tools                    |
    | `QueryEngineTool`            | `llamaindex`           | Wrap query engines as agent tools                |
    | `getResponseSynthesizer`     | `llamaindex`           | Factory for response synthesis modes             |
    | `MarkdownNodeParser`         | `llamaindex`           | Structure-aware markdown splitting               |
    | `openai` / `OpenAIEmbedding` | `@llamaindex/openai`   | OpenAI LLM and embedding provider                |
    | `ollama`                     | `@llamaindex/ollama`   | Ollama local model provider                      |
    | `agent` / `multiAgent`       | `@llamaindex/workflow` | Agent and multi-agent creation                   |
    | `agentStreamEvent`           | `@llamaindex/workflow` | Typed event filter for streaming                 |
    
    ---
    
    ## Settings Singleton
    
    ```typescript
    import { Settings, SentenceSplitter } from "llamaindex";
    import { openai, OpenAIEmbedding } from "@llamaindex/openai";
    
    Settings.llm = openai({ model: "gpt-4o" });
    Settings.embedModel = new OpenAIEmbedding({ model: "text-embedding-3-small" });
    Settings.nodeParser = new SentenceSplitter({ chunkSize: 512 });
    Settings.chunkSize = 512; // Shortcut for node parser chunk size
    Settings.chunkOverlap = 50; // Shortcut for node parser overlap
    ```
    
    ### Settings Properties
    
    | Property          | Type              | Default               | Purpose                           |
    | ----------------- | ----------------- | --------------------- | --------------------------------- |
    | `llm`             | `LLM`             | OpenAI (lazy)         | Language model for generation     |
    | `embedModel`      | `BaseEmbedding`   | OpenAI ada-002 (lazy) | Embedding model for vector index  |
    | `nodeParser`      | `NodeParser`      | SentenceSplitter      | Document chunking strategy        |
    | `chunkSize`       | `number`          | 1024                  | Token count per chunk             |
    | `chunkOverlap`    | `number`          | 20                    | Token overlap between chunks      |
    | `callbackManager` | `CallbackManager` | empty                 | Event callbacks for observability |
    
    ---
    
    ## VectorStoreIndex Methods
    
    ```typescript
    // Create from documents (embeds all chunks)
    const index = await VectorStoreIndex.fromDocuments(documents, {
      storageContext?,    // StorageContext for persistence
      serviceContext?,    // Deprecated -- use Settings instead
    });
    
    // Load from persisted storage
    const index = await VectorStoreIndex.init({
      storageContext,     // Required: StorageContext with persistDir
    });
    
    // Get query engine
    const queryEngine = index.asQueryEngine({
      similarityTopK?: number,     // Number of results to retrieve (default: 2)
      responseSynthesizer?: ResponseSynthesizer,
    });
    
    // Get retriever
    const retriever = index.asRetriever({
      similarityTopK?: number,     // Number of results to retrieve
    });
    
    // Get chat engine
    const chatEngine = index.asChatEngine();
    ```
    
    ---
    
    ## Query Engine Interface
    
    ```typescript
    // Basic query
    const response = await queryEngine.query({ query: "Your question" });
    response.message.content; // string -- the answer
    response.sourceNodes; // NodeWithScore[] -- retrieved chunks with relevance scores
    
    // Streaming query
    const stream = await queryEngine.query({
      query: "Your question",
      stream: true,
    });
    for await (const chunk of stream) {
      process.stdout.write(chunk.message.content);
    }
    ```
    
    ---
    
    ## Chat Engine Interface
    
    ```typescript
    // Basic chat (maintains conversation history)
    const response = await chatEngine.chat({ message: "Hello" });
    response.message.content; // string -- the reply
    
    // Streaming chat
    const stream = await chatEngine.chat({ message: "Hello", stream: true });
    for await (const chunk of stream) {
      process.stdout.write(chunk.message.content);
    }
    
    // Reset conversation
    chatEngine.reset();
    ```
    
    ---
    
    ## Response Synthesizer Modes
    
    | Mode                | Behavior                                     | Best For            |
    | ------------------- | -------------------------------------------- | ------------------- |
    | `compact` (default) | Stuffs chunks into prompt, refines if needed | General purpose     |
    | `refine`            | Sequential refinement through each chunk     | Detailed answers    |
    | `tree_summarize`    | Recursive tree summarization                 | Summarization tasks |
    | `multi_modal`       | Handles text + images/audio                  | Multi-modal queries |
    
    ```typescript
    import { getResponseSynthesizer, responseModeSchema } from "llamaindex";
    
    const synthesizer = getResponseSynthesizer(responseModeSchema.Enum.compact);
    ```
    
    ---
    
    ## Agent API
    
    ```typescript
    import { agent, multiAgent, agentStreamEvent } from "@llamaindex/workflow";
    import { tool } from "llamaindex";
    import { z } from "zod";
    
    // Define a tool
    const myTool = tool({
      name: "toolName",
      description: "What this tool does (max 125 chars)",
      parameters: z.object({ /* Zod schema */ }),
      execute: async (input) => { /* implementation */ },
    });
    
    // Single agent
    const myAgent = agent({
      tools: [myTool],
      llm?: LLM,                    // Override Settings.llm
      systemPrompt?: string,
    });
    
    // Run
    const result = await myAgent.run("prompt");
    result.data;                     // Response data
    
    // Stream
    for await (const event of myAgent.runStream("prompt")) {
      if (agentStreamEvent.include(event)) {
        process.stdout.write(event.data.delta);
      }
    }
    
    // Multi-agent
    const agents = multiAgent({
      agents: [agentA, agentB],
      rootAgent: agentA,
    });
    ```
    
    ---
    
    ## SimpleDirectoryReader Options
    
    ```typescript
    const reader = new SimpleDirectoryReader();
    const documents = await reader.loadData({
      directoryPath: "./data",             // Required: path to directory
      numWorkers?: number,                 // Concurrent workers (default: 1, max: 9)
      overrideReader?: BaseReader,         // Use one reader for all files
      fileExtToReader?: Record<string, BaseReader>,  // Map extensions to readers
      defaultReader?: BaseReader,          // Fallback reader (default: TextFileReader)
    });
    ```
    
    ### Supported File Types (Default)
    
    | Extension        | Reader         |
    | ---------------- | -------------- |
    | `.txt`           | TextFileReader |
    | `.pdf`           | PDFReader      |
    | `.csv`           | CSVReader      |
    | `.md`            | MarkdownReader |
    | `.docx`          | DocxReader     |
    | `.html` / `.htm` | HTMLReader     |
    | `.jpg` / `.png`  | ImageReader    |
    
    ---
    
    ## Node Parsers
    
    | Parser               | Import                         | Use Case                                   |
    | -------------------- | ------------------------------ | ------------------------------------------ |
    | `SentenceSplitter`   | `llamaindex`                   | General text (default)                     |
    | `MarkdownNodeParser` | `llamaindex`                   | Markdown with header hierarchy             |
    | `CodeSplitter`       | `@llamaindex/node-parser/code` | Source code (AST-aware, needs tree-sitter) |
    
    ---
    
    ## LLM Provider Configuration
    
    ### OpenAI
    
    ```typescript
    import { openai, OpenAIEmbedding } from "@llamaindex/openai";
    Settings.llm = openai({ model: "gpt-4o", apiKey: process.env.OPENAI_API_KEY });
    Settings.embedModel = new OpenAIEmbedding({ model: "text-embedding-3-small" });
    ```
    
    ### Ollama (Local)
    
    ```typescript
    import { ollama } from "@llamaindex/ollama";
    Settings.llm = ollama({ model: "mixtral:8x7b" });
    ```
    
    ### Anthropic
    
    ```typescript
    import { anthropic } from "@llamaindex/anthropic";
    Settings.llm = anthropic({ model: "claude-sonnet-4-20250514" });
    ```
    
    ---
    
    ## TypeScript Configuration
    
    ```json
    {
      "compilerOptions": {
        "target": "ES2022",
        "module": "ESNext",
        "moduleResolution": "bundler",
        "esModuleInterop": true,
        "skipLibCheck": true,
        "strict": true,
        "lib": ["ES2022", "DOM.AsyncIterable"]
      }
    }
    ```
    
    **Required settings:**
    
    - `"moduleResolution": "bundler"` or `"nodenext"` -- classic `"node"` will fail
    - `"DOM.AsyncIterable"` in `lib` -- needed for Web Stream API types
    - Node.js >= 20
    
    ---
    
    ## Environment Variables
    
    | Variable            | Provider  | Purpose                                                      |
    | ------------------- | --------- | ------------------------------------------------------------ |
    | `OPENAI_API_KEY`    | OpenAI    | API key (auto-detected by `@llamaindex/openai`)              |
    | `ANTHROPIC_API_KEY` | Anthropic | API key                                                      |
    | `OLLAMA_HOST`       | Ollama    | Custom Ollama server URL (default: `http://localhost:11434`) |
    
    ---
    
    ## Storage Context
    
    ```typescript
    import { storageContextFromDefaults } from "llamaindex";
    
    // Create with persistence
    const ctx = await storageContextFromDefaults({ persistDir: "./storage" });
    
    // Files created in persistDir:
    // - docstore.json        (document metadata)
    // - index_store.json     (index structure)
    // - vector_store.json    (embeddings)
    // - graph_store.json     (graph relationships)
    ```
    
  • SKILL.md 18.7 KB
    ---
    name: ai-orchestration-llamaindex
    description: LlamaIndex.TS data framework for RAG, indexing, retrieval, query engines, chat engines, and agentic workflows in TypeScript
    ---
    
    # LlamaIndex.TS Patterns
    
    > **Quick Guide:** LlamaIndex.TS is a data framework for building context-aware LLM applications in TypeScript. Use `Settings` singleton to configure LLM and embedding models globally. Load documents with `SimpleDirectoryReader`, chunk with `SentenceSplitter`, index with `VectorStoreIndex.fromDocuments()`, and query with `index.asQueryEngine()`. For agents, use `agent()` from `@llamaindex/workflow` with `tool()` definitions using Zod schemas. All core operations are async -- every function returns a Promise. The `llamaindex` package re-exports most things, but LLM providers require separate packages like `@llamaindex/openai` or `@llamaindex/ollama`.
    
    ---
    
    <critical_requirements>
    
    ## CRITICAL: Before Using This Skill
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST configure `Settings.llm` and `Settings.embedModel` before any indexing or querying -- the Settings singleton is lazily initialized and defaults to OpenAI, which will fail without an API key)**
    
    **(You MUST await all LlamaIndex operations -- `fromDocuments()`, `asQueryEngine()`, `query()`, `chat()`, `loadData()` are ALL async)**
    
    **(You MUST install provider packages separately -- `@llamaindex/openai`, `@llamaindex/ollama`, `@llamaindex/anthropic` are NOT included in the base `llamaindex` package)**
    
    **(You MUST use `storageContextFromDefaults({ persistDir })` to persist indexes -- without persistence, indexes are rebuilt from scratch on every restart)**
    
    **(You MUST never hardcode API keys -- use environment variables and `dotenv/config`)**
    
    </critical_requirements>
    
    ---
    
    **Auto-detection:** LlamaIndex, llamaindex, VectorStoreIndex, SimpleDirectoryReader, Settings.llm, Settings.embedModel, asQueryEngine, asChatEngine, ContextChatEngine, SentenceSplitter, storageContextFromDefaults, @llamaindex/openai, @llamaindex/ollama, @llamaindex/workflow, FunctionTool, QueryEngineTool, agentStreamEvent
    
    **When to use:**
    
    - Building RAG (Retrieval-Augmented Generation) applications with custom documents
    - Loading, chunking, and indexing documents for LLM consumption
    - Creating query engines that answer questions from indexed data
    - Building chat interfaces with conversation memory over your data
    - Implementing agentic RAG with tool-calling agents that query indexes
    - Working with multiple data sources (files, PDFs, markdown, code)
    - Persisting vector indexes to avoid re-indexing on every restart
    
    **Key patterns covered:**
    
    - Settings singleton for LLM and embedding model configuration
    - Document loading with SimpleDirectoryReader and custom readers
    - VectorStoreIndex creation, persistence, and querying
    - Query engines and chat engines
    - Agent creation with `agent()` and `tool()` using Zod schemas
    - Text splitting and chunking strategies
    - Streaming responses from query and chat engines
    - Storage context and index persistence
    
    **When NOT to use:**
    
    - Simple one-shot LLM calls without document context -- use the LLM provider SDK directly
    - Applications that only need embeddings without indexing -- use the embedding API directly
    - Client-side / browser applications -- LlamaIndex.TS is server-side focused (Node.js >= 20)
    
    ---
    
    ## Examples Index
    
    - [Core: Setup, Indexing & Querying](examples/core.md) -- Settings config, document loading, VectorStoreIndex, query engines, persistence
    - [Agents & Tools](examples/agents.md) -- FunctionTool, agent(), multi-agent workflows, QueryEngineTool
    - [Chat & Streaming](examples/chat-streaming.md) -- Chat engines, ContextChatEngine, streaming responses
    - [Ingestion & Splitting](examples/ingestion.md) -- Text splitters, node parsers, ingestion pipeline, custom readers
    - [Quick API Reference](reference.md) -- Package map, method signatures, response modes, model providers
    
    ---
    
    <philosophy>
    
    ## Philosophy
    
    LlamaIndex.TS is a **data framework** -- its core value proposition is connecting your data to LLMs through indexing, retrieval, and synthesis. It sits between raw LLM APIs and full application frameworks.
    
    **Core principles:**
    
    1. **Context engineering** -- Inject the right data into the LLM prompt at the right time. This drives RAG, agent memory, extraction, and summarization.
    2. **Modular provider system** -- LLM providers, embedding models, vector stores, and readers are separate packages you compose. The base `llamaindex` package provides the framework; providers are installed separately.
    3. **Settings singleton** -- Global configuration for LLM, embedding model, node parser, and other shared resources. Set once, used everywhere. Override locally when needed.
    4. **Async-first design** -- Every I/O operation is async. Document loading, indexing, querying, and chat all return Promises.
    5. **Index as the core abstraction** -- Documents are loaded, split into nodes, embedded, and stored in an index. Queries retrieve relevant nodes and synthesize responses.
    
    **When to use LlamaIndex.TS:**
    
    - You have documents/data that need to be indexed for LLM consumption
    - You want structured RAG pipelines with configurable retrieval and synthesis
    - You need agentic RAG where agents query multiple indexes with tools
    - You want persistence and incremental updates to your index
    
    **When NOT to use:**
    
    - Simple LLM calls without data context -- use the provider SDK directly
    - Browser-only applications -- LlamaIndex.TS requires Node.js >= 20
    - You only need embeddings -- use the embedding API directly
    
    </philosophy>
    
    ---
    
    <patterns>
    
    ## Core Patterns
    
    ### Pattern 1: Settings Configuration
    
    The `Settings` singleton configures LLM, embedding model, and node parser globally. Set it once at application startup before any indexing or querying.
    
    ```typescript
    import { Settings } from "llamaindex";
    import { openai, OpenAIEmbedding } from "@llamaindex/openai";
    
    // Configure at app startup -- before any index operations
    Settings.llm = openai({ model: "gpt-4o" });
    Settings.embedModel = new OpenAIEmbedding({ model: "text-embedding-3-small" });
    ```
    
    **Why good:** Single configuration point, provider packages are explicit imports, model names are visible
    
    ```typescript
    // BAD: No Settings configuration, relying on implicit defaults
    import { VectorStoreIndex, SimpleDirectoryReader } from "llamaindex";
    
    // This will silently try to use OpenAI with OPENAI_API_KEY from env
    // Fails with cryptic error if key is missing
    const documents = await new SimpleDirectoryReader().loadData({
      directoryPath: "./data",
    });
    const index = await VectorStoreIndex.fromDocuments(documents);
    ```
    
    **Why bad:** Implicit defaults make failures confusing, no explicit provider, no model selection
    
    **See:** [examples/core.md](examples/core.md) for local LLM setup with Ollama, Anthropic configuration, and embedding model options
    
    ---
    
    ### Pattern 2: Document Loading and Indexing
    
    Load documents, create a vector index, and query it. This is the canonical RAG pipeline.
    
    ```typescript
    import { SimpleDirectoryReader, VectorStoreIndex, Settings } from "llamaindex";
    import { openai, OpenAIEmbedding } from "@llamaindex/openai";
    
    Settings.llm = openai({ model: "gpt-4o" });
    Settings.embedModel = new OpenAIEmbedding({ model: "text-embedding-3-small" });
    
    // Load all supported files from a directory
    const documents = await new SimpleDirectoryReader().loadData({
      directoryPath: "./data",
    });
    
    // Create vector index -- embeds and stores all document chunks
    const index = await VectorStoreIndex.fromDocuments(documents);
    
    // Query the index
    const queryEngine = index.asQueryEngine();
    const response = await queryEngine.query({ query: "What is the main topic?" });
    console.log(response.message.content);
    ```
    
    **Why good:** Complete pipeline in minimal code, explicit Settings, clear data flow
    
    **See:** [examples/core.md](examples/core.md) for persistence, custom readers, and advanced indexing options
    
    ---
    
    ### Pattern 3: Index Persistence
    
    Persist indexes to disk to avoid re-indexing on every restart.
    
    ```typescript
    import {
      VectorStoreIndex,
      storageContextFromDefaults,
      SimpleDirectoryReader,
    } from "llamaindex";
    
    const PERSIST_DIR = "./storage";
    
    // First run: create and persist
    const storageContext = await storageContextFromDefaults({
      persistDir: PERSIST_DIR,
    });
    const documents = await new SimpleDirectoryReader().loadData({
      directoryPath: "./data",
    });
    const index = await VectorStoreIndex.fromDocuments(documents, {
      storageContext,
    });
    
    // Subsequent runs: load from storage
    const loadedStorageContext = await storageContextFromDefaults({
      persistDir: PERSIST_DIR,
    });
    const loadedIndex = await VectorStoreIndex.init({
      storageContext: loadedStorageContext,
    });
    ```
    
    **Why good:** Named constant for path, separate create vs load paths, storage context reuse
    
    ```typescript
    // BAD: Rebuilding index on every request
    async function handleQuery(question: string) {
      const docs = await new SimpleDirectoryReader().loadData({
        directoryPath: "./data",
      });
      const index = await VectorStoreIndex.fromDocuments(docs); // Expensive!
      const engine = index.asQueryEngine();
      return engine.query({ query: question });
    }
    ```
    
    **Why bad:** Re-indexes all documents on every call, wastes time and API credits on re-embedding
    
    **See:** [examples/core.md](examples/core.md) for load-or-create pattern
    
    ---
    
    ### Pattern 4: Agents with Tool Definitions
    
    Create agents that use tools defined with Zod schemas. Use `agent()` from `@llamaindex/workflow`.
    
    ```typescript
    import { tool, Settings } from "llamaindex";
    import { agent, agentStreamEvent } from "@llamaindex/workflow";
    import { openai } from "@llamaindex/openai";
    import { z } from "zod";
    
    Settings.llm = openai({ model: "gpt-4o" });
    
    const weatherTool = tool({
      name: "getWeather",
      description: "Get current weather for a city",
      parameters: z.object({
        city: z.string({ description: "City name" }),
      }),
      execute: async ({ city }) => {
        // Your weather API call here
        return { temperature: 22, condition: "sunny" };
      },
    });
    
    const myAgent = agent({ tools: [weatherTool] });
    const result = await myAgent.run("What's the weather in Paris?");
    console.log(result.data);
    ```
    
    **Why good:** Zod schema for type-safe parameters, description guides the LLM, async execute function
    
    **See:** [examples/agents.md](examples/agents.md) for multi-agent workflows, QueryEngineTool, streaming agents
    
    ---
    
    ### Pattern 5: Chat Engine
    
    Build conversational interfaces over your indexed data with conversation memory.
    
    ```typescript
    import {
      VectorStoreIndex,
      ContextChatEngine,
      SimpleDirectoryReader,
    } from "llamaindex";
    
    const documents = await new SimpleDirectoryReader().loadData({
      directoryPath: "./data",
    });
    const index = await VectorStoreIndex.fromDocuments(documents);
    const retriever = index.asRetriever({ similarityTopK: 3 });
    
    const chatEngine = new ContextChatEngine({ retriever });
    
    // Multi-turn conversation -- chat engine maintains history
    const response1 = await chatEngine.chat({ message: "What is LlamaIndex?" });
    console.log(response1.message.content);
    
    const response2 = await chatEngine.chat({
      message: "How does it handle streaming?",
    });
    console.log(response2.message.content);
    ```
    
    **Why good:** Retriever-based context injection, automatic conversation history, multi-turn support
    
    **See:** [examples/chat-streaming.md](examples/chat-streaming.md) for streaming chat, system prompts, chat history management
    
    ---
    
    ### Pattern 6: Streaming Responses
    
    Stream responses for user-facing applications.
    
    ```typescript
    import { agentStreamEvent } from "@llamaindex/workflow";
    
    // Agent streaming
    const events = myAgent.runStream("Tell me about TypeScript");
    for await (const event of events) {
      if (agentStreamEvent.include(event)) {
        process.stdout.write(event.data.delta);
      }
    }
    
    // Query engine streaming
    const response = await queryEngine.query({
      query: "Summarize the document",
      stream: true,
    });
    for await (const chunk of response) {
      process.stdout.write(chunk.message.content);
    }
    ```
    
    **Why good:** Event-based agent streaming with typed filters, query engine streaming with for-await
    
    **See:** [examples/chat-streaming.md](examples/chat-streaming.md) for response synthesizer streaming, chat engine streaming
    
    ---
    
    ### Pattern 7: Text Splitting and Node Parsing
    
    Configure how documents are chunked before indexing.
    
    ```typescript
    import { SentenceSplitter, Settings } from "llamaindex";
    
    const CHUNK_SIZE = 512;
    const CHUNK_OVERLAP = 50;
    
    // Set globally via Settings
    Settings.nodeParser = new SentenceSplitter({
      chunkSize: CHUNK_SIZE,
      chunkOverlap: CHUNK_OVERLAP,
    });
    
    // Or use standalone
    const splitter = new SentenceSplitter({ chunkSize: CHUNK_SIZE });
    const texts = splitter.splitText("Your long document text here...");
    ```
    
    **Why good:** Named constants for chunk parameters, global vs standalone usage shown, sentence-aware splitting
    
    ```typescript
    // BAD: Using default chunk size without considering document characteristics
    const index = await VectorStoreIndex.fromDocuments(documents);
    // Default chunk size may be too large for short Q&A or too small for long narratives
    ```
    
    **Why bad:** Default chunk size (1024 tokens) may not suit your data, causes poor retrieval quality
    
    **See:** [examples/ingestion.md](examples/ingestion.md) for MarkdownNodeParser, CodeSplitter, custom chunk strategies
    
    </patterns>
    
    ---
    
    <decision_framework>
    
    ## Decision Framework
    
    ### Which Index Type to Use
    
    ```
    What is your use case?
    +-- Semantic search over documents -> VectorStoreIndex (most common)
    +-- Summarization of all documents -> SummaryIndex
    +-- Both search AND summarization -> Create both, use as separate tools in an agent
    +-- Hierarchical document structure -> Use MarkdownNodeParser + VectorStoreIndex
    ```
    
    ### Query Engine vs Chat Engine vs Agent
    
    ```
    How should users interact with your data?
    +-- Single question, single answer -> Query Engine (index.asQueryEngine())
    +-- Multi-turn conversation -> Chat Engine (ContextChatEngine)
    +-- Multiple tools/indexes + reasoning -> Agent (agent() from @llamaindex/workflow)
    +-- Complex multi-step workflow -> Multi-agent with handoffs
    ```
    
    ### Which LLM Provider
    
    ```
    Which LLM provider are you using?
    +-- OpenAI -> npm install @llamaindex/openai
    +-- Anthropic -> npm install @llamaindex/anthropic
    +-- Local (Ollama) -> npm install @llamaindex/ollama
    +-- Groq -> npm install @llamaindex/groq
    +-- Google Gemini -> npm install @llamaindex/gemini
    ```
    
    ### Chunk Size Selection
    
    ```
    What kind of documents are you indexing?
    +-- Short Q&A pairs -> chunkSize: 256-512
    +-- Technical documentation -> chunkSize: 512-1024
    +-- Long narratives/reports -> chunkSize: 1024-2048
    +-- Code files -> Use CodeSplitter (AST-aware)
    +-- Markdown -> Use MarkdownNodeParser (structure-aware)
    ```
    
    </decision_framework>
    
    ---
    
    <red_flags>
    
    ## RED FLAGS
    
    **High Priority Issues:**
    
    - Not configuring `Settings.llm` before indexing/querying -- defaults to OpenAI, fails silently without API key
    - Forgetting to `await` async operations -- `fromDocuments()`, `query()`, `chat()` all return Promises
    - Rebuilding indexes on every request instead of persisting with `storageContextFromDefaults`
    - Hardcoding API keys instead of using environment variables
    - Installing only `llamaindex` without provider packages (`@llamaindex/openai`, etc.)
    
    **Medium Priority Issues:**
    
    - Using default chunk size (1024) without considering document characteristics -- causes poor retrieval
    - Not setting `similarityTopK` on retrievers -- default may return too few or too many results
    - Ignoring the response `sourceNodes` -- they contain the retrieved context for debugging and citations
    - Creating a new `SimpleDirectoryReader` per request instead of caching the loaded documents
    - Not handling the case where `response.message.content` might be empty on retrieval failure
    
    **Common Mistakes:**
    
    - Confusing `asQueryEngine()` (single question) with `ContextChatEngine` (multi-turn conversation)
    - Using `VectorStoreIndex.fromDocuments()` when you should use `VectorStoreIndex.init()` to load from storage
    - Importing `openai` from `llamaindex` instead of `@llamaindex/openai` -- the `llamaindex` package may re-export some things but provider-specific imports are more reliable
    - Passing `messages` array to `query()` -- query engines take `{ query: string }`, not a messages array
    - Using `index.asQueryEngine()` multiple times instead of storing the engine reference
    
    **Gotchas & Edge Cases:**
    
    - `Settings` is a global singleton -- setting it in one module affects all others. Override locally by passing `llm` directly to constructors when you need different models for different operations.
    - `SimpleDirectoryReader` only works on Node.js -- it uses `fs` internally. For edge/serverless, load documents differently or use LlamaParse.
    - `storageContextFromDefaults` creates four JSON files in the persist directory (`docstore.json`, `graph_store.json`, `index_store.json`, `vector_store.json`). If any are corrupted, delete the directory and re-index.
    - Node.js >= 20 is required. Some modules use Web Stream API (`ReadableStream`, `WritableStream`), so add `"DOM.AsyncIterable"` to `tsconfig.json` `lib` if you get type errors.
    - `tsconfig.json` must use `"moduleResolution": "bundler"` or `"nodenext"` -- the classic `"node"` resolution will fail to resolve LlamaIndex sub-packages.
    - Default tokenizer is slow -- install `gpt-tokenizer` for 60x faster tokenization.
    - `SentenceSplitter` chunk size is in tokens, not characters. A 512-token chunk is roughly 2000 characters.
    - The `llamaindex` package is large (~2MB+). For production, consider importing specific sub-packages to reduce bundle size.
    - `VectorStoreIndex.fromDocuments()` makes embedding API calls for every chunk. For large document sets, this can be expensive. Monitor costs.
    - Chat engine conversation history grows unbounded -- implement history pruning for long-running sessions.
    
    </red_flags>
    
    ---
    
    <critical_reminders>
    
    ## CRITICAL REMINDERS
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST configure `Settings.llm` and `Settings.embedModel` before any indexing or querying -- the Settings singleton is lazily initialized and defaults to OpenAI, which will fail without an API key)**
    
    **(You MUST await all LlamaIndex operations -- `fromDocuments()`, `asQueryEngine()`, `query()`, `chat()`, `loadData()` are ALL async)**
    
    **(You MUST install provider packages separately -- `@llamaindex/openai`, `@llamaindex/ollama`, `@llamaindex/anthropic` are NOT included in the base `llamaindex` package)**
    
    **(You MUST use `storageContextFromDefaults({ persistDir })` to persist indexes -- without persistence, indexes are rebuilt from scratch on every restart)**
    
    **(You MUST never hardcode API keys -- use environment variables and `dotenv/config`)**
    
    **Failure to follow these rules will produce broken RAG pipelines, wasted embedding API credits, or cryptic runtime errors.**
    
    </critical_reminders>
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related