ai-orchestration-llamaindex
LlamaIndex.TS data framework for RAG, indexing, retrieval, query engines, chat engines, and agentic workflows in TypeScript
Install
npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/ai-orchestration-llamaindex/skills/ai-orchestration-llamaindex
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
git clone https://github.com/agents-inc/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
LlamaIndex.TS Patterns
Quick Guide: LlamaIndex.TS is a data framework for building context-aware LLM applications in TypeScript. Use
Settingssingleton to configure LLM and embedding models globally. Load documents withSimpleDirectoryReader, chunk withSentenceSplitter, index withVectorStoreIndex.fromDocuments(), and query withindex.asQueryEngine(). For agents, useagent()from@llamaindex/workflowwithtool()definitions using Zod schemas. All core operations are async -- every function returns a Promise. Thellamaindexpackage re-exports most things, but LLM providers require separate packages like@llamaindex/openaior@llamaindex/ollama.
<critical_requirements>
CRITICAL: Before Using This Skill
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST configure Settings.llm and Settings.embedModel before any indexing or querying -- the Settings singleton is lazily initialized and defaults to OpenAI, which will fail without an API key)
(You MUST await all LlamaIndex operations -- fromDocuments(), asQueryEngine(), query(), chat(), loadData() are ALL async)
(You MUST install provider packages separately -- @llamaindex/openai, @llamaindex/ollama, @llamaindex/anthropic are NOT included in the base llamaindex package)
(You MUST use storageContextFromDefaults({ persistDir }) to persist indexes -- without persistence, indexes are rebuilt from scratch on every restart)
(You MUST never hardcode API keys -- use environment variables and dotenv/config)
</critical_requirements>
Auto-detection: LlamaIndex, llamaindex, VectorStoreIndex, SimpleDirectoryReader, Settings.llm, Settings.embedModel, asQueryEngine, asChatEngine, ContextChatEngine, SentenceSplitter, storageContextFromDefaults, @llamaindex/openai, @llamaindex/ollama, @llamaindex/workflow, FunctionTool, QueryEngineTool, agentStreamEvent
When to use:
- Building RAG (Retrieval-Augmented Generation) applications with custom documents
- Loading, chunking, and indexing documents for LLM consumption
- Creating query engines that answer questions from indexed data
- Building chat interfaces with conversation memory over your data
- Implementing agentic RAG with tool-calling agents that query indexes
- Working with multiple data sources (files, PDFs, markdown, code)
- Persisting vector indexes to avoid re-indexing on every restart
Key patterns covered:
- Settings singleton for LLM and embedding model configuration
- Document loading with SimpleDirectoryReader and custom readers
- VectorStoreIndex creation, persistence, and querying
- Query engines and chat engines
- Agent creation with
agent()andtool()using Zod schemas - Text splitting and chunking strategies
- Streaming responses from query and chat engines
- Storage context and index persistence
When NOT to use:
- Simple one-shot LLM calls without document context -- use the LLM provider SDK directly
- Applications that only need embeddings without indexing -- use the embedding API directly
- Client-side / browser applications -- LlamaIndex.TS is server-side focused (Node.js >= 20)
Examples Index
- Core: Setup, Indexing & Querying -- Settings config, document loading, VectorStoreIndex, query engines, persistence
- Agents & Tools -- FunctionTool, agent(), multi-agent workflows, QueryEngineTool
- Chat & Streaming -- Chat engines, ContextChatEngine, streaming responses
- Ingestion & Splitting -- Text splitters, node parsers, ingestion pipeline, custom readers
- Quick API Reference -- Package map, method signatures, response modes, model providers
<decision_framework>
Decision Framework
Which Index Type to Use
What is your use case?
+-- Semantic search over documents -> VectorStoreIndex (most common)
+-- Summarization of all documents -> SummaryIndex
+-- Both search AND summarization -> Create both, use as separate tools in an agent
+-- Hierarchical document structure -> Use MarkdownNodeParser + VectorStoreIndex
Query Engine vs Chat Engine vs Agent
How should users interact with your data?
+-- Single question, single answer -> Query Engine (index.asQueryEngine())
+-- Multi-turn conversation -> Chat Engine (ContextChatEngine)
+-- Multiple tools/indexes + reasoning -> Agent (agent() from @llamaindex/workflow)
+-- Complex multi-step workflow -> Multi-agent with handoffs
Which LLM Provider
Which LLM provider are you using?
+-- OpenAI -> npm install @llamaindex/openai
+-- Anthropic -> npm install @llamaindex/anthropic
+-- Local (Ollama) -> npm install @llamaindex/ollama
+-- Groq -> npm install @llamaindex/groq
+-- Google Gemini -> npm install @llamaindex/gemini
Chunk Size Selection
What kind of documents are you indexing?
+-- Short Q&A pairs -> chunkSize: 256-512
+-- Technical documentation -> chunkSize: 512-1024
+-- Long narratives/reports -> chunkSize: 1024-2048
+-- Code files -> Use CodeSplitter (AST-aware)
+-- Markdown -> Use MarkdownNodeParser (structure-aware)
</decision_framework>
<red_flags>
RED FLAGS
High Priority Issues:
- Not configuring
Settings.llmbefore indexing/querying -- defaults to OpenAI, fails silently without API key - Forgetting to
awaitasync operations --fromDocuments(),query(),chat()all return Promises - Rebuilding indexes on every request instead of persisting with
storageContextFromDefaults - Hardcoding API keys instead of using environment variables
- Installing only
llamaindexwithout provider packages (@llamaindex/openai, etc.)
Medium Priority Issues:
- Using default chunk size (1024) without considering document characteristics -- causes poor retrieval
- Not setting
similarityTopKon retrievers -- default may return too few or too many results - Ignoring the response
sourceNodes-- they contain the retrieved context for debugging and citations - Creating a new
SimpleDirectoryReaderper request instead of caching the loaded documents - Not handling the case where
response.message.contentmight be empty on retrieval failure
Common Mistakes:
- Confusing
asQueryEngine()(single question) withContextChatEngine(multi-turn conversation) - Using
VectorStoreIndex.fromDocuments()when you should useVectorStoreIndex.init()to load from storage - Importing
openaifromllamaindexinstead of@llamaindex/openai-- thellamaindexpackage may re-export some things but provider-specific imports are more reliable - Passing
messagesarray toquery()-- query engines take{ query: string }, not a messages array - Using
index.asQueryEngine()multiple times instead of storing the engine reference
Gotchas & Edge Cases:
Settingsis a global singleton -- setting it in one module affects all others. Override locally by passingllmdirectly to constructors when you need different models for different operations.SimpleDirectoryReaderonly works on Node.js -- it usesfsinternally. For edge/serverless, load documents differently or use LlamaParse.storageContextFromDefaultscreates four JSON files in the persist directory (docstore.json,graph_store.json,index_store.json,vector_store.json). If any are corrupted, delete the directory and re-index.- Node.js >= 20 is required. Some modules use Web Stream API (
ReadableStream,WritableStream), so add"DOM.AsyncIterable"totsconfig.jsonlibif you get type errors. tsconfig.jsonmust use"moduleResolution": "bundler"or"nodenext"-- the classic"node"resolution will fail to resolve LlamaIndex sub-packages.- Default tokenizer is slow -- install
gpt-tokenizerfor 60x faster tokenization. SentenceSplitterchunk size is in tokens, not characters. A 512-token chunk is roughly 2000 characters.- The
llamaindexpackage is large (~2MB+). For production, consider importing specific sub-packages to reduce bundle size. VectorStoreIndex.fromDocuments()makes embedding API calls for every chunk. For large document sets, this can be expensive. Monitor costs.- Chat engine conversation history grows unbounded -- implement history pruning for long-running sessions.
</red_flags>
<critical_reminders>
CRITICAL REMINDERS
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST configure Settings.llm and Settings.embedModel before any indexing or querying -- the Settings singleton is lazily initialized and defaults to OpenAI, which will fail without an API key)
(You MUST await all LlamaIndex operations -- fromDocuments(), asQueryEngine(), query(), chat(), loadData() are ALL async)
(You MUST install provider packages separately -- @llamaindex/openai, @llamaindex/ollama, @llamaindex/anthropic are NOT included in the base llamaindex package)
(You MUST use storageContextFromDefaults({ persistDir }) to persist indexes -- without persistence, indexes are rebuilt from scratch on every restart)
(You MUST never hardcode API keys -- use environment variables and dotenv/config)
Failure to follow these rules will produce broken RAG pipelines, wasted embedding API credits, or cryptic runtime errors.
</critical_reminders>
Files (skills)
-
examples
-
agents.md 7.3 KB
# LlamaIndex.TS -- Agents & Tools Examples > Agent creation, tool definitions, multi-agent workflows, and QueryEngineTool patterns. See [core.md](core.md) for basic setup. **Prerequisites:** Understand Settings configuration and VectorStoreIndex from [core.md](core.md). **Related examples:** - [core.md](core.md) -- Setup, indexing, query engines - [chat-streaming.md](chat-streaming.md) -- Chat engines, streaming --- ## Basic Agent with Custom Tools ```typescript // agents/weather-agent.ts import "dotenv/config"; import { tool, Settings } from "llamaindex"; import { agent } from "@llamaindex/workflow"; import { openai } from "@llamaindex/openai"; import { z } from "zod"; Settings.llm = openai({ model: "gpt-4o" }); const getWeather = tool({ name: "getWeather", description: "Get current weather for a city", parameters: z.object({ city: z.string({ description: "City name (e.g. 'Paris')" }), unit: z .enum(["celsius", "fahrenheit"]) .default("celsius") .describe("Temperature unit"), }), execute: async ({ city, unit }) => { // Replace with actual API call return { city, temperature: unit === "celsius" ? 22 : 72, condition: "sunny", }; }, }); const weatherAgent = agent({ tools: [getWeather], }); const result = await weatherAgent.run("What's the weather in Paris?"); console.log(result.data); ``` **Why good:** Zod schema with descriptions guides the model, default values reduce friction, typed execute function --- ## Agent with Streaming Output ```typescript import { agentStreamEvent } from "@llamaindex/workflow"; const events = weatherAgent.runStream( "Compare weather in Paris and Tokyo, recommend which to visit", ); for await (const event of events) { if (agentStreamEvent.include(event)) { process.stdout.write(event.data.delta); } } console.log(); // Final newline ``` **Why good:** Type-safe event filtering, progressive output, simple for-await pattern --- ## Agent with Structured Output ```typescript import { z } from "zod"; const travelRecommendation = z.object({ city: z.string(), reason: z.string(), bestSeason: z.string(), averageCost: z.number(), }); const result = await weatherAgent.run("Recommend a city to visit in summer", { responseFormat: travelRecommendation, }); // result.data.object is typed as { city, reason, bestSeason, averageCost } console.log(result.data.object); ``` **Why good:** Zod schema validates and types the response, structured data extraction --- ## RAG Agent with QueryEngineTool ```typescript // agents/rag-agent.ts import { VectorStoreIndex, SimpleDirectoryReader, QueryEngineTool, Settings, } from "llamaindex"; import { agent } from "@llamaindex/workflow"; import { openai, OpenAIEmbedding } from "@llamaindex/openai"; Settings.llm = openai({ model: "gpt-4o" }); Settings.embedModel = new OpenAIEmbedding({ model: "text-embedding-3-small" }); // Create index from documents const documents = await new SimpleDirectoryReader().loadData({ directoryPath: "./data/product-docs", }); const index = await VectorStoreIndex.fromDocuments(documents); // Wrap query engine as a tool const docSearchTool = new QueryEngineTool({ queryEngine: index.asQueryEngine(), metadata: { name: "product_docs", description: "Search product documentation for features, pricing, and technical specs", }, }); const ragAgent = agent({ tools: [docSearchTool], systemPrompt: "You are a product support agent. Use the product_docs tool to answer questions accurately. Cite sources.", }); const response = await ragAgent.run("What are the pricing tiers?"); console.log(response.data); ``` **Why good:** QueryEngineTool bridges RAG and agents, descriptive metadata for accurate tool routing, system prompt sets behavior --- ## Multi-Agent Workflow ```typescript // agents/multi-agent.ts import { tool, Settings, QueryEngineTool, VectorStoreIndex } from "llamaindex"; import { agent, multiAgent } from "@llamaindex/workflow"; import { openai } from "@llamaindex/openai"; import { z } from "zod"; Settings.llm = openai({ model: "gpt-4o" }); // Technical support agent const techAgent = agent({ name: "TechSupport", description: "Handles technical questions about APIs and integrations", tools: [techDocsTool], // QueryEngineTool over technical docs }); // Billing agent const billingAgent = agent({ name: "Billing", description: "Handles billing, pricing, and subscription questions", tools: [billingDocsTool], // QueryEngineTool over billing docs }); // Router agent -- delegates to specialists const routerAgent = agent({ name: "Router", description: "Routes customer questions to the right specialist", tools: [], canHandoffTo: [techAgent, billingAgent], }); // Create multi-agent system const system = multiAgent({ agents: [routerAgent, techAgent, billingAgent], rootAgent: routerAgent, }); const result = await system.run("How do I update my payment method?"); console.log(result.data); // Router identifies this as billing -> hands off to Billing agent ``` **Why good:** Named agents with descriptions for routing, `canHandoffTo` for delegation, root agent as entry point --- ## Agent with Multiple Tool Types ```typescript import { tool, FunctionTool, QueryEngineTool } from "llamaindex"; import { agent } from "@llamaindex/workflow"; import { z } from "zod"; // Tool from Zod schema (recommended for new tools) const calculateTool = tool({ name: "calculate", description: "Perform arithmetic calculations", parameters: z.object({ expression: z.string({ description: "Math expression like '2 + 2'" }), }), execute: async ({ expression }) => { // Use a safe math parser, never eval() return { result: "4" }; }, }); // FunctionTool wrapping an existing function function lookupUser(params: { userId: string }): { name: string; email: string; } { return { name: "Alice", email: "alice@example.com" }; } const userLookupTool = new FunctionTool(lookupUser, { name: "lookupUser", description: "Look up user details by ID", parameters: { type: "object", properties: { userId: { type: "string", description: "User ID" }, }, required: ["userId"], }, }); // QueryEngineTool wrapping an index const knowledgeTool = new QueryEngineTool({ queryEngine: index.asQueryEngine(), metadata: { name: "knowledge_base", description: "Search internal knowledge base for company policies", }, }); const supportAgent = agent({ tools: [calculateTool, userLookupTool, knowledgeTool], }); ``` **Why good:** Shows all three tool types, Zod-based `tool()` preferred for new tools, FunctionTool for legacy functions, QueryEngineTool for RAG --- ## Agent Event Monitoring ```typescript import { agentToolCallEvent, agentStreamEvent } from "@llamaindex/workflow"; const events = myAgent.runStream("Research the best practices for RAG"); for await (const event of events) { // Track tool calls if (agentToolCallEvent.include(event)) { console.log(`Tool called: ${event.data.toolName}`); console.log(` Args: ${JSON.stringify(event.data.toolKwargs)}`); } // Stream text output if (agentStreamEvent.include(event)) { process.stdout.write(event.data.delta); } } ``` **Why good:** Typed event discrimination, tool call visibility for debugging, concurrent streaming output --- _For core setup, see [core.md](core.md). For chat and streaming details, see [chat-streaming.md](chat-streaming.md)._ -
chat-streaming.md 5.5 KB
# LlamaIndex.TS -- Chat Engines & Streaming Examples > Conversational interfaces over indexed data, streaming responses, and system prompt configuration. See [core.md](core.md) for basic setup. **Prerequisites:** Understand Settings configuration and VectorStoreIndex from [core.md](core.md). **Related examples:** - [core.md](core.md) -- Setup, indexing, query engines - [agents.md](agents.md) -- Agent creation, tools - [ingestion.md](ingestion.md) -- Text splitters, custom readers --- ## ContextChatEngine -- Basic Usage ```typescript // chat/context-chat.ts import { VectorStoreIndex, ContextChatEngine, SimpleDirectoryReader, Settings, } from "llamaindex"; import { openai, OpenAIEmbedding } from "@llamaindex/openai"; Settings.llm = openai({ model: "gpt-4o" }); Settings.embedModel = new OpenAIEmbedding({ model: "text-embedding-3-small" }); const SIMILARITY_TOP_K = 3; const documents = await new SimpleDirectoryReader().loadData({ directoryPath: "./data", }); const index = await VectorStoreIndex.fromDocuments(documents); const retriever = index.asRetriever({ similarityTopK: SIMILARITY_TOP_K }); const chatEngine = new ContextChatEngine({ retriever, systemPrompt: "You are a helpful assistant. Answer questions using the provided context. " + "If the context does not contain the answer, say so explicitly.", }); // First turn const response1 = await chatEngine.chat({ message: "What is the return policy?", }); console.log("Bot:", response1.message.content); // Follow-up (chat engine remembers history) const response2 = await chatEngine.chat({ message: "Does that apply to digital products too?", }); console.log("Bot:", response2.message.content); ``` **Why good:** Custom system prompt sets behavior, retriever injects relevant context per turn, follow-up questions use conversation history --- ## Streaming Chat Engine ```typescript const chatEngine = new ContextChatEngine({ retriever, systemPrompt: "You are a concise assistant.", }); // Stream the response const stream = await chatEngine.chat({ message: "Explain the benefits of TypeScript", stream: true, }); for await (const chunk of stream) { process.stdout.write(chunk.message.content); } console.log(); // Final newline ``` **Why good:** Progressive output for user-facing chat, same engine supports both streaming and non-streaming --- ## Query Engine Streaming ```typescript const queryEngine = index.asQueryEngine(); // Stream a query response const stream = await queryEngine.query({ query: "Summarize the key findings", stream: true, }); for await (const chunk of stream) { process.stdout.write(chunk.message.content); } ``` **Why good:** Same streaming pattern as chat engine, `stream: true` option on query --- ## Chat with Custom Response Synthesizer ```typescript import { ContextChatEngine, getResponseSynthesizer, responseModeSchema, } from "llamaindex"; // Use tree_summarize for summarization-heavy conversations const synthesizer = getResponseSynthesizer( responseModeSchema.Enum.tree_summarize, ); const chatEngine = new ContextChatEngine({ retriever, responseSynthesizer: synthesizer, systemPrompt: "Summarize information concisely.", }); ``` **Why good:** Explicit response mode selection, tree_summarize is optimal for summarization tasks --- ## Chat History Management ```typescript const chatEngine = new ContextChatEngine({ retriever }); // Chat maintains internal history await chatEngine.chat({ message: "What products do you offer?" }); await chatEngine.chat({ message: "Tell me more about the premium plan." }); // Reset history when starting a new conversation chatEngine.reset(); await chatEngine.chat({ message: "I have a billing question." }); ``` **Why good:** Explicit reset for new conversations, prevents context bleed between sessions --- ## Response Synthesizer Streaming (Standalone) ```typescript import { getResponseSynthesizer, responseModeSchema } from "llamaindex"; const synthesizer = getResponseSynthesizer(responseModeSchema.Enum.compact); // Synthesize from nodes directly const response = await synthesizer.synthesize({ query: "What are the main themes?", nodesWithScore: retrievedNodes, stream: true, }); for await (const chunk of response) { process.stdout.write(chunk.message.content); } ``` **Why good:** Direct synthesizer usage for custom pipelines, bypasses query engine when you have your own retrieval --- ## Server-Sent Events (SSE) Streaming Pattern ```typescript // Integrate with your HTTP framework's request/response API import { ContextChatEngine } from "llamaindex"; // Pre-initialize chat engine at startup (not per-request) let chatEngine: ContextChatEngine; async function handleChatStream(message: string): Promise<ReadableStream> { return new ReadableStream({ async start(controller) { const encoder = new TextEncoder(); const stream = await chatEngine.chat({ message, stream: true }); for await (const chunk of stream) { const content = chunk.message.content; if (content) { controller.enqueue( encoder.encode(`data: ${JSON.stringify({ content })}\n\n`), ); } } controller.enqueue(encoder.encode("data: [DONE]\n\n")); controller.close(); }, }); } // Return ReadableStream with Content-Type: text/event-stream from your route handler ``` **Why good:** Framework-agnostic ReadableStream, SSE-compatible format, pre-initialized engine avoids per-request overhead --- _For core setup and indexing, see [core.md](core.md). For agent patterns, see [agents.md](agents.md)._ -
core.md 6.3 KB
# LlamaIndex.TS -- Setup, Indexing & Querying Examples > Core patterns for Settings configuration, document loading, VectorStoreIndex, query engines, and persistence. See [SKILL.md](../SKILL.md) for concepts and decisions. **Related examples:** - [agents.md](agents.md) -- Agent creation, tools, multi-agent workflows - [chat-streaming.md](chat-streaming.md) -- Chat engines, streaming responses - [ingestion.md](ingestion.md) -- Text splitters, node parsers, custom readers --- ## Settings Configuration -- OpenAI ```typescript // lib/settings.ts import "dotenv/config"; import { Settings } from "llamaindex"; import { openai, OpenAIEmbedding } from "@llamaindex/openai"; Settings.llm = openai({ model: "gpt-4o" }); Settings.embedModel = new OpenAIEmbedding({ model: "text-embedding-3-small" }); export { Settings }; ``` **Why good:** Centralized config in a single module, import triggers setup, explicit provider --- ## Settings Configuration -- Ollama (Local) ```typescript // lib/settings.ts import { Settings, SentenceSplitter } from "llamaindex"; import { ollama } from "@llamaindex/ollama"; import { HuggingFaceEmbedding } from "@llamaindex/huggingface"; const CHUNK_SIZE = 512; const CHUNK_OVERLAP = 50; // Local LLM -- no API key needed Settings.llm = ollama({ model: "mixtral:8x7b" }); // Local embeddings -- no API key needed Settings.embedModel = new HuggingFaceEmbedding({ modelType: "BAAI/bge-small-en-v1.5", }); Settings.nodeParser = new SentenceSplitter({ chunkSize: CHUNK_SIZE, chunkOverlap: CHUNK_OVERLAP, }); ``` **Why good:** Fully local setup, no API keys, named constants for chunk params, explicit embedding model --- ## Basic RAG Pipeline ```typescript // rag.ts import "dotenv/config"; import { Settings, SimpleDirectoryReader, VectorStoreIndex } from "llamaindex"; import { openai, OpenAIEmbedding } from "@llamaindex/openai"; Settings.llm = openai({ model: "gpt-4o" }); Settings.embedModel = new OpenAIEmbedding({ model: "text-embedding-3-small" }); async function main() { // Step 1: Load documents const documents = await new SimpleDirectoryReader().loadData({ directoryPath: "./data", }); // Step 2: Create vector index (embeds all chunks) const index = await VectorStoreIndex.fromDocuments(documents); // Step 3: Create query engine const queryEngine = index.asQueryEngine(); // Step 4: Query const response = await queryEngine.query({ query: "What are the main topics in these documents?", }); console.log("Answer:", response.message.content); // Step 5: Inspect source nodes for debugging/citations for (const node of response.sourceNodes ?? []) { console.log( `Score: ${node.score}, Text: ${node.node.getText().slice(0, 100)}...`, ); } } main().catch(console.error); ``` **Why good:** Complete pipeline, error handling via catch, source node inspection for debugging --- ## Persisted Index -- Load or Create Pattern ```typescript // index-manager.ts import { VectorStoreIndex, storageContextFromDefaults, SimpleDirectoryReader, } from "llamaindex"; import { existsSync } from "node:fs"; const PERSIST_DIR = "./storage"; const DATA_DIR = "./data"; async function getOrCreateIndex(): Promise<VectorStoreIndex> { const storageContext = await storageContextFromDefaults({ persistDir: PERSIST_DIR, }); if (existsSync(PERSIST_DIR)) { // Load existing index console.log("Loading index from storage..."); return await VectorStoreIndex.init({ storageContext }); } // Create new index console.log("Creating new index..."); const documents = await new SimpleDirectoryReader().loadData({ directoryPath: DATA_DIR, }); return await VectorStoreIndex.fromDocuments(documents, { storageContext }); } const index = await getOrCreateIndex(); const queryEngine = index.asQueryEngine(); ``` **Why good:** Checks for existing index first, avoids re-embedding, named constants for paths --- ## Custom Retriever Configuration ```typescript const SIMILARITY_TOP_K = 5; // Configure retriever with custom top-k const retriever = index.asRetriever({ similarityTopK: SIMILARITY_TOP_K, }); // Use retriever directly for retrieval without synthesis const nodes = await retriever.retrieve({ query: "TypeScript generics" }); for (const node of nodes) { console.log( `[${node.score?.toFixed(3)}] ${node.node.getText().slice(0, 200)}`, ); } // Or use with query engine const queryEngine = index.asQueryEngine({ similarityTopK: SIMILARITY_TOP_K, }); ``` **Why good:** Named constant for top-k, direct retriever access for debugging, score inspection --- ## Multiple Indexes for Different Data Sources ```typescript import { VectorStoreIndex, SimpleDirectoryReader, QueryEngineTool, } from "llamaindex"; // Separate indexes for different document types const techDocs = await new SimpleDirectoryReader().loadData({ directoryPath: "./data/technical", }); const techIndex = await VectorStoreIndex.fromDocuments(techDocs); const hrDocs = await new SimpleDirectoryReader().loadData({ directoryPath: "./data/hr-policies", }); const hrIndex = await VectorStoreIndex.fromDocuments(hrDocs); // Wrap as tools for agent use const techTool = new QueryEngineTool({ queryEngine: techIndex.asQueryEngine(), metadata: { name: "technical_docs", description: "Search technical documentation for API references and guides", }, }); const hrTool = new QueryEngineTool({ queryEngine: hrIndex.asQueryEngine(), metadata: { name: "hr_policies", description: "Search HR policies for leave, benefits, and company rules", }, }); ``` **Why good:** Separate indexes for domain separation, descriptive tool metadata guides the LLM --- ## Response with Source Citations ```typescript const response = await queryEngine.query({ query: "What is the refund policy?", }); console.log("Answer:", response.message.content); console.log("\nSources:"); for (const node of response.sourceNodes ?? []) { const metadata = node.node.metadata; const score = node.score?.toFixed(3) ?? "N/A"; const preview = node.node.getText().slice(0, 150); console.log( ` - [${score}] ${metadata.file_path ?? "unknown"}: ${preview}...`, ); } ``` **Why good:** Extracts file path from metadata, shows relevance score, truncates preview text --- _For agent patterns, see [agents.md](agents.md). For chat and streaming, see [chat-streaming.md](chat-streaming.md). For API reference, see [reference.md](../reference.md)._ -
ingestion.md 6.6 KB
# LlamaIndex.TS -- Ingestion & Text Splitting Examples > Text splitters, node parsers, custom readers, and ingestion patterns. See [core.md](core.md) for basic setup. **Prerequisites:** Understand Settings configuration and VectorStoreIndex from [core.md](core.md). **Related examples:** - [core.md](core.md) -- Setup, indexing, query engines - [agents.md](agents.md) -- Agent creation, tools - [chat-streaming.md](chat-streaming.md) -- Chat engines, streaming --- ## SentenceSplitter Configuration ```typescript import { SentenceSplitter, Settings } from "llamaindex"; const CHUNK_SIZE = 512; const CHUNK_OVERLAP = 50; // Set globally -- affects all indexing operations Settings.nodeParser = new SentenceSplitter({ chunkSize: CHUNK_SIZE, chunkOverlap: CHUNK_OVERLAP, }); ``` **Why good:** Named constants, sentence-aware splitting preserves meaning, overlap ensures context continuity at boundaries ```typescript // BAD: Magic numbers, no overlap Settings.nodeParser = new SentenceSplitter({ chunkSize: 100 }); ``` **Why bad:** Tiny chunks lose context, no overlap means important sentences get split across chunks, magic number --- ## Standalone Text Splitting ```typescript import { SentenceSplitter } from "llamaindex"; const CHUNK_SIZE = 256; const splitter = new SentenceSplitter({ chunkSize: CHUNK_SIZE }); // Split a single text const chunks = splitter.splitText( "LlamaIndex is a data framework for building LLM applications. " + "It provides tools for document loading, text splitting, indexing, " + "and querying. The framework supports multiple LLM providers.", ); console.log(`Split into ${chunks.length} chunks`); chunks.forEach((chunk, i) => console.log(`Chunk ${i}: ${chunk}`)); ``` **Why good:** Useful for testing chunk sizes before indexing, named constant, standalone usage without Settings --- ## MarkdownNodeParser for Structured Documents ```typescript import { MarkdownNodeParser } from "llamaindex"; import { MarkdownReader } from "@llamaindex/readers/markdown"; const reader = new MarkdownReader(); const documents = await reader.loadData("./docs/api-reference.md"); const parser = new MarkdownNodeParser(); const nodes = parser(documents); // Nodes preserve markdown structure in metadata for (const node of nodes) { console.log("Headers:", node.metadata); // e.g. { 'Header 1': 'API Reference', 'Header 2': 'Authentication' } console.log("Text:", node.getText().slice(0, 100)); } ``` **Why good:** Structure-aware parsing, header hierarchy preserved in metadata, enables hierarchical retrieval --- ## CodeSplitter for Source Code ```typescript import { CodeSplitter } from "@llamaindex/node-parser/code"; import Parser from "tree-sitter"; import TS from "tree-sitter-typescript"; const MAX_CHARS = 1500; const parser = new Parser(); parser.setLanguage(TS.typescript); const codeSplitter = new CodeSplitter({ getParser: () => parser, maxChars: MAX_CHARS, }); // Splits code at function/class boundaries, not arbitrary character positions const chunks = codeSplitter.splitText(` export function calculateTotal(items: Item[]): number { return items.reduce((sum, item) => sum + item.price * item.quantity, 0); } export function formatCurrency(amount: number): string { return new Intl.NumberFormat("en-US", { style: "currency", currency: "USD", }).format(amount); } `); ``` **Why good:** AST-aware splitting preserves function boundaries, named constant for max chars, language-specific parsing --- ## Custom Reader for API Data ```typescript import { Document } from "llamaindex"; import type { BaseReader } from "llamaindex"; class ApiDocumentReader implements BaseReader { private readonly baseUrl: string; constructor(baseUrl: string) { this.baseUrl = baseUrl; } async loadData(): Promise<Document[]> { const response = await fetch(`${this.baseUrl}/documents`); const items = (await response.json()) as Array<{ id: string; title: string; content: string; }>; return items.map( (item) => new Document({ text: item.content, metadata: { source: this.baseUrl, title: item.title, id: item.id, }, }), ); } } // Usage const reader = new ApiDocumentReader("https://api.example.com"); const documents = await reader.loadData(); ``` **Why good:** Implements BaseReader interface, metadata for source tracking, typed API response --- ## SimpleDirectoryReader with Custom File Handlers ```typescript import { SimpleDirectoryReader } from "llamaindex"; import { PDFReader } from "@llamaindex/readers/pdf"; const NUM_WORKERS = 4; const reader = new SimpleDirectoryReader(); const documents = await reader.loadData({ directoryPath: "./data/mixed-files", numWorkers: NUM_WORKERS, fileExtToReader: { // Override default PDF reader with a specific one ".pdf": new PDFReader(), }, defaultReader: undefined, // Skip unsupported file types instead of using TextFileReader }); console.log(`Loaded ${documents.length} documents`); ``` **Why good:** Parallel loading with workers, custom reader mapping, explicit skip for unsupported types --- ## Chunk Size Selection Guide ```typescript import { SentenceSplitter, Settings } from "llamaindex"; // Short Q&A pairs, FAQs const FAQ_CHUNK_SIZE = 256; const FAQ_CHUNK_OVERLAP = 25; // Technical documentation const DOCS_CHUNK_SIZE = 512; const DOCS_CHUNK_OVERLAP = 50; // Long narratives, reports const REPORT_CHUNK_SIZE = 1024; const REPORT_CHUNK_OVERLAP = 100; // Choose based on your document type Settings.nodeParser = new SentenceSplitter({ chunkSize: DOCS_CHUNK_SIZE, chunkOverlap: DOCS_CHUNK_OVERLAP, }); ``` **Why good:** Named constants with descriptive names, chunk overlap proportional to chunk size (~10%) --- ## Document Metadata for Filtering ```typescript import { Document, VectorStoreIndex, MetadataFilter } from "llamaindex"; const documents = [ new Document({ text: "Product A costs $99 per month...", metadata: { category: "pricing", product: "A" }, }), new Document({ text: "Product B includes enterprise features...", metadata: { category: "features", product: "B" }, }), ]; const index = await VectorStoreIndex.fromDocuments(documents); // Query with metadata filtering (when using compatible vector stores) const queryEngine = index.asQueryEngine({ preFilters: { filters: [{ key: "category", value: "pricing", operator: "==" }], }, }); ``` **Why good:** Metadata enables filtered retrieval, reduces noise in results, category-based routing --- _For core patterns, see [core.md](core.md). For agent patterns, see [agents.md](agents.md). For API reference, see [reference.md](../reference.md)._
-
-
reference.md 10.9 KB
# LlamaIndex.TS Quick Reference > Package map, method signatures, response modes, and model providers. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples. --- ## Package Installation ```bash # Core framework (always required) npm install llamaindex # LLM providers (install one or more) npm install @llamaindex/openai # OpenAI (GPT-4o, GPT-5, embeddings) npm install @llamaindex/anthropic # Anthropic (Claude) npm install @llamaindex/ollama # Local models via Ollama npm install @llamaindex/groq # Groq npm install @llamaindex/gemini # Google Gemini # Agent workflows npm install @llamaindex/workflow # agent(), multiAgent(), agentStreamEvent # Readers (optional -- SimpleDirectoryReader is in core) npm install @llamaindex/readers # Additional file readers # Embedding providers npm install @llamaindex/huggingface # Local embeddings (BAAI/bge-small-en-v1.5) # Tools npm install @llamaindex/tools # Built-in tools (wiki, MCP) # Validation npm install zod # For tool parameter schemas ``` --- ## Import Map | Import | Package | Purpose | | ---------------------------- | ---------------------- | ------------------------------------------------ | | `Settings` | `llamaindex` | Global singleton for LLM, embedding, node parser | | `VectorStoreIndex` | `llamaindex` | Vector-based document index | | `SummaryIndex` | `llamaindex` | Full-document summary index | | `SimpleDirectoryReader` | `llamaindex` | Load files from a directory | | `SentenceSplitter` | `llamaindex` | Sentence-aware text chunking | | `ContextChatEngine` | `llamaindex` | Chat engine with retrieval context | | `storageContextFromDefaults` | `llamaindex` | Storage context for persistence | | `tool` | `llamaindex` | Define tools with Zod schemas | | `FunctionTool` | `llamaindex` | Wrap functions as agent tools | | `QueryEngineTool` | `llamaindex` | Wrap query engines as agent tools | | `getResponseSynthesizer` | `llamaindex` | Factory for response synthesis modes | | `MarkdownNodeParser` | `llamaindex` | Structure-aware markdown splitting | | `openai` / `OpenAIEmbedding` | `@llamaindex/openai` | OpenAI LLM and embedding provider | | `ollama` | `@llamaindex/ollama` | Ollama local model provider | | `agent` / `multiAgent` | `@llamaindex/workflow` | Agent and multi-agent creation | | `agentStreamEvent` | `@llamaindex/workflow` | Typed event filter for streaming | --- ## Settings Singleton ```typescript import { Settings, SentenceSplitter } from "llamaindex"; import { openai, OpenAIEmbedding } from "@llamaindex/openai"; Settings.llm = openai({ model: "gpt-4o" }); Settings.embedModel = new OpenAIEmbedding({ model: "text-embedding-3-small" }); Settings.nodeParser = new SentenceSplitter({ chunkSize: 512 }); Settings.chunkSize = 512; // Shortcut for node parser chunk size Settings.chunkOverlap = 50; // Shortcut for node parser overlap ``` ### Settings Properties | Property | Type | Default | Purpose | | ----------------- | ----------------- | --------------------- | --------------------------------- | | `llm` | `LLM` | OpenAI (lazy) | Language model for generation | | `embedModel` | `BaseEmbedding` | OpenAI ada-002 (lazy) | Embedding model for vector index | | `nodeParser` | `NodeParser` | SentenceSplitter | Document chunking strategy | | `chunkSize` | `number` | 1024 | Token count per chunk | | `chunkOverlap` | `number` | 20 | Token overlap between chunks | | `callbackManager` | `CallbackManager` | empty | Event callbacks for observability | --- ## VectorStoreIndex Methods ```typescript // Create from documents (embeds all chunks) const index = await VectorStoreIndex.fromDocuments(documents, { storageContext?, // StorageContext for persistence serviceContext?, // Deprecated -- use Settings instead }); // Load from persisted storage const index = await VectorStoreIndex.init({ storageContext, // Required: StorageContext with persistDir }); // Get query engine const queryEngine = index.asQueryEngine({ similarityTopK?: number, // Number of results to retrieve (default: 2) responseSynthesizer?: ResponseSynthesizer, }); // Get retriever const retriever = index.asRetriever({ similarityTopK?: number, // Number of results to retrieve }); // Get chat engine const chatEngine = index.asChatEngine(); ``` --- ## Query Engine Interface ```typescript // Basic query const response = await queryEngine.query({ query: "Your question" }); response.message.content; // string -- the answer response.sourceNodes; // NodeWithScore[] -- retrieved chunks with relevance scores // Streaming query const stream = await queryEngine.query({ query: "Your question", stream: true, }); for await (const chunk of stream) { process.stdout.write(chunk.message.content); } ``` --- ## Chat Engine Interface ```typescript // Basic chat (maintains conversation history) const response = await chatEngine.chat({ message: "Hello" }); response.message.content; // string -- the reply // Streaming chat const stream = await chatEngine.chat({ message: "Hello", stream: true }); for await (const chunk of stream) { process.stdout.write(chunk.message.content); } // Reset conversation chatEngine.reset(); ``` --- ## Response Synthesizer Modes | Mode | Behavior | Best For | | ------------------- | -------------------------------------------- | ------------------- | | `compact` (default) | Stuffs chunks into prompt, refines if needed | General purpose | | `refine` | Sequential refinement through each chunk | Detailed answers | | `tree_summarize` | Recursive tree summarization | Summarization tasks | | `multi_modal` | Handles text + images/audio | Multi-modal queries | ```typescript import { getResponseSynthesizer, responseModeSchema } from "llamaindex"; const synthesizer = getResponseSynthesizer(responseModeSchema.Enum.compact); ``` --- ## Agent API ```typescript import { agent, multiAgent, agentStreamEvent } from "@llamaindex/workflow"; import { tool } from "llamaindex"; import { z } from "zod"; // Define a tool const myTool = tool({ name: "toolName", description: "What this tool does (max 125 chars)", parameters: z.object({ /* Zod schema */ }), execute: async (input) => { /* implementation */ }, }); // Single agent const myAgent = agent({ tools: [myTool], llm?: LLM, // Override Settings.llm systemPrompt?: string, }); // Run const result = await myAgent.run("prompt"); result.data; // Response data // Stream for await (const event of myAgent.runStream("prompt")) { if (agentStreamEvent.include(event)) { process.stdout.write(event.data.delta); } } // Multi-agent const agents = multiAgent({ agents: [agentA, agentB], rootAgent: agentA, }); ``` --- ## SimpleDirectoryReader Options ```typescript const reader = new SimpleDirectoryReader(); const documents = await reader.loadData({ directoryPath: "./data", // Required: path to directory numWorkers?: number, // Concurrent workers (default: 1, max: 9) overrideReader?: BaseReader, // Use one reader for all files fileExtToReader?: Record<string, BaseReader>, // Map extensions to readers defaultReader?: BaseReader, // Fallback reader (default: TextFileReader) }); ``` ### Supported File Types (Default) | Extension | Reader | | ---------------- | -------------- | | `.txt` | TextFileReader | | `.pdf` | PDFReader | | `.csv` | CSVReader | | `.md` | MarkdownReader | | `.docx` | DocxReader | | `.html` / `.htm` | HTMLReader | | `.jpg` / `.png` | ImageReader | --- ## Node Parsers | Parser | Import | Use Case | | -------------------- | ------------------------------ | ------------------------------------------ | | `SentenceSplitter` | `llamaindex` | General text (default) | | `MarkdownNodeParser` | `llamaindex` | Markdown with header hierarchy | | `CodeSplitter` | `@llamaindex/node-parser/code` | Source code (AST-aware, needs tree-sitter) | --- ## LLM Provider Configuration ### OpenAI ```typescript import { openai, OpenAIEmbedding } from "@llamaindex/openai"; Settings.llm = openai({ model: "gpt-4o", apiKey: process.env.OPENAI_API_KEY }); Settings.embedModel = new OpenAIEmbedding({ model: "text-embedding-3-small" }); ``` ### Ollama (Local) ```typescript import { ollama } from "@llamaindex/ollama"; Settings.llm = ollama({ model: "mixtral:8x7b" }); ``` ### Anthropic ```typescript import { anthropic } from "@llamaindex/anthropic"; Settings.llm = anthropic({ model: "claude-sonnet-4-20250514" }); ``` --- ## TypeScript Configuration ```json { "compilerOptions": { "target": "ES2022", "module": "ESNext", "moduleResolution": "bundler", "esModuleInterop": true, "skipLibCheck": true, "strict": true, "lib": ["ES2022", "DOM.AsyncIterable"] } } ``` **Required settings:** - `"moduleResolution": "bundler"` or `"nodenext"` -- classic `"node"` will fail - `"DOM.AsyncIterable"` in `lib` -- needed for Web Stream API types - Node.js >= 20 --- ## Environment Variables | Variable | Provider | Purpose | | ------------------- | --------- | ------------------------------------------------------------ | | `OPENAI_API_KEY` | OpenAI | API key (auto-detected by `@llamaindex/openai`) | | `ANTHROPIC_API_KEY` | Anthropic | API key | | `OLLAMA_HOST` | Ollama | Custom Ollama server URL (default: `http://localhost:11434`) | --- ## Storage Context ```typescript import { storageContextFromDefaults } from "llamaindex"; // Create with persistence const ctx = await storageContextFromDefaults({ persistDir: "./storage" }); // Files created in persistDir: // - docstore.json (document metadata) // - index_store.json (index structure) // - vector_store.json (embeddings) // - graph_store.json (graph relationships) ``` -
SKILL.md 18.7 KB
--- name: ai-orchestration-llamaindex description: LlamaIndex.TS data framework for RAG, indexing, retrieval, query engines, chat engines, and agentic workflows in TypeScript --- # LlamaIndex.TS Patterns > **Quick Guide:** LlamaIndex.TS is a data framework for building context-aware LLM applications in TypeScript. Use `Settings` singleton to configure LLM and embedding models globally. Load documents with `SimpleDirectoryReader`, chunk with `SentenceSplitter`, index with `VectorStoreIndex.fromDocuments()`, and query with `index.asQueryEngine()`. For agents, use `agent()` from `@llamaindex/workflow` with `tool()` definitions using Zod schemas. All core operations are async -- every function returns a Promise. The `llamaindex` package re-exports most things, but LLM providers require separate packages like `@llamaindex/openai` or `@llamaindex/ollama`. --- <critical_requirements> ## CRITICAL: Before Using This Skill > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST configure `Settings.llm` and `Settings.embedModel` before any indexing or querying -- the Settings singleton is lazily initialized and defaults to OpenAI, which will fail without an API key)** **(You MUST await all LlamaIndex operations -- `fromDocuments()`, `asQueryEngine()`, `query()`, `chat()`, `loadData()` are ALL async)** **(You MUST install provider packages separately -- `@llamaindex/openai`, `@llamaindex/ollama`, `@llamaindex/anthropic` are NOT included in the base `llamaindex` package)** **(You MUST use `storageContextFromDefaults({ persistDir })` to persist indexes -- without persistence, indexes are rebuilt from scratch on every restart)** **(You MUST never hardcode API keys -- use environment variables and `dotenv/config`)** </critical_requirements> --- **Auto-detection:** LlamaIndex, llamaindex, VectorStoreIndex, SimpleDirectoryReader, Settings.llm, Settings.embedModel, asQueryEngine, asChatEngine, ContextChatEngine, SentenceSplitter, storageContextFromDefaults, @llamaindex/openai, @llamaindex/ollama, @llamaindex/workflow, FunctionTool, QueryEngineTool, agentStreamEvent **When to use:** - Building RAG (Retrieval-Augmented Generation) applications with custom documents - Loading, chunking, and indexing documents for LLM consumption - Creating query engines that answer questions from indexed data - Building chat interfaces with conversation memory over your data - Implementing agentic RAG with tool-calling agents that query indexes - Working with multiple data sources (files, PDFs, markdown, code) - Persisting vector indexes to avoid re-indexing on every restart **Key patterns covered:** - Settings singleton for LLM and embedding model configuration - Document loading with SimpleDirectoryReader and custom readers - VectorStoreIndex creation, persistence, and querying - Query engines and chat engines - Agent creation with `agent()` and `tool()` using Zod schemas - Text splitting and chunking strategies - Streaming responses from query and chat engines - Storage context and index persistence **When NOT to use:** - Simple one-shot LLM calls without document context -- use the LLM provider SDK directly - Applications that only need embeddings without indexing -- use the embedding API directly - Client-side / browser applications -- LlamaIndex.TS is server-side focused (Node.js >= 20) --- ## Examples Index - [Core: Setup, Indexing & Querying](examples/core.md) -- Settings config, document loading, VectorStoreIndex, query engines, persistence - [Agents & Tools](examples/agents.md) -- FunctionTool, agent(), multi-agent workflows, QueryEngineTool - [Chat & Streaming](examples/chat-streaming.md) -- Chat engines, ContextChatEngine, streaming responses - [Ingestion & Splitting](examples/ingestion.md) -- Text splitters, node parsers, ingestion pipeline, custom readers - [Quick API Reference](reference.md) -- Package map, method signatures, response modes, model providers --- <philosophy> ## Philosophy LlamaIndex.TS is a **data framework** -- its core value proposition is connecting your data to LLMs through indexing, retrieval, and synthesis. It sits between raw LLM APIs and full application frameworks. **Core principles:** 1. **Context engineering** -- Inject the right data into the LLM prompt at the right time. This drives RAG, agent memory, extraction, and summarization. 2. **Modular provider system** -- LLM providers, embedding models, vector stores, and readers are separate packages you compose. The base `llamaindex` package provides the framework; providers are installed separately. 3. **Settings singleton** -- Global configuration for LLM, embedding model, node parser, and other shared resources. Set once, used everywhere. Override locally when needed. 4. **Async-first design** -- Every I/O operation is async. Document loading, indexing, querying, and chat all return Promises. 5. **Index as the core abstraction** -- Documents are loaded, split into nodes, embedded, and stored in an index. Queries retrieve relevant nodes and synthesize responses. **When to use LlamaIndex.TS:** - You have documents/data that need to be indexed for LLM consumption - You want structured RAG pipelines with configurable retrieval and synthesis - You need agentic RAG where agents query multiple indexes with tools - You want persistence and incremental updates to your index **When NOT to use:** - Simple LLM calls without data context -- use the provider SDK directly - Browser-only applications -- LlamaIndex.TS requires Node.js >= 20 - You only need embeddings -- use the embedding API directly </philosophy> --- <patterns> ## Core Patterns ### Pattern 1: Settings Configuration The `Settings` singleton configures LLM, embedding model, and node parser globally. Set it once at application startup before any indexing or querying. ```typescript import { Settings } from "llamaindex"; import { openai, OpenAIEmbedding } from "@llamaindex/openai"; // Configure at app startup -- before any index operations Settings.llm = openai({ model: "gpt-4o" }); Settings.embedModel = new OpenAIEmbedding({ model: "text-embedding-3-small" }); ``` **Why good:** Single configuration point, provider packages are explicit imports, model names are visible ```typescript // BAD: No Settings configuration, relying on implicit defaults import { VectorStoreIndex, SimpleDirectoryReader } from "llamaindex"; // This will silently try to use OpenAI with OPENAI_API_KEY from env // Fails with cryptic error if key is missing const documents = await new SimpleDirectoryReader().loadData({ directoryPath: "./data", }); const index = await VectorStoreIndex.fromDocuments(documents); ``` **Why bad:** Implicit defaults make failures confusing, no explicit provider, no model selection **See:** [examples/core.md](examples/core.md) for local LLM setup with Ollama, Anthropic configuration, and embedding model options --- ### Pattern 2: Document Loading and Indexing Load documents, create a vector index, and query it. This is the canonical RAG pipeline. ```typescript import { SimpleDirectoryReader, VectorStoreIndex, Settings } from "llamaindex"; import { openai, OpenAIEmbedding } from "@llamaindex/openai"; Settings.llm = openai({ model: "gpt-4o" }); Settings.embedModel = new OpenAIEmbedding({ model: "text-embedding-3-small" }); // Load all supported files from a directory const documents = await new SimpleDirectoryReader().loadData({ directoryPath: "./data", }); // Create vector index -- embeds and stores all document chunks const index = await VectorStoreIndex.fromDocuments(documents); // Query the index const queryEngine = index.asQueryEngine(); const response = await queryEngine.query({ query: "What is the main topic?" }); console.log(response.message.content); ``` **Why good:** Complete pipeline in minimal code, explicit Settings, clear data flow **See:** [examples/core.md](examples/core.md) for persistence, custom readers, and advanced indexing options --- ### Pattern 3: Index Persistence Persist indexes to disk to avoid re-indexing on every restart. ```typescript import { VectorStoreIndex, storageContextFromDefaults, SimpleDirectoryReader, } from "llamaindex"; const PERSIST_DIR = "./storage"; // First run: create and persist const storageContext = await storageContextFromDefaults({ persistDir: PERSIST_DIR, }); const documents = await new SimpleDirectoryReader().loadData({ directoryPath: "./data", }); const index = await VectorStoreIndex.fromDocuments(documents, { storageContext, }); // Subsequent runs: load from storage const loadedStorageContext = await storageContextFromDefaults({ persistDir: PERSIST_DIR, }); const loadedIndex = await VectorStoreIndex.init({ storageContext: loadedStorageContext, }); ``` **Why good:** Named constant for path, separate create vs load paths, storage context reuse ```typescript // BAD: Rebuilding index on every request async function handleQuery(question: string) { const docs = await new SimpleDirectoryReader().loadData({ directoryPath: "./data", }); const index = await VectorStoreIndex.fromDocuments(docs); // Expensive! const engine = index.asQueryEngine(); return engine.query({ query: question }); } ``` **Why bad:** Re-indexes all documents on every call, wastes time and API credits on re-embedding **See:** [examples/core.md](examples/core.md) for load-or-create pattern --- ### Pattern 4: Agents with Tool Definitions Create agents that use tools defined with Zod schemas. Use `agent()` from `@llamaindex/workflow`. ```typescript import { tool, Settings } from "llamaindex"; import { agent, agentStreamEvent } from "@llamaindex/workflow"; import { openai } from "@llamaindex/openai"; import { z } from "zod"; Settings.llm = openai({ model: "gpt-4o" }); const weatherTool = tool({ name: "getWeather", description: "Get current weather for a city", parameters: z.object({ city: z.string({ description: "City name" }), }), execute: async ({ city }) => { // Your weather API call here return { temperature: 22, condition: "sunny" }; }, }); const myAgent = agent({ tools: [weatherTool] }); const result = await myAgent.run("What's the weather in Paris?"); console.log(result.data); ``` **Why good:** Zod schema for type-safe parameters, description guides the LLM, async execute function **See:** [examples/agents.md](examples/agents.md) for multi-agent workflows, QueryEngineTool, streaming agents --- ### Pattern 5: Chat Engine Build conversational interfaces over your indexed data with conversation memory. ```typescript import { VectorStoreIndex, ContextChatEngine, SimpleDirectoryReader, } from "llamaindex"; const documents = await new SimpleDirectoryReader().loadData({ directoryPath: "./data", }); const index = await VectorStoreIndex.fromDocuments(documents); const retriever = index.asRetriever({ similarityTopK: 3 }); const chatEngine = new ContextChatEngine({ retriever }); // Multi-turn conversation -- chat engine maintains history const response1 = await chatEngine.chat({ message: "What is LlamaIndex?" }); console.log(response1.message.content); const response2 = await chatEngine.chat({ message: "How does it handle streaming?", }); console.log(response2.message.content); ``` **Why good:** Retriever-based context injection, automatic conversation history, multi-turn support **See:** [examples/chat-streaming.md](examples/chat-streaming.md) for streaming chat, system prompts, chat history management --- ### Pattern 6: Streaming Responses Stream responses for user-facing applications. ```typescript import { agentStreamEvent } from "@llamaindex/workflow"; // Agent streaming const events = myAgent.runStream("Tell me about TypeScript"); for await (const event of events) { if (agentStreamEvent.include(event)) { process.stdout.write(event.data.delta); } } // Query engine streaming const response = await queryEngine.query({ query: "Summarize the document", stream: true, }); for await (const chunk of response) { process.stdout.write(chunk.message.content); } ``` **Why good:** Event-based agent streaming with typed filters, query engine streaming with for-await **See:** [examples/chat-streaming.md](examples/chat-streaming.md) for response synthesizer streaming, chat engine streaming --- ### Pattern 7: Text Splitting and Node Parsing Configure how documents are chunked before indexing. ```typescript import { SentenceSplitter, Settings } from "llamaindex"; const CHUNK_SIZE = 512; const CHUNK_OVERLAP = 50; // Set globally via Settings Settings.nodeParser = new SentenceSplitter({ chunkSize: CHUNK_SIZE, chunkOverlap: CHUNK_OVERLAP, }); // Or use standalone const splitter = new SentenceSplitter({ chunkSize: CHUNK_SIZE }); const texts = splitter.splitText("Your long document text here..."); ``` **Why good:** Named constants for chunk parameters, global vs standalone usage shown, sentence-aware splitting ```typescript // BAD: Using default chunk size without considering document characteristics const index = await VectorStoreIndex.fromDocuments(documents); // Default chunk size may be too large for short Q&A or too small for long narratives ``` **Why bad:** Default chunk size (1024 tokens) may not suit your data, causes poor retrieval quality **See:** [examples/ingestion.md](examples/ingestion.md) for MarkdownNodeParser, CodeSplitter, custom chunk strategies </patterns> --- <decision_framework> ## Decision Framework ### Which Index Type to Use ``` What is your use case? +-- Semantic search over documents -> VectorStoreIndex (most common) +-- Summarization of all documents -> SummaryIndex +-- Both search AND summarization -> Create both, use as separate tools in an agent +-- Hierarchical document structure -> Use MarkdownNodeParser + VectorStoreIndex ``` ### Query Engine vs Chat Engine vs Agent ``` How should users interact with your data? +-- Single question, single answer -> Query Engine (index.asQueryEngine()) +-- Multi-turn conversation -> Chat Engine (ContextChatEngine) +-- Multiple tools/indexes + reasoning -> Agent (agent() from @llamaindex/workflow) +-- Complex multi-step workflow -> Multi-agent with handoffs ``` ### Which LLM Provider ``` Which LLM provider are you using? +-- OpenAI -> npm install @llamaindex/openai +-- Anthropic -> npm install @llamaindex/anthropic +-- Local (Ollama) -> npm install @llamaindex/ollama +-- Groq -> npm install @llamaindex/groq +-- Google Gemini -> npm install @llamaindex/gemini ``` ### Chunk Size Selection ``` What kind of documents are you indexing? +-- Short Q&A pairs -> chunkSize: 256-512 +-- Technical documentation -> chunkSize: 512-1024 +-- Long narratives/reports -> chunkSize: 1024-2048 +-- Code files -> Use CodeSplitter (AST-aware) +-- Markdown -> Use MarkdownNodeParser (structure-aware) ``` </decision_framework> --- <red_flags> ## RED FLAGS **High Priority Issues:** - Not configuring `Settings.llm` before indexing/querying -- defaults to OpenAI, fails silently without API key - Forgetting to `await` async operations -- `fromDocuments()`, `query()`, `chat()` all return Promises - Rebuilding indexes on every request instead of persisting with `storageContextFromDefaults` - Hardcoding API keys instead of using environment variables - Installing only `llamaindex` without provider packages (`@llamaindex/openai`, etc.) **Medium Priority Issues:** - Using default chunk size (1024) without considering document characteristics -- causes poor retrieval - Not setting `similarityTopK` on retrievers -- default may return too few or too many results - Ignoring the response `sourceNodes` -- they contain the retrieved context for debugging and citations - Creating a new `SimpleDirectoryReader` per request instead of caching the loaded documents - Not handling the case where `response.message.content` might be empty on retrieval failure **Common Mistakes:** - Confusing `asQueryEngine()` (single question) with `ContextChatEngine` (multi-turn conversation) - Using `VectorStoreIndex.fromDocuments()` when you should use `VectorStoreIndex.init()` to load from storage - Importing `openai` from `llamaindex` instead of `@llamaindex/openai` -- the `llamaindex` package may re-export some things but provider-specific imports are more reliable - Passing `messages` array to `query()` -- query engines take `{ query: string }`, not a messages array - Using `index.asQueryEngine()` multiple times instead of storing the engine reference **Gotchas & Edge Cases:** - `Settings` is a global singleton -- setting it in one module affects all others. Override locally by passing `llm` directly to constructors when you need different models for different operations. - `SimpleDirectoryReader` only works on Node.js -- it uses `fs` internally. For edge/serverless, load documents differently or use LlamaParse. - `storageContextFromDefaults` creates four JSON files in the persist directory (`docstore.json`, `graph_store.json`, `index_store.json`, `vector_store.json`). If any are corrupted, delete the directory and re-index. - Node.js >= 20 is required. Some modules use Web Stream API (`ReadableStream`, `WritableStream`), so add `"DOM.AsyncIterable"` to `tsconfig.json` `lib` if you get type errors. - `tsconfig.json` must use `"moduleResolution": "bundler"` or `"nodenext"` -- the classic `"node"` resolution will fail to resolve LlamaIndex sub-packages. - Default tokenizer is slow -- install `gpt-tokenizer` for 60x faster tokenization. - `SentenceSplitter` chunk size is in tokens, not characters. A 512-token chunk is roughly 2000 characters. - The `llamaindex` package is large (~2MB+). For production, consider importing specific sub-packages to reduce bundle size. - `VectorStoreIndex.fromDocuments()` makes embedding API calls for every chunk. For large document sets, this can be expensive. Monitor costs. - Chat engine conversation history grows unbounded -- implement history pruning for long-running sessions. </red_flags> --- <critical_reminders> ## CRITICAL REMINDERS > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST configure `Settings.llm` and `Settings.embedModel` before any indexing or querying -- the Settings singleton is lazily initialized and defaults to OpenAI, which will fail without an API key)** **(You MUST await all LlamaIndex operations -- `fromDocuments()`, `asQueryEngine()`, `query()`, `chat()`, `loadData()` are ALL async)** **(You MUST install provider packages separately -- `@llamaindex/openai`, `@llamaindex/ollama`, `@llamaindex/anthropic` are NOT included in the base `llamaindex` package)** **(You MUST use `storageContextFromDefaults({ persistDir })` to persist indexes -- without persistence, indexes are rebuilt from scratch on every restart)** **(You MUST never hardcode API keys -- use environment variables and `dotenv/config`)** **Failure to follow these rules will produce broken RAG pipelines, wasted embedding API credits, or cryptic runtime errors.** </critical_reminders>
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.