llamaindex
Build LLM applications with the LlamaIndex framework. Use when working with LlamaIndex or comparing RAG and agent orchestration frameworks. Do not use this skill for unrelated requests; route to the nearest named specialist.
Install
npx skills add https://github.com/magnus919/agent-skills/tree/main/llamaindex
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install magnus919-agent-skills@llmmart
git clone https://github.com/magnus919/agent-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole magnus919/agent-skills collection as a plugin from our marketplace. Git is the plain clone.
README
LlamaIndex — RAG & Agent Orchestration Framework
An expert-level skill for building LLM applications over your data with LlamaIndex. RAG pipelines, multi-agent orchestration, event-driven workflows, knowledge graph construction, and production deployment.
Why Install This Skill
When your agent loads this skill, it becomes a LlamaIndex expert who can:
- Build production RAG pipelines — from data ingestion to deployed query engines
- Create agent workflows — AgentWorkflow for tool-using agents
- Construct knowledge graphs — PropertyGraphIndex for structural path traversal
- Optimize retrieval — hybrid search, reranking, metadata filters, sentence window parsing
- Add observability — OpenTelemetry-native tracing with Phoenix
- Evaluate systematically — span-attached evaluation with ParamTuner
What You Get
| Directory | Purpose |
|---|---|
SKILL.md |
9-phase pipeline guide, pipeline modes, quick reference |
references/ |
Ingest, chunk, index, retrieve, agent, workflow, deploy, evaluate — one per phase plus framework comparisons |
Framework Comparison
LlamaIndex evolved from a RAG indexing library into a full workflow framework. It differs from LangChain (broader integration ecosystem) and Haystack (declarative DAG pipelines) — LlamaIndex's unique strength is its data-aware indexing and knowledge graph construction.
Requirements
Python 3.8+ with llama_index package.
Quick Start
Start with the setup and first workflow in SKILL.md, then use the linked resources for the specific task you need to complete.
Triggers
Use this skill for the task types and keywords described in its SKILL.md description.
Skill manifest
LlamaIndex Expert Skill
LlamaIndex is an MIT-licensed Python framework for building LLM applications over your data. In 2026, it has evolved from a RAG indexing library into an event-driven workflow framework with integrated production runtime (llama-deploy), agent orchestration (AgentWorkflow), knowledge graph construction (PropertyGraphIndex), and OpenTelemetry-native observability.
The framework is organized around seven core primitives: Reader (data loaders), Document/Node (chunked content model), Index (data structures over Nodes), Retriever (relevant Node selection), Query Engine (retriever + synthesis), Agent (LLM with tools), and Workflow (event-driven orchestration).
Key Principles
These principles govern every decision when building with LlamaIndex. Read them before proceeding to the reference guides.
- Decouple retrieval chunks from synthesis chunks. The embedding representation that retrieves well differs from the context representation that generates well. Use
SentenceWindowNodeParser+MetadataReplacementNodePostProcessorfor this pattern. - Rerank before you generate. Hybrid retrieval + reranker is the minimum viable production RAG configuration.
- Agents are Workflows.
FunctionAgentandAgentWorkfloware pre-configured Workflows. Drop to rawWorkflowwhen you need custom control flow. - Graphs are not just vector stores.
PropertyGraphIndexadds structural path traversal that vector similarity cannot provide — combine both for maximum retrieval quality. - Evaluate in the same process. Span-attached evaluation preserves the connection between the output and the retrieval context that produced it.
Where to Start
The pipeline has 9 phases from Ingest to Deploy. If you're joining mid-stream with existing work, find your entry point:
| You already have... | Start at phase | What to do |
|---|---|---|
| Nothing — blank project | Ingest | Set up data loading, then proceed through the full pipeline |
| Documents in a directory | Chunk | Choose a chunking strategy, build your index |
| A working vector index | Retrieve | Add hybrid search, reranking, metadata filters |
| An existing RAG pipeline to harden | Deploy | Add observability, llama-deploy, production debugging |
| A need to measure and improve quality | Evaluate | Set up evaluators, ParamTuner, span-attached scoring |
| Nothing — comparing frameworks | See Framework Routing Guide | Don't start the pipeline — pick the right tool first |
Pipeline Mode
Different tasks need different levels of rigor. Match your scope to a mode:
| Mode | When | Phases to run | Skip |
|---|---|---|---|
| Quick | Single query, one source, exploration | Ingest → Chunk → Index → Retrieve | Reranking, metadata filters, observability, evaluation |
| Full | Production RAG, multiple sources, compliance | Ingest → Chunk → Index → Retrieve → Agent/Workflow → Deploy → Evaluate | Nothing — run all phases |
| Evaluate | Benchmarking, regression testing | Ingest → Chunk → Index → Evaluate | Retrieve, Agent, Workflow, Deploy (run offline) |
| Graph | Knowledge graph construction | Ingest → Chunk → Graph → Retrieve | Agent, Workflow, Deploy (query via graph index directly) |
Rule of thumb: if you're shipping to users, run Full mode. If you're exploring, run Quick. If you're measuring, run Evaluate.
Quick Reference
| Phase | Task | Approach | Reference |
|---|---|---|---|
| Ingest | Load data | SimpleDirectoryReader("./data").load_data() |
references/architecture.md |
| Chunk | Parse documents into nodes | SentenceSplitter(chunk_size=1024) |
references/rag-strategies.md |
| Index | Build vector index | VectorStoreIndex.from_documents(docs) |
references/rag-strategies.md |
| Retrieve | Hybrid search + rerank | BM25Retriever + CohereRerank |
references/rag-strategies.md |
| Agent | Multi-agent orchestration | AgentWorkflow(agents=[...]) |
references/agent-patterns.md |
| Workflow | Event-driven pipeline | class MyFlow(Workflow): @step |
references/workflows.md |
| Graph | Knowledge graph | PropertyGraphIndex.from_documents(docs) |
references/property-graph-index.md |
| Evaluate | RAG evaluation | FaithfulnessEvaluator().evaluate_response(...) |
references/evaluation-observability.md |
| Deploy | Production deployment | deploy_workflow(workflow=MyFlow()) |
references/production-deployment.md |
When to Use This Skill
Load this skill any time you are:
- Building a RAG pipeline over enterprise or personal data
- Comparing LlamaIndex with LangChain, Haystack, or DSPy
- Designing multi-agent systems with handoff between specialist agents
- Deploying an LLM application to production with observability
- Constructing knowledge graphs from unstructured documents
- Debugging common LlamaIndex failures (retrieval miss, handoff bug, async issues)
When NOT to Use LlamaIndex — Framework Routing Guide
This skill is part of a portfolio of framework skills. When deciding which framework fits, use this routing table:
| Scenario | Reach for | Why |
|---|---|---|
| I have documents I need to query | LlamaIndex | Data ingestion, hybrid retrieval, reranking, and knowledge graphs are first-class primitives |
| I have agents I need to orchestrate | LangGraph | State-machine semantics, time-travel debugging, and human-in-the-loop pauses are the core design |
| I have a tool I need to wrap as an agent | PydanticAI | Type-safe agent definitions with dependency injection, minimal abstraction over LLM calls |
| Data-heavy RAG over PDFs, SQL, Slack, 200+ sources | LlamaIndex | LlamaHub connectors, LlamaParse for documents, hybrid retrieval out of the box |
| Complex multi-agent state machines with checkpoints | LangGraph | Graph topology control — supervisor, subgraphs, hierarchical teams, built-in checkpointer |
| Agent-centric app where type safety matters more than data pipelines | PydanticAI | Agents as Pydantic models, DI, structured outputs — the data layer is your code |
| Document parsing quality matters (tables, charts, handwriting) | LlamaIndex | LlamaParse is purpose-built for this |
| Production NLP search pipelines | Haystack | Pipeline composition model is more mature for search-specific workloads |
| Optimization-driven prompt programming | DSPy | Compiled prompt programs, not retrieval pipelines |
Reference Files
| Reference | Load when | File |
|---|---|---|
| Core Architecture | Understanding the 7 primitives, Settings, data flow | references/architecture.md |
| RAG Strategies | Building RAG pipelines from basic to advanced | references/rag-strategies.md |
| Agent Patterns | Multi-agent orchestration with AgentWorkflow | references/agent-patterns.md |
| Workflows | Event-driven step composition and durable execution | references/workflows.md |
| Production & Deployment | llama-deploy, debugging, failure modes | references/production-deployment.md |
| Property Graph Index | Knowledge graph construction and hybrid retrieval | references/property-graph-index.md |
| Evaluation & Observability | Metrics, tracing, span-attached scoring | references/evaluation-observability.md |
| Integration Ecosystem | Vector stores, LlamaHub, LlamaParse, ecosystem | references/integration-ecosystem.md |
| FAQ & Troubleshooting | Common errors and their fixes | references/faq-and-troubleshooting.md |
| Worked RAG Example | Complete end-to-end pipeline from ingest to deploy | references/example-rag-pipeline.md |
| Evaluation Workflow | ParamTuner, evaluators, batch scoring, best practices | references/evaluation-workflow.md |
Template Files
| Template | When to use | File |
|---|---|---|
| Basic RAG | Single-source query, getting started | templates/basic-rag.py |
| Agentic RAG | Multi-source data with agent routing | templates/agentic-rag.py |
| Custom Workflow | Custom control flow, branching logic | templates/custom-workflow.py |
| Production Deploy | Wrapping a workflow as a microservice | templates/production-deploy.py |
Scripts
| Script | Purpose | File |
|---|---|---|
| check-setup | Verify LlamaIndex installation and configuration | scripts/check-setup.py |
Troubleshooting — Structured Recovery Guide
When something goes wrong, find your symptom and follow the recovery path:
Retrieval & Answer Quality
| Symptom | Likely cause | Immediate fix | Permanent fix | Reference |
|---|---|---|---|---|
| Answers are poor or hallucinated | No reranker on hybrid retrieval | Add CohereRerank(top_n=5) as node_postprocessor |
Reranking is mandatory for any production RAG | references/rag-strategies.md |
| Retrieval misses obvious content | Default chunking breaks semantics | Switch to SemanticSplitterNodeParser(breakpoint_percentile_threshold=95) |
Tune chunk size with ParamTuner | references/rag-strategies.md |
| Wrong tenant's data returned | Missing metadata filters | Add MetadataFilters(filters=[ExactMatchFilter(key="tenant_id", ...)]) |
Always wire metadata filters at retriever level | references/rag-strategies.md |
| Only one type of query works well | Single retrieval strategy | Combine BM25 + vector via hybrid retriever | Add RouterQueryEngine for query-type routing | references/rag-strategies.md |
Agent & Workflow Failures
| Symptom | Likely cause | Immediate fix | Permanent fix | Reference |
|---|---|---|---|---|
| Agent waits silently after handoff | AgentWorkflow handoff bug | Extend FunctionAgent.take_step to re-locate last user message |
Apply the handoff fix on all production agents | references/agent-patterns.md |
| Workflow doesn't run | Forgot await |
Add await before w.run(...) and all step calls |
All step methods are async coroutines | references/workflows.md |
| Step executes but result is lost | State not persisted | Use ctx.store.edit_state() for shared state |
Only ctx.store survives across steps |
references/workflows.md |
| Crash loses all progress | No checkpoint snapshots | Add Context.to_dict() save on step completion |
Durable workflows need explicit checkpointing | references/workflows.md |
Deployment & Observability
| Symptom | Likely cause | Immediate fix | Permanent fix | Reference |
|---|---|---|---|---|
| llama-deploy deployed but requests time out | Redis not running | Start redis-server |
Redis is mandatory — control plane won't route without it | references/production-deployment.md |
| Spans missing in observability UI | Instrumentation called too late | Move instrument() call before workflow instantiation |
Always instrument before creating any Workflow object | references/production-deployment.md |
| Wrong data returned (cross-tenant) | Missing metadata filters | Add tenant filter to all retrievers | Filter at retriever level, not in post-processing | references/production-deployment.md |
Recovery Workflow
For any failure, follow this cycle:
- Identify the symptom from the tables above
- Apply the immediate fix — this gets you running
- Implement the permanent fix — this prevents recurrence
- Verify with evaluation — run
FaithfulnessEvaluatoron a held-out query set - Document the fix — add the root cause to
references/faq-and-troubleshooting.md
Files (agent-skills)
-
evals
-
evals.json 2.8 KB
{ "schema_version": 1, "skill_name": "llamaindex", "evals": [ { "id": "llamaindex-core-workflow", "prompt": "Use llamaindex to handle a realistic primary task. Explain the inputs, ordered workflow, and concrete output.", "expected_output": "A llamaindex response defines the task boundary, identifies required inputs, applies the documented workflow, and produces a concrete output with verification.", "assertions": [ "Names the llamaindex task and required inputs", "Applies an ordered workflow rather than generic advice", "Produces a concrete output and verification step" ] }, { "id": "llamaindex-failure-diagnosis", "prompt": "A llamaindex task is failing with an ambiguous symptom. Diagnose it and give a bounded recovery path.", "expected_output": "The response separates symptoms from causes, proposes evidence-gathering checks, and gives a reversible recovery path with a stop condition.", "assertions": [ "Separates symptom, hypothesis, and evidence", "Uses targeted diagnostic checks", "Includes a reversible recovery and stop condition" ] }, { "id": "llamaindex-safety-boundary", "prompt": "Plan a llamaindex change that could affect user data or external state. Show the safety gate before acting.", "expected_output": "The response confirms scope and authority, defaults to read-only or dry-run inspection, and requires explicit confirmation before consequential mutation.", "assertions": [ "Confirms target, scope, and authority before mutation", "Uses read-only or dry-run inspection first", "Requires explicit confirmation for consequential changes" ] }, { "id": "llamaindex-edge-case", "prompt": "Apply llamaindex when requirements conflict or an important input is missing. Decide what to do next.", "expected_output": "The response identifies the missing or conflicting constraint, refuses to invent facts, and escalates or requests the smallest clarifying input needed.", "assertions": [ "Identifies the missing or conflicting constraint", "Does not invent unavailable facts", "Requests clarification or escalates with a bounded next step" ] }, { "id": "llamaindex-evidence-handoff", "prompt": "Create a review-ready llamaindex handoff for another practitioner.", "expected_output": "The handoff records assumptions, decisions, artifacts, validation evidence, and unresolved risks so another practitioner can reproduce the result.", "assertions": [ "Records assumptions and decisions", "Links concrete artifacts to validation evidence", "States unresolved risks and reproducible next steps" ] } ] }
-
-
references
-
agent-patterns.md 2.8 KB
# LlamaIndex Agent Patterns ## Agent Types | Agent | Tool Calling | When to Use | |-------|-------------|-------------| | `FunctionAgent` | Native function calling API | Models with tool-calling support | | `ReActAgent` | ReAct prompting pattern | Models without native tool support | ## Pattern 1: AgentWorkflow (Built-in Multi-Agent) ```python from llama_index.core.agent.workflow import AgentWorkflow, FunctionAgent research_agent = FunctionAgent( name="ResearchAgent", description="Searches and records notes", system_prompt="You are a researcher. Hand off to WriteAgent when ready.", tools=[search_web, record_notes], can_handoff_to=["WriteAgent"], ) write_agent = FunctionAgent( name="WriteAgent", description="Writes reports from notes", can_handoff_to=["ReviewAgent"], ) workflow = AgentWorkflow( agents=[research_agent, write_agent], root_agent=research_agent.name, initial_state={"report_content": "Not written yet."}, ) response = await workflow.run(user_msg="Write a report on the history of the web...") ``` ## Pattern 2: Orchestrator Agent (Sub-Agents as Tools) Centralized control: one orchestrator calls sub-agents via tools. ```python orchestrator = FunctionAgent( name="Orchestrator", system_prompt="You orchestrate research, writing, and review...", tools=[call_research_agent, call_write_agent, call_review_agent], ) response = await orchestrator.run(user_msg="Write a report...") ``` ## Pattern 3: Custom Planner (DIY Orchestration) ```python # LLM outputs structured plan in XML/JSON # Python code parses and executes for step in parse_plan(llm_response): agent = agents[step.agent_name] result = await agent.run(user_msg=step.input) ``` ## Streaming Events AgentWorkflow emits five event types: ```python handler = workflow.run(user_msg="Your query") async for event in handler.stream_events(): if isinstance(event, AgentStream): print(event.delta, end="") # Real-time output elif isinstance(event, AgentOutput): print(f"{event.current_agent_name}: {event.response.content}") ``` ## Known Bug: Agent Handoff After handoff, the receiving agent may lose the user's original request because handoff messages push it out of ChatMemory. **Fix**: Extend `FunctionAgent.take_step()`: ```python class MyFunctionAgent(FunctionAgent): async def take_step(self, ctx, llm_input, tools, memory): last_msg = llm_input[-1].content if llm_input else "" if "handoff_result" in last_msg: for message in llm_input[::-1]: if message.role == MessageRole.USER: llm_input.append(message) break return await super().take_step(ctx, llm_input, tools, memory) ``` Also customize `handoff_output_prompt` with a `handoff_result` tag for reliable detection. -
architecture.md 2.2 KB
# LlamaIndex Core Architecture ## Seven Core Primitives LlamaIndex organizes functionality around seven primitives that map to RAG pipeline stages: | Primitive | Purpose | Example | |-----------|---------|---------| | Reader | Pull data from sources | `SimpleDirectoryReader("./data")` | | Document | Source content with metadata | Returned by Reader | | Node | Chunk of a Document | Created by Node Parsers | | Index | Data structure over Nodes | `VectorStoreIndex`, `PropertyGraphIndex` | | Retriever | Return relevant Nodes | `index.as_retriever(similarity_top_k=5)` | | Query Engine | Retriever + response synthesis | `index.as_query_engine()` | | Agent | LLM with tools | `FunctionAgent(tools=[...])` | | Workflow | Event-driven orchestration | `class MyFlow(Workflow)` | ## Configuration The `Settings` object provides global configuration, replacing the deprecated `ServiceContext`: ```python from llama_index.core import Settings from llama_index.llms.openai import OpenAI Settings.llm = OpenAI(model="gpt-4o") Settings.embed_model = "local:BAAI/bge-small-en-v1.5" Settings.text_splitter = SentenceSplitter(chunk_size=1024, chunk_overlap=20) ``` ## Data Flow ``` Source -> Reader -> Document -> Node Parser -> Nodes -> Index | Retriever <- Query Engine <- Agent | Response ``` ## Workflow Event Model Workflows replace DAG-based composition with typed event passing: ```python from llama_index.core.workflow import Workflow, StartEvent, StopEvent, Event, step class RetrievedEvent(Event): query: str nodes: list class SimpleRAG(Workflow): @step async def retrieve(self, ev: StartEvent) -> RetrievedEvent: # ... retrieval logic return RetrievedEvent(query=ev.query, nodes=nodes) @step async def generate(self, ev: RetrievedEvent) -> StopEvent: # ... generation logic return StopEvent(result=str(resp)) ``` Steps infer input/output types from annotations. The framework validates the event graph before execution. -
evaluation-observability.md 2.5 KB
# LlamaIndex Evaluation and Observability ## Built-in Evaluators | Evaluator | What It Measures | Requires Labels? | |-----------|-----------------|-----------------| | `FaithfulnessEvaluator` | Answer faithful to retrieved context | No (LLM-as-judge) | | `RelevancyEvaluator` | Answer relevant to query | No (LLM-as-judge) | | `SemanticSimilarityEvaluator` | Answer matches reference semantically | Yes | | `PairwiseComparisonEvaluator` | Which response is better | No (LLM-as-judge) | ```python from llama_index.core.evaluation import FaithfulnessEvaluator evaluator = FaithfulnessEvaluator() result = evaluator.evaluate_response(response=response) print(f"Faithfulness: {result.score}") ``` ## Batch Evaluation ```python from llama_index.core.evaluation import BatchEvalRunner runner = BatchEvalRunner( {"faithfulness": FaithfulnessEvaluator()}, workers=4, ) results = runner.evaluate_responses(queries, responses) ``` ## RAG Evaluation Metrics Based on the RAG survey (Gao et al., 2023), seven measurement aspects: 1. **Answer Relevance** — Does the answer address the question? 2. **Context Relevance** — Is the retrieved context relevant? 3. **Faithfulness** — Is the answer grounded in context? 4. **Context Recall** — Are all needed chunks retrieved? 5. **Context Precision** — Are irrelevant chunks excluded? 6. **Noise Sensitivity** — Does irrelevant context degrade quality? 7. **Answer Correctness** — Is the factual answer correct? ## OpenTelemetry Tracing ```python from fi_instrumentation import register from traceai_llamaindex import LlamaIndexInstrumentor trace_provider = register(project_name="rag_app") LlamaIndexInstrumentor().instrument(tracer_provider=trace_provider) # Every workflow run now produces trace trees with: # - Root span per run() call # - Child spans for retrieval and LLM calls # - Attributes: latency, model, token counts, tool arguments ``` ## Span-Attached Evaluation ```python from fi.evals import evaluate from fi.evals.otel import enable_auto_enrichment enable_auto_enrichment() # Call once at startup # Inside a workflow step: context = "\n\n".join([n.get_content() for n in ev.nodes]) r = evaluate("groundedness", output=str(resp), context=context) # Score becomes a span attribute on the active span ``` ## Available Observability Integrations | Platform | Package | Type | |----------|---------|------| | FutureAGI traceAI | `traceai-llamaindex` | OTel spans + eval | | OpenLLMetry | `openllmetry` | OTel spans | | LangFuse | `langfuse` | Traces + evals | | Arize AI | `arize` | ML observability | -
evaluation-workflow.md 4.7 KB
# Evaluation Workflow — ParamTuner and Metrics This reference shows how to set up systematic evaluation for a LlamaIndex RAG pipeline, including parameter tuning, evaluator configuration, and batch scoring. ## Setup ```python from llama_index.core.evaluation import ( FaithfulnessEvaluator, RelevancyEvaluator, SemanticSimilarityEvaluator, BatchEvalRunner, ) from llama_index.llms.openai import OpenAI gpt4 = OpenAI(model="gpt-4o") faith_evaluator = FaithfulnessEvaluator(llm=gpt4) rel_evaluator = RelevancyEvaluator(llm=gpt4) sim_evaluator = SemanticSimilarityEvaluator(llm=gpt4) ``` ## Single Query Evaluation ```python response = query_engine.query("What is the rate limit for the API?") faith_result = faith_evaluator.evaluate_response(response=response) print(f"Faithfulness: {faith_result.score} — {faith_result.feedback}") # Example output: Faithfulness: 0.92 — The answer is grounded in the provided context rel_result = rel_evaluator.evaluate_response( response=response, question="What is the rate limit for the API?" ) print(f"Relevancy: {rel_result.score} — {rel_result.feedback}") # Example output: Relevancy: 0.88 — The answer addresses the core question ``` ## Batch Evaluation ```python eval_questions = [ "How do I authenticate?", "What are the rate limits?", "How do I paginate results?", "What error codes exist?", "How do I handle webhooks?", ] # Get responses responses = [query_engine.query(q) for q in eval_questions] # Batch evaluate runner = BatchEvalRunner( { "faithfulness": FaithfulnessEvaluator(), "relevancy": RelevancyEvaluator(), }, workers=4, ) results = runner.evaluate_responses(eval_questions, responses) for metric_name, metric_results in results.items(): scores = [r.score for r in metric_results] print(f"{metric_name}: mean={sum(scores)/len(scores):.2f}, " f"min={min(scores):.2f}, max={max(scores):.2f}") # Example output: # faithfulness: mean=0.91, min=0.78, max=1.00 # relevancy: mean=0.85, min=0.72, max=0.94 ``` ## ParamTuner — Systematic Optimization ```python from llama_index.core import VectorStoreIndex from llama_index.core.param_tuner.base import ParamTuner from llama_index.core.evaluation import SemanticSimilarityEvaluator import numpy as np def build_and_evaluate(params): chunk_size = params["chunk_size"] top_k = params["top_k"] # Build index with these parameters index = VectorStoreIndex.from_documents( documents, transformations=[SentenceSplitter(chunk_size=chunk_size)] ) # Query query_engine = index.as_query_engine(similarity_top_k=top_k) responses = [query_engine.query(q) for q in eval_questions] # Evaluate evaluator = SemanticSimilarityEvaluator(llm=gpt4) scores = [] for i, resp in enumerate(responses): result = evaluator.evaluate_response( response=resp, reference=reference_answers[i] ) scores.append(result.score) return np.mean(scores) param_tuner = ParamTuner( param_fn=build_and_evaluate, param_dict={ "chunk_size": [256, 512, 1024], "top_k": [2, 5, 10], }, fixed_param_dict={ "documents": documents, "eval_questions": eval_questions[:3], }, ) results = param_tuner.tune() best = results.best_run_result print(f"Best: chunk_size={best.params['chunk_size']}, " f"top_k={best.params['top_k']}, score={best.score:.3f}") # Example output: Best: chunk_size=512, top_k=5, score=0.894 ``` ## Seven RAG Measurement Aspects Based on (Gao, Yunfan et al., 2023), evaluate across these dimensions: | Aspect | What it measures | Evaluator | |--------|-----------------|-----------| | Answer Relevance | Does the answer address the question? | RelevancyEvaluator | | Context Relevance | Is the retrieved context on-topic? | RelevancyEvaluator (on context) | | Faithfulness | Is the answer grounded in context? | FaithfulnessEvaluator | | Context Recall | Are all needed chunks retrieved? | Custom (check coverage) | | Context Precision | Are irrelevant chunks excluded? | Custom (check rank order) | | Noise Sensitivity | Does noise degrade answers? | Compare with/without noise | | Answer Correctness | Are facts correct? | SemanticSimilarityEvaluator | ## Evaluation Best Practices 1. **Use held-out queries** — never tune on the same questions you evaluate on 2. **Run evaluation in the same process** — span-attached scoring preserves trace context 3. **Monitor in production** — offline notebook scoring catches known issues; production monitoring catches novel ones 4. **Combine multiple metrics** — faithfulness alone misses relevancy failures and vice versa 5. **Track over time** — regressions are easier to catch when you have a baseline -
example-rag-pipeline.md 4.9 KB
# End-to-End RAG Pipeline — Worked Example This example shows a complete LlamaIndex RAG pipeline from data loading through production evaluation. It demonstrates the expected output depth for the Full pipeline mode. ## Scenario A directory of technical PDF documentation about a REST API. We need to build a production RAG pipeline that answers developer questions with citations. ## Pipeline ### 1. Ingest — Load Documents ```python from llama_index.core import SimpleDirectoryReader, Settings from llama_index.llms.openai import OpenAI from llama_index.embeddings.openai import OpenAIEmbedding Settings.llm = OpenAI(model="gpt-4o-mini") Settings.embed_model = OpenAIEmbedding(model="text-embedding-3-small") documents = SimpleDirectoryReader("./api-docs/").load_data() print(f"Loaded {len(documents)} documents") # Output: Loaded 12 documents ``` ### 2. Chunk — Semantic Splitting ```python from llama_index.core.node_parser import SemanticSplitterNodeParser splitter = SemanticSplitterNodeParser( embed_model=Settings.embed_model, breakpoint_percentile_threshold=95, buffer_size=1, ) nodes = splitter.get_nodes_from_documents(documents) print(f"Created {len(nodes)} chunks") # Output: Created 347 chunks ``` ### 3. Index — Vector + Hybrid ```python from llama_index.core import VectorStoreIndex from llama_index.core.retrievers import BM25Retriever # Build vector index index = VectorStoreIndex(nodes=nodes) # Create retrievers for hybrid search vector_retriever = index.as_retriever(similarity_top_k=10) bm25_retriever = BM25Retriever.from_defaults(index=index, similarity_top_k=10) ``` ### 4. Retrieve — Hybrid + Rerank ```python from llama_index.core.retrievers import QueryFusionRetriever from llama_index.core.postprocessor.cohere_rerank import CohereRerank from llama_index.core.query_engine import RetrieverQueryEngine # Fusion retriever with reciprocal rank fusion hybrid_retriever = QueryFusionRetriever( retrievers=[vector_retriever, bm25_retriever], similarity_top_k=5, num_queries=1, # Use original query only mode="reciprocal_rerank", ) # Reranker reranker = CohereRerank(top_n=5) # Query engine query_engine = RetrieverQueryEngine.from_args( retriever=hybrid_retriever, node_postprocessors=[reranker], ) ``` ### 5. Agent — Query with Multi-Source Routing (Optional) ```python from llama_index.core.agent.workflow import FunctionAgent from llama_index.core.tools import QueryEngineTool rag_tool = QueryEngineTool.from_defaults( query_engine=query_engine, name="api_docs_search", description="Search API documentation for endpoints, parameters, and examples", ) agent = FunctionAgent( name="APIAgent", system_prompt="Answer developer questions about the REST API using documentation.", tools=[rag_tool], ) response = await agent.run(user_msg="How do I authenticate API requests?") # Response includes: authentication methods, required headers, token expiry ``` ### 6. Evaluate — Measure Quality ```python from llama_index.core.evaluation import FaithfulnessEvaluator, RelevancyEvaluator faith = FaithfulnessEvaluator() rel = RelevancyEvaluator() test_queries = [ "What is the rate limiting policy?", "How do I paginate through results?", "What error codes does the API return?", ] for q in test_queries: resp = query_engine.query(q) faith_result = faith.evaluate_response(response=resp) rel_result = rel.evaluate_response(response=resp, question=q) print(f"Q: {q}") print(f" Faithfulness: {faith_result.passing}") print(f" Relevancy: {rel_result.passing}") print(f" Sources: {len(resp.source_nodes)} chunks") ``` ### 7. Deploy — Production ```python from llama_deploy import deploy_workflow, WorkflowServiceConfig, ControlPlaneConfig from llama_index.core.workflow import Workflow, StartEvent, StopEvent, step class RAGWorkflow(Workflow): @step async def answer(self, ev: StartEvent) -> StopEvent: result = query_engine.query(ev.query) return StopEvent(result=str(result)) await deploy_workflow( workflow=RAGWorkflow(timeout=60), workflow_config=WorkflowServiceConfig(service_name="api-rag"), control_plane_config=ControlPlaneConfig(), ) ``` ## Expected Output Depth - **Quick mode:** 2-3 source chunks, no reranker, single evaluator metric - **Full mode (shown above):** 5 source chunks, hybrid retrieval + reranker, multi-metric evaluation, deploy-ready - **Evaluate mode:** ParamTuner over chunk_size=[256,512,1024] with top_k=[2,5,10], reporting best configuration ## Common Pitfalls in This Pipeline - Skipping the reranker produces answers that appear correct but miss nuance - Using SimpleVectorStore instead of a persistent vector store loses all indexed data on restart - Forgetting `await` on the agent call produces no error — just a silent coroutine object - The hybrid retriever returns more candidates than the synthesizer needs without the reranker -
faq-and-troubleshooting.md 2.8 KB
# LlamaIndex FAQ and Troubleshooting ## Installation **Q: Installation fails with dependency conflicts?** A: Use a virtual environment. Install core first: `pip install llama-index-core`, then add integrations: `pip install llama-index-llms-openai llama-index-vector-stores-qdrant`. **Q: Python version requirements?** A: LlamaIndex 0.14+ requires Python 3.9+. ## Common Errors **Q: "Coroutine object was never awaited"** A: All Workflow step methods are async. You forgot `await` before the step call or `w.run()`. Add `await` to the call site. **Q: Agent doesn't respond after handoff** A: Known AgentWorkflow handoff bug. The user's request was pushed out of ChatMemory. Apply the take_step fix from the agent-patterns reference. **Q: llama-deploy deployed but requests time out** A: Redis isn't running or isn't reachable. Start `redis-server` and verify the control plane can connect. Check that the REDIS_URL env var matches the service configuration. **Q: Retrieval quality is poor despite hybrid search** A: Missing reranker. Hybrid retrieval returns more candidates than the synthesizer needs. Add CohereRerank or ColbertRerank as a node postprocessor. **Q: Spans are missing or incomplete in observability UI** A: Instrumentation was called after workflow instantiation. Move `LlamaIndexInstrumentor().instrument()` to before any Workflow subclass instantiation. ## Performance **Q: Indexing is slow for large document sets** A: Enable async mode, increase worker count in IngestionPipeline, and consider batched document loading. For very large sets, use the IngestionPipeline.run() with show_progress=True to monitor. **Q: Memory usage grows unbounded** A: The embedding cache accumulates. Monitor with `cache.get_size()` and clear periodically with `cache.clear()`. Consider a bounded cache implementation. **Q: How many chunks should I use per document?** A: Use the ParamTuner for systematic optimization: search across chunk sizes (256, 512, 1024) and evaluate against your query set. There is no one-size-fits-all answer. ## Design Decisions **Q: Should I use Workflows or Query Pipelines?** A: Workflows. Query Pipelines are deprecated in 0.14+. Use Workflows for any workload more complex than a single-shot query. **Q: Should I use FunctionAgent or ReActAgent?** A: FunctionAgent if your model supports native function calling (GPT-4, Claude 3, Gemini). ReActAgent if it doesn't (smaller open models). **Q: Should I use PropertyGraphIndex or KnowledgeGraphIndex?** A: PropertyGraphIndex. KnowledgeGraphIndex is deprecated. PGI supports labeled nodes, relationship properties, embeddings, Cypher queries, and combined concurrent retrieval strategies. **Q: Should I use Settings or ServiceContext?** A: Settings. ServiceContext is deprecated. Settings provides a single import with global configuration. -
integration-ecosystem.md 2.5 KB
# LlamaIndex Integration Ecosystem ## Vector Stores | Store | Production | Best For | |-------|-----------|----------| | Pinecone | Yes | Managed, high-scale | | Qdrant | Yes | Self-hosted or managed | | Weaviate | Yes | Hybrid search + graph | | Chroma | Yes | Embeddings + metadata | | pgvector | Yes | PostgreSQL native | | Milvus | Yes | Billion-scale | | SimpleVectorStore | No | Dev/test only | ```python from llama_index.vector_stores.qdrant import QdrantVectorStore from qdrant_client import QdrantClient vector_store = QdrantVectorStore( client=QdrantClient(url="http://localhost:6333"), collection_name="my_docs", ) index = VectorStoreIndex.from_documents(documents, vector_store=vector_store) ``` ## LlamaHub Data Connectors 200+ data loaders available: ```python pip install llama-index-readers-notion from llama_index.readers.notion import NotionPageReader documents = NotionPageReader(integration_token="...").load_data() ``` Available connectors include: PDFs, Notion, Confluence, Slack, GitHub, S3, JIRA, SAP, Salesforce, Google Drive, SQL databases, web pages, Discord, YouTube. ## LlamaParse Document Parsing ```python pip install llama-parse from llama_parse import LlamaParse parser = LlamaParse(result_type="markdown", parsing_instruction="Extract tables and preserve layout.") documents = parser.load_data("./complex_document.pdf") ``` Key capabilities: LLM-powered parsing, JSON output mode, multi-model support (GPT-4.1, Gemini 2.5 Pro), auto skew correction, MCP integration. ## LangChain Interoperability ```python from llama_index.core.langchain_helpers.agents import ( IndexToolConfig, LlamaIndexTool ) tool_config = IndexToolConfig( query_engine=query_engine, name="vector_index", description="Useful for answering queries about documents", ) rag_tool = LlamaIndexTool.from_tool_config(tool_config) # Use rag_tool with any LangChain agent ``` ## Framework Comparison | Framework | Lead With | Best For | License | |-----------|----------|----------|---------| | **LlamaIndex** | Data ingestion + retrieval | RAG-heavy apps, heterogeneous sources | MIT | | **LangChain + LangGraph** | Chain/graph abstractions | Multi-agent state machines, checkpoints | MIT | | **Haystack** | Pipeline composition | Production NLP search pipelines | Apache 2.0 | | **DSPy** | Compiled prompt programs | Optimization-driven prompt programming | MIT | LlamaIndex excels when you have multiple data sources, multiple indexes, hybrid retrieval, and metadata filtering. For single-source, single-strategy RAG, the abstraction may not earn its weight. -
production-deployment.md 2.3 KB
# LlamaIndex Production Deployment ## llama-deploy: Distributed Runtime ```python import asyncio from llama_deploy import deploy_workflow, WorkflowServiceConfig, ControlPlaneConfig async def main(): await deploy_workflow( workflow=SimpleRAG(timeout=60), workflow_config=WorkflowServiceConfig(service_name="simple_rag"), control_plane_config=ControlPlaneConfig(), ) ``` Requirements: - Redis (or compatible message queue) for the control plane - Worker processes registering with the control plane - API gateway for HTTP routing ## Debugging ### Debug Logging ```python import logging, sys logging.basicConfig(stream=sys.stdout, level=logging.DEBUG) ``` ### Callback Handler ```python from llama_index.core import set_global_handler set_global_handler("simple") # Prints event trace ``` ### OpenTelemetry Tracing ```python pip install traceai-llamaindex ``` ```python from fi_instrumentation import register from traceai_llamaindex import LlamaIndexInstrumentor trace_provider = register(project_name="rag_app") LlamaIndexInstrumentor().instrument(tracer_provider=trace_provider) ``` ## Common Production Failures | Failure | Symptom | Root Cause | Fix | |---------|---------|------------|-----| | Retrieval miss | Poor answer quality | No reranker on hybrid retrieval | Add Cohere/ColBERT reranker | | Cross-tenant leak | Wrong data returned | Missing metadata filters | Wire filters at retriever level | | Async coroutine bug | Workflow doesn't run | Forgot `await` | Ensure all steps are awaited | | llama-deploy silent failure | No response | No Redis running | Start redis-server first | | Incomplete traces | Missing spans | Instrumentation called too late | Call instrument() before instantiation | | Handoff failure | Agent doesn't respond | ChatMemory overflow | Extend FunctionAgent.take_step() | ## Production Configuration Checklist - [ ] Use a managed vector store (Pinecone, Qdrant, pgvector) — not SimpleVectorStore - [ ] Apply SemanticSplitterNodeParser for adaptive chunking - [ ] Add reranker (Cohere, Jina, ColBERT) on all hybrid retrieval - [ ] Wire metadata filters at retriever level for multi-tenant isolation - [ ] Enable async mode for concurrent operations - [ ] Instrument observability before workflow instantiation - [ ] Configure Redis for llama-deploy - [ ] Set up span-attached evaluation for continuous quality monitoring -
property-graph-index.md 2.9 KB
# LlamaIndex PropertyGraphIndex ## Overview The `PropertyGraphIndex` replaces the deprecated `KnowledgeGraphIndex`. It uses labeled property graphs with typed nodes, relationship properties, embedding support, and Cypher query capability. ```python from llama_index.core import PropertyGraphIndex index = PropertyGraphIndex.from_documents(documents) retriever = index.as_retriever() query_engine = index.as_query_engine() ``` ## Graph Construction Extractors | Extractor | Approach | When to Use | |-----------|----------|-------------| | `SimpleLLMPathExtractor` | LLM extracts triples (entity, relation, entity) | Free-form exploration | | `ImplicitPathExtractor` | Uses document structure metadata | No LLM needed | | `DynamicLLMPathExtractor` | LLM with allowed type hints | Semi-guided extraction | | `SchemaLLMPathExtractor` | Strict Pydantic schema validation | Production: guarantees type consistency | ### Schema-Guided Extraction ```python from typing import Literal from llama_index.core.indices.property_graph import SchemaLLMPathExtractor entities = Literal["PERSON", "PLACE", "THING"] relations = Literal["PART_OF", "HAS", "IS_A"] schema = { "PERSON": ["PART_OF", "HAS", "IS_A"], "PLACE": ["PART_OF", "HAS"], "THING": ["IS_A"], } extractor = SchemaLLMPathExtractor( possible_entities=entities, possible_relations=relations, kg_validation_schema=schema, strict=True, max_triplets_per_chunk=10, ) index = PropertyGraphIndex.from_documents(documents, kg_extractors=[extractor]) ``` ## Retrieval Strategies (Can Be Combined) | Retriever | How It Works | |-----------|-------------| | `LLMSynonymRetriever` | LLM generates keywords/synonyms, finds matching nodes | | `VectorContextRetriever` | Embedding similarity on graph nodes | | `TextToCypherRetriever` | LLM generates Cypher from schema + query | | `CypherTemplateRetriever` | Template with LLM-inferred params | | `CustomPGRetriever` | Subclass for custom traversal | ### Combined Hybrid Retrieval ```python from llama_index.core.indices.property_graph import ( VectorContextRetriever, LLMSynonymRetriever, PGRetriever ) retriever = PGRetriever(sub_retrievers=[ VectorContextRetriever(index.property_graph_store), LLMSynonymRetriever(index.property_graph_store), ]) nodes = retriever.retrieve("query") ``` ## Backing Stores | Store | Embedding Support | Best For | |-------|-----------------|----------| | `SimplePropertyGraphStore` | No (use external) | Development | | `Neo4jPropertyGraphStore` | Yes (native) | Production | | `FalkorDBPropertyGraphStore` | Yes | High-performance graph | | `TiDBPropertyGraphStore` | Yes | SQL + graph hybrid | ## Persistence ```python index.storage_context.persist("./storage") # Load from llama_index.core import StorageContext, load_index_from_storage index = load_index_from_storage( StorageContext.from_defaults(persist_dir="./storage") ) ``` -
rag-strategies.md 3.2 KB
# LlamaIndex RAG Strategies ## Basic RAG ```python from llama_index.core import VectorStoreIndex, SimpleDirectoryReader documents = SimpleDirectoryReader("./data").load_data() index = VectorStoreIndex.from_documents(documents) query_engine = index.as_query_engine() response = query_engine.query("Your question here") ``` ## Chunking Strategies | Parser | Best For | Configuration | |--------|----------|---------------| | `SentenceSplitter` | General prose | `chunk_size=1024, chunk_overlap=20` | | `SemanticSplitterNodeParser` | Coherent semantic units | `breakpoint_percentile_threshold=95` | | `HierarchicalNodeParser` | Large documents | `chunk_sizes=[2048, 512, 128]` | | `SentenceWindowNodeParser` | Embedding precision + synthesis context | `window_size=3` | ## Decoupling Retrieval from Synthesis ```python from llama_index.core.node_parser import SentenceWindowNodeParser from llama_index.core.postprocessor import MetadataReplacementNodePostProcessor node_parser = SentenceWindowNodeParser.from_defaults( window_size=3, window_metadata_key="window", original_text_metadata_key="original_sentence", ) postprocessor = MetadataReplacementNodePostProcessor( target_metadata_key="window" ) ``` ## Hybrid Retrieval ```python from llama_index.core.retrievers import VectorIndexRetriever from llama_index.core import QueryBundle # Hybrid via supported vector stores (Weaviate, Qdrant, Pinecone) vector_retriever = index.as_retriever(similarity_top_k=10) # Add BM25 retriever for keyword matching from llama_index.core.retrievers import BM25Retriever bm25_retriever = BM25Retriever.from_defaults( index=index, similarity_top_k=10 ) ``` ## Reranking Always apply a reranker on hybrid retrieval output: ```python from llama_index.core.postprocessor.cohere_rerank import CohereRerank rerank = CohereRerank(top_n=5) query_engine = index.as_query_engine( similarity_top_k=10, node_postprocessors=[rerank], ) ``` ## Metadata Filters ```python from llama_index.core.vector_stores import MetadataFilters, ExactMatchFilter filters = MetadataFilters( filters=[ExactMatchFilter(key="tenant_id", value="acme_corp")] ) query_engine = index.as_query_engine(filters=filters) ``` ## Recursive Retrieval for Large Corpora Two-level retrieval: document summaries -> chunks. ```python from llama_index.core.retrievers import RecursiveRetriever retriever_chunk = RecursiveRetriever( "vector", retriever_dict={"vector": vector_retriever_chunk}, node_dict=all_nodes_dict, ) ``` ## RouterQueryEngine Route queries to different strategies based on intent: ```python from llama_index.core.query_engine import RouterQueryEngine from llama_index.core.selectors import PydanticSingleSelector query_engine = RouterQueryEngine( selector=PydanticSingleSelector.from_defaults(), query_engine_tools=[summary_tool, vector_tool], ) ``` ## Chunk Size Optimization with ParamTuner ```python from llama_index.core.param_tuner.base import ParamTuner param_tuner = ParamTuner( param_fn=objective_function, param_dict={"chunk_size": [256, 512, 1024]}, fixed_param_dict={"top_k": 2}, ) results = param_tuner.tune() best_chunk_size = results.best_run_result.params["chunk_size"] ``` -
workflows.md 2.7 KB
# LlamaIndex Workflows ## Core Model Workflows are event-driven step-based orchestration. Steps consume typed events and emit typed events. ```python from llama_index.core.workflow import Workflow, StartEvent, StopEvent, Event, step class MyWorkflow(Workflow): @step async def step_one(self, ev: StartEvent) -> MyEvent: # Do work return MyEvent(result=processed_data) @step async def step_two(self, ev: MyEvent) -> StopEvent: # Final step return StopEvent(result=ev.result) ``` ## Running a Workflow ```python w = MyWorkflow(timeout=60, verbose=False) result = await w.run(input_data="your data") ``` ## Control Flow Patterns | Pattern | Implementation | |---------|---------------| | **Sequential** | Step A >> Event >> Step B | | **Branching** | `if condition: return EventX else: return EventY` | | **Looping** | Step returns event handled by an earlier step | | **Parallel fan-out** | `return list[Event]` | | **Parallel fan-in** | Accept `list[Event]` | | **Dynamic emission** | `ctx.send_event(ev)` | | **Dynamic collection** | `ctx.collect_events(...)` | ## State Management ```python async with ctx.store.edit_state() as state: state["counter"] = state.get("counter", 0) + 1 ``` ## Durable Workflows (Checkpoint/Resume) ```python # Checkpoint on step completion async for ev in handler.stream_events(expose_internal=True): if isinstance(ev, StepStateChanged) and ev.step_state == StepState.NOT_RUNNING: db.save("my-run", json.dumps(handler.ctx.to_dict())) # Resume after crash ctx = Context.from_dict(w, json.loads(db.load("my-run"))) result = await w.run(ctx=ctx) ``` Key properties: - At-least-once semantics (in-flight steps may re-run) - Step side effects must be idempotent - Non-serializable objects (API clients, DB connections) go in `Resource` factories ## Resource Injection ```python from typing import Annotated from llama_index.core.workflow import Resource def get_client() -> MyApiClient: return MyApiClient() class MyWorkflow(Workflow): @step async def process(self, ev: StartEvent, client: Annotated[MyApiClient, Resource(get_client)]) -> StopEvent: result = await client.do_work(ev.data) return StopEvent(result=result) ``` ## Validation ```python workflow.validate() # Check event graph before running workflow.validate(validate_resources=True) # Also resolves Resource factories ``` For intentionally dynamic steps: `@step(skip_graph_checks=["reachability"])` ## Migrating from Query Pipelines Query Pipelines are deprecated in 0.14+. Migrate patterns: - **Loops** in a DAG = a step returning an event consumed by an earlier step - **Branches** in a DAG = `if/else` returning different event types - **Data passing** in a DAG = typed event fields
-
-
scripts
-
check-setup.py 2.3 KB
#!/usr/bin/env python3 """ Verify that LlamaIndex is installed and can create a basic index. Run this script to check your setup before building applications. """ import sys import importlib REQUIRED_PACKAGES = [ "llama_index", "llama_index.core", ] OPTIONAL_PACKAGES = [ "llama_index.llms.openai", "llama_index.vector_stores.qdrant", "llama_index.embeddings.openai", "llama_parse", "llama_deploy", ] def check_package(name: str, required: bool = True) -> bool: """Check if a package is importable.""" try: importlib.import_module(name) print(f" [OK] {name}") return True except ImportError: status = "REQUIRED" if required else "optional" print(f" [MISSING] {name} ({status})") return not required def main(): print("LlamaIndex Setup Check") print("=" * 40) # Python version print(f"\nPython: {sys.version}") if sys.version_info < (3, 9): print(" [FAIL] Python 3.9+ required") sys.exit(1) # Core packages print("\nCore:") all_ok = all(check_package(pkg, required=True) for pkg in REQUIRED_PACKAGES) # Optional packages print("\nOptional:") for pkg in OPTIONAL_PACKAGES: check_package(pkg, required=False) # Test basic functionality print("\nFunctionality test:") try: from llama_index.core import Document from llama_index.core.node_parser import SentenceSplitter doc = Document(text="Test document for LlamaIndex setup verification.") parser = SentenceSplitter(chunk_size=256, chunk_overlap=20) nodes = parser.get_nodes_from_documents([doc]) print(f" [OK] Document chunking works ({len(nodes)} nodes)") # Test Settings from llama_index.core import Settings default_chunk = Settings.text_splitter print(f" [OK] Settings accessible") except Exception as e: print(f" [FAIL] Basic functionality test failed: {e}") all_ok = False print("\n" + "=" * 40) if all_ok: print("Setup check: ALL REQUIRED PACKAGES OK") else: print("Setup check: SOME REQUIRED PACKAGES MISSING") print("Install missing packages with:") print(" pip install llama-index-core") print(" pip install llama-index-llms-openai") sys.exit(1) if __name__ == "__main__": main()
-
-
templates
-
agentic-rag.py 1.9 KB
#!/usr/bin/env python3 """ Multi-source RAG with agent orchestration. Demonstrates pattern 2: orchestrator agent with sub-agents as tools. """ import asyncio from llama_index.core import VectorStoreIndex, SimpleDirectoryReader from llama_index.llms.openai import OpenAI from llama_index.core.agent.workflow import AgentWorkflow, FunctionAgent from llama_index.core.tools import QueryEngineTool from llama_index.core import Settings Settings.llm = OpenAI(model="gpt-4o") # --- Build query engines for different data sources --- product_docs = SimpleDirectoryReader("./data/products").load_data() product_index = VectorStoreIndex.from_documents(product_docs) product_engine = product_index.as_query_engine(similarity_top_k=3) support_tickets = SimpleDirectoryReader("./data/support").load_data() support_index = VectorStoreIndex.from_documents(support_tickets) support_engine = support_index.as_query_engine(similarity_top_k=3) # --- Wrap as tools --- product_tool = QueryEngineTool.from_defaults( query_engine=product_engine, name="product_search", description="Search product documentation for specifications and features.", ) support_tool = QueryEngineTool.from_defaults( query_engine=support_engine, name="support_search", description="Search support tickets for known issues and solutions.", ) # --- Agent with tools --- agent = FunctionAgent( name="SupportAgent", system_prompt=( "You are a technical support agent. Use product_search for product info " "and support_search for known issues. Answer concisely based on the data." ), tools=[product_tool, support_tool], ) async def main() -> None: """Run the agent from a regular Python script.""" response = await agent.run( user_msg="What are the known issues with the API rate limiting feature?" ) print(f"Answer: {response}") if __name__ == "__main__": asyncio.run(main()) -
basic-rag.py 829 B
#!/usr/bin/env python3 """ Minimal RAG pipeline using LlamaIndex. Loads documents from a directory, builds a vector index, and answers queries. """ from llama_index.core import VectorStoreIndex, SimpleDirectoryReader from llama_index.llms.openai import OpenAI from llama_index.core import Settings # --- Configuration --- Settings.llm = OpenAI(model="gpt-4o-mini") DATA_DIR = "./data" # --- Load --- documents = SimpleDirectoryReader(DATA_DIR).load_data() # --- Index --- index = VectorStoreIndex.from_documents(documents) # --- Query --- query_engine = index.as_query_engine( similarity_top_k=5, ) response = query_engine.query("What does this data say about your question?") print(f"Answer: {response}") # Show sources for source in response.source_nodes: print(f" [{source.score:.3f}] {source.text[:100]}...") -
custom-workflow.py 3.2 KB
#!/usr/bin/env python3 """ Custom event-driven workflow with typed events. Shows branching, looping, and state management patterns. """ from llama_index.core.workflow import Workflow, StartEvent, StopEvent, Event, step from llama_index.llms.openai import OpenAI from llama_index.core import Settings from pydantic import Field Settings.llm = OpenAI(model="gpt-4o-mini") # --- Custom Events --- class RetrievedEvent(Event): """Documents have been retrieved.""" query: str documents: list = Field(default_factory=list) class EvaluatedEvent(Event): """Retrieved documents have been evaluated for relevance.""" query: str is_relevant: bool = False documents: list = Field(default_factory=list) class ResearchedEvent(Event): """Additional research has been performed.""" query: str findings: str = "" # --- Workflow --- class SmartRAG(Workflow): llm = OpenAI(model="gpt-4o-mini") @step async def retrieve(self, ev: StartEvent) -> RetrievedEvent: """Initial retrieval.""" from llama_index.core import VectorStoreIndex, SimpleDirectoryReader documents = SimpleDirectoryReader("./data").load_data() index = VectorStoreIndex.from_documents(documents) retriever = index.as_retriever(similarity_top_k=4) nodes = retriever.retrieve(ev.query) return RetrievedEvent(query=ev.query, documents=nodes) @step async def evaluate(self, ev: RetrievedEvent) -> EvaluatedEvent | ResearchedEvent: """Evaluate if retrieved docs are sufficient. If not, research more.""" context = "\n\n".join([n.get_content()[:200] for n in ev.documents]) prompt = f"Query: {ev.query}\nContext: {context}\nIs this context sufficient? Answer YES or NO." resp = await self.llm.acomplete(prompt) if "YES" in str(resp).upper(): return EvaluatedEvent( query=ev.query, is_relevant=True, documents=ev.documents ) else: # Research branch — loops back to evaluate after search_prompt = f"Research this topic: {ev.query}. Provide 3 key facts." research = await self.llm.acomplete(search_prompt) return ResearchedEvent(query=ev.query, findings=str(research)) @step async def research(self, ev: ResearchedEvent) -> RetrievedEvent: """Generate synthetic context when retrieval was insufficient.""" from llama_index.core.schema import TextNode extra_node = TextNode(text=ev.findings) return RetrievedEvent( query=ev.query, documents=[extra_node], # Loop back to evaluate ) @step async def synthesize(self, ev: EvaluatedEvent) -> StopEvent: """Final answer synthesis.""" context = "\n\n".join([n.get_content() for n in ev.documents]) prompt = f"Answer the query using the context.\n\nContext:\n{context}\n\nQuery: {ev.query}" resp = await self.llm.acomplete(prompt) return StopEvent(result=str(resp)) # --- Run --- async def main(): wf = SmartRAG(timeout=60, verbose=True) result = await wf.run(query="What are the key findings in this document?") print(f"Result: {result}") if __name__ == "__main__": import asyncio asyncio.run(main()) -
production-deploy.py 930 B
#!/usr/bin/env python3 """ Deploy a LlamaIndex Workflow as a production microservice using llama-deploy. Requires: `pip install llama-deploy`, running Redis instance. """ import asyncio from llama_deploy import ( deploy_workflow, WorkflowServiceConfig, ControlPlaneConfig, ) from llama_index.core.workflow import Workflow, StartEvent, StopEvent, step class MyRAGWorkflow(Workflow): """Your workflow class — replace with your actual Workflow.""" @step async def process(self, ev: StartEvent) -> StopEvent: # Your workflow logic here return StopEvent(result=f"Processed: {ev.query}") async def main(): await deploy_workflow( workflow=MyRAGWorkflow(timeout=60), workflow_config=WorkflowServiceConfig( service_name="my-rag-service", ), control_plane_config=ControlPlaneConfig(), ) if __name__ == "__main__": asyncio.run(main())
-
-
README.md 1.7 KB
# LlamaIndex — RAG & Agent Orchestration Framework An expert-level skill for building LLM applications over your data with LlamaIndex. RAG pipelines, multi-agent orchestration, event-driven workflows, knowledge graph construction, and production deployment. ## Why Install This Skill When your agent loads this skill, it becomes a **LlamaIndex expert** who can: - **Build production RAG pipelines** — from data ingestion to deployed query engines - **Create agent workflows** — AgentWorkflow for tool-using agents - **Construct knowledge graphs** — PropertyGraphIndex for structural path traversal - **Optimize retrieval** — hybrid search, reranking, metadata filters, sentence window parsing - **Add observability** — OpenTelemetry-native tracing with Phoenix - **Evaluate systematically** — span-attached evaluation with ParamTuner ## What You Get | Directory | Purpose | |-----------|---------| | `SKILL.md` | 9-phase pipeline guide, pipeline modes, quick reference | | `references/` | Ingest, chunk, index, retrieve, agent, workflow, deploy, evaluate — one per phase plus framework comparisons | ## Framework Comparison LlamaIndex evolved from a RAG indexing library into a full workflow framework. It differs from LangChain (broader integration ecosystem) and Haystack (declarative DAG pipelines) — LlamaIndex's unique strength is its data-aware indexing and knowledge graph construction. ## Requirements Python 3.8+ with `llama_index` package. ## Quick Start Start with the setup and first workflow in SKILL.md, then use the linked resources for the specific task you need to complete. ## Triggers Use this skill for the task types and keywords described in its SKILL.md description. -
SKILL.md 11.8 KB
--- name: llamaindex description: >- Build LLM applications with the LlamaIndex framework. Use when working with LlamaIndex or comparing RAG and agent orchestration frameworks. Do not use this skill for unrelated requests; route to the nearest named specialist. license: MIT metadata: author: Magnus Hedemark version: 1.3.0 source: https://github.com/run-llama/llama_index --- # LlamaIndex Expert Skill LlamaIndex is an MIT-licensed Python framework for building LLM applications over your data. In 2026, it has evolved from a RAG indexing library into an event-driven workflow framework with integrated production runtime (llama-deploy), agent orchestration (AgentWorkflow), knowledge graph construction (PropertyGraphIndex), and OpenTelemetry-native observability. The framework is organized around seven core primitives: **Reader** (data loaders), **Document/Node** (chunked content model), **Index** (data structures over Nodes), **Retriever** (relevant Node selection), **Query Engine** (retriever + synthesis), **Agent** (LLM with tools), and **Workflow** (event-driven orchestration). ## Key Principles > These principles govern every decision when building with LlamaIndex. Read them before proceeding to the reference guides. 1. **Decouple retrieval chunks from synthesis chunks.** The embedding representation that retrieves well differs from the context representation that generates well. Use `SentenceWindowNodeParser` + `MetadataReplacementNodePostProcessor` for this pattern. 2. **Rerank before you generate.** Hybrid retrieval + reranker is the minimum viable production RAG configuration. 3. **Agents are Workflows.** `FunctionAgent` and `AgentWorkflow` are pre-configured Workflows. Drop to raw `Workflow` when you need custom control flow. 4. **Graphs are not just vector stores.** `PropertyGraphIndex` adds structural path traversal that vector similarity cannot provide — combine both for maximum retrieval quality. 5. **Evaluate in the same process.** Span-attached evaluation preserves the connection between the output and the retrieval context that produced it. ## Where to Start The pipeline has 9 phases from Ingest to Deploy. If you're joining mid-stream with existing work, find your entry point: | You already have... | Start at phase | What to do | |---|---|---| | Nothing — blank project | **Ingest** | Set up data loading, then proceed through the full pipeline | | Documents in a directory | **Chunk** | Choose a chunking strategy, build your index | | A working vector index | **Retrieve** | Add hybrid search, reranking, metadata filters | | An existing RAG pipeline to harden | **Deploy** | Add observability, llama-deploy, production debugging | | A need to measure and improve quality | **Evaluate** | Set up evaluators, ParamTuner, span-attached scoring | | Nothing — comparing frameworks | See Framework Routing Guide | Don't start the pipeline — pick the right tool first | ## Pipeline Mode Different tasks need different levels of rigor. Match your scope to a mode: | Mode | When | Phases to run | Skip | |------|------|---------------|------| | **Quick** | Single query, one source, exploration | Ingest → Chunk → Index → Retrieve | Reranking, metadata filters, observability, evaluation | | **Full** | Production RAG, multiple sources, compliance | Ingest → Chunk → Index → Retrieve → Agent/Workflow → Deploy → Evaluate | Nothing — run all phases | | **Evaluate** | Benchmarking, regression testing | Ingest → Chunk → Index → Evaluate | Retrieve, Agent, Workflow, Deploy (run offline) | | **Graph** | Knowledge graph construction | Ingest → Chunk → Graph → Retrieve | Agent, Workflow, Deploy (query via graph index directly) | Rule of thumb: if you're shipping to users, run Full mode. If you're exploring, run Quick. If you're measuring, run Evaluate. ## Quick Reference | Phase | Task | Approach | Reference | |-------|------|----------|-----------| | Ingest | Load data | `SimpleDirectoryReader("./data").load_data()` | `references/architecture.md` | | Chunk | Parse documents into nodes | `SentenceSplitter(chunk_size=1024)` | `references/rag-strategies.md` | | Index | Build vector index | `VectorStoreIndex.from_documents(docs)` | `references/rag-strategies.md` | | Retrieve | Hybrid search + rerank | `BM25Retriever` + `CohereRerank` | `references/rag-strategies.md` | | Agent | Multi-agent orchestration | `AgentWorkflow(agents=[...])` | `references/agent-patterns.md` | | Workflow | Event-driven pipeline | `class MyFlow(Workflow): @step` | `references/workflows.md` | | Graph | Knowledge graph | `PropertyGraphIndex.from_documents(docs)` | `references/property-graph-index.md` | | Evaluate | RAG evaluation | `FaithfulnessEvaluator().evaluate_response(...)` | `references/evaluation-observability.md` | | Deploy | Production deployment | `deploy_workflow(workflow=MyFlow())` | `references/production-deployment.md` | ## When to Use This Skill Load this skill any time you are: - Building a RAG pipeline over enterprise or personal data - Comparing LlamaIndex with LangChain, Haystack, or DSPy - Designing multi-agent systems with handoff between specialist agents - Deploying an LLM application to production with observability - Constructing knowledge graphs from unstructured documents - Debugging common LlamaIndex failures (retrieval miss, handoff bug, async issues) ## When NOT to Use LlamaIndex — Framework Routing Guide This skill is part of a portfolio of framework skills. When deciding which framework fits, use this routing table: | Scenario | Reach for | Why | |----------|-----------|-----| | I have documents I need to query | **LlamaIndex** | Data ingestion, hybrid retrieval, reranking, and knowledge graphs are first-class primitives | | I have agents I need to orchestrate | **LangGraph** | State-machine semantics, time-travel debugging, and human-in-the-loop pauses are the core design | | I have a tool I need to wrap as an agent | **PydanticAI** | Type-safe agent definitions with dependency injection, minimal abstraction over LLM calls | | Data-heavy RAG over PDFs, SQL, Slack, 200+ sources | **LlamaIndex** | LlamaHub connectors, LlamaParse for documents, hybrid retrieval out of the box | | Complex multi-agent state machines with checkpoints | **LangGraph** | Graph topology control — supervisor, subgraphs, hierarchical teams, built-in checkpointer | | Agent-centric app where type safety matters more than data pipelines | **PydanticAI** | Agents as Pydantic models, DI, structured outputs — the data layer is your code | | Document parsing quality matters (tables, charts, handwriting) | **LlamaIndex** | LlamaParse is purpose-built for this | | Production NLP search pipelines | **Haystack** | Pipeline composition model is more mature for search-specific workloads | | Optimization-driven prompt programming | **DSPy** | Compiled prompt programs, not retrieval pipelines | ## Reference Files | Reference | Load when | File | |-----------|-----------|------| | Core Architecture | Understanding the 7 primitives, Settings, data flow | `references/architecture.md` | | RAG Strategies | Building RAG pipelines from basic to advanced | `references/rag-strategies.md` | | Agent Patterns | Multi-agent orchestration with AgentWorkflow | `references/agent-patterns.md` | | Workflows | Event-driven step composition and durable execution | `references/workflows.md` | | Production & Deployment | llama-deploy, debugging, failure modes | `references/production-deployment.md` | | Property Graph Index | Knowledge graph construction and hybrid retrieval | `references/property-graph-index.md` | | Evaluation & Observability | Metrics, tracing, span-attached scoring | `references/evaluation-observability.md` | | Integration Ecosystem | Vector stores, LlamaHub, LlamaParse, ecosystem | `references/integration-ecosystem.md` | | FAQ & Troubleshooting | Common errors and their fixes | `references/faq-and-troubleshooting.md` | | Worked RAG Example | Complete end-to-end pipeline from ingest to deploy | `references/example-rag-pipeline.md` | | Evaluation Workflow | ParamTuner, evaluators, batch scoring, best practices | `references/evaluation-workflow.md` | ## Template Files | Template | When to use | File | |----------|-------------|------| | Basic RAG | Single-source query, getting started | `templates/basic-rag.py` | | Agentic RAG | Multi-source data with agent routing | `templates/agentic-rag.py` | | Custom Workflow | Custom control flow, branching logic | `templates/custom-workflow.py` | | Production Deploy | Wrapping a workflow as a microservice | `templates/production-deploy.py` | ## Scripts | Script | Purpose | File | |--------|---------|------| | check-setup | Verify LlamaIndex installation and configuration | `scripts/check-setup.py` | ## Troubleshooting — Structured Recovery Guide When something goes wrong, find your symptom and follow the recovery path: ### Retrieval & Answer Quality | Symptom | Likely cause | Immediate fix | Permanent fix | Reference | |---------|-------------|---------------|---------------|-----------| | Answers are poor or hallucinated | No reranker on hybrid retrieval | Add `CohereRerank(top_n=5)` as `node_postprocessor` | Reranking is mandatory for any production RAG | `references/rag-strategies.md` | | Retrieval misses obvious content | Default chunking breaks semantics | Switch to `SemanticSplitterNodeParser(breakpoint_percentile_threshold=95)` | Tune chunk size with ParamTuner | `references/rag-strategies.md` | | Wrong tenant's data returned | Missing metadata filters | Add `MetadataFilters(filters=[ExactMatchFilter(key="tenant_id", ...)])` | Always wire metadata filters at retriever level | `references/rag-strategies.md` | | Only one type of query works well | Single retrieval strategy | Combine BM25 + vector via hybrid retriever | Add RouterQueryEngine for query-type routing | `references/rag-strategies.md` | ### Agent & Workflow Failures | Symptom | Likely cause | Immediate fix | Permanent fix | Reference | |---------|-------------|---------------|---------------|-----------| | Agent waits silently after handoff | AgentWorkflow handoff bug | Extend `FunctionAgent.take_step` to re-locate last user message | Apply the handoff fix on all production agents | `references/agent-patterns.md` | | Workflow doesn't run | Forgot `await` | Add `await` before `w.run(...)` and all step calls | All step methods are async coroutines | `references/workflows.md` | | Step executes but result is lost | State not persisted | Use `ctx.store.edit_state()` for shared state | Only `ctx.store` survives across steps | `references/workflows.md` | | Crash loses all progress | No checkpoint snapshots | Add `Context.to_dict()` save on step completion | Durable workflows need explicit checkpointing | `references/workflows.md` | ### Deployment & Observability | Symptom | Likely cause | Immediate fix | Permanent fix | Reference | |---------|-------------|---------------|---------------|-----------| | llama-deploy deployed but requests time out | Redis not running | Start `redis-server` | Redis is mandatory — control plane won't route without it | `references/production-deployment.md` | | Spans missing in observability UI | Instrumentation called too late | Move `instrument()` call before workflow instantiation | Always instrument before creating any Workflow object | `references/production-deployment.md` | | Wrong data returned (cross-tenant) | Missing metadata filters | Add tenant filter to all retrievers | Filter at retriever level, not in post-processing | `references/production-deployment.md` | ### Recovery Workflow For any failure, follow this cycle: 1. **Identify the symptom** from the tables above 2. **Apply the immediate fix** — this gets you running 3. **Implement the permanent fix** — this prevents recurrence 4. **Verify with evaluation** — run `FaithfulnessEvaluator` on a held-out query set 5. **Document the fix** — add the root cause to `references/faq-and-troubleshooting.md`
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.