Claude Skill

llamaindex

Build LLM applications with the LlamaIndex framework. Use when working with LlamaIndex or comparing RAG and agent orchestration frameworks. Do not use this skill for unrelated requests; route to the nearest named specialist.

LLM Mart · 0 points · 13 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download magnus919-agent-skills-llamaindex-addad86.zip · 26 KB
Part of magnus919/agent-skills — 145 skills

Install

skills CLI npx skills add https://github.com/magnus919/agent-skills/tree/main/llamaindex
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install magnus919-agent-skills@llmmart
Git git clone https://github.com/magnus919/agent-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole magnus919/agent-skills collection as a plugin from our marketplace. Git is the plain clone.

README

LlamaIndex — RAG & Agent Orchestration Framework

An expert-level skill for building LLM applications over your data with LlamaIndex. RAG pipelines, multi-agent orchestration, event-driven workflows, knowledge graph construction, and production deployment.

Why Install This Skill

When your agent loads this skill, it becomes a LlamaIndex expert who can:

  • Build production RAG pipelines — from data ingestion to deployed query engines
  • Create agent workflows — AgentWorkflow for tool-using agents
  • Construct knowledge graphs — PropertyGraphIndex for structural path traversal
  • Optimize retrieval — hybrid search, reranking, metadata filters, sentence window parsing
  • Add observability — OpenTelemetry-native tracing with Phoenix
  • Evaluate systematically — span-attached evaluation with ParamTuner

What You Get

Directory Purpose
SKILL.md 9-phase pipeline guide, pipeline modes, quick reference
references/ Ingest, chunk, index, retrieve, agent, workflow, deploy, evaluate — one per phase plus framework comparisons

Framework Comparison

LlamaIndex evolved from a RAG indexing library into a full workflow framework. It differs from LangChain (broader integration ecosystem) and Haystack (declarative DAG pipelines) — LlamaIndex's unique strength is its data-aware indexing and knowledge graph construction.

Requirements

Python 3.8+ with llama_index package.

Quick Start

Start with the setup and first workflow in SKILL.md, then use the linked resources for the specific task you need to complete.

Triggers

Use this skill for the task types and keywords described in its SKILL.md description.

Skill manifest

LlamaIndex Expert Skill

LlamaIndex is an MIT-licensed Python framework for building LLM applications over your data. In 2026, it has evolved from a RAG indexing library into an event-driven workflow framework with integrated production runtime (llama-deploy), agent orchestration (AgentWorkflow), knowledge graph construction (PropertyGraphIndex), and OpenTelemetry-native observability.

The framework is organized around seven core primitives: Reader (data loaders), Document/Node (chunked content model), Index (data structures over Nodes), Retriever (relevant Node selection), Query Engine (retriever + synthesis), Agent (LLM with tools), and Workflow (event-driven orchestration).

Key Principles

These principles govern every decision when building with LlamaIndex. Read them before proceeding to the reference guides.

  1. Decouple retrieval chunks from synthesis chunks. The embedding representation that retrieves well differs from the context representation that generates well. Use SentenceWindowNodeParser + MetadataReplacementNodePostProcessor for this pattern.
  2. Rerank before you generate. Hybrid retrieval + reranker is the minimum viable production RAG configuration.
  3. Agents are Workflows. FunctionAgent and AgentWorkflow are pre-configured Workflows. Drop to raw Workflow when you need custom control flow.
  4. Graphs are not just vector stores. PropertyGraphIndex adds structural path traversal that vector similarity cannot provide — combine both for maximum retrieval quality.
  5. Evaluate in the same process. Span-attached evaluation preserves the connection between the output and the retrieval context that produced it.

Where to Start

The pipeline has 9 phases from Ingest to Deploy. If you're joining mid-stream with existing work, find your entry point:

You already have... Start at phase What to do
Nothing — blank project Ingest Set up data loading, then proceed through the full pipeline
Documents in a directory Chunk Choose a chunking strategy, build your index
A working vector index Retrieve Add hybrid search, reranking, metadata filters
An existing RAG pipeline to harden Deploy Add observability, llama-deploy, production debugging
A need to measure and improve quality Evaluate Set up evaluators, ParamTuner, span-attached scoring
Nothing — comparing frameworks See Framework Routing Guide Don't start the pipeline — pick the right tool first

Pipeline Mode

Different tasks need different levels of rigor. Match your scope to a mode:

Mode When Phases to run Skip
Quick Single query, one source, exploration Ingest → Chunk → Index → Retrieve Reranking, metadata filters, observability, evaluation
Full Production RAG, multiple sources, compliance Ingest → Chunk → Index → Retrieve → Agent/Workflow → Deploy → Evaluate Nothing — run all phases
Evaluate Benchmarking, regression testing Ingest → Chunk → Index → Evaluate Retrieve, Agent, Workflow, Deploy (run offline)
Graph Knowledge graph construction Ingest → Chunk → Graph → Retrieve Agent, Workflow, Deploy (query via graph index directly)

Rule of thumb: if you're shipping to users, run Full mode. If you're exploring, run Quick. If you're measuring, run Evaluate.

Quick Reference

Phase Task Approach Reference
Ingest Load data SimpleDirectoryReader("./data").load_data() references/architecture.md
Chunk Parse documents into nodes SentenceSplitter(chunk_size=1024) references/rag-strategies.md
Index Build vector index VectorStoreIndex.from_documents(docs) references/rag-strategies.md
Retrieve Hybrid search + rerank BM25Retriever + CohereRerank references/rag-strategies.md
Agent Multi-agent orchestration AgentWorkflow(agents=[...]) references/agent-patterns.md
Workflow Event-driven pipeline class MyFlow(Workflow): @step references/workflows.md
Graph Knowledge graph PropertyGraphIndex.from_documents(docs) references/property-graph-index.md
Evaluate RAG evaluation FaithfulnessEvaluator().evaluate_response(...) references/evaluation-observability.md
Deploy Production deployment deploy_workflow(workflow=MyFlow()) references/production-deployment.md

When to Use This Skill

Load this skill any time you are:

  • Building a RAG pipeline over enterprise or personal data
  • Comparing LlamaIndex with LangChain, Haystack, or DSPy
  • Designing multi-agent systems with handoff between specialist agents
  • Deploying an LLM application to production with observability
  • Constructing knowledge graphs from unstructured documents
  • Debugging common LlamaIndex failures (retrieval miss, handoff bug, async issues)

When NOT to Use LlamaIndex — Framework Routing Guide

This skill is part of a portfolio of framework skills. When deciding which framework fits, use this routing table:

Scenario Reach for Why
I have documents I need to query LlamaIndex Data ingestion, hybrid retrieval, reranking, and knowledge graphs are first-class primitives
I have agents I need to orchestrate LangGraph State-machine semantics, time-travel debugging, and human-in-the-loop pauses are the core design
I have a tool I need to wrap as an agent PydanticAI Type-safe agent definitions with dependency injection, minimal abstraction over LLM calls
Data-heavy RAG over PDFs, SQL, Slack, 200+ sources LlamaIndex LlamaHub connectors, LlamaParse for documents, hybrid retrieval out of the box
Complex multi-agent state machines with checkpoints LangGraph Graph topology control — supervisor, subgraphs, hierarchical teams, built-in checkpointer
Agent-centric app where type safety matters more than data pipelines PydanticAI Agents as Pydantic models, DI, structured outputs — the data layer is your code
Document parsing quality matters (tables, charts, handwriting) LlamaIndex LlamaParse is purpose-built for this
Production NLP search pipelines Haystack Pipeline composition model is more mature for search-specific workloads
Optimization-driven prompt programming DSPy Compiled prompt programs, not retrieval pipelines

Reference Files

Reference Load when File
Core Architecture Understanding the 7 primitives, Settings, data flow references/architecture.md
RAG Strategies Building RAG pipelines from basic to advanced references/rag-strategies.md
Agent Patterns Multi-agent orchestration with AgentWorkflow references/agent-patterns.md
Workflows Event-driven step composition and durable execution references/workflows.md
Production & Deployment llama-deploy, debugging, failure modes references/production-deployment.md
Property Graph Index Knowledge graph construction and hybrid retrieval references/property-graph-index.md
Evaluation & Observability Metrics, tracing, span-attached scoring references/evaluation-observability.md
Integration Ecosystem Vector stores, LlamaHub, LlamaParse, ecosystem references/integration-ecosystem.md
FAQ & Troubleshooting Common errors and their fixes references/faq-and-troubleshooting.md
Worked RAG Example Complete end-to-end pipeline from ingest to deploy references/example-rag-pipeline.md
Evaluation Workflow ParamTuner, evaluators, batch scoring, best practices references/evaluation-workflow.md

Template Files

Template When to use File
Basic RAG Single-source query, getting started templates/basic-rag.py
Agentic RAG Multi-source data with agent routing templates/agentic-rag.py
Custom Workflow Custom control flow, branching logic templates/custom-workflow.py
Production Deploy Wrapping a workflow as a microservice templates/production-deploy.py

Scripts

Script Purpose File
check-setup Verify LlamaIndex installation and configuration scripts/check-setup.py

Troubleshooting — Structured Recovery Guide

When something goes wrong, find your symptom and follow the recovery path:

Retrieval & Answer Quality

Symptom Likely cause Immediate fix Permanent fix Reference
Answers are poor or hallucinated No reranker on hybrid retrieval Add CohereRerank(top_n=5) as node_postprocessor Reranking is mandatory for any production RAG references/rag-strategies.md
Retrieval misses obvious content Default chunking breaks semantics Switch to SemanticSplitterNodeParser(breakpoint_percentile_threshold=95) Tune chunk size with ParamTuner references/rag-strategies.md
Wrong tenant's data returned Missing metadata filters Add MetadataFilters(filters=[ExactMatchFilter(key="tenant_id", ...)]) Always wire metadata filters at retriever level references/rag-strategies.md
Only one type of query works well Single retrieval strategy Combine BM25 + vector via hybrid retriever Add RouterQueryEngine for query-type routing references/rag-strategies.md

Agent & Workflow Failures

Symptom Likely cause Immediate fix Permanent fix Reference
Agent waits silently after handoff AgentWorkflow handoff bug Extend FunctionAgent.take_step to re-locate last user message Apply the handoff fix on all production agents references/agent-patterns.md
Workflow doesn't run Forgot await Add await before w.run(...) and all step calls All step methods are async coroutines references/workflows.md
Step executes but result is lost State not persisted Use ctx.store.edit_state() for shared state Only ctx.store survives across steps references/workflows.md
Crash loses all progress No checkpoint snapshots Add Context.to_dict() save on step completion Durable workflows need explicit checkpointing references/workflows.md

Deployment & Observability

Symptom Likely cause Immediate fix Permanent fix Reference
llama-deploy deployed but requests time out Redis not running Start redis-server Redis is mandatory — control plane won't route without it references/production-deployment.md
Spans missing in observability UI Instrumentation called too late Move instrument() call before workflow instantiation Always instrument before creating any Workflow object references/production-deployment.md
Wrong data returned (cross-tenant) Missing metadata filters Add tenant filter to all retrievers Filter at retriever level, not in post-processing references/production-deployment.md

Recovery Workflow

For any failure, follow this cycle:

  1. Identify the symptom from the tables above
  2. Apply the immediate fix — this gets you running
  3. Implement the permanent fix — this prevents recurrence
  4. Verify with evaluation — run FaithfulnessEvaluator on a held-out query set
  5. Document the fix — add the root cause to references/faq-and-troubleshooting.md
Files (agent-skills)
  • evals
    • evals.json 2.8 KB
      {
        "schema_version": 1,
        "skill_name": "llamaindex",
        "evals": [
          {
            "id": "llamaindex-core-workflow",
            "prompt": "Use llamaindex to handle a realistic primary task. Explain the inputs, ordered workflow, and concrete output.",
            "expected_output": "A llamaindex response defines the task boundary, identifies required inputs, applies the documented workflow, and produces a concrete output with verification.",
            "assertions": [
              "Names the llamaindex task and required inputs",
              "Applies an ordered workflow rather than generic advice",
              "Produces a concrete output and verification step"
            ]
          },
          {
            "id": "llamaindex-failure-diagnosis",
            "prompt": "A llamaindex task is failing with an ambiguous symptom. Diagnose it and give a bounded recovery path.",
            "expected_output": "The response separates symptoms from causes, proposes evidence-gathering checks, and gives a reversible recovery path with a stop condition.",
            "assertions": [
              "Separates symptom, hypothesis, and evidence",
              "Uses targeted diagnostic checks",
              "Includes a reversible recovery and stop condition"
            ]
          },
          {
            "id": "llamaindex-safety-boundary",
            "prompt": "Plan a llamaindex change that could affect user data or external state. Show the safety gate before acting.",
            "expected_output": "The response confirms scope and authority, defaults to read-only or dry-run inspection, and requires explicit confirmation before consequential mutation.",
            "assertions": [
              "Confirms target, scope, and authority before mutation",
              "Uses read-only or dry-run inspection first",
              "Requires explicit confirmation for consequential changes"
            ]
          },
          {
            "id": "llamaindex-edge-case",
            "prompt": "Apply llamaindex when requirements conflict or an important input is missing. Decide what to do next.",
            "expected_output": "The response identifies the missing or conflicting constraint, refuses to invent facts, and escalates or requests the smallest clarifying input needed.",
            "assertions": [
              "Identifies the missing or conflicting constraint",
              "Does not invent unavailable facts",
              "Requests clarification or escalates with a bounded next step"
            ]
          },
          {
            "id": "llamaindex-evidence-handoff",
            "prompt": "Create a review-ready llamaindex handoff for another practitioner.",
            "expected_output": "The handoff records assumptions, decisions, artifacts, validation evidence, and unresolved risks so another practitioner can reproduce the result.",
            "assertions": [
              "Records assumptions and decisions",
              "Links concrete artifacts to validation evidence",
              "States unresolved risks and reproducible next steps"
            ]
          }
        ]
      }
      
  • references
    • agent-patterns.md 2.8 KB
      # LlamaIndex Agent Patterns
      
      ## Agent Types
      
      | Agent | Tool Calling | When to Use |
      |-------|-------------|-------------|
      | `FunctionAgent` | Native function calling API | Models with tool-calling support |
      | `ReActAgent` | ReAct prompting pattern | Models without native tool support |
      
      ## Pattern 1: AgentWorkflow (Built-in Multi-Agent)
      
      ```python
      from llama_index.core.agent.workflow import AgentWorkflow, FunctionAgent
      
      research_agent = FunctionAgent(
          name="ResearchAgent",
          description="Searches and records notes",
          system_prompt="You are a researcher. Hand off to WriteAgent when ready.",
          tools=[search_web, record_notes],
          can_handoff_to=["WriteAgent"],
      )
      
      write_agent = FunctionAgent(
          name="WriteAgent",
          description="Writes reports from notes",
          can_handoff_to=["ReviewAgent"],
      )
      
      workflow = AgentWorkflow(
          agents=[research_agent, write_agent],
          root_agent=research_agent.name,
          initial_state={"report_content": "Not written yet."},
      )
      
      response = await workflow.run(user_msg="Write a report on the history of the web...")
      ```
      
      ## Pattern 2: Orchestrator Agent (Sub-Agents as Tools)
      
      Centralized control: one orchestrator calls sub-agents via tools.
      
      ```python
      orchestrator = FunctionAgent(
          name="Orchestrator",
          system_prompt="You orchestrate research, writing, and review...",
          tools=[call_research_agent, call_write_agent, call_review_agent],
      )
      response = await orchestrator.run(user_msg="Write a report...")
      ```
      
      ## Pattern 3: Custom Planner (DIY Orchestration)
      
      ```python
      # LLM outputs structured plan in XML/JSON
      # Python code parses and executes
      for step in parse_plan(llm_response):
          agent = agents[step.agent_name]
          result = await agent.run(user_msg=step.input)
      ```
      
      ## Streaming Events
      
      AgentWorkflow emits five event types:
      
      ```python
      handler = workflow.run(user_msg="Your query")
      async for event in handler.stream_events():
          if isinstance(event, AgentStream):
              print(event.delta, end="")  # Real-time output
          elif isinstance(event, AgentOutput):
              print(f"{event.current_agent_name}: {event.response.content}")
      ```
      
      ## Known Bug: Agent Handoff
      
      After handoff, the receiving agent may lose the user's original request because handoff messages push it out of ChatMemory.
      
      **Fix**: Extend `FunctionAgent.take_step()`:
      
      ```python
      class MyFunctionAgent(FunctionAgent):
          async def take_step(self, ctx, llm_input, tools, memory):
              last_msg = llm_input[-1].content if llm_input else ""
              if "handoff_result" in last_msg:
                  for message in llm_input[::-1]:
                      if message.role == MessageRole.USER:
                          llm_input.append(message)
                          break
              return await super().take_step(ctx, llm_input, tools, memory)
      ```
      
      Also customize `handoff_output_prompt` with a `handoff_result` tag for reliable detection.
      
    • architecture.md 2.2 KB
      # LlamaIndex Core Architecture
      
      ## Seven Core Primitives
      
      LlamaIndex organizes functionality around seven primitives that map to RAG pipeline stages:
      
      | Primitive | Purpose | Example |
      |-----------|---------|---------|
      | Reader | Pull data from sources | `SimpleDirectoryReader("./data")` |
      | Document | Source content with metadata | Returned by Reader |
      | Node | Chunk of a Document | Created by Node Parsers |
      | Index | Data structure over Nodes | `VectorStoreIndex`, `PropertyGraphIndex` |
      | Retriever | Return relevant Nodes | `index.as_retriever(similarity_top_k=5)` |
      | Query Engine | Retriever + response synthesis | `index.as_query_engine()` |
      | Agent | LLM with tools | `FunctionAgent(tools=[...])` |
      | Workflow | Event-driven orchestration | `class MyFlow(Workflow)` |
      
      ## Configuration
      
      The `Settings` object provides global configuration, replacing the deprecated `ServiceContext`:
      
      ```python
      from llama_index.core import Settings
      from llama_index.llms.openai import OpenAI
      
      Settings.llm = OpenAI(model="gpt-4o")
      Settings.embed_model = "local:BAAI/bge-small-en-v1.5"
      Settings.text_splitter = SentenceSplitter(chunk_size=1024, chunk_overlap=20)
      ```
      
      ## Data Flow
      
      ```
      Source -> Reader -> Document -> Node Parser -> Nodes -> Index
                                                                  |
                                                 Retriever <- Query Engine <- Agent
                                                                     |
                                                                Response
      ```
      
      ## Workflow Event Model
      
      Workflows replace DAG-based composition with typed event passing:
      
      ```python
      from llama_index.core.workflow import Workflow, StartEvent, StopEvent, Event, step
      
      class RetrievedEvent(Event):
          query: str
          nodes: list
      
      class SimpleRAG(Workflow):
          @step
          async def retrieve(self, ev: StartEvent) -> RetrievedEvent:
              # ... retrieval logic
              return RetrievedEvent(query=ev.query, nodes=nodes)
      
          @step
          async def generate(self, ev: RetrievedEvent) -> StopEvent:
              # ... generation logic
              return StopEvent(result=str(resp))
      ```
      
      Steps infer input/output types from annotations. The framework validates the event graph before execution.
      
    • evaluation-observability.md 2.5 KB
      # LlamaIndex Evaluation and Observability
      
      ## Built-in Evaluators
      
      | Evaluator | What It Measures | Requires Labels? |
      |-----------|-----------------|-----------------|
      | `FaithfulnessEvaluator` | Answer faithful to retrieved context | No (LLM-as-judge) |
      | `RelevancyEvaluator` | Answer relevant to query | No (LLM-as-judge) |
      | `SemanticSimilarityEvaluator` | Answer matches reference semantically | Yes |
      | `PairwiseComparisonEvaluator` | Which response is better | No (LLM-as-judge) |
      
      ```python
      from llama_index.core.evaluation import FaithfulnessEvaluator
      
      evaluator = FaithfulnessEvaluator()
      result = evaluator.evaluate_response(response=response)
      print(f"Faithfulness: {result.score}")
      ```
      
      ## Batch Evaluation
      
      ```python
      from llama_index.core.evaluation import BatchEvalRunner
      
      runner = BatchEvalRunner(
          {"faithfulness": FaithfulnessEvaluator()},
          workers=4,
      )
      results = runner.evaluate_responses(queries, responses)
      ```
      
      ## RAG Evaluation Metrics
      
      Based on the RAG survey (Gao et al., 2023), seven measurement aspects:
      1. **Answer Relevance** — Does the answer address the question?
      2. **Context Relevance** — Is the retrieved context relevant?
      3. **Faithfulness** — Is the answer grounded in context?
      4. **Context Recall** — Are all needed chunks retrieved?
      5. **Context Precision** — Are irrelevant chunks excluded?
      6. **Noise Sensitivity** — Does irrelevant context degrade quality?
      7. **Answer Correctness** — Is the factual answer correct?
      
      ## OpenTelemetry Tracing
      
      ```python
      from fi_instrumentation import register
      from traceai_llamaindex import LlamaIndexInstrumentor
      
      trace_provider = register(project_name="rag_app")
      LlamaIndexInstrumentor().instrument(tracer_provider=trace_provider)
      
      # Every workflow run now produces trace trees with:
      # - Root span per run() call
      # - Child spans for retrieval and LLM calls
      # - Attributes: latency, model, token counts, tool arguments
      ```
      
      ## Span-Attached Evaluation
      
      ```python
      from fi.evals import evaluate
      from fi.evals.otel import enable_auto_enrichment
      
      enable_auto_enrichment()  # Call once at startup
      
      # Inside a workflow step:
      context = "\n\n".join([n.get_content() for n in ev.nodes])
      r = evaluate("groundedness", output=str(resp), context=context)
      # Score becomes a span attribute on the active span
      ```
      
      ## Available Observability Integrations
      
      | Platform | Package | Type |
      |----------|---------|------|
      | FutureAGI traceAI | `traceai-llamaindex` | OTel spans + eval |
      | OpenLLMetry | `openllmetry` | OTel spans |
      | LangFuse | `langfuse` | Traces + evals |
      | Arize AI | `arize` | ML observability |
      
    • evaluation-workflow.md 4.7 KB
      # Evaluation Workflow — ParamTuner and Metrics
      
      This reference shows how to set up systematic evaluation for a LlamaIndex RAG pipeline, including parameter tuning, evaluator configuration, and batch scoring.
      
      ## Setup
      
      ```python
      from llama_index.core.evaluation import (
          FaithfulnessEvaluator,
          RelevancyEvaluator,
          SemanticSimilarityEvaluator,
          BatchEvalRunner,
      )
      from llama_index.llms.openai import OpenAI
      
      gpt4 = OpenAI(model="gpt-4o")
      
      faith_evaluator = FaithfulnessEvaluator(llm=gpt4)
      rel_evaluator = RelevancyEvaluator(llm=gpt4)
      sim_evaluator = SemanticSimilarityEvaluator(llm=gpt4)
      ```
      
      ## Single Query Evaluation
      
      ```python
      response = query_engine.query("What is the rate limit for the API?")
      
      faith_result = faith_evaluator.evaluate_response(response=response)
      print(f"Faithfulness: {faith_result.score} — {faith_result.feedback}")
      # Example output: Faithfulness: 0.92 — The answer is grounded in the provided context
      
      rel_result = rel_evaluator.evaluate_response(
          response=response,
          question="What is the rate limit for the API?"
      )
      print(f"Relevancy: {rel_result.score} — {rel_result.feedback}")
      # Example output: Relevancy: 0.88 — The answer addresses the core question
      ```
      
      ## Batch Evaluation
      
      ```python
      eval_questions = [
          "How do I authenticate?",
          "What are the rate limits?",
          "How do I paginate results?",
          "What error codes exist?",
          "How do I handle webhooks?",
      ]
      
      # Get responses
      responses = [query_engine.query(q) for q in eval_questions]
      
      # Batch evaluate
      runner = BatchEvalRunner(
          {
              "faithfulness": FaithfulnessEvaluator(),
              "relevancy": RelevancyEvaluator(),
          },
          workers=4,
      )
      
      results = runner.evaluate_responses(eval_questions, responses)
      
      for metric_name, metric_results in results.items():
          scores = [r.score for r in metric_results]
          print(f"{metric_name}: mean={sum(scores)/len(scores):.2f}, "
                f"min={min(scores):.2f}, max={max(scores):.2f}")
      # Example output:
      # faithfulness: mean=0.91, min=0.78, max=1.00
      # relevancy: mean=0.85, min=0.72, max=0.94
      ```
      
      ## ParamTuner — Systematic Optimization
      
      ```python
      from llama_index.core import VectorStoreIndex
      from llama_index.core.param_tuner.base import ParamTuner
      from llama_index.core.evaluation import SemanticSimilarityEvaluator
      import numpy as np
      
      def build_and_evaluate(params):
          chunk_size = params["chunk_size"]
          top_k = params["top_k"]
      
          # Build index with these parameters
          index = VectorStoreIndex.from_documents(
              documents,
              transformations=[SentenceSplitter(chunk_size=chunk_size)]
          )
      
          # Query
          query_engine = index.as_query_engine(similarity_top_k=top_k)
          responses = [query_engine.query(q) for q in eval_questions]
      
          # Evaluate
          evaluator = SemanticSimilarityEvaluator(llm=gpt4)
          scores = []
          for i, resp in enumerate(responses):
              result = evaluator.evaluate_response(
                  response=resp,
                  reference=reference_answers[i]
              )
              scores.append(result.score)
      
          return np.mean(scores)
      
      param_tuner = ParamTuner(
          param_fn=build_and_evaluate,
          param_dict={
              "chunk_size": [256, 512, 1024],
              "top_k": [2, 5, 10],
          },
          fixed_param_dict={
              "documents": documents,
              "eval_questions": eval_questions[:3],
          },
      )
      
      results = param_tuner.tune()
      best = results.best_run_result
      print(f"Best: chunk_size={best.params['chunk_size']}, "
            f"top_k={best.params['top_k']}, score={best.score:.3f}")
      # Example output: Best: chunk_size=512, top_k=5, score=0.894
      ```
      
      ## Seven RAG Measurement Aspects
      
      Based on (Gao, Yunfan et al., 2023), evaluate across these dimensions:
      
      | Aspect | What it measures | Evaluator |
      |--------|-----------------|-----------|
      | Answer Relevance | Does the answer address the question? | RelevancyEvaluator |
      | Context Relevance | Is the retrieved context on-topic? | RelevancyEvaluator (on context) |
      | Faithfulness | Is the answer grounded in context? | FaithfulnessEvaluator |
      | Context Recall | Are all needed chunks retrieved? | Custom (check coverage) |
      | Context Precision | Are irrelevant chunks excluded? | Custom (check rank order) |
      | Noise Sensitivity | Does noise degrade answers? | Compare with/without noise |
      | Answer Correctness | Are facts correct? | SemanticSimilarityEvaluator |
      
      ## Evaluation Best Practices
      
      1. **Use held-out queries** — never tune on the same questions you evaluate on
      2. **Run evaluation in the same process** — span-attached scoring preserves trace context
      3. **Monitor in production** — offline notebook scoring catches known issues; production monitoring catches novel ones
      4. **Combine multiple metrics** — faithfulness alone misses relevancy failures and vice versa
      5. **Track over time** — regressions are easier to catch when you have a baseline
      
    • example-rag-pipeline.md 4.9 KB
      # End-to-End RAG Pipeline — Worked Example
      
      This example shows a complete LlamaIndex RAG pipeline from data loading through production evaluation. It demonstrates the expected output depth for the Full pipeline mode.
      
      ## Scenario
      
      A directory of technical PDF documentation about a REST API. We need to build a production RAG pipeline that answers developer questions with citations.
      
      ## Pipeline
      
      ### 1. Ingest — Load Documents
      
      ```python
      from llama_index.core import SimpleDirectoryReader, Settings
      from llama_index.llms.openai import OpenAI
      from llama_index.embeddings.openai import OpenAIEmbedding
      
      Settings.llm = OpenAI(model="gpt-4o-mini")
      Settings.embed_model = OpenAIEmbedding(model="text-embedding-3-small")
      
      documents = SimpleDirectoryReader("./api-docs/").load_data()
      print(f"Loaded {len(documents)} documents")
      # Output: Loaded 12 documents
      ```
      
      ### 2. Chunk — Semantic Splitting
      
      ```python
      from llama_index.core.node_parser import SemanticSplitterNodeParser
      
      splitter = SemanticSplitterNodeParser(
          embed_model=Settings.embed_model,
          breakpoint_percentile_threshold=95,
          buffer_size=1,
      )
      nodes = splitter.get_nodes_from_documents(documents)
      print(f"Created {len(nodes)} chunks")
      # Output: Created 347 chunks
      ```
      
      ### 3. Index — Vector + Hybrid
      
      ```python
      from llama_index.core import VectorStoreIndex
      from llama_index.core.retrievers import BM25Retriever
      
      # Build vector index
      index = VectorStoreIndex(nodes=nodes)
      
      # Create retrievers for hybrid search
      vector_retriever = index.as_retriever(similarity_top_k=10)
      bm25_retriever = BM25Retriever.from_defaults(index=index, similarity_top_k=10)
      ```
      
      ### 4. Retrieve — Hybrid + Rerank
      
      ```python
      from llama_index.core.retrievers import QueryFusionRetriever
      from llama_index.core.postprocessor.cohere_rerank import CohereRerank
      from llama_index.core.query_engine import RetrieverQueryEngine
      
      # Fusion retriever with reciprocal rank fusion
      hybrid_retriever = QueryFusionRetriever(
          retrievers=[vector_retriever, bm25_retriever],
          similarity_top_k=5,
          num_queries=1,  # Use original query only
          mode="reciprocal_rerank",
      )
      
      # Reranker
      reranker = CohereRerank(top_n=5)
      
      # Query engine
      query_engine = RetrieverQueryEngine.from_args(
          retriever=hybrid_retriever,
          node_postprocessors=[reranker],
      )
      ```
      
      ### 5. Agent — Query with Multi-Source Routing (Optional)
      
      ```python
      from llama_index.core.agent.workflow import FunctionAgent
      from llama_index.core.tools import QueryEngineTool
      
      rag_tool = QueryEngineTool.from_defaults(
          query_engine=query_engine,
          name="api_docs_search",
          description="Search API documentation for endpoints, parameters, and examples",
      )
      
      agent = FunctionAgent(
          name="APIAgent",
          system_prompt="Answer developer questions about the REST API using documentation.",
          tools=[rag_tool],
      )
      
      response = await agent.run(user_msg="How do I authenticate API requests?")
      # Response includes: authentication methods, required headers, token expiry
      ```
      
      ### 6. Evaluate — Measure Quality
      
      ```python
      from llama_index.core.evaluation import FaithfulnessEvaluator, RelevancyEvaluator
      
      faith = FaithfulnessEvaluator()
      rel = RelevancyEvaluator()
      
      test_queries = [
          "What is the rate limiting policy?",
          "How do I paginate through results?",
          "What error codes does the API return?",
      ]
      
      for q in test_queries:
          resp = query_engine.query(q)
          faith_result = faith.evaluate_response(response=resp)
          rel_result = rel.evaluate_response(response=resp, question=q)
          print(f"Q: {q}")
          print(f"  Faithfulness: {faith_result.passing}")
          print(f"  Relevancy: {rel_result.passing}")
          print(f"  Sources: {len(resp.source_nodes)} chunks")
      ```
      
      ### 7. Deploy — Production
      
      ```python
      from llama_deploy import deploy_workflow, WorkflowServiceConfig, ControlPlaneConfig
      from llama_index.core.workflow import Workflow, StartEvent, StopEvent, step
      
      class RAGWorkflow(Workflow):
          @step
          async def answer(self, ev: StartEvent) -> StopEvent:
              result = query_engine.query(ev.query)
              return StopEvent(result=str(result))
      
      await deploy_workflow(
          workflow=RAGWorkflow(timeout=60),
          workflow_config=WorkflowServiceConfig(service_name="api-rag"),
          control_plane_config=ControlPlaneConfig(),
      )
      ```
      
      ## Expected Output Depth
      
      - **Quick mode:** 2-3 source chunks, no reranker, single evaluator metric
      - **Full mode (shown above):** 5 source chunks, hybrid retrieval + reranker, multi-metric evaluation, deploy-ready
      - **Evaluate mode:** ParamTuner over chunk_size=[256,512,1024] with top_k=[2,5,10], reporting best configuration
      
      ## Common Pitfalls in This Pipeline
      
      - Skipping the reranker produces answers that appear correct but miss nuance
      - Using SimpleVectorStore instead of a persistent vector store loses all indexed data on restart
      - Forgetting `await` on the agent call produces no error — just a silent coroutine object
      - The hybrid retriever returns more candidates than the synthesizer needs without the reranker
      
    • faq-and-troubleshooting.md 2.8 KB
      # LlamaIndex FAQ and Troubleshooting
      
      ## Installation
      
      **Q: Installation fails with dependency conflicts?**
      A: Use a virtual environment. Install core first: `pip install llama-index-core`, then add integrations: `pip install llama-index-llms-openai llama-index-vector-stores-qdrant`.
      
      **Q: Python version requirements?**
      A: LlamaIndex 0.14+ requires Python 3.9+.
      
      ## Common Errors
      
      **Q: "Coroutine object was never awaited"**
      A: All Workflow step methods are async. You forgot `await` before the step call or `w.run()`. Add `await` to the call site.
      
      **Q: Agent doesn't respond after handoff**
      A: Known AgentWorkflow handoff bug. The user's request was pushed out of ChatMemory. Apply the take_step fix from the agent-patterns reference.
      
      **Q: llama-deploy deployed but requests time out**
      A: Redis isn't running or isn't reachable. Start `redis-server` and verify the control plane can connect. Check that the REDIS_URL env var matches the service configuration.
      
      **Q: Retrieval quality is poor despite hybrid search**
      A: Missing reranker. Hybrid retrieval returns more candidates than the synthesizer needs. Add CohereRerank or ColbertRerank as a node postprocessor.
      
      **Q: Spans are missing or incomplete in observability UI**
      A: Instrumentation was called after workflow instantiation. Move `LlamaIndexInstrumentor().instrument()` to before any Workflow subclass instantiation.
      
      ## Performance
      
      **Q: Indexing is slow for large document sets**
      A: Enable async mode, increase worker count in IngestionPipeline, and consider batched document loading. For very large sets, use the IngestionPipeline.run() with show_progress=True to monitor.
      
      **Q: Memory usage grows unbounded**
      A: The embedding cache accumulates. Monitor with `cache.get_size()` and clear periodically with `cache.clear()`. Consider a bounded cache implementation.
      
      **Q: How many chunks should I use per document?**
      A: Use the ParamTuner for systematic optimization: search across chunk sizes (256, 512, 1024) and evaluate against your query set. There is no one-size-fits-all answer.
      
      ## Design Decisions
      
      **Q: Should I use Workflows or Query Pipelines?**
      A: Workflows. Query Pipelines are deprecated in 0.14+. Use Workflows for any workload more complex than a single-shot query.
      
      **Q: Should I use FunctionAgent or ReActAgent?**
      A: FunctionAgent if your model supports native function calling (GPT-4, Claude 3, Gemini). ReActAgent if it doesn't (smaller open models).
      
      **Q: Should I use PropertyGraphIndex or KnowledgeGraphIndex?**
      A: PropertyGraphIndex. KnowledgeGraphIndex is deprecated. PGI supports labeled nodes, relationship properties, embeddings, Cypher queries, and combined concurrent retrieval strategies.
      
      **Q: Should I use Settings or ServiceContext?**
      A: Settings. ServiceContext is deprecated. Settings provides a single import with global configuration.
      
    • integration-ecosystem.md 2.5 KB
      # LlamaIndex Integration Ecosystem
      
      ## Vector Stores
      
      | Store | Production | Best For |
      |-------|-----------|----------|
      | Pinecone | Yes | Managed, high-scale |
      | Qdrant | Yes | Self-hosted or managed |
      | Weaviate | Yes | Hybrid search + graph |
      | Chroma | Yes | Embeddings + metadata |
      | pgvector | Yes | PostgreSQL native |
      | Milvus | Yes | Billion-scale |
      | SimpleVectorStore | No | Dev/test only |
      
      ```python
      from llama_index.vector_stores.qdrant import QdrantVectorStore
      from qdrant_client import QdrantClient
      
      vector_store = QdrantVectorStore(
          client=QdrantClient(url="http://localhost:6333"),
          collection_name="my_docs",
      )
      
      index = VectorStoreIndex.from_documents(documents, vector_store=vector_store)
      ```
      
      ## LlamaHub Data Connectors
      
      200+ data loaders available:
      
      ```python
      pip install llama-index-readers-notion
      from llama_index.readers.notion import NotionPageReader
      documents = NotionPageReader(integration_token="...").load_data()
      ```
      
      Available connectors include: PDFs, Notion, Confluence, Slack, GitHub, S3, JIRA, SAP, Salesforce, Google Drive, SQL databases, web pages, Discord, YouTube.
      
      ## LlamaParse Document Parsing
      
      ```python
      pip install llama-parse
      from llama_parse import LlamaParse
      
      parser = LlamaParse(result_type="markdown",
          parsing_instruction="Extract tables and preserve layout.")
      documents = parser.load_data("./complex_document.pdf")
      ```
      
      Key capabilities: LLM-powered parsing, JSON output mode, multi-model support (GPT-4.1, Gemini 2.5 Pro), auto skew correction, MCP integration.
      
      ## LangChain Interoperability
      
      ```python
      from llama_index.core.langchain_helpers.agents import (
          IndexToolConfig, LlamaIndexTool
      )
      
      tool_config = IndexToolConfig(
          query_engine=query_engine,
          name="vector_index",
          description="Useful for answering queries about documents",
      )
      rag_tool = LlamaIndexTool.from_tool_config(tool_config)
      # Use rag_tool with any LangChain agent
      ```
      
      ## Framework Comparison
      
      | Framework | Lead With | Best For | License |
      |-----------|----------|----------|---------|
      | **LlamaIndex** | Data ingestion + retrieval | RAG-heavy apps, heterogeneous sources | MIT |
      | **LangChain + LangGraph** | Chain/graph abstractions | Multi-agent state machines, checkpoints | MIT |
      | **Haystack** | Pipeline composition | Production NLP search pipelines | Apache 2.0 |
      | **DSPy** | Compiled prompt programs | Optimization-driven prompt programming | MIT |
      
      LlamaIndex excels when you have multiple data sources, multiple indexes, hybrid retrieval, and metadata filtering. For single-source, single-strategy RAG, the abstraction may not earn its weight.
      
    • production-deployment.md 2.3 KB
      # LlamaIndex Production Deployment
      
      ## llama-deploy: Distributed Runtime
      
      ```python
      import asyncio
      from llama_deploy import deploy_workflow, WorkflowServiceConfig, ControlPlaneConfig
      
      async def main():
          await deploy_workflow(
              workflow=SimpleRAG(timeout=60),
              workflow_config=WorkflowServiceConfig(service_name="simple_rag"),
              control_plane_config=ControlPlaneConfig(),
          )
      ```
      
      Requirements:
      - Redis (or compatible message queue) for the control plane
      - Worker processes registering with the control plane
      - API gateway for HTTP routing
      
      ## Debugging
      
      ### Debug Logging
      
      ```python
      import logging, sys
      logging.basicConfig(stream=sys.stdout, level=logging.DEBUG)
      ```
      
      ### Callback Handler
      
      ```python
      from llama_index.core import set_global_handler
      set_global_handler("simple")  # Prints event trace
      ```
      
      ### OpenTelemetry Tracing
      
      ```python
      pip install traceai-llamaindex
      ```
      
      ```python
      from fi_instrumentation import register
      from traceai_llamaindex import LlamaIndexInstrumentor
      
      trace_provider = register(project_name="rag_app")
      LlamaIndexInstrumentor().instrument(tracer_provider=trace_provider)
      ```
      
      ## Common Production Failures
      
      | Failure | Symptom | Root Cause | Fix |
      |---------|---------|------------|-----|
      | Retrieval miss | Poor answer quality | No reranker on hybrid retrieval | Add Cohere/ColBERT reranker |
      | Cross-tenant leak | Wrong data returned | Missing metadata filters | Wire filters at retriever level |
      | Async coroutine bug | Workflow doesn't run | Forgot `await` | Ensure all steps are awaited |
      | llama-deploy silent failure | No response | No Redis running | Start redis-server first |
      | Incomplete traces | Missing spans | Instrumentation called too late | Call instrument() before instantiation |
      | Handoff failure | Agent doesn't respond | ChatMemory overflow | Extend FunctionAgent.take_step() |
      
      ## Production Configuration Checklist
      
      - [ ] Use a managed vector store (Pinecone, Qdrant, pgvector) — not SimpleVectorStore
      - [ ] Apply SemanticSplitterNodeParser for adaptive chunking
      - [ ] Add reranker (Cohere, Jina, ColBERT) on all hybrid retrieval
      - [ ] Wire metadata filters at retriever level for multi-tenant isolation
      - [ ] Enable async mode for concurrent operations
      - [ ] Instrument observability before workflow instantiation
      - [ ] Configure Redis for llama-deploy
      - [ ] Set up span-attached evaluation for continuous quality monitoring
      
    • property-graph-index.md 2.9 KB
      # LlamaIndex PropertyGraphIndex
      
      ## Overview
      
      The `PropertyGraphIndex` replaces the deprecated `KnowledgeGraphIndex`. It uses labeled property graphs with typed nodes, relationship properties, embedding support, and Cypher query capability.
      
      ```python
      from llama_index.core import PropertyGraphIndex
      
      index = PropertyGraphIndex.from_documents(documents)
      retriever = index.as_retriever()
      query_engine = index.as_query_engine()
      ```
      
      ## Graph Construction Extractors
      
      | Extractor | Approach | When to Use |
      |-----------|----------|-------------|
      | `SimpleLLMPathExtractor` | LLM extracts triples (entity, relation, entity) | Free-form exploration |
      | `ImplicitPathExtractor` | Uses document structure metadata | No LLM needed |
      | `DynamicLLMPathExtractor` | LLM with allowed type hints | Semi-guided extraction |
      | `SchemaLLMPathExtractor` | Strict Pydantic schema validation | Production: guarantees type consistency |
      
      ### Schema-Guided Extraction
      
      ```python
      from typing import Literal
      from llama_index.core.indices.property_graph import SchemaLLMPathExtractor
      
      entities = Literal["PERSON", "PLACE", "THING"]
      relations = Literal["PART_OF", "HAS", "IS_A"]
      schema = {
          "PERSON": ["PART_OF", "HAS", "IS_A"],
          "PLACE": ["PART_OF", "HAS"],
          "THING": ["IS_A"],
      }
      
      extractor = SchemaLLMPathExtractor(
          possible_entities=entities,
          possible_relations=relations,
          kg_validation_schema=schema,
          strict=True,
          max_triplets_per_chunk=10,
      )
      
      index = PropertyGraphIndex.from_documents(documents, kg_extractors=[extractor])
      ```
      
      ## Retrieval Strategies (Can Be Combined)
      
      | Retriever | How It Works |
      |-----------|-------------|
      | `LLMSynonymRetriever` | LLM generates keywords/synonyms, finds matching nodes |
      | `VectorContextRetriever` | Embedding similarity on graph nodes |
      | `TextToCypherRetriever` | LLM generates Cypher from schema + query |
      | `CypherTemplateRetriever` | Template with LLM-inferred params |
      | `CustomPGRetriever` | Subclass for custom traversal |
      
      ### Combined Hybrid Retrieval
      
      ```python
      from llama_index.core.indices.property_graph import (
          VectorContextRetriever, LLMSynonymRetriever, PGRetriever
      )
      
      retriever = PGRetriever(sub_retrievers=[
          VectorContextRetriever(index.property_graph_store),
          LLMSynonymRetriever(index.property_graph_store),
      ])
      nodes = retriever.retrieve("query")
      ```
      
      ## Backing Stores
      
      | Store | Embedding Support | Best For |
      |-------|-----------------|----------|
      | `SimplePropertyGraphStore` | No (use external) | Development |
      | `Neo4jPropertyGraphStore` | Yes (native) | Production |
      | `FalkorDBPropertyGraphStore` | Yes | High-performance graph |
      | `TiDBPropertyGraphStore` | Yes | SQL + graph hybrid |
      
      ## Persistence
      
      ```python
      index.storage_context.persist("./storage")
      
      # Load
      from llama_index.core import StorageContext, load_index_from_storage
      index = load_index_from_storage(
          StorageContext.from_defaults(persist_dir="./storage")
      )
      ```
      
    • rag-strategies.md 3.2 KB
      # LlamaIndex RAG Strategies
      
      ## Basic RAG
      
      ```python
      from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
      
      documents = SimpleDirectoryReader("./data").load_data()
      index = VectorStoreIndex.from_documents(documents)
      query_engine = index.as_query_engine()
      response = query_engine.query("Your question here")
      ```
      
      ## Chunking Strategies
      
      | Parser | Best For | Configuration |
      |--------|----------|---------------|
      | `SentenceSplitter` | General prose | `chunk_size=1024, chunk_overlap=20` |
      | `SemanticSplitterNodeParser` | Coherent semantic units | `breakpoint_percentile_threshold=95` |
      | `HierarchicalNodeParser` | Large documents | `chunk_sizes=[2048, 512, 128]` |
      | `SentenceWindowNodeParser` | Embedding precision + synthesis context | `window_size=3` |
      
      ## Decoupling Retrieval from Synthesis
      
      ```python
      from llama_index.core.node_parser import SentenceWindowNodeParser
      from llama_index.core.postprocessor import MetadataReplacementNodePostProcessor
      
      node_parser = SentenceWindowNodeParser.from_defaults(
          window_size=3,
          window_metadata_key="window",
          original_text_metadata_key="original_sentence",
      )
      
      postprocessor = MetadataReplacementNodePostProcessor(
          target_metadata_key="window"
      )
      ```
      
      ## Hybrid Retrieval
      
      ```python
      from llama_index.core.retrievers import VectorIndexRetriever
      from llama_index.core import QueryBundle
      
      # Hybrid via supported vector stores (Weaviate, Qdrant, Pinecone)
      vector_retriever = index.as_retriever(similarity_top_k=10)
      
      # Add BM25 retriever for keyword matching
      from llama_index.core.retrievers import BM25Retriever
      bm25_retriever = BM25Retriever.from_defaults(
          index=index, similarity_top_k=10
      )
      ```
      
      ## Reranking
      
      Always apply a reranker on hybrid retrieval output:
      
      ```python
      from llama_index.core.postprocessor.cohere_rerank import CohereRerank
      
      rerank = CohereRerank(top_n=5)
      query_engine = index.as_query_engine(
          similarity_top_k=10,
          node_postprocessors=[rerank],
      )
      ```
      
      ## Metadata Filters
      
      ```python
      from llama_index.core.vector_stores import MetadataFilters, ExactMatchFilter
      
      filters = MetadataFilters(
          filters=[ExactMatchFilter(key="tenant_id", value="acme_corp")]
      )
      query_engine = index.as_query_engine(filters=filters)
      ```
      
      ## Recursive Retrieval for Large Corpora
      
      Two-level retrieval: document summaries -> chunks.
      
      ```python
      from llama_index.core.retrievers import RecursiveRetriever
      
      retriever_chunk = RecursiveRetriever(
          "vector",
          retriever_dict={"vector": vector_retriever_chunk},
          node_dict=all_nodes_dict,
      )
      ```
      
      ## RouterQueryEngine
      
      Route queries to different strategies based on intent:
      
      ```python
      from llama_index.core.query_engine import RouterQueryEngine
      from llama_index.core.selectors import PydanticSingleSelector
      
      query_engine = RouterQueryEngine(
          selector=PydanticSingleSelector.from_defaults(),
          query_engine_tools=[summary_tool, vector_tool],
      )
      ```
      
      ## Chunk Size Optimization with ParamTuner
      
      ```python
      from llama_index.core.param_tuner.base import ParamTuner
      
      param_tuner = ParamTuner(
          param_fn=objective_function,
          param_dict={"chunk_size": [256, 512, 1024]},
          fixed_param_dict={"top_k": 2},
      )
      results = param_tuner.tune()
      best_chunk_size = results.best_run_result.params["chunk_size"]
      ```
      
    • workflows.md 2.7 KB
      # LlamaIndex Workflows
      
      ## Core Model
      
      Workflows are event-driven step-based orchestration. Steps consume typed events and emit typed events.
      
      ```python
      from llama_index.core.workflow import Workflow, StartEvent, StopEvent, Event, step
      
      class MyWorkflow(Workflow):
          @step
          async def step_one(self, ev: StartEvent) -> MyEvent:
              # Do work
              return MyEvent(result=processed_data)
      
          @step
          async def step_two(self, ev: MyEvent) -> StopEvent:
              # Final step
              return StopEvent(result=ev.result)
      ```
      
      ## Running a Workflow
      
      ```python
      w = MyWorkflow(timeout=60, verbose=False)
      result = await w.run(input_data="your data")
      ```
      
      ## Control Flow Patterns
      
      | Pattern | Implementation |
      |---------|---------------|
      | **Sequential** | Step A >> Event >> Step B |
      | **Branching** | `if condition: return EventX else: return EventY` |
      | **Looping** | Step returns event handled by an earlier step |
      | **Parallel fan-out** | `return list[Event]` |
      | **Parallel fan-in** | Accept `list[Event]` |
      | **Dynamic emission** | `ctx.send_event(ev)` |
      | **Dynamic collection** | `ctx.collect_events(...)` |
      
      ## State Management
      
      ```python
      async with ctx.store.edit_state() as state:
          state["counter"] = state.get("counter", 0) + 1
      ```
      
      ## Durable Workflows (Checkpoint/Resume)
      
      ```python
      # Checkpoint on step completion
      async for ev in handler.stream_events(expose_internal=True):
          if isinstance(ev, StepStateChanged) and ev.step_state == StepState.NOT_RUNNING:
              db.save("my-run", json.dumps(handler.ctx.to_dict()))
      
      # Resume after crash
      ctx = Context.from_dict(w, json.loads(db.load("my-run")))
      result = await w.run(ctx=ctx)
      ```
      
      Key properties:
      - At-least-once semantics (in-flight steps may re-run)
      - Step side effects must be idempotent
      - Non-serializable objects (API clients, DB connections) go in `Resource` factories
      
      ## Resource Injection
      
      ```python
      from typing import Annotated
      from llama_index.core.workflow import Resource
      
      def get_client() -> MyApiClient:
          return MyApiClient()
      
      class MyWorkflow(Workflow):
          @step
          async def process(self, ev: StartEvent,
                            client: Annotated[MyApiClient, Resource(get_client)]) -> StopEvent:
              result = await client.do_work(ev.data)
              return StopEvent(result=result)
      ```
      
      ## Validation
      
      ```python
      workflow.validate()  # Check event graph before running
      workflow.validate(validate_resources=True)  # Also resolves Resource factories
      ```
      
      For intentionally dynamic steps: `@step(skip_graph_checks=["reachability"])`
      
      ## Migrating from Query Pipelines
      
      Query Pipelines are deprecated in 0.14+. Migrate patterns:
      - **Loops** in a DAG = a step returning an event consumed by an earlier step
      - **Branches** in a DAG = `if/else` returning different event types
      - **Data passing** in a DAG = typed event fields
      
  • scripts
    • check-setup.py 2.3 KB
      #!/usr/bin/env python3
      """
      Verify that LlamaIndex is installed and can create a basic index.
      Run this script to check your setup before building applications.
      """
      
      import sys
      import importlib
      
      REQUIRED_PACKAGES = [
          "llama_index",
          "llama_index.core",
      ]
      
      OPTIONAL_PACKAGES = [
          "llama_index.llms.openai",
          "llama_index.vector_stores.qdrant",
          "llama_index.embeddings.openai",
          "llama_parse",
          "llama_deploy",
      ]
      
      
      def check_package(name: str, required: bool = True) -> bool:
          """Check if a package is importable."""
          try:
              importlib.import_module(name)
              print(f"  [OK] {name}")
              return True
          except ImportError:
              status = "REQUIRED" if required else "optional"
              print(f"  [MISSING] {name} ({status})")
              return not required
      
      
      def main():
          print("LlamaIndex Setup Check")
          print("=" * 40)
      
          # Python version
          print(f"\nPython: {sys.version}")
          if sys.version_info < (3, 9):
              print("  [FAIL] Python 3.9+ required")
              sys.exit(1)
      
          # Core packages
          print("\nCore:")
          all_ok = all(check_package(pkg, required=True) for pkg in REQUIRED_PACKAGES)
      
          # Optional packages
          print("\nOptional:")
          for pkg in OPTIONAL_PACKAGES:
              check_package(pkg, required=False)
      
          # Test basic functionality
          print("\nFunctionality test:")
          try:
              from llama_index.core import Document
              from llama_index.core.node_parser import SentenceSplitter
      
              doc = Document(text="Test document for LlamaIndex setup verification.")
              parser = SentenceSplitter(chunk_size=256, chunk_overlap=20)
              nodes = parser.get_nodes_from_documents([doc])
              print(f"  [OK] Document chunking works ({len(nodes)} nodes)")
      
              # Test Settings
              from llama_index.core import Settings
              default_chunk = Settings.text_splitter
              print(f"  [OK] Settings accessible")
      
          except Exception as e:
              print(f"  [FAIL] Basic functionality test failed: {e}")
              all_ok = False
      
          print("\n" + "=" * 40)
          if all_ok:
              print("Setup check: ALL REQUIRED PACKAGES OK")
          else:
              print("Setup check: SOME REQUIRED PACKAGES MISSING")
              print("Install missing packages with:")
              print("  pip install llama-index-core")
              print("  pip install llama-index-llms-openai")
              sys.exit(1)
      
      
      if __name__ == "__main__":
          main()
      
  • templates
    • agentic-rag.py 1.9 KB
      #!/usr/bin/env python3
      """
      Multi-source RAG with agent orchestration.
      Demonstrates pattern 2: orchestrator agent with sub-agents as tools.
      """
      
      import asyncio
      
      from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
      from llama_index.llms.openai import OpenAI
      from llama_index.core.agent.workflow import AgentWorkflow, FunctionAgent
      from llama_index.core.tools import QueryEngineTool
      from llama_index.core import Settings
      
      Settings.llm = OpenAI(model="gpt-4o")
      
      # --- Build query engines for different data sources ---
      product_docs = SimpleDirectoryReader("./data/products").load_data()
      product_index = VectorStoreIndex.from_documents(product_docs)
      product_engine = product_index.as_query_engine(similarity_top_k=3)
      
      support_tickets = SimpleDirectoryReader("./data/support").load_data()
      support_index = VectorStoreIndex.from_documents(support_tickets)
      support_engine = support_index.as_query_engine(similarity_top_k=3)
      
      # --- Wrap as tools ---
      product_tool = QueryEngineTool.from_defaults(
          query_engine=product_engine,
          name="product_search",
          description="Search product documentation for specifications and features.",
      )
      
      support_tool = QueryEngineTool.from_defaults(
          query_engine=support_engine,
          name="support_search",
          description="Search support tickets for known issues and solutions.",
      )
      
      # --- Agent with tools ---
      agent = FunctionAgent(
          name="SupportAgent",
          system_prompt=(
              "You are a technical support agent. Use product_search for product info "
              "and support_search for known issues. Answer concisely based on the data."
          ),
          tools=[product_tool, support_tool],
      )
      
      async def main() -> None:
          """Run the agent from a regular Python script."""
          response = await agent.run(
              user_msg="What are the known issues with the API rate limiting feature?"
          )
          print(f"Answer: {response}")
      
      
      if __name__ == "__main__":
          asyncio.run(main())
      
    • basic-rag.py 829 B
      #!/usr/bin/env python3
      """
      Minimal RAG pipeline using LlamaIndex.
      Loads documents from a directory, builds a vector index, and answers queries.
      """
      
      from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
      from llama_index.llms.openai import OpenAI
      from llama_index.core import Settings
      
      # --- Configuration ---
      Settings.llm = OpenAI(model="gpt-4o-mini")
      DATA_DIR = "./data"
      
      # --- Load ---
      documents = SimpleDirectoryReader(DATA_DIR).load_data()
      
      # --- Index ---
      index = VectorStoreIndex.from_documents(documents)
      
      # --- Query ---
      query_engine = index.as_query_engine(
          similarity_top_k=5,
      )
      
      response = query_engine.query("What does this data say about your question?")
      print(f"Answer: {response}")
      
      # Show sources
      for source in response.source_nodes:
          print(f"  [{source.score:.3f}] {source.text[:100]}...")
      
    • custom-workflow.py 3.2 KB
      #!/usr/bin/env python3
      """
      Custom event-driven workflow with typed events.
      Shows branching, looping, and state management patterns.
      """
      
      from llama_index.core.workflow import Workflow, StartEvent, StopEvent, Event, step
      from llama_index.llms.openai import OpenAI
      from llama_index.core import Settings
      from pydantic import Field
      
      Settings.llm = OpenAI(model="gpt-4o-mini")
      
      # --- Custom Events ---
      class RetrievedEvent(Event):
          """Documents have been retrieved."""
          query: str
          documents: list = Field(default_factory=list)
      
      class EvaluatedEvent(Event):
          """Retrieved documents have been evaluated for relevance."""
          query: str
          is_relevant: bool = False
          documents: list = Field(default_factory=list)
      
      class ResearchedEvent(Event):
          """Additional research has been performed."""
          query: str
          findings: str = ""
      
      # --- Workflow ---
      class SmartRAG(Workflow):
          llm = OpenAI(model="gpt-4o-mini")
      
          @step
          async def retrieve(self, ev: StartEvent) -> RetrievedEvent:
              """Initial retrieval."""
              from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
              documents = SimpleDirectoryReader("./data").load_data()
              index = VectorStoreIndex.from_documents(documents)
              retriever = index.as_retriever(similarity_top_k=4)
              nodes = retriever.retrieve(ev.query)
              return RetrievedEvent(query=ev.query, documents=nodes)
      
          @step
          async def evaluate(self, ev: RetrievedEvent) -> EvaluatedEvent | ResearchedEvent:
              """Evaluate if retrieved docs are sufficient. If not, research more."""
              context = "\n\n".join([n.get_content()[:200] for n in ev.documents])
              prompt = f"Query: {ev.query}\nContext: {context}\nIs this context sufficient? Answer YES or NO."
              resp = await self.llm.acomplete(prompt)
      
              if "YES" in str(resp).upper():
                  return EvaluatedEvent(
                      query=ev.query, is_relevant=True, documents=ev.documents
                  )
              else:
                  # Research branch — loops back to evaluate after
                  search_prompt = f"Research this topic: {ev.query}. Provide 3 key facts."
                  research = await self.llm.acomplete(search_prompt)
                  return ResearchedEvent(query=ev.query, findings=str(research))
      
          @step
          async def research(self, ev: ResearchedEvent) -> RetrievedEvent:
              """Generate synthetic context when retrieval was insufficient."""
              from llama_index.core.schema import TextNode
              extra_node = TextNode(text=ev.findings)
              return RetrievedEvent(
                  query=ev.query,
                  documents=[extra_node],  # Loop back to evaluate
              )
      
          @step
          async def synthesize(self, ev: EvaluatedEvent) -> StopEvent:
              """Final answer synthesis."""
              context = "\n\n".join([n.get_content() for n in ev.documents])
              prompt = f"Answer the query using the context.\n\nContext:\n{context}\n\nQuery: {ev.query}"
              resp = await self.llm.acomplete(prompt)
              return StopEvent(result=str(resp))
      
      # --- Run ---
      async def main():
          wf = SmartRAG(timeout=60, verbose=True)
          result = await wf.run(query="What are the key findings in this document?")
          print(f"Result: {result}")
      
      if __name__ == "__main__":
          import asyncio
          asyncio.run(main())
      
    • production-deploy.py 930 B
      #!/usr/bin/env python3
      """
      Deploy a LlamaIndex Workflow as a production microservice using llama-deploy.
      Requires: `pip install llama-deploy`, running Redis instance.
      """
      
      import asyncio
      from llama_deploy import (
          deploy_workflow,
          WorkflowServiceConfig,
          ControlPlaneConfig,
      )
      from llama_index.core.workflow import Workflow, StartEvent, StopEvent, step
      
      
      class MyRAGWorkflow(Workflow):
          """Your workflow class — replace with your actual Workflow."""
      
          @step
          async def process(self, ev: StartEvent) -> StopEvent:
              # Your workflow logic here
              return StopEvent(result=f"Processed: {ev.query}")
      
      
      async def main():
          await deploy_workflow(
              workflow=MyRAGWorkflow(timeout=60),
              workflow_config=WorkflowServiceConfig(
                  service_name="my-rag-service",
              ),
              control_plane_config=ControlPlaneConfig(),
          )
      
      
      if __name__ == "__main__":
          asyncio.run(main())
      
  • README.md 1.7 KB
    # LlamaIndex — RAG & Agent Orchestration Framework
    
    An expert-level skill for building LLM applications over your data with LlamaIndex. RAG pipelines, multi-agent orchestration, event-driven workflows, knowledge graph construction, and production deployment.
    
    ## Why Install This Skill
    
    When your agent loads this skill, it becomes a **LlamaIndex expert** who can:
    
    - **Build production RAG pipelines** — from data ingestion to deployed query engines
    - **Create agent workflows** — AgentWorkflow for tool-using agents
    - **Construct knowledge graphs** — PropertyGraphIndex for structural path traversal
    - **Optimize retrieval** — hybrid search, reranking, metadata filters, sentence window parsing
    - **Add observability** — OpenTelemetry-native tracing with Phoenix
    - **Evaluate systematically** — span-attached evaluation with ParamTuner
    
    ## What You Get
    
    | Directory | Purpose |
    |-----------|---------|
    | `SKILL.md` | 9-phase pipeline guide, pipeline modes, quick reference |
    | `references/` | Ingest, chunk, index, retrieve, agent, workflow, deploy, evaluate — one per phase plus framework comparisons |
    
    ## Framework Comparison
    
    LlamaIndex evolved from a RAG indexing library into a full workflow framework. It differs from LangChain (broader integration ecosystem) and Haystack (declarative DAG pipelines) — LlamaIndex's unique strength is its data-aware indexing and knowledge graph construction.
    
    ## Requirements
    
    Python 3.8+ with `llama_index` package.
    
    
    ## Quick Start
    
    Start with the setup and first workflow in SKILL.md, then use the linked resources for the specific task you need to complete.
    
    
    ## Triggers
    
    Use this skill for the task types and keywords described in its SKILL.md description.
    
  • SKILL.md 11.8 KB
    ---
    name: llamaindex
    description: >-
      Build LLM applications with the LlamaIndex framework. Use when working with LlamaIndex
      or comparing RAG and agent orchestration frameworks. Do not use this skill for unrelated
      requests; route to the nearest named specialist.
    license: MIT
    metadata:
      author: Magnus Hedemark
      version: 1.3.0
      source: https://github.com/run-llama/llama_index
    ---
    
    # LlamaIndex Expert Skill
    
    LlamaIndex is an MIT-licensed Python framework for building LLM applications over your data. In 2026, it has evolved from a RAG indexing library into an event-driven workflow framework with integrated production runtime (llama-deploy), agent orchestration (AgentWorkflow), knowledge graph construction (PropertyGraphIndex), and OpenTelemetry-native observability.
    
    The framework is organized around seven core primitives: **Reader** (data loaders), **Document/Node** (chunked content model), **Index** (data structures over Nodes), **Retriever** (relevant Node selection), **Query Engine** (retriever + synthesis), **Agent** (LLM with tools), and **Workflow** (event-driven orchestration).
    
    ## Key Principles
    
    > These principles govern every decision when building with LlamaIndex. Read them before proceeding to the reference guides.
    
    1. **Decouple retrieval chunks from synthesis chunks.** The embedding representation that retrieves well differs from the context representation that generates well. Use `SentenceWindowNodeParser` + `MetadataReplacementNodePostProcessor` for this pattern.
    2. **Rerank before you generate.** Hybrid retrieval + reranker is the minimum viable production RAG configuration.
    3. **Agents are Workflows.** `FunctionAgent` and `AgentWorkflow` are pre-configured Workflows. Drop to raw `Workflow` when you need custom control flow.
    4. **Graphs are not just vector stores.** `PropertyGraphIndex` adds structural path traversal that vector similarity cannot provide — combine both for maximum retrieval quality.
    5. **Evaluate in the same process.** Span-attached evaluation preserves the connection between the output and the retrieval context that produced it.
    
    ## Where to Start
    
    The pipeline has 9 phases from Ingest to Deploy. If you're joining mid-stream with existing work, find your entry point:
    
    | You already have... | Start at phase | What to do |
    |---|---|---|
    | Nothing — blank project | **Ingest** | Set up data loading, then proceed through the full pipeline |
    | Documents in a directory | **Chunk** | Choose a chunking strategy, build your index |
    | A working vector index | **Retrieve** | Add hybrid search, reranking, metadata filters |
    | An existing RAG pipeline to harden | **Deploy** | Add observability, llama-deploy, production debugging |
    | A need to measure and improve quality | **Evaluate** | Set up evaluators, ParamTuner, span-attached scoring |
    | Nothing — comparing frameworks | See Framework Routing Guide | Don't start the pipeline — pick the right tool first |
    
    ## Pipeline Mode
    
    Different tasks need different levels of rigor. Match your scope to a mode:
    
    | Mode | When | Phases to run | Skip |
    |------|------|---------------|------|
    | **Quick** | Single query, one source, exploration | Ingest → Chunk → Index → Retrieve | Reranking, metadata filters, observability, evaluation |
    | **Full** | Production RAG, multiple sources, compliance | Ingest → Chunk → Index → Retrieve → Agent/Workflow → Deploy → Evaluate | Nothing — run all phases |
    | **Evaluate** | Benchmarking, regression testing | Ingest → Chunk → Index → Evaluate | Retrieve, Agent, Workflow, Deploy (run offline) |
    | **Graph** | Knowledge graph construction | Ingest → Chunk → Graph → Retrieve | Agent, Workflow, Deploy (query via graph index directly) |
    
    Rule of thumb: if you're shipping to users, run Full mode. If you're exploring, run Quick. If you're measuring, run Evaluate.
    
    ## Quick Reference
    
    | Phase | Task | Approach | Reference |
    |-------|------|----------|-----------|
    | Ingest | Load data | `SimpleDirectoryReader("./data").load_data()` | `references/architecture.md` |
    | Chunk | Parse documents into nodes | `SentenceSplitter(chunk_size=1024)` | `references/rag-strategies.md` |
    | Index | Build vector index | `VectorStoreIndex.from_documents(docs)` | `references/rag-strategies.md` |
    | Retrieve | Hybrid search + rerank | `BM25Retriever` + `CohereRerank` | `references/rag-strategies.md` |
    | Agent | Multi-agent orchestration | `AgentWorkflow(agents=[...])` | `references/agent-patterns.md` |
    | Workflow | Event-driven pipeline | `class MyFlow(Workflow): @step` | `references/workflows.md` |
    | Graph | Knowledge graph | `PropertyGraphIndex.from_documents(docs)` | `references/property-graph-index.md` |
    | Evaluate | RAG evaluation | `FaithfulnessEvaluator().evaluate_response(...)` | `references/evaluation-observability.md` |
    | Deploy | Production deployment | `deploy_workflow(workflow=MyFlow())` | `references/production-deployment.md` |
    
    ## When to Use This Skill
    
    Load this skill any time you are:
    - Building a RAG pipeline over enterprise or personal data
    - Comparing LlamaIndex with LangChain, Haystack, or DSPy
    - Designing multi-agent systems with handoff between specialist agents
    - Deploying an LLM application to production with observability
    - Constructing knowledge graphs from unstructured documents
    - Debugging common LlamaIndex failures (retrieval miss, handoff bug, async issues)
    
    ## When NOT to Use LlamaIndex — Framework Routing Guide
    
    This skill is part of a portfolio of framework skills. When deciding which framework fits, use this routing table:
    
    | Scenario | Reach for | Why |
    |----------|-----------|-----|
    | I have documents I need to query | **LlamaIndex** | Data ingestion, hybrid retrieval, reranking, and knowledge graphs are first-class primitives |
    | I have agents I need to orchestrate | **LangGraph** | State-machine semantics, time-travel debugging, and human-in-the-loop pauses are the core design |
    | I have a tool I need to wrap as an agent | **PydanticAI** | Type-safe agent definitions with dependency injection, minimal abstraction over LLM calls |
    | Data-heavy RAG over PDFs, SQL, Slack, 200+ sources | **LlamaIndex** | LlamaHub connectors, LlamaParse for documents, hybrid retrieval out of the box |
    | Complex multi-agent state machines with checkpoints | **LangGraph** | Graph topology control — supervisor, subgraphs, hierarchical teams, built-in checkpointer |
    | Agent-centric app where type safety matters more than data pipelines | **PydanticAI** | Agents as Pydantic models, DI, structured outputs — the data layer is your code |
    | Document parsing quality matters (tables, charts, handwriting) | **LlamaIndex** | LlamaParse is purpose-built for this |
    | Production NLP search pipelines | **Haystack** | Pipeline composition model is more mature for search-specific workloads |
    | Optimization-driven prompt programming | **DSPy** | Compiled prompt programs, not retrieval pipelines |
    
    ## Reference Files
    
    | Reference | Load when | File |
    |-----------|-----------|------|
    | Core Architecture | Understanding the 7 primitives, Settings, data flow | `references/architecture.md` |
    | RAG Strategies | Building RAG pipelines from basic to advanced | `references/rag-strategies.md` |
    | Agent Patterns | Multi-agent orchestration with AgentWorkflow | `references/agent-patterns.md` |
    | Workflows | Event-driven step composition and durable execution | `references/workflows.md` |
    | Production & Deployment | llama-deploy, debugging, failure modes | `references/production-deployment.md` |
    | Property Graph Index | Knowledge graph construction and hybrid retrieval | `references/property-graph-index.md` |
    | Evaluation & Observability | Metrics, tracing, span-attached scoring | `references/evaluation-observability.md` |
    | Integration Ecosystem | Vector stores, LlamaHub, LlamaParse, ecosystem | `references/integration-ecosystem.md` |
    | FAQ & Troubleshooting | Common errors and their fixes | `references/faq-and-troubleshooting.md` |
    | Worked RAG Example | Complete end-to-end pipeline from ingest to deploy | `references/example-rag-pipeline.md` |
    | Evaluation Workflow | ParamTuner, evaluators, batch scoring, best practices | `references/evaluation-workflow.md` |
    
    ## Template Files
    
    | Template | When to use | File |
    |----------|-------------|------|
    | Basic RAG | Single-source query, getting started | `templates/basic-rag.py` |
    | Agentic RAG | Multi-source data with agent routing | `templates/agentic-rag.py` |
    | Custom Workflow | Custom control flow, branching logic | `templates/custom-workflow.py` |
    | Production Deploy | Wrapping a workflow as a microservice | `templates/production-deploy.py` |
    
    ## Scripts
    
    | Script | Purpose | File |
    |--------|---------|------|
    | check-setup | Verify LlamaIndex installation and configuration | `scripts/check-setup.py` |
    
    ## Troubleshooting — Structured Recovery Guide
    
    When something goes wrong, find your symptom and follow the recovery path:
    
    ### Retrieval & Answer Quality
    
    | Symptom | Likely cause | Immediate fix | Permanent fix | Reference |
    |---------|-------------|---------------|---------------|-----------|
    | Answers are poor or hallucinated | No reranker on hybrid retrieval | Add `CohereRerank(top_n=5)` as `node_postprocessor` | Reranking is mandatory for any production RAG | `references/rag-strategies.md` |
    | Retrieval misses obvious content | Default chunking breaks semantics | Switch to `SemanticSplitterNodeParser(breakpoint_percentile_threshold=95)` | Tune chunk size with ParamTuner | `references/rag-strategies.md` |
    | Wrong tenant's data returned | Missing metadata filters | Add `MetadataFilters(filters=[ExactMatchFilter(key="tenant_id", ...)])` | Always wire metadata filters at retriever level | `references/rag-strategies.md` |
    | Only one type of query works well | Single retrieval strategy | Combine BM25 + vector via hybrid retriever | Add RouterQueryEngine for query-type routing | `references/rag-strategies.md` |
    
    ### Agent & Workflow Failures
    
    | Symptom | Likely cause | Immediate fix | Permanent fix | Reference |
    |---------|-------------|---------------|---------------|-----------|
    | Agent waits silently after handoff | AgentWorkflow handoff bug | Extend `FunctionAgent.take_step` to re-locate last user message | Apply the handoff fix on all production agents | `references/agent-patterns.md` |
    | Workflow doesn't run | Forgot `await` | Add `await` before `w.run(...)` and all step calls | All step methods are async coroutines | `references/workflows.md` |
    | Step executes but result is lost | State not persisted | Use `ctx.store.edit_state()` for shared state | Only `ctx.store` survives across steps | `references/workflows.md` |
    | Crash loses all progress | No checkpoint snapshots | Add `Context.to_dict()` save on step completion | Durable workflows need explicit checkpointing | `references/workflows.md` |
    
    ### Deployment & Observability
    
    | Symptom | Likely cause | Immediate fix | Permanent fix | Reference |
    |---------|-------------|---------------|---------------|-----------|
    | llama-deploy deployed but requests time out | Redis not running | Start `redis-server` | Redis is mandatory — control plane won't route without it | `references/production-deployment.md` |
    | Spans missing in observability UI | Instrumentation called too late | Move `instrument()` call before workflow instantiation | Always instrument before creating any Workflow object | `references/production-deployment.md` |
    | Wrong data returned (cross-tenant) | Missing metadata filters | Add tenant filter to all retrievers | Filter at retriever level, not in post-processing | `references/production-deployment.md` |
    
    ### Recovery Workflow
    
    For any failure, follow this cycle:
    1. **Identify the symptom** from the tables above
    2. **Apply the immediate fix** — this gets you running
    3. **Implement the permanent fix** — this prevents recurrence
    4. **Verify with evaluation** — run `FaithfulnessEvaluator` on a held-out query set
    5. **Document the fix** — add the root cause to `references/faq-and-troubleshooting.md`
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related