Claude Skill

ai-observability-langfuse

LLM observability with Langfuse — OpenTelemetry-based tracing, evaluations, prompt management, datasets, and production best practices

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download agents-inc-skills-dist_plugins_ai-observability-langfuse_skills_ai-observability-langfuse-3a51ef5.zip · 21 KB
Part of agents-inc/skills — 130 skills

Install

skills CLI npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/ai-observability-langfuse/skills/ai-observability-langfuse
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
Git git clone https://github.com/agents-inc/skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Langfuse Observability Patterns

Quick Guide: Use the Langfuse TypeScript SDK (built on OpenTelemetry) to add observability to LLM applications. Install @langfuse/tracing, @langfuse/otel, and @opentelemetry/sdk-node for core tracing. Use startActiveObservation() for automatic context propagation or observe() to wrap functions. Use @langfuse/openai with observeOpenAI() for zero-config OpenAI tracing. Use LangfuseClient from @langfuse/client for prompt management, scores, and datasets. Always call forceFlush() or sdk.shutdown() in short-lived processes.


<critical_requirements>

CRITICAL: Before Using This Skill

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST import and register instrumentation.ts at the top of your entry point BEFORE any other imports -- OpenTelemetry must instrument modules before they are loaded)

(You MUST call forceFlush() or sdk.shutdown() in short-lived processes (serverless, scripts, CLI tools) -- events are batched and will be lost without explicit flushing)

(You MUST use @langfuse/openai with observeOpenAI() for OpenAI SDK tracing -- do NOT manually create generation observations for OpenAI calls when the wrapper handles it automatically)

(You MUST set LANGFUSE_SECRET_KEY, LANGFUSE_PUBLIC_KEY, and LANGFUSE_BASE_URL via environment variables -- never hardcode credentials)

(You MUST use startActiveObservation() or observe() for nested tracing -- manual startObservation() requires explicit .end() calls and does NOT propagate context automatically)

</critical_requirements>


Auto-detection: Langfuse, langfuse, @langfuse/tracing, @langfuse/otel, @langfuse/client, @langfuse/openai, LangfuseSpanProcessor, LangfuseClient, startActiveObservation, startObservation, observeOpenAI, langfuse.score, langfuse.prompt, langfuse.dataset, LANGFUSE_SECRET_KEY, LANGFUSE_PUBLIC_KEY, forceFlush

When to use:

  • Adding observability and tracing to LLM application code (any provider)
  • Wrapping OpenAI SDK calls for automatic token/cost tracking
  • Managing prompt templates with versioning, labels, and variable compilation
  • Evaluating LLM output quality with scores (numeric, categorical, boolean)
  • Running experiments against datasets for regression testing
  • Tracking sessions, users, and metadata across multi-turn conversations
  • Monitoring LLM costs and token usage in production

Key patterns covered:

  • OpenTelemetry setup with LangfuseSpanProcessor
  • Tracing with startActiveObservation, observe, and manual startObservation
  • Observation types (span, generation, agent, tool, retriever, evaluator, embedding, chain, guardrail)
  • OpenAI SDK auto-instrumentation with observeOpenAI()
  • Prompt management (get, compile, text vs chat prompts, versioning)
  • Scores and evaluations (numeric, categorical, boolean)
  • Datasets and experiments for testing
  • Flush, shutdown, and lifecycle management

When NOT to use:

  • You only need basic console.log debugging -- Langfuse is for structured production observability
  • You want provider-specific tracing built into an AI SDK -- check if your framework has native observability
  • You need APM/infrastructure monitoring (CPU, memory, HTTP latency) -- use a general-purpose observability tool

Examples Index

  • Core: Setup & Configuration -- OpenTelemetry setup, instrumentation file, client init, flush/shutdown
  • Tracing -- startActiveObservation, observe, manual tracing, nesting, observation types, metadata
  • OpenAI Integration -- observeOpenAI wrapper, streaming, token tracking, custom attributes
  • Prompt Management -- getPrompt, compile, text vs chat, versioning, caching
  • Scores & Datasets -- Numeric/categorical/boolean scores, datasets, experiments
  • Quick API Reference -- Package index, environment variables, observation types, score methods




<decision_framework>

Decision Framework

Which Packages to Install

What do you need?
+-- Tracing LLM calls?
|   +-- YES -> npm install @langfuse/tracing @langfuse/otel @opentelemetry/sdk-node
|   +-- Also using OpenAI SDK?
|       +-- YES -> npm install @langfuse/openai
+-- Prompt management, scores, or datasets?
|   +-- YES -> npm install @langfuse/client
+-- Both tracing AND client features?
    +-- YES -> Install all: @langfuse/tracing @langfuse/otel @opentelemetry/sdk-node @langfuse/client

Which Tracing Method to Use

How do you want to instrument?
+-- Wrapping a function? -> observe() (declarative, auto-captures inputs/outputs)
+-- Block of code with nesting? -> startActiveObservation() (context propagation, auto-end)
+-- Need manual start/end control? -> startObservation() (requires explicit .end())
+-- OpenAI SDK calls? -> observeOpenAI() (zero-config auto-instrumentation)
+-- Update active span without reference? -> updateActiveObservation()

Which Observation Type (asType)

What is this observation?
+-- LLM call (prompt -> completion) -> "generation"
+-- AI agent decision-making step -> "agent"
+-- External API or function call -> "tool"
+-- Vector store or DB retrieval -> "retriever"
+-- Quality assessment step -> "evaluator"
+-- Embedding creation -> "embedding"
+-- Link between application steps -> "chain"
+-- Content safety / jailbreak check -> "guardrail"
+-- Generic duration operation -> "span" (default)
+-- Point-in-time event -> "event"

</decision_framework>


<red_flags>

RED FLAGS

High Priority Issues:

  • Not importing instrumentation.ts before other modules (auto-instrumentation silently fails)
  • Exiting short-lived processes without forceFlush() or sdk.shutdown() (events are silently lost)
  • Hardcoding LANGFUSE_SECRET_KEY or LANGFUSE_PUBLIC_KEY in source code (credential exposure)
  • Using manual generation observations when observeOpenAI() would handle it automatically (duplicated effort, less accurate data)
  • Using startObservation() without calling .end() (observation stays open indefinitely)

Medium Priority Issues:

  • Not setting stream_options: { include_usage: true } on OpenAI streaming calls (token counts missing from observeOpenAI traces)
  • Forgetting to call langfuse.score.flush() in short-lived processes (scores are batched and may be lost)
  • Using startObservation() when startActiveObservation() would work (no automatic context propagation or auto-end)
  • Not using asType on observations (all observations appear as generic spans, losing semantic meaning)
  • Not setting LANGFUSE_BASE_URL for self-hosted instances (defaults to cloud.langfuse.com)

Common Mistakes:

  • Importing @langfuse/openai without setting up the OTel NodeSDK first -- the OpenAI wrapper requires OTel context to send traces
  • Confusing LangfuseClient (from @langfuse/client, for prompts/scores/datasets) with the OTel tracing functions (from @langfuse/tracing)
  • Using prompt.compile() without matching all {{variable}} placeholders -- unmatched variables remain as literal {{name}} in output
  • Calling langfuse.score.create() with a value of type string for NUMERIC scores or number for CATEGORICAL scores (type mismatch)
  • Running dataset experiments without OTel setup -- experiment tasks run inside startActiveObservation which requires OTel

Gotchas & Edge Cases:

  • observeOpenAI() does NOT support the OpenAI Assistants API -- only Chat Completions and Responses API
  • The SDK's default span filter only exports Langfuse and GenAI spans. If you use a custom instrumentation library, you must configure shouldExportSpan to include it.
  • LangfuseClient.prompt.get() caches prompts with a default TTL. If you update a prompt and don't see changes, set cacheTtlSeconds: 0 to bypass caching.
  • Boolean scores use float values (0 or 1), not JavaScript booleans (true/false).
  • Self-hosted Langfuse requires platform version >= 3.95.0 for TypeScript SDK v4 compatibility.
  • score.create() is fire-and-forget (synchronous) -- it queues the score for batched delivery. You only need await on flush().
  • Dataset names with slashes (evaluation/qa-dataset) must be URL-encoded when used as path parameters.
  • The v4+ SDK is a complete rewrite from v3 -- Langfuse class, trace(), span(), generation() from v3 are replaced by OTel-based APIs.

</red_flags>


<critical_reminders>

CRITICAL REMINDERS

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST import and register instrumentation.ts at the top of your entry point BEFORE any other imports -- OpenTelemetry must instrument modules before they are loaded)

(You MUST call forceFlush() or sdk.shutdown() in short-lived processes (serverless, scripts, CLI tools) -- events are batched and will be lost without explicit flushing)

(You MUST use @langfuse/openai with observeOpenAI() for OpenAI SDK tracing -- do NOT manually create generation observations for OpenAI calls when the wrapper handles it automatically)

(You MUST set LANGFUSE_SECRET_KEY, LANGFUSE_PUBLIC_KEY, and LANGFUSE_BASE_URL via environment variables -- never hardcode credentials)

(You MUST use startActiveObservation() or observe() for nested tracing -- manual startObservation() requires explicit .end() calls and does NOT propagate context automatically)

Failure to follow these rules will produce silent data loss, missing traces, or credential exposure in LLM observability.

</critical_reminders>

Files (skills)
  • examples
    • core.md 6.3 KB
      # Langfuse -- Setup & Configuration Examples
      
      > OpenTelemetry initialization, environment config, sampling, data masking, multi-project setup, and lifecycle management. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [tracing.md](tracing.md) -- Tracing with observations, nesting, metadata
      - [openai-integration.md](openai-integration.md) -- OpenAI SDK auto-instrumentation
      - [prompt-management.md](prompt-management.md) -- Prompt management and versioning
      - [scores-datasets.md](scores-datasets.md) -- Scores, datasets, experiments
      
      ---
      
      ## Basic Setup
      
      ```typescript
      // instrumentation.ts
      import { NodeSDK } from "@opentelemetry/sdk-node";
      import { LangfuseSpanProcessor } from "@langfuse/otel";
      
      const sdk = new NodeSDK({
        spanProcessors: [new LangfuseSpanProcessor()],
      });
      
      sdk.start();
      
      export { sdk };
      ```
      
      ```bash
      # .env
      LANGFUSE_SECRET_KEY="sk-lf-..."
      LANGFUSE_PUBLIC_KEY="pk-lf-..."
      LANGFUSE_BASE_URL="https://cloud.langfuse.com"   # EU region
      # LANGFUSE_BASE_URL="https://us.cloud.langfuse.com"  # US region
      ```
      
      ```typescript
      // index.ts -- instrumentation MUST be imported first
      import { sdk } from "./instrumentation";
      import { startActiveObservation } from "@langfuse/tracing";
      
      async function main() {
        await startActiveObservation("my-trace", async (span) => {
          span.update({
            input: "Hello, Langfuse!",
            output: "Trace captured successfully.",
          });
        });
      }
      
      main().finally(() => sdk.shutdown());
      ```
      
      ---
      
      ## Client Initialization
      
      ```typescript
      // lib/langfuse.ts
      import { LangfuseClient } from "@langfuse/client";
      
      // Reads credentials from environment variables automatically
      const langfuse = new LangfuseClient();
      
      export { langfuse };
      ```
      
      ```typescript
      // lib/langfuse.ts -- explicit configuration
      import { LangfuseClient } from "@langfuse/client";
      
      const TIMEOUT_MS = 10_000;
      
      const langfuse = new LangfuseClient({
        publicKey: process.env.LANGFUSE_PUBLIC_KEY,
        secretKey: process.env.LANGFUSE_SECRET_KEY,
        baseUrl: process.env.LANGFUSE_BASE_URL,
        timeout: TIMEOUT_MS,
      });
      
      export { langfuse };
      ```
      
      ---
      
      ## Sampling for High-Volume Applications
      
      ```typescript
      // instrumentation.ts -- sample 20% of traces
      import { NodeSDK } from "@opentelemetry/sdk-node";
      import { TraceIdRatioBasedSampler } from "@opentelemetry/sdk-trace-base";
      import { LangfuseSpanProcessor } from "@langfuse/otel";
      
      const SAMPLE_RATE = 0.2;
      
      const sdk = new NodeSDK({
        sampler: new TraceIdRatioBasedSampler(SAMPLE_RATE),
        spanProcessors: [new LangfuseSpanProcessor()],
      });
      
      sdk.start();
      
      export { sdk };
      ```
      
      Or via environment variable: `LANGFUSE_SAMPLE_RATE=0.2`
      
      ---
      
      ## Data Masking
      
      ```typescript
      // instrumentation.ts -- redact PII before sending to Langfuse
      import { NodeSDK } from "@opentelemetry/sdk-node";
      import { LangfuseSpanProcessor } from "@langfuse/otel";
      
      const CREDIT_CARD_PATTERN = /\b\d{4}[- ]?\d{4}[- ]?\d{4}[- ]?\d{4}\b/g;
      const EMAIL_PATTERN = /\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b/g;
      
      const processor = new LangfuseSpanProcessor({
        mask: ({ data }) =>
          data
            .replace(CREDIT_CARD_PATTERN, "***MASKED_CC***")
            .replace(EMAIL_PATTERN, "***MASKED_EMAIL***"),
      });
      
      const sdk = new NodeSDK({
        spanProcessors: [processor],
      });
      
      sdk.start();
      
      export { sdk };
      ```
      
      ---
      
      ## Span Filtering
      
      ```typescript
      // instrumentation.ts -- only export Langfuse + custom spans
      import { NodeSDK } from "@opentelemetry/sdk-node";
      import { LangfuseSpanProcessor, isDefaultExportSpan } from "@langfuse/otel";
      
      const processor = new LangfuseSpanProcessor({
        shouldExportSpan: ({ otelSpan }) =>
          isDefaultExportSpan(otelSpan) ||
          otelSpan.instrumentationScope.name.startsWith("my-app"),
      });
      
      const sdk = new NodeSDK({
        spanProcessors: [processor],
      });
      
      sdk.start();
      
      export { sdk };
      ```
      
      ---
      
      ## Isolated TracerProvider
      
      Separate Langfuse tracing from other observability backends:
      
      ```typescript
      // instrumentation.ts -- isolated provider
      import { NodeTracerProvider } from "@opentelemetry/sdk-trace-node";
      import { LangfuseSpanProcessor } from "@langfuse/otel";
      import { setLangfuseTracerProvider } from "@langfuse/tracing";
      
      const langfuseProvider = new NodeTracerProvider({
        spanProcessors: [new LangfuseSpanProcessor()],
      });
      
      setLangfuseTracerProvider(langfuseProvider);
      ```
      
      ---
      
      ## Multi-Project Setup
      
      ```typescript
      // instrumentation.ts -- send spans to multiple Langfuse projects
      import { NodeSDK } from "@opentelemetry/sdk-node";
      import { LangfuseSpanProcessor } from "@langfuse/otel";
      
      const sdk = new NodeSDK({
        spanProcessors: [
          new LangfuseSpanProcessor({
            publicKey: "pk-lf-project-1",
            secretKey: "sk-lf-project-1",
          }),
          new LangfuseSpanProcessor({
            publicKey: "pk-lf-project-2",
            secretKey: "sk-lf-project-2",
          }),
        ],
      });
      
      sdk.start();
      
      export { sdk };
      ```
      
      ---
      
      ## Debug Logging
      
      ```bash
      # Via environment variable
      LANGFUSE_LOG_LEVEL=DEBUG
      
      # Or via code
      ```
      
      ```typescript
      import { configureGlobalLogger } from "@langfuse/core";
      
      configureGlobalLogger({ level: "DEBUG" });
      ```
      
      ---
      
      ## Serverless / Short-Lived Process Pattern
      
      For serverless environments, use `exportMode: "immediate"` on the span processor to export spans immediately rather than batching:
      
      ```typescript
      // instrumentation.ts (serverless variant)
      import { NodeSDK } from "@opentelemetry/sdk-node";
      import { LangfuseSpanProcessor } from "@langfuse/otel";
      
      export const langfuseSpanProcessor = new LangfuseSpanProcessor({
        exportMode: "immediate",
      });
      
      export const sdk = new NodeSDK({
        spanProcessors: [langfuseSpanProcessor],
      });
      
      sdk.start();
      ```
      
      ```typescript
      import { sdk } from "./instrumentation";
      import { startActiveObservation } from "@langfuse/tracing";
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      
      export async function handler(event: unknown) {
        try {
          return await startActiveObservation("lambda-handler", async (span) => {
            span.update({ input: event });
            const result = await processEvent(event);
            span.update({ output: result });
      
            // Score the result
            langfuse.score.activeTrace({
              name: "success",
              value: 1,
              dataType: "BOOLEAN",
            });
      
            return result;
          });
        } finally {
          // CRITICAL: flush before Lambda freezes
          await langfuse.score.flush();
          await sdk.shutdown();
        }
      }
      ```
      
      ---
      
      _For tracing patterns, see [tracing.md](tracing.md). For API reference, see [reference.md](../reference.md)._
      
    • openai-integration.md 6 KB
      # Langfuse -- OpenAI Integration Examples
      
      > Auto-instrumentation with `observeOpenAI()`, streaming token tracking, custom attributes, and nested tracing. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Setup, environment config, flush/shutdown
      - [tracing.md](tracing.md) -- Manual tracing patterns
      - [prompt-management.md](prompt-management.md) -- Prompt management
      - [scores-datasets.md](scores-datasets.md) -- Scores and datasets
      
      ---
      
      ## Basic OpenAI Wrapper
      
      ```typescript
      // lib/openai.ts
      import OpenAI from "openai";
      import { observeOpenAI } from "@langfuse/openai";
      
      // Wrap the OpenAI client -- all calls automatically traced
      const openai = observeOpenAI(new OpenAI());
      
      export { openai };
      ```
      
      ```typescript
      // usage.ts
      import { openai } from "./lib/openai";
      
      const completion = await openai.chat.completions.create({
        model: "gpt-4o",
        messages: [
          { role: "developer", content: "You are a helpful assistant." },
          { role: "user", content: "Explain async/await in TypeScript." },
        ],
      });
      
      console.log(completion.choices[0].message.content);
      // Langfuse automatically captures: model, tokens, cost, latency
      ```
      
      ---
      
      ## Custom Trace Attributes
      
      Pass metadata to control how the generation appears in Langfuse.
      
      ```typescript
      import OpenAI from "openai";
      import { observeOpenAI } from "@langfuse/openai";
      
      const openai = observeOpenAI(new OpenAI(), {
        generationName: "intent-classifier",
        sessionId: "session-abc-123",
        userId: "user-456",
        tags: ["production", "classifier-v2"],
        generationMetadata: {
          feature: "intent-classification",
          version: "2.1",
        },
      });
      
      const result = await openai.chat.completions.create({
        model: "gpt-4o",
        messages: [{ role: "user", content: "I want to book a flight" }],
      });
      ```
      
      ---
      
      ## Streaming with Token Tracking
      
      For streaming calls, set `stream_options.include_usage` to capture token counts.
      
      ```typescript
      import { openai } from "./lib/openai";
      
      const stream = await openai.chat.completions.create({
        model: "gpt-4o",
        messages: [{ role: "user", content: "Tell me a story." }],
        stream: true,
        stream_options: { include_usage: true }, // REQUIRED for token tracking on streams
      });
      
      for await (const chunk of stream) {
        const content = chunk.choices[0]?.delta?.content;
        if (content) process.stdout.write(content);
      }
      // Token usage automatically captured by observeOpenAI
      ```
      
      **Without `stream_options: { include_usage: true }`**, OpenAI does not return token counts for streaming responses, and Langfuse will show `null` for token usage.
      
      ---
      
      ## Nested Inside Manual Traces
      
      Combine `observeOpenAI` with manual tracing for richer context.
      
      ```typescript
      import OpenAI from "openai";
      import { observeOpenAI } from "@langfuse/openai";
      import { startActiveObservation } from "@langfuse/tracing";
      
      const openai = observeOpenAI(new OpenAI());
      
      async function summarizeArticle(article: string): Promise<string> {
        return await startActiveObservation("summarize-article", async (span) => {
          span.update({ input: { articleLength: article.length } });
      
          // OpenAI call is automatically a child of "summarize-article"
          const completion = await openai.chat.completions.create({
            model: "gpt-4o",
            messages: [
              {
                role: "developer",
                content: "Summarize the following article concisely.",
              },
              { role: "user", content: article },
            ],
          });
      
          const summary = completion.choices[0].message.content ?? "";
          span.update({ output: { summary, summaryLength: summary.length } });
          return summary;
        });
      }
      ```
      
      ---
      
      ## Multi-Step Pipeline with OpenAI
      
      ```typescript
      import OpenAI from "openai";
      import { observeOpenAI } from "@langfuse/openai";
      import { startActiveObservation, propagateAttributes } from "@langfuse/tracing";
      
      const openai = observeOpenAI(new OpenAI());
      
      async function chatPipeline(
        userId: string,
        sessionId: string,
        message: string,
      ): Promise<string> {
        return await propagateAttributes(
          { userId, sessionId, tags: ["chat-v3"] },
          async () => {
            return await startActiveObservation("chat-pipeline", async (rootSpan) => {
              rootSpan.update({ input: { message } });
      
              // Step 1: Classify intent
              const intentResult = await startActiveObservation(
                "classify-intent",
                async () => {
                  const completion = await openai.chat.completions.create({
                    model: "gpt-4o",
                    messages: [
                      {
                        role: "developer",
                        content:
                          "Classify the user intent as: question, command, or greeting.",
                      },
                      { role: "user", content: message },
                    ],
                    temperature: 0,
                  });
                  return completion.choices[0].message.content ?? "unknown";
                },
                { asType: "agent" },
              );
      
              // Step 2: Generate response based on intent
              const response = await startActiveObservation(
                "generate-response",
                async () => {
                  const completion = await openai.chat.completions.create({
                    model: "gpt-4o",
                    messages: [
                      {
                        role: "developer",
                        content: `User intent: ${intentResult}. Respond appropriately.`,
                      },
                      { role: "user", content: message },
                    ],
                  });
                  return completion.choices[0].message.content ?? "";
                },
                { asType: "generation" },
              );
      
              rootSpan.update({ output: { intent: intentResult, response } });
              return response;
            });
          },
        );
      }
      ```
      
      ---
      
      ## Limitations
      
      - `observeOpenAI()` does **NOT** support the OpenAI Assistants API (server-side state is incompatible)
      - Requires OpenAI SDK version >= 4.0.0
      - Requires OTel `NodeSDK` to be initialized before using the wrapper
      - Token counts on streaming calls require explicit `stream_options: { include_usage: true }`
      
      ---
      
      _For tracing patterns, see [tracing.md](tracing.md). For setup, see [core.md](core.md)._
      
    • prompt-management.md 5.3 KB
      # Langfuse -- Prompt Management Examples
      
      > Fetching prompts, compiling variables, text vs chat prompts, versioning, labels, caching, and linking prompts to traces. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Setup, environment config, flush/shutdown
      - [tracing.md](tracing.md) -- Manual tracing patterns
      - [openai-integration.md](openai-integration.md) -- OpenAI auto-instrumentation
      - [scores-datasets.md](scores-datasets.md) -- Scores and datasets
      
      ---
      
      ## Fetching and Compiling Text Prompts
      
      ```typescript
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      
      // Fetch a text prompt (production label by default)
      const prompt = await langfuse.prompt.get("summarize-article");
      
      // Compile with variables -- replaces {{variable}} placeholders
      const compiled = prompt.compile({
        topic: "climate change",
        length: "200 words",
      });
      // -> "Write a 200 words summary about climate change."
      
      // Use compiled prompt with your LLM
      const completion = await openai.chat.completions.create({
        model: "gpt-4o",
        messages: [
          { role: "developer", content: compiled },
          { role: "user", content: articleText },
        ],
      });
      ```
      
      ---
      
      ## Chat Prompts
      
      Chat prompts return an array of message objects ready for LLM APIs.
      
      ```typescript
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      
      // Fetch a chat prompt
      const chatPrompt = await langfuse.prompt.get("customer-support", {
        type: "chat",
      });
      
      // Compile with variables
      const messages = chatPrompt.compile({
        userName: "Alice",
        productName: "Pro Plan",
      });
      // -> [
      //   { role: "system", content: "You are helping Alice with their Pro Plan." },
      //   { role: "user", content: "{{user_message}}" }
      // ]
      
      // Use directly with OpenAI
      const completion = await openai.chat.completions.create({
        model: "gpt-4o",
        messages: [
          ...messages,
          { role: "user", content: "How do I cancel my subscription?" },
        ],
      });
      ```
      
      ---
      
      ## Prompt Versioning and Labels
      
      ```typescript
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      
      // Default: fetches "production" label
      const prodPrompt = await langfuse.prompt.get("classifier");
      
      // Fetch a specific label (for A/B testing or staging)
      const stagingPrompt = await langfuse.prompt.get("classifier", {
        label: "staging",
      });
      
      // Fetch a specific version number
      const v2Prompt = await langfuse.prompt.get("classifier", {
        version: 2,
      });
      ```
      
      ---
      
      ## Cache Control
      
      The SDK caches prompts by default. Control caching behavior for development or frequently updated prompts.
      
      ```typescript
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      
      // Bypass cache -- always fetch from server
      const freshPrompt = await langfuse.prompt.get("classifier", {
        cacheTtlSeconds: 0,
      });
      
      // Custom cache TTL (in seconds)
      const CACHE_TTL_SECONDS = 300;
      const cachedPrompt = await langfuse.prompt.get("classifier", {
        cacheTtlSeconds: CACHE_TTL_SECONDS,
      });
      ```
      
      **Note:** When a cached prompt expires, the SDK returns the stale prompt immediately while refreshing in the background. This prevents latency spikes from cache misses.
      
      ---
      
      ## Creating Prompts Programmatically
      
      ```typescript
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      
      // Create a text prompt
      await langfuse.prompt.create({
        name: "summarize-article",
        prompt: "Write a {{length}} summary about {{topic}}. Focus on key facts.",
        type: "text",
      });
      
      // Create a chat prompt
      await langfuse.prompt.create({
        name: "customer-support",
        prompt: [
          {
            role: "system",
            content: "You are a support agent helping {{userName}}.",
          },
          { role: "user", content: "{{user_message}}" },
        ],
        type: "chat",
      });
      ```
      
      ---
      
      ## Linking Prompts to Traces
      
      When using `observeOpenAI`, pass the `langfusePrompt` option to link a prompt version to the generation trace.
      
      ```typescript
      import OpenAI from "openai";
      import { observeOpenAI } from "@langfuse/openai";
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      const openai = observeOpenAI(new OpenAI());
      
      async function classifyWithManagedPrompt(text: string): Promise<string> {
        const prompt = await langfuse.prompt.get("classifier");
        const compiled = prompt.compile({ text });
      
        // Link prompt version to the trace for tracking
        const trackedOpenai = observeOpenAI(new OpenAI(), {
          langfusePrompt: prompt,
          generationName: "classify-text",
        });
      
        const completion = await trackedOpenai.chat.completions.create({
          model: "gpt-4o",
          messages: [{ role: "user", content: compiled }],
        });
      
        return completion.choices[0].message.content ?? "";
      }
      ```
      
      ---
      
      ## Prompt with Fallback
      
      For guaranteed availability, provide a fallback prompt in case the Langfuse API is unreachable.
      
      ```typescript
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      
      const FALLBACK_PROMPT = "Summarize the following text concisely: {{text}}";
      
      const prompt = await langfuse.prompt.get("summarizer", {
        fallback: FALLBACK_PROMPT,
      });
      
      // If Langfuse API is down, compile() uses the fallback template
      const compiled = prompt.compile({ text: articleContent });
      ```
      
      ---
      
      _For tracing patterns, see [tracing.md](tracing.md). For setup, see [core.md](core.md)._
      
    • scores-datasets.md 7.8 KB
      # Langfuse -- Scores & Datasets Examples
      
      > Creating scores (numeric, categorical, boolean), attaching to traces/observations, datasets, experiments, and evaluation workflows. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Setup, environment config, flush/shutdown
      - [tracing.md](tracing.md) -- Manual tracing patterns
      - [openai-integration.md](openai-integration.md) -- OpenAI auto-instrumentation
      - [prompt-management.md](prompt-management.md) -- Prompt management
      
      ---
      
      ## Score Types
      
      ```typescript
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      
      // Numeric score (float value)
      langfuse.score.create({
        traceId: "trace-abc-123",
        name: "relevance",
        value: 0.92,
        dataType: "NUMERIC",
      });
      
      // Categorical score (string value)
      langfuse.score.create({
        traceId: "trace-abc-123",
        name: "tone",
        value: "professional",
        dataType: "CATEGORICAL",
      });
      
      // Boolean score (0 or 1 as float, NOT true/false)
      langfuse.score.create({
        traceId: "trace-abc-123",
        name: "contains-pii",
        value: 0,
        dataType: "BOOLEAN",
      });
      ```
      
      **Note:** Boolean scores use `0` (false) or `1` (true) as float values, not JavaScript `true`/`false`.
      
      ---
      
      ## Scoring Specific Observations
      
      ```typescript
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      
      // Score a specific generation within a trace
      langfuse.score.create({
        traceId: "trace-abc-123",
        observationId: "obs-def-456",
        name: "accuracy",
        value: 0.88,
        dataType: "NUMERIC",
      });
      ```
      
      ---
      
      ## Scoring Active Observations
      
      Score the currently active OTel span without needing trace/observation IDs.
      
      ```typescript
      import { startActiveObservation } from "@langfuse/tracing";
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      
      await startActiveObservation("process-query", async (span) => {
        span.update({ input: { query: "What is AI?" } });
      
        const result = await processQuery("What is AI?");
      
        // Score the active observation (this span)
        langfuse.score.activeObservation({
          name: "quality",
          value: 0.95,
          dataType: "NUMERIC",
        });
      
        // Score the active trace (root trace)
        langfuse.score.activeTrace({
          name: "user-satisfaction",
          value: "satisfied",
          dataType: "CATEGORICAL",
        });
      
        span.update({ output: { result } });
      });
      
      // Flush in short-lived processes
      await langfuse.score.flush();
      ```
      
      ---
      
      ## User Feedback Scores
      
      Capture user feedback (thumbs up/down, ratings) as scores.
      
      ```typescript
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      
      // User thumbs up/down
      function handleFeedback(traceId: string, isPositive: boolean) {
        langfuse.score.create({
          traceId,
          name: "user-feedback",
          value: isPositive ? 1 : 0,
          dataType: "BOOLEAN",
          comment: isPositive
            ? "User liked the response"
            : "User disliked the response",
        });
      }
      
      // User star rating (1-5)
      function handleRating(traceId: string, rating: number) {
        langfuse.score.create({
          traceId,
          name: "user-rating",
          value: rating,
          dataType: "NUMERIC",
        });
      }
      ```
      
      ---
      
      ## Session-Level Scores
      
      Score an entire session (multiple traces).
      
      ```typescript
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      
      // Score a session
      langfuse.score.create({
        sessionId: "session-xyz-789",
        name: "session-quality",
        value: "excellent",
        dataType: "CATEGORICAL",
      });
      ```
      
      ---
      
      ## Creating Datasets
      
      ```typescript
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      
      // Create a dataset
      await langfuse.api.datasets.create({
        name: "qa-benchmark",
        description: "Question-answer pairs for regression testing",
        metadata: {
          author: "Alice",
          created: "2025-10-01",
          type: "benchmark",
        },
      });
      
      // Add items to the dataset
      await langfuse.api.datasetItems.create({
        datasetName: "qa-benchmark",
        input: {
          question: "What is the capital of France?",
        },
        expectedOutput: {
          answer: "Paris",
        },
        metadata: {
          difficulty: "easy",
          category: "geography",
        },
      });
      
      await langfuse.api.datasetItems.create({
        datasetName: "qa-benchmark",
        input: {
          question: "Explain quantum entanglement in simple terms.",
        },
        expectedOutput: {
          answer:
            "Quantum entanglement is when two particles are connected so that measuring one instantly affects the other, regardless of distance.",
        },
        metadata: {
          difficulty: "hard",
          category: "physics",
        },
      });
      ```
      
      ---
      
      ## Running Experiments
      
      Run your LLM function against a dataset and collect results.
      
      ```typescript
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      
      // Get the dataset
      const dataset = await langfuse.dataset.get("qa-benchmark");
      
      // Run experiment with a task function
      const result = await dataset.runExperiment({
        name: "gpt-4o-baseline-v1",
        description: "Baseline experiment with gpt-4o",
        task: async ({ item }) => {
          // Your LLM function
          const completion = await openai.chat.completions.create({
            model: "gpt-4o",
            messages: [
              { role: "developer", content: "Answer the question accurately." },
              { role: "user", content: item.input.question },
            ],
          });
          return completion.choices[0].message.content;
        },
      });
      
      console.log(`Experiment: ${result.runName}`);
      console.log(`Results: ${result.itemResults.length} items processed`);
      console.log(`URL: ${result.datasetRunUrl}`);
      ```
      
      ---
      
      ## Experiments with Evaluators
      
      Add evaluators to automatically score experiment results.
      
      ```typescript
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      
      const dataset = await langfuse.dataset.get("qa-benchmark");
      
      const result = await dataset.runExperiment({
        name: "gpt-4o-with-eval",
        task: async ({ item }) => {
          const completion = await openai.chat.completions.create({
            model: "gpt-4o",
            messages: [{ role: "user", content: item.input.question }],
          });
          return completion.choices[0].message.content;
        },
        evaluators: [
          // Per-item evaluator
          async ({ input, output, expectedOutput }) => {
            const isCorrect = output
              ?.toLowerCase()
              .includes(expectedOutput?.answer?.toLowerCase() ?? "");
            return {
              name: "contains-answer",
              value: isCorrect ? 1 : 0,
              dataType: "BOOLEAN" as const,
            };
          },
        ],
      });
      ```
      
      ---
      
      ## Linking Dataset Items to Production Traces
      
      Connect dataset items to real production traces for context.
      
      ```typescript
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      
      // Link a dataset item to a source trace
      await langfuse.api.datasetItems.create({
        datasetName: "qa-benchmark",
        input: { question: "What is the capital of France?" },
        expectedOutput: { answer: "Paris" },
        sourceTraceId: "trace-from-production-123",
        sourceObservationId: "obs-456",
      });
      ```
      
      ---
      
      ## Versioned Datasets
      
      Run experiments against specific dataset snapshots.
      
      ```typescript
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      
      // Get dataset at a specific version timestamp
      const versionTimestamp = new Date("2025-12-15T06:30:00").toISOString();
      const versionedDataset = await langfuse.dataset.get("qa-benchmark", {
        version: versionTimestamp,
      });
      
      const result = await versionedDataset.runExperiment({
        name: "baseline-on-v1",
        description: "Running against December 2025 snapshot",
        task: async ({ item }) => {
          return await myLLMFunction(item.input);
        },
      });
      ```
      
      ---
      
      ## Archiving Dataset Items
      
      ```typescript
      import { LangfuseClient } from "@langfuse/client";
      
      const langfuse = new LangfuseClient();
      
      // Archive a dataset item by ID
      await langfuse.api.datasetItems.create({
        id: "item-abc-123",
        status: "ARCHIVED",
      });
      ```
      
      ---
      
      _For tracing patterns, see [tracing.md](tracing.md). For setup, see [core.md](core.md)._
      
    • tracing.md 9.9 KB
      # Langfuse -- Tracing Examples
      
      > Observation creation, nesting, context propagation, observation types, metadata, sessions, and user tracking. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Setup, environment config, flush/shutdown
      - [openai-integration.md](openai-integration.md) -- OpenAI auto-instrumentation
      - [prompt-management.md](prompt-management.md) -- Prompt management
      - [scores-datasets.md](scores-datasets.md) -- Scores and datasets
      
      ---
      
      ## startActiveObservation -- Context Manager
      
      The primary tracing method. Creates an observation, sets it as active context, and ends it automatically.
      
      ```typescript
      import { startActiveObservation } from "@langfuse/tracing";
      
      async function handleChatMessage(message: string): Promise<string> {
        return await startActiveObservation("chat-handler", async (span) => {
          span.update({ input: { message } });
      
          const intent = await classifyIntent(message);
          const response = await generateResponse(intent, message);
      
          span.update({ output: { response, intent } });
          return response;
        });
      }
      ```
      
      ---
      
      ## Nested Observations
      
      Child observations automatically inherit parent context when created during the parent's active scope.
      
      ```typescript
      import { startActiveObservation } from "@langfuse/tracing";
      
      async function ragPipeline(query: string): Promise<string> {
        return await startActiveObservation("rag-pipeline", async (rootSpan) => {
          rootSpan.update({ input: { query } });
      
          // Child 1: Retrieve documents (typed as retriever)
          const docs = await startActiveObservation(
            "retrieve-docs",
            async (span) => {
              span.update({ input: { query } });
              const results = await vectorStore.search(query);
              span.update({ output: { documentCount: results.length } });
              return results;
            },
            { asType: "retriever" },
          );
      
          // Child 2: Generate answer (typed as generation)
          const answer = await startActiveObservation(
            "generate-answer",
            async (span) => {
              span.update({
                input: { query, contextDocs: docs.length },
                model: "gpt-4o",
              });
              const result = await callLLM(query, docs);
              span.update({ output: { answer: result } });
              return result;
            },
            { asType: "generation" },
          );
      
          rootSpan.update({ output: { answer } });
          return answer;
        });
      }
      ```
      
      ---
      
      ## observe() Function Wrapper
      
      Wraps functions for declarative tracing. Inputs and outputs are captured automatically.
      
      ```typescript
      import { observe } from "@langfuse/tracing";
      
      const classifyIntent = observe(
        async (query: string): Promise<string> => {
          const completion = await openai.chat.completions.create({
            model: "gpt-4o",
            messages: [
              { role: "system", content: "Classify the user intent." },
              { role: "user", content: query },
            ],
          });
          return completion.choices[0].message.content ?? "unknown";
        },
        { name: "classify-intent", asType: "generation" },
      );
      
      const searchDocuments = observe(
        async (query: string): Promise<Document[]> => {
          return await vectorStore.similaritySearch(query);
        },
        { name: "search-documents", asType: "retriever" },
      );
      
      // Usage -- both functions are automatically traced
      const intent = await classifyIntent("Book a flight to Paris");
      const docs = await searchDocuments("Paris flights availability");
      ```
      
      ### Controlling Input/Output Capture
      
      ```typescript
      // Disable input capture (for sensitive data)
      const secureFn = observe(
        async (sensitiveInput: string) => {
          /* ... */
        },
        { name: "secure-fn", captureInput: false },
      );
      
      // Disable output capture
      const largeFn = observe(
        async (input: string) => {
          /* returns large object */
        },
        { name: "large-fn", captureOutput: false },
      );
      ```
      
      ---
      
      ## Manual Observations with startObservation
      
      For cases where you need explicit control over observation lifecycle.
      
      ```typescript
      import { startObservation } from "@langfuse/tracing";
      
      async function processItem(item: WorkItem): Promise<void> {
        const span = startObservation(
          "process-item",
          { input: { itemId: item.id } },
          { asType: "tool" },
        );
      
        try {
          const result = await performWork(item);
          span.update({ output: { result } });
        } catch (error) {
          span.update({
            output: {
              error: error instanceof Error ? error.message : "Unknown error",
            },
          });
          throw error;
        } finally {
          span.end(); // REQUIRED -- must explicitly end
        }
      }
      ```
      
      ---
      
      ## Token Usage on Manual Generations
      
      When manually creating generation observations (not using `observeOpenAI`), report token counts via `usageDetails`:
      
      ```typescript
      import { startObservation } from "@langfuse/tracing";
      
      const generation = startObservation(
        "llm-call",
        {
          model: "gpt-4o",
          input: [{ role: "user", content: "What is the capital of France?" }],
        },
        { asType: "generation" },
      );
      
      // ... perform LLM call ...
      
      generation.update({
        output: { content: "The capital of France is Paris." },
        usageDetails: {
          input: 10,
          output: 5,
          total: 15,
        },
      });
      
      generation.end();
      ```
      
      **Note:** `total` is automatically calculated if omitted. You can also add custom token metrics (e.g., `cache_read_input_tokens`).
      
      ---
      
      ## Observation Types
      
      Use `asType` to give semantic meaning to observations.
      
      ```typescript
      import { startActiveObservation } from "@langfuse/tracing";
      
      // Agent step -- decision-making
      await startActiveObservation(
        "plan-next-action",
        async (span) => {
          span.update({ input: { state: currentState } });
          const action = await agent.decide(currentState);
          span.update({ output: { action } });
        },
        { asType: "agent" },
      );
      
      // Tool call -- external API
      await startActiveObservation(
        "call-weather-api",
        async (span) => {
          span.update({ input: { location: "Paris" } });
          const weather = await weatherAPI.get("Paris");
          span.update({ output: weather });
        },
        { asType: "tool" },
      );
      
      // Retriever -- vector search
      await startActiveObservation(
        "vector-search",
        async (span) => {
          span.update({ input: { query: "climate change" } });
          const docs = await vectorStore.search("climate change");
          span.update({ output: { count: docs.length } });
        },
        { asType: "retriever" },
      );
      
      // Evaluator -- quality check
      await startActiveObservation(
        "check-relevance",
        async (span) => {
          span.update({ input: { answer, query } });
          const score = await evaluateRelevance(answer, query);
          span.update({ output: { score } });
        },
        { asType: "evaluator" },
      );
      
      // Embedding call
      await startActiveObservation(
        "create-embedding",
        async (span) => {
          span.update({
            input: { text: "Hello world" },
            model: "text-embedding-3-small",
          });
          const embedding = await createEmbedding("Hello world");
          span.update({ output: { dimensions: embedding.length } });
        },
        { asType: "embedding" },
      );
      ```
      
      ---
      
      ## User and Session Tracking
      
      Use `propagateAttributes()` to attach user, session, and metadata to all nested observations. It wraps a callback -- all observations created inside the callback inherit the attributes.
      
      ```typescript
      import { startActiveObservation, propagateAttributes } from "@langfuse/tracing";
      
      async function handleUserRequest(
        userId: string,
        sessionId: string,
        query: string,
      ) {
        return await propagateAttributes(
          {
            userId,
            sessionId,
            tags: ["production", "v2"],
            metadata: { environment: "prod", region: "eu-west-1" },
          },
          async () => {
            return await startActiveObservation("user-request", async (span) => {
              span.update({ input: { query } });
              const result = await processQuery(query);
              span.update({ output: { result } });
              return result;
            });
          },
        );
      }
      ```
      
      ---
      
      ## Updating Active Observation Without Reference
      
      ```typescript
      import {
        startActiveObservation,
        updateActiveObservation,
      } from "@langfuse/tracing";
      
      async function deeplyNestedFunction() {
        // Update the currently active observation from anywhere in the call stack
        updateActiveObservation({
          metadata: { step: "validation", validated: true },
        });
      }
      
      await startActiveObservation("outer", async () => {
        await deeplyNestedFunction(); // Updates "outer" observation's metadata
      });
      ```
      
      ---
      
      ## Trace and Span ID Access
      
      ```typescript
      import {
        startActiveObservation,
        getActiveTraceId,
        getActiveSpanId,
        createTraceId,
      } from "@langfuse/tracing";
      
      await startActiveObservation("my-trace", async () => {
        const traceId = getActiveTraceId();
        const spanId = getActiveSpanId();
        console.log(`Trace: ${traceId}, Span: ${spanId}`);
      });
      
      // Deterministic trace ID for correlation with external systems
      const traceId = createTraceId("order-12345");
      ```
      
      ---
      
      ## Time to First Token (TTFT) Tracking
      
      ```typescript
      import { startActiveObservation } from "@langfuse/tracing";
      
      await startActiveObservation(
        "llm-call",
        async (span) => {
          span.update({ input: { prompt: "Tell me a story" }, model: "gpt-4o" });
      
          let firstTokenReceived = false;
          const stream = await openai.chat.completions.create({
            model: "gpt-4o",
            messages: [{ role: "user", content: "Tell me a story" }],
            stream: true,
          });
      
          for await (const chunk of stream) {
            if (!firstTokenReceived) {
              span.update({ completionStartTime: new Date().toISOString() });
              firstTokenReceived = true;
            }
            // Process chunk...
          }
        },
        { asType: "generation" },
      );
      ```
      
      ---
      
      ## Cross-Service Context Propagation
      
      Use `asBaggage: true` to propagate context across HTTP boundaries via W3C Baggage headers.
      
      ```typescript
      import { propagateAttributes } from "@langfuse/tracing";
      
      // In the upstream service -- wraps a callback
      await propagateAttributes(
        { userId: "user-123", sessionId: "session-456" },
        async () => {
          // All observations created here inherit userId and sessionId
          await handleRequest();
        },
        { asBaggage: true }, // Propagates via HTTP headers
      );
      ```
      
      ---
      
      _For setup and configuration, see [core.md](core.md). For API reference, see [reference.md](../reference.md)._
      
  • reference.md 9.6 KB
    # Langfuse Quick Reference
    
    > Package index, environment variables, tracing API, observation types, score methods, and client API. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples.
    
    ---
    
    ## Package Installation
    
    ```bash
    # Core tracing (always required for observability)
    npm install @langfuse/tracing @langfuse/otel @opentelemetry/sdk-node
    
    # OpenAI auto-instrumentation (if using OpenAI SDK)
    npm install @langfuse/openai
    
    # Client API (prompts, scores, datasets)
    npm install @langfuse/client
    ```
    
    ### Package Overview
    
    | Package                   | Purpose                                                                      |
    | ------------------------- | ---------------------------------------------------------------------------- |
    | `@langfuse/tracing`       | `startActiveObservation`, `observe`, `startObservation`, context propagation |
    | `@langfuse/otel`          | `LangfuseSpanProcessor` for OTel NodeSDK                                     |
    | `@opentelemetry/sdk-node` | OpenTelemetry SDK (peer dependency)                                          |
    | `@langfuse/client`        | `LangfuseClient` for prompts, scores, datasets                               |
    | `@langfuse/openai`        | `observeOpenAI()` wrapper for OpenAI SDK                                     |
    | `@langfuse/core`          | Shared utilities and types (transitive)                                      |
    
    ---
    
    ## Environment Variables
    
    | Variable                  | Required | Default                      | Description                           |
    | ------------------------- | -------- | ---------------------------- | ------------------------------------- |
    | `LANGFUSE_SECRET_KEY`     | Yes      | --                           | Secret API key (`sk-lf-...`)          |
    | `LANGFUSE_PUBLIC_KEY`     | Yes      | --                           | Public API key (`pk-lf-...`)          |
    | `LANGFUSE_BASE_URL`       | Yes      | `https://cloud.langfuse.com` | Langfuse instance URL                 |
    | `LANGFUSE_SAMPLE_RATE`    | No       | `1.0`                        | Trace sampling rate (0.0-1.0)         |
    | `LANGFUSE_LOG_LEVEL`      | No       | `WARN`                       | SDK log level (DEBUG/INFO/WARN/ERROR) |
    | `LANGFUSE_FLUSH_AT`       | No       | `10`                         | Flush after N queued events           |
    | `LANGFUSE_FLUSH_INTERVAL` | No       | `1` (seconds)                | Flush interval in seconds             |
    | `LANGFUSE_TIMEOUT`        | No       | `5000` (ms)                  | HTTP request timeout                  |
    
    ---
    
    ## Tracing API (`@langfuse/tracing`)
    
    ### Context Managers
    
    ```typescript
    import {
      startActiveObservation,
      startObservation,
      observe,
      updateActiveObservation,
      propagateAttributes,
      getActiveTraceId,
      getActiveSpanId,
      createTraceId,
    } from "@langfuse/tracing";
    
    // Auto-context, auto-end
    await startActiveObservation("name", async (span) => {
      span.update({ input: {...}, output: {...} });
    }, { asType: "generation" });
    
    // Manual control (requires explicit .end())
    const span = startObservation("name", { input: {...} }, { asType: "tool" });
    span.update({ output: {...} });
    span.end();
    
    // Function wrapper (auto-captures inputs/outputs)
    const fn = observe(async (input: string) => { return output; }, {
      name: "fn-name",
      asType: "agent",
    });
    
    // Update active span without reference
    updateActiveObservation({ metadata: { key: "value" } });
    
    // Propagate attributes to all nested observations (wraps a callback)
    await propagateAttributes(
      { userId: "user-1", sessionId: "session-1", tags: ["prod"] },
      async () => { /* observations created here inherit attributes */ },
    );
    
    // Get current trace/span IDs
    const traceId = getActiveTraceId();
    const spanId = getActiveSpanId();
    
    // Deterministic trace ID from seed
    const traceId = createTraceId("my-correlation-key");
    ```
    
    ### Span Update Properties
    
    ```typescript
    span.update({
      input: any, // Observation input
      output: any, // Observation output
      metadata: Record<string, any>, // Custom key-value metadata
      model: string, // LLM model name (for generations)
      completionStartTime: string, // ISO timestamp for TTFT tracking
    });
    ```
    
    ---
    
    ## Observation Types (`asType`)
    
    | Type           | Semantic Meaning                     | Dashboard Icon |
    | -------------- | ------------------------------------ | -------------- |
    | `"span"`       | Generic duration (default)           | --             |
    | `"event"`      | Point-in-time occurrence             | --             |
    | `"generation"` | LLM call (prompt/completion, tokens) | AI icon        |
    | `"agent"`      | AI agent decision step               | Agent icon     |
    | `"tool"`       | External tool/API call               | Tool icon      |
    | `"chain"`      | Link between application steps       | Chain icon     |
    | `"retriever"`  | Vector store / DB retrieval          | Search icon    |
    | `"evaluator"`  | Quality assessment function          | Check icon     |
    | `"embedding"`  | Embedding model call                 | Vector icon    |
    | `"guardrail"`  | Content safety / jailbreak check     | Shield icon    |
    
    ---
    
    ## OpenAI Integration (`@langfuse/openai`)
    
    ```typescript
    import OpenAI from "openai";
    import { observeOpenAI } from "@langfuse/openai";
    
    // Wrap OpenAI client -- all calls automatically traced
    const openai = observeOpenAI(new OpenAI());
    
    // Optional: custom trace attributes per call
    const openai = observeOpenAI(new OpenAI(), {
      generationName: "classify-intent",
      sessionId: "session-123",
      userId: "user-456",
      tags: ["production"],
      generationMetadata: { feature: "intent-classifier" },
    });
    ```
    
    **Auto-captured data:** model, token counts (prompt/completion), estimated cost, latency, time-to-first-token (streaming), errors, function calls.
    
    **Stream token tracking:** Set `stream_options: { include_usage: true }` on OpenAI streaming calls.
    
    ---
    
    ## Client API (`@langfuse/client`)
    
    ```typescript
    import { LangfuseClient } from "@langfuse/client";
    const langfuse = new LangfuseClient();
    ```
    
    ### Prompt Manager (`langfuse.prompt`)
    
    ```typescript
    // Get text prompt (production label by default)
    const prompt = await langfuse.prompt.get("name");
    const compiled = prompt.compile({ var1: "value1" });
    
    // Get chat prompt
    const chatPrompt = await langfuse.prompt.get("name", { type: "chat" });
    const messages = chatPrompt.compile({ var1: "value1" });
    
    // Specific version or label
    await langfuse.prompt.get("name", { version: 2 });
    await langfuse.prompt.get("name", { label: "staging" });
    
    // Bypass cache
    await langfuse.prompt.get("name", { cacheTtlSeconds: 0 });
    
    // Create prompt
    await langfuse.prompt.create({
      name: "name",
      prompt: "Template {{var}}",
      type: "text",
    });
    ```
    
    ### Score Manager (`langfuse.score`)
    
    ```typescript
    // By trace ID
    langfuse.score.create({
      traceId,
      name: "quality",
      value: 0.9,
      dataType: "NUMERIC",
    });
    langfuse.score.create({
      traceId,
      name: "label",
      value: "good",
      dataType: "CATEGORICAL",
    });
    langfuse.score.create({
      traceId,
      name: "passed",
      value: 1,
      dataType: "BOOLEAN",
    });
    
    // By observation ID
    langfuse.score.create({
      traceId,
      observationId,
      name: "accuracy",
      value: 0.8,
      dataType: "NUMERIC",
    });
    
    // By active OTel span
    langfuse.score.activeObservation({
      name: "quality",
      value: 0.9,
      dataType: "NUMERIC",
    });
    langfuse.score.activeTrace({
      name: "quality",
      value: 0.9,
      dataType: "NUMERIC",
    });
    
    // By OTel span reference
    langfuse.score.observation(
      { otelSpan },
      { name: "quality", value: 0.9, dataType: "NUMERIC" },
    );
    langfuse.score.trace(
      { otelSpan },
      { name: "quality", value: 0.9, dataType: "NUMERIC" },
    );
    
    // Flush pending scores
    await langfuse.score.flush();
    ```
    
    ### Dataset Manager (`langfuse.dataset`)
    
    ```typescript
    // Get dataset
    const dataset = await langfuse.dataset.get("dataset-name");
    
    // Run experiment
    const result = await dataset.runExperiment({
      name: "experiment-v1",
      task: async ({ item }) => {
        const output = await myLLMFunction(item.input);
        return output;
      },
    });
    ```
    
    ### Direct API Access (`langfuse.api`)
    
    ```typescript
    // Create dataset
    await langfuse.api.datasets.create({ name: "my-dataset" });
    
    // Add dataset item
    await langfuse.api.datasetItems.create({
      datasetName: "my-dataset",
      input: { text: "hello" },
      expectedOutput: { text: "world" },
    });
    
    // Get trace
    await langfuse.api.trace.get("trace-id");
    
    // Get observation
    await langfuse.api.observations.get("observation-id");
    ```
    
    ### Utility Methods
    
    ```typescript
    // Get dashboard URL for a trace
    langfuse.getTraceUrl("trace-id");
    
    // Flush all pending client data (scores, prompts)
    await langfuse.flush();
    
    // Graceful shutdown
    await langfuse.shutdown();
    ```
    
    ### Flushing OTel Spans (Short-Lived Processes)
    
    ```typescript
    // Flush pending spans via the processor
    await langfuseSpanProcessor.forceFlush();
    
    // Or shutdown the entire OTel SDK (flushes + closes)
    await sdk.shutdown();
    ```
    
    ---
    
    ## LangfuseSpanProcessor Configuration
    
    ```typescript
    import { LangfuseSpanProcessor, isDefaultExportSpan } from "@langfuse/otel";
    
    const processor = new LangfuseSpanProcessor({
      // Credentials (optional -- falls back to env vars)
      publicKey: "pk-lf-...",
      secretKey: "sk-lf-...",
      baseUrl: "https://cloud.langfuse.com",
    
      // Data masking
      mask: ({ data }) => data.replace(/\b\d{16}\b/g, "***MASKED***"),
    
      // Span filtering
      shouldExportSpan: ({ otelSpan }) =>
        isDefaultExportSpan(otelSpan) ||
        otelSpan.instrumentationScope.name.startsWith("my-app"),
    });
    ```
    
    ---
    
    ## Score Data Types
    
    | Type          | Value Type        | Example                 |
    | ------------- | ----------------- | ----------------------- |
    | `NUMERIC`     | `number` (float)  | `0.95`, `42`, `-1.5`    |
    | `CATEGORICAL` | `string`          | `"good"`, `"incorrect"` |
    | `BOOLEAN`     | `number` (0 or 1) | `0` (false), `1` (true) |
    
  • SKILL.md 19.5 KB
    ---
    name: ai-observability-langfuse
    description: LLM observability with Langfuse — OpenTelemetry-based tracing, evaluations, prompt management, datasets, and production best practices
    ---
    
    # Langfuse Observability Patterns
    
    > **Quick Guide:** Use the Langfuse TypeScript SDK (built on OpenTelemetry) to add observability to LLM applications. Install `@langfuse/tracing`, `@langfuse/otel`, and `@opentelemetry/sdk-node` for core tracing. Use `startActiveObservation()` for automatic context propagation or `observe()` to wrap functions. Use `@langfuse/openai` with `observeOpenAI()` for zero-config OpenAI tracing. Use `LangfuseClient` from `@langfuse/client` for prompt management, scores, and datasets. Always call `forceFlush()` or `sdk.shutdown()` in short-lived processes.
    
    ---
    
    <critical_requirements>
    
    ## CRITICAL: Before Using This Skill
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST import and register `instrumentation.ts` at the top of your entry point BEFORE any other imports -- OpenTelemetry must instrument modules before they are loaded)**
    
    **(You MUST call `forceFlush()` or `sdk.shutdown()` in short-lived processes (serverless, scripts, CLI tools) -- events are batched and will be lost without explicit flushing)**
    
    **(You MUST use `@langfuse/openai` with `observeOpenAI()` for OpenAI SDK tracing -- do NOT manually create generation observations for OpenAI calls when the wrapper handles it automatically)**
    
    **(You MUST set `LANGFUSE_SECRET_KEY`, `LANGFUSE_PUBLIC_KEY`, and `LANGFUSE_BASE_URL` via environment variables -- never hardcode credentials)**
    
    **(You MUST use `startActiveObservation()` or `observe()` for nested tracing -- manual `startObservation()` requires explicit `.end()` calls and does NOT propagate context automatically)**
    
    </critical_requirements>
    
    ---
    
    **Auto-detection:** Langfuse, langfuse, @langfuse/tracing, @langfuse/otel, @langfuse/client, @langfuse/openai, LangfuseSpanProcessor, LangfuseClient, startActiveObservation, startObservation, observeOpenAI, langfuse.score, langfuse.prompt, langfuse.dataset, LANGFUSE_SECRET_KEY, LANGFUSE_PUBLIC_KEY, forceFlush
    
    **When to use:**
    
    - Adding observability and tracing to LLM application code (any provider)
    - Wrapping OpenAI SDK calls for automatic token/cost tracking
    - Managing prompt templates with versioning, labels, and variable compilation
    - Evaluating LLM output quality with scores (numeric, categorical, boolean)
    - Running experiments against datasets for regression testing
    - Tracking sessions, users, and metadata across multi-turn conversations
    - Monitoring LLM costs and token usage in production
    
    **Key patterns covered:**
    
    - OpenTelemetry setup with `LangfuseSpanProcessor`
    - Tracing with `startActiveObservation`, `observe`, and manual `startObservation`
    - Observation types (span, generation, agent, tool, retriever, evaluator, embedding, chain, guardrail)
    - OpenAI SDK auto-instrumentation with `observeOpenAI()`
    - Prompt management (get, compile, text vs chat prompts, versioning)
    - Scores and evaluations (numeric, categorical, boolean)
    - Datasets and experiments for testing
    - Flush, shutdown, and lifecycle management
    
    **When NOT to use:**
    
    - You only need basic `console.log` debugging -- Langfuse is for structured production observability
    - You want provider-specific tracing built into an AI SDK -- check if your framework has native observability
    - You need APM/infrastructure monitoring (CPU, memory, HTTP latency) -- use a general-purpose observability tool
    
    ---
    
    ## Examples Index
    
    - [Core: Setup & Configuration](examples/core.md) -- OpenTelemetry setup, instrumentation file, client init, flush/shutdown
    - [Tracing](examples/tracing.md) -- startActiveObservation, observe, manual tracing, nesting, observation types, metadata
    - [OpenAI Integration](examples/openai-integration.md) -- observeOpenAI wrapper, streaming, token tracking, custom attributes
    - [Prompt Management](examples/prompt-management.md) -- getPrompt, compile, text vs chat, versioning, caching
    - [Scores & Datasets](examples/scores-datasets.md) -- Numeric/categorical/boolean scores, datasets, experiments
    - [Quick API Reference](reference.md) -- Package index, environment variables, observation types, score methods
    
    ---
    
    <philosophy>
    
    ## Philosophy
    
    Langfuse provides **open-source LLM observability** built on OpenTelemetry. The SDK (v4+, August 2025) is a ground-up rewrite using OTel as the tracing backbone, meaning traces integrate naturally with the broader observability ecosystem.
    
    **Core principles:**
    
    1. **OpenTelemetry-native** -- Built on OTel spans and context propagation. Langfuse observations are wrappers around OTel spans with LLM-specific attributes (model, tokens, cost). This means any OTel-compatible instrumentation library works alongside Langfuse.
    2. **Zero-latency tracing** -- All trace events are queued locally and flushed in background batches. Your application's response time is not affected by observability.
    3. **Modular packages** -- `@langfuse/tracing` for instrumentation, `@langfuse/client` for prompts/scores/datasets, `@langfuse/openai` for OpenAI auto-instrumentation. Install only what you need.
    4. **Context-first** -- `startActiveObservation()` automatically propagates parent-child relationships. Nested observations inherit context without manual ID threading.
    5. **Observation types** -- LLM-specific types (`generation`, `agent`, `tool`, `retriever`, `evaluator`, `embedding`) provide semantic meaning to traces, enabling richer dashboard views and filtering.
    
    </philosophy>
    
    ---
    
    <patterns>
    
    ## Core Patterns
    
    ### Pattern 1: OpenTelemetry Setup
    
    Create an `instrumentation.ts` file and import it at the top of your entry point.
    
    ```typescript
    // instrumentation.ts
    import { NodeSDK } from "@opentelemetry/sdk-node";
    import { LangfuseSpanProcessor } from "@langfuse/otel";
    
    const sdk = new NodeSDK({
      spanProcessors: [new LangfuseSpanProcessor()],
    });
    
    sdk.start();
    
    export { sdk };
    ```
    
    ```typescript
    // index.ts -- import instrumentation FIRST
    import "./instrumentation";
    // All other imports AFTER instrumentation
    import { startActiveObservation } from "@langfuse/tracing";
    ```
    
    **Why good:** OTel must instrument modules before they are loaded; importing instrumentation first ensures all subsequent imports are traced automatically
    
    ```typescript
    // BAD: importing instrumentation after other modules
    import { startActiveObservation } from "@langfuse/tracing";
    import "./instrumentation"; // TOO LATE -- tracing won't capture earlier imports
    ```
    
    **Why bad:** Auto-instrumentation of LLM SDKs requires OTel to be initialized before those modules are imported
    
    **See:** [examples/core.md](examples/core.md) for environment variables, sampling, masking, and production configuration
    
    ---
    
    ### Pattern 2: Tracing with startActiveObservation
    
    The primary instrumentation pattern. Creates an observation, makes it the active context, and automatically ends it when the callback completes.
    
    ```typescript
    import { startActiveObservation } from "@langfuse/tracing";
    
    async function handleRequest(query: string): Promise<string> {
      return await startActiveObservation("handle-request", async (span) => {
        span.update({ input: { query } });
    
        // Nested observation -- automatically becomes a child
        const result = await startActiveObservation(
          "process-query",
          async (child) => {
            child.update({ input: { query } });
            const answer = await callLLM(query);
            child.update({ output: { answer } });
            return answer;
          },
        );
    
        span.update({ output: { result } });
        return result;
      });
    }
    ```
    
    **Why good:** Automatic context propagation, automatic end on callback completion, nesting creates parent-child hierarchy without manual ID management
    
    ```typescript
    // BAD: using startObservation without ending it
    import { startObservation } from "@langfuse/tracing";
    
    const span = startObservation("my-span");
    await doWork();
    // span.end() never called -- observation stays open forever
    ```
    
    **Why bad:** Manual `startObservation` requires explicit `.end()` calls; forgetting creates open-ended observations
    
    **See:** [examples/tracing.md](examples/tracing.md) for observe wrapper, observation types, metadata, and manual tracing
    
    ---
    
    ### Pattern 3: The observe() Wrapper
    
    Wraps a function to automatically capture inputs, outputs, timings, and errors.
    
    ```typescript
    import { observe } from "@langfuse/tracing";
    
    const classifyIntent = observe(
      async (query: string) => {
        const result = await callLLM(query);
        return result.intent;
      },
      { name: "classify-intent", asType: "generation" },
    );
    
    // Usage -- automatically traced
    const intent = await classifyIntent("Book a flight to Paris");
    ```
    
    **Why good:** Declarative tracing, inputs/outputs captured automatically, `asType` tags the observation type for richer dashboard filtering
    
    ---
    
    ### Pattern 4: OpenAI Auto-Instrumentation
    
    Use `observeOpenAI()` to wrap the OpenAI client for automatic tracing of all calls.
    
    ```typescript
    import OpenAI from "openai";
    import { observeOpenAI } from "@langfuse/openai";
    
    const openai = observeOpenAI(new OpenAI());
    
    // All calls automatically traced with model, tokens, cost
    const completion = await openai.chat.completions.create({
      model: "gpt-4o",
      messages: [{ role: "user", content: "Hello" }],
    });
    ```
    
    **Why good:** Zero manual instrumentation, captures model name, token counts, estimated costs, latency, and streaming metrics automatically
    
    ```typescript
    // BAD: manually creating generation observations for OpenAI calls
    await startActiveObservation("openai-call", async (span) => {
      const result = await rawOpenai.chat.completions.create({ ... });
      span.update({
        model: "gpt-4o",
        input: messages,
        output: result.choices[0].message.content,
      });
    }, { asType: "generation" });
    ```
    
    **Why bad:** `observeOpenAI` handles all of this automatically with more accurate token/cost data; manual tracking is error-prone and duplicates effort
    
    **See:** [examples/openai-integration.md](examples/openai-integration.md) for streaming, custom attributes, and token tracking on streams
    
    ---
    
    ### Pattern 5: Prompt Management
    
    Fetch versioned prompts, compile with variables, and link to traces.
    
    ```typescript
    import { LangfuseClient } from "@langfuse/client";
    
    const langfuse = new LangfuseClient();
    
    // Fetch a text prompt (production label by default)
    const prompt = await langfuse.prompt.get("summarize-article");
    const compiled = prompt.compile({ topic: "AI safety", length: "brief" });
    // -> "Write a brief summary about AI safety."
    
    // Fetch a chat prompt
    const chatPrompt = await langfuse.prompt.get("assistant-v2", { type: "chat" });
    const messages = chatPrompt.compile({ userName: "Alice" });
    // -> [{ role: "system", content: "You are helping Alice..." }, ...]
    ```
    
    **Why good:** Centralized prompt management with versioning, labels for A/B testing, variable compilation, and built-in caching
    
    **See:** [examples/prompt-management.md](examples/prompt-management.md) for versioning, labels, cache control, and linking prompts to traces
    
    ---
    
    ### Pattern 6: Scores and Evaluations
    
    Attach quality measurements to traces and observations.
    
    ```typescript
    import { LangfuseClient } from "@langfuse/client";
    
    const langfuse = new LangfuseClient();
    
    // Numeric score
    langfuse.score.create({
      traceId: "trace-123",
      name: "relevance",
      value: 0.95,
      dataType: "NUMERIC",
    });
    
    // Categorical score
    langfuse.score.create({
      traceId: "trace-123",
      name: "quality",
      value: "good",
      dataType: "CATEGORICAL",
    });
    
    // Boolean score (0 or 1)
    langfuse.score.create({
      traceId: "trace-123",
      name: "contains-hallucination",
      value: 0,
      dataType: "BOOLEAN",
    });
    
    // Score a specific observation within a trace
    langfuse.score.create({
      traceId: "trace-123",
      observationId: "obs-456",
      name: "accuracy",
      value: 0.88,
      dataType: "NUMERIC",
    });
    
    // Flush in short-lived processes
    await langfuse.score.flush();
    ```
    
    **Why good:** Three data types cover all evaluation needs, scores attach at trace or observation level, fire-and-forget API with batching
    
    **See:** [examples/scores-datasets.md](examples/scores-datasets.md) for active observation scoring, session scores, datasets, and experiments
    
    ---
    
    ### Pattern 7: Flush and Shutdown
    
    Always flush in short-lived processes. The SDK batches events and sends them asynchronously.
    
    ```typescript
    import { sdk } from "./instrumentation";
    import { LangfuseClient } from "@langfuse/client";
    import { LangfuseSpanProcessor } from "@langfuse/otel";
    
    const langfuse = new LangfuseClient();
    
    async function main() {
      // ... do work ...
    
      // Flush scores
      await langfuse.score.flush();
    
      // Shutdown OTel SDK (flushes all pending spans)
      await sdk.shutdown();
    }
    
    main();
    ```
    
    **Why good:** Explicit flush/shutdown ensures all events are sent before the process exits; without this, data is silently lost in serverless and scripts
    
    ```typescript
    // BAD: exiting without flushing
    async function handler() {
      await startActiveObservation("my-trace", async (span) => {
        span.update({ output: "done" });
      });
      // Process exits -- batched events never sent
    }
    ```
    
    **Why bad:** Langfuse batches events locally; if the process exits before the flush interval, events are lost
    
    </patterns>
    
    ---
    
    <performance>
    
    ## Performance Optimization
    
    ### Sampling for High-Volume Applications
    
    Reduce costs by sampling a subset of traces:
    
    ```typescript
    import { TraceIdRatioBasedSampler } from "@opentelemetry/sdk-trace-base";
    
    const sdk = new NodeSDK({
      sampler: new TraceIdRatioBasedSampler(0.2), // Sample 20% of traces
      spanProcessors: [new LangfuseSpanProcessor()],
    });
    ```
    
    Or via environment variable: `LANGFUSE_SAMPLE_RATE=0.2`
    
    ### Key Optimization Patterns
    
    - **Batch flush tuning** -- Configure `LANGFUSE_FLUSH_AT` (default 10) and `LANGFUSE_FLUSH_INTERVAL` (default 1s) for your workload
    - **Span filtering** -- Use `shouldExportSpan` on `LangfuseSpanProcessor` to drop noisy non-LLM spans
    - **Data masking** -- Redact PII before transmission with the `mask` option to avoid storing sensitive data
    - **Stream token tracking** -- Set `stream_options: { include_usage: true }` on OpenAI streaming calls so `observeOpenAI` captures token counts
    
    </performance>
    
    ---
    
    <decision_framework>
    
    ## Decision Framework
    
    ### Which Packages to Install
    
    ```
    What do you need?
    +-- Tracing LLM calls?
    |   +-- YES -> npm install @langfuse/tracing @langfuse/otel @opentelemetry/sdk-node
    |   +-- Also using OpenAI SDK?
    |       +-- YES -> npm install @langfuse/openai
    +-- Prompt management, scores, or datasets?
    |   +-- YES -> npm install @langfuse/client
    +-- Both tracing AND client features?
        +-- YES -> Install all: @langfuse/tracing @langfuse/otel @opentelemetry/sdk-node @langfuse/client
    ```
    
    ### Which Tracing Method to Use
    
    ```
    How do you want to instrument?
    +-- Wrapping a function? -> observe() (declarative, auto-captures inputs/outputs)
    +-- Block of code with nesting? -> startActiveObservation() (context propagation, auto-end)
    +-- Need manual start/end control? -> startObservation() (requires explicit .end())
    +-- OpenAI SDK calls? -> observeOpenAI() (zero-config auto-instrumentation)
    +-- Update active span without reference? -> updateActiveObservation()
    ```
    
    ### Which Observation Type (asType)
    
    ```
    What is this observation?
    +-- LLM call (prompt -> completion) -> "generation"
    +-- AI agent decision-making step -> "agent"
    +-- External API or function call -> "tool"
    +-- Vector store or DB retrieval -> "retriever"
    +-- Quality assessment step -> "evaluator"
    +-- Embedding creation -> "embedding"
    +-- Link between application steps -> "chain"
    +-- Content safety / jailbreak check -> "guardrail"
    +-- Generic duration operation -> "span" (default)
    +-- Point-in-time event -> "event"
    ```
    
    </decision_framework>
    
    ---
    
    <red_flags>
    
    ## RED FLAGS
    
    **High Priority Issues:**
    
    - Not importing `instrumentation.ts` before other modules (auto-instrumentation silently fails)
    - Exiting short-lived processes without `forceFlush()` or `sdk.shutdown()` (events are silently lost)
    - Hardcoding `LANGFUSE_SECRET_KEY` or `LANGFUSE_PUBLIC_KEY` in source code (credential exposure)
    - Using manual generation observations when `observeOpenAI()` would handle it automatically (duplicated effort, less accurate data)
    - Using `startObservation()` without calling `.end()` (observation stays open indefinitely)
    
    **Medium Priority Issues:**
    
    - Not setting `stream_options: { include_usage: true }` on OpenAI streaming calls (token counts missing from `observeOpenAI` traces)
    - Forgetting to call `langfuse.score.flush()` in short-lived processes (scores are batched and may be lost)
    - Using `startObservation()` when `startActiveObservation()` would work (no automatic context propagation or auto-end)
    - Not using `asType` on observations (all observations appear as generic spans, losing semantic meaning)
    - Not setting `LANGFUSE_BASE_URL` for self-hosted instances (defaults to cloud.langfuse.com)
    
    **Common Mistakes:**
    
    - Importing `@langfuse/openai` without setting up the OTel `NodeSDK` first -- the OpenAI wrapper requires OTel context to send traces
    - Confusing `LangfuseClient` (from `@langfuse/client`, for prompts/scores/datasets) with the OTel tracing functions (from `@langfuse/tracing`)
    - Using `prompt.compile()` without matching all `{{variable}}` placeholders -- unmatched variables remain as literal `{{name}}` in output
    - Calling `langfuse.score.create()` with a `value` of type `string` for `NUMERIC` scores or `number` for `CATEGORICAL` scores (type mismatch)
    - Running dataset experiments without OTel setup -- experiment tasks run inside `startActiveObservation` which requires OTel
    
    **Gotchas & Edge Cases:**
    
    - `observeOpenAI()` does NOT support the OpenAI Assistants API -- only Chat Completions and Responses API
    - The SDK's default span filter only exports Langfuse and GenAI spans. If you use a custom instrumentation library, you must configure `shouldExportSpan` to include it.
    - `LangfuseClient.prompt.get()` caches prompts with a default TTL. If you update a prompt and don't see changes, set `cacheTtlSeconds: 0` to bypass caching.
    - Boolean scores use float values (`0` or `1`), not JavaScript booleans (`true`/`false`).
    - Self-hosted Langfuse requires platform version >= 3.95.0 for TypeScript SDK v4 compatibility.
    - `score.create()` is fire-and-forget (synchronous) -- it queues the score for batched delivery. You only need `await` on `flush()`.
    - Dataset names with slashes (`evaluation/qa-dataset`) must be URL-encoded when used as path parameters.
    - The v4+ SDK is a complete rewrite from v3 -- `Langfuse` class, `trace()`, `span()`, `generation()` from v3 are replaced by OTel-based APIs.
    
    </red_flags>
    
    ---
    
    <critical_reminders>
    
    ## CRITICAL REMINDERS
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST import and register `instrumentation.ts` at the top of your entry point BEFORE any other imports -- OpenTelemetry must instrument modules before they are loaded)**
    
    **(You MUST call `forceFlush()` or `sdk.shutdown()` in short-lived processes (serverless, scripts, CLI tools) -- events are batched and will be lost without explicit flushing)**
    
    **(You MUST use `@langfuse/openai` with `observeOpenAI()` for OpenAI SDK tracing -- do NOT manually create generation observations for OpenAI calls when the wrapper handles it automatically)**
    
    **(You MUST set `LANGFUSE_SECRET_KEY`, `LANGFUSE_PUBLIC_KEY`, and `LANGFUSE_BASE_URL` via environment variables -- never hardcode credentials)**
    
    **(You MUST use `startActiveObservation()` or `observe()` for nested tracing -- manual `startObservation()` requires explicit `.end()` calls and does NOT propagate context automatically)**
    
    **Failure to follow these rules will produce silent data loss, missing traces, or credential exposure in LLM observability.**
    
    </critical_reminders>
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related