ai-observability-langfuse
LLM observability with Langfuse — OpenTelemetry-based tracing, evaluations, prompt management, datasets, and production best practices
Install
npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/ai-observability-langfuse/skills/ai-observability-langfuse
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
git clone https://github.com/agents-inc/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Langfuse Observability Patterns
Quick Guide: Use the Langfuse TypeScript SDK (built on OpenTelemetry) to add observability to LLM applications. Install
@langfuse/tracing,@langfuse/otel, and@opentelemetry/sdk-nodefor core tracing. UsestartActiveObservation()for automatic context propagation orobserve()to wrap functions. Use@langfuse/openaiwithobserveOpenAI()for zero-config OpenAI tracing. UseLangfuseClientfrom@langfuse/clientfor prompt management, scores, and datasets. Always callforceFlush()orsdk.shutdown()in short-lived processes.
<critical_requirements>
CRITICAL: Before Using This Skill
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST import and register instrumentation.ts at the top of your entry point BEFORE any other imports -- OpenTelemetry must instrument modules before they are loaded)
(You MUST call forceFlush() or sdk.shutdown() in short-lived processes (serverless, scripts, CLI tools) -- events are batched and will be lost without explicit flushing)
(You MUST use @langfuse/openai with observeOpenAI() for OpenAI SDK tracing -- do NOT manually create generation observations for OpenAI calls when the wrapper handles it automatically)
(You MUST set LANGFUSE_SECRET_KEY, LANGFUSE_PUBLIC_KEY, and LANGFUSE_BASE_URL via environment variables -- never hardcode credentials)
(You MUST use startActiveObservation() or observe() for nested tracing -- manual startObservation() requires explicit .end() calls and does NOT propagate context automatically)
</critical_requirements>
Auto-detection: Langfuse, langfuse, @langfuse/tracing, @langfuse/otel, @langfuse/client, @langfuse/openai, LangfuseSpanProcessor, LangfuseClient, startActiveObservation, startObservation, observeOpenAI, langfuse.score, langfuse.prompt, langfuse.dataset, LANGFUSE_SECRET_KEY, LANGFUSE_PUBLIC_KEY, forceFlush
When to use:
- Adding observability and tracing to LLM application code (any provider)
- Wrapping OpenAI SDK calls for automatic token/cost tracking
- Managing prompt templates with versioning, labels, and variable compilation
- Evaluating LLM output quality with scores (numeric, categorical, boolean)
- Running experiments against datasets for regression testing
- Tracking sessions, users, and metadata across multi-turn conversations
- Monitoring LLM costs and token usage in production
Key patterns covered:
- OpenTelemetry setup with
LangfuseSpanProcessor - Tracing with
startActiveObservation,observe, and manualstartObservation - Observation types (span, generation, agent, tool, retriever, evaluator, embedding, chain, guardrail)
- OpenAI SDK auto-instrumentation with
observeOpenAI() - Prompt management (get, compile, text vs chat prompts, versioning)
- Scores and evaluations (numeric, categorical, boolean)
- Datasets and experiments for testing
- Flush, shutdown, and lifecycle management
When NOT to use:
- You only need basic
console.logdebugging -- Langfuse is for structured production observability - You want provider-specific tracing built into an AI SDK -- check if your framework has native observability
- You need APM/infrastructure monitoring (CPU, memory, HTTP latency) -- use a general-purpose observability tool
Examples Index
- Core: Setup & Configuration -- OpenTelemetry setup, instrumentation file, client init, flush/shutdown
- Tracing -- startActiveObservation, observe, manual tracing, nesting, observation types, metadata
- OpenAI Integration -- observeOpenAI wrapper, streaming, token tracking, custom attributes
- Prompt Management -- getPrompt, compile, text vs chat, versioning, caching
- Scores & Datasets -- Numeric/categorical/boolean scores, datasets, experiments
- Quick API Reference -- Package index, environment variables, observation types, score methods
<decision_framework>
Decision Framework
Which Packages to Install
What do you need?
+-- Tracing LLM calls?
| +-- YES -> npm install @langfuse/tracing @langfuse/otel @opentelemetry/sdk-node
| +-- Also using OpenAI SDK?
| +-- YES -> npm install @langfuse/openai
+-- Prompt management, scores, or datasets?
| +-- YES -> npm install @langfuse/client
+-- Both tracing AND client features?
+-- YES -> Install all: @langfuse/tracing @langfuse/otel @opentelemetry/sdk-node @langfuse/client
Which Tracing Method to Use
How do you want to instrument?
+-- Wrapping a function? -> observe() (declarative, auto-captures inputs/outputs)
+-- Block of code with nesting? -> startActiveObservation() (context propagation, auto-end)
+-- Need manual start/end control? -> startObservation() (requires explicit .end())
+-- OpenAI SDK calls? -> observeOpenAI() (zero-config auto-instrumentation)
+-- Update active span without reference? -> updateActiveObservation()
Which Observation Type (asType)
What is this observation?
+-- LLM call (prompt -> completion) -> "generation"
+-- AI agent decision-making step -> "agent"
+-- External API or function call -> "tool"
+-- Vector store or DB retrieval -> "retriever"
+-- Quality assessment step -> "evaluator"
+-- Embedding creation -> "embedding"
+-- Link between application steps -> "chain"
+-- Content safety / jailbreak check -> "guardrail"
+-- Generic duration operation -> "span" (default)
+-- Point-in-time event -> "event"
</decision_framework>
<red_flags>
RED FLAGS
High Priority Issues:
- Not importing
instrumentation.tsbefore other modules (auto-instrumentation silently fails) - Exiting short-lived processes without
forceFlush()orsdk.shutdown()(events are silently lost) - Hardcoding
LANGFUSE_SECRET_KEYorLANGFUSE_PUBLIC_KEYin source code (credential exposure) - Using manual generation observations when
observeOpenAI()would handle it automatically (duplicated effort, less accurate data) - Using
startObservation()without calling.end()(observation stays open indefinitely)
Medium Priority Issues:
- Not setting
stream_options: { include_usage: true }on OpenAI streaming calls (token counts missing fromobserveOpenAItraces) - Forgetting to call
langfuse.score.flush()in short-lived processes (scores are batched and may be lost) - Using
startObservation()whenstartActiveObservation()would work (no automatic context propagation or auto-end) - Not using
asTypeon observations (all observations appear as generic spans, losing semantic meaning) - Not setting
LANGFUSE_BASE_URLfor self-hosted instances (defaults to cloud.langfuse.com)
Common Mistakes:
- Importing
@langfuse/openaiwithout setting up the OTelNodeSDKfirst -- the OpenAI wrapper requires OTel context to send traces - Confusing
LangfuseClient(from@langfuse/client, for prompts/scores/datasets) with the OTel tracing functions (from@langfuse/tracing) - Using
prompt.compile()without matching all{{variable}}placeholders -- unmatched variables remain as literal{{name}}in output - Calling
langfuse.score.create()with avalueof typestringforNUMERICscores ornumberforCATEGORICALscores (type mismatch) - Running dataset experiments without OTel setup -- experiment tasks run inside
startActiveObservationwhich requires OTel
Gotchas & Edge Cases:
observeOpenAI()does NOT support the OpenAI Assistants API -- only Chat Completions and Responses API- The SDK's default span filter only exports Langfuse and GenAI spans. If you use a custom instrumentation library, you must configure
shouldExportSpanto include it. LangfuseClient.prompt.get()caches prompts with a default TTL. If you update a prompt and don't see changes, setcacheTtlSeconds: 0to bypass caching.- Boolean scores use float values (
0or1), not JavaScript booleans (true/false). - Self-hosted Langfuse requires platform version >= 3.95.0 for TypeScript SDK v4 compatibility.
score.create()is fire-and-forget (synchronous) -- it queues the score for batched delivery. You only needawaitonflush().- Dataset names with slashes (
evaluation/qa-dataset) must be URL-encoded when used as path parameters. - The v4+ SDK is a complete rewrite from v3 --
Langfuseclass,trace(),span(),generation()from v3 are replaced by OTel-based APIs.
</red_flags>
<critical_reminders>
CRITICAL REMINDERS
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST import and register instrumentation.ts at the top of your entry point BEFORE any other imports -- OpenTelemetry must instrument modules before they are loaded)
(You MUST call forceFlush() or sdk.shutdown() in short-lived processes (serverless, scripts, CLI tools) -- events are batched and will be lost without explicit flushing)
(You MUST use @langfuse/openai with observeOpenAI() for OpenAI SDK tracing -- do NOT manually create generation observations for OpenAI calls when the wrapper handles it automatically)
(You MUST set LANGFUSE_SECRET_KEY, LANGFUSE_PUBLIC_KEY, and LANGFUSE_BASE_URL via environment variables -- never hardcode credentials)
(You MUST use startActiveObservation() or observe() for nested tracing -- manual startObservation() requires explicit .end() calls and does NOT propagate context automatically)
Failure to follow these rules will produce silent data loss, missing traces, or credential exposure in LLM observability.
</critical_reminders>
Files (skills)
-
examples
-
core.md 6.3 KB
# Langfuse -- Setup & Configuration Examples > OpenTelemetry initialization, environment config, sampling, data masking, multi-project setup, and lifecycle management. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [tracing.md](tracing.md) -- Tracing with observations, nesting, metadata - [openai-integration.md](openai-integration.md) -- OpenAI SDK auto-instrumentation - [prompt-management.md](prompt-management.md) -- Prompt management and versioning - [scores-datasets.md](scores-datasets.md) -- Scores, datasets, experiments --- ## Basic Setup ```typescript // instrumentation.ts import { NodeSDK } from "@opentelemetry/sdk-node"; import { LangfuseSpanProcessor } from "@langfuse/otel"; const sdk = new NodeSDK({ spanProcessors: [new LangfuseSpanProcessor()], }); sdk.start(); export { sdk }; ``` ```bash # .env LANGFUSE_SECRET_KEY="sk-lf-..." LANGFUSE_PUBLIC_KEY="pk-lf-..." LANGFUSE_BASE_URL="https://cloud.langfuse.com" # EU region # LANGFUSE_BASE_URL="https://us.cloud.langfuse.com" # US region ``` ```typescript // index.ts -- instrumentation MUST be imported first import { sdk } from "./instrumentation"; import { startActiveObservation } from "@langfuse/tracing"; async function main() { await startActiveObservation("my-trace", async (span) => { span.update({ input: "Hello, Langfuse!", output: "Trace captured successfully.", }); }); } main().finally(() => sdk.shutdown()); ``` --- ## Client Initialization ```typescript // lib/langfuse.ts import { LangfuseClient } from "@langfuse/client"; // Reads credentials from environment variables automatically const langfuse = new LangfuseClient(); export { langfuse }; ``` ```typescript // lib/langfuse.ts -- explicit configuration import { LangfuseClient } from "@langfuse/client"; const TIMEOUT_MS = 10_000; const langfuse = new LangfuseClient({ publicKey: process.env.LANGFUSE_PUBLIC_KEY, secretKey: process.env.LANGFUSE_SECRET_KEY, baseUrl: process.env.LANGFUSE_BASE_URL, timeout: TIMEOUT_MS, }); export { langfuse }; ``` --- ## Sampling for High-Volume Applications ```typescript // instrumentation.ts -- sample 20% of traces import { NodeSDK } from "@opentelemetry/sdk-node"; import { TraceIdRatioBasedSampler } from "@opentelemetry/sdk-trace-base"; import { LangfuseSpanProcessor } from "@langfuse/otel"; const SAMPLE_RATE = 0.2; const sdk = new NodeSDK({ sampler: new TraceIdRatioBasedSampler(SAMPLE_RATE), spanProcessors: [new LangfuseSpanProcessor()], }); sdk.start(); export { sdk }; ``` Or via environment variable: `LANGFUSE_SAMPLE_RATE=0.2` --- ## Data Masking ```typescript // instrumentation.ts -- redact PII before sending to Langfuse import { NodeSDK } from "@opentelemetry/sdk-node"; import { LangfuseSpanProcessor } from "@langfuse/otel"; const CREDIT_CARD_PATTERN = /\b\d{4}[- ]?\d{4}[- ]?\d{4}[- ]?\d{4}\b/g; const EMAIL_PATTERN = /\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b/g; const processor = new LangfuseSpanProcessor({ mask: ({ data }) => data .replace(CREDIT_CARD_PATTERN, "***MASKED_CC***") .replace(EMAIL_PATTERN, "***MASKED_EMAIL***"), }); const sdk = new NodeSDK({ spanProcessors: [processor], }); sdk.start(); export { sdk }; ``` --- ## Span Filtering ```typescript // instrumentation.ts -- only export Langfuse + custom spans import { NodeSDK } from "@opentelemetry/sdk-node"; import { LangfuseSpanProcessor, isDefaultExportSpan } from "@langfuse/otel"; const processor = new LangfuseSpanProcessor({ shouldExportSpan: ({ otelSpan }) => isDefaultExportSpan(otelSpan) || otelSpan.instrumentationScope.name.startsWith("my-app"), }); const sdk = new NodeSDK({ spanProcessors: [processor], }); sdk.start(); export { sdk }; ``` --- ## Isolated TracerProvider Separate Langfuse tracing from other observability backends: ```typescript // instrumentation.ts -- isolated provider import { NodeTracerProvider } from "@opentelemetry/sdk-trace-node"; import { LangfuseSpanProcessor } from "@langfuse/otel"; import { setLangfuseTracerProvider } from "@langfuse/tracing"; const langfuseProvider = new NodeTracerProvider({ spanProcessors: [new LangfuseSpanProcessor()], }); setLangfuseTracerProvider(langfuseProvider); ``` --- ## Multi-Project Setup ```typescript // instrumentation.ts -- send spans to multiple Langfuse projects import { NodeSDK } from "@opentelemetry/sdk-node"; import { LangfuseSpanProcessor } from "@langfuse/otel"; const sdk = new NodeSDK({ spanProcessors: [ new LangfuseSpanProcessor({ publicKey: "pk-lf-project-1", secretKey: "sk-lf-project-1", }), new LangfuseSpanProcessor({ publicKey: "pk-lf-project-2", secretKey: "sk-lf-project-2", }), ], }); sdk.start(); export { sdk }; ``` --- ## Debug Logging ```bash # Via environment variable LANGFUSE_LOG_LEVEL=DEBUG # Or via code ``` ```typescript import { configureGlobalLogger } from "@langfuse/core"; configureGlobalLogger({ level: "DEBUG" }); ``` --- ## Serverless / Short-Lived Process Pattern For serverless environments, use `exportMode: "immediate"` on the span processor to export spans immediately rather than batching: ```typescript // instrumentation.ts (serverless variant) import { NodeSDK } from "@opentelemetry/sdk-node"; import { LangfuseSpanProcessor } from "@langfuse/otel"; export const langfuseSpanProcessor = new LangfuseSpanProcessor({ exportMode: "immediate", }); export const sdk = new NodeSDK({ spanProcessors: [langfuseSpanProcessor], }); sdk.start(); ``` ```typescript import { sdk } from "./instrumentation"; import { startActiveObservation } from "@langfuse/tracing"; import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); export async function handler(event: unknown) { try { return await startActiveObservation("lambda-handler", async (span) => { span.update({ input: event }); const result = await processEvent(event); span.update({ output: result }); // Score the result langfuse.score.activeTrace({ name: "success", value: 1, dataType: "BOOLEAN", }); return result; }); } finally { // CRITICAL: flush before Lambda freezes await langfuse.score.flush(); await sdk.shutdown(); } } ``` --- _For tracing patterns, see [tracing.md](tracing.md). For API reference, see [reference.md](../reference.md)._ -
openai-integration.md 6 KB
# Langfuse -- OpenAI Integration Examples > Auto-instrumentation with `observeOpenAI()`, streaming token tracking, custom attributes, and nested tracing. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Setup, environment config, flush/shutdown - [tracing.md](tracing.md) -- Manual tracing patterns - [prompt-management.md](prompt-management.md) -- Prompt management - [scores-datasets.md](scores-datasets.md) -- Scores and datasets --- ## Basic OpenAI Wrapper ```typescript // lib/openai.ts import OpenAI from "openai"; import { observeOpenAI } from "@langfuse/openai"; // Wrap the OpenAI client -- all calls automatically traced const openai = observeOpenAI(new OpenAI()); export { openai }; ``` ```typescript // usage.ts import { openai } from "./lib/openai"; const completion = await openai.chat.completions.create({ model: "gpt-4o", messages: [ { role: "developer", content: "You are a helpful assistant." }, { role: "user", content: "Explain async/await in TypeScript." }, ], }); console.log(completion.choices[0].message.content); // Langfuse automatically captures: model, tokens, cost, latency ``` --- ## Custom Trace Attributes Pass metadata to control how the generation appears in Langfuse. ```typescript import OpenAI from "openai"; import { observeOpenAI } from "@langfuse/openai"; const openai = observeOpenAI(new OpenAI(), { generationName: "intent-classifier", sessionId: "session-abc-123", userId: "user-456", tags: ["production", "classifier-v2"], generationMetadata: { feature: "intent-classification", version: "2.1", }, }); const result = await openai.chat.completions.create({ model: "gpt-4o", messages: [{ role: "user", content: "I want to book a flight" }], }); ``` --- ## Streaming with Token Tracking For streaming calls, set `stream_options.include_usage` to capture token counts. ```typescript import { openai } from "./lib/openai"; const stream = await openai.chat.completions.create({ model: "gpt-4o", messages: [{ role: "user", content: "Tell me a story." }], stream: true, stream_options: { include_usage: true }, // REQUIRED for token tracking on streams }); for await (const chunk of stream) { const content = chunk.choices[0]?.delta?.content; if (content) process.stdout.write(content); } // Token usage automatically captured by observeOpenAI ``` **Without `stream_options: { include_usage: true }`**, OpenAI does not return token counts for streaming responses, and Langfuse will show `null` for token usage. --- ## Nested Inside Manual Traces Combine `observeOpenAI` with manual tracing for richer context. ```typescript import OpenAI from "openai"; import { observeOpenAI } from "@langfuse/openai"; import { startActiveObservation } from "@langfuse/tracing"; const openai = observeOpenAI(new OpenAI()); async function summarizeArticle(article: string): Promise<string> { return await startActiveObservation("summarize-article", async (span) => { span.update({ input: { articleLength: article.length } }); // OpenAI call is automatically a child of "summarize-article" const completion = await openai.chat.completions.create({ model: "gpt-4o", messages: [ { role: "developer", content: "Summarize the following article concisely.", }, { role: "user", content: article }, ], }); const summary = completion.choices[0].message.content ?? ""; span.update({ output: { summary, summaryLength: summary.length } }); return summary; }); } ``` --- ## Multi-Step Pipeline with OpenAI ```typescript import OpenAI from "openai"; import { observeOpenAI } from "@langfuse/openai"; import { startActiveObservation, propagateAttributes } from "@langfuse/tracing"; const openai = observeOpenAI(new OpenAI()); async function chatPipeline( userId: string, sessionId: string, message: string, ): Promise<string> { return await propagateAttributes( { userId, sessionId, tags: ["chat-v3"] }, async () => { return await startActiveObservation("chat-pipeline", async (rootSpan) => { rootSpan.update({ input: { message } }); // Step 1: Classify intent const intentResult = await startActiveObservation( "classify-intent", async () => { const completion = await openai.chat.completions.create({ model: "gpt-4o", messages: [ { role: "developer", content: "Classify the user intent as: question, command, or greeting.", }, { role: "user", content: message }, ], temperature: 0, }); return completion.choices[0].message.content ?? "unknown"; }, { asType: "agent" }, ); // Step 2: Generate response based on intent const response = await startActiveObservation( "generate-response", async () => { const completion = await openai.chat.completions.create({ model: "gpt-4o", messages: [ { role: "developer", content: `User intent: ${intentResult}. Respond appropriately.`, }, { role: "user", content: message }, ], }); return completion.choices[0].message.content ?? ""; }, { asType: "generation" }, ); rootSpan.update({ output: { intent: intentResult, response } }); return response; }); }, ); } ``` --- ## Limitations - `observeOpenAI()` does **NOT** support the OpenAI Assistants API (server-side state is incompatible) - Requires OpenAI SDK version >= 4.0.0 - Requires OTel `NodeSDK` to be initialized before using the wrapper - Token counts on streaming calls require explicit `stream_options: { include_usage: true }` --- _For tracing patterns, see [tracing.md](tracing.md). For setup, see [core.md](core.md)._ -
prompt-management.md 5.3 KB
# Langfuse -- Prompt Management Examples > Fetching prompts, compiling variables, text vs chat prompts, versioning, labels, caching, and linking prompts to traces. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Setup, environment config, flush/shutdown - [tracing.md](tracing.md) -- Manual tracing patterns - [openai-integration.md](openai-integration.md) -- OpenAI auto-instrumentation - [scores-datasets.md](scores-datasets.md) -- Scores and datasets --- ## Fetching and Compiling Text Prompts ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Fetch a text prompt (production label by default) const prompt = await langfuse.prompt.get("summarize-article"); // Compile with variables -- replaces {{variable}} placeholders const compiled = prompt.compile({ topic: "climate change", length: "200 words", }); // -> "Write a 200 words summary about climate change." // Use compiled prompt with your LLM const completion = await openai.chat.completions.create({ model: "gpt-4o", messages: [ { role: "developer", content: compiled }, { role: "user", content: articleText }, ], }); ``` --- ## Chat Prompts Chat prompts return an array of message objects ready for LLM APIs. ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Fetch a chat prompt const chatPrompt = await langfuse.prompt.get("customer-support", { type: "chat", }); // Compile with variables const messages = chatPrompt.compile({ userName: "Alice", productName: "Pro Plan", }); // -> [ // { role: "system", content: "You are helping Alice with their Pro Plan." }, // { role: "user", content: "{{user_message}}" } // ] // Use directly with OpenAI const completion = await openai.chat.completions.create({ model: "gpt-4o", messages: [ ...messages, { role: "user", content: "How do I cancel my subscription?" }, ], }); ``` --- ## Prompt Versioning and Labels ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Default: fetches "production" label const prodPrompt = await langfuse.prompt.get("classifier"); // Fetch a specific label (for A/B testing or staging) const stagingPrompt = await langfuse.prompt.get("classifier", { label: "staging", }); // Fetch a specific version number const v2Prompt = await langfuse.prompt.get("classifier", { version: 2, }); ``` --- ## Cache Control The SDK caches prompts by default. Control caching behavior for development or frequently updated prompts. ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Bypass cache -- always fetch from server const freshPrompt = await langfuse.prompt.get("classifier", { cacheTtlSeconds: 0, }); // Custom cache TTL (in seconds) const CACHE_TTL_SECONDS = 300; const cachedPrompt = await langfuse.prompt.get("classifier", { cacheTtlSeconds: CACHE_TTL_SECONDS, }); ``` **Note:** When a cached prompt expires, the SDK returns the stale prompt immediately while refreshing in the background. This prevents latency spikes from cache misses. --- ## Creating Prompts Programmatically ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Create a text prompt await langfuse.prompt.create({ name: "summarize-article", prompt: "Write a {{length}} summary about {{topic}}. Focus on key facts.", type: "text", }); // Create a chat prompt await langfuse.prompt.create({ name: "customer-support", prompt: [ { role: "system", content: "You are a support agent helping {{userName}}.", }, { role: "user", content: "{{user_message}}" }, ], type: "chat", }); ``` --- ## Linking Prompts to Traces When using `observeOpenAI`, pass the `langfusePrompt` option to link a prompt version to the generation trace. ```typescript import OpenAI from "openai"; import { observeOpenAI } from "@langfuse/openai"; import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); const openai = observeOpenAI(new OpenAI()); async function classifyWithManagedPrompt(text: string): Promise<string> { const prompt = await langfuse.prompt.get("classifier"); const compiled = prompt.compile({ text }); // Link prompt version to the trace for tracking const trackedOpenai = observeOpenAI(new OpenAI(), { langfusePrompt: prompt, generationName: "classify-text", }); const completion = await trackedOpenai.chat.completions.create({ model: "gpt-4o", messages: [{ role: "user", content: compiled }], }); return completion.choices[0].message.content ?? ""; } ``` --- ## Prompt with Fallback For guaranteed availability, provide a fallback prompt in case the Langfuse API is unreachable. ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); const FALLBACK_PROMPT = "Summarize the following text concisely: {{text}}"; const prompt = await langfuse.prompt.get("summarizer", { fallback: FALLBACK_PROMPT, }); // If Langfuse API is down, compile() uses the fallback template const compiled = prompt.compile({ text: articleContent }); ``` --- _For tracing patterns, see [tracing.md](tracing.md). For setup, see [core.md](core.md)._ -
scores-datasets.md 7.8 KB
# Langfuse -- Scores & Datasets Examples > Creating scores (numeric, categorical, boolean), attaching to traces/observations, datasets, experiments, and evaluation workflows. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Setup, environment config, flush/shutdown - [tracing.md](tracing.md) -- Manual tracing patterns - [openai-integration.md](openai-integration.md) -- OpenAI auto-instrumentation - [prompt-management.md](prompt-management.md) -- Prompt management --- ## Score Types ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Numeric score (float value) langfuse.score.create({ traceId: "trace-abc-123", name: "relevance", value: 0.92, dataType: "NUMERIC", }); // Categorical score (string value) langfuse.score.create({ traceId: "trace-abc-123", name: "tone", value: "professional", dataType: "CATEGORICAL", }); // Boolean score (0 or 1 as float, NOT true/false) langfuse.score.create({ traceId: "trace-abc-123", name: "contains-pii", value: 0, dataType: "BOOLEAN", }); ``` **Note:** Boolean scores use `0` (false) or `1` (true) as float values, not JavaScript `true`/`false`. --- ## Scoring Specific Observations ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Score a specific generation within a trace langfuse.score.create({ traceId: "trace-abc-123", observationId: "obs-def-456", name: "accuracy", value: 0.88, dataType: "NUMERIC", }); ``` --- ## Scoring Active Observations Score the currently active OTel span without needing trace/observation IDs. ```typescript import { startActiveObservation } from "@langfuse/tracing"; import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); await startActiveObservation("process-query", async (span) => { span.update({ input: { query: "What is AI?" } }); const result = await processQuery("What is AI?"); // Score the active observation (this span) langfuse.score.activeObservation({ name: "quality", value: 0.95, dataType: "NUMERIC", }); // Score the active trace (root trace) langfuse.score.activeTrace({ name: "user-satisfaction", value: "satisfied", dataType: "CATEGORICAL", }); span.update({ output: { result } }); }); // Flush in short-lived processes await langfuse.score.flush(); ``` --- ## User Feedback Scores Capture user feedback (thumbs up/down, ratings) as scores. ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // User thumbs up/down function handleFeedback(traceId: string, isPositive: boolean) { langfuse.score.create({ traceId, name: "user-feedback", value: isPositive ? 1 : 0, dataType: "BOOLEAN", comment: isPositive ? "User liked the response" : "User disliked the response", }); } // User star rating (1-5) function handleRating(traceId: string, rating: number) { langfuse.score.create({ traceId, name: "user-rating", value: rating, dataType: "NUMERIC", }); } ``` --- ## Session-Level Scores Score an entire session (multiple traces). ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Score a session langfuse.score.create({ sessionId: "session-xyz-789", name: "session-quality", value: "excellent", dataType: "CATEGORICAL", }); ``` --- ## Creating Datasets ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Create a dataset await langfuse.api.datasets.create({ name: "qa-benchmark", description: "Question-answer pairs for regression testing", metadata: { author: "Alice", created: "2025-10-01", type: "benchmark", }, }); // Add items to the dataset await langfuse.api.datasetItems.create({ datasetName: "qa-benchmark", input: { question: "What is the capital of France?", }, expectedOutput: { answer: "Paris", }, metadata: { difficulty: "easy", category: "geography", }, }); await langfuse.api.datasetItems.create({ datasetName: "qa-benchmark", input: { question: "Explain quantum entanglement in simple terms.", }, expectedOutput: { answer: "Quantum entanglement is when two particles are connected so that measuring one instantly affects the other, regardless of distance.", }, metadata: { difficulty: "hard", category: "physics", }, }); ``` --- ## Running Experiments Run your LLM function against a dataset and collect results. ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Get the dataset const dataset = await langfuse.dataset.get("qa-benchmark"); // Run experiment with a task function const result = await dataset.runExperiment({ name: "gpt-4o-baseline-v1", description: "Baseline experiment with gpt-4o", task: async ({ item }) => { // Your LLM function const completion = await openai.chat.completions.create({ model: "gpt-4o", messages: [ { role: "developer", content: "Answer the question accurately." }, { role: "user", content: item.input.question }, ], }); return completion.choices[0].message.content; }, }); console.log(`Experiment: ${result.runName}`); console.log(`Results: ${result.itemResults.length} items processed`); console.log(`URL: ${result.datasetRunUrl}`); ``` --- ## Experiments with Evaluators Add evaluators to automatically score experiment results. ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); const dataset = await langfuse.dataset.get("qa-benchmark"); const result = await dataset.runExperiment({ name: "gpt-4o-with-eval", task: async ({ item }) => { const completion = await openai.chat.completions.create({ model: "gpt-4o", messages: [{ role: "user", content: item.input.question }], }); return completion.choices[0].message.content; }, evaluators: [ // Per-item evaluator async ({ input, output, expectedOutput }) => { const isCorrect = output ?.toLowerCase() .includes(expectedOutput?.answer?.toLowerCase() ?? ""); return { name: "contains-answer", value: isCorrect ? 1 : 0, dataType: "BOOLEAN" as const, }; }, ], }); ``` --- ## Linking Dataset Items to Production Traces Connect dataset items to real production traces for context. ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Link a dataset item to a source trace await langfuse.api.datasetItems.create({ datasetName: "qa-benchmark", input: { question: "What is the capital of France?" }, expectedOutput: { answer: "Paris" }, sourceTraceId: "trace-from-production-123", sourceObservationId: "obs-456", }); ``` --- ## Versioned Datasets Run experiments against specific dataset snapshots. ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Get dataset at a specific version timestamp const versionTimestamp = new Date("2025-12-15T06:30:00").toISOString(); const versionedDataset = await langfuse.dataset.get("qa-benchmark", { version: versionTimestamp, }); const result = await versionedDataset.runExperiment({ name: "baseline-on-v1", description: "Running against December 2025 snapshot", task: async ({ item }) => { return await myLLMFunction(item.input); }, }); ``` --- ## Archiving Dataset Items ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Archive a dataset item by ID await langfuse.api.datasetItems.create({ id: "item-abc-123", status: "ARCHIVED", }); ``` --- _For tracing patterns, see [tracing.md](tracing.md). For setup, see [core.md](core.md)._ -
tracing.md 9.9 KB
# Langfuse -- Tracing Examples > Observation creation, nesting, context propagation, observation types, metadata, sessions, and user tracking. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Setup, environment config, flush/shutdown - [openai-integration.md](openai-integration.md) -- OpenAI auto-instrumentation - [prompt-management.md](prompt-management.md) -- Prompt management - [scores-datasets.md](scores-datasets.md) -- Scores and datasets --- ## startActiveObservation -- Context Manager The primary tracing method. Creates an observation, sets it as active context, and ends it automatically. ```typescript import { startActiveObservation } from "@langfuse/tracing"; async function handleChatMessage(message: string): Promise<string> { return await startActiveObservation("chat-handler", async (span) => { span.update({ input: { message } }); const intent = await classifyIntent(message); const response = await generateResponse(intent, message); span.update({ output: { response, intent } }); return response; }); } ``` --- ## Nested Observations Child observations automatically inherit parent context when created during the parent's active scope. ```typescript import { startActiveObservation } from "@langfuse/tracing"; async function ragPipeline(query: string): Promise<string> { return await startActiveObservation("rag-pipeline", async (rootSpan) => { rootSpan.update({ input: { query } }); // Child 1: Retrieve documents (typed as retriever) const docs = await startActiveObservation( "retrieve-docs", async (span) => { span.update({ input: { query } }); const results = await vectorStore.search(query); span.update({ output: { documentCount: results.length } }); return results; }, { asType: "retriever" }, ); // Child 2: Generate answer (typed as generation) const answer = await startActiveObservation( "generate-answer", async (span) => { span.update({ input: { query, contextDocs: docs.length }, model: "gpt-4o", }); const result = await callLLM(query, docs); span.update({ output: { answer: result } }); return result; }, { asType: "generation" }, ); rootSpan.update({ output: { answer } }); return answer; }); } ``` --- ## observe() Function Wrapper Wraps functions for declarative tracing. Inputs and outputs are captured automatically. ```typescript import { observe } from "@langfuse/tracing"; const classifyIntent = observe( async (query: string): Promise<string> => { const completion = await openai.chat.completions.create({ model: "gpt-4o", messages: [ { role: "system", content: "Classify the user intent." }, { role: "user", content: query }, ], }); return completion.choices[0].message.content ?? "unknown"; }, { name: "classify-intent", asType: "generation" }, ); const searchDocuments = observe( async (query: string): Promise<Document[]> => { return await vectorStore.similaritySearch(query); }, { name: "search-documents", asType: "retriever" }, ); // Usage -- both functions are automatically traced const intent = await classifyIntent("Book a flight to Paris"); const docs = await searchDocuments("Paris flights availability"); ``` ### Controlling Input/Output Capture ```typescript // Disable input capture (for sensitive data) const secureFn = observe( async (sensitiveInput: string) => { /* ... */ }, { name: "secure-fn", captureInput: false }, ); // Disable output capture const largeFn = observe( async (input: string) => { /* returns large object */ }, { name: "large-fn", captureOutput: false }, ); ``` --- ## Manual Observations with startObservation For cases where you need explicit control over observation lifecycle. ```typescript import { startObservation } from "@langfuse/tracing"; async function processItem(item: WorkItem): Promise<void> { const span = startObservation( "process-item", { input: { itemId: item.id } }, { asType: "tool" }, ); try { const result = await performWork(item); span.update({ output: { result } }); } catch (error) { span.update({ output: { error: error instanceof Error ? error.message : "Unknown error", }, }); throw error; } finally { span.end(); // REQUIRED -- must explicitly end } } ``` --- ## Token Usage on Manual Generations When manually creating generation observations (not using `observeOpenAI`), report token counts via `usageDetails`: ```typescript import { startObservation } from "@langfuse/tracing"; const generation = startObservation( "llm-call", { model: "gpt-4o", input: [{ role: "user", content: "What is the capital of France?" }], }, { asType: "generation" }, ); // ... perform LLM call ... generation.update({ output: { content: "The capital of France is Paris." }, usageDetails: { input: 10, output: 5, total: 15, }, }); generation.end(); ``` **Note:** `total` is automatically calculated if omitted. You can also add custom token metrics (e.g., `cache_read_input_tokens`). --- ## Observation Types Use `asType` to give semantic meaning to observations. ```typescript import { startActiveObservation } from "@langfuse/tracing"; // Agent step -- decision-making await startActiveObservation( "plan-next-action", async (span) => { span.update({ input: { state: currentState } }); const action = await agent.decide(currentState); span.update({ output: { action } }); }, { asType: "agent" }, ); // Tool call -- external API await startActiveObservation( "call-weather-api", async (span) => { span.update({ input: { location: "Paris" } }); const weather = await weatherAPI.get("Paris"); span.update({ output: weather }); }, { asType: "tool" }, ); // Retriever -- vector search await startActiveObservation( "vector-search", async (span) => { span.update({ input: { query: "climate change" } }); const docs = await vectorStore.search("climate change"); span.update({ output: { count: docs.length } }); }, { asType: "retriever" }, ); // Evaluator -- quality check await startActiveObservation( "check-relevance", async (span) => { span.update({ input: { answer, query } }); const score = await evaluateRelevance(answer, query); span.update({ output: { score } }); }, { asType: "evaluator" }, ); // Embedding call await startActiveObservation( "create-embedding", async (span) => { span.update({ input: { text: "Hello world" }, model: "text-embedding-3-small", }); const embedding = await createEmbedding("Hello world"); span.update({ output: { dimensions: embedding.length } }); }, { asType: "embedding" }, ); ``` --- ## User and Session Tracking Use `propagateAttributes()` to attach user, session, and metadata to all nested observations. It wraps a callback -- all observations created inside the callback inherit the attributes. ```typescript import { startActiveObservation, propagateAttributes } from "@langfuse/tracing"; async function handleUserRequest( userId: string, sessionId: string, query: string, ) { return await propagateAttributes( { userId, sessionId, tags: ["production", "v2"], metadata: { environment: "prod", region: "eu-west-1" }, }, async () => { return await startActiveObservation("user-request", async (span) => { span.update({ input: { query } }); const result = await processQuery(query); span.update({ output: { result } }); return result; }); }, ); } ``` --- ## Updating Active Observation Without Reference ```typescript import { startActiveObservation, updateActiveObservation, } from "@langfuse/tracing"; async function deeplyNestedFunction() { // Update the currently active observation from anywhere in the call stack updateActiveObservation({ metadata: { step: "validation", validated: true }, }); } await startActiveObservation("outer", async () => { await deeplyNestedFunction(); // Updates "outer" observation's metadata }); ``` --- ## Trace and Span ID Access ```typescript import { startActiveObservation, getActiveTraceId, getActiveSpanId, createTraceId, } from "@langfuse/tracing"; await startActiveObservation("my-trace", async () => { const traceId = getActiveTraceId(); const spanId = getActiveSpanId(); console.log(`Trace: ${traceId}, Span: ${spanId}`); }); // Deterministic trace ID for correlation with external systems const traceId = createTraceId("order-12345"); ``` --- ## Time to First Token (TTFT) Tracking ```typescript import { startActiveObservation } from "@langfuse/tracing"; await startActiveObservation( "llm-call", async (span) => { span.update({ input: { prompt: "Tell me a story" }, model: "gpt-4o" }); let firstTokenReceived = false; const stream = await openai.chat.completions.create({ model: "gpt-4o", messages: [{ role: "user", content: "Tell me a story" }], stream: true, }); for await (const chunk of stream) { if (!firstTokenReceived) { span.update({ completionStartTime: new Date().toISOString() }); firstTokenReceived = true; } // Process chunk... } }, { asType: "generation" }, ); ``` --- ## Cross-Service Context Propagation Use `asBaggage: true` to propagate context across HTTP boundaries via W3C Baggage headers. ```typescript import { propagateAttributes } from "@langfuse/tracing"; // In the upstream service -- wraps a callback await propagateAttributes( { userId: "user-123", sessionId: "session-456" }, async () => { // All observations created here inherit userId and sessionId await handleRequest(); }, { asBaggage: true }, // Propagates via HTTP headers ); ``` --- _For setup and configuration, see [core.md](core.md). For API reference, see [reference.md](../reference.md)._
-
-
reference.md 9.6 KB
# Langfuse Quick Reference > Package index, environment variables, tracing API, observation types, score methods, and client API. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples. --- ## Package Installation ```bash # Core tracing (always required for observability) npm install @langfuse/tracing @langfuse/otel @opentelemetry/sdk-node # OpenAI auto-instrumentation (if using OpenAI SDK) npm install @langfuse/openai # Client API (prompts, scores, datasets) npm install @langfuse/client ``` ### Package Overview | Package | Purpose | | ------------------------- | ---------------------------------------------------------------------------- | | `@langfuse/tracing` | `startActiveObservation`, `observe`, `startObservation`, context propagation | | `@langfuse/otel` | `LangfuseSpanProcessor` for OTel NodeSDK | | `@opentelemetry/sdk-node` | OpenTelemetry SDK (peer dependency) | | `@langfuse/client` | `LangfuseClient` for prompts, scores, datasets | | `@langfuse/openai` | `observeOpenAI()` wrapper for OpenAI SDK | | `@langfuse/core` | Shared utilities and types (transitive) | --- ## Environment Variables | Variable | Required | Default | Description | | ------------------------- | -------- | ---------------------------- | ------------------------------------- | | `LANGFUSE_SECRET_KEY` | Yes | -- | Secret API key (`sk-lf-...`) | | `LANGFUSE_PUBLIC_KEY` | Yes | -- | Public API key (`pk-lf-...`) | | `LANGFUSE_BASE_URL` | Yes | `https://cloud.langfuse.com` | Langfuse instance URL | | `LANGFUSE_SAMPLE_RATE` | No | `1.0` | Trace sampling rate (0.0-1.0) | | `LANGFUSE_LOG_LEVEL` | No | `WARN` | SDK log level (DEBUG/INFO/WARN/ERROR) | | `LANGFUSE_FLUSH_AT` | No | `10` | Flush after N queued events | | `LANGFUSE_FLUSH_INTERVAL` | No | `1` (seconds) | Flush interval in seconds | | `LANGFUSE_TIMEOUT` | No | `5000` (ms) | HTTP request timeout | --- ## Tracing API (`@langfuse/tracing`) ### Context Managers ```typescript import { startActiveObservation, startObservation, observe, updateActiveObservation, propagateAttributes, getActiveTraceId, getActiveSpanId, createTraceId, } from "@langfuse/tracing"; // Auto-context, auto-end await startActiveObservation("name", async (span) => { span.update({ input: {...}, output: {...} }); }, { asType: "generation" }); // Manual control (requires explicit .end()) const span = startObservation("name", { input: {...} }, { asType: "tool" }); span.update({ output: {...} }); span.end(); // Function wrapper (auto-captures inputs/outputs) const fn = observe(async (input: string) => { return output; }, { name: "fn-name", asType: "agent", }); // Update active span without reference updateActiveObservation({ metadata: { key: "value" } }); // Propagate attributes to all nested observations (wraps a callback) await propagateAttributes( { userId: "user-1", sessionId: "session-1", tags: ["prod"] }, async () => { /* observations created here inherit attributes */ }, ); // Get current trace/span IDs const traceId = getActiveTraceId(); const spanId = getActiveSpanId(); // Deterministic trace ID from seed const traceId = createTraceId("my-correlation-key"); ``` ### Span Update Properties ```typescript span.update({ input: any, // Observation input output: any, // Observation output metadata: Record<string, any>, // Custom key-value metadata model: string, // LLM model name (for generations) completionStartTime: string, // ISO timestamp for TTFT tracking }); ``` --- ## Observation Types (`asType`) | Type | Semantic Meaning | Dashboard Icon | | -------------- | ------------------------------------ | -------------- | | `"span"` | Generic duration (default) | -- | | `"event"` | Point-in-time occurrence | -- | | `"generation"` | LLM call (prompt/completion, tokens) | AI icon | | `"agent"` | AI agent decision step | Agent icon | | `"tool"` | External tool/API call | Tool icon | | `"chain"` | Link between application steps | Chain icon | | `"retriever"` | Vector store / DB retrieval | Search icon | | `"evaluator"` | Quality assessment function | Check icon | | `"embedding"` | Embedding model call | Vector icon | | `"guardrail"` | Content safety / jailbreak check | Shield icon | --- ## OpenAI Integration (`@langfuse/openai`) ```typescript import OpenAI from "openai"; import { observeOpenAI } from "@langfuse/openai"; // Wrap OpenAI client -- all calls automatically traced const openai = observeOpenAI(new OpenAI()); // Optional: custom trace attributes per call const openai = observeOpenAI(new OpenAI(), { generationName: "classify-intent", sessionId: "session-123", userId: "user-456", tags: ["production"], generationMetadata: { feature: "intent-classifier" }, }); ``` **Auto-captured data:** model, token counts (prompt/completion), estimated cost, latency, time-to-first-token (streaming), errors, function calls. **Stream token tracking:** Set `stream_options: { include_usage: true }` on OpenAI streaming calls. --- ## Client API (`@langfuse/client`) ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); ``` ### Prompt Manager (`langfuse.prompt`) ```typescript // Get text prompt (production label by default) const prompt = await langfuse.prompt.get("name"); const compiled = prompt.compile({ var1: "value1" }); // Get chat prompt const chatPrompt = await langfuse.prompt.get("name", { type: "chat" }); const messages = chatPrompt.compile({ var1: "value1" }); // Specific version or label await langfuse.prompt.get("name", { version: 2 }); await langfuse.prompt.get("name", { label: "staging" }); // Bypass cache await langfuse.prompt.get("name", { cacheTtlSeconds: 0 }); // Create prompt await langfuse.prompt.create({ name: "name", prompt: "Template {{var}}", type: "text", }); ``` ### Score Manager (`langfuse.score`) ```typescript // By trace ID langfuse.score.create({ traceId, name: "quality", value: 0.9, dataType: "NUMERIC", }); langfuse.score.create({ traceId, name: "label", value: "good", dataType: "CATEGORICAL", }); langfuse.score.create({ traceId, name: "passed", value: 1, dataType: "BOOLEAN", }); // By observation ID langfuse.score.create({ traceId, observationId, name: "accuracy", value: 0.8, dataType: "NUMERIC", }); // By active OTel span langfuse.score.activeObservation({ name: "quality", value: 0.9, dataType: "NUMERIC", }); langfuse.score.activeTrace({ name: "quality", value: 0.9, dataType: "NUMERIC", }); // By OTel span reference langfuse.score.observation( { otelSpan }, { name: "quality", value: 0.9, dataType: "NUMERIC" }, ); langfuse.score.trace( { otelSpan }, { name: "quality", value: 0.9, dataType: "NUMERIC" }, ); // Flush pending scores await langfuse.score.flush(); ``` ### Dataset Manager (`langfuse.dataset`) ```typescript // Get dataset const dataset = await langfuse.dataset.get("dataset-name"); // Run experiment const result = await dataset.runExperiment({ name: "experiment-v1", task: async ({ item }) => { const output = await myLLMFunction(item.input); return output; }, }); ``` ### Direct API Access (`langfuse.api`) ```typescript // Create dataset await langfuse.api.datasets.create({ name: "my-dataset" }); // Add dataset item await langfuse.api.datasetItems.create({ datasetName: "my-dataset", input: { text: "hello" }, expectedOutput: { text: "world" }, }); // Get trace await langfuse.api.trace.get("trace-id"); // Get observation await langfuse.api.observations.get("observation-id"); ``` ### Utility Methods ```typescript // Get dashboard URL for a trace langfuse.getTraceUrl("trace-id"); // Flush all pending client data (scores, prompts) await langfuse.flush(); // Graceful shutdown await langfuse.shutdown(); ``` ### Flushing OTel Spans (Short-Lived Processes) ```typescript // Flush pending spans via the processor await langfuseSpanProcessor.forceFlush(); // Or shutdown the entire OTel SDK (flushes + closes) await sdk.shutdown(); ``` --- ## LangfuseSpanProcessor Configuration ```typescript import { LangfuseSpanProcessor, isDefaultExportSpan } from "@langfuse/otel"; const processor = new LangfuseSpanProcessor({ // Credentials (optional -- falls back to env vars) publicKey: "pk-lf-...", secretKey: "sk-lf-...", baseUrl: "https://cloud.langfuse.com", // Data masking mask: ({ data }) => data.replace(/\b\d{16}\b/g, "***MASKED***"), // Span filtering shouldExportSpan: ({ otelSpan }) => isDefaultExportSpan(otelSpan) || otelSpan.instrumentationScope.name.startsWith("my-app"), }); ``` --- ## Score Data Types | Type | Value Type | Example | | ------------- | ----------------- | ----------------------- | | `NUMERIC` | `number` (float) | `0.95`, `42`, `-1.5` | | `CATEGORICAL` | `string` | `"good"`, `"incorrect"` | | `BOOLEAN` | `number` (0 or 1) | `0` (false), `1` (true) | -
SKILL.md 19.5 KB
--- name: ai-observability-langfuse description: LLM observability with Langfuse — OpenTelemetry-based tracing, evaluations, prompt management, datasets, and production best practices --- # Langfuse Observability Patterns > **Quick Guide:** Use the Langfuse TypeScript SDK (built on OpenTelemetry) to add observability to LLM applications. Install `@langfuse/tracing`, `@langfuse/otel`, and `@opentelemetry/sdk-node` for core tracing. Use `startActiveObservation()` for automatic context propagation or `observe()` to wrap functions. Use `@langfuse/openai` with `observeOpenAI()` for zero-config OpenAI tracing. Use `LangfuseClient` from `@langfuse/client` for prompt management, scores, and datasets. Always call `forceFlush()` or `sdk.shutdown()` in short-lived processes. --- <critical_requirements> ## CRITICAL: Before Using This Skill > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST import and register `instrumentation.ts` at the top of your entry point BEFORE any other imports -- OpenTelemetry must instrument modules before they are loaded)** **(You MUST call `forceFlush()` or `sdk.shutdown()` in short-lived processes (serverless, scripts, CLI tools) -- events are batched and will be lost without explicit flushing)** **(You MUST use `@langfuse/openai` with `observeOpenAI()` for OpenAI SDK tracing -- do NOT manually create generation observations for OpenAI calls when the wrapper handles it automatically)** **(You MUST set `LANGFUSE_SECRET_KEY`, `LANGFUSE_PUBLIC_KEY`, and `LANGFUSE_BASE_URL` via environment variables -- never hardcode credentials)** **(You MUST use `startActiveObservation()` or `observe()` for nested tracing -- manual `startObservation()` requires explicit `.end()` calls and does NOT propagate context automatically)** </critical_requirements> --- **Auto-detection:** Langfuse, langfuse, @langfuse/tracing, @langfuse/otel, @langfuse/client, @langfuse/openai, LangfuseSpanProcessor, LangfuseClient, startActiveObservation, startObservation, observeOpenAI, langfuse.score, langfuse.prompt, langfuse.dataset, LANGFUSE_SECRET_KEY, LANGFUSE_PUBLIC_KEY, forceFlush **When to use:** - Adding observability and tracing to LLM application code (any provider) - Wrapping OpenAI SDK calls for automatic token/cost tracking - Managing prompt templates with versioning, labels, and variable compilation - Evaluating LLM output quality with scores (numeric, categorical, boolean) - Running experiments against datasets for regression testing - Tracking sessions, users, and metadata across multi-turn conversations - Monitoring LLM costs and token usage in production **Key patterns covered:** - OpenTelemetry setup with `LangfuseSpanProcessor` - Tracing with `startActiveObservation`, `observe`, and manual `startObservation` - Observation types (span, generation, agent, tool, retriever, evaluator, embedding, chain, guardrail) - OpenAI SDK auto-instrumentation with `observeOpenAI()` - Prompt management (get, compile, text vs chat prompts, versioning) - Scores and evaluations (numeric, categorical, boolean) - Datasets and experiments for testing - Flush, shutdown, and lifecycle management **When NOT to use:** - You only need basic `console.log` debugging -- Langfuse is for structured production observability - You want provider-specific tracing built into an AI SDK -- check if your framework has native observability - You need APM/infrastructure monitoring (CPU, memory, HTTP latency) -- use a general-purpose observability tool --- ## Examples Index - [Core: Setup & Configuration](examples/core.md) -- OpenTelemetry setup, instrumentation file, client init, flush/shutdown - [Tracing](examples/tracing.md) -- startActiveObservation, observe, manual tracing, nesting, observation types, metadata - [OpenAI Integration](examples/openai-integration.md) -- observeOpenAI wrapper, streaming, token tracking, custom attributes - [Prompt Management](examples/prompt-management.md) -- getPrompt, compile, text vs chat, versioning, caching - [Scores & Datasets](examples/scores-datasets.md) -- Numeric/categorical/boolean scores, datasets, experiments - [Quick API Reference](reference.md) -- Package index, environment variables, observation types, score methods --- <philosophy> ## Philosophy Langfuse provides **open-source LLM observability** built on OpenTelemetry. The SDK (v4+, August 2025) is a ground-up rewrite using OTel as the tracing backbone, meaning traces integrate naturally with the broader observability ecosystem. **Core principles:** 1. **OpenTelemetry-native** -- Built on OTel spans and context propagation. Langfuse observations are wrappers around OTel spans with LLM-specific attributes (model, tokens, cost). This means any OTel-compatible instrumentation library works alongside Langfuse. 2. **Zero-latency tracing** -- All trace events are queued locally and flushed in background batches. Your application's response time is not affected by observability. 3. **Modular packages** -- `@langfuse/tracing` for instrumentation, `@langfuse/client` for prompts/scores/datasets, `@langfuse/openai` for OpenAI auto-instrumentation. Install only what you need. 4. **Context-first** -- `startActiveObservation()` automatically propagates parent-child relationships. Nested observations inherit context without manual ID threading. 5. **Observation types** -- LLM-specific types (`generation`, `agent`, `tool`, `retriever`, `evaluator`, `embedding`) provide semantic meaning to traces, enabling richer dashboard views and filtering. </philosophy> --- <patterns> ## Core Patterns ### Pattern 1: OpenTelemetry Setup Create an `instrumentation.ts` file and import it at the top of your entry point. ```typescript // instrumentation.ts import { NodeSDK } from "@opentelemetry/sdk-node"; import { LangfuseSpanProcessor } from "@langfuse/otel"; const sdk = new NodeSDK({ spanProcessors: [new LangfuseSpanProcessor()], }); sdk.start(); export { sdk }; ``` ```typescript // index.ts -- import instrumentation FIRST import "./instrumentation"; // All other imports AFTER instrumentation import { startActiveObservation } from "@langfuse/tracing"; ``` **Why good:** OTel must instrument modules before they are loaded; importing instrumentation first ensures all subsequent imports are traced automatically ```typescript // BAD: importing instrumentation after other modules import { startActiveObservation } from "@langfuse/tracing"; import "./instrumentation"; // TOO LATE -- tracing won't capture earlier imports ``` **Why bad:** Auto-instrumentation of LLM SDKs requires OTel to be initialized before those modules are imported **See:** [examples/core.md](examples/core.md) for environment variables, sampling, masking, and production configuration --- ### Pattern 2: Tracing with startActiveObservation The primary instrumentation pattern. Creates an observation, makes it the active context, and automatically ends it when the callback completes. ```typescript import { startActiveObservation } from "@langfuse/tracing"; async function handleRequest(query: string): Promise<string> { return await startActiveObservation("handle-request", async (span) => { span.update({ input: { query } }); // Nested observation -- automatically becomes a child const result = await startActiveObservation( "process-query", async (child) => { child.update({ input: { query } }); const answer = await callLLM(query); child.update({ output: { answer } }); return answer; }, ); span.update({ output: { result } }); return result; }); } ``` **Why good:** Automatic context propagation, automatic end on callback completion, nesting creates parent-child hierarchy without manual ID management ```typescript // BAD: using startObservation without ending it import { startObservation } from "@langfuse/tracing"; const span = startObservation("my-span"); await doWork(); // span.end() never called -- observation stays open forever ``` **Why bad:** Manual `startObservation` requires explicit `.end()` calls; forgetting creates open-ended observations **See:** [examples/tracing.md](examples/tracing.md) for observe wrapper, observation types, metadata, and manual tracing --- ### Pattern 3: The observe() Wrapper Wraps a function to automatically capture inputs, outputs, timings, and errors. ```typescript import { observe } from "@langfuse/tracing"; const classifyIntent = observe( async (query: string) => { const result = await callLLM(query); return result.intent; }, { name: "classify-intent", asType: "generation" }, ); // Usage -- automatically traced const intent = await classifyIntent("Book a flight to Paris"); ``` **Why good:** Declarative tracing, inputs/outputs captured automatically, `asType` tags the observation type for richer dashboard filtering --- ### Pattern 4: OpenAI Auto-Instrumentation Use `observeOpenAI()` to wrap the OpenAI client for automatic tracing of all calls. ```typescript import OpenAI from "openai"; import { observeOpenAI } from "@langfuse/openai"; const openai = observeOpenAI(new OpenAI()); // All calls automatically traced with model, tokens, cost const completion = await openai.chat.completions.create({ model: "gpt-4o", messages: [{ role: "user", content: "Hello" }], }); ``` **Why good:** Zero manual instrumentation, captures model name, token counts, estimated costs, latency, and streaming metrics automatically ```typescript // BAD: manually creating generation observations for OpenAI calls await startActiveObservation("openai-call", async (span) => { const result = await rawOpenai.chat.completions.create({ ... }); span.update({ model: "gpt-4o", input: messages, output: result.choices[0].message.content, }); }, { asType: "generation" }); ``` **Why bad:** `observeOpenAI` handles all of this automatically with more accurate token/cost data; manual tracking is error-prone and duplicates effort **See:** [examples/openai-integration.md](examples/openai-integration.md) for streaming, custom attributes, and token tracking on streams --- ### Pattern 5: Prompt Management Fetch versioned prompts, compile with variables, and link to traces. ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Fetch a text prompt (production label by default) const prompt = await langfuse.prompt.get("summarize-article"); const compiled = prompt.compile({ topic: "AI safety", length: "brief" }); // -> "Write a brief summary about AI safety." // Fetch a chat prompt const chatPrompt = await langfuse.prompt.get("assistant-v2", { type: "chat" }); const messages = chatPrompt.compile({ userName: "Alice" }); // -> [{ role: "system", content: "You are helping Alice..." }, ...] ``` **Why good:** Centralized prompt management with versioning, labels for A/B testing, variable compilation, and built-in caching **See:** [examples/prompt-management.md](examples/prompt-management.md) for versioning, labels, cache control, and linking prompts to traces --- ### Pattern 6: Scores and Evaluations Attach quality measurements to traces and observations. ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Numeric score langfuse.score.create({ traceId: "trace-123", name: "relevance", value: 0.95, dataType: "NUMERIC", }); // Categorical score langfuse.score.create({ traceId: "trace-123", name: "quality", value: "good", dataType: "CATEGORICAL", }); // Boolean score (0 or 1) langfuse.score.create({ traceId: "trace-123", name: "contains-hallucination", value: 0, dataType: "BOOLEAN", }); // Score a specific observation within a trace langfuse.score.create({ traceId: "trace-123", observationId: "obs-456", name: "accuracy", value: 0.88, dataType: "NUMERIC", }); // Flush in short-lived processes await langfuse.score.flush(); ``` **Why good:** Three data types cover all evaluation needs, scores attach at trace or observation level, fire-and-forget API with batching **See:** [examples/scores-datasets.md](examples/scores-datasets.md) for active observation scoring, session scores, datasets, and experiments --- ### Pattern 7: Flush and Shutdown Always flush in short-lived processes. The SDK batches events and sends them asynchronously. ```typescript import { sdk } from "./instrumentation"; import { LangfuseClient } from "@langfuse/client"; import { LangfuseSpanProcessor } from "@langfuse/otel"; const langfuse = new LangfuseClient(); async function main() { // ... do work ... // Flush scores await langfuse.score.flush(); // Shutdown OTel SDK (flushes all pending spans) await sdk.shutdown(); } main(); ``` **Why good:** Explicit flush/shutdown ensures all events are sent before the process exits; without this, data is silently lost in serverless and scripts ```typescript // BAD: exiting without flushing async function handler() { await startActiveObservation("my-trace", async (span) => { span.update({ output: "done" }); }); // Process exits -- batched events never sent } ``` **Why bad:** Langfuse batches events locally; if the process exits before the flush interval, events are lost </patterns> --- <performance> ## Performance Optimization ### Sampling for High-Volume Applications Reduce costs by sampling a subset of traces: ```typescript import { TraceIdRatioBasedSampler } from "@opentelemetry/sdk-trace-base"; const sdk = new NodeSDK({ sampler: new TraceIdRatioBasedSampler(0.2), // Sample 20% of traces spanProcessors: [new LangfuseSpanProcessor()], }); ``` Or via environment variable: `LANGFUSE_SAMPLE_RATE=0.2` ### Key Optimization Patterns - **Batch flush tuning** -- Configure `LANGFUSE_FLUSH_AT` (default 10) and `LANGFUSE_FLUSH_INTERVAL` (default 1s) for your workload - **Span filtering** -- Use `shouldExportSpan` on `LangfuseSpanProcessor` to drop noisy non-LLM spans - **Data masking** -- Redact PII before transmission with the `mask` option to avoid storing sensitive data - **Stream token tracking** -- Set `stream_options: { include_usage: true }` on OpenAI streaming calls so `observeOpenAI` captures token counts </performance> --- <decision_framework> ## Decision Framework ### Which Packages to Install ``` What do you need? +-- Tracing LLM calls? | +-- YES -> npm install @langfuse/tracing @langfuse/otel @opentelemetry/sdk-node | +-- Also using OpenAI SDK? | +-- YES -> npm install @langfuse/openai +-- Prompt management, scores, or datasets? | +-- YES -> npm install @langfuse/client +-- Both tracing AND client features? +-- YES -> Install all: @langfuse/tracing @langfuse/otel @opentelemetry/sdk-node @langfuse/client ``` ### Which Tracing Method to Use ``` How do you want to instrument? +-- Wrapping a function? -> observe() (declarative, auto-captures inputs/outputs) +-- Block of code with nesting? -> startActiveObservation() (context propagation, auto-end) +-- Need manual start/end control? -> startObservation() (requires explicit .end()) +-- OpenAI SDK calls? -> observeOpenAI() (zero-config auto-instrumentation) +-- Update active span without reference? -> updateActiveObservation() ``` ### Which Observation Type (asType) ``` What is this observation? +-- LLM call (prompt -> completion) -> "generation" +-- AI agent decision-making step -> "agent" +-- External API or function call -> "tool" +-- Vector store or DB retrieval -> "retriever" +-- Quality assessment step -> "evaluator" +-- Embedding creation -> "embedding" +-- Link between application steps -> "chain" +-- Content safety / jailbreak check -> "guardrail" +-- Generic duration operation -> "span" (default) +-- Point-in-time event -> "event" ``` </decision_framework> --- <red_flags> ## RED FLAGS **High Priority Issues:** - Not importing `instrumentation.ts` before other modules (auto-instrumentation silently fails) - Exiting short-lived processes without `forceFlush()` or `sdk.shutdown()` (events are silently lost) - Hardcoding `LANGFUSE_SECRET_KEY` or `LANGFUSE_PUBLIC_KEY` in source code (credential exposure) - Using manual generation observations when `observeOpenAI()` would handle it automatically (duplicated effort, less accurate data) - Using `startObservation()` without calling `.end()` (observation stays open indefinitely) **Medium Priority Issues:** - Not setting `stream_options: { include_usage: true }` on OpenAI streaming calls (token counts missing from `observeOpenAI` traces) - Forgetting to call `langfuse.score.flush()` in short-lived processes (scores are batched and may be lost) - Using `startObservation()` when `startActiveObservation()` would work (no automatic context propagation or auto-end) - Not using `asType` on observations (all observations appear as generic spans, losing semantic meaning) - Not setting `LANGFUSE_BASE_URL` for self-hosted instances (defaults to cloud.langfuse.com) **Common Mistakes:** - Importing `@langfuse/openai` without setting up the OTel `NodeSDK` first -- the OpenAI wrapper requires OTel context to send traces - Confusing `LangfuseClient` (from `@langfuse/client`, for prompts/scores/datasets) with the OTel tracing functions (from `@langfuse/tracing`) - Using `prompt.compile()` without matching all `{{variable}}` placeholders -- unmatched variables remain as literal `{{name}}` in output - Calling `langfuse.score.create()` with a `value` of type `string` for `NUMERIC` scores or `number` for `CATEGORICAL` scores (type mismatch) - Running dataset experiments without OTel setup -- experiment tasks run inside `startActiveObservation` which requires OTel **Gotchas & Edge Cases:** - `observeOpenAI()` does NOT support the OpenAI Assistants API -- only Chat Completions and Responses API - The SDK's default span filter only exports Langfuse and GenAI spans. If you use a custom instrumentation library, you must configure `shouldExportSpan` to include it. - `LangfuseClient.prompt.get()` caches prompts with a default TTL. If you update a prompt and don't see changes, set `cacheTtlSeconds: 0` to bypass caching. - Boolean scores use float values (`0` or `1`), not JavaScript booleans (`true`/`false`). - Self-hosted Langfuse requires platform version >= 3.95.0 for TypeScript SDK v4 compatibility. - `score.create()` is fire-and-forget (synchronous) -- it queues the score for batched delivery. You only need `await` on `flush()`. - Dataset names with slashes (`evaluation/qa-dataset`) must be URL-encoded when used as path parameters. - The v4+ SDK is a complete rewrite from v3 -- `Langfuse` class, `trace()`, `span()`, `generation()` from v3 are replaced by OTel-based APIs. </red_flags> --- <critical_reminders> ## CRITICAL REMINDERS > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST import and register `instrumentation.ts` at the top of your entry point BEFORE any other imports -- OpenTelemetry must instrument modules before they are loaded)** **(You MUST call `forceFlush()` or `sdk.shutdown()` in short-lived processes (serverless, scripts, CLI tools) -- events are batched and will be lost without explicit flushing)** **(You MUST use `@langfuse/openai` with `observeOpenAI()` for OpenAI SDK tracing -- do NOT manually create generation observations for OpenAI calls when the wrapper handles it automatically)** **(You MUST set `LANGFUSE_SECRET_KEY`, `LANGFUSE_PUBLIC_KEY`, and `LANGFUSE_BASE_URL` via environment variables -- never hardcode credentials)** **(You MUST use `startActiveObservation()` or `observe()` for nested tracing -- manual `startObservation()` requires explicit `.end()` calls and does NOT propagate context automatically)** **Failure to follow these rules will produce silent data loss, missing traces, or credential exposure in LLM observability.** </critical_reminders>
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.