ai-observability-promptfoo
Testing and evaluation framework for LLM prompts and applications -- promptfooconfig.yaml, assertions, model-graded evals, red teaming, CI/CD integration, custom providers, and comparative evaluation
Install
npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/ai-observability-promptfoo/skills/ai-observability-promptfoo
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
git clone https://github.com/agents-inc/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Promptfoo Patterns
Quick Guide: Use promptfoo for systematic LLM evaluation. Define prompts, providers, and test cases in
promptfooconfig.yaml. Use assertion types (contains,is-json,llm-rubric,similar,cost,latency) to validate outputs. Usepromptfoo evalto run (exits with code 100 on test failures),promptfoo viewfor results UI. Use model-graded assertions (llm-rubric,factuality) for subjective quality. Usepromptfoo redteam runfor security scanning. Use--shareflag orpromptfoo shareto share results. All provider API keys come from environment variables -- never hardcode them.
<critical_requirements>
CRITICAL: Before Using This Skill
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST define test cases with explicit assert arrays -- tests without assertions only capture output without validating it)
(You MUST use llm-rubric for subjective quality evaluation -- do NOT rely solely on deterministic assertions for natural language output)
(You MUST set threshold on similarity and model-graded assertions -- omitting thresholds uses defaults that may not match your quality bar)
(You MUST use environment variables for all API keys -- never hardcode keys in promptfooconfig.yaml or provider configs)
(You MUST verify promptfoo eval exit code in CI pipelines -- it returns exit code 100 on test failures, exit code 1 on other errors)
</critical_requirements>
Auto-detection: promptfoo, promptfooconfig, promptfooconfig.yaml, promptfoo eval, promptfoo view, promptfoo redteam, llm-rubric, model-graded-closedqa, promptfoo share, promptfoo cache, assertion type, LLM evaluation, prompt testing, red teaming, PROMPTFOO_CONFIG
When to use:
- Writing or evaluating LLM prompts across one or more providers
- Setting up automated test suites for LLM-powered features
- Comparing model outputs side-by-side (GPT vs Claude vs Gemini)
- Running model-graded evaluations (LLM-as-a-judge)
- Red teaming LLM applications for security vulnerabilities
- Integrating LLM quality gates into CI/CD pipelines
- Validating structured output (JSON, function calls) from LLMs
Key patterns covered:
promptfooconfig.yamlstructure (prompts, providers, tests, defaultTest)- Assertion types (deterministic, model-graded, performance)
- Custom TypeScript providers
- Red teaming configuration (plugins, strategies)
- CI/CD integration with GitHub Actions
- Programmatic API (
evaluate()function) - Result sharing and caching
When NOT to use:
- Unit testing application code (use your test runner)
- Load testing / benchmarking API throughput (use a load testing tool)
- Runtime monitoring of production LLM calls (use observability tooling)
Examples Index
- Core: Config & Assertions -- promptfooconfig.yaml structure, providers, prompts, test cases, assertion types
- Model-Graded & Advanced Assertions -- llm-rubric, factuality, similar, context evaluation, custom assertions
- Red Teaming -- Security scanning, plugins, strategies, presets
- Custom Providers & Programmatic API -- TypeScript providers, evaluate() function, CI/CD integration
- Quick API Reference -- CLI commands, assertion type table, provider IDs, red team plugins
<decision_framework>
Decision Framework
Which Assertion Type to Use
What are you validating?
+-- Exact or structural match?
| +-- Exact text -> equals
| +-- Contains substring -> contains / icontains
| +-- Regex pattern -> regex
| +-- Valid JSON -> is-json
| +-- Valid function call -> is-valid-openai-tools-call
| +-- Cost under budget -> cost (with threshold)
| +-- Response time -> latency (with threshold)
+-- Subjective quality?
| +-- General quality criteria -> llm-rubric
| +-- Factual accuracy against ground truth -> factuality
| +-- Semantic similarity -> similar (with threshold)
| +-- Closed-domain QA accuracy -> model-graded-closedqa
| +-- RAG context fidelity -> context-faithfulness
+-- Custom logic?
+-- JavaScript function -> javascript
+-- Python function -> python
+-- External service -> webhook
When to Use Red Teaming vs Eval
What are you testing?
+-- Prompt quality and correctness?
| +-- Use promptfoo eval with test cases and assertions
+-- Security vulnerabilities?
| +-- Use promptfoo redteam run with plugins and strategies
+-- Both?
+-- Run eval for quality, redteam for security -- separate configs or sections
Provider Selection
How does your LLM integration work?
+-- Direct API call to OpenAI/Anthropic/etc?
| +-- Use built-in provider: openai:gpt-4o, anthropic:messages:claude-sonnet-4-6
+-- Custom pipeline (RAG, agents, middleware)?
| +-- Use custom TypeScript provider: file://providers/my-app.ts
+-- HTTP endpoint?
| +-- Use HTTP provider: id: https://api.example.com/chat
+-- Multiple providers to compare?
+-- List all in providers array -- promptfoo runs tests against each
</decision_framework>
<red_flags>
RED FLAGS
High Priority Issues:
- Tests without
assertarrays (output is captured but never validated -- tests always "pass") - Not checking
promptfoo evalexit code in CI (promptfoo evalexits 100 on test failures -- ensure your CI pipeline treats non-zero exit codes as failures) - Hardcoded API keys in
promptfooconfig.yaml(use environment variables) - Using
llm-rubricfor checks thatis-jsonorcontainscan do deterministically (wastes money and adds non-determinism) - Red teaming without
purpose(generic attacks miss application-specific vulnerabilities)
Medium Priority Issues:
- Missing
thresholdonsimilarassertions (default may not match your quality bar) - Not caching in CI (every run makes full API calls -- expensive and slow)
- Using
model-graded-closedqawhenllm-rubricwould be simpler (closedqa is for specific ground-truth QA) - Not setting
provideron model-graded assertions (uses default which may not be the grader you want) - Running red team with default
numTests: 5in production scans (too few for comprehensive coverage)
Common Mistakes:
- Confusing
prompts(the LLM prompt templates) withtests(the evaluation cases) -- prompts define what to send, tests define what to check - Using
equalsfor natural language output (LLM output is non-deterministic, usellm-rubricorsimilar) - Forgetting
{{variable}}syntax in prompts (promptfoo uses Nunjucks templating, not${variable}) - Putting assertions in
defaultTestthat should only apply to specific tests (assertions indefaultTestapply to ALL tests) - Using
file://paths without the prefix (promptfoo treats bare paths as literal strings, not file references)
Gotchas & Edge Cases:
promptfoo evalcaches LLM responses by default -- usepromptfoo cache clearor--no-cacheto force fresh calls--shareuploads results to promptfoo's servers -- do not use with sensitive data unless self-hosting- Red team
strategieswrappluginsoutput -- a plugin generates the malicious content, a strategy delivers it (e.g., via jailbreak encoding) defaultTest.assertmerges with per-test assertions, it does not replace them -- both arrays run- CSV test files map column headers to variable names -- header
inputbecomes{{input}}in prompts transformin test options runs JavaScript on the output before assertions -- useful for extracting JSON from markdown-wrapped responses- Provider configs in YAML use
config:key for model parameters (temperature,max_tokens), not top-level fields - The
weightproperty on assertions affects scoring in the results UI but does not change pass/fail behavior
</red_flags>
<critical_reminders>
CRITICAL REMINDERS
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST define test cases with explicit assert arrays -- tests without assertions only capture output without validating it)
(You MUST use llm-rubric for subjective quality evaluation -- do NOT rely solely on deterministic assertions for natural language output)
(You MUST set threshold on similarity and model-graded assertions -- omitting thresholds uses defaults that may not match your quality bar)
(You MUST use environment variables for all API keys -- never hardcode keys in promptfooconfig.yaml or provider configs)
(You MUST verify promptfoo eval exit code in CI pipelines -- it returns exit code 100 on test failures, exit code 1 on other errors)
Failure to follow these rules will produce untested, insecure, or falsely-passing LLM evaluation pipelines.
</critical_reminders>
Files (skills)
-
examples
-
core.md 7.9 KB
# Promptfoo -- Config & Assertions Examples > Core configuration patterns, prompt definitions, provider setup, test cases, and assertion types. See [SKILL.md](../SKILL.md) for decision guidance. **Related examples:** - [model-graded.md](model-graded.md) -- LLM-as-judge assertions, factuality, context evaluation - [red-teaming.md](red-teaming.md) -- Security scanning, plugins, strategies - [custom-providers.md](custom-providers.md) -- TypeScript providers, programmatic API, CI/CD --- ## Minimal Configuration ```yaml # promptfooconfig.yaml prompts: - "Answer this question: {{question}}" providers: - openai:gpt-4o tests: - vars: question: "What is 2 + 2?" assert: - type: contains value: "4" ``` --- ## Multi-Provider Comparison ```yaml # promptfooconfig.yaml description: "Compare models for customer support" prompts: - file://prompts/support-agent.txt providers: - id: openai:gpt-4o label: "GPT-4o" config: temperature: 0.3 - id: anthropic:messages:claude-sonnet-4-6 label: "Claude Sonnet" config: temperature: 0.3 - id: openai:gpt-4.1-mini label: "GPT-4.1 Mini" config: temperature: 0.3 tests: - vars: customer_query: "I want to return my order" assert: - type: llm-rubric value: "Response is empathetic, offers clear return instructions, and asks for order number" - type: not-contains value: "I cannot" - type: latency threshold: 3000 - type: cost threshold: 0.05 ``` --- ## Prompts From Files ```yaml # promptfooconfig.yaml prompts: # Plain text file with Nunjucks variables - file://prompts/chat.txt # JSON messages array (Chat Completions format) - file://prompts/chat.json # Multiple prompt variants to compare - file://prompts/v1-concise.txt - file://prompts/v2-detailed.txt ``` ```text # prompts/chat.txt You are a helpful customer support agent for {{company_name}}. Answer the following question: {{question}} Be concise and friendly. ``` ```json // prompts/chat.json [ { "role": "system", "content": "You are a helpful assistant for {{company_name}}." }, { "role": "user", "content": "{{question}}" } ] ``` --- ## Test Cases From Files ```yaml # promptfooconfig.yaml prompts: - "Translate to {{language}}: {{input}}" providers: - openai:gpt-4o # Load tests from external files tests: - file://tests/translation-tests.yaml - file://tests/edge-cases.yaml # Or from CSV tests: file://tests/cases.csv ``` ```yaml # tests/translation-tests.yaml - vars: language: French input: "Hello world" assert: - type: icontains value: "bonjour" - type: llm-rubric value: "Natural French translation" - vars: language: Spanish input: "Good morning" assert: - type: icontains value: "buenos" ``` ```csv # tests/cases.csv -- column headers become variable names language,input,__expected French,Hello world,bonjour Spanish,Good morning,buenos German,Thank you,danke ``` --- ## Default Test (Shared Assertions) ```yaml # promptfooconfig.yaml # defaultTest applies to ALL test cases defaultTest: vars: system_prompt: "You are a helpful coding assistant." assert: # These assertions run on every test - type: not-contains value: "I'm just an AI" - type: cost threshold: 0.05 - type: latency threshold: 5000 tests: - vars: question: "Explain async/await" assert: # These run IN ADDITION to defaultTest assertions - type: icontains value: "await" - vars: question: "What is a closure?" assert: - type: llm-rubric value: "Explains closures with a practical example" ``` --- ## All Deterministic Assertion Types ### String Matching ```yaml assert: # Exact match - type: equals value: "Hello, World!" # Substring (case-sensitive) - type: contains value: "Hello" # Substring (case-insensitive) - type: icontains value: "hello" # Must NOT contain - type: not-contains value: "error" # Must NOT equal - type: not-equals value: "I don't know" # Starts with prefix - type: starts-with value: "Sure" # Contains at least one of these - type: contains-any value: - "option A" - "option B" - "option C" # Contains ALL of these - type: contains-all value: - "step 1" - "step 2" - "step 3" # Regex match - type: regex value: "\\d{3}-\\d{3}-\\d{4}" # phone number pattern # Must NOT match regex - type: not-regex value: "(?i)error|fail|exception" ``` ### Structured Output ```yaml assert: # Output is valid JSON - type: is-json # Output contains valid JSON (may have surrounding text) - type: contains-json # Output is valid HTML - type: is-html # Output is valid XML - type: is-xml # Output is valid SQL - type: is-sql # Valid OpenAI function call format - type: is-valid-openai-function-call # Valid OpenAI tools call format - type: is-valid-openai-tools-call # Model refused to answer - type: is-refusal ``` ### Performance & Cost ```yaml assert: # Response cost under budget - type: cost threshold: 0.01 # max $0.01 per call # Response time under limit - type: latency threshold: 2000 # max 2 seconds (milliseconds) # Perplexity score - type: perplexity threshold: 50 ``` ### Text Similarity ```yaml assert: # Edit distance (Levenshtein) - type: levenshtein value: "expected output text" threshold: 5 # max 5 character edits # ROUGE-N overlap - type: rouge-n value: "reference text for comparison" threshold: 0.7 # BLEU score - type: bleu value: "reference translation" threshold: 0.5 ``` --- ## Transform (Pre-Process Output) ````yaml tests: - vars: input: "Generate a JSON report" options: # Extract JSON from markdown code blocks before assertions run transform: | const match = output.match(/```json\n([\s\S]*?)\n```/); return match ? match[1] : output; assert: - type: is-json ```` --- ## Assertion Sets (Partial Pass) ```yaml assert: # At least 2 out of 3 assertions must pass - type: assert-set threshold: 0.67 assert: - type: icontains value: "hello" - type: icontains value: "bonjour" - type: icontains value: "hola" ``` --- ## Weighted Assertions ```yaml assert: # Accuracy is more important than speed - type: llm-rubric value: "Response is factually correct" weight: 3.0 metric: "accuracy" - type: latency threshold: 3000 weight: 1.0 metric: "speed" - type: cost threshold: 0.02 weight: 1.0 metric: "cost" ``` --- ## YAML References (Reusable Assertions) ```yaml # Define reusable assertion templates assertionTemplates: qualityCheck: type: llm-rubric value: "Response is helpful, accurate, and professionally written" noHallucination: type: not-contains value: "I'm not sure but" budgetCheck: type: cost threshold: 0.02 tests: - vars: question: "What is TypeScript?" assert: - $ref: "#/assertionTemplates/qualityCheck" - $ref: "#/assertionTemplates/noHallucination" - $ref: "#/assertionTemplates/budgetCheck" - vars: question: "Explain React hooks" assert: - $ref: "#/assertionTemplates/qualityCheck" - $ref: "#/assertionTemplates/budgetCheck" ``` --- ## Variable Types ```yaml tests: # Simple string - vars: question: "What is TypeScript?" # File content as variable - vars: context: file://data/context.txt question: "Summarize the above" # Array (creates one test per value) - vars: language: - French - Spanish - German input: "Hello world" # Nested object - vars: user: name: "Alice" role: "admin" ``` --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._ -
custom-providers.md 8.6 KB
# Promptfoo -- Custom Providers & Programmatic API Examples > TypeScript custom providers, inline function providers, programmatic evaluation, and CI/CD integration patterns. See [core.md](core.md) for basic configuration. **Related examples:** - [core.md](core.md) -- Config structure, assertions - [model-graded.md](model-graded.md) -- Model-graded assertions - [red-teaming.md](red-teaming.md) -- Security scanning --- ## TypeScript Custom Provider Use when your LLM integration involves custom logic (RAG, agents, middleware). ```typescript // providers/rag-provider.ts import type { ApiProvider, ProviderOptions, ProviderResponse, CallApiContextParams, } from "promptfoo"; // NOTE: default export required by promptfoo's file:// provider loader export default class RagProvider implements ApiProvider { private config: Record<string, unknown>; constructor(options: ProviderOptions) { this.config = options.config || {}; } id(): string { return "rag-provider"; } async callApi( prompt: string, context?: CallApiContextParams, ): Promise<ProviderResponse> { // 1. Retrieve relevant documents const docs = await retrieveDocuments(prompt); // 2. Build augmented prompt const augmentedPrompt = `Context:\n${docs.join("\n")}\n\nQuestion: ${prompt}`; // 3. Call LLM const response = await callLLM(augmentedPrompt); return { output: response.text, tokenUsage: { total: response.totalTokens, prompt: response.promptTokens, completion: response.completionTokens, }, cost: response.cost, metadata: { docsRetrieved: docs.length, retrievalLatencyMs: response.retrievalTime, }, }; } } ``` ```yaml # promptfooconfig.yaml prompts: - "{{question}}" providers: - file://providers/rag-provider.ts tests: - vars: question: "What is our refund policy?" assert: - type: context-faithfulness threshold: 0.9 provider: openai:gpt-4o - type: icontains value: "30 days" ``` --- ## Provider With Config From YAML Pass configuration from YAML to your TypeScript provider. ```yaml # promptfooconfig.yaml providers: - id: file://providers/rag-provider.ts label: "RAG (top-3)" config: topK: 3 temperature: 0.2 maxTokens: 500 - id: file://providers/rag-provider.ts label: "RAG (top-5)" config: topK: 5 temperature: 0.2 maxTokens: 500 ``` ```typescript // providers/rag-provider.ts // NOTE: default export required by promptfoo's file:// provider loader export default class RagProvider implements ApiProvider { private topK: number; constructor(options: ProviderOptions) { this.topK = (options.config?.topK as number) || 3; } // ...use this.topK in callApi } ``` --- ## Inline Function Provider For simple cases, define the provider inline without a separate file. ```typescript // eval.ts import promptfoo from "promptfoo"; const results = await promptfoo.evaluate({ prompts: ["Summarize: {{text}}"], providers: [ // Inline provider function async (prompt: string, context) => { const response = await fetch("https://api.myapp.com/summarize", { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify({ text: prompt }), }); const data = await response.json(); return { output: data.summary }; }, ], tests: [ { vars: { text: "A long article about climate change..." }, assert: [{ type: "llm-rubric", value: "Concise and accurate summary" }], }, ], }); ``` --- ## Programmatic API (evaluate) Run evaluations from TypeScript code. Useful for integration tests or scheduled jobs. ```typescript // scripts/run-eval.ts import promptfoo from "promptfoo"; import type { EvaluateSummary } from "promptfoo"; const MAX_CONCURRENCY = 5; const results: EvaluateSummary = await promptfoo.evaluate( { prompts: [ "Translate to {{language}}: {{text}}", "You are a translator. Convert this to {{language}}: {{text}}", ], providers: ["openai:gpt-4o", "anthropic:messages:claude-sonnet-4-6"], tests: [ { vars: { language: "French", text: "Hello world" }, assert: [ { type: "icontains", value: "bonjour" }, { type: "cost", threshold: 0.01 }, ], }, { vars: { language: "Spanish", text: "Good morning" }, assert: [ { type: "icontains", value: "buenos" }, { type: "llm-rubric", value: "Natural translation, not word-for-word", }, ], }, ], writeLatestResults: true, sharing: true, }, { maxConcurrency: MAX_CONCURRENCY }, ); // Access results const totalTests = results.stats.successes + results.stats.failures; console.log("Total tests:", totalTests); console.log("Pass rate:", results.stats.successes / totalTests); if (results.shareableUrl) { console.log("Results:", results.shareableUrl); } ``` --- ## CI/CD: GitHub Actions Workflow ```yaml # .github/workflows/llm-eval.yml name: LLM Evaluation on: pull_request: paths: - "prompts/**" - "promptfooconfig.yaml" - "src/ai/**" jobs: evaluate: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - uses: actions/setup-node@v4 with: node-version: "22" # Cache LLM responses to save cost - uses: actions/cache@v4 with: path: ~/.cache/promptfoo key: ${{ runner.os }}-promptfoo-v1 restore-keys: | ${{ runner.os }}-promptfoo- - name: Run LLM evaluation env: OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }} ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} run: | npx promptfoo@latest eval \ -o results.json \ -o report.html \ --share - name: Upload report if: always() uses: actions/upload-artifact@v4 with: name: llm-eval-report path: report.html ``` --- ## CI/CD: Using promptfoo-action ```yaml # .github/workflows/llm-eval.yml name: Prompt Evaluation on: pull_request: paths: - "prompts/**" jobs: evaluate: runs-on: ubuntu-latest permissions: pull-requests: write steps: - uses: actions/cache@v4 with: path: ~/.cache/promptfoo key: ${{ runner.os }}-promptfoo-v1 - uses: promptfoo/promptfoo-action@v1 with: openai-api-key: ${{ secrets.OPENAI_API_KEY }} github-token: ${{ secrets.GITHUB_TOKEN }} prompts: "prompts/**/*.json" config: "promptfooconfig.yaml" cache-path: ~/.cache/promptfoo ``` --- ## CI/CD: npm Script With Quality Gate ```json // package.json { "scripts": { "test:llm": "promptfoo eval", "test:llm:view": "promptfoo view", "test:llm:redteam": "promptfoo redteam run", "test:llm:share": "promptfoo eval --share" } } ``` ```bash # Quality gate with pass rate threshold RESULTS=$(npx promptfoo@latest eval -o results.json 2>&1) PASS_RATE=$(jq '.results.stats.successes / (.results.stats.successes + .results.stats.failures) * 100' results.json) MIN_PASS_RATE=95 if (( $(echo "$PASS_RATE < $MIN_PASS_RATE" | bc -l) )); then echo "Quality gate failed: ${PASS_RATE}% < ${MIN_PASS_RATE}%" exit 1 fi echo "Quality gate passed: ${PASS_RATE}%" ``` --- ## HTTP Provider (No Code) Test any HTTP endpoint without writing a custom provider. ```yaml # promptfooconfig.yaml providers: - id: https://api.myapp.com/chat label: "My App API" config: method: POST headers: Authorization: "Bearer {{env.APP_API_KEY}}" Content-Type: "application/json" body: message: "{{prompt}}" user_id: "eval-user" responseParser: "json.choices[0].message.content" tests: - vars: input: "What are your business hours?" assert: - type: icontains value: "9" - type: llm-rubric value: "Provides specific business hours" ``` --- ## Sharing Results ```bash # Share after eval npx promptfoo@latest eval --share # Share most recent results npx promptfoo@latest share # Output to file for archival npx promptfoo@latest eval -o results.json -o report.html ``` --- ## Cache Management ```bash # Clear all cached responses npx promptfoo@latest cache clear # Run without cache (fresh LLM calls) npx promptfoo@latest eval --no-cache # Disable cache via env PROMPTFOO_DISABLE_CACHE=1 npx promptfoo@latest eval ``` --- _For basic configuration, see [core.md](core.md). For API reference, see [reference.md](../reference.md)._ -
model-graded.md 8.3 KB
# Promptfoo -- Model-Graded & Advanced Assertions Examples > LLM-as-judge evaluation, factuality checks, semantic similarity, RAG context evaluation, and custom grading. See [core.md](core.md) for basic assertions. **Related examples:** - [core.md](core.md) -- Config structure, deterministic assertions - [red-teaming.md](red-teaming.md) -- Security scanning - [custom-providers.md](custom-providers.md) -- TypeScript providers, programmatic API --- ## llm-rubric (General Quality) The most versatile model-graded assertion. Evaluates output against arbitrary criteria using an LLM judge. ```yaml tests: - vars: question: "Explain quantum computing to a 10-year-old" assert: - type: llm-rubric value: > The explanation uses simple analogies, avoids jargon like 'superposition' or 'entanglement', and is no longer than 3 paragraphs provider: openai:gpt-4o - vars: question: "Write a professional email declining a meeting" assert: - type: llm-rubric value: > Email is polite but firm, suggests an alternative (reschedule or async update), and is under 150 words ``` **When to use:** Subjective quality criteria that cannot be expressed as string matching or regex --- ## llm-rubric With Custom Grading Provider ```yaml # Use a cheaper model for grading to reduce eval costs defaultTest: options: provider: openai:gpt-4o-mini # Default grader for all model-graded assertions tests: - vars: question: "Explain recursion" assert: - type: llm-rubric value: "Includes a base case and recursive case in the explanation" # Uses defaultTest.options.provider (gpt-4o-mini) for grading - type: llm-rubric value: "Code example compiles and demonstrates recursion correctly" provider: openai:gpt-4o # Override: use stronger model for code evaluation ``` --- ## factuality (Ground Truth Accuracy) Checks whether the output is factually consistent with provided ground truth. ```yaml tests: - vars: question: "What is the capital of France?" assert: - type: factuality value: "The capital of France is Paris. Paris has a population of approximately 2.1 million in the city proper." provider: openai:gpt-4o - vars: question: "When was Python created?" assert: - type: factuality value: "Python was created by Guido van Rossum and first released in 1991." provider: openai:gpt-4o ``` **When to use:** Validating factual accuracy against known ground truth -- stricter than `llm-rubric` --- ## model-graded-closedqa (Closed-Domain QA) Uses OpenAI's public evals prompt for closed-domain question answering. Checks if the answer correctly addresses the question given a reference answer. ```yaml tests: - vars: question: "What year was the Eiffel Tower completed?" assert: - type: model-graded-closedqa value: "The Eiffel Tower was completed in 1889." provider: openai:gpt-4o ``` **When to use:** Question-answer pairs with definitive correct answers -- more specific than `llm-rubric` --- ## similar (Semantic Similarity) Compares output to a reference using embedding-based cosine similarity. Does not require an LLM call for grading -- uses embeddings. ```yaml tests: - vars: input: "Bonjour le monde" assert: - type: similar value: "Hello world" threshold: 0.8 # 0.0 to 1.0 -- higher = more similar required - vars: input: "What is machine learning?" assert: - type: similar value: "Machine learning is a subset of AI that enables systems to learn from data" threshold: 0.7 ``` **When to use:** Comparing semantic meaning when exact wording varies -- cheaper than `llm-rubric` --- ## RAG Context Evaluation Evaluate retrieval-augmented generation quality across multiple dimensions. ```yaml tests: - vars: question: "What is our refund policy?" context: "Refunds are available within 30 days of purchase for unused items with original receipt." assert: # Is the answer faithful to the retrieved context? - type: context-faithfulness threshold: 0.9 provider: openai:gpt-4o # Does the context contain the information needed to answer? - type: context-recall value: "Refunds require original receipt and must be within 30 days" threshold: 0.8 provider: openai:gpt-4o # Is the retrieved context relevant to the question? - type: context-relevance threshold: 0.8 provider: openai:gpt-4o # Is the answer relevant to the question asked? - type: answer-relevance threshold: 0.8 provider: openai:gpt-4o ``` **When to use:** Evaluating RAG pipeline quality -- faithfulness prevents hallucination, recall ensures coverage --- ## Custom JavaScript Assertion For evaluation logic that no built-in assertion handles. ```yaml tests: - vars: question: "Generate a JSON user profile" assert: # Inline JavaScript - type: javascript value: | const parsed = JSON.parse(output); const hasRequiredFields = parsed.name && parsed.email && parsed.role; const validEmail = parsed.email.includes('@'); return { pass: hasRequiredFields && validEmail, score: hasRequiredFields && validEmail ? 1.0 : 0.0, reason: hasRequiredFields ? (validEmail ? 'Valid profile' : 'Invalid email format') : 'Missing required fields (name, email, role)', }; # From file - type: javascript value: file://assertions/validate-profile.js ``` ```javascript // assertions/validate-profile.js module.exports = (output, context) => { const { vars } = context; const parsed = JSON.parse(output); const checks = [ { name: "has name", pass: Boolean(parsed.name) }, { name: "has email", pass: Boolean(parsed.email) }, { name: "valid role", pass: ["admin", "user", "viewer"].includes(parsed.role), }, ]; const passed = checks.filter((c) => c.pass); const failed = checks.filter((c) => !c.pass); return { pass: failed.length === 0, score: passed.length / checks.length, reason: failed.length > 0 ? `Failed: ${failed.map((c) => c.name).join(", ")}` : "All checks passed", }; }; ``` --- ## Custom Python Assertion ```yaml assert: - type: python value: file://assertions/validate.py ``` ```python # assertions/validate.py import json def get_assert(output, context): try: data = json.loads(output) has_name = bool(data.get("name")) has_email = bool(data.get("email")) return { "pass": has_name and has_email, "score": 1.0 if (has_name and has_email) else 0.0, "reason": "Valid" if (has_name and has_email) else "Missing fields", } except json.JSONDecodeError: return { "pass": False, "score": 0.0, "reason": "Output is not valid JSON", } ``` --- ## Combining Deterministic and Model-Graded Best practice: use deterministic assertions for structure, model-graded for quality. ```yaml tests: - vars: task: "Generate a product description for a laptop" assert: # Structure checks (fast, free, deterministic) - type: is-json - type: javascript value: | const parsed = JSON.parse(output); return { pass: parsed.title && parsed.description && parsed.price, score: 1.0, reason: 'Has required fields', }; # Quality checks (LLM-graded, slower, costs money) - type: llm-rubric value: > Description is compelling, highlights key features, and is appropriate for an e-commerce listing provider: openai:gpt-4o # Performance checks - type: cost threshold: 0.02 - type: latency threshold: 3000 ``` --- ## Negating Model-Graded Assertions ```yaml assert: # Output must NOT be a refusal - type: not-is-refusal # Output must NOT be similar to a known bad response - type: not-similar value: "I cannot help with that request" threshold: 0.8 ``` --- _For basic config and deterministic assertions, see [core.md](core.md). For API reference, see [reference.md](../reference.md)._ -
red-teaming.md 5.1 KB
# Promptfoo -- Red Teaming Examples > Security scanning for LLM applications -- plugins, strategies, presets, multi-turn attacks, and custom policies. See [core.md](core.md) for basic configuration. **Related examples:** - [core.md](core.md) -- Config structure, assertions - [model-graded.md](model-graded.md) -- Model-graded assertions - [custom-providers.md](custom-providers.md) -- Custom providers for testing your application --- ## Basic Red Team Configuration ```yaml # promptfooconfig.yaml targets: - openai:gpt-4o redteam: purpose: "Customer support chatbot for an online electronics store" numTests: 10 plugins: - harmful - pii - hallucination strategies: - jailbreak - prompt-injection ``` Run with: ```bash npx promptfoo@latest redteam run ``` --- ## Comprehensive Security Scan ```yaml # promptfooconfig.yaml targets: - id: openai:gpt-4o label: "Production chatbot" config: temperature: 0.3 redteam: purpose: > Healthcare appointment scheduling assistant. Must not provide medical advice, share patient data, or schedule appointments outside business hours (9am-5pm EST). numTests: 25 plugins: # Privacy - pii # Harmful content - harmful # Security - prompt-extraction - shell-injection - sql-injection # Misinformation - hallucination - contracts - excessive-agency - competitor-endorsement # Custom policy - id: policy config: policy: "Must not provide medical advice or diagnoses" strategies: - jailbreak - prompt-injection - crescendo # Multi-turn escalation - base64 # Encoded payloads - leetspeak # Obfuscated text ``` --- ## Using OWASP and NIST Presets ```yaml # promptfooconfig.yaml targets: - file://providers/my-app.ts redteam: purpose: "Financial advisor chatbot for a retail bank" numTests: 15 plugins: # OWASP LLM Top 10 coverage - owasp:llm:01 # Prompt injection - owasp:llm:02 # Insecure output handling - owasp:llm:06 # Sensitive information disclosure - owasp:llm:09 # Overreliance # NIST AI RMF measures - nist:ai:measure # Additional custom checks - contracts - excessive-agency strategies: - jailbreak - prompt-injection ``` --- ## Custom Target (Your Application) Red team your actual application endpoint, not just the raw model. ```yaml # promptfooconfig.yaml targets: # Custom HTTP endpoint - id: https://api.myapp.com/chat label: "My App API" config: method: POST headers: Authorization: "Bearer {{env.APP_API_KEY}}" Content-Type: "application/json" body: message: "{{prompt}}" session_id: "red-team-{{uuid}}" # Or custom TypeScript provider - file://providers/my-app-provider.ts redteam: purpose: "Internal HR assistant for employee benefits questions" numTests: 20 plugins: - pii - harmful - prompt-extraction strategies: - jailbreak - prompt-injection - crescendo ``` --- ## Multi-Turn Red Teaming Test resistance to escalation attacks that build across conversation turns. ```yaml # promptfooconfig.yaml targets: - openai:gpt-4o redteam: purpose: "Children's educational tutor for ages 8-12" numTests: 15 plugins: - harmful - pii strategies: # Multi-turn attacks - crescendo # Gradually escalates across turns - goat # Generative offensive agent tester - id: crescendo config: maxTurns: 8 # Override default turn limit ``` --- ## Plugin-Specific Configuration ```yaml redteam: purpose: "Legal document summarization tool" plugins: # Default config - harmful # Custom test count per plugin - id: pii numTests: 20 # More PII tests (higher risk area) # Custom policy plugin - id: policy numTests: 15 config: policy: > Must not draft legal documents. Must not provide legal advice. Must clearly state it is a summarization tool only. # Custom topic plugin - id: topic config: topic: "competitor products and pricing" policy: "Must not discuss competitor products" strategies: - jailbreak - prompt-injection ``` --- ## Red Team With Custom Grading ```yaml redteam: purpose: "Product recommendation engine" numTests: 10 provider: openai:gpt-4o # Model used to generate attacks AND grade results # Custom grading examples to calibrate the judge graderExamples: - output: "I'd recommend checking out our competitor's product instead" pass: false reason: "Endorsed a competitor product" - output: "Based on your needs, I recommend our Premium Plan" pass: true reason: "Appropriate product recommendation" plugins: - competitor-endorsement - contracts strategies: - jailbreak ``` --- ## Viewing Red Team Results ```bash # Run the scan npx promptfoo@latest redteam run # View results in browser npx promptfoo@latest redteam report # Share results npx promptfoo@latest redteam run --share ``` --- _For basic configuration, see [core.md](core.md). For assertion types, see [reference.md](../reference.md)._
-
-
reference.md 11.2 KB
# Promptfoo Quick Reference > CLI commands, assertion type table, provider IDs, red team plugins, and configuration keys. See [SKILL.md](SKILL.md) for core patterns and [examples/](examples/) for code examples. --- ## Installation ```bash # Run directly (no install) npx promptfoo@latest eval # Global install npm install -g promptfoo # Project dependency npm install --save-dev promptfoo ``` --- ## CLI Commands | Command | Description | | --------------------------------------- | ------------------------------------------- | | `promptfoo init` | Create a new promptfooconfig.yaml | | `promptfoo eval` | Run evaluation (exits 100 on test failures) | | `promptfoo eval -c path/to/config.yaml` | Use specific config file | | `promptfoo eval --no-cache` | Skip cache, force fresh LLM calls | | `promptfoo eval -o results.json` | Output results to file | | `promptfoo eval --share` | Generate shareable URL | | `promptfoo view` | Open results web UI | | `promptfoo share` | Share most recent results | | `promptfoo cache clear` | Clear cached LLM responses | | `promptfoo redteam run` | Run red team security scan | | `promptfoo redteam generate` | Generate red team test cases only | | `promptfoo redteam report` | View red team results | ### Common Flags | Flag | Description | | ------------------------- | ----------------------------------------------------- | | `-c, --config` | Path to config file (default: `promptfooconfig.yaml`) | | `-o, --output` | Output file path (supports `.json`, `.html`, `.xml`) | | `--no-cache` | Disable response caching | | `--share` | Upload results and print shareable URL | | `-j, --max-concurrency` | Max parallel provider calls | | `--table-cell-max-length` | Max chars in table output cells | | `--env-file` | Path to `.env` file for API keys | --- ## Configuration Keys ### Top-Level ```yaml description: string # Project description prompts: string[] | object[] # Prompt templates providers: string[] | object[] # LLM providers tests: object[] | string # Test cases (inline or file://) defaultTest: object # Default vars/assertions for all tests outputPath: string # Results output path sharing: boolean | object # Enable sharing ``` ### Provider Object ```yaml providers: - id: openai:gpt-4o # Provider identifier label: "GPT-4o" # Display name in results config: temperature: 0.7 max_tokens: 1000 top_p: 1 ``` ### Test Object ```yaml tests: - description: "Test name" # Optional label vars: # Template variables input: "Hello" assert: # Assertions array - type: contains value: "hello" options: transform: "output.trim()" # Pre-process output provider: openai:gpt-4o # Override provider for this test ``` ### Default Test ```yaml defaultTest: vars: system_prompt: "You are a helpful assistant" assert: - type: not-contains value: "I cannot" options: provider: openai:gpt-4o ``` --- ## Assertion Types ### Deterministic | Type | Value | Description | | ------------------------------- | -------- | ------------------------------ | | `equals` | string | Exact match | | `contains` | string | Substring match | | `icontains` | string | Case-insensitive substring | | `not-contains` | string | Must not contain | | `not-equals` | string | Must not equal | | `contains-any` | string[] | Contains at least one | | `contains-all` | string[] | Contains all | | `icontains-any` | string[] | Case-insensitive, at least one | | `icontains-all` | string[] | Case-insensitive, all | | `starts-with` | string | Prefix match | | `regex` | string | Regular expression match | | `not-regex` | string | Must not match regex | | `is-json` | -- | Valid JSON | | `contains-json` | -- | Contains valid JSON | | `is-html` | -- | Valid HTML | | `is-xml` | -- | Valid XML | | `is-sql` | -- | Valid SQL | | `is-valid-openai-tools-call` | -- | Valid OpenAI tools call | | `is-valid-openai-function-call` | -- | Valid OpenAI function call | | `is-refusal` | -- | Model refused to answer | ### Performance | Type | Threshold | Description | | ------------ | ------------ | -------------------- | | `cost` | number (USD) | Max cost per call | | `latency` | number (ms) | Max response time | | `perplexity` | number | Max perplexity score | ### Text Similarity | Type | Threshold | Description | | ------------- | ------------------ | -------------------------------- | | `levenshtein` | number (max edits) | Edit distance | | `similar` | number (0-1) | Cosine similarity via embeddings | | `rouge-n` | number (0-1) | ROUGE-N overlap score | | `bleu` | number (0-1) | BLEU translation score | ### Model-Graded | Type | Value | Description | | ----------------------- | --------------- | --------------------------------- | | `llm-rubric` | criteria string | General LLM-as-judge | | `factuality` | ground truth | Factual accuracy check | | `model-graded-closedqa` | ground truth | Closed-domain QA accuracy | | `answer-relevance` | -- | Answer relevance to question | | `context-faithfulness` | -- | RAG: answer faithful to context | | `context-recall` | -- | RAG: context covers ground truth | | `context-relevance` | -- | RAG: context relevant to question | | `classifier` | criteria | Classification evaluation | | `select-best` | -- | Pick best output across providers | ### Custom | Type | Value | Description | | ------------ | ----------------- | ---------------------------- | | `javascript` | code or `file://` | Custom JS assertion | | `python` | code or `file://` | Custom Python assertion | | `webhook` | URL | External validation endpoint | ### Assertion Properties ```yaml assert: - type: llm-rubric value: "criteria" # Expected value / criteria threshold: 0.8 # Pass/fail threshold (0-1) weight: 2.0 # Scoring weight in results UI metric: "quality" # Label for UI aggregation provider: openai:gpt-4o # Grading model (model-graded only) transform: "output.trim()" # Pre-process output before assertion ``` --- ## Built-in Provider IDs ### OpenAI ``` openai:gpt-4o openai:gpt-4o-mini openai:o4-mini openai:gpt-4.1 openai:gpt-4.1-mini ``` ### Anthropic ``` anthropic:messages:claude-sonnet-4-6 anthropic:messages:claude-opus-4-6 anthropic:messages:claude-sonnet-4-5-latest anthropic:messages:claude-3-5-haiku-latest ``` ### Google ``` vertex:gemini-2.0-flash vertex:gemini-2.5-pro ``` ### Other ``` ollama:llama3 ollama:mistral huggingface:text-generation:MODEL_NAME ``` ### Custom ```yaml # TypeScript/JavaScript file - file://providers/my-provider.ts # HTTP endpoint - id: https://api.example.com/chat config: method: POST headers: Authorization: "Bearer {{env.API_KEY}}" body: prompt: "{{prompt}}" ``` --- ## Red Team Configuration ### Plugins (Attack Generators) | Plugin | Category | Description | | ------------------------ | -------------- | ---------------------------- | | `harmful` | Criminal | Harmful content generation | | `pii` | Privacy | PII extraction attempts | | `prompt-extraction` | Security | System prompt extraction | | `hallucination` | Misinformation | Hallucinated information | | `contracts` | Misinformation | Unauthorized commitments | | `excessive-agency` | Misinformation | Taking unauthorized actions | | `competitor-endorsement` | Misinformation | Endorsing competitors | | `hijacking` | Security | Topic/task hijacking | | `overreliance` | Misinformation | Over-reliance on user claims | | `shell-injection` | Security | OS command injection | | `sql-injection` | Security | SQL injection attempts | ### Strategies (Delivery Methods) | Strategy | Description | | ------------------ | -------------------------- | | `jailbreak` | Jailbreak prompt templates | | `prompt-injection` | Direct prompt injection | | `crescendo` | Multi-turn escalation | | `goat` | Generative offensive agent | | `base64` | Base64-encoded payloads | | `rot13` | ROT13-encoded payloads | | `leetspeak` | Leetspeak encoding | | `iterative` | Iterative refinement | ### Presets ```yaml redteam: plugins: - owasp:llm:01 # OWASP LLM Top 10 - nist:ai:measure # NIST AI RMF ``` --- ## Environment Variables | Variable | Purpose | | --------------------------- | -------------------------------- | | `OPENAI_API_KEY` | OpenAI provider API key | | `ANTHROPIC_API_KEY` | Anthropic provider API key | | `GOOGLE_API_KEY` | Google / Vertex AI API key | | `PROMPTFOO_CONFIG` | Override config file path | | `PROMPTFOO_CACHE_PATH` | Custom cache directory | | `PROMPTFOO_DISABLE_CACHE` | Disable caching (`1` to disable) | | `PROMPTFOO_SHARING_APP_URL` | Custom sharing server URL | --- ## Nunjucks Template Syntax ```yaml # Variable substitution prompts: - "Translate to {{language}}: {{input}}" # Conditional - "{% if context %}Context: {{context}}{% endif %}\n{{question}}" # Loop - "{% for item in items %}{{item}}\n{% endfor %}" # Environment variable - "file://{{ env.PROMPT_DIR }}/prompt.txt" ``` > For YAML references (reusable assertion blocks), see [examples/core.md](examples/core.md#yaml-references-reusable-assertions). -
SKILL.md 16.8 KB
--- name: ai-observability-promptfoo description: Testing and evaluation framework for LLM prompts and applications -- promptfooconfig.yaml, assertions, model-graded evals, red teaming, CI/CD integration, custom providers, and comparative evaluation --- # Promptfoo Patterns > **Quick Guide:** Use promptfoo for systematic LLM evaluation. Define prompts, providers, and test cases in `promptfooconfig.yaml`. Use assertion types (`contains`, `is-json`, `llm-rubric`, `similar`, `cost`, `latency`) to validate outputs. Use `promptfoo eval` to run (exits with code 100 on test failures), `promptfoo view` for results UI. Use model-graded assertions (`llm-rubric`, `factuality`) for subjective quality. Use `promptfoo redteam run` for security scanning. Use `--share` flag or `promptfoo share` to share results. All provider API keys come from environment variables -- never hardcode them. --- <critical_requirements> ## CRITICAL: Before Using This Skill > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST define test cases with explicit `assert` arrays -- tests without assertions only capture output without validating it)** **(You MUST use `llm-rubric` for subjective quality evaluation -- do NOT rely solely on deterministic assertions for natural language output)** **(You MUST set `threshold` on similarity and model-graded assertions -- omitting thresholds uses defaults that may not match your quality bar)** **(You MUST use environment variables for all API keys -- never hardcode keys in promptfooconfig.yaml or provider configs)** **(You MUST verify `promptfoo eval` exit code in CI pipelines -- it returns exit code 100 on test failures, exit code 1 on other errors)** </critical_requirements> --- **Auto-detection:** promptfoo, promptfooconfig, promptfooconfig.yaml, promptfoo eval, promptfoo view, promptfoo redteam, llm-rubric, model-graded-closedqa, promptfoo share, promptfoo cache, assertion type, LLM evaluation, prompt testing, red teaming, PROMPTFOO_CONFIG **When to use:** - Writing or evaluating LLM prompts across one or more providers - Setting up automated test suites for LLM-powered features - Comparing model outputs side-by-side (GPT vs Claude vs Gemini) - Running model-graded evaluations (LLM-as-a-judge) - Red teaming LLM applications for security vulnerabilities - Integrating LLM quality gates into CI/CD pipelines - Validating structured output (JSON, function calls) from LLMs **Key patterns covered:** - `promptfooconfig.yaml` structure (prompts, providers, tests, defaultTest) - Assertion types (deterministic, model-graded, performance) - Custom TypeScript providers - Red teaming configuration (plugins, strategies) - CI/CD integration with GitHub Actions - Programmatic API (`evaluate()` function) - Result sharing and caching **When NOT to use:** - Unit testing application code (use your test runner) - Load testing / benchmarking API throughput (use a load testing tool) - Runtime monitoring of production LLM calls (use observability tooling) --- ## Examples Index - [Core: Config & Assertions](examples/core.md) -- promptfooconfig.yaml structure, providers, prompts, test cases, assertion types - [Model-Graded & Advanced Assertions](examples/model-graded.md) -- llm-rubric, factuality, similar, context evaluation, custom assertions - [Red Teaming](examples/red-teaming.md) -- Security scanning, plugins, strategies, presets - [Custom Providers & Programmatic API](examples/custom-providers.md) -- TypeScript providers, evaluate() function, CI/CD integration - [Quick API Reference](reference.md) -- CLI commands, assertion type table, provider IDs, red team plugins --- <philosophy> ## Philosophy Promptfoo brings **test-driven development to LLM applications**. Instead of manually checking outputs, you define expected behaviors as assertions and run them systematically across prompts and providers. **Core principles:** 1. **Declarative test definitions** -- YAML config over imperative test scripts. Define prompts, providers, test cases, and assertions in `promptfooconfig.yaml`. No code required for standard evaluations. 2. **Assertion-driven validation** -- Every test case should have assertions. Deterministic assertions (`contains`, `is-json`, `equals`) for structured output; model-graded assertions (`llm-rubric`, `factuality`) for subjective quality. 3. **Comparative evaluation** -- Run the same tests across multiple providers or prompt variants simultaneously. The results matrix shows which combination performs best. 4. **Shift-left LLM testing** -- Catch prompt regressions in CI before they reach production. `promptfoo eval` exits with code 100 on test failures, making it a natural CI quality gate. 5. **Red teaming as a first-class concern** -- Security scanning for prompt injection, PII leakage, harmful content, and jailbreak vulnerabilities is built in, not bolted on. </philosophy> --- <patterns> ## Core Patterns ### Pattern 1: Basic Configuration Every promptfoo project starts with `promptfooconfig.yaml`. Three required sections: `prompts`, `providers`, `tests`. ```yaml # promptfooconfig.yaml description: "Translation quality evaluation" prompts: - "Convert the following to {{language}}: {{input}}" providers: - openai:gpt-4o - anthropic:messages:claude-sonnet-4-6 tests: - vars: language: French input: Hello world assert: - type: icontains value: "bonjour" - type: llm-rubric value: "Output is a natural French translation, not word-for-word" ``` **Why good:** Declarative config, multi-provider comparison, both deterministic and model-graded assertions ```yaml # BAD: Tests without assertions tests: - vars: language: French input: Hello world # No assert array -- output is captured but never validated ``` **Why bad:** Tests without assertions only log output, they never fail -- you lose the entire point of automated evaluation **See:** [examples/core.md](examples/core.md) for prompts from files, provider config, defaultTest, variable loading from CSV --- ### Pattern 2: Deterministic Assertions Use for outputs with predictable, verifiable structure. ```yaml assert: # String matching - type: contains value: "error" - type: icontains # case-insensitive value: "success" - type: not-contains value: "internal server error" - type: starts-with value: "{" - type: regex value: "\\d{4}-\\d{2}-\\d{2}" # date pattern # Structured output - type: is-json - type: contains-json - type: is-valid-openai-tools-call # Performance - type: cost threshold: 0.01 # max $0.01 per call - type: latency threshold: 5000 # max 5 seconds ``` **Why good:** Fast, deterministic, no LLM cost for evaluation, catches structural regressions immediately ```yaml # BAD: Using llm-rubric for JSON validation assert: - type: llm-rubric value: "Output must be valid JSON" ``` **Why bad:** Expensive (requires LLM call), slower, non-deterministic -- `is-json` does this deterministically for free **See:** [examples/core.md](examples/core.md) for all deterministic assertion types with examples --- ### Pattern 3: Model-Graded Assertions Use for subjective quality where deterministic checks cannot capture intent. ```yaml assert: - type: llm-rubric value: "Response is helpful, accurate, and conversational in tone" provider: openai:gpt-4o - type: factuality value: "The capital of France is Paris. It has a population of ~2.1 million." provider: openai:gpt-4o - type: similar value: "The weather in Paris is sunny today" threshold: 0.8 - type: model-graded-closedqa value: "Paris is the capital of France" provider: openai:gpt-4o ``` **Why good:** Evaluates subjective quality that deterministic assertions cannot capture, configurable grading provider ```yaml # BAD: No threshold on similar assertion assert: - type: similar value: "expected output" # Missing threshold -- uses default which may be too lenient or strict ``` **Why bad:** Default similarity threshold may not match your quality bar, always set it explicitly **See:** [examples/model-graded.md](examples/model-graded.md) for llm-rubric with custom providers, context evaluation, factuality, custom grading prompts --- ### Pattern 4: Red Teaming Use `redteam` section to scan for security vulnerabilities. ```yaml # promptfooconfig.yaml targets: - openai:gpt-4o redteam: purpose: "Customer support chatbot for an e-commerce platform" numTests: 10 plugins: - harmful - pii - contracts - hallucination - prompt-extraction strategies: - jailbreak - prompt-injection ``` **Why good:** Declarative security scanning, purpose provides context for realistic attacks, composable plugins and strategies ```yaml # BAD: Red team without purpose redteam: plugins: - harmful # Missing purpose -- attacks will be generic and less effective ``` **Why bad:** Without `purpose`, the red team generator creates generic attacks that miss application-specific vulnerabilities **See:** [examples/red-teaming.md](examples/red-teaming.md) for presets (OWASP, NIST), advanced strategies, multi-turn attacks --- ### Pattern 5: Custom TypeScript Provider Use when your LLM integration is not a direct API call (RAG pipelines, agent chains, custom middleware). ```typescript // providers/my-app.ts import type { ApiProvider, ProviderOptions, ProviderResponse, CallApiContextParams, } from "promptfoo"; // NOTE: default export required by promptfoo's file:// provider loader export default class MyAppProvider implements ApiProvider { private config: Record<string, unknown>; constructor(options: ProviderOptions) { this.config = options.config || {}; } id(): string { return "my-app-provider"; } async callApi( prompt: string, context?: CallApiContextParams, ): Promise<ProviderResponse> { // Call your application's LLM pipeline const result = await myApp.processQuery(prompt); return { output: result.answer, tokenUsage: { total: result.totalTokens, prompt: result.promptTokens, completion: result.completionTokens, }, cost: result.cost, }; } } ``` ```yaml # promptfooconfig.yaml providers: - file://providers/my-app.ts ``` **Why good:** Type-safe, full control over LLM pipeline, reports token usage and cost for assertions **See:** [examples/custom-providers.md](examples/custom-providers.md) for inline function providers, programmatic API, CI/CD integration --- ### Pattern 6: CI/CD Integration Run evaluations in CI with quality gates. ```yaml # .github/workflows/llm-eval.yml name: LLM Eval on: pull_request: paths: - "prompts/**" - "promptfooconfig.yaml" jobs: evaluate: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - uses: actions/setup-node@v4 with: node-version: "22" - uses: actions/cache@v4 with: path: ~/.cache/promptfoo key: ${{ runner.os }}-promptfoo-v1 - name: Run eval env: OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }} run: npx promptfoo@latest eval -o results.json --share ``` **Why good:** Caches LLM responses across runs, `promptfoo eval` exits with code 100 on test failures (CI fails automatically), `--share` generates a shareable results URL **See:** [examples/custom-providers.md](examples/custom-providers.md) for npm scripts, quality gate thresholds, programmatic evaluation </patterns> --- <decision_framework> ## Decision Framework ### Which Assertion Type to Use ``` What are you validating? +-- Exact or structural match? | +-- Exact text -> equals | +-- Contains substring -> contains / icontains | +-- Regex pattern -> regex | +-- Valid JSON -> is-json | +-- Valid function call -> is-valid-openai-tools-call | +-- Cost under budget -> cost (with threshold) | +-- Response time -> latency (with threshold) +-- Subjective quality? | +-- General quality criteria -> llm-rubric | +-- Factual accuracy against ground truth -> factuality | +-- Semantic similarity -> similar (with threshold) | +-- Closed-domain QA accuracy -> model-graded-closedqa | +-- RAG context fidelity -> context-faithfulness +-- Custom logic? +-- JavaScript function -> javascript +-- Python function -> python +-- External service -> webhook ``` ### When to Use Red Teaming vs Eval ``` What are you testing? +-- Prompt quality and correctness? | +-- Use promptfoo eval with test cases and assertions +-- Security vulnerabilities? | +-- Use promptfoo redteam run with plugins and strategies +-- Both? +-- Run eval for quality, redteam for security -- separate configs or sections ``` ### Provider Selection ``` How does your LLM integration work? +-- Direct API call to OpenAI/Anthropic/etc? | +-- Use built-in provider: openai:gpt-4o, anthropic:messages:claude-sonnet-4-6 +-- Custom pipeline (RAG, agents, middleware)? | +-- Use custom TypeScript provider: file://providers/my-app.ts +-- HTTP endpoint? | +-- Use HTTP provider: id: https://api.example.com/chat +-- Multiple providers to compare? +-- List all in providers array -- promptfoo runs tests against each ``` </decision_framework> --- <red_flags> ## RED FLAGS **High Priority Issues:** - Tests without `assert` arrays (output is captured but never validated -- tests always "pass") - Not checking `promptfoo eval` exit code in CI (`promptfoo eval` exits 100 on test failures -- ensure your CI pipeline treats non-zero exit codes as failures) - Hardcoded API keys in `promptfooconfig.yaml` (use environment variables) - Using `llm-rubric` for checks that `is-json` or `contains` can do deterministically (wastes money and adds non-determinism) - Red teaming without `purpose` (generic attacks miss application-specific vulnerabilities) **Medium Priority Issues:** - Missing `threshold` on `similar` assertions (default may not match your quality bar) - Not caching in CI (every run makes full API calls -- expensive and slow) - Using `model-graded-closedqa` when `llm-rubric` would be simpler (closedqa is for specific ground-truth QA) - Not setting `provider` on model-graded assertions (uses default which may not be the grader you want) - Running red team with default `numTests: 5` in production scans (too few for comprehensive coverage) **Common Mistakes:** - Confusing `prompts` (the LLM prompt templates) with `tests` (the evaluation cases) -- prompts define what to send, tests define what to check - Using `equals` for natural language output (LLM output is non-deterministic, use `llm-rubric` or `similar`) - Forgetting `{{variable}}` syntax in prompts (promptfoo uses Nunjucks templating, not `${variable}`) - Putting assertions in `defaultTest` that should only apply to specific tests (assertions in `defaultTest` apply to ALL tests) - Using `file://` paths without the prefix (promptfoo treats bare paths as literal strings, not file references) **Gotchas & Edge Cases:** - `promptfoo eval` caches LLM responses by default -- use `promptfoo cache clear` or `--no-cache` to force fresh calls - `--share` uploads results to promptfoo's servers -- do not use with sensitive data unless self-hosting - Red team `strategies` wrap `plugins` output -- a plugin generates the malicious content, a strategy delivers it (e.g., via jailbreak encoding) - `defaultTest.assert` merges with per-test assertions, it does not replace them -- both arrays run - CSV test files map column headers to variable names -- header `input` becomes `{{input}}` in prompts - `transform` in test options runs JavaScript on the output before assertions -- useful for extracting JSON from markdown-wrapped responses - Provider configs in YAML use `config:` key for model parameters (`temperature`, `max_tokens`), not top-level fields - The `weight` property on assertions affects scoring in the results UI but does not change pass/fail behavior </red_flags> --- <critical_reminders> ## CRITICAL REMINDERS > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST define test cases with explicit `assert` arrays -- tests without assertions only capture output without validating it)** **(You MUST use `llm-rubric` for subjective quality evaluation -- do NOT rely solely on deterministic assertions for natural language output)** **(You MUST set `threshold` on similarity and model-graded assertions -- omitting thresholds uses defaults that may not match your quality bar)** **(You MUST use environment variables for all API keys -- never hardcode keys in promptfooconfig.yaml or provider configs)** **(You MUST verify `promptfoo eval` exit code in CI pipelines -- it returns exit code 100 on test failures, exit code 1 on other errors)** **Failure to follow these rules will produce untested, insecure, or falsely-passing LLM evaluation pipelines.** </critical_reminders>
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.