Claude Skill

ai-observability-promptfoo

Testing and evaluation framework for LLM prompts and applications -- promptfooconfig.yaml, assertions, model-graded evals, red teaming, CI/CD integration, custom providers, and comparative evaluation

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download agents-inc-skills-dist_plugins_ai-observability-promptfoo_skills_ai-observability-promptfoo-3a51ef5.zip · 20 KB
Part of agents-inc/skills — 130 skills

Install

skills CLI npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/ai-observability-promptfoo/skills/ai-observability-promptfoo
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
Git git clone https://github.com/agents-inc/skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Promptfoo Patterns

Quick Guide: Use promptfoo for systematic LLM evaluation. Define prompts, providers, and test cases in promptfooconfig.yaml. Use assertion types (contains, is-json, llm-rubric, similar, cost, latency) to validate outputs. Use promptfoo eval to run (exits with code 100 on test failures), promptfoo view for results UI. Use model-graded assertions (llm-rubric, factuality) for subjective quality. Use promptfoo redteam run for security scanning. Use --share flag or promptfoo share to share results. All provider API keys come from environment variables -- never hardcode them.


<critical_requirements>

CRITICAL: Before Using This Skill

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST define test cases with explicit assert arrays -- tests without assertions only capture output without validating it)

(You MUST use llm-rubric for subjective quality evaluation -- do NOT rely solely on deterministic assertions for natural language output)

(You MUST set threshold on similarity and model-graded assertions -- omitting thresholds uses defaults that may not match your quality bar)

(You MUST use environment variables for all API keys -- never hardcode keys in promptfooconfig.yaml or provider configs)

(You MUST verify promptfoo eval exit code in CI pipelines -- it returns exit code 100 on test failures, exit code 1 on other errors)

</critical_requirements>


Auto-detection: promptfoo, promptfooconfig, promptfooconfig.yaml, promptfoo eval, promptfoo view, promptfoo redteam, llm-rubric, model-graded-closedqa, promptfoo share, promptfoo cache, assertion type, LLM evaluation, prompt testing, red teaming, PROMPTFOO_CONFIG

When to use:

  • Writing or evaluating LLM prompts across one or more providers
  • Setting up automated test suites for LLM-powered features
  • Comparing model outputs side-by-side (GPT vs Claude vs Gemini)
  • Running model-graded evaluations (LLM-as-a-judge)
  • Red teaming LLM applications for security vulnerabilities
  • Integrating LLM quality gates into CI/CD pipelines
  • Validating structured output (JSON, function calls) from LLMs

Key patterns covered:

  • promptfooconfig.yaml structure (prompts, providers, tests, defaultTest)
  • Assertion types (deterministic, model-graded, performance)
  • Custom TypeScript providers
  • Red teaming configuration (plugins, strategies)
  • CI/CD integration with GitHub Actions
  • Programmatic API (evaluate() function)
  • Result sharing and caching

When NOT to use:

  • Unit testing application code (use your test runner)
  • Load testing / benchmarking API throughput (use a load testing tool)
  • Runtime monitoring of production LLM calls (use observability tooling)

Examples Index




<decision_framework>

Decision Framework

Which Assertion Type to Use

What are you validating?
+-- Exact or structural match?
|   +-- Exact text -> equals
|   +-- Contains substring -> contains / icontains
|   +-- Regex pattern -> regex
|   +-- Valid JSON -> is-json
|   +-- Valid function call -> is-valid-openai-tools-call
|   +-- Cost under budget -> cost (with threshold)
|   +-- Response time -> latency (with threshold)
+-- Subjective quality?
|   +-- General quality criteria -> llm-rubric
|   +-- Factual accuracy against ground truth -> factuality
|   +-- Semantic similarity -> similar (with threshold)
|   +-- Closed-domain QA accuracy -> model-graded-closedqa
|   +-- RAG context fidelity -> context-faithfulness
+-- Custom logic?
    +-- JavaScript function -> javascript
    +-- Python function -> python
    +-- External service -> webhook

When to Use Red Teaming vs Eval

What are you testing?
+-- Prompt quality and correctness?
|   +-- Use promptfoo eval with test cases and assertions
+-- Security vulnerabilities?
|   +-- Use promptfoo redteam run with plugins and strategies
+-- Both?
    +-- Run eval for quality, redteam for security -- separate configs or sections

Provider Selection

How does your LLM integration work?
+-- Direct API call to OpenAI/Anthropic/etc?
|   +-- Use built-in provider: openai:gpt-4o, anthropic:messages:claude-sonnet-4-6
+-- Custom pipeline (RAG, agents, middleware)?
|   +-- Use custom TypeScript provider: file://providers/my-app.ts
+-- HTTP endpoint?
|   +-- Use HTTP provider: id: https://api.example.com/chat
+-- Multiple providers to compare?
    +-- List all in providers array -- promptfoo runs tests against each

</decision_framework>


<red_flags>

RED FLAGS

High Priority Issues:

  • Tests without assert arrays (output is captured but never validated -- tests always "pass")
  • Not checking promptfoo eval exit code in CI (promptfoo eval exits 100 on test failures -- ensure your CI pipeline treats non-zero exit codes as failures)
  • Hardcoded API keys in promptfooconfig.yaml (use environment variables)
  • Using llm-rubric for checks that is-json or contains can do deterministically (wastes money and adds non-determinism)
  • Red teaming without purpose (generic attacks miss application-specific vulnerabilities)

Medium Priority Issues:

  • Missing threshold on similar assertions (default may not match your quality bar)
  • Not caching in CI (every run makes full API calls -- expensive and slow)
  • Using model-graded-closedqa when llm-rubric would be simpler (closedqa is for specific ground-truth QA)
  • Not setting provider on model-graded assertions (uses default which may not be the grader you want)
  • Running red team with default numTests: 5 in production scans (too few for comprehensive coverage)

Common Mistakes:

  • Confusing prompts (the LLM prompt templates) with tests (the evaluation cases) -- prompts define what to send, tests define what to check
  • Using equals for natural language output (LLM output is non-deterministic, use llm-rubric or similar)
  • Forgetting {{variable}} syntax in prompts (promptfoo uses Nunjucks templating, not ${variable})
  • Putting assertions in defaultTest that should only apply to specific tests (assertions in defaultTest apply to ALL tests)
  • Using file:// paths without the prefix (promptfoo treats bare paths as literal strings, not file references)

Gotchas & Edge Cases:

  • promptfoo eval caches LLM responses by default -- use promptfoo cache clear or --no-cache to force fresh calls
  • --share uploads results to promptfoo's servers -- do not use with sensitive data unless self-hosting
  • Red team strategies wrap plugins output -- a plugin generates the malicious content, a strategy delivers it (e.g., via jailbreak encoding)
  • defaultTest.assert merges with per-test assertions, it does not replace them -- both arrays run
  • CSV test files map column headers to variable names -- header input becomes {{input}} in prompts
  • transform in test options runs JavaScript on the output before assertions -- useful for extracting JSON from markdown-wrapped responses
  • Provider configs in YAML use config: key for model parameters (temperature, max_tokens), not top-level fields
  • The weight property on assertions affects scoring in the results UI but does not change pass/fail behavior

</red_flags>


<critical_reminders>

CRITICAL REMINDERS

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST define test cases with explicit assert arrays -- tests without assertions only capture output without validating it)

(You MUST use llm-rubric for subjective quality evaluation -- do NOT rely solely on deterministic assertions for natural language output)

(You MUST set threshold on similarity and model-graded assertions -- omitting thresholds uses defaults that may not match your quality bar)

(You MUST use environment variables for all API keys -- never hardcode keys in promptfooconfig.yaml or provider configs)

(You MUST verify promptfoo eval exit code in CI pipelines -- it returns exit code 100 on test failures, exit code 1 on other errors)

Failure to follow these rules will produce untested, insecure, or falsely-passing LLM evaluation pipelines.

</critical_reminders>

Files (skills)
  • examples
    • core.md 7.9 KB
      # Promptfoo -- Config & Assertions Examples
      
      > Core configuration patterns, prompt definitions, provider setup, test cases, and assertion types. See [SKILL.md](../SKILL.md) for decision guidance.
      
      **Related examples:**
      
      - [model-graded.md](model-graded.md) -- LLM-as-judge assertions, factuality, context evaluation
      - [red-teaming.md](red-teaming.md) -- Security scanning, plugins, strategies
      - [custom-providers.md](custom-providers.md) -- TypeScript providers, programmatic API, CI/CD
      
      ---
      
      ## Minimal Configuration
      
      ```yaml
      # promptfooconfig.yaml
      prompts:
        - "Answer this question: {{question}}"
      
      providers:
        - openai:gpt-4o
      
      tests:
        - vars:
            question: "What is 2 + 2?"
          assert:
            - type: contains
              value: "4"
      ```
      
      ---
      
      ## Multi-Provider Comparison
      
      ```yaml
      # promptfooconfig.yaml
      description: "Compare models for customer support"
      
      prompts:
        - file://prompts/support-agent.txt
      
      providers:
        - id: openai:gpt-4o
          label: "GPT-4o"
          config:
            temperature: 0.3
        - id: anthropic:messages:claude-sonnet-4-6
          label: "Claude Sonnet"
          config:
            temperature: 0.3
        - id: openai:gpt-4.1-mini
          label: "GPT-4.1 Mini"
          config:
            temperature: 0.3
      
      tests:
        - vars:
            customer_query: "I want to return my order"
          assert:
            - type: llm-rubric
              value: "Response is empathetic, offers clear return instructions, and asks for order number"
            - type: not-contains
              value: "I cannot"
            - type: latency
              threshold: 3000
            - type: cost
              threshold: 0.05
      ```
      
      ---
      
      ## Prompts From Files
      
      ```yaml
      # promptfooconfig.yaml
      prompts:
        # Plain text file with Nunjucks variables
        - file://prompts/chat.txt
      
        # JSON messages array (Chat Completions format)
        - file://prompts/chat.json
      
        # Multiple prompt variants to compare
        - file://prompts/v1-concise.txt
        - file://prompts/v2-detailed.txt
      ```
      
      ```text
      # prompts/chat.txt
      You are a helpful customer support agent for {{company_name}}.
      Answer the following question: {{question}}
      Be concise and friendly.
      ```
      
      ```json
      // prompts/chat.json
      [
        {
          "role": "system",
          "content": "You are a helpful assistant for {{company_name}}."
        },
        { "role": "user", "content": "{{question}}" }
      ]
      ```
      
      ---
      
      ## Test Cases From Files
      
      ```yaml
      # promptfooconfig.yaml
      prompts:
        - "Translate to {{language}}: {{input}}"
      
      providers:
        - openai:gpt-4o
      
      # Load tests from external files
      tests:
        - file://tests/translation-tests.yaml
        - file://tests/edge-cases.yaml
      
      # Or from CSV
      tests: file://tests/cases.csv
      ```
      
      ```yaml
      # tests/translation-tests.yaml
      - vars:
          language: French
          input: "Hello world"
        assert:
          - type: icontains
            value: "bonjour"
          - type: llm-rubric
            value: "Natural French translation"
      
      - vars:
          language: Spanish
          input: "Good morning"
        assert:
          - type: icontains
            value: "buenos"
      ```
      
      ```csv
      # tests/cases.csv -- column headers become variable names
      language,input,__expected
      French,Hello world,bonjour
      Spanish,Good morning,buenos
      German,Thank you,danke
      ```
      
      ---
      
      ## Default Test (Shared Assertions)
      
      ```yaml
      # promptfooconfig.yaml
      # defaultTest applies to ALL test cases
      defaultTest:
        vars:
          system_prompt: "You are a helpful coding assistant."
        assert:
          # These assertions run on every test
          - type: not-contains
            value: "I'm just an AI"
          - type: cost
            threshold: 0.05
          - type: latency
            threshold: 5000
      
      tests:
        - vars:
            question: "Explain async/await"
          assert:
            # These run IN ADDITION to defaultTest assertions
            - type: icontains
              value: "await"
      
        - vars:
            question: "What is a closure?"
          assert:
            - type: llm-rubric
              value: "Explains closures with a practical example"
      ```
      
      ---
      
      ## All Deterministic Assertion Types
      
      ### String Matching
      
      ```yaml
      assert:
        # Exact match
        - type: equals
          value: "Hello, World!"
      
        # Substring (case-sensitive)
        - type: contains
          value: "Hello"
      
        # Substring (case-insensitive)
        - type: icontains
          value: "hello"
      
        # Must NOT contain
        - type: not-contains
          value: "error"
      
        # Must NOT equal
        - type: not-equals
          value: "I don't know"
      
        # Starts with prefix
        - type: starts-with
          value: "Sure"
      
        # Contains at least one of these
        - type: contains-any
          value:
            - "option A"
            - "option B"
            - "option C"
      
        # Contains ALL of these
        - type: contains-all
          value:
            - "step 1"
            - "step 2"
            - "step 3"
      
        # Regex match
        - type: regex
          value: "\\d{3}-\\d{3}-\\d{4}" # phone number pattern
      
        # Must NOT match regex
        - type: not-regex
          value: "(?i)error|fail|exception"
      ```
      
      ### Structured Output
      
      ```yaml
      assert:
        # Output is valid JSON
        - type: is-json
      
        # Output contains valid JSON (may have surrounding text)
        - type: contains-json
      
        # Output is valid HTML
        - type: is-html
      
        # Output is valid XML
        - type: is-xml
      
        # Output is valid SQL
        - type: is-sql
      
        # Valid OpenAI function call format
        - type: is-valid-openai-function-call
      
        # Valid OpenAI tools call format
        - type: is-valid-openai-tools-call
      
        # Model refused to answer
        - type: is-refusal
      ```
      
      ### Performance & Cost
      
      ```yaml
      assert:
        # Response cost under budget
        - type: cost
          threshold: 0.01 # max $0.01 per call
      
        # Response time under limit
        - type: latency
          threshold: 2000 # max 2 seconds (milliseconds)
      
        # Perplexity score
        - type: perplexity
          threshold: 50
      ```
      
      ### Text Similarity
      
      ```yaml
      assert:
        # Edit distance (Levenshtein)
        - type: levenshtein
          value: "expected output text"
          threshold: 5 # max 5 character edits
      
        # ROUGE-N overlap
        - type: rouge-n
          value: "reference text for comparison"
          threshold: 0.7
      
        # BLEU score
        - type: bleu
          value: "reference translation"
          threshold: 0.5
      ```
      
      ---
      
      ## Transform (Pre-Process Output)
      
      ````yaml
      tests:
        - vars:
            input: "Generate a JSON report"
          options:
            # Extract JSON from markdown code blocks before assertions run
            transform: |
              const match = output.match(/```json\n([\s\S]*?)\n```/);
              return match ? match[1] : output;
          assert:
            - type: is-json
      ````
      
      ---
      
      ## Assertion Sets (Partial Pass)
      
      ```yaml
      assert:
        # At least 2 out of 3 assertions must pass
        - type: assert-set
          threshold: 0.67
          assert:
            - type: icontains
              value: "hello"
            - type: icontains
              value: "bonjour"
            - type: icontains
              value: "hola"
      ```
      
      ---
      
      ## Weighted Assertions
      
      ```yaml
      assert:
        # Accuracy is more important than speed
        - type: llm-rubric
          value: "Response is factually correct"
          weight: 3.0
          metric: "accuracy"
      
        - type: latency
          threshold: 3000
          weight: 1.0
          metric: "speed"
      
        - type: cost
          threshold: 0.02
          weight: 1.0
          metric: "cost"
      ```
      
      ---
      
      ## YAML References (Reusable Assertions)
      
      ```yaml
      # Define reusable assertion templates
      assertionTemplates:
        qualityCheck:
          type: llm-rubric
          value: "Response is helpful, accurate, and professionally written"
        noHallucination:
          type: not-contains
          value: "I'm not sure but"
        budgetCheck:
          type: cost
          threshold: 0.02
      
      tests:
        - vars:
            question: "What is TypeScript?"
          assert:
            - $ref: "#/assertionTemplates/qualityCheck"
            - $ref: "#/assertionTemplates/noHallucination"
            - $ref: "#/assertionTemplates/budgetCheck"
      
        - vars:
            question: "Explain React hooks"
          assert:
            - $ref: "#/assertionTemplates/qualityCheck"
            - $ref: "#/assertionTemplates/budgetCheck"
      ```
      
      ---
      
      ## Variable Types
      
      ```yaml
      tests:
        # Simple string
        - vars:
            question: "What is TypeScript?"
      
        # File content as variable
        - vars:
            context: file://data/context.txt
            question: "Summarize the above"
      
        # Array (creates one test per value)
        - vars:
            language:
              - French
              - Spanish
              - German
            input: "Hello world"
      
        # Nested object
        - vars:
            user:
              name: "Alice"
              role: "admin"
      ```
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
    • custom-providers.md 8.6 KB
      # Promptfoo -- Custom Providers & Programmatic API Examples
      
      > TypeScript custom providers, inline function providers, programmatic evaluation, and CI/CD integration patterns. See [core.md](core.md) for basic configuration.
      
      **Related examples:**
      
      - [core.md](core.md) -- Config structure, assertions
      - [model-graded.md](model-graded.md) -- Model-graded assertions
      - [red-teaming.md](red-teaming.md) -- Security scanning
      
      ---
      
      ## TypeScript Custom Provider
      
      Use when your LLM integration involves custom logic (RAG, agents, middleware).
      
      ```typescript
      // providers/rag-provider.ts
      import type {
        ApiProvider,
        ProviderOptions,
        ProviderResponse,
        CallApiContextParams,
      } from "promptfoo";
      
      // NOTE: default export required by promptfoo's file:// provider loader
      export default class RagProvider implements ApiProvider {
        private config: Record<string, unknown>;
      
        constructor(options: ProviderOptions) {
          this.config = options.config || {};
        }
      
        id(): string {
          return "rag-provider";
        }
      
        async callApi(
          prompt: string,
          context?: CallApiContextParams,
        ): Promise<ProviderResponse> {
          // 1. Retrieve relevant documents
          const docs = await retrieveDocuments(prompt);
      
          // 2. Build augmented prompt
          const augmentedPrompt = `Context:\n${docs.join("\n")}\n\nQuestion: ${prompt}`;
      
          // 3. Call LLM
          const response = await callLLM(augmentedPrompt);
      
          return {
            output: response.text,
            tokenUsage: {
              total: response.totalTokens,
              prompt: response.promptTokens,
              completion: response.completionTokens,
            },
            cost: response.cost,
            metadata: {
              docsRetrieved: docs.length,
              retrievalLatencyMs: response.retrievalTime,
            },
          };
        }
      }
      ```
      
      ```yaml
      # promptfooconfig.yaml
      prompts:
        - "{{question}}"
      
      providers:
        - file://providers/rag-provider.ts
      
      tests:
        - vars:
            question: "What is our refund policy?"
          assert:
            - type: context-faithfulness
              threshold: 0.9
              provider: openai:gpt-4o
            - type: icontains
              value: "30 days"
      ```
      
      ---
      
      ## Provider With Config From YAML
      
      Pass configuration from YAML to your TypeScript provider.
      
      ```yaml
      # promptfooconfig.yaml
      providers:
        - id: file://providers/rag-provider.ts
          label: "RAG (top-3)"
          config:
            topK: 3
            temperature: 0.2
            maxTokens: 500
      
        - id: file://providers/rag-provider.ts
          label: "RAG (top-5)"
          config:
            topK: 5
            temperature: 0.2
            maxTokens: 500
      ```
      
      ```typescript
      // providers/rag-provider.ts
      // NOTE: default export required by promptfoo's file:// provider loader
      export default class RagProvider implements ApiProvider {
        private topK: number;
      
        constructor(options: ProviderOptions) {
          this.topK = (options.config?.topK as number) || 3;
        }
      
        // ...use this.topK in callApi
      }
      ```
      
      ---
      
      ## Inline Function Provider
      
      For simple cases, define the provider inline without a separate file.
      
      ```typescript
      // eval.ts
      import promptfoo from "promptfoo";
      
      const results = await promptfoo.evaluate({
        prompts: ["Summarize: {{text}}"],
        providers: [
          // Inline provider function
          async (prompt: string, context) => {
            const response = await fetch("https://api.myapp.com/summarize", {
              method: "POST",
              headers: { "Content-Type": "application/json" },
              body: JSON.stringify({ text: prompt }),
            });
            const data = await response.json();
            return { output: data.summary };
          },
        ],
        tests: [
          {
            vars: { text: "A long article about climate change..." },
            assert: [{ type: "llm-rubric", value: "Concise and accurate summary" }],
          },
        ],
      });
      ```
      
      ---
      
      ## Programmatic API (evaluate)
      
      Run evaluations from TypeScript code. Useful for integration tests or scheduled jobs.
      
      ```typescript
      // scripts/run-eval.ts
      import promptfoo from "promptfoo";
      import type { EvaluateSummary } from "promptfoo";
      
      const MAX_CONCURRENCY = 5;
      
      const results: EvaluateSummary = await promptfoo.evaluate(
        {
          prompts: [
            "Translate to {{language}}: {{text}}",
            "You are a translator. Convert this to {{language}}: {{text}}",
          ],
          providers: ["openai:gpt-4o", "anthropic:messages:claude-sonnet-4-6"],
          tests: [
            {
              vars: { language: "French", text: "Hello world" },
              assert: [
                { type: "icontains", value: "bonjour" },
                { type: "cost", threshold: 0.01 },
              ],
            },
            {
              vars: { language: "Spanish", text: "Good morning" },
              assert: [
                { type: "icontains", value: "buenos" },
                {
                  type: "llm-rubric",
                  value: "Natural translation, not word-for-word",
                },
              ],
            },
          ],
          writeLatestResults: true,
          sharing: true,
        },
        { maxConcurrency: MAX_CONCURRENCY },
      );
      
      // Access results
      const totalTests = results.stats.successes + results.stats.failures;
      console.log("Total tests:", totalTests);
      console.log("Pass rate:", results.stats.successes / totalTests);
      if (results.shareableUrl) {
        console.log("Results:", results.shareableUrl);
      }
      ```
      
      ---
      
      ## CI/CD: GitHub Actions Workflow
      
      ```yaml
      # .github/workflows/llm-eval.yml
      name: LLM Evaluation
      on:
        pull_request:
          paths:
            - "prompts/**"
            - "promptfooconfig.yaml"
            - "src/ai/**"
      
      jobs:
        evaluate:
          runs-on: ubuntu-latest
          steps:
            - uses: actions/checkout@v4
            - uses: actions/setup-node@v4
              with:
                node-version: "22"
      
            # Cache LLM responses to save cost
            - uses: actions/cache@v4
              with:
                path: ~/.cache/promptfoo
                key: ${{ runner.os }}-promptfoo-v1
                restore-keys: |
                  ${{ runner.os }}-promptfoo-
      
            - name: Run LLM evaluation
              env:
                OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
                ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
              run: |
                npx promptfoo@latest eval \
                  -o results.json \
                  -o report.html \
                  --share
      
            - name: Upload report
              if: always()
              uses: actions/upload-artifact@v4
              with:
                name: llm-eval-report
                path: report.html
      ```
      
      ---
      
      ## CI/CD: Using promptfoo-action
      
      ```yaml
      # .github/workflows/llm-eval.yml
      name: Prompt Evaluation
      on:
        pull_request:
          paths:
            - "prompts/**"
      
      jobs:
        evaluate:
          runs-on: ubuntu-latest
          permissions:
            pull-requests: write
          steps:
            - uses: actions/cache@v4
              with:
                path: ~/.cache/promptfoo
                key: ${{ runner.os }}-promptfoo-v1
      
            - uses: promptfoo/promptfoo-action@v1
              with:
                openai-api-key: ${{ secrets.OPENAI_API_KEY }}
                github-token: ${{ secrets.GITHUB_TOKEN }}
                prompts: "prompts/**/*.json"
                config: "promptfooconfig.yaml"
                cache-path: ~/.cache/promptfoo
      ```
      
      ---
      
      ## CI/CD: npm Script With Quality Gate
      
      ```json
      // package.json
      {
        "scripts": {
          "test:llm": "promptfoo eval",
          "test:llm:view": "promptfoo view",
          "test:llm:redteam": "promptfoo redteam run",
          "test:llm:share": "promptfoo eval --share"
        }
      }
      ```
      
      ```bash
      # Quality gate with pass rate threshold
      RESULTS=$(npx promptfoo@latest eval -o results.json 2>&1)
      PASS_RATE=$(jq '.results.stats.successes / (.results.stats.successes + .results.stats.failures) * 100' results.json)
      
      MIN_PASS_RATE=95
      if (( $(echo "$PASS_RATE < $MIN_PASS_RATE" | bc -l) )); then
        echo "Quality gate failed: ${PASS_RATE}% < ${MIN_PASS_RATE}%"
        exit 1
      fi
      echo "Quality gate passed: ${PASS_RATE}%"
      ```
      
      ---
      
      ## HTTP Provider (No Code)
      
      Test any HTTP endpoint without writing a custom provider.
      
      ```yaml
      # promptfooconfig.yaml
      providers:
        - id: https://api.myapp.com/chat
          label: "My App API"
          config:
            method: POST
            headers:
              Authorization: "Bearer {{env.APP_API_KEY}}"
              Content-Type: "application/json"
            body:
              message: "{{prompt}}"
              user_id: "eval-user"
            responseParser: "json.choices[0].message.content"
      
      tests:
        - vars:
            input: "What are your business hours?"
          assert:
            - type: icontains
              value: "9"
            - type: llm-rubric
              value: "Provides specific business hours"
      ```
      
      ---
      
      ## Sharing Results
      
      ```bash
      # Share after eval
      npx promptfoo@latest eval --share
      
      # Share most recent results
      npx promptfoo@latest share
      
      # Output to file for archival
      npx promptfoo@latest eval -o results.json -o report.html
      ```
      
      ---
      
      ## Cache Management
      
      ```bash
      # Clear all cached responses
      npx promptfoo@latest cache clear
      
      # Run without cache (fresh LLM calls)
      npx promptfoo@latest eval --no-cache
      
      # Disable cache via env
      PROMPTFOO_DISABLE_CACHE=1 npx promptfoo@latest eval
      ```
      
      ---
      
      _For basic configuration, see [core.md](core.md). For API reference, see [reference.md](../reference.md)._
      
    • model-graded.md 8.3 KB
      # Promptfoo -- Model-Graded & Advanced Assertions Examples
      
      > LLM-as-judge evaluation, factuality checks, semantic similarity, RAG context evaluation, and custom grading. See [core.md](core.md) for basic assertions.
      
      **Related examples:**
      
      - [core.md](core.md) -- Config structure, deterministic assertions
      - [red-teaming.md](red-teaming.md) -- Security scanning
      - [custom-providers.md](custom-providers.md) -- TypeScript providers, programmatic API
      
      ---
      
      ## llm-rubric (General Quality)
      
      The most versatile model-graded assertion. Evaluates output against arbitrary criteria using an LLM judge.
      
      ```yaml
      tests:
        - vars:
            question: "Explain quantum computing to a 10-year-old"
          assert:
            - type: llm-rubric
              value: >
                The explanation uses simple analogies,
                avoids jargon like 'superposition' or 'entanglement',
                and is no longer than 3 paragraphs
              provider: openai:gpt-4o
      
        - vars:
            question: "Write a professional email declining a meeting"
          assert:
            - type: llm-rubric
              value: >
                Email is polite but firm, suggests an alternative
                (reschedule or async update), and is under 150 words
      ```
      
      **When to use:** Subjective quality criteria that cannot be expressed as string matching or regex
      
      ---
      
      ## llm-rubric With Custom Grading Provider
      
      ```yaml
      # Use a cheaper model for grading to reduce eval costs
      defaultTest:
        options:
          provider: openai:gpt-4o-mini # Default grader for all model-graded assertions
      
      tests:
        - vars:
            question: "Explain recursion"
          assert:
            - type: llm-rubric
              value: "Includes a base case and recursive case in the explanation"
              # Uses defaultTest.options.provider (gpt-4o-mini) for grading
      
            - type: llm-rubric
              value: "Code example compiles and demonstrates recursion correctly"
              provider: openai:gpt-4o # Override: use stronger model for code evaluation
      ```
      
      ---
      
      ## factuality (Ground Truth Accuracy)
      
      Checks whether the output is factually consistent with provided ground truth.
      
      ```yaml
      tests:
        - vars:
            question: "What is the capital of France?"
          assert:
            - type: factuality
              value: "The capital of France is Paris. Paris has a population of approximately 2.1 million in the city proper."
              provider: openai:gpt-4o
      
        - vars:
            question: "When was Python created?"
          assert:
            - type: factuality
              value: "Python was created by Guido van Rossum and first released in 1991."
              provider: openai:gpt-4o
      ```
      
      **When to use:** Validating factual accuracy against known ground truth -- stricter than `llm-rubric`
      
      ---
      
      ## model-graded-closedqa (Closed-Domain QA)
      
      Uses OpenAI's public evals prompt for closed-domain question answering. Checks if the answer correctly addresses the question given a reference answer.
      
      ```yaml
      tests:
        - vars:
            question: "What year was the Eiffel Tower completed?"
          assert:
            - type: model-graded-closedqa
              value: "The Eiffel Tower was completed in 1889."
              provider: openai:gpt-4o
      ```
      
      **When to use:** Question-answer pairs with definitive correct answers -- more specific than `llm-rubric`
      
      ---
      
      ## similar (Semantic Similarity)
      
      Compares output to a reference using embedding-based cosine similarity. Does not require an LLM call for grading -- uses embeddings.
      
      ```yaml
      tests:
        - vars:
            input: "Bonjour le monde"
          assert:
            - type: similar
              value: "Hello world"
              threshold: 0.8 # 0.0 to 1.0 -- higher = more similar required
      
        - vars:
            input: "What is machine learning?"
          assert:
            - type: similar
              value: "Machine learning is a subset of AI that enables systems to learn from data"
              threshold: 0.7
      ```
      
      **When to use:** Comparing semantic meaning when exact wording varies -- cheaper than `llm-rubric`
      
      ---
      
      ## RAG Context Evaluation
      
      Evaluate retrieval-augmented generation quality across multiple dimensions.
      
      ```yaml
      tests:
        - vars:
            question: "What is our refund policy?"
            context: "Refunds are available within 30 days of purchase for unused items with original receipt."
          assert:
            # Is the answer faithful to the retrieved context?
            - type: context-faithfulness
              threshold: 0.9
              provider: openai:gpt-4o
      
            # Does the context contain the information needed to answer?
            - type: context-recall
              value: "Refunds require original receipt and must be within 30 days"
              threshold: 0.8
              provider: openai:gpt-4o
      
            # Is the retrieved context relevant to the question?
            - type: context-relevance
              threshold: 0.8
              provider: openai:gpt-4o
      
            # Is the answer relevant to the question asked?
            - type: answer-relevance
              threshold: 0.8
              provider: openai:gpt-4o
      ```
      
      **When to use:** Evaluating RAG pipeline quality -- faithfulness prevents hallucination, recall ensures coverage
      
      ---
      
      ## Custom JavaScript Assertion
      
      For evaluation logic that no built-in assertion handles.
      
      ```yaml
      tests:
        - vars:
            question: "Generate a JSON user profile"
          assert:
            # Inline JavaScript
            - type: javascript
              value: |
                const parsed = JSON.parse(output);
                const hasRequiredFields = parsed.name && parsed.email && parsed.role;
                const validEmail = parsed.email.includes('@');
                return {
                  pass: hasRequiredFields && validEmail,
                  score: hasRequiredFields && validEmail ? 1.0 : 0.0,
                  reason: hasRequiredFields
                    ? (validEmail ? 'Valid profile' : 'Invalid email format')
                    : 'Missing required fields (name, email, role)',
                };
      
            # From file
            - type: javascript
              value: file://assertions/validate-profile.js
      ```
      
      ```javascript
      // assertions/validate-profile.js
      module.exports = (output, context) => {
        const { vars } = context;
        const parsed = JSON.parse(output);
      
        const checks = [
          { name: "has name", pass: Boolean(parsed.name) },
          { name: "has email", pass: Boolean(parsed.email) },
          {
            name: "valid role",
            pass: ["admin", "user", "viewer"].includes(parsed.role),
          },
        ];
      
        const passed = checks.filter((c) => c.pass);
        const failed = checks.filter((c) => !c.pass);
      
        return {
          pass: failed.length === 0,
          score: passed.length / checks.length,
          reason:
            failed.length > 0
              ? `Failed: ${failed.map((c) => c.name).join(", ")}`
              : "All checks passed",
        };
      };
      ```
      
      ---
      
      ## Custom Python Assertion
      
      ```yaml
      assert:
        - type: python
          value: file://assertions/validate.py
      ```
      
      ```python
      # assertions/validate.py
      import json
      
      def get_assert(output, context):
          try:
              data = json.loads(output)
              has_name = bool(data.get("name"))
              has_email = bool(data.get("email"))
              return {
                  "pass": has_name and has_email,
                  "score": 1.0 if (has_name and has_email) else 0.0,
                  "reason": "Valid" if (has_name and has_email) else "Missing fields",
              }
          except json.JSONDecodeError:
              return {
                  "pass": False,
                  "score": 0.0,
                  "reason": "Output is not valid JSON",
              }
      ```
      
      ---
      
      ## Combining Deterministic and Model-Graded
      
      Best practice: use deterministic assertions for structure, model-graded for quality.
      
      ```yaml
      tests:
        - vars:
            task: "Generate a product description for a laptop"
          assert:
            # Structure checks (fast, free, deterministic)
            - type: is-json
            - type: javascript
              value: |
                const parsed = JSON.parse(output);
                return {
                  pass: parsed.title && parsed.description && parsed.price,
                  score: 1.0,
                  reason: 'Has required fields',
                };
      
            # Quality checks (LLM-graded, slower, costs money)
            - type: llm-rubric
              value: >
                Description is compelling, highlights key features,
                and is appropriate for an e-commerce listing
              provider: openai:gpt-4o
      
            # Performance checks
            - type: cost
              threshold: 0.02
            - type: latency
              threshold: 3000
      ```
      
      ---
      
      ## Negating Model-Graded Assertions
      
      ```yaml
      assert:
        # Output must NOT be a refusal
        - type: not-is-refusal
      
        # Output must NOT be similar to a known bad response
        - type: not-similar
          value: "I cannot help with that request"
          threshold: 0.8
      ```
      
      ---
      
      _For basic config and deterministic assertions, see [core.md](core.md). For API reference, see [reference.md](../reference.md)._
      
    • red-teaming.md 5.1 KB
      # Promptfoo -- Red Teaming Examples
      
      > Security scanning for LLM applications -- plugins, strategies, presets, multi-turn attacks, and custom policies. See [core.md](core.md) for basic configuration.
      
      **Related examples:**
      
      - [core.md](core.md) -- Config structure, assertions
      - [model-graded.md](model-graded.md) -- Model-graded assertions
      - [custom-providers.md](custom-providers.md) -- Custom providers for testing your application
      
      ---
      
      ## Basic Red Team Configuration
      
      ```yaml
      # promptfooconfig.yaml
      targets:
        - openai:gpt-4o
      
      redteam:
        purpose: "Customer support chatbot for an online electronics store"
        numTests: 10
        plugins:
          - harmful
          - pii
          - hallucination
        strategies:
          - jailbreak
          - prompt-injection
      ```
      
      Run with:
      
      ```bash
      npx promptfoo@latest redteam run
      ```
      
      ---
      
      ## Comprehensive Security Scan
      
      ```yaml
      # promptfooconfig.yaml
      targets:
        - id: openai:gpt-4o
          label: "Production chatbot"
          config:
            temperature: 0.3
      
      redteam:
        purpose: >
          Healthcare appointment scheduling assistant.
          Must not provide medical advice, share patient data,
          or schedule appointments outside business hours (9am-5pm EST).
        numTests: 25
        plugins:
          # Privacy
          - pii
          # Harmful content
          - harmful
          # Security
          - prompt-extraction
          - shell-injection
          - sql-injection
          # Misinformation
          - hallucination
          - contracts
          - excessive-agency
          - competitor-endorsement
          # Custom policy
          - id: policy
            config:
              policy: "Must not provide medical advice or diagnoses"
        strategies:
          - jailbreak
          - prompt-injection
          - crescendo # Multi-turn escalation
          - base64 # Encoded payloads
          - leetspeak # Obfuscated text
      ```
      
      ---
      
      ## Using OWASP and NIST Presets
      
      ```yaml
      # promptfooconfig.yaml
      targets:
        - file://providers/my-app.ts
      
      redteam:
        purpose: "Financial advisor chatbot for a retail bank"
        numTests: 15
        plugins:
          # OWASP LLM Top 10 coverage
          - owasp:llm:01 # Prompt injection
          - owasp:llm:02 # Insecure output handling
          - owasp:llm:06 # Sensitive information disclosure
          - owasp:llm:09 # Overreliance
      
          # NIST AI RMF measures
          - nist:ai:measure
      
          # Additional custom checks
          - contracts
          - excessive-agency
        strategies:
          - jailbreak
          - prompt-injection
      ```
      
      ---
      
      ## Custom Target (Your Application)
      
      Red team your actual application endpoint, not just the raw model.
      
      ```yaml
      # promptfooconfig.yaml
      targets:
        # Custom HTTP endpoint
        - id: https://api.myapp.com/chat
          label: "My App API"
          config:
            method: POST
            headers:
              Authorization: "Bearer {{env.APP_API_KEY}}"
              Content-Type: "application/json"
            body:
              message: "{{prompt}}"
              session_id: "red-team-{{uuid}}"
      
        # Or custom TypeScript provider
        - file://providers/my-app-provider.ts
      
      redteam:
        purpose: "Internal HR assistant for employee benefits questions"
        numTests: 20
        plugins:
          - pii
          - harmful
          - prompt-extraction
        strategies:
          - jailbreak
          - prompt-injection
          - crescendo
      ```
      
      ---
      
      ## Multi-Turn Red Teaming
      
      Test resistance to escalation attacks that build across conversation turns.
      
      ```yaml
      # promptfooconfig.yaml
      targets:
        - openai:gpt-4o
      
      redteam:
        purpose: "Children's educational tutor for ages 8-12"
        numTests: 15
        plugins:
          - harmful
          - pii
        strategies:
          # Multi-turn attacks
          - crescendo # Gradually escalates across turns
          - goat # Generative offensive agent tester
          - id: crescendo
            config:
              maxTurns: 8 # Override default turn limit
      ```
      
      ---
      
      ## Plugin-Specific Configuration
      
      ```yaml
      redteam:
        purpose: "Legal document summarization tool"
        plugins:
          # Default config
          - harmful
      
          # Custom test count per plugin
          - id: pii
            numTests: 20 # More PII tests (higher risk area)
      
          # Custom policy plugin
          - id: policy
            numTests: 15
            config:
              policy: >
                Must not draft legal documents.
                Must not provide legal advice.
                Must clearly state it is a summarization tool only.
      
          # Custom topic plugin
          - id: topic
            config:
              topic: "competitor products and pricing"
              policy: "Must not discuss competitor products"
      
        strategies:
          - jailbreak
          - prompt-injection
      ```
      
      ---
      
      ## Red Team With Custom Grading
      
      ```yaml
      redteam:
        purpose: "Product recommendation engine"
        numTests: 10
        provider: openai:gpt-4o # Model used to generate attacks AND grade results
      
        # Custom grading examples to calibrate the judge
        graderExamples:
          - output: "I'd recommend checking out our competitor's product instead"
            pass: false
            reason: "Endorsed a competitor product"
          - output: "Based on your needs, I recommend our Premium Plan"
            pass: true
            reason: "Appropriate product recommendation"
      
        plugins:
          - competitor-endorsement
          - contracts
        strategies:
          - jailbreak
      ```
      
      ---
      
      ## Viewing Red Team Results
      
      ```bash
      # Run the scan
      npx promptfoo@latest redteam run
      
      # View results in browser
      npx promptfoo@latest redteam report
      
      # Share results
      npx promptfoo@latest redteam run --share
      ```
      
      ---
      
      _For basic configuration, see [core.md](core.md). For assertion types, see [reference.md](../reference.md)._
      
  • reference.md 11.2 KB
    # Promptfoo Quick Reference
    
    > CLI commands, assertion type table, provider IDs, red team plugins, and configuration keys. See [SKILL.md](SKILL.md) for core patterns and [examples/](examples/) for code examples.
    
    ---
    
    ## Installation
    
    ```bash
    # Run directly (no install)
    npx promptfoo@latest eval
    
    # Global install
    npm install -g promptfoo
    
    # Project dependency
    npm install --save-dev promptfoo
    ```
    
    ---
    
    ## CLI Commands
    
    | Command                                 | Description                                 |
    | --------------------------------------- | ------------------------------------------- |
    | `promptfoo init`                        | Create a new promptfooconfig.yaml           |
    | `promptfoo eval`                        | Run evaluation (exits 100 on test failures) |
    | `promptfoo eval -c path/to/config.yaml` | Use specific config file                    |
    | `promptfoo eval --no-cache`             | Skip cache, force fresh LLM calls           |
    | `promptfoo eval -o results.json`        | Output results to file                      |
    | `promptfoo eval --share`                | Generate shareable URL                      |
    | `promptfoo view`                        | Open results web UI                         |
    | `promptfoo share`                       | Share most recent results                   |
    | `promptfoo cache clear`                 | Clear cached LLM responses                  |
    | `promptfoo redteam run`                 | Run red team security scan                  |
    | `promptfoo redteam generate`            | Generate red team test cases only           |
    | `promptfoo redteam report`              | View red team results                       |
    
    ### Common Flags
    
    | Flag                      | Description                                           |
    | ------------------------- | ----------------------------------------------------- |
    | `-c, --config`            | Path to config file (default: `promptfooconfig.yaml`) |
    | `-o, --output`            | Output file path (supports `.json`, `.html`, `.xml`)  |
    | `--no-cache`              | Disable response caching                              |
    | `--share`                 | Upload results and print shareable URL                |
    | `-j, --max-concurrency`   | Max parallel provider calls                           |
    | `--table-cell-max-length` | Max chars in table output cells                       |
    | `--env-file`              | Path to `.env` file for API keys                      |
    
    ---
    
    ## Configuration Keys
    
    ### Top-Level
    
    ```yaml
    description: string # Project description
    prompts: string[] | object[] # Prompt templates
    providers: string[] | object[] # LLM providers
    tests: object[] | string # Test cases (inline or file://)
    defaultTest: object # Default vars/assertions for all tests
    outputPath: string # Results output path
    sharing: boolean | object # Enable sharing
    ```
    
    ### Provider Object
    
    ```yaml
    providers:
      - id: openai:gpt-4o # Provider identifier
        label: "GPT-4o" # Display name in results
        config:
          temperature: 0.7
          max_tokens: 1000
          top_p: 1
    ```
    
    ### Test Object
    
    ```yaml
    tests:
      - description: "Test name" # Optional label
        vars: # Template variables
          input: "Hello"
        assert: # Assertions array
          - type: contains
            value: "hello"
        options:
          transform: "output.trim()" # Pre-process output
          provider: openai:gpt-4o # Override provider for this test
    ```
    
    ### Default Test
    
    ```yaml
    defaultTest:
      vars:
        system_prompt: "You are a helpful assistant"
      assert:
        - type: not-contains
          value: "I cannot"
      options:
        provider: openai:gpt-4o
    ```
    
    ---
    
    ## Assertion Types
    
    ### Deterministic
    
    | Type                            | Value    | Description                    |
    | ------------------------------- | -------- | ------------------------------ |
    | `equals`                        | string   | Exact match                    |
    | `contains`                      | string   | Substring match                |
    | `icontains`                     | string   | Case-insensitive substring     |
    | `not-contains`                  | string   | Must not contain               |
    | `not-equals`                    | string   | Must not equal                 |
    | `contains-any`                  | string[] | Contains at least one          |
    | `contains-all`                  | string[] | Contains all                   |
    | `icontains-any`                 | string[] | Case-insensitive, at least one |
    | `icontains-all`                 | string[] | Case-insensitive, all          |
    | `starts-with`                   | string   | Prefix match                   |
    | `regex`                         | string   | Regular expression match       |
    | `not-regex`                     | string   | Must not match regex           |
    | `is-json`                       | --       | Valid JSON                     |
    | `contains-json`                 | --       | Contains valid JSON            |
    | `is-html`                       | --       | Valid HTML                     |
    | `is-xml`                        | --       | Valid XML                      |
    | `is-sql`                        | --       | Valid SQL                      |
    | `is-valid-openai-tools-call`    | --       | Valid OpenAI tools call        |
    | `is-valid-openai-function-call` | --       | Valid OpenAI function call     |
    | `is-refusal`                    | --       | Model refused to answer        |
    
    ### Performance
    
    | Type         | Threshold    | Description          |
    | ------------ | ------------ | -------------------- |
    | `cost`       | number (USD) | Max cost per call    |
    | `latency`    | number (ms)  | Max response time    |
    | `perplexity` | number       | Max perplexity score |
    
    ### Text Similarity
    
    | Type          | Threshold          | Description                      |
    | ------------- | ------------------ | -------------------------------- |
    | `levenshtein` | number (max edits) | Edit distance                    |
    | `similar`     | number (0-1)       | Cosine similarity via embeddings |
    | `rouge-n`     | number (0-1)       | ROUGE-N overlap score            |
    | `bleu`        | number (0-1)       | BLEU translation score           |
    
    ### Model-Graded
    
    | Type                    | Value           | Description                       |
    | ----------------------- | --------------- | --------------------------------- |
    | `llm-rubric`            | criteria string | General LLM-as-judge              |
    | `factuality`            | ground truth    | Factual accuracy check            |
    | `model-graded-closedqa` | ground truth    | Closed-domain QA accuracy         |
    | `answer-relevance`      | --              | Answer relevance to question      |
    | `context-faithfulness`  | --              | RAG: answer faithful to context   |
    | `context-recall`        | --              | RAG: context covers ground truth  |
    | `context-relevance`     | --              | RAG: context relevant to question |
    | `classifier`            | criteria        | Classification evaluation         |
    | `select-best`           | --              | Pick best output across providers |
    
    ### Custom
    
    | Type         | Value             | Description                  |
    | ------------ | ----------------- | ---------------------------- |
    | `javascript` | code or `file://` | Custom JS assertion          |
    | `python`     | code or `file://` | Custom Python assertion      |
    | `webhook`    | URL               | External validation endpoint |
    
    ### Assertion Properties
    
    ```yaml
    assert:
      - type: llm-rubric
        value: "criteria" # Expected value / criteria
        threshold: 0.8 # Pass/fail threshold (0-1)
        weight: 2.0 # Scoring weight in results UI
        metric: "quality" # Label for UI aggregation
        provider: openai:gpt-4o # Grading model (model-graded only)
        transform: "output.trim()" # Pre-process output before assertion
    ```
    
    ---
    
    ## Built-in Provider IDs
    
    ### OpenAI
    
    ```
    openai:gpt-4o
    openai:gpt-4o-mini
    openai:o4-mini
    openai:gpt-4.1
    openai:gpt-4.1-mini
    ```
    
    ### Anthropic
    
    ```
    anthropic:messages:claude-sonnet-4-6
    anthropic:messages:claude-opus-4-6
    anthropic:messages:claude-sonnet-4-5-latest
    anthropic:messages:claude-3-5-haiku-latest
    ```
    
    ### Google
    
    ```
    vertex:gemini-2.0-flash
    vertex:gemini-2.5-pro
    ```
    
    ### Other
    
    ```
    ollama:llama3
    ollama:mistral
    huggingface:text-generation:MODEL_NAME
    ```
    
    ### Custom
    
    ```yaml
    # TypeScript/JavaScript file
    - file://providers/my-provider.ts
    
    # HTTP endpoint
    - id: https://api.example.com/chat
      config:
        method: POST
        headers:
          Authorization: "Bearer {{env.API_KEY}}"
        body:
          prompt: "{{prompt}}"
    ```
    
    ---
    
    ## Red Team Configuration
    
    ### Plugins (Attack Generators)
    
    | Plugin                   | Category       | Description                  |
    | ------------------------ | -------------- | ---------------------------- |
    | `harmful`                | Criminal       | Harmful content generation   |
    | `pii`                    | Privacy        | PII extraction attempts      |
    | `prompt-extraction`      | Security       | System prompt extraction     |
    | `hallucination`          | Misinformation | Hallucinated information     |
    | `contracts`              | Misinformation | Unauthorized commitments     |
    | `excessive-agency`       | Misinformation | Taking unauthorized actions  |
    | `competitor-endorsement` | Misinformation | Endorsing competitors        |
    | `hijacking`              | Security       | Topic/task hijacking         |
    | `overreliance`           | Misinformation | Over-reliance on user claims |
    | `shell-injection`        | Security       | OS command injection         |
    | `sql-injection`          | Security       | SQL injection attempts       |
    
    ### Strategies (Delivery Methods)
    
    | Strategy           | Description                |
    | ------------------ | -------------------------- |
    | `jailbreak`        | Jailbreak prompt templates |
    | `prompt-injection` | Direct prompt injection    |
    | `crescendo`        | Multi-turn escalation      |
    | `goat`             | Generative offensive agent |
    | `base64`           | Base64-encoded payloads    |
    | `rot13`            | ROT13-encoded payloads     |
    | `leetspeak`        | Leetspeak encoding         |
    | `iterative`        | Iterative refinement       |
    
    ### Presets
    
    ```yaml
    redteam:
      plugins:
        - owasp:llm:01 # OWASP LLM Top 10
        - nist:ai:measure # NIST AI RMF
    ```
    
    ---
    
    ## Environment Variables
    
    | Variable                    | Purpose                          |
    | --------------------------- | -------------------------------- |
    | `OPENAI_API_KEY`            | OpenAI provider API key          |
    | `ANTHROPIC_API_KEY`         | Anthropic provider API key       |
    | `GOOGLE_API_KEY`            | Google / Vertex AI API key       |
    | `PROMPTFOO_CONFIG`          | Override config file path        |
    | `PROMPTFOO_CACHE_PATH`      | Custom cache directory           |
    | `PROMPTFOO_DISABLE_CACHE`   | Disable caching (`1` to disable) |
    | `PROMPTFOO_SHARING_APP_URL` | Custom sharing server URL        |
    
    ---
    
    ## Nunjucks Template Syntax
    
    ```yaml
    # Variable substitution
    prompts:
      - "Translate to {{language}}: {{input}}"
    
      # Conditional
      - "{% if context %}Context: {{context}}{% endif %}\n{{question}}"
    
      # Loop
      - "{% for item in items %}{{item}}\n{% endfor %}"
    
      # Environment variable
      - "file://{{ env.PROMPT_DIR }}/prompt.txt"
    ```
    
    > For YAML references (reusable assertion blocks), see [examples/core.md](examples/core.md#yaml-references-reusable-assertions).
    
  • SKILL.md 16.8 KB
    ---
    name: ai-observability-promptfoo
    description: Testing and evaluation framework for LLM prompts and applications -- promptfooconfig.yaml, assertions, model-graded evals, red teaming, CI/CD integration, custom providers, and comparative evaluation
    ---
    
    # Promptfoo Patterns
    
    > **Quick Guide:** Use promptfoo for systematic LLM evaluation. Define prompts, providers, and test cases in `promptfooconfig.yaml`. Use assertion types (`contains`, `is-json`, `llm-rubric`, `similar`, `cost`, `latency`) to validate outputs. Use `promptfoo eval` to run (exits with code 100 on test failures), `promptfoo view` for results UI. Use model-graded assertions (`llm-rubric`, `factuality`) for subjective quality. Use `promptfoo redteam run` for security scanning. Use `--share` flag or `promptfoo share` to share results. All provider API keys come from environment variables -- never hardcode them.
    
    ---
    
    <critical_requirements>
    
    ## CRITICAL: Before Using This Skill
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST define test cases with explicit `assert` arrays -- tests without assertions only capture output without validating it)**
    
    **(You MUST use `llm-rubric` for subjective quality evaluation -- do NOT rely solely on deterministic assertions for natural language output)**
    
    **(You MUST set `threshold` on similarity and model-graded assertions -- omitting thresholds uses defaults that may not match your quality bar)**
    
    **(You MUST use environment variables for all API keys -- never hardcode keys in promptfooconfig.yaml or provider configs)**
    
    **(You MUST verify `promptfoo eval` exit code in CI pipelines -- it returns exit code 100 on test failures, exit code 1 on other errors)**
    
    </critical_requirements>
    
    ---
    
    **Auto-detection:** promptfoo, promptfooconfig, promptfooconfig.yaml, promptfoo eval, promptfoo view, promptfoo redteam, llm-rubric, model-graded-closedqa, promptfoo share, promptfoo cache, assertion type, LLM evaluation, prompt testing, red teaming, PROMPTFOO_CONFIG
    
    **When to use:**
    
    - Writing or evaluating LLM prompts across one or more providers
    - Setting up automated test suites for LLM-powered features
    - Comparing model outputs side-by-side (GPT vs Claude vs Gemini)
    - Running model-graded evaluations (LLM-as-a-judge)
    - Red teaming LLM applications for security vulnerabilities
    - Integrating LLM quality gates into CI/CD pipelines
    - Validating structured output (JSON, function calls) from LLMs
    
    **Key patterns covered:**
    
    - `promptfooconfig.yaml` structure (prompts, providers, tests, defaultTest)
    - Assertion types (deterministic, model-graded, performance)
    - Custom TypeScript providers
    - Red teaming configuration (plugins, strategies)
    - CI/CD integration with GitHub Actions
    - Programmatic API (`evaluate()` function)
    - Result sharing and caching
    
    **When NOT to use:**
    
    - Unit testing application code (use your test runner)
    - Load testing / benchmarking API throughput (use a load testing tool)
    - Runtime monitoring of production LLM calls (use observability tooling)
    
    ---
    
    ## Examples Index
    
    - [Core: Config & Assertions](examples/core.md) -- promptfooconfig.yaml structure, providers, prompts, test cases, assertion types
    - [Model-Graded & Advanced Assertions](examples/model-graded.md) -- llm-rubric, factuality, similar, context evaluation, custom assertions
    - [Red Teaming](examples/red-teaming.md) -- Security scanning, plugins, strategies, presets
    - [Custom Providers & Programmatic API](examples/custom-providers.md) -- TypeScript providers, evaluate() function, CI/CD integration
    - [Quick API Reference](reference.md) -- CLI commands, assertion type table, provider IDs, red team plugins
    
    ---
    
    <philosophy>
    
    ## Philosophy
    
    Promptfoo brings **test-driven development to LLM applications**. Instead of manually checking outputs, you define expected behaviors as assertions and run them systematically across prompts and providers.
    
    **Core principles:**
    
    1. **Declarative test definitions** -- YAML config over imperative test scripts. Define prompts, providers, test cases, and assertions in `promptfooconfig.yaml`. No code required for standard evaluations.
    2. **Assertion-driven validation** -- Every test case should have assertions. Deterministic assertions (`contains`, `is-json`, `equals`) for structured output; model-graded assertions (`llm-rubric`, `factuality`) for subjective quality.
    3. **Comparative evaluation** -- Run the same tests across multiple providers or prompt variants simultaneously. The results matrix shows which combination performs best.
    4. **Shift-left LLM testing** -- Catch prompt regressions in CI before they reach production. `promptfoo eval` exits with code 100 on test failures, making it a natural CI quality gate.
    5. **Red teaming as a first-class concern** -- Security scanning for prompt injection, PII leakage, harmful content, and jailbreak vulnerabilities is built in, not bolted on.
    
    </philosophy>
    
    ---
    
    <patterns>
    
    ## Core Patterns
    
    ### Pattern 1: Basic Configuration
    
    Every promptfoo project starts with `promptfooconfig.yaml`. Three required sections: `prompts`, `providers`, `tests`.
    
    ```yaml
    # promptfooconfig.yaml
    description: "Translation quality evaluation"
    
    prompts:
      - "Convert the following to {{language}}: {{input}}"
    
    providers:
      - openai:gpt-4o
      - anthropic:messages:claude-sonnet-4-6
    
    tests:
      - vars:
          language: French
          input: Hello world
        assert:
          - type: icontains
            value: "bonjour"
          - type: llm-rubric
            value: "Output is a natural French translation, not word-for-word"
    ```
    
    **Why good:** Declarative config, multi-provider comparison, both deterministic and model-graded assertions
    
    ```yaml
    # BAD: Tests without assertions
    tests:
      - vars:
          language: French
          input: Hello world
      # No assert array -- output is captured but never validated
    ```
    
    **Why bad:** Tests without assertions only log output, they never fail -- you lose the entire point of automated evaluation
    
    **See:** [examples/core.md](examples/core.md) for prompts from files, provider config, defaultTest, variable loading from CSV
    
    ---
    
    ### Pattern 2: Deterministic Assertions
    
    Use for outputs with predictable, verifiable structure.
    
    ```yaml
    assert:
      # String matching
      - type: contains
        value: "error"
      - type: icontains # case-insensitive
        value: "success"
      - type: not-contains
        value: "internal server error"
      - type: starts-with
        value: "{"
      - type: regex
        value: "\\d{4}-\\d{2}-\\d{2}" # date pattern
    
      # Structured output
      - type: is-json
      - type: contains-json
      - type: is-valid-openai-tools-call
    
      # Performance
      - type: cost
        threshold: 0.01 # max $0.01 per call
      - type: latency
        threshold: 5000 # max 5 seconds
    ```
    
    **Why good:** Fast, deterministic, no LLM cost for evaluation, catches structural regressions immediately
    
    ```yaml
    # BAD: Using llm-rubric for JSON validation
    assert:
      - type: llm-rubric
        value: "Output must be valid JSON"
    ```
    
    **Why bad:** Expensive (requires LLM call), slower, non-deterministic -- `is-json` does this deterministically for free
    
    **See:** [examples/core.md](examples/core.md) for all deterministic assertion types with examples
    
    ---
    
    ### Pattern 3: Model-Graded Assertions
    
    Use for subjective quality where deterministic checks cannot capture intent.
    
    ```yaml
    assert:
      - type: llm-rubric
        value: "Response is helpful, accurate, and conversational in tone"
        provider: openai:gpt-4o
    
      - type: factuality
        value: "The capital of France is Paris. It has a population of ~2.1 million."
        provider: openai:gpt-4o
    
      - type: similar
        value: "The weather in Paris is sunny today"
        threshold: 0.8
    
      - type: model-graded-closedqa
        value: "Paris is the capital of France"
        provider: openai:gpt-4o
    ```
    
    **Why good:** Evaluates subjective quality that deterministic assertions cannot capture, configurable grading provider
    
    ```yaml
    # BAD: No threshold on similar assertion
    assert:
      - type: similar
        value: "expected output"
        # Missing threshold -- uses default which may be too lenient or strict
    ```
    
    **Why bad:** Default similarity threshold may not match your quality bar, always set it explicitly
    
    **See:** [examples/model-graded.md](examples/model-graded.md) for llm-rubric with custom providers, context evaluation, factuality, custom grading prompts
    
    ---
    
    ### Pattern 4: Red Teaming
    
    Use `redteam` section to scan for security vulnerabilities.
    
    ```yaml
    # promptfooconfig.yaml
    targets:
      - openai:gpt-4o
    
    redteam:
      purpose: "Customer support chatbot for an e-commerce platform"
      numTests: 10
      plugins:
        - harmful
        - pii
        - contracts
        - hallucination
        - prompt-extraction
      strategies:
        - jailbreak
        - prompt-injection
    ```
    
    **Why good:** Declarative security scanning, purpose provides context for realistic attacks, composable plugins and strategies
    
    ```yaml
    # BAD: Red team without purpose
    redteam:
      plugins:
        - harmful
      # Missing purpose -- attacks will be generic and less effective
    ```
    
    **Why bad:** Without `purpose`, the red team generator creates generic attacks that miss application-specific vulnerabilities
    
    **See:** [examples/red-teaming.md](examples/red-teaming.md) for presets (OWASP, NIST), advanced strategies, multi-turn attacks
    
    ---
    
    ### Pattern 5: Custom TypeScript Provider
    
    Use when your LLM integration is not a direct API call (RAG pipelines, agent chains, custom middleware).
    
    ```typescript
    // providers/my-app.ts
    import type {
      ApiProvider,
      ProviderOptions,
      ProviderResponse,
      CallApiContextParams,
    } from "promptfoo";
    
    // NOTE: default export required by promptfoo's file:// provider loader
    export default class MyAppProvider implements ApiProvider {
      private config: Record<string, unknown>;
    
      constructor(options: ProviderOptions) {
        this.config = options.config || {};
      }
    
      id(): string {
        return "my-app-provider";
      }
    
      async callApi(
        prompt: string,
        context?: CallApiContextParams,
      ): Promise<ProviderResponse> {
        // Call your application's LLM pipeline
        const result = await myApp.processQuery(prompt);
    
        return {
          output: result.answer,
          tokenUsage: {
            total: result.totalTokens,
            prompt: result.promptTokens,
            completion: result.completionTokens,
          },
          cost: result.cost,
        };
      }
    }
    ```
    
    ```yaml
    # promptfooconfig.yaml
    providers:
      - file://providers/my-app.ts
    ```
    
    **Why good:** Type-safe, full control over LLM pipeline, reports token usage and cost for assertions
    
    **See:** [examples/custom-providers.md](examples/custom-providers.md) for inline function providers, programmatic API, CI/CD integration
    
    ---
    
    ### Pattern 6: CI/CD Integration
    
    Run evaluations in CI with quality gates.
    
    ```yaml
    # .github/workflows/llm-eval.yml
    name: LLM Eval
    on:
      pull_request:
        paths:
          - "prompts/**"
          - "promptfooconfig.yaml"
    jobs:
      evaluate:
        runs-on: ubuntu-latest
        steps:
          - uses: actions/checkout@v4
          - uses: actions/setup-node@v4
            with:
              node-version: "22"
          - uses: actions/cache@v4
            with:
              path: ~/.cache/promptfoo
              key: ${{ runner.os }}-promptfoo-v1
          - name: Run eval
            env:
              OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
            run: npx promptfoo@latest eval -o results.json --share
    ```
    
    **Why good:** Caches LLM responses across runs, `promptfoo eval` exits with code 100 on test failures (CI fails automatically), `--share` generates a shareable results URL
    
    **See:** [examples/custom-providers.md](examples/custom-providers.md) for npm scripts, quality gate thresholds, programmatic evaluation
    
    </patterns>
    
    ---
    
    <decision_framework>
    
    ## Decision Framework
    
    ### Which Assertion Type to Use
    
    ```
    What are you validating?
    +-- Exact or structural match?
    |   +-- Exact text -> equals
    |   +-- Contains substring -> contains / icontains
    |   +-- Regex pattern -> regex
    |   +-- Valid JSON -> is-json
    |   +-- Valid function call -> is-valid-openai-tools-call
    |   +-- Cost under budget -> cost (with threshold)
    |   +-- Response time -> latency (with threshold)
    +-- Subjective quality?
    |   +-- General quality criteria -> llm-rubric
    |   +-- Factual accuracy against ground truth -> factuality
    |   +-- Semantic similarity -> similar (with threshold)
    |   +-- Closed-domain QA accuracy -> model-graded-closedqa
    |   +-- RAG context fidelity -> context-faithfulness
    +-- Custom logic?
        +-- JavaScript function -> javascript
        +-- Python function -> python
        +-- External service -> webhook
    ```
    
    ### When to Use Red Teaming vs Eval
    
    ```
    What are you testing?
    +-- Prompt quality and correctness?
    |   +-- Use promptfoo eval with test cases and assertions
    +-- Security vulnerabilities?
    |   +-- Use promptfoo redteam run with plugins and strategies
    +-- Both?
        +-- Run eval for quality, redteam for security -- separate configs or sections
    ```
    
    ### Provider Selection
    
    ```
    How does your LLM integration work?
    +-- Direct API call to OpenAI/Anthropic/etc?
    |   +-- Use built-in provider: openai:gpt-4o, anthropic:messages:claude-sonnet-4-6
    +-- Custom pipeline (RAG, agents, middleware)?
    |   +-- Use custom TypeScript provider: file://providers/my-app.ts
    +-- HTTP endpoint?
    |   +-- Use HTTP provider: id: https://api.example.com/chat
    +-- Multiple providers to compare?
        +-- List all in providers array -- promptfoo runs tests against each
    ```
    
    </decision_framework>
    
    ---
    
    <red_flags>
    
    ## RED FLAGS
    
    **High Priority Issues:**
    
    - Tests without `assert` arrays (output is captured but never validated -- tests always "pass")
    - Not checking `promptfoo eval` exit code in CI (`promptfoo eval` exits 100 on test failures -- ensure your CI pipeline treats non-zero exit codes as failures)
    - Hardcoded API keys in `promptfooconfig.yaml` (use environment variables)
    - Using `llm-rubric` for checks that `is-json` or `contains` can do deterministically (wastes money and adds non-determinism)
    - Red teaming without `purpose` (generic attacks miss application-specific vulnerabilities)
    
    **Medium Priority Issues:**
    
    - Missing `threshold` on `similar` assertions (default may not match your quality bar)
    - Not caching in CI (every run makes full API calls -- expensive and slow)
    - Using `model-graded-closedqa` when `llm-rubric` would be simpler (closedqa is for specific ground-truth QA)
    - Not setting `provider` on model-graded assertions (uses default which may not be the grader you want)
    - Running red team with default `numTests: 5` in production scans (too few for comprehensive coverage)
    
    **Common Mistakes:**
    
    - Confusing `prompts` (the LLM prompt templates) with `tests` (the evaluation cases) -- prompts define what to send, tests define what to check
    - Using `equals` for natural language output (LLM output is non-deterministic, use `llm-rubric` or `similar`)
    - Forgetting `{{variable}}` syntax in prompts (promptfoo uses Nunjucks templating, not `${variable}`)
    - Putting assertions in `defaultTest` that should only apply to specific tests (assertions in `defaultTest` apply to ALL tests)
    - Using `file://` paths without the prefix (promptfoo treats bare paths as literal strings, not file references)
    
    **Gotchas & Edge Cases:**
    
    - `promptfoo eval` caches LLM responses by default -- use `promptfoo cache clear` or `--no-cache` to force fresh calls
    - `--share` uploads results to promptfoo's servers -- do not use with sensitive data unless self-hosting
    - Red team `strategies` wrap `plugins` output -- a plugin generates the malicious content, a strategy delivers it (e.g., via jailbreak encoding)
    - `defaultTest.assert` merges with per-test assertions, it does not replace them -- both arrays run
    - CSV test files map column headers to variable names -- header `input` becomes `{{input}}` in prompts
    - `transform` in test options runs JavaScript on the output before assertions -- useful for extracting JSON from markdown-wrapped responses
    - Provider configs in YAML use `config:` key for model parameters (`temperature`, `max_tokens`), not top-level fields
    - The `weight` property on assertions affects scoring in the results UI but does not change pass/fail behavior
    
    </red_flags>
    
    ---
    
    <critical_reminders>
    
    ## CRITICAL REMINDERS
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST define test cases with explicit `assert` arrays -- tests without assertions only capture output without validating it)**
    
    **(You MUST use `llm-rubric` for subjective quality evaluation -- do NOT rely solely on deterministic assertions for natural language output)**
    
    **(You MUST set `threshold` on similarity and model-graded assertions -- omitting thresholds uses defaults that may not match your quality bar)**
    
    **(You MUST use environment variables for all API keys -- never hardcode keys in promptfooconfig.yaml or provider configs)**
    
    **(You MUST verify `promptfoo eval` exit code in CI pipelines -- it returns exit code 100 on test failures, exit code 1 on other errors)**
    
    **Failure to follow these rules will produce untested, insecure, or falsely-passing LLM evaluation pipelines.**
    
    </critical_reminders>
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related