Claude Skill

agent-eval

Use when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual recall) or agent trajectories (tool correctness, completion), or picking an eval framework. NOT build

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies

#agents #ai #evals

Virus-scanned Reviewed automatically before listing.

Full trust report

Download ericrisco-rsc-harness-skills_agent-eval-953fef5.zip · 15 KB
Part of ericrisco/rsc-harness — 46 skills

Install

skills CLI npx skills add https://github.com/ericrisco/rsc-harness/tree/main/skills/agent-eval
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ericrisco-rsc-harness@llmmart
Git git clone https://github.com/ericrisco/rsc-harness.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ericrisco/rsc-harness collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Measure agent quality you can defend and gate on

Turn "the agent feels better" into a number you can put in a PR check. You own the eval dataset, the scorer mix, the LLM-as-judge calibration, and the block-on-regression CI gate — framework-neutral, provider-neutral.

Do NOT use — route instead

The ask Route to Why it is not this skill
Build the agent loop, tools, RAG plumbing building-agents It builds the system; you score it. They cross-link.
"Make the answers shorter / rewrite the prompt" prompt-engineering Evals say it is worse; that skill changes the words. You never edit the prompt.
pytest/jest on deterministic functions testing-py / testing-web Assert-equals on pure code, not stochastic outputs scored by a judge.
Dashboards / tracing of live production traffic observability Online monitoring; you are offline + pre-merge.
Red-team, jailbreak, prompt injection agent-safety Adversarial coverage, not quality measurement.
Per-token cost budgets and accounting cost-tracking You report cost-per-task as one metric; the discipline lives there.
A/B stats on product/funnel metrics ab-testing Web experiments, not offline model comparison on a fixed set.

The eval anatomy

Every framework instantiates the same five-stage pipeline. Learn it once; the tool is a detail.

dataset ──▶ runner ──▶ scorers ──▶ metrics ──▶ gate
(JSONL    (calls the   (det / judge  (aggregate +  (pass/fail
golden    system per   / human)      bootstrap CI) exit code)
set)      case)

DeepEval, Inspect AI, and promptfoo are all just opinionated wrappers around this. If you understand the stages you can switch tools without relearning the craft.

Build the dataset first

Build your own golden set — a public leaderboard number is not your number, because identical model weights swing SWE-Bench Verified by 10–20 points just by changing the harness. Measure your task on your data. The dataset is the asset; everything else is replaceable. Rules:

  • 50–200 hand-labeled cases per failure mode, not per total. Coverage of how the system fails beats raw volume. 80 real failure cases > 1000 generic ones.
  • Never synthetic-only. A set the model wrote will not surface the model's blind spots. Mine real traffic / tickets / transcripts and hand-label.
  • Version it in git as JSONL, a first-class reviewed asset — same as code. Diffs are reviewable; relabels are auditable.
  • Decontaminate. The eval set must not appear in training data or few-shot examples, or the score is a memorization artifact, not a capability.

Case schema — one JSON object per line:

{"id":"refund-001","input":"Where is my refund for order 4821?","expected":"States refunds take 5-7 business days and asks for nothing already on file","context":["policy: refunds 5-7 business days"],"meta":{"failure_mode":"hallucinated_policy","source":"ticket#4821"}}
{"id":"refund-002","input":"Cancel my subscription and refund this month","expected":"Cancels, refunds prorated amount, confirms no future charge","context":["policy: prorated refund on cancel"],"meta":{"failure_mode":"missed_tool_call","source":"ticket#5190"}}

failure_mode in meta is what lets you slice metrics by mode and find which kind of bug regressed — not just that the aggregate dropped.

Bad: "generate 1000 test questions with GPT and use those." Good: "80 real failure-mode cases pulled from support tickets, hand-labeled, tagged by failure mode."

Choose the scorer — the 60/30/10 mix

Reach for the cheapest scorer that correlates with human judgment. Default mix:

Share Scorer kind Use for Why
~60% Deterministic — exact match, regex, JSON-schema validation, latency threshold Anything with a checkable shape: format, required fields, a known string, a budget Free, instant, zero drift. Never spend a judge call on something a regex settles.
~30% LLM-as-judge — G-Eval, DAG, custom Python scorer Meaning: is this answer faithful, relevant, helpful Only where correctness is semantic. Costs money and can drift — so calibrate it.
~10% Human-in-the-loop Genuinely ambiguous cases the judge disagrees on The ground truth you calibrate the judge against.

One Scorer protocol, two implementations behind it — deterministic and judge are interchangeable to the runner:

from typing import Protocol

class Scorer(Protocol):
    name: str
    def score(self, case: dict, output: str) -> float: ...  # 0.0–1.0

class JsonSchemaScorer:
    name = "schema_valid"
    def score(self, case, output):  # deterministic, free, no drift
        import json
        try:
            json.loads(output)
            return 1.0
        except ValueError:
            return 0.0

class FaithfulnessJudge:
    name = "faithfulness"
    def __init__(self, judge_model): self.judge = judge_model
    def score(self, case, output):  # judge only where meaning matters
        return self.judge.rate(case["context"], output)  # see judge-design.md

LLM-as-judge you can trust

A score you do not trust is worse than no score: an uncalibrated judge gives false confidence, which is more dangerous than admitted ignorance. Each rule, with its why:

  • Judge model ≥ system under test. A weaker judge cannot reliably rank a stronger system — it scores noise.
  • The rubric must force a written rationale before the score. Rationale-first judging is what pushes judge–human agreement to ~85% — higher than two humans agree with each other. A bare number is a vibe with a decimal point.
  • Pairwise beats pointwise for stability. "Is A or B better?" is more reproducible than "rate A from 1–10," which inflates and clusters at 8–9.
  • Swap positions and average. Judges favor whichever answer came first; run A-then-B and B-then-A to cancel position bias.
  • Calibrate against human gold and report the agreement before you gate anything on the judge. Not a formality — this is the step that makes every number downstream defensible.

Bad judge prompt: "Rate this answer 1–10." → everything lands 8–9, useless. Good: "Compare answer A and answer B against the reference. First write one sentence on each per the rubric, then output the better label." → forces reasoning, gives a stable signal.

Full rubric templates (pointwise + pairwise), the position-swap harness, the calibration script (agreement / Cohen's kappa vs human gold), G-Eval vs DAG, and the judge bias catalog (length, position, self-preference) with mitigations live in references/judge-design.md.

Agent and RAG scorers

Score the path, not only the destination. Beyond exact/judge:

RAG (DeepEval / RAGAS names):

  • Faithfulness — does the answer only claim what the retrieved context supports? Catches hallucination.
  • Answer relevancy — does it actually address the question, or drift?
  • Contextual recall / precision — did retrieval fetch the right chunks, and not bury them in noise? Separates a retrieval bug from a generation bug.

Agent:

  • Tool correctness — right tool, right arguments, right order.
  • Task completion / goal accuracy — did it finish the job, not just produce plausible text.
  • Trajectory scoring — grade the sequence of steps. A correct final answer from a wrong path will fail differently next time; only trajectory scoring catches it.

The system side of these (how the loop and tools are built) is ../building-agents/SKILL.md; a common system-under-test is ../chatbot/SKILL.md.

The regression gate

Gate policy: block on regression vs a committed baseline, not on an absolute threshold. An absolute threshold flaps CI on judge noise and gives no signal on drift; "did this PR make a tracked metric worse than main?" is the question that matters.

  • Compute a bootstrap confidence interval on each metric so judge noise alone does not fail the build — only a drop beyond the CI counts.
  • The runner writes eval-report.json (metrics, per-failure-mode slices, baseline, pass/fail) and exits non-zero on a real regression so the merge is blocked.
import json, sys

def gate(current: dict, baseline: dict, margin: float = 0.0) -> int:
    regressed = []
    for metric, score in current.items():
        if metric in baseline and score < baseline[metric] - margin:
            regressed.append((metric, baseline[metric], score))
    report = {"metrics": current, "baseline": baseline, "regressed": regressed,
              "passed": not regressed}
    with open("eval-report.json", "w") as f:
        json.dump(report, f, indent=2)
    if regressed:
        for m, b, c in regressed:
            print(f"REGRESSION {m}: {b:.3f} -> {c:.3f}", file=sys.stderr)
        return 1
    return 0

sys.exit(gate(run_eval(), json.load(open("eval-baseline.json"))))

The complete provider-neutral runner (JSONL loader, scorer registry, bootstrap-CI metrics), the GitHub Actions workflow, and side-by-side DeepEval-pytest + Inspect-AI Task/Solver/Scorer versions of the same eval live in references/runner-and-gate.md.

Framework cheat-sheet

Pick by where the eval runs and what it must do. Versions as of 2026-06 — re-verify, they rot.

Tool What it is Reach for it when
DeepEval v4.0.3 pytest-native, 50+ metrics, Decision-Graph (DAG) logic Your CI is Python/pytest and you want metrics that read like tests.
Inspect AI v0.3.225 (UK AISI) dataset→Task→Solver→Scorer, bootstrap CIs, first-class tool-use & trajectory logging, 200+ pre-built evals Multi-provider, safety-adjacent, or you need real trajectory scoring.
promptfoo (acquired by OpenAI 2026-03) CLI + YAML, strong pre-deploy + red-team across 50+ vuln types Config-driven pre-deploy checks; route the red-team half to agent-safety.
Braintrust / LangSmith / Phoenix v16.0.0 platforms: annotation, regression tracking, dashboards You need human annotation queues and historical regression tracking.

The two-tool pattern is normal, not over-engineering: a light CI gate (DeepEval / RAGAS / promptfoo) plus a platform (Braintrust / LangSmith / Arize) for annotation and history. They share data; different jobs.

Anti-patterns

Anti-pattern Why it bites Do instead
Vibes-gating ("feels better, merge it") No artifact to defend or reproduce Gate on a number from a committed dataset
Synthetic-only dataset Model-written cases miss the model's blind spots Hand-label real traffic by failure mode
Uncalibrated judge Confident wrong scores; worse than none Report agreement vs human gold first
Judge weaker than system Cannot rank a stronger system; scores noise Judge model ≥ system under test
Absolute-threshold gate Flaps CI on judge noise, blind to drift Block on regression vs baseline + bootstrap CI
Shipping on a leaderboard number Harness effect = 10–20pt swing Build your own golden set
Scoring only the final answer A right answer from a wrong path regresses later Score the trajectory too
Never relabeling drifted gold Stale "truth" silently rots the gate Review and relabel the golden set on a schedule

Project grounding

If the workspace has a 02-DOCS/ harness, record the eval policy in 02-DOCS/wiki/stack/evals.md: dataset location, scorer mix, gate baseline file, judge model, and the failure modes covered. Follow the harness wiki-article-template.md (type: stack) and index it in 02-DOCS/wiki/index.md. This is recorded, not gated — skip silently if there is no harness.

verify.sh

scripts/verify.sh is read-only and tool-detecting. It validates that every *.jsonl golden set in the project parses and that each line carries the required id, input, expected keys; checks the shape of any eval-report.json; and runs ruff / mypy on example Python and markdownlint on docs when those tools are installed. Every missing tool prints a yellow WARN and is skipped — never a failure. An empty or clean target exits 0.

Files (rsc-harness)
  • evals
    • cases.yaml 3.3 KB
      skill: agent-eval
      
      should_trigger:
        - prompt: "Is the new system prompt actually better than the old one, or am I just fooling myself?"
          why: "Before/after on a fixed dataset is the core job. Non-obvious: uses no eval vocabulary at all."
        - prompt: "Build me a golden set and a CI check that fails the PR if our support bot's answer quality drops."
          why: "Dataset construction plus a block-on-regression gate, stated explicitly."
        - prompt: "Our LLM judge gives every answer a 9 out of 10 — the scores are useless."
          why: "Judge calibration / inflation; the fix is rationale-forcing rubric and pairwise over pointwise. Non-obvious symptom phrasing, no mention of 'eval'."
        - prompt: "Score our RAG pipeline for faithfulness and whether it actually uses the retrieved context."
          why: "RAG metrics — faithfulness and contextual recall — are squarely this skill."
        - prompt: "I need to know if the agent took the right tool path, not just whether the final answer looks ok."
          why: "Trajectory and tool-correctness scoring. Non-obvious: framed as behavior, not as 'eval'."
        - prompt: "Necesito medir si el agente mejoró con el modelo nuevo antes de hacer merge."
          why: "Spanish — before/after measurement plus a pre-merge gate."
        - prompt: "Should we use DeepEval or Inspect AI for our eval suite, and how do we wire the gate?"
          why: "Framework choice plus gate wiring; the cheat-sheet and runner are the answer."
      
      should_not_trigger:
        - prompt: "Rewrite this system prompt so the answers come out shorter and punchier."
          route_to: prompt-engineering
          why: "Changing the words, not measuring quality. Evals say it is worse; that skill changes it."
        - prompt: "Build the agent loop and tool registry for our new assistant."
          route_to: building-agents
          why: "Building the system under test, not scoring it."
        - prompt: "Add a dashboard showing token cost and latency of our live LLM traffic."
          route_to: observability
          why: "Online monitoring of production, not offline pre-merge evaluation."
        - prompt: "Write pytest unit tests for our data-cleaning functions."
          route_to: testing-py
          why: "Deterministic code with assert-equals, not stochastic outputs scored by a judge."
        - prompt: "Red-team our chatbot for jailbreaks and prompt injection."
          route_to: agent-safety
          why: "Adversarial safety coverage, not quality measurement on a golden set."
      
      capability:
        - scenario: "Set up an eval plus a regression gate for a RAG support bot that is regressing silently between releases."
          must_include:
            - "Builds a versioned JSONL golden set of 50-200 hand-labeled cases per failure mode, decontaminated, not synthetic-only"
            - "Picks a scorer mix: deterministic first (schema/regex/latency), faithfulness and answer-relevancy judge where meaning matters"
            - "Calibrates the LLM-as-judge against human labels and reports agreement (raw and/or Cohen's kappa) before trusting it"
            - "Gate policy is block-on-regression vs a committed baseline with a bootstrap CI, not an absolute threshold"
            - "Runner emits eval-report.json and exits non-zero to block the merge"
            - "Names at least one concrete framework (DeepEval, Inspect AI, or promptfoo) and fits it to the case"
            - "Records the eval policy in 02-DOCS/wiki/stack/evals.md if a harness exists — recorded, not gated"
      
    • README.md 847 B
      # Evals for agent-eval
      
      `cases.yaml` is the trigger and capability spec for this skill. It is read by the catalog's
      skill-eval harness, not by a standalone runner here. `should_trigger` lists prompts the skill
      must claim (including non-obvious and Spanish phrasings); `should_not_trigger` lists prompts
      that must route to a named sibling instead; `capability` is a rubric a graded run must satisfy.
      
      To check by hand: read each `should_trigger` prompt and confirm the description in `SKILL.md`
      would plausibly fire on it; read each `should_not_trigger` prompt and confirm the `route_to`
      sibling is the better home and is a real catalog id. For the `capability` case, draft the
      skill's answer and confirm every `must_include` bullet is covered. If your harness scores these
      automatically, point it at this file with `skill: agent-eval` as the key.
      
  • references
    • judge-design.md 5.1 KB
      # LLM-as-judge design and calibration
      
      Depth offloaded from `SKILL.md`. A judge is a scorer, so the one rule binds hardest here:
      **calibrate against human gold and report agreement before you trust a single judge score.**
      Judge–human agreement reaches ~85% in 2026 (higher than two humans on the same task) — but
      only with a capable judge model and a rubric that forces a written rationale.
      
      ## Pointwise rubric template
      
      Pointwise (rate one answer) is convenient but inflates and clusters at 8–9. Use it only when
      you cannot run pairwise, and always force the rationale first.
      
      ```text
      You are grading an answer against a reference. The dimension is FAITHFULNESS:
      every claim in the answer must be supported by the provided context.
      
      Context:
      {context}
      
      Answer:
      {answer}
      
      Steps (do them in order):
      1. List each factual claim in the answer.
      2. For each claim, mark SUPPORTED or UNSUPPORTED against the context.
      3. Only then output a JSON object: {"rationale": "...", "score": <0.0-1.0>}
         where score = supported_claims / total_claims.
      ```
      
      The order matters: a rubric that asks for the number first gets a vibe; asking for the
      claim-by-claim breakdown first forces the reasoning that earns the ~85% agreement.
      
      ## Pairwise rubric template (preferred)
      
      "Is A or B better?" is more reproducible than an absolute rating. Use it for before/after and
      A/B comparisons.
      
      ```text
      Compare two answers, A and B, against the reference for HELPFULNESS.
      
      Reference: {reference}
      Answer A: {a}
      Answer B: {b}
      
      1. One sentence on what A does well/badly vs the reference.
      2. One sentence on what B does well/badly vs the reference.
      3. Output JSON: {"rationale": "...", "winner": "A" | "B" | "tie"}
      ```
      
      ## Position-swap harness (kills position bias)
      
      Judges favor whichever answer appears first. Run each comparison twice with positions swapped
      and only count a clear win as a win.
      
      ```python
      def pairwise(judge, reference, a, b):
          """Return 'a', 'b', or 'tie' after canceling position bias."""
          first  = judge.compare(reference, a, b)   # A in slot 1
          second = judge.compare(reference, b, a)   # A in slot 2 -> remap
          remap = {"A": "b", "B": "a", "tie": "tie"}
          r1 = {"A": "a", "B": "b", "tie": "tie"}[first]
          r2 = remap[second]
          if r1 == r2:
              return r1                # consistent across positions -> trust it
          return "tie"                 # judge flipped with position -> not a real signal
      ```
      
      A judge that flips when you swap positions has told you it cannot separate the two — record a
      tie, do not pick the first-slot winner.
      
      ## Calibration script (agreement and Cohen's kappa vs human gold)
      
      You need a small human-labeled gold slice (~10% of cases, the ambiguous ones). Run the judge
      on it and measure how often it agrees with the humans. Report this number in `eval-report.json`
      and refuse to gate if it is low.
      
      ```python
      def cohen_kappa(human: list[int], judge: list[int]) -> float:
          """Agreement corrected for chance. 1.0 perfect, 0 chance-level."""
          n = len(human)
          po = sum(h == j for h, j in zip(human, judge)) / n        # raw agreement
          labels = set(human) | set(judge)
          pe = sum((human.count(k) / n) * (judge.count(k) / n) for k in labels)
          return (po - pe) / (1 - pe) if pe != 1 else 1.0
      
      def calibration_report(human, judge):
          raw = sum(h == j for h, j in zip(human, judge)) / len(human)
          return {"raw_agreement": round(raw, 3),
                  "cohen_kappa": round(cohen_kappa(human, judge), 3)}
      ```
      
      Rules of thumb for the kappa: < 0.4 the judge is unusable, fix the rubric or model; 0.4–0.6
      marginal, widen the human slice; > 0.6 with raw agreement ~0.85 is the working zone. These are
      thresholds for *trusting the judge*, not for gating the system.
      
      ## G-Eval vs DAG
      
      - **G-Eval** — you give the judge an evaluation-criteria sentence; it generates the
        chain-of-thought steps and a weighted score. Fast to author, good for fuzzy semantic
        dimensions (coherence, helpfulness). Less reproducible for hard rules.
      - **DAG (Decision Graph)** — you author an explicit decision tree of yes/no checks; the score
        is deterministic given the answers. Use for compliance-style "must / must-not" criteria where
        you want auditability over flexibility. DeepEval v4 ships this as first-class.
      
      Pick G-Eval for taste, DAG for rules.
      
      ## Judge bias catalog and mitigations
      
      | Bias | Symptom | Mitigation |
      | --- | --- | --- |
      | Position | Favors the first answer shown | Position-swap harness above; count only consistent wins |
      | Length | Scores longer answers higher regardless of quality | Add "ignore length; reward density" to the rubric; spot-check long losers |
      | Self-preference | A model prefers text in its own style | Use a different model family as judge than the system under test |
      | Verbosity-of-rationale | Long rationale read as more correct | Score the claim breakdown, not the prose |
      | Anchoring on reference wording | Penalizes correct paraphrases | Grade meaning vs reference, not surface overlap; test with known-good paraphrases |
      
      Re-run the calibration slice whenever you change the judge model or rubric — a judge swap is a
      scorer change, and an uncalibrated scorer is back to a number you cannot defend.
      
    • runner-and-gate.md 5.8 KB
      # The runner, the gate, and framework equivalents
      
      Depth offloaded from `SKILL.md`. A provider-neutral runner you can drop into any repo, the CI
      wiring, and the same eval expressed in DeepEval and Inspect AI so you can switch tools without
      relearning the craft.
      
      ## The five stages, in code
      
      `dataset → runner → scorers → metrics → gate`. Keep them separable so each can be swapped.
      
      ### Dataset loader (JSONL, decontaminated, versioned)
      
      ```python
      import json
      from pathlib import Path
      
      REQUIRED = {"id", "input", "expected"}
      
      def load_cases(path: str) -> list[dict]:
          cases = []
          for i, line in enumerate(Path(path).read_text().splitlines(), 1):
              line = line.strip()
              if not line:
                  continue
              case = json.loads(line)
              missing = REQUIRED - case.keys()
              if missing:
                  raise ValueError(f"{path}:{i} missing keys {missing}")
              cases.append(case)
          return cases
      ```
      
      ### Scorer registry
      
      Every scorer satisfies one protocol — deterministic and judge are interchangeable. See
      `judge-design.md` for the judge implementations.
      
      ```python
      from dataclasses import dataclass
      from typing import Callable
      
      @dataclass
      class Scorer:
          name: str
          fn: Callable[[dict, str], float]   # (case, output) -> 0.0..1.0
      
      def exact_match(case, output):
          return 1.0 if output.strip() == case["expected"].strip() else 0.0
      
      def latency_under(threshold_s):
          def _fn(case, output):
              return 1.0 if case["meta"].get("latency_s", 0) <= threshold_s else 0.0
          return _fn
      
      REGISTRY = [
          Scorer("exact_match", exact_match),
          Scorer("latency_2s", latency_under(2.0)),
          # Scorer("faithfulness", FaithfulnessJudge(judge_model).score),  # 30% judge
      ]
      ```
      
      ### Runner + bootstrap-CI metrics
      
      The bootstrap CI is what stops judge noise from flapping the gate: only a drop beyond the
      interval counts as a real regression.
      
      ```python
      import random, statistics
      
      def run(system, cases, scorers):
          rows = []
          for case in cases:
              output = system(case["input"])               # the system under test
              row = {"id": case["id"],
                     "failure_mode": case["meta"].get("failure_mode", "none")}
              for s in scorers:
                  row[s.name] = s.fn(case, output)
              rows.append(row)
          return rows
      
      def metrics(rows, scorers, n_boot=1000):
          out = {}
          for s in scorers:
              vals = [r[s.name] for r in rows]
              mean = statistics.fmean(vals)
              boots = [statistics.fmean(random.choices(vals, k=len(vals)))
                       for _ in range(n_boot)]
              boots.sort()
              out[s.name] = {"mean": round(mean, 4),
                             "ci_low": round(boots[int(0.025 * n_boot)], 4),
                             "ci_high": round(boots[int(0.975 * n_boot)], 4)}
          return out
      ```
      
      ### Gate (block on regression vs committed baseline)
      
      ```python
      import json, sys
      
      def gate(current: dict, baseline: dict, report_path="eval-report.json") -> int:
          regressed = []
          for metric, m in current.items():
              base = baseline.get(metric, {}).get("mean")
              # Real regression: the current CI is entirely below the baseline mean.
              if base is not None and m["ci_high"] < base:
                  regressed.append({"metric": metric, "baseline": base,
                                    "now": m["mean"], "ci_high": m["ci_high"]})
          report = {"metrics": current, "baseline": baseline,
                    "regressed": regressed, "passed": not regressed}
          Path(report_path).write_text(json.dumps(report, indent=2))
          for r in regressed:
              print(f"REGRESSION {r['metric']}: {r['baseline']} -> {r['now']}",
                    file=sys.stderr)
          return 1 if regressed else 0
      ```
      
      Update the baseline deliberately — commit a new `eval-baseline.json` in the PR that *intends*
      to move the metric, so the move is reviewed, not silent.
      
      ## GitHub Actions wiring
      
      ```yaml
      name: eval-gate
      on: [pull_request]
      jobs:
        eval:
          runs-on: ubuntu-latest
          steps:
            - uses: actions/checkout@v4
            - uses: actions/setup-python@v5
              with: { python-version: "3.12" }
            - run: pip install -r evals/requirements.txt
            - name: Run eval gate
              env:
                OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
              run: python evals/run.py          # exits non-zero on regression -> blocks merge
            - uses: actions/upload-artifact@v4
              if: always()
              with: { name: eval-report, path: eval-report.json }
      ```
      
      The non-zero exit is the gate. Upload the report `if: always()` so a failed run still shows
      *which* metric and failure mode regressed.
      
      ## Same eval in DeepEval (pytest-native)
      
      ```python
      from deepeval import assert_test
      from deepeval.test_case import LLMTestCase
      from deepeval.metrics import FaithfulnessMetric
      
      def test_faithfulness():
          case = LLMTestCase(
              input="Where is my refund for order 4821?",
              actual_output=system("Where is my refund for order 4821?"),
              retrieval_context=["policy: refunds 5-7 business days"],
          )
          assert_test(case, [FaithfulnessMetric(threshold=0.8)])
      ```
      
      DeepEval v4.0.3 reads like tests and runs under pytest; reach for it when CI is already Python.
      
      ## Same eval in Inspect AI (Task / Solver / Scorer)
      
      ```python
      from inspect_ai import Task, task
      from inspect_ai.dataset import json_dataset
      from inspect_ai.solver import generate
      from inspect_ai.scorer import model_graded_qa
      
      @task
      def refund_quality():
          return Task(
              dataset=json_dataset("evals/cases.jsonl"),
              solver=generate(),
              scorer=model_graded_qa(),   # judge with rationale; bootstrap CIs built in
          )
      # inspect eval refund_quality.py --model openai/gpt-... -> pass/fail + CIs
      ```
      
      Inspect AI v0.3.225 (UK AISI) gives first-class tool-use and trajectory logging plus bootstrap
      CIs out of the box; reach for it when you are multi-provider or need real trajectory scoring.
      Both express the identical `dataset → scorer → gate` pipeline — the tool is a detail.
      
  • scripts
    • verify.sh 5.8 KB
      #!/usr/bin/env bash
      set -euo pipefail
      
      # verify.sh — agent-eval skill gate. Run from your PROJECT root.
      #
      # What it does (read-only, idempotent, NEVER calls an LLM or network):
      #   1. Discovers *.jsonl golden sets (skips vendor dirs); checks every line parses
      #      and carries the required keys id, input, expected. Hard fail on a bad line.
      #   2. If an eval-report.json exists, validates its shape (a "passed" boolean and a
      #      "metrics" object). Hard fail only on malformed JSON / missing shape.
      #   3. Optional lint, advisory only: ruff + mypy on example Python, markdownlint on docs.
      #
      # Exit non-zero ONLY on a malformed JSONL line, a missing required key, or a malformed
      # eval-report.json. Every optional lint and every missing tool is a yellow WARN/skip.
      # An empty or clean target exits 0. Stock macOS bash 3.2 (no mapfile / associative arrays).
      
      YELLOW=$'\033[33m'; GREEN=$'\033[32m'; RED=$'\033[31m'; NC=$'\033[0m'
      EXIT=0
      skip() { printf '%s[skip]%s %s\n' "$YELLOW" "$NC" "$*"; }
      note() { printf '%s[warn]%s %s\n' "$YELLOW" "$NC" "$*"; }
      ok()   { printf '%s[ok]%s %s\n'   "$GREEN"  "$NC" "$*"; }
      err()  { printf '%s[fail]%s %s\n' "$RED"    "$NC" "$*"; EXIT=1; }
      
      ROOT="$(pwd)"
      
      # Need a JSON-aware tool for the structural checks. Prefer python3, fall back to jq.
      JSON_TOOL=""
      if command -v python3 >/dev/null 2>&1; then JSON_TOOL="python3"
      elif command -v jq >/dev/null 2>&1; then JSON_TOOL="jq"; fi
      
      # --- 1. golden-set JSONL validation -----------------------------------------
      JSONL_FILES=()
      while IFS= read -r -d '' f; do
        JSONL_FILES+=("$f")
      done < <(
        find "$ROOT" \
          \( -path '*/node_modules/*' -o -path '*/.git/*' -o -path '*/vendor/*' \
             -o -path '*/.venv/*' -o -path '*/dist/*' \) -prune -o \
          -type f -name '*.jsonl' -print0 2>/dev/null
      )
      
      if [ "${#JSONL_FILES[@]}" -eq 0 ]; then
        skip "no *.jsonl golden sets found under $ROOT — nothing to validate"
      elif [ -z "$JSON_TOOL" ]; then
        skip "neither python3 nor jq found — cannot validate JSONL, skipping"
      else
        for f in "${JSONL_FILES[@]}"; do
          if [ "$JSON_TOOL" = "python3" ]; then
            msg="$(python3 - "$f" <<'PY'
      import json, sys
      path = sys.argv[1]
      req = {"id", "input", "expected"}
      bad = []
      with open(path, encoding="utf-8") as fh:
          for n, line in enumerate(fh, 1):
              s = line.strip()
              if not s:
                  continue
              try:
                  obj = json.loads(s)
              except ValueError as e:
                  bad.append(f"line {n}: invalid JSON ({e})"); continue
              if not isinstance(obj, dict):
                  bad.append(f"line {n}: not a JSON object"); continue
              missing = req - obj.keys()
              if missing:
                  bad.append(f"line {n}: missing keys {sorted(missing)}")
      print("\n".join(bad), end="")
      PY
      )" || msg="parser crashed on $f"
            if [ -n "$msg" ]; then
              err "$f:"; printf '  %s\n' "$msg"
            else
              ok "$f valid (id/input/expected present on every line)"
            fi
          else
            # jq path: each line must be an object with the three keys.
            if jq -e 'has("id") and has("input") and has("expected")' "$f" >/dev/null 2>&1; then
              ok "$f valid (jq: required keys present)"
            else
              err "$f: a line is not an object or is missing id/input/expected (jq)"
            fi
          fi
        done
      fi
      
      # --- 2. eval-report.json shape ------------------------------------------------
      REPORTS=()
      while IFS= read -r -d '' f; do
        REPORTS+=("$f")
      done < <(
        find "$ROOT" \
          \( -path '*/node_modules/*' -o -path '*/.git/*' -o -path '*/.venv/*' \) -prune -o \
          -type f -name 'eval-report.json' -print0 2>/dev/null
      )
      
      if [ "${#REPORTS[@]}" -eq 0 ]; then
        skip "no eval-report.json found — gate output not validated"
      elif [ -z "$JSON_TOOL" ]; then
        skip "no python3/jq — cannot validate eval-report.json shape"
      else
        for f in "${REPORTS[@]}"; do
          if [ "$JSON_TOOL" = "python3" ]; then
            if python3 - "$f" <<'PY'
      import json, sys
      try:
          r = json.load(open(sys.argv[1], encoding="utf-8"))
      except ValueError as e:
          print(e); sys.exit(1)
      ok = isinstance(r, dict) and isinstance(r.get("passed"), bool) \
           and isinstance(r.get("metrics"), dict)
      sys.exit(0 if ok else 1)
      PY
            then ok "$f shape ok (passed: bool, metrics: object)"
            else err "$f: malformed or missing 'passed'(bool)/'metrics'(object)"; fi
          else
            if jq -e 'type=="object" and (.passed|type=="boolean") and (.metrics|type=="object")' "$f" >/dev/null 2>&1; then
              ok "$f shape ok (jq)"
            else err "$f: malformed or missing passed/metrics (jq)"; fi
          fi
        done
      fi
      
      # --- 3. optional lint (advisory only) ----------------------------------------
      PY_FILES=()
      while IFS= read -r -d '' f; do
        PY_FILES+=("$f")
      done < <(
        find "$ROOT" \
          \( -path '*/node_modules/*' -o -path '*/.git/*' -o -path '*/.venv/*' \) -prune -o \
          -type f -name '*.py' -print0 2>/dev/null
      )
      
      if [ "${#PY_FILES[@]}" -gt 0 ]; then
        if command -v ruff >/dev/null 2>&1; then
          if ruff check "${PY_FILES[@]}" >/dev/null 2>&1; then ok "ruff clean"; else note "ruff reported lint findings (advisory)"; fi
        else skip "ruff not installed — skipping Python lint"; fi
        if command -v mypy >/dev/null 2>&1; then
          mypy "${PY_FILES[@]}" >/dev/null 2>&1 && ok "mypy clean" || note "mypy reported type findings (advisory)"
        else skip "mypy not installed — skipping type check"; fi
      else
        skip "no *.py files — skipping ruff/mypy"
      fi
      
      MD_FILES=()
      while IFS= read -r -d '' f; do
        MD_FILES+=("$f")
      done < <(
        find "$ROOT" \
          \( -path '*/node_modules/*' -o -path '*/.git/*' \) -prune -o \
          -type f -name '*.md' -print0 2>/dev/null
      )
      if [ "${#MD_FILES[@]}" -gt 0 ] && command -v markdownlint >/dev/null 2>&1; then
        markdownlint "${MD_FILES[@]}" >/dev/null 2>&1 && ok "markdownlint clean" || note "markdownlint findings (advisory)"
      else
        skip "markdownlint not installed or no docs — skipping doc lint"
      fi
      
      printf '\n'
      if [ "$EXIT" -eq 0 ]; then ok "verify.sh passed"; else err "verify.sh found failures"; fi
      exit "$EXIT"
      
  • SKILL.md 12.6 KB
    ---
    name: agent-eval
    description: "Use when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual recall) or agent trajectories (tool correctness, completion), or picking an eval framework. NOT building the agent loop, tools or RAG plumbing (that is `building-agents`)."
    tags: [evals, llm, agents, llm-as-judge, regression-gate, ai]
    recommends: [building-agents, prompt-engineering, observability]
    origin: risco
    ---
    
    # Measure agent quality you can defend and gate on
    
    Turn "the agent feels better" into a number you can put in a PR check. You own the eval dataset, the scorer mix, the LLM-as-judge calibration, and the block-on-regression CI gate — framework-neutral, provider-neutral.
    
    ## Do NOT use — route instead
    
    | The ask | Route to | Why it is not this skill |
    | --- | --- | --- |
    | Build the agent loop, tools, RAG plumbing | `building-agents` | It builds the system; you score it. They cross-link. |
    | "Make the answers shorter / rewrite the prompt" | `prompt-engineering` | Evals say it is worse; that skill changes the words. You never edit the prompt. |
    | pytest/jest on deterministic functions | `testing-py` / `testing-web` | Assert-equals on pure code, not stochastic outputs scored by a judge. |
    | Dashboards / tracing of live production traffic | `observability` | Online monitoring; you are offline + pre-merge. |
    | Red-team, jailbreak, prompt injection | `agent-safety` | Adversarial coverage, not quality measurement. |
    | Per-token cost budgets and accounting | `cost-tracking` | You report cost-per-task as one metric; the discipline lives there. |
    | A/B stats on product/funnel metrics | `ab-testing` | Web experiments, not offline model comparison on a fixed set. |
    
    ## The eval anatomy
    
    Every framework instantiates the same five-stage pipeline. Learn it once; the tool is a detail.
    
    ```text
    dataset ──▶ runner ──▶ scorers ──▶ metrics ──▶ gate
    (JSONL    (calls the   (det / judge  (aggregate +  (pass/fail
    golden    system per   / human)      bootstrap CI) exit code)
    set)      case)
    ```
    
    DeepEval, Inspect AI, and promptfoo are all just opinionated wrappers around this. If you understand the stages you can switch tools without relearning the craft.
    
    ## Build the dataset first
    
    Build your own golden set — a public leaderboard number is not your number, because identical model weights swing SWE-Bench Verified by 10–20 points just by changing the harness. Measure *your* task on *your* data. The dataset is the asset; everything else is replaceable. Rules:
    
    - **50–200 hand-labeled cases per failure mode**, not per total. Coverage of how the system fails beats raw volume. 80 real failure cases > 1000 generic ones.
    - **Never synthetic-only.** A set the model wrote will not surface the model's blind spots. Mine real traffic / tickets / transcripts and hand-label.
    - **Version it in git as JSONL**, a first-class reviewed asset — same as code. Diffs are reviewable; relabels are auditable.
    - **Decontaminate.** The eval set must not appear in training data or few-shot examples, or the score is a memorization artifact, not a capability.
    
    Case schema — one JSON object per line:
    
    ```jsonl
    {"id":"refund-001","input":"Where is my refund for order 4821?","expected":"States refunds take 5-7 business days and asks for nothing already on file","context":["policy: refunds 5-7 business days"],"meta":{"failure_mode":"hallucinated_policy","source":"ticket#4821"}}
    {"id":"refund-002","input":"Cancel my subscription and refund this month","expected":"Cancels, refunds prorated amount, confirms no future charge","context":["policy: prorated refund on cancel"],"meta":{"failure_mode":"missed_tool_call","source":"ticket#5190"}}
    ```
    
    `failure_mode` in `meta` is what lets you slice metrics by mode and find *which* kind of bug regressed — not just that the aggregate dropped.
    
    > Bad: "generate 1000 test questions with GPT and use those." Good: "80 real failure-mode cases pulled from support tickets, hand-labeled, tagged by failure mode."
    
    ## Choose the scorer — the 60/30/10 mix
    
    Reach for the cheapest scorer that correlates with human judgment. Default mix:
    
    | Share | Scorer kind | Use for | Why |
    | --- | --- | --- | --- |
    | ~60% | Deterministic — exact match, regex, JSON-schema validation, latency threshold | Anything with a checkable shape: format, required fields, a known string, a budget | Free, instant, zero drift. Never spend a judge call on something a regex settles. |
    | ~30% | LLM-as-judge — G-Eval, DAG, custom Python scorer | Meaning: is this answer faithful, relevant, helpful | Only where correctness is semantic. Costs money and can drift — so calibrate it. |
    | ~10% | Human-in-the-loop | Genuinely ambiguous cases the judge disagrees on | The ground truth you calibrate the judge against. |
    
    One `Scorer` protocol, two implementations behind it — deterministic and judge are interchangeable to the runner:
    
    ```python
    from typing import Protocol
    
    class Scorer(Protocol):
        name: str
        def score(self, case: dict, output: str) -> float: ...  # 0.0–1.0
    
    class JsonSchemaScorer:
        name = "schema_valid"
        def score(self, case, output):  # deterministic, free, no drift
            import json
            try:
                json.loads(output)
                return 1.0
            except ValueError:
                return 0.0
    
    class FaithfulnessJudge:
        name = "faithfulness"
        def __init__(self, judge_model): self.judge = judge_model
        def score(self, case, output):  # judge only where meaning matters
            return self.judge.rate(case["context"], output)  # see judge-design.md
    ```
    
    ## LLM-as-judge you can trust
    
    A score you do not trust is worse than no score: an uncalibrated judge gives false confidence, which is more dangerous than admitted ignorance. Each rule, with its why:
    
    - **Judge model ≥ system under test.** A weaker judge cannot reliably rank a stronger system — it scores noise.
    - **The rubric must force a written rationale before the score.** Rationale-first judging is what pushes judge–human agreement to ~85% — higher than two humans agree with each other. A bare number is a vibe with a decimal point.
    - **Pairwise beats pointwise for stability.** "Is A or B better?" is more reproducible than "rate A from 1–10," which inflates and clusters at 8–9.
    - **Swap positions and average.** Judges favor whichever answer came first; run A-then-B and B-then-A to cancel position bias.
    - **Calibrate against human gold and report the agreement before you gate anything on the judge.** Not a formality — this is the step that makes every number downstream defensible.
    
    > Bad judge prompt: "Rate this answer 1–10." → everything lands 8–9, useless.
    > Good: "Compare answer A and answer B against the reference. First write one sentence on each per the rubric, then output the better label." → forces reasoning, gives a stable signal.
    
    Full rubric templates (pointwise + pairwise), the position-swap harness, the calibration script (agreement / Cohen's kappa vs human gold), G-Eval vs DAG, and the judge bias catalog (length, position, self-preference) with mitigations live in **[references/judge-design.md](references/judge-design.md)**.
    
    ## Agent and RAG scorers
    
    Score the path, not only the destination. Beyond exact/judge:
    
    **RAG** (DeepEval / RAGAS names):
    - **Faithfulness** — does the answer only claim what the retrieved context supports? Catches hallucination.
    - **Answer relevancy** — does it actually address the question, or drift?
    - **Contextual recall / precision** — did retrieval fetch the right chunks, and not bury them in noise? Separates a retrieval bug from a generation bug.
    
    **Agent:**
    - **Tool correctness** — right tool, right arguments, right order.
    - **Task completion / goal accuracy** — did it finish the job, not just produce plausible text.
    - **Trajectory scoring** — grade the sequence of steps. A correct final answer from a wrong path will fail differently next time; only trajectory scoring catches it.
    
    The system side of these (how the loop and tools are built) is `../building-agents/SKILL.md`; a common system-under-test is `../chatbot/SKILL.md`.
    
    ## The regression gate
    
    > Gate policy: **block on regression vs a committed baseline, not on an absolute threshold.** An absolute threshold flaps CI on judge noise and gives no signal on drift; "did this PR make a tracked metric worse than `main`?" is the question that matters.
    
    - Compute a **bootstrap confidence interval** on each metric so judge noise alone does not fail the build — only a drop beyond the CI counts.
    - The runner writes `eval-report.json` (metrics, per-failure-mode slices, baseline, pass/fail) and **exits non-zero** on a real regression so the merge is blocked.
    
    ```python
    import json, sys
    
    def gate(current: dict, baseline: dict, margin: float = 0.0) -> int:
        regressed = []
        for metric, score in current.items():
            if metric in baseline and score < baseline[metric] - margin:
                regressed.append((metric, baseline[metric], score))
        report = {"metrics": current, "baseline": baseline, "regressed": regressed,
                  "passed": not regressed}
        with open("eval-report.json", "w") as f:
            json.dump(report, f, indent=2)
        if regressed:
            for m, b, c in regressed:
                print(f"REGRESSION {m}: {b:.3f} -> {c:.3f}", file=sys.stderr)
            return 1
        return 0
    
    sys.exit(gate(run_eval(), json.load(open("eval-baseline.json"))))
    ```
    
    The complete provider-neutral runner (JSONL loader, scorer registry, bootstrap-CI metrics), the GitHub Actions workflow, and side-by-side DeepEval-pytest + Inspect-AI Task/Solver/Scorer versions of the same eval live in **[references/runner-and-gate.md](references/runner-and-gate.md)**.
    
    ## Framework cheat-sheet
    
    Pick by where the eval runs and what it must do. Versions as of 2026-06 — re-verify, they rot.
    
    | Tool | What it is | Reach for it when |
    | --- | --- | --- |
    | **DeepEval** v4.0.3 | pytest-native, 50+ metrics, Decision-Graph (DAG) logic | Your CI is Python/pytest and you want metrics that read like tests. |
    | **Inspect AI** v0.3.225 (UK AISI) | dataset→Task→Solver→Scorer, bootstrap CIs, first-class tool-use & trajectory logging, 200+ pre-built evals | Multi-provider, safety-adjacent, or you need real trajectory scoring. |
    | **promptfoo** (acquired by OpenAI 2026-03) | CLI + YAML, strong pre-deploy + red-team across 50+ vuln types | Config-driven pre-deploy checks; route the red-team half to `agent-safety`. |
    | **Braintrust / LangSmith / Phoenix** v16.0.0 | platforms: annotation, regression tracking, dashboards | You need human annotation queues and historical regression tracking. |
    
    > The two-tool pattern is normal, not over-engineering: a light CI gate (DeepEval / RAGAS / promptfoo) **plus** a platform (Braintrust / LangSmith / Arize) for annotation and history. They share data; different jobs.
    
    ## Anti-patterns
    
    | Anti-pattern | Why it bites | Do instead |
    | --- | --- | --- |
    | Vibes-gating ("feels better, merge it") | No artifact to defend or reproduce | Gate on a number from a committed dataset |
    | Synthetic-only dataset | Model-written cases miss the model's blind spots | Hand-label real traffic by failure mode |
    | Uncalibrated judge | Confident wrong scores; worse than none | Report agreement vs human gold first |
    | Judge weaker than system | Cannot rank a stronger system; scores noise | Judge model ≥ system under test |
    | Absolute-threshold gate | Flaps CI on judge noise, blind to drift | Block on regression vs baseline + bootstrap CI |
    | Shipping on a leaderboard number | Harness effect = 10–20pt swing | Build your own golden set |
    | Scoring only the final answer | A right answer from a wrong path regresses later | Score the trajectory too |
    | Never relabeling drifted gold | Stale "truth" silently rots the gate | Review and relabel the golden set on a schedule |
    
    ## Project grounding
    
    If the workspace has a `02-DOCS/` harness, record the eval policy in `02-DOCS/wiki/stack/evals.md`: dataset location, scorer mix, gate baseline file, judge model, and the failure modes covered. Follow the harness [`wiki-article-template.md`](../harness/references/wiki-article-template.md) (`type: stack`) and index it in `02-DOCS/wiki/index.md`. This is **recorded, not gated** — skip silently if there is no harness.
    
    ## verify.sh
    
    `scripts/verify.sh` is read-only and tool-detecting. It validates that every `*.jsonl` golden set in the project parses and that each line carries the required `id`, `input`, `expected` keys; checks the shape of any `eval-report.json`; and runs `ruff` / `mypy` on example Python and `markdownlint` on docs when those tools are installed. Every missing tool prints a yellow WARN and is skipped — never a failure. An empty or clean target exits 0.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related